System
A remote care system using voice and image recognition with AI analysis addresses the challenges of caring for elderly individuals alone by providing nutritional guidance, medication reminders, and home safety monitoring, improving their quality of life and reducing caregiver burden.
Patent Information
- Application Number
- JP2024130410
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-06
- Publication Date
- 2026-02-19
AI Technical Summary
Existing systems struggle to provide comprehensive care for elderly individuals living alone, particularly in managing their health and safety, due to a lack of real-time monitoring and appropriate response capabilities, exacerbated by caregiver shortages.
A remote care system utilizing voice and image recognition technologies, combined with a generative AI model, to analyze data from elderly individuals, providing nutritional guidance, medication reminders, dementia detection, and home safety monitoring.
Enhances the quality of life for elderly individuals by ensuring appropriate care is delivered remotely, reducing caregiver burden through real-time health and safety monitoring.
Smart Images

Figure 2026028112000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] The current social challenges of an increase in elderly people living alone and a shortage of caregivers make it difficult to provide appropriate care services. Therefore, there is a need for a remote care system that supports the daily lives of elderly people, thereby improving their quality of life and reducing the burden on caregivers. However, conventional systems have had difficulty grasping the real-time status of elderly people and providing appropriate responses. The present invention aims to solve these problems and provide a system that effectively supports the lives and manages the health of elderly people. [Means for solving the problem]
[0005] The remote care system of the present invention includes the following means for monitoring and supporting the elderly's living conditions. Voice data of the elderly is acquired by a voice recognition means, and image data of the elderly is acquired by an image recognition means. Next, this voice data and image data are analyzed by an analysis means using a generative AI model, and a notification means is provided for providing necessary advice and reminders based on the analysis results. The analysis means also has the function of analyzing the elderly's dietary content and evaluating nutritional balance and calories, as well as the function of analyzing medication intake and generating reminders to prevent medication forgetting. This improves the quality of life of the elderly and makes it easier for family members and care service providers living in remote locations to understand the elderly's condition.
[0006] The "voice recognition means" is a device or system that has the function of acquiring voice data from the elderly person and converting it into text data.
[0007] The "image recognition means" is a device or system that has the function of acquiring image data of an elderly person and analyzing the content of that image data.
[0008] A "generative AI model" is an artificial intelligence model used to analyze acquired audio and image data and output results.
[0009] "Analysis means" refers to a device or system that has the function of analyzing voice data and image data using a generative AI model and assessing the condition of an elderly person.
[0010] The "notification means" is a device or system that has the function of transmitting advice and reminders generated based on the analysis results to the elderly.
[0011] "Nutritional balance" is an indicator that shows the appropriate distribution of nutrients in the diet of elderly people.
[0012] "Calories" is a unit that indicates the amount of energy contained in food, and is an important factor to evaluate in the health management of elderly people.
[0013] A "reminder" is a notification that prompts the elderly to take some action based on a specific time or situation.
[0014] "Medication status" refers to information such as which medications an elderly person has taken or needs to take, and when. [Brief explanation of the drawings]
[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0017] First, the terms used in the following description will be explained.
[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0023] [First embodiment]
[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0036] This remote care system consists of elderly end users, and a server and terminal (robot) located in a remote location. A specific embodiment of the system is described below. The program code itself is not provided, but the processing flow and each step are described in detail.
[0037] System configuration
[0038] 1. Terminal configuration
[0039] Speech recognition means: The device is equipped with a microphone to collect conversations with the elderly in real time, and speech recognition technology is used to convert the captured speech into text data.
[0040] Image recognition method: The device is equipped with a camera that captures images of the elderly person's daily life, especially when they eat or take medicine. The image data is sent to a server in real time.
[0041] 2. Server-side configuration
[0042] Analysis method: The server uses a generative AI model to analyze the voice and image data sent from the device. This analysis allows for a real-time assessment of the elderly person's health and living conditions.
[0043] Notification method: Based on the analysis results, appropriate advice and reminders are generated and sent to the device, which then communicates this to the elderly via voice.
[0044] Meal management function for the elderly
[0045] Overview: The dietary management function monitors whether the elderly are eating a balanced diet.
[0046] Terminal: Ask the senior, "What did you have for breakfast today?"
[0047] User: "I had bread and milk."
[0048] Device: Converts voice into text data and simultaneously takes pictures of the meal using a camera.
[0049] Server: Analyzes the received text and image data and evaluates the calorie and nutritional balance of the meal.
[0050] Server: Generates advice as needed, such as "Today's breakfast appears to be low in calories. We recommend eating some fruits and vegetables at your next meal."
[0051] Device: Provides advice to the elderly via voice.
[0052] Medication management function
[0053] Overview: Medication management helps seniors take their medications appropriately.
[0054] Server: Detects when it's time to take your medicine and generates a reminder saying, "It's time to take your medicine. Are you ready?"
[0055] Device: Provides voice reminders to seniors.
[0056] User: "Yes, I'll take my medicine."
[0057] Device: Activate the camera and record the scene of taking the medicine.
[0058] Server: Analyzes image data and verifies whether the medication has been taken.
[0059] Server: If the medication is confirmed, generate a confirmation message saying "Your medication was taken correctly" and set the next reminder.
[0060] Terminal: Notify the elderly person of a confirmation message.
[0061] Early detection of dementia
[0062] Overview: The early dementia detection function detects signs of dementia from the everyday conversations of elderly people.
[0063] Device: Initiate everyday conversations such as "Tell me how your day was."
[0064] User: Talk about everyday events.
[0065] Terminal: Converts the voice into text data and sends it to the server.
[0066] Server: Analyzes conversation data and evaluates language patterns and emotional states, for example, detecting vocabulary declines and changes in emotional tone.
[0067] Server: If signs of dementia are detected, it generates an alert such as, "Recently, there has been a decline in vocabulary in conversations."
[0068] Device: Notifies the senior of the alert and, if necessary, notifies caregivers and family members.
[0069] As described above, this remote care system provides comprehensive support for the daily lives of elderly people and makes it easier for caregivers and family members in remote locations to understand the condition of the elderly, thereby improving the quality of life of the elderly and reducing the burden on caregivers.
[0070] The processing flow will be explained below.
[0071] Meal management function for the elderly
[0072] Step 1:
[0073] The device asks the senior, "What did you have for breakfast today?"
[0074] Step 2:
[0075] The user responds, "I had bread and milk."
[0076] Step 3:
[0077] The terminal uses a voice recognition system to convert the user's speech into text data.
[0078] Step 4:
[0079] The device activates the camera and takes pictures of the elderly person eating.
[0080] Step 5:
[0081] The terminal transmits the acquired text data and image data to the server.
[0082] Step 6:
[0083] The server analyzes the received text data and extracts the meal details.
[0084] Step 7:
[0085] The server analyzes the received image data and calculates the calories and nutritional balance of the meal contents.
[0086] Step 8:
[0087] Based on the analysis results, the server generates advice such as, "Today's breakfast appears to be low in calories. We recommend that you eat fruits and vegetables at your next meal."
[0088] Step 9:
[0089] The server sends the generated advice to the terminal.
[0090] Step 10:
[0091] The device will provide advice to the elderly via voice.
[0092] Medication management function
[0093] Step 1:
[0094] The server recognizes pre-set medication times and generates reminders.
[0095] Step 2:
[0096] The device will then send a voice reminder to the elderly person saying, "It's time to take your medicine. Are you ready?"
[0097] Step 3:
[0098] The user responds, "Yes, I'll take my medicine."
[0099] Step 4:
[0100] The terminal uses a voice recognition system to convert the user's response into text data.
[0101] Step 5:
[0102] The device activates the camera and takes a picture of the elderly person taking their medicine.
[0103] Step 6:
[0104] The video data acquired by the terminal is transmitted to the server.
[0105] Step 7:
[0106] The server analyzes the video data and confirms that the medication has been taken.
[0107] Step 8:
[0108] The server generates a confirmation message saying "Your medication was taken correctly" and sets the next reminder.
[0109] Step 9:
[0110] The server sends a confirmation message to the terminal.
[0111] Step 10:
[0112] The device will then send a confirmation message to the elderly person.
[0113] Early detection of dementia
[0114] Step 1:
[0115] The device begins a casual conversation with the elderly person, asking, "Tell me how your day was today."
[0116] Step 2:
[0117] Users talk about everyday events.
[0118] Step 3:
[0119] The terminal uses a voice recognition system to convert the user's speech into text data.
[0120] Step 4:
[0121] The terminal transmits the acquired text data to the server.
[0122] Step 5:
[0123] The server analyzes the text data and evaluates the elderly person's language patterns and emotional state.
[0124] Step 6:
[0125] The server detects signs of dementia based on the analysis results.
[0126] Step 7:
[0127] The server generates an alert such as, "Recently, we have noticed a decline in vocabulary in conversations."
[0128] Step 8:
[0129] The server generates alerts and sends them to the device, which also notifies family members and caregivers.
[0130] Step 9:
[0131] The device will notify the elderly person of the alert.
[0132] These processing steps are designed to ensure that the lives of the elderly are managed effectively and to reduce the burden on caregivers.
[0133] Example 1
[0134] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0135] It is difficult for elderly people to maintain appropriate lifestyle habits and monitor their health. In particular, it is important for elderly people to manage the nutritional balance of their diet, their medication status, and even to detect dementia early, but it is difficult for them to do these things alone. It is also not easy for caregivers and family members in remote locations to understand the condition of elderly people. Conventional systems lack the functionality to comprehensively support the lives of elderly people.
[0136] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0137] In this invention, the server includes a voice recognition means for acquiring voice data of the elderly person, an image recognition means for acquiring image data of the elderly person, an analysis means using a generative AI model to analyze the voice data and image data, a notification means for notifying the elderly person of necessary advice and reminders based on the analysis results, a means for detecting signs of dementia from the elderly person's everyday conversation and generating an alert, a means for reminding the elderly person to take their medicine and confirming that they have taken it, and a means for monitoring the elderly person's diet to evaluate nutritional balance and provide advice. This makes it possible to comprehensively support the elderly's living conditions and enable caregivers and family members in remote locations to easily understand the elderly person's condition.
[0138] The "voice recognition means" is a device or technology for capturing the voice of the elderly person in real time and converting the voice data into text data.
[0139] "Image recognition means" refers to a device or technology for photographing the elderly person's daily life and acquiring image data.
[0140] A "generative AI model" is an artificial intelligence model that analyzes voice and image data and generates appropriate advice and reminders based on the results.
[0141] The "analysis means" is a means of evaluating the health status and living conditions of elderly people using a generative AI model, using the acquired voice and image data.
[0142] The "notification means" is a means of notifying the elderly of advice and reminders generated based on the analysis results via voice or other means.
[0143] The "means for detecting signs of dementia" is a method for analyzing data on everyday conversations of elderly people and detecting early signs of dementia from factors such as a decrease in vocabulary and changes in emotional tone.
[0144] "Reminder measures" are measures that notify elderly people when it is time to take their medicine and help them ensure that they take their medicine.
[0145] A "dietary monitoring method" is a method for monitoring the dietary content of elderly people, evaluating their nutritional balance and calories, and providing necessary advice.
[0146] This invention is a remote care system for monitoring and supporting the living conditions of elderly people. The system consists of an elderly person in a remote location, a server, and a terminal (robot). This system is configured and operates as follows to monitor the elderly person's daily behavior and health status and provide necessary advice and reminders.
[0147] System configuration
[0148] Terminal configuration
[0149] Speech recognition method: The device is equipped with a microphone to collect conversations with the elderly in real time. The captured speech is converted into text data using speech recognition technology (e.g., Google Cloud Speech-to-Text).
[0150] Image recognition method: The device is equipped with a camera that captures images of the elderly person's daily life, especially when they eat or take medicine. The image data is sent to a server in real time.
[0151] Server-side configuration
[0152] Analysis method: The server uses a generative AI model (e.g., OpenAI GPT-4) to analyze the voice and image data sent from the device. This analysis allows for a real-time assessment of the elderly person's health and living conditions.
[0153] Notification method: Based on the analysis results, appropriate advice and reminders are generated and sent to the device, which then communicates this to the elderly via voice.
[0154] Meal management function for the elderly
[0155] Overview: The dietary management function monitors whether the elderly are eating a balanced diet.
[0156] Terminal: Ask the senior, "What did you have for breakfast today?"
[0157] User: "I had bread and milk."
[0158] Device: Converts voice into text data and simultaneously takes pictures of the meal using a camera.
[0159] Server: Analyzes the received text and image data and evaluates the calorie and nutritional balance of the meal.
[0160] Server: Generates advice as needed, such as "Today's breakfast appears to be low in calories. We recommend eating some fruits and vegetables at your next meal."
[0161] Device: Provides advice to the elderly via voice.
[0162] Specific prompt examples:
[0163] "If an elderly person eats bread and milk for breakfast, generate a sentence that evaluates the nutritional balance and advises them to add more fruits and vegetables to their next meal if the calories are low."
[0164] Medication management function
[0165] Overview: Medication management helps seniors take their medications appropriately.
[0166] Server: Detects when it's time to take your medicine and generates a reminder saying, "It's time to take your medicine. Are you ready?"
[0167] Device: Provides voice reminders to seniors.
[0168] User: "Yes, I'll take my medicine."
[0169] Device: Activate the camera and record the scene of taking the medicine.
[0170] Server: Analyzes image data and verifies whether the medication has been taken.
[0171] Server: If the medication is confirmed, generate a confirmation message saying "Your medication was taken correctly" and set the next reminder.
[0172] Terminal: Notify the elderly person of a confirmation message.
[0173] Specific prompt examples:
[0174] "Generate a reminder when it's time for the senior to take their medication, and then create a sentence to confirm if they actually took their medication."
[0175] Early detection of dementia
[0176] Overview: The early dementia detection function detects signs of dementia from the everyday conversations of elderly people.
[0177] Device: Initiate everyday conversations such as "Tell me how your day was."
[0178] User: Talk about everyday events.
[0179] Terminal: Converts the voice into text data and sends it to the server.
[0180] Server: Analyzes conversation data and evaluates language patterns and emotional states, for example, detecting vocabulary declines and changes in emotional tone.
[0181] Server: If signs of dementia are detected, it generates an alert such as, "Recently, there has been a decline in vocabulary in conversations."
[0182] Device: Notifies the senior of the alert and, if necessary, notifies caregivers and family members.
[0183] Specific prompt examples:
[0184] "To detect signs of dementia in everyday conversations of elderly people, create sentences that analyze vocabulary decline and changes in emotional tone."
[0185] As described above, this remote care system provides comprehensive support for the daily lives of elderly people and makes it easier for caregivers and family members in remote locations to understand the condition of the elderly, thereby improving the quality of life of the elderly and reducing the burden on caregivers.
[0186] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0187] Meal management function for the elderly
[0188] Step 1:
[0189] The device asks the senior, "What did you have for breakfast today?"
[0190] Input: Breakfast Question
[0191] Output: Waiting for response from the elderly person
[0192] Step 2:
[0193] The user responds, "I had bread and milk."
[0194] Input: Elderly person's voice response
[0195] Output: Audio data
[0196] Step 3:
[0197] The device converts the voice into text data and simultaneously captures the meal with a camera.
[0198] Input: Audio data, video of elderly people eating
[0199] Output: Text data, image data
[0200] What it does: The microphone captures your voice and converts it into text using a speech recognition engine. The camera captures footage of you eating.
[0201] Step 4:
[0202] The server receives the text and image data and analyzes it using a generative AI model.
[0203] Input: Text data, image data
[0204] Output: Calorie and nutritional balance evaluation results of the meal
[0205] How it works: The server extracts meal items from text data and analyzes the meal contents using an image recognition algorithm. The generative AI model calculates nutritional balance and calories.
[0206] Step 5:
[0207] The server generates advice based on the analysis results.
[0208] Input: Evaluation results of calorie and nutritional balance of meals
[0209] Output: Advice message
[0210] Specific behavior: The server generates advice such as "Today's breakfast is low in calories. We recommend that you eat fruits and vegetables at your next meal."
[0211] Step 6:
[0212] The device will provide advice to the elderly via voice.
[0213] Input: Advice message
[0214] Output: Voice notification to the elderly
[0215] Specific operation: Advice messages are synthesized and conveyed to the elderly.
[0216] Medication management function
[0217] Step 1:
[0218] The server detects when it's time to take your medicine and generates a reminder saying, "It's time to take your medicine. Are you ready?"
[0219] Input: Medication schedule
[0220] Output: Reminder message
[0221] What it does: The server periodically checks the schedule and generates a reminder message when it's time to take your medication.
[0222] Step 2:
[0223] The device will notify the elderly of the reminder via voice.
[0224] Input: Reminder message
[0225] Output: Voice notification to the elderly
[0226] Specific actions: Reminder messages are delivered to elderly people via voice synthesis.
[0227] Step 3:
[0228] The user responds, "Yes, I'll take my medicine."
[0229] Input: Elderly person's voice response
[0230] Output: Response audio data
[0231] Step 4:
[0232] The device activates the camera and captures the scene of the elderly person taking their medicine.
[0233] Input: Video of elderly person taking medication
[0234] Output: Image data of the medication scene
[0235] Specific operation: The camera captures the elderly person taking their medicine and saves the image data on the device.
[0236] Step 5:
[0237] The server analyzes the image data and verifies whether the medication has been taken.
[0238] Input: Image data of a medication scene
[0239] Output: Medication confirmation result
[0240] What it does: The server uses image recognition technology to verify that the medication was taken correctly.
[0241] Step 6:
[0242] The server generates a confirmation message and sets the next reminder.
[0243] Input: Medication confirmation result
[0244] Output: Confirmation message, next reminder setting
[0245] Specific behavior: If the medication is confirmed, the server generates a confirmation message saying "Your medication was taken correctly" and sets the next reminder.
[0246] Step 7:
[0247] The device will then send a confirmation message to the elderly person.
[0248] Input: Confirmation message
[0249] Output: Voice notification to the elderly
[0250] Specific operation: A confirmation message is synthesized and conveyed to the elderly person.
[0251] Early detection of dementia
[0252] Step 1:
[0253] The device will begin a casual conversation such as, "Tell me how your day was today."
[0254] Input: Daily conversation question
[0255] Output: Waiting for response from the elderly person
[0256] Step 2:
[0257] Users talk about everyday events.
[0258] Input: Elderly person's voice
[0259] Output: Audio data
[0260] Step 3:
[0261] The device converts the voice into text data and sends it to the server.
[0262] Input: Audio data
[0263] Output: Text data
[0264] What it does: The microphone captures your voice and the speech recognition engine converts it into text.
[0265] Step 4:
[0266] The server analyzes the conversation data and evaluates language patterns and emotional states.
[0267] Input: Text data
[0268] Output: Evaluation result
[0269] Specific operation: The server uses the generative AI model to detect vocabulary loss, changes in emotional tone, etc.
[0270] Step 5:
[0271] If the server detects signs of dementia, it generates an alert message.
[0272] Input: Evaluation result
[0273] Output: Alert message
[0274] Specific behavior: If signs of dementia are detected, an alert will be generated, such as "Recently, there has been a decline in vocabulary in conversations."
[0275] Step 6:
[0276] The device notifies the elderly person of the alert and, if necessary, notifies caregivers and family members.
[0277] Input: Alert message
[0278] Output: Audio notification to the senior and caregiver / family
[0279] Specific operation: An alert message is synthesized and conveyed to the elderly. At the same time, an alert is sent to caregivers and family members.
[0280] The above are the specific processing steps and their detailed operations. These steps improve the quality of life for the elderly and enable appropriate care to be provided remotely.
[0281] (Application example 1)
[0282] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0283] Remote care systems that support the lives of the elderly are limited in their management functions to the health status of the elderly, and do not address comprehensive home safety. There is a need for a system that comprehensively monitors security aspects such as managing entry and exit within the home and detecting fires, gas leaks, and suspicious individuals. The objective of this invention is to provide a comprehensive support system that not only manages the health of the elderly, but also monitors the safety status of the home.
[0284] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0285] In this invention, the server includes an analysis means, a notification means, and a safety monitoring means, which makes it possible not only to grasp and support the living conditions of the elderly, but also to comprehensively monitor the safety of the home.
[0286] A "remote care system" is a system that monitors and supports the living conditions of elderly people from a remote location.
[0287] The "voice recognition means" is a device that has the function of acquiring voice data from the elderly person and converting it into text data in real time.
[0288] The "image recognition means" is a device that has the function of acquiring image data of an elderly person and analyzing the content of that image data.
[0289] A "generative AI model" is an artificial intelligence model that analyzes acquired voice data and image data and generates the necessary information.
[0290] The "analysis means" is a function that analyzes voice data and image data to evaluate the living conditions of the elderly person.
[0291] "Notification means" is a function that notifies elderly people of advice and reminders based on the analysis results.
[0292] The "safety monitoring means" is a device that has the function of monitoring the safety status of the home, detecting abnormalities, and notifying them.
[0293] "Nutritional balance" is an indicator of whether nutrients are evenly distributed in a particular diet.
[0294] A "calorie" is a unit that indicates the amount of energy ingested through food.
[0295] A "reminder" is a notification that encourages seniors to take a specific action.
[0296] This remote care system is a system for monitoring and supporting the living conditions and home safety conditions of elderly people. A specific embodiment of the system will be described below.
[0297] System configuration
[0298] The server and user terminals (smartphones) work together to monitor the lives of the elderly and the safety of their homes using various sensor devices. This system includes the following hardware and software:
[0299] Hardware
[0300] Smartphone: microphone, camera
[0301] Sensors: Door and window sensors, smoke detectors, gas leak detectors
[0302] software
[0303] Speech recognition API: Used to convert voice data into text data. Typical examples include Google Speech-to-Text and Apple Siri.
[0304] Image recognition API: Used to analyze image data. Representative examples include Google Cloud Vision and Amazon Rekognition.
[0305] Server: Platforms that run generative AI models for data analysis include Google Cloud and AWS (Amazon Web Services).
[0306] System Operation
[0307] The server uses voice recognition, image recognition, and safety monitoring means to notify the user based on the analysis results. Each of these means is described in detail below.
[0308] Voice recognition means
[0309] The microphone on the user's device collects the elderly's voice and converts it into text data in real time using a speech recognition API. For example, if an elderly person says "help me," it is converted into text data and sent to the server.
[0310] Image Recognition Method
[0311] The camera on the user's device captures images of the elderly's living conditions and home situation, and the images are analyzed using an image recognition API. Examples include detecting fires, gas leaks, and suspicious people. It also includes capturing images of elderly people taking medicine, and analyzing the image data on a server.
[0312] Safety monitoring means
[0313] Sensor devices such as door and window sensors, smoke detectors, and gas leak detectors monitor the safety status of the home and notify the server if an abnormality is detected. The server analyzes the abnormality and issues a warning to the user in real time.
[0314] Server analysis method
[0315] The server analyzes the voice and image data using a generative AI model, and based on the analysis results, generates appropriate advice and reminders and sends them to the user's device.
[0316] Notification means
[0317] Based on the analysis results, the system will provide voice reminders and warnings to the elderly, such as specific instructions such as "There is a fire. Please close the doors of each room."
[0318] Specific examples
[0319] In case of fire detection
[0320] 1. The microphone detects the sound of a fire alarm.
[0321] 2. The voice recognition API converts the message "Fire has broken out" into text data.
[0322] 3. Analyze fire data on the server.
[0323] 4. An "emergency fire alert" notification will be sent to your smartphone.
[0324] 5. An audio guide will be given to residents to "close the doors of each room."
[0325] Prompt Sentence Examples
[0326] "Design a program that uses a voice recognition system to detect fire alarm sounds in real time and notify a smartphone of that information. If an abnormal sound is detected, provide voice guidance on first aid measures for the fire."
[0327] This will make it possible to realize a system that comprehensively supports the lives and home safety of the elderly.
[0328] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0329] Step 1:
[0330] The microphone on the user's device captures the elderly person's voice data, which is then sent to a voice recognition API and converted into text data in real time.
[0331] Input: Elderly voice data
[0332] Output: Text data
[0333] Specific operation: When an elderly person says "help me," their voice is picked up by a microphone and converted into text data using a voice recognition API.
[0334] Step 2:
[0335] The camera on the user's device captures image data of the elderly person, which is then sent to an image recognition API for real-time analysis.
[0336] Input: Image data of elderly people
[0337] Output: Analysis data
[0338] Specific operation: The camera captures scenes of elderly people eating or taking medicine, and the image data is analyzed using an image recognition API.
[0339] Step 3:
[0340] The server receives the voice and image data sent from the user's device and analyzes the data using a generative AI model.
[0341] Input: Text data, analysis data
[0342] Output: Analysis results
[0343] Specific operation: The voice data is converted into text and the analysis results of the image data are sent to a server, and the generated AI model analyzes the elderly person's living conditions and health condition.
[0344] Step 4:
[0345] The server generates appropriate advice and reminders based on the analysis results.
[0346] Input: Analysis results
[0347] Output: Advice, reminder
[0348] Specific behavior: For example, if an elderly person's breakfast is analyzed to be insufficient in calories, advice such as "We recommend that you eat fruits and vegetables at your next meal" will be generated.
[0349] Step 5:
[0350] The server transmits the generated advice and reminders to the user terminal using a notification means.
[0351] Input: Advice, Reminder
[0352] Output: Notification data
[0353] Specific operation: The generated advice and reminders are sent to the user's device as push notifications, and the elderly person receives notifications such as "It's time to take their medicine."
[0354] Step 6:
[0355] Sensors on the user's device (doors, windows, smoke, gas leaks) monitor the safety status of the home, and if an abnormality is detected, the data is sent to the server.
[0356] Input: Sensor data
[0357] Output: Anomaly detection data
[0358] Specific operation: If a door or window is opened or closed suspiciously, or if a fire or gas leak is detected, the information is sent to the server.
[0359] Step 7:
[0360] The server analyzes the anomaly detection data, generates necessary warnings, and sends them to the user terminal.
[0361] Input: Anomaly detection data
[0362] Output: Warning notification data
[0363] Specific operation: For example, a warning message such as "A fire has been detected. Please close the doors of each room" is generated and sent to the user terminal.
[0364] Step 8:
[0365] The user terminal notifies the elderly person of the received warning notification data by voice in real time.
[0366] Input: Alert notification data
[0367] Output: Audio notification
[0368] Specific actions: The elderly person will be notified by voice, "A fire has broken out. Please close the doors of each room."
[0369] This will enable comprehensive monitoring of the elderly's living conditions and home safety, and provide appropriate support at the appropriate time.
[0370] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0371] This invention is a remote care system that monitors the living conditions of elderly people and provides appropriate support, and is equipped with an emotion engine that analyzes voice data and image data of the elderly person and recognizes the emotional state of the user. Specific embodiments of the invention will be described below.
[0372] System configuration
[0373] 1. Terminal configuration
[0374] Voice recognition means: The device is equipped with a microphone to collect conversations with the elderly in real time. The collected voice data is converted into text data by a voice recognition system.
[0375] Image recognition means: The device is equipped with a camera that takes pictures of, for example, an elderly person eating or taking medicine. The captured image data is sent to a server in real time.
[0376] 2. Server-side configuration
[0377] Analysis method: The server is equipped with a generative AI model with advanced analytical capabilities that analyzes the voice and image data sent from the device. This analysis allows for a real-time assessment of the elderly person's health and living conditions.
[0378] Emotion engine: The server is also equipped with an emotion engine that recognizes the emotional state of the elderly by analyzing voice data.
[0379] 3. Means of notification
[0380] Advice and reminders: Based on the analysis results and emotional state, advice and reminders are generated for the elderly. This information is sent from the server to the device, which then notifies the elderly by voice.
[0381] Meal management function for the elderly
[0382] Overview: The dietary management function monitors whether the elderly are eating a balanced diet and provides advice as needed.
[0383] Terminal: Ask the senior, "What did you have for breakfast today?"
[0384] User: "I had bread and milk."
[0385] Device: A voice recognition system converts the user's speech into text data. A camera captures the meal.
[0386] Server: Analyzes text and image data to evaluate the calorie and nutritional balance of meal content. Also analyzes the user's emotional state using an emotion engine.
[0387] Server: Generate advice such as, "Your breakfast today seems low in calories. We recommend adding more fruits and vegetables to your next meal."
[0388] Device: Provides advice to the elderly via voice.
[0389] Medication management function
[0390] Abstract: Medication management is a function that helps elderly people take their medications appropriately.
[0391] Server: Detects medication times and generates reminders.
[0392] Device: Reminds seniors, "It's time to take your medicine. Are you ready?"
[0393] User: "Yes, I'll drink it now."
[0394] Device: A voice recognition system converts responses into text data, and a camera captures the process of taking the medicine.
[0395] Server: Analyzes video data and confirms whether the user has taken the medication. An emotion engine also analyzes the user's emotional state.
[0396] Server: Generates a confirmation message saying "Your medication was taken correctly" and sets the next reminder.
[0397] Terminal: Notify the elderly person of a confirmation message.
[0398] Early detection of dementia
[0399] Overview: The early dementia detection function detects signs of dementia from the everyday conversations of elderly people.
[0400] Device: Begins a conversation with "Tell me how your day was."
[0401] User: Talk about everyday events.
[0402] Terminal: The speech recognition system converts the conversation into text data and sends it to the server.
[0403] Server: Analyzes text data and evaluates language patterns and emotional states. An emotion engine also analyzes emotional fluctuations.
[0404] Server: Generate an alert such as, "Recently, we have noticed a decline in vocabulary in conversations."
[0405] Device: Notifies the senior of the alert and, if necessary, notifies caregivers and family members.
[0406] In this way, the remote care system of the present invention supports the lives of elderly people in a variety of ways, and by combining it with an emotion engine, it is possible to provide more accurate advice and appropriate care.
[0407] The processing flow will be explained below.
[0408] Meal management function for the elderly
[0409] Step 1:
[0410] The device asks the senior, "What did you have for breakfast today?"
[0411] Step 2:
[0412] The user responds, "I had bread and milk."
[0413] Step 3:
[0414] The terminal uses a voice recognition system to convert the user's speech into text data.
[0415] Step 4:
[0416] The device activates the camera and takes pictures of the elderly person eating.
[0417] Step 5:
[0418] The terminal transmits the acquired text data and image data to the server.
[0419] Step 6:
[0420] The server analyzes the received text data and extracts the meal details.
[0421] Step 7:
[0422] The server analyzes the received image data and calculates the calories and nutritional balance of the meal contents.
[0423] Step 8:
[0424] The server uses an emotion engine to recognize the user's emotional state from the voice data.
[0425] Step 9:
[0426] Based on the analysis results and emotional state, the server generates advice such as, "Today's breakfast seems low in calories. I recommend eating fruits and vegetables at your next meal."
[0427] Step 10:
[0428] The server sends the generated advice to the terminal.
[0429] Step 11:
[0430] The device will provide advice to the elderly via voice.
[0431] Medication management function
[0432] Step 1:
[0433] The server recognizes pre-set medication times and generates reminders.
[0434] Step 2:
[0435] The device will then send a voice reminder to the elderly person saying, "It's time to take your medicine. Are you ready?"
[0436] Step 3:
[0437] The user responds, "Yes, I'll take my medicine."
[0438] Step 4:
[0439] The terminal uses a voice recognition system to convert the user's response into text data.
[0440] Step 5:
[0441] The device activates the camera and takes a picture of the elderly person taking their medicine.
[0442] Step 6:
[0443] The video data acquired by the terminal is transmitted to the server.
[0444] Step 7:
[0445] The server analyzes the video data and confirms that the medication has been taken.
[0446] Step 8:
[0447] The server uses an emotion engine to analyze the user's emotional state.
[0448] Step 9:
[0449] The server generates a confirmation message saying "Your medication was taken correctly" and sets the next reminder.
[0450] Step 10:
[0451] The server sends a confirmation message to the terminal.
[0452] Step 11:
[0453] The device will then send a confirmation message to the elderly person.
[0454] Early detection of dementia
[0455] Step 1:
[0456] The device begins a casual conversation by asking, "Tell me how your day was today."
[0457] Step 2:
[0458] Users talk about everyday events.
[0459] Step 3:
[0460] The terminal uses a voice recognition system to convert the user's speech into text data.
[0461] Step 4:
[0462] The terminal transmits the acquired text data to the server.
[0463] Step 5:
[0464] The server analyzes the text data and evaluates the elderly person's language patterns.
[0465] Step 6:
[0466] The server uses an emotion engine to analyze the user's emotional state from their speech.
[0467] Step 7:
[0468] Based on the analysis results and emotional state, the server generates an alert such as, "Recently, there has been a decrease in vocabulary in conversations."
[0469] Step 8:
[0470] The server generates an alert and sends it to the device.
[0471] Step 9:
[0472] The device will notify the elderly person of the alert.
[0473] Step 10:
[0474] The server will also notify caregivers and family members of the alerts as needed.
[0475] These processing steps enable the remote care system of the present invention to monitor the user's living conditions with high accuracy and provide appropriate care and support. In addition, the emotion engine enables flexible responses according to the user's emotional state, improving the user's quality of life and reducing the burden on caregivers.
[0476] Example 2
[0477] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0478] For elderly people to maintain independent lifestyles, it is important to properly monitor their daily health and emotional states and provide timely advice and reminders. However, current systems have difficulty effectively analyzing the subject's voice and images and providing appropriate advice and reminders that take their emotional state into account. Furthermore, the lack of personalized support tailored to each individual's health condition and specific lifestyle circumstances makes it difficult for elderly people to maintain appropriate diets and medication intake.
[0479] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0480] In this invention, the server includes an analysis means that uses a generative AI model to analyze voice data and image data and evaluate health conditions and living situations, an emotion engine that recognizes emotional states from voice data, a means that analyzes dietary content and medication status to generate advice and reminders necessary for the elderly, and a notification means that notifies the elderly of the generated advice and reminders by voice. This enables comprehensive analysis using the voices and images of the elderly, and makes it possible to provide appropriate, personalized advice and reminders.
[0481] "Speech recognition means" is a technology that collects speech and converts it into text data.
[0482] "Image recognition means" is a technology that collects images and sends the data to a server.
[0483] A "generative AI model" is an artificial intelligence technology that analyzes voice and image data to assess the health and living conditions of elderly people.
[0484] "Analysis means" refers to the process of analyzing and evaluating audio and image data using a generative AI model.
[0485] "Emotion engine" is a technology that recognizes the user's emotional state from voice data.
[0486] The "notification means" is a technology that notifies the elderly of the generated advice and reminders by voice.
[0487] A "reminder" is a notification that helps seniors remember to take certain actions, such as taking medicine.
[0488] "Text data" refers to character data converted from speech by speech recognition means.
[0489] The present invention is a remote care system that monitors the living conditions of elderly people and provides appropriate support, and is equipped with an emotion engine that analyzes voice data and image data of the elderly person and recognizes the emotional state of the user. Specific embodiments of the present invention will be described below.
[0490] System configuration
[0491] This system consists of a voice recognition means, an image recognition means, an analysis means, an emotion engine, and a notification means.
[0492] Voice recognition means
[0493] The device is equipped with a microphone that collects conversations with elderly people in real time. The collected voice data is converted into text data through a voice recognition system, allowing the user's speech to be captured as text information.
[0494] Image Recognition Method
[0495] The device is equipped with a camera that takes pictures of the elderly person eating and taking their medicine. The captured image data is sent to a server in real time, allowing for a visual record of the elderly person's behavior.
[0496] Analysis means
[0497] The server is equipped with a generative AI model with advanced analytical capabilities that analyzes the voice and image data sent from the device. This analysis makes it possible to evaluate the health and living conditions of elderly people in real time. The generative AI model uses deep learning and natural language processing techniques, for example, to perform highly accurate analysis of the data.
[0498] Emotion Engine
[0499] The server's emotion engine has the ability to recognize the user's emotional state from voice data, allowing it to assess not just their physical state but also their psychological state. Emotion analysis is performed based on the tone of voice and the choice of words used.
[0500] Notification means
[0501] Based on the analysis results and emotional state, the server generates advice and reminders for the elderly, which are then sent to the device, which then notifies the elderly by voice.
[0502] Meal management function
[0503] Overview: The dietary management function monitors whether the elderly are eating a balanced diet and provides advice as needed.
[0504] The device asks the senior, "What did you have for breakfast today?"
[0505] The user responds, "I had bread and milk."
[0506] The device uses a voice recognition system to convert the user's speech into text data and a camera to capture the meal.
[0507] The server analyzes text and image data to evaluate the calorie and nutritional balance of the meal, and also analyzes the user's emotional state using an emotion engine.
[0508] The server generates advice such as, "Today's breakfast seems low in calories. We recommend adding more fruits and vegetables to your next meal."
[0509] The device will provide advice to the elderly via voice.
[0510] Example prompt sentence:
[0511] Prompts for older adults regarding dietary management functions:
[0512] Please explain the analysis process and advice generation steps when an elderly person answers "I had bread and milk" to the question "What did you have for breakfast today?"
[0513] Medication management function
[0514] Abstract: Medication management is a function that helps elderly people take their medications appropriately.
[0515] The server detects when it's time to take the medication and generates a reminder.
[0516] The device will remind the elderly person, "It's time to take your medicine. Are you ready?"
[0517] The user responds, "Yes, I'll drink it now."
[0518] The device uses a voice recognition system to convert responses into text data and a camera to record the patient taking the medicine.
[0519] The server analyzes the video data to confirm whether the user has taken the medicine, and also analyzes the user's emotional state using an emotion engine.
[0520] The server generates a confirmation message saying "Your medication was taken correctly" and sets the next reminder.
[0521] The device will then send a confirmation message to the elderly person.
[0522] Example prompt sentence:
[0523] Prompts regarding medication management functions for seniors:
[0524] Please explain the analysis process and reminder generation steps when an elderly person responds "Yes, I'll take it right away" to the reminder "It's time to take your medicine. Are you ready?"
[0525] Early detection of dementia
[0526] Overview: The early dementia detection function detects signs of dementia from the everyday conversations of elderly people.
[0527] The device begins a casual conversation by asking, "Tell me how your day was today."
[0528] Users talk about everyday events.
[0529] The terminal uses a voice recognition system to convert the conversation into text data and send it to the server.
[0530] The server analyzes the text data, assessing language patterns and emotional states, and also analyzes emotional fluctuations using an emotion engine.
[0531] The server generates an alert such as, "Recently, we have noticed a decline in vocabulary in conversations."
[0532] The device will notify the senior of the alert and, if necessary, their caregiver or family member.
[0533] Example prompt sentence:
[0534] Here are some prompts for early dementia detection:
[0535] Please explain the steps for analyzing and generating alerts after an elderly person talks about their daily life when asked, "Tell me how your day was today."
[0536] In this way, the remote care system of the present invention supports the lives of elderly people in a variety of ways, and by combining it with an emotion engine, it is possible to provide more accurate advice and appropriate care.
[0537] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0538] Meal management function processing steps
[0539] Step 1:
[0540] Voice input and data collection (terminal)
[0541] The device uses a microphone to ask the elderly person, "What did you have for breakfast today?" The elderly person responds, "I had bread and milk." This voice input is collected.
[0542] Input: Voice input for the elderly
[0543] Output: Audio data collected by microphone
[0544] Step 2:
[0545] Speech recognition and data conversion (terminal)
[0546] The device uses a voice recognition system to convert voice data into text data, and also uses a camera to record the elderly person's mealtimes.
[0547] Input: Audio data, video data of meal
[0548] Output: Text data, image data
[0549] Step 3:
[0550] Data transmission (terminal)
[0551] The terminal encrypts the converted text data and the captured image data and transmits them to the server in real time.
[0552] Input: Text data, image data
[0553] Output: Data sent to the server
[0554] Step 4:
[0555] Data analysis (server)
[0556] The server uses the generative AI model to analyze the received text and image data and evaluate the calories and nutritional balance of the meal contents.
[0557] Input: Text data, image data
[0558] Output: Calorie and nutritional balance assessment results
[0559] Step 5:
[0560] Sentiment analysis (server)
[0561] The server's emotion engine recognizes the user's emotional state from the voice data.
[0562] Input: Text data
[0563] Output: User's emotional state
[0564] Step 6:
[0565] Advice Generation (Server)
[0566] The server generates advice based on the meal evaluation and emotional state, such as "We recommend adding more fruits and vegetables to your next meal."
[0567] Input: Calorie and nutritional balance assessment results, user's emotional state
[0568] Output: The generated advice
[0569] Step 7:
[0570] Notifications (device)
[0571] The device notifies the elderly person of the advice received from the server via voice.
[0572] Input: Advice from the server
[0573] Output: Voice notification to the elderly
[0574] Medication management function processing steps
[0575] Step 1:
[0576] Medication time detection (server)
[0577] The server detects the preset time for taking the medication.
[0578] Input: Pre-set medication time
[0579] Output: Start generating medication reminders
[0580] Step 2:
[0581] Reminder generation (server)
[0582] The server generates a reminder: "It's time to take your medicine. Are you ready?"
[0583] Input: Medication time detection
[0584] Output: The generated reminder
[0585] Step 3:
[0586] Reminder notification (device)
[0587] The device will notify the elderly of reminders via voice.
[0588] Input: Reminder from server
[0589] Output: Voice notification to the elderly
[0590] Step 4:
[0591] User response (terminal)
[0592] The user responds, "Yes, I'll drink it now." The device uses a voice recognition system to convert the response into text data.
[0593] Input: User's voice response
[0594] Output: Text data
[0595] Step 5:
[0596] Medication confirmation (terminal)
[0597] The device uses a camera to record the patient taking the medicine.
[0598] Input: User's medication actions
[0599] Output: Video data
[0600] Step 6:
[0601] Data transmission (terminal)
[0602] The converted text data and video data are sent to the server.
[0603] Input: Text data, video data
[0604] Output: Data sent to the server
[0605] Step 7:
[0606] Data analysis (server)
[0607] The server analyzes the video data to confirm whether the user has taken the medicine, and also analyzes the user's emotional state using an emotion engine.
[0608] Input: Text data, video data
[0609] Output: Medication confirmation result, user's emotional state
[0610] Step 8:
[0611] Confirmation message generation (server)
[0612] The server generates a confirmation message saying "Your medication was taken correctly" and sets the next reminder.
[0613] Input: Medication confirmation result, user's emotional state
[0614] Output: Generated confirmation message, next reminder
[0615] Step 9:
[0616] Confirmation notification (terminal)
[0617] The device will then provide a confirmation message to the elderly person via voice.
[0618] Input: Confirmation message from the server
[0619] Output: Voice notification to the elderly
[0620] Dementia early detection function processing steps
[0621] Step 1:
[0622] Starting a daily conversation (terminal)
[0623] The device speaks to the elderly person, asking, "Tell us how your day was today."
[0624] Input: Pre-configured question prompt
[0625] Output: Start a conversation with the elderly person
[0626] Step 2:
[0627] User response (terminal)
[0628] Users talk about everyday events, and the device converts the conversation into text data using a voice recognition system.
[0629] Input: User's daily conversation
[0630] Output: Text data
[0631] Step 3:
[0632] Data transmission (terminal)
[0633] The converted text data is sent to the server.
[0634] Input: Text data
[0635] Output: Data sent to the server
[0636] Step 4:
[0637] Data analysis (server)
[0638] The server analyzes the text data, assessing language patterns and emotional states, and also analyzes emotional fluctuations using an emotion engine.
[0639] Input: Text data
[0640] Output: Language pattern evaluation results, emotional state evaluation results
[0641] Step 5:
[0642] Alert Generation (Server)
[0643] The server generates an alert such as, "Recently, we have noticed a decline in vocabulary in conversations."
[0644] Input: Language pattern evaluation results, emotional state evaluation results
[0645] Output: The generated alert
[0646] Step 6:
[0647] Alert notification (terminal)
[0648] The device will notify the senior of the alert, and if necessary, a caregiver or family member.
[0649] Input: Alert from the server
[0650] Output: Notification to seniors and caregivers
[0651] (Application example 2)
[0652] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0653] It is difficult to monitor the health and emotional states of elderly people in real time and provide appropriate advice and reminders when they go about their daily lives, especially when they use self-driving vehicles. Furthermore, conventional remote care systems do not provide support tailored to the elderly's situation in self-driving vehicles, which may reduce the elderly's sense of safety and security. Therefore, there is a need for a remote care system that can monitor the health and emotional states of elderly people in real time and provide appropriate support when they use self-driving vehicles.
[0654] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0655] In this invention, the server includes a voice recognition means for acquiring voice data of the elderly person, an image recognition means for acquiring image data of the elderly person, an analysis means using a generative AI model to analyze the voice data and image data, a notification means for notifying the elderly person of necessary advice and reminders based on the analysis results, a means for monitoring the elderly person's health and emotional state while in the autonomously driven vehicle, a means for transmitting the data collected by the above means to the server and analyzing it, a means for providing the elderly person with advice regarding their living conditions in the autonomously driven vehicle based on the analysis results, and a means for notifying the elderly person of the advice visually and audibly. This makes it possible to support elderly people in living safely and healthily when using autonomously driven vehicles.
[0656] The "voice recognition means" is a means for acquiring and analyzing voice data of the elderly person.
[0657] The "image recognition means" is a means for acquiring image data of an elderly person and analyzing it.
[0658] A "generative AI model" is an advanced artificial intelligence model that analyzes audio and image data and is used to assess the health and emotional state of an elderly person.
[0659] The "analysis means" is a means for assessing the health and emotional state of an elderly person using audio and image data.
[0660] The "notification means" is a means for visually and audibly notifying the elderly of necessary advice and reminders based on the analysis results.
[0661] An "automated vehicle" is a vehicle that is driven automatically and can be used by seniors.
[0662] "Health status" refers to the physical condition and physical function of an elderly person.
[0663] "Emotional state" refers to the psychological state and emotional movements of elderly people.
[0664] The "means for transmitting data to a server" is a means for transmitting voice data and image data collected from elderly people to a server.
[0665] "Visual and audio notification means" refers to means for conveying advice to the elderly person through a visual display and audio.
[0666] The present invention relates to a remote care system for elderly people to live safe and healthy lives when using autonomous vehicles. An embodiment of the present invention will be specifically described.
[0667] System configuration
[0668] Terminal configuration
[0669] 1. Voice recognition means:
[0670] The device is equipped with a microphone that is used to capture voice data from the elderly, which is then analyzed in real time.
[0671] 2. Image Recognition Methods:
[0672] The device is equipped with a camera that is used to capture image data of the elderly person and send it to a server. For example, the device captures the elderly person's facial expressions and body movements.
[0673] Server-side configuration
[0674] 1. Analysis method:
[0675] The server is equipped with a generative AI model that analyzes the voice and image data sent from the device, and this analysis evaluates the health and emotional state of the elderly person.
[0676] 2. Means of notification:
[0677] Based on the analysis results, the server generates necessary advice and reminders for the elderly, which are then sent to the device, which then notifies the elderly by voice or display.
[0678] Hardware and software used
[0679] Hardware:
[0680] Microphone: Used to collect voice data from the elderly.
[0681] Camera: Used to capture image data of the elderly.
[0682] Self-driving vehicles: Used to monitor the health and emotional state of elderly people in real time while they are in the vehicle.
[0683] software:
[0684] speech_recognition library: Used to achieve speech recognition.
[0685] OpenCV library: Used to realize image processing.
[0686] TensorFlow: Used to realize a generative AI model for elderly emotion engine analysis.
[0687] gTTS: Used to generate audio notifications.
[0688] requests: Used to realize data transmission to the server.
[0689] System operation example
[0690] 1. While the elderly person is riding in a self-driving vehicle:
[0691] Microphones and cameras capture the elderly person's voice and image data in real time, which is automatically sent to a server.
[0692] The server uses a generative AI model to analyze audio and image data to assess the health and emotional state of the elderly.
[0693] Based on the assessment results, necessary advice and reminders are generated and sent to the elderly via the device. For example, if a person does not eat enough breakfast, they will receive a voice message saying, "We recommend adding more fruits and vegetables to your next meal."
[0694] Prompt Sentence Examples
[0695] After the elderly person answers questions about health care in the self-driving car, please send the image and voice data to a server and generate a Python program that will provide appropriate advice and reminders.
[0696] As described above, the remote care system of the present invention provides multifaceted support to enable elderly people to use self-driving vehicles with peace of mind.
[0697] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0698] Step 1: Acquire audio data
[0699] The device uses a microphone to capture the elderly person's voice data. The input is the elderly person's voice, which is collected in real time. The device stores the collected voice data either directly or temporarily.
[0700] Step 2: Acquiring image data
[0701] The device uses a camera to acquire image data (facial expressions and body movements) of the elderly. The input is a real-time image of the elderly acquired through the camera, which is captured at high resolution. The device stores the captured image data either directly or temporarily.
[0702] Step 3: Send data to the server
[0703] The terminal sends the collected voice and image data to the server. The input is the data acquired in steps 1 and 2 above, and this is sent to the server via the network. The output is the data sent to the server itself.
[0704] Step 4: Analyzing the audio data
[0705] The server analyzes the received voice data. The input is the voice data sent from the device, which is analyzed using a generative AI model. The content of the elderly person's speech and their emotional state are extracted from the voice data and output as text data.
[0706] Step 5: Analyzing the image data
[0707] The server analyzes the received image data. The input is image data sent from the device, which is analyzed using a generative AI model. The image data is used to evaluate the elderly person's facial expressions, body movements, health condition, etc., and is output as numerical data or evaluation results.
[0708] Step 6: Integrating the analysis results
[0709] The server integrates the results of the voice data analysis and the image data analysis to evaluate the elderly person's comprehensive health and emotional state. The input is the analysis results of steps 4 and 5, which are integrated to comprehensively grasp the elderly person's current condition. The output is the integrated evaluation result.
[0710] Step 7: Generate advice and reminders
[0711] The server generates necessary advice and reminders for the elderly based on the integrated evaluation results. The input is the integrated evaluation results, and appropriate messages are generated based on them. The output is text data of the generated advice and reminders.
[0712] Step 8: Implementing Notification
[0713] The device notifies the elderly of advice and reminders received from the server. The input is text data sent from the server, which is notified visually and audibly. Specifically, messages are displayed on the display and read aloud using gTTS. The output is the advice and reminders notified to the elderly.
[0714] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0715] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0716] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0717] [Second embodiment]
[0718] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0719] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0720] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0721] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0722] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0723] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0724] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0725] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0726] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0727] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0728] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0729] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0730] This remote care system consists of elderly end users, and a server and terminal (robot) located in a remote location. A specific embodiment of the system is described below. The program code itself is not provided, but the processing flow and each step are described in detail.
[0731] System configuration
[0732] 1. Terminal configuration
[0733] Speech recognition means: The device is equipped with a microphone to collect conversations with the elderly in real time, and speech recognition technology is used to convert the captured speech into text data.
[0734] Image recognition method: The device is equipped with a camera that captures images of the elderly person's daily life, especially when they eat or take medicine. The image data is sent to a server in real time.
[0735] 2. Server-side configuration
[0736] Analysis method: The server uses a generative AI model to analyze the voice and image data sent from the device. This analysis allows for a real-time assessment of the elderly person's health and living conditions.
[0737] Notification method: Based on the analysis results, appropriate advice and reminders are generated and sent to the device, which then communicates this to the elderly via voice.
[0738] Meal management function for the elderly
[0739] Overview: The dietary management function monitors whether the elderly are eating a balanced diet.
[0740] Terminal: Ask the senior, "What did you have for breakfast today?"
[0741] User: "I had bread and milk."
[0742] Device: Converts voice into text data and simultaneously takes pictures of the meal using a camera.
[0743] Server: Analyzes the received text and image data and evaluates the calorie and nutritional balance of the meal.
[0744] Server: Generates advice as needed, such as "Today's breakfast appears to be low in calories. We recommend eating some fruits and vegetables at your next meal."
[0745] Device: Provides advice to the elderly via voice.
[0746] Medication management function
[0747] Overview: Medication management helps seniors take their medications appropriately.
[0748] Server: Detects when it's time to take your medicine and generates a reminder saying, "It's time to take your medicine. Are you ready?"
[0749] Device: Provides voice reminders to seniors.
[0750] User: "Yes, I'll take my medicine."
[0751] Device: Activate the camera and record the scene of taking the medicine.
[0752] Server: Analyzes image data and verifies whether the medication has been taken.
[0753] Server: If the medication is confirmed, generate a confirmation message saying "Your medication was taken correctly" and set the next reminder.
[0754] Terminal: Notify the elderly person of a confirmation message.
[0755] Early detection of dementia
[0756] Overview: The early dementia detection function detects signs of dementia from the everyday conversations of elderly people.
[0757] Device: Initiate everyday conversations such as "Tell me how your day was."
[0758] User: Talk about everyday events.
[0759] Terminal: Converts the voice into text data and sends it to the server.
[0760] Server: Analyzes conversation data and evaluates language patterns and emotional states, for example, detecting vocabulary declines and changes in emotional tone.
[0761] Server: If signs of dementia are detected, it generates an alert such as, "Recently, there has been a decline in vocabulary in conversations."
[0762] Device: Notifies the senior of the alert and, if necessary, notifies caregivers and family members.
[0763] As described above, this remote care system provides comprehensive support for the daily lives of elderly people and makes it easier for caregivers and family members in remote locations to understand the condition of the elderly, thereby improving the quality of life of the elderly and reducing the burden on caregivers.
[0764] The processing flow will be explained below.
[0765] Meal management function for the elderly
[0766] Step 1:
[0767] The device asks the senior, "What did you have for breakfast today?"
[0768] Step 2:
[0769] The user responds, "I had bread and milk."
[0770] Step 3:
[0771] The terminal uses a voice recognition system to convert the user's speech into text data.
[0772] Step 4:
[0773] The device activates the camera and takes pictures of the elderly person eating.
[0774] Step 5:
[0775] The terminal transmits the acquired text data and image data to the server.
[0776] Step 6:
[0777] The server analyzes the received text data and extracts the meal details.
[0778] Step 7:
[0779] The server analyzes the received image data and calculates the calories and nutritional balance of the meal contents.
[0780] Step 8:
[0781] Based on the analysis results, the server generates advice such as, "Today's breakfast appears to be low in calories. We recommend that you eat fruits and vegetables at your next meal."
[0782] Step 9:
[0783] The server sends the generated advice to the terminal.
[0784] Step 10:
[0785] The device will provide advice to the elderly via voice.
[0786] Medication management function
[0787] Step 1:
[0788] The server recognizes pre-set medication times and generates reminders.
[0789] Step 2:
[0790] The device will then send a voice reminder to the elderly person saying, "It's time to take your medicine. Are you ready?"
[0791] Step 3:
[0792] The user responds, "Yes, I'll take my medicine."
[0793] Step 4:
[0794] The terminal uses a voice recognition system to convert the user's response into text data.
[0795] Step 5:
[0796] The device activates the camera and takes a picture of the elderly person taking their medicine.
[0797] Step 6:
[0798] The video data acquired by the terminal is transmitted to the server.
[0799] Step 7:
[0800] The server analyzes the video data and confirms that the medication has been taken.
[0801] Step 8:
[0802] The server generates a confirmation message saying "Your medication was taken correctly" and sets the next reminder.
[0803] Step 9:
[0804] The server sends a confirmation message to the terminal.
[0805] Step 10:
[0806] The device will then send a confirmation message to the elderly person.
[0807] Early detection of dementia
[0808] Step 1:
[0809] The device begins a casual conversation with the elderly person, asking, "Tell me how your day was today."
[0810] Step 2:
[0811] Users talk about everyday events.
[0812] Step 3:
[0813] The terminal uses a voice recognition system to convert the user's speech into text data.
[0814] Step 4:
[0815] The terminal transmits the acquired text data to the server.
[0816] Step 5:
[0817] The server analyzes the text data and evaluates the elderly person's language patterns and emotional state.
[0818] Step 6:
[0819] The server detects signs of dementia based on the analysis results.
[0820] Step 7:
[0821] The server generates an alert such as, "Recently, we have noticed a decline in vocabulary in conversations."
[0822] Step 8:
[0823] The server generates alerts and sends them to the device, which also notifies family members and caregivers.
[0824] Step 9:
[0825] The device will notify the elderly person of the alert.
[0826] These processing steps are designed to ensure that the lives of the elderly are managed effectively and to reduce the burden on caregivers.
[0827] Example 1
[0828] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0829] It is difficult for elderly people to maintain appropriate lifestyle habits and monitor their health. In particular, it is important for elderly people to manage the nutritional balance of their diet, their medication status, and even to detect dementia early, but it is difficult for them to do these things alone. It is also not easy for caregivers and family members in remote locations to understand the condition of elderly people. Conventional systems lack the functionality to comprehensively support the lives of elderly people.
[0830] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0831] In this invention, the server includes a voice recognition means for acquiring voice data of the elderly person, an image recognition means for acquiring image data of the elderly person, an analysis means using a generative AI model to analyze the voice data and image data, a notification means for notifying the elderly person of necessary advice and reminders based on the analysis results, a means for detecting signs of dementia from the elderly person's everyday conversation and generating an alert, a means for reminding the elderly person to take their medicine and confirming that they have taken it, and a means for monitoring the elderly person's diet to evaluate nutritional balance and provide advice. This makes it possible to comprehensively support the elderly's living conditions and enable caregivers and family members in remote locations to easily understand the elderly person's condition.
[0832] The "voice recognition means" is a device or technology for capturing the voice of the elderly person in real time and converting the voice data into text data.
[0833] "Image recognition means" refers to a device or technology for photographing the elderly person's daily life and acquiring image data.
[0834] A "generative AI model" is an artificial intelligence model that analyzes voice and image data and generates appropriate advice and reminders based on the results.
[0835] The "analysis means" is a means of evaluating the health status and living conditions of elderly people using a generative AI model, using the acquired voice and image data.
[0836] The "notification means" is a means of notifying the elderly of advice and reminders generated based on the analysis results via voice or other means.
[0837] The "means for detecting signs of dementia" is a method for analyzing data on everyday conversations of elderly people and detecting early signs of dementia from factors such as a decrease in vocabulary and changes in emotional tone.
[0838] "Reminder measures" are measures that notify elderly people when it is time to take their medicine and help them ensure that they take their medicine.
[0839] A "dietary monitoring method" is a method for monitoring the dietary content of elderly people, evaluating their nutritional balance and calories, and providing necessary advice.
[0840] This invention is a remote care system for monitoring and supporting the living conditions of elderly people. The system consists of an elderly person in a remote location, a server, and a terminal (robot). This system is configured and operates as follows to monitor the elderly person's daily behavior and health status and provide necessary advice and reminders.
[0841] System configuration
[0842] Terminal configuration
[0843] Speech recognition method: The device is equipped with a microphone to collect conversations with the elderly in real time. The captured speech is converted into text data using speech recognition technology (e.g., Google Cloud Speech-to-Text).
[0844] Image recognition method: The device is equipped with a camera that captures images of the elderly person's daily life, especially when they eat or take medicine. The image data is sent to a server in real time.
[0845] Server-side configuration
[0846] Analysis method: The server uses a generative AI model (e.g., OpenAI GPT-4) to analyze the voice and image data sent from the device. This analysis allows for a real-time assessment of the elderly person's health and living conditions.
[0847] Notification method: Based on the analysis results, appropriate advice and reminders are generated and sent to the device, which then communicates this to the elderly via voice.
[0848] Meal management function for the elderly
[0849] Overview: The dietary management function monitors whether the elderly are eating a balanced diet.
[0850] Terminal: Ask the senior, "What did you have for breakfast today?"
[0851] User: "I had bread and milk."
[0852] Device: Converts voice into text data and simultaneously takes pictures of the meal using a camera.
[0853] Server: Analyzes the received text and image data and evaluates the calorie and nutritional balance of the meal.
[0854] Server: Generates advice as needed, such as "Today's breakfast appears to be low in calories. We recommend eating some fruits and vegetables at your next meal."
[0855] Device: Provides advice to the elderly via voice.
[0856] Specific prompt examples:
[0857] "If an elderly person eats bread and milk for breakfast, generate a sentence that evaluates the nutritional balance and advises them to add more fruits and vegetables to their next meal if the calories are low."
[0858] Medication management function
[0859] Overview: Medication management helps seniors take their medications appropriately.
[0860] Server: Detects when it's time to take your medicine and generates a reminder saying, "It's time to take your medicine. Are you ready?"
[0861] Device: Provides voice reminders to seniors.
[0862] User: "Yes, I'll take my medicine."
[0863] Device: Activate the camera and record the scene of taking the medicine.
[0864] Server: Analyzes image data and verifies whether the medication has been taken.
[0865] Server: If the medication is confirmed, generate a confirmation message saying "Your medication was taken correctly" and set the next reminder.
[0866] Terminal: Notify the elderly person of a confirmation message.
[0867] Specific prompt examples:
[0868] "Generate a reminder when it's time for the senior to take their medication, and then create a sentence to confirm if they actually took their medication."
[0869] Early detection of dementia
[0870] Overview: The early dementia detection function detects signs of dementia from the everyday conversations of elderly people.
[0871] Device: Initiate everyday conversations such as "Tell me how your day was."
[0872] User: Talk about everyday events.
[0873] Terminal: Converts the voice into text data and sends it to the server.
[0874] Server: Analyzes conversation data and evaluates language patterns and emotional states, for example, detecting vocabulary declines and changes in emotional tone.
[0875] Server: If signs of dementia are detected, it generates an alert such as, "Recently, there has been a decline in vocabulary in conversations."
[0876] Device: Notifies the senior of the alert and, if necessary, notifies caregivers and family members.
[0877] Specific prompt examples:
[0878] "To detect signs of dementia in everyday conversations of elderly people, create sentences that analyze vocabulary decline and changes in emotional tone."
[0879] As described above, this remote care system provides comprehensive support for the daily lives of elderly people and makes it easier for caregivers and family members in remote locations to understand the condition of the elderly, thereby improving the quality of life of the elderly and reducing the burden on caregivers.
[0880] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0881] Meal management function for the elderly
[0882] Step 1:
[0883] The device asks the senior, "What did you have for breakfast today?"
[0884] Input: Breakfast Question
[0885] Output: Waiting for response from the elderly person
[0886] Step 2:
[0887] The user responds, "I had bread and milk."
[0888] Input: Elderly person's voice response
[0889] Output: Audio data
[0890] Step 3:
[0891] The device converts the voice into text data and simultaneously captures the meal with a camera.
[0892] Input: Audio data, video of elderly people eating
[0893] Output: Text data, image data
[0894] What it does: The microphone captures your voice and converts it into text using a speech recognition engine. The camera captures footage of you eating.
[0895] Step 4:
[0896] The server receives the text and image data and analyzes it using a generative AI model.
[0897] Input: Text data, image data
[0898] Output: Calorie and nutritional balance evaluation results of the meal
[0899] How it works: The server extracts meal items from text data and analyzes the meal contents using an image recognition algorithm. The generative AI model calculates nutritional balance and calories.
[0900] Step 5:
[0901] The server generates advice based on the analysis results.
[0902] Input: Evaluation results of calorie and nutritional balance of meals
[0903] Output: Advice message
[0904] Specific behavior: The server generates advice such as "Today's breakfast is low in calories. We recommend that you eat fruits and vegetables at your next meal."
[0905] Step 6:
[0906] The device will provide advice to the elderly via voice.
[0907] Input: Advice message
[0908] Output: Voice notification to the elderly
[0909] Specific operation: Advice messages are synthesized and conveyed to the elderly.
[0910] Medication management function
[0911] Step 1:
[0912] The server detects when it's time to take your medicine and generates a reminder saying, "It's time to take your medicine. Are you ready?"
[0913] Input: Medication schedule
[0914] Output: Reminder message
[0915] What it does: The server periodically checks the schedule and generates a reminder message when it's time to take your medication.
[0916] Step 2:
[0917] The device will notify the elderly of the reminder via voice.
[0918] Input: Reminder message
[0919] Output: Voice notification to the elderly
[0920] Specific actions: Reminder messages are delivered to elderly people via voice synthesis.
[0921] Step 3:
[0922] The user responds, "Yes, I'll take my medicine."
[0923] Input: Elderly person's voice response
[0924] Output: Response audio data
[0925] Step 4:
[0926] The device activates the camera and captures the scene of the elderly person taking their medicine.
[0927] Input: Video of elderly person taking medication
[0928] Output: Image data of the medication scene
[0929] Specific operation: The camera captures the elderly person taking their medicine and saves the image data on the device.
[0930] Step 5:
[0931] The server analyzes the image data and verifies whether the medication has been taken.
[0932] Input: Image data of a medication scene
[0933] Output: Medication confirmation result
[0934] What it does: The server uses image recognition technology to verify that the medication was taken correctly.
[0935] Step 6:
[0936] The server generates a confirmation message and sets the next reminder.
[0937] Input: Medication confirmation result
[0938] Output: Confirmation message, next reminder setting
[0939] Specific behavior: If the medication is confirmed, the server generates a confirmation message saying "Your medication was taken correctly" and sets the next reminder.
[0940] Step 7:
[0941] The device will then send a confirmation message to the elderly person.
[0942] Input: Confirmation message
[0943] Output: Voice notification to the elderly
[0944] Specific operation: A confirmation message is synthesized and conveyed to the elderly person.
[0945] Early detection of dementia
[0946] Step 1:
[0947] The device will begin a casual conversation such as, "Tell me how your day was today."
[0948] Input: Daily conversation question
[0949] Output: Waiting for response from the elderly person
[0950] Step 2:
[0951] Users talk about everyday events.
[0952] Input: Elderly person's voice
[0953] Output: Audio data
[0954] Step 3:
[0955] The device converts the voice into text data and sends it to the server.
[0956] Input: Audio data
[0957] Output: Text data
[0958] What it does: The microphone captures your voice and the speech recognition engine converts it into text.
[0959] Step 4:
[0960] The server analyzes the conversation data and evaluates language patterns and emotional states.
[0961] Input: Text data
[0962] Output: Evaluation result
[0963] Specific operation: The server uses the generative AI model to detect vocabulary loss, changes in emotional tone, etc.
[0964] Step 5:
[0965] If the server detects signs of dementia, it generates an alert message.
[0966] Input: Evaluation result
[0967] Output: Alert message
[0968] Specific behavior: If signs of dementia are detected, an alert will be generated, such as "Recently, there has been a decline in vocabulary in conversations."
[0969] Step 6:
[0970] The device notifies the elderly person of the alert and, if necessary, notifies caregivers and family members.
[0971] Input: Alert message
[0972] Output: Audio notification to the senior and caregiver / family
[0973] Specific operation: An alert message is synthesized and conveyed to the elderly. At the same time, an alert is sent to caregivers and family members.
[0974] The above are the specific processing steps and their detailed operations. These steps improve the quality of life for the elderly and enable appropriate care to be provided remotely.
[0975] (Application example 1)
[0976] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0977] Remote care systems that support the lives of the elderly are limited in their management functions to the health status of the elderly, and do not address comprehensive home safety. There is a need for a system that comprehensively monitors security aspects such as managing entry and exit within the home and detecting fires, gas leaks, and suspicious individuals. The objective of this invention is to provide a comprehensive support system that not only manages the health of the elderly, but also monitors the safety status of the home.
[0978] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0979] In this invention, the server includes an analysis means, a notification means, and a safety monitoring means, which makes it possible not only to grasp and support the living conditions of the elderly, but also to comprehensively monitor the safety of the home.
[0980] A "remote care system" is a system that monitors and supports the living conditions of elderly people from a remote location.
[0981] The "voice recognition means" is a device that has the function of acquiring voice data from the elderly person and converting it into text data in real time.
[0982] The "image recognition means" is a device that has the function of acquiring image data of an elderly person and analyzing the content of that image data.
[0983] A "generative AI model" is an artificial intelligence model that analyzes acquired voice data and image data and generates the necessary information.
[0984] The "analysis means" is a function that analyzes voice data and image data to evaluate the living conditions of the elderly person.
[0985] "Notification means" is a function that notifies elderly people of advice and reminders based on the analysis results.
[0986] The "safety monitoring means" is a device that has the function of monitoring the safety status of the home, detecting abnormalities, and notifying them.
[0987] "Nutritional balance" is an indicator of whether nutrients are evenly distributed in a particular diet.
[0988] A "calorie" is a unit that indicates the amount of energy ingested through food.
[0989] A "reminder" is a notification that encourages seniors to take a specific action.
[0990] This remote care system is a system for monitoring and supporting the living conditions and home safety conditions of elderly people. A specific embodiment of the system will be described below.
[0991] System configuration
[0992] The server and user terminals (smartphones) work together to monitor the lives of the elderly and the safety of their homes using various sensor devices. This system includes the following hardware and software:
[0993] Hardware
[0994] Smartphone: microphone, camera
[0995] Sensors: Door and window sensors, smoke detectors, gas leak detectors
[0996] software
[0997] Speech recognition API: Used to convert voice data into text data. Typical examples include Google Speech-to-Text and Apple Siri.
[0998] Image recognition API: Used to analyze image data. Representative examples include Google Cloud Vision and Amazon Rekognition.
[0999] Server: Platforms that run generative AI models for data analysis include Google Cloud and AWS (Amazon Web Services).
[1000] System Operation
[1001] The server uses voice recognition, image recognition, and safety monitoring means to notify the user based on the analysis results. Each of these means is described in detail below.
[1002] Voice recognition means
[1003] The microphone on the user's device collects the elderly's voice and converts it into text data in real time using a speech recognition API. For example, if an elderly person says "help me," it is converted into text data and sent to the server.
[1004] Image Recognition Method
[1005] The camera on the user's device captures images of the elderly's living conditions and home situation, and the images are analyzed using an image recognition API. Examples include detecting fires, gas leaks, and suspicious people. It also includes capturing images of elderly people taking medicine, and analyzing the image data on a server.
[1006] Safety monitoring means
[1007] Sensor devices such as door and window sensors, smoke detectors, and gas leak detectors monitor the safety status of the home and notify the server if an abnormality is detected. The server analyzes the abnormality and issues a warning to the user in real time.
[1008] Server analysis method
[1009] The server analyzes the voice and image data using a generative AI model, and based on the analysis results, generates appropriate advice and reminders and sends them to the user's device.
[1010] Notification means
[1011] Based on the analysis results, the system will provide voice reminders and warnings to the elderly, such as specific instructions such as "There is a fire. Please close the doors of each room."
[1012] Specific examples
[1013] In case of fire detection
[1014] 1. The microphone detects the sound of a fire alarm.
[1015] 2. The voice recognition API converts the message "Fire has broken out" into text data.
[1016] 3. Analyze fire data on the server.
[1017] 4. An "emergency fire alert" notification will be sent to your smartphone.
[1018] 5. An audio guide will be given to residents to "close the doors of each room."
[1019] Prompt Sentence Examples
[1020] "Design a program that uses a voice recognition system to detect fire alarm sounds in real time and notify a smartphone of that information. If an abnormal sound is detected, provide voice guidance on first aid measures for the fire."
[1021] This will make it possible to realize a system that comprehensively supports the lives and home safety of the elderly.
[1022] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1023] Step 1:
[1024] The microphone on the user's device captures the elderly person's voice data, which is then sent to a voice recognition API and converted into text data in real time.
[1025] Input: Elderly voice data
[1026] Output: Text data
[1027] Specific operation: When an elderly person says "help me," their voice is picked up by a microphone and converted into text data using a voice recognition API.
[1028] Step 2:
[1029] The camera on the user's device captures image data of the elderly person, which is then sent to an image recognition API for real-time analysis.
[1030] Input: Image data of elderly people
[1031] Output: Analysis data
[1032] Specific operation: The camera captures scenes of elderly people eating or taking medicine, and the image data is analyzed using an image recognition API.
[1033] Step 3:
[1034] The server receives the voice and image data sent from the user's device and analyzes the data using a generative AI model.
[1035] Input: Text data, analysis data
[1036] Output: Analysis results
[1037] Specific operation: The voice data is converted into text and the analysis results of the image data are sent to a server, and the generated AI model analyzes the elderly person's living conditions and health condition.
[1038] Step 4:
[1039] The server generates appropriate advice and reminders based on the analysis results.
[1040] Input: Analysis results
[1041] Output: Advice, reminder
[1042] Specific behavior: For example, if an elderly person's breakfast is analyzed to be insufficient in calories, advice such as "We recommend that you eat fruits and vegetables at your next meal" will be generated.
[1043] Step 5:
[1044] The server transmits the generated advice and reminders to the user terminal using a notification means.
[1045] Input: Advice, Reminder
[1046] Output: Notification data
[1047] Specific operation: The generated advice and reminders are sent to the user's device as push notifications, and the elderly person receives notifications such as "It's time to take their medicine."
[1048] Step 6:
[1049] Sensors on the user's device (doors, windows, smoke, gas leaks) monitor the safety status of the home, and if an abnormality is detected, the data is sent to the server.
[1050] Input: Sensor data
[1051] Output: Anomaly detection data
[1052] Specific operation: If a door or window is opened or closed suspiciously, or if a fire or gas leak is detected, the information is sent to the server.
[1053] Step 7:
[1054] The server analyzes the anomaly detection data, generates necessary warnings, and sends them to the user terminal.
[1055] Input: Anomaly detection data
[1056] Output: Warning notification data
[1057] Specific operation: For example, a warning message such as "A fire has been detected. Please close the doors of each room" is generated and sent to the user terminal.
[1058] Step 8:
[1059] The user terminal notifies the elderly person of the received warning notification data by voice in real time.
[1060] Input: Alert notification data
[1061] Output: Audio notification
[1062] Specific actions: The elderly person will be notified by voice, "A fire has broken out. Please close the doors of each room."
[1063] This will enable comprehensive monitoring of the elderly's living conditions and home safety, and provide appropriate support at the appropriate time.
[1064] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1065] This invention is a remote care system that monitors the living conditions of elderly people and provides appropriate support, and is equipped with an emotion engine that analyzes voice data and image data of the elderly person and recognizes the emotional state of the user. Specific embodiments of the invention will be described below.
[1066] System configuration
[1067] 1. Terminal configuration
[1068] Voice recognition means: The device is equipped with a microphone to collect conversations with the elderly in real time. The collected voice data is converted into text data by a voice recognition system.
[1069] Image recognition means: The device is equipped with a camera that takes pictures of, for example, an elderly person eating or taking medicine. The captured image data is sent to a server in real time.
[1070] 2. Server-side configuration
[1071] Analysis method: The server is equipped with a generative AI model with advanced analytical capabilities that analyzes the voice and image data sent from the device. This analysis allows for a real-time assessment of the elderly person's health and living conditions.
[1072] Emotion engine: The server is also equipped with an emotion engine that recognizes the emotional state of the elderly by analyzing voice data.
[1073] 3. Means of notification
[1074] Advice and reminders: Based on the analysis results and emotional state, advice and reminders are generated for the elderly. This information is sent from the server to the device, which then notifies the elderly by voice.
[1075] Meal management function for the elderly
[1076] Overview: The dietary management function monitors whether the elderly are eating a balanced diet and provides advice as needed.
[1077] Terminal: Ask the senior, "What did you have for breakfast today?"
[1078] User: "I had bread and milk."
[1079] Device: A voice recognition system converts the user's speech into text data. A camera captures the meal.
[1080] Server: Analyzes text and image data to evaluate the calorie and nutritional balance of meal content. Also analyzes the user's emotional state using an emotion engine.
[1081] Server: Generate advice such as, "Your breakfast today seems low in calories. We recommend adding more fruits and vegetables to your next meal."
[1082] Device: Provides advice to the elderly via voice.
[1083] Medication management function
[1084] Abstract: Medication management is a function that helps elderly people take their medications appropriately.
[1085] Server: Detects medication times and generates reminders.
[1086] Device: Reminds seniors, "It's time to take your medicine. Are you ready?"
[1087] User: "Yes, I'll drink it now."
[1088] Device: A voice recognition system converts responses into text data, and a camera captures the process of taking the medicine.
[1089] Server: Analyzes video data and confirms whether the user has taken the medication. An emotion engine also analyzes the user's emotional state.
[1090] Server: Generates a confirmation message saying "Your medication was taken correctly" and sets the next reminder.
[1091] Terminal: Notify the elderly person of a confirmation message.
[1092] Early detection of dementia
[1093] Overview: The early dementia detection function detects signs of dementia from the everyday conversations of elderly people.
[1094] Device: Begins a conversation with "Tell me how your day was."
[1095] User: Talk about everyday events.
[1096] Terminal: The speech recognition system converts the conversation into text data and sends it to the server.
[1097] Server: Analyzes text data and evaluates language patterns and emotional states. An emotion engine also analyzes emotional fluctuations.
[1098] Server: Generate an alert such as, "Recently, we have noticed a decline in vocabulary in conversations."
[1099] Device: Notifies the senior of the alert and, if necessary, notifies caregivers and family members.
[1100] In this way, the remote care system of the present invention supports the lives of elderly people in a variety of ways, and by combining it with an emotion engine, it is possible to provide more accurate advice and appropriate care.
[1101] The processing flow will be explained below.
[1102] Meal management function for the elderly
[1103] Step 1:
[1104] The device asks the senior, "What did you have for breakfast today?"
[1105] Step 2:
[1106] The user responds, "I had bread and milk."
[1107] Step 3:
[1108] The terminal uses a voice recognition system to convert the user's speech into text data.
[1109] Step 4:
[1110] The device activates the camera and takes pictures of the elderly person eating.
[1111] Step 5:
[1112] The terminal transmits the acquired text data and image data to the server.
[1113] Step 6:
[1114] The server analyzes the received text data and extracts the meal details.
[1115] Step 7:
[1116] The server analyzes the received image data and calculates the calories and nutritional balance of the meal contents.
[1117] Step 8:
[1118] The server uses an emotion engine to recognize the user's emotional state from the voice data.
[1119] Step 9:
[1120] Based on the analysis results and emotional state, the server generates advice such as, "Today's breakfast seems low in calories. I recommend eating fruits and vegetables at your next meal."
[1121] Step 10:
[1122] The server sends the generated advice to the terminal.
[1123] Step 11:
[1124] The device will provide advice to the elderly via voice.
[1125] Medication management function
[1126] Step 1:
[1127] The server recognizes pre-set medication times and generates reminders.
[1128] Step 2:
[1129] The device will then send a voice reminder to the elderly person saying, "It's time to take your medicine. Are you ready?"
[1130] Step 3:
[1131] The user responds, "Yes, I'll take my medicine."
[1132] Step 4:
[1133] The terminal uses a voice recognition system to convert the user's response into text data.
[1134] Step 5:
[1135] The device activates the camera and takes a picture of the elderly person taking their medicine.
[1136] Step 6:
[1137] The video data acquired by the terminal is transmitted to the server.
[1138] Step 7:
[1139] The server analyzes the video data and confirms that the medication has been taken.
[1140] Step 8:
[1141] The server uses an emotion engine to analyze the user's emotional state.
[1142] Step 9:
[1143] The server generates a confirmation message saying "Your medication was taken correctly" and sets the next reminder.
[1144] Step 10:
[1145] The server sends a confirmation message to the terminal.
[1146] Step 11:
[1147] The device will then send a confirmation message to the elderly person.
[1148] Early detection of dementia
[1149] Step 1:
[1150] The device begins a casual conversation by asking, "Tell me how your day was today."
[1151] Step 2:
[1152] Users talk about everyday events.
[1153] Step 3:
[1154] The terminal uses a voice recognition system to convert the user's speech into text data.
[1155] Step 4:
[1156] The terminal transmits the acquired text data to the server.
[1157] Step 5:
[1158] The server analyzes the text data and evaluates the elderly person's language patterns.
[1159] Step 6:
[1160] The server uses an emotion engine to analyze the user's emotional state from their speech.
[1161] Step 7:
[1162] Based on the analysis results and emotional state, the server generates an alert such as, "Recently, there has been a decrease in vocabulary in conversations."
[1163] Step 8:
[1164] The server generates an alert and sends it to the device.
[1165] Step 9:
[1166] The device will notify the elderly person of the alert.
[1167] Step 10:
[1168] The server will also notify caregivers and family members of the alerts as needed.
[1169] These processing steps enable the remote care system of the present invention to monitor the user's living conditions with high accuracy and provide appropriate care and support. In addition, the emotion engine enables flexible responses according to the user's emotional state, improving the user's quality of life and reducing the burden on caregivers.
[1170] Example 2
[1171] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1172] For elderly people to maintain independent lifestyles, it is important to properly monitor their daily health and emotional states and provide timely advice and reminders. However, current systems have difficulty effectively analyzing the subject's voice and images and providing appropriate advice and reminders that take their emotional state into account. Furthermore, the lack of personalized support tailored to each individual's health condition and specific lifestyle circumstances makes it difficult for elderly people to maintain appropriate diets and medication intake.
[1173] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1174] In this invention, the server includes an analysis means that uses a generative AI model to analyze voice data and image data and evaluate health conditions and living situations, an emotion engine that recognizes emotional states from voice data, a means that analyzes dietary content and medication status to generate advice and reminders necessary for the elderly, and a notification means that notifies the elderly of the generated advice and reminders by voice. This enables comprehensive analysis using the voices and images of the elderly, and makes it possible to provide appropriate, personalized advice and reminders.
[1175] "Speech recognition means" is a technology that collects speech and converts it into text data.
[1176] "Image recognition means" is a technology that collects images and sends the data to a server.
[1177] A "generative AI model" is an artificial intelligence technology that analyzes voice and image data to assess the health and living conditions of elderly people.
[1178] "Analysis means" refers to the process of analyzing and evaluating audio and image data using a generative AI model.
[1179] "Emotion engine" is a technology that recognizes the user's emotional state from voice data.
[1180] The "notification means" is a technology that notifies the elderly of the generated advice and reminders by voice.
[1181] A "reminder" is a notification that helps seniors remember to take certain actions, such as taking medicine.
[1182] "Text data" refers to character data converted from speech by speech recognition means.
[1183] The present invention is a remote care system that monitors the living conditions of elderly people and provides appropriate support, and is equipped with an emotion engine that analyzes voice data and image data of the elderly person and recognizes the emotional state of the user. Specific embodiments of the present invention will be described below.
[1184] System configuration
[1185] This system consists of a voice recognition means, an image recognition means, an analysis means, an emotion engine, and a notification means.
[1186] Voice recognition means
[1187] The device is equipped with a microphone that collects conversations with elderly people in real time. The collected voice data is converted into text data through a voice recognition system, allowing the user's speech to be captured as text information.
[1188] Image Recognition Method
[1189] The device is equipped with a camera that takes pictures of the elderly person eating and taking their medicine. The captured image data is sent to a server in real time, allowing for a visual record of the elderly person's behavior.
[1190] Analysis means
[1191] The server is equipped with a generative AI model with advanced analytical capabilities that analyzes the voice and image data sent from the device. This analysis makes it possible to evaluate the health and living conditions of elderly people in real time. The generative AI model performs highly accurate analysis of the data, using techniques such as deep learning and natural language processing.
[1192] Emotion Engine
[1193] The server's emotion engine has the ability to recognize the user's emotional state from voice data, allowing it to assess not just their physical state but also their psychological state. Emotion analysis is performed based on the tone of voice and the choice of words used.
[1194] Notification means
[1195] Based on the analysis results and emotional state, the server generates advice and reminders for the elderly, which are then sent to the device, which then notifies the elderly by voice.
[1196] Meal management function
[1197] Overview: The dietary management function monitors whether the elderly are eating a balanced diet and provides advice as needed.
[1198] The device asks the senior, "What did you have for breakfast today?"
[1199] The user responds, "I had bread and milk."
[1200] The device uses a voice recognition system to convert the user's speech into text data and a camera to capture the meal.
[1201] The server analyzes text and image data to evaluate the calorie and nutritional balance of the meal, and also analyzes the user's emotional state using an emotion engine.
[1202] The server generates advice such as, "Today's breakfast seems low in calories. We recommend adding more fruits and vegetables to your next meal."
[1203] The device will provide advice to the elderly via voice.
[1204] Example prompt sentence:
[1205] Prompts for older adults regarding dietary management functions:
[1206] Please explain the analysis process and advice generation steps when an elderly person answers "I had bread and milk" to the question "What did you have for breakfast today?"
[1207] Medication management function
[1208] Abstract: Medication management is a function that helps elderly people take their medications appropriately.
[1209] The server detects when it's time to take the medication and generates a reminder.
[1210] The device will remind the elderly person, "It's time to take your medicine. Are you ready?"
[1211] The user responds, "Yes, I'll drink it now."
[1212] The device uses a voice recognition system to convert responses into text data and a camera to record the patient taking the medicine.
[1213] The server analyzes the video data to confirm whether the user has taken the medicine, and also analyzes the user's emotional state using an emotion engine.
[1214] The server generates a confirmation message saying "Your medication was taken correctly" and sets the next reminder.
[1215] The device will then send a confirmation message to the elderly person.
[1216] Example prompt sentence:
[1217] Prompts regarding medication management functions for seniors:
[1218] Please explain the analysis process and reminder generation steps when an elderly person responds "Yes, I'll take it right away" to the reminder "It's time to take your medicine. Are you ready?"
[1219] Early detection of dementia
[1220] Overview: The early dementia detection function detects signs of dementia from the everyday conversations of elderly people.
[1221] The device begins a casual conversation by asking, "Tell me how your day was today."
[1222] Users talk about everyday events.
[1223] The terminal uses a voice recognition system to convert the conversation into text data and send it to the server.
[1224] The server analyzes the text data, assessing language patterns and emotional states, and also analyzes emotional fluctuations using an emotion engine.
[1225] The server generates an alert such as, "Recently, we have noticed a decline in vocabulary in conversations."
[1226] The device will notify the senior of the alert and, if necessary, their caregiver or family member.
[1227] Example prompt sentence:
[1228] Here are some prompts for early dementia detection:
[1229] Please explain the steps for analyzing and generating alerts after an elderly person talks about their daily life when asked, "Tell me how your day was today."
[1230] In this way, the remote care system of the present invention supports the lives of elderly people in a variety of ways, and by combining it with an emotion engine, it is possible to provide more accurate advice and appropriate care.
[1231] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1232] Meal management function processing steps
[1233] Step 1:
[1234] Voice input and data collection (terminal)
[1235] The device uses a microphone to ask the elderly person, "What did you have for breakfast today?" The elderly person responds, "I had bread and milk." This voice input is collected.
[1236] Input: Voice input for the elderly
[1237] Output: Audio data collected by microphone
[1238] Step 2:
[1239] Speech recognition and data conversion (terminal)
[1240] The device uses a voice recognition system to convert voice data into text data, and also uses a camera to record the elderly person's mealtimes.
[1241] Input: Audio data, video data of meal
[1242] Output: Text data, image data
[1243] Step 3:
[1244] Data transmission (terminal)
[1245] The terminal encrypts the converted text data and the captured image data and transmits them to the server in real time.
[1246] Input: Text data, image data
[1247] Output: Data sent to the server
[1248] Step 4:
[1249] Data analysis (server)
[1250] The server uses the generative AI model to analyze the received text and image data and evaluate the calories and nutritional balance of the meal contents.
[1251] Input: Text data, image data
[1252] Output: Calorie and nutritional balance assessment results
[1253] Step 5:
[1254] Sentiment analysis (server)
[1255] The server's emotion engine recognizes the user's emotional state from the voice data.
[1256] Input: Text data
[1257] Output: User's emotional state
[1258] Step 6:
[1259] Advice Generation (Server)
[1260] The server generates advice based on the meal evaluation and emotional state, such as "We recommend adding more fruits and vegetables to your next meal."
[1261] Input: Calorie and nutritional balance assessment results, user's emotional state
[1262] Output: The generated advice
[1263] Step 7:
[1264] Notifications (device)
[1265] The device notifies the elderly person of the advice received from the server via voice.
[1266] Input: Advice from the server
[1267] Output: Voice notification to the elderly
[1268] Medication management function processing steps
[1269] Step 1:
[1270] Medication time detection (server)
[1271] The server detects the preset time for taking the medication.
[1272] Input: Pre-set medication time
[1273] Output: Start generating medication reminders
[1274] Step 2:
[1275] Reminder generation (server)
[1276] The server generates a reminder: "It's time to take your medicine. Are you ready?"
[1277] Input: Medication time detection
[1278] Output: The generated reminder
[1279] Step 3:
[1280] Reminder notification (device)
[1281] The device will notify the elderly of reminders via voice.
[1282] Input: Reminder from server
[1283] Output: Voice notification to the elderly
[1284] Step 4:
[1285] User response (terminal)
[1286] The user responds, "Yes, I'll drink it now." The device uses a voice recognition system to convert the response into text data.
[1287] Input: User's voice response
[1288] Output: Text data
[1289] Step 5:
[1290] Medication confirmation (terminal)
[1291] The device uses a camera to record the patient taking the medicine.
[1292] Input: User's medication actions
[1293] Output: Video data
[1294] Step 6:
[1295] Data transmission (terminal)
[1296] The converted text data and video data are sent to the server.
[1297] Input: Text data, video data
[1298] Output: Data sent to the server
[1299] Step 7:
[1300] Data analysis (server)
[1301] The server analyzes the video data to confirm whether the user has taken the medicine, and also analyzes the user's emotional state using an emotion engine.
[1302] Input: Text data, video data
[1303] Output: Medication confirmation result, user's emotional state
[1304] Step 8:
[1305] Confirmation message generation (server)
[1306] The server generates a confirmation message saying "Your medication was taken correctly" and sets the next reminder.
[1307] Input: Medication confirmation result, user's emotional state
[1308] Output: Generated confirmation message, next reminder
[1309] Step 9:
[1310] Confirmation notification (terminal)
[1311] The device will then provide a confirmation message to the elderly person via voice.
[1312] Input: Confirmation message from the server
[1313] Output: Voice notification to the elderly
[1314] Dementia early detection function processing steps
[1315] Step 1:
[1316] Starting a daily conversation (terminal)
[1317] The device speaks to the elderly person, asking, "Tell us how your day was today."
[1318] Input: Pre-configured question prompt
[1319] Output: Start a conversation with the elderly person
[1320] Step 2:
[1321] User response (terminal)
[1322] Users talk about everyday events, and the device converts the conversation into text data using a voice recognition system.
[1323] Input: User's daily conversation
[1324] Output: Text data
[1325] Step 3:
[1326] Data transmission (terminal)
[1327] The converted text data is sent to the server.
[1328] Input: Text data
[1329] Output: Data sent to the server
[1330] Step 4:
[1331] Data analysis (server)
[1332] The server analyzes the text data, assessing language patterns and emotional states, and also analyzes emotional fluctuations using an emotion engine.
[1333] Input: Text data
[1334] Output: Language pattern evaluation results, emotional state evaluation results
[1335] Step 5:
[1336] Alert Generation (Server)
[1337] The server generates an alert such as, "Recently, we have noticed a decline in vocabulary in conversations."
[1338] Input: Language pattern evaluation results, emotional state evaluation results
[1339] Output: The generated alert
[1340] Step 6:
[1341] Alert notification (terminal)
[1342] The device will notify the senior of the alert, and if necessary, a caregiver or family member.
[1343] Input: Alert from the server
[1344] Output: Notification to seniors and caregivers
[1345] (Application example 2)
[1346] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1347] It is difficult to monitor the health and emotional states of elderly people in real time and provide appropriate advice and reminders when they go about their daily lives, especially when they use self-driving vehicles. Furthermore, conventional remote care systems do not provide support tailored to the elderly's situation in self-driving vehicles, which may reduce the elderly's sense of safety and security. Therefore, there is a need for a remote care system that can monitor the health and emotional states of elderly people in real time and provide appropriate support when they use self-driving vehicles.
[1348] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1349] In this invention, the server includes a voice recognition means for acquiring voice data of the elderly person, an image recognition means for acquiring image data of the elderly person, an analysis means using a generative AI model to analyze the voice data and image data, a notification means for notifying the elderly person of necessary advice and reminders based on the analysis results, a means for monitoring the elderly person's health and emotional state while in the autonomously driven vehicle, a means for transmitting the data collected by the above means to the server and analyzing it, a means for providing the elderly person with advice regarding their living conditions in the autonomously driven vehicle based on the analysis results, and a means for notifying the elderly person of the advice visually and audibly. This makes it possible to support elderly people in living safely and healthily when using autonomously driven vehicles.
[1350] The "voice recognition means" is a means for acquiring and analyzing voice data of the elderly person.
[1351] The "image recognition means" is a means for acquiring image data of an elderly person and analyzing it.
[1352] A "generative AI model" is an advanced artificial intelligence model that analyzes audio and image data and is used to assess the health and emotional state of an elderly person.
[1353] The "analysis means" is a means for assessing the health and emotional state of an elderly person using audio and image data.
[1354] The "notification means" is a means for visually and audibly notifying the elderly of necessary advice and reminders based on the analysis results.
[1355] An "automated vehicle" is a vehicle that is driven automatically and can be used by seniors.
[1356] "Health status" refers to the physical condition and physical function of an elderly person.
[1357] "Emotional state" refers to the psychological state and emotional movements of elderly people.
[1358] The "means for transmitting data to a server" is a means for transmitting voice data and image data collected from elderly people to a server.
[1359] "Visual and audio notification means" refers to means for conveying advice to the elderly person through a visual display and audio.
[1360] The present invention relates to a remote care system for elderly people to live safe and healthy lives when using autonomous vehicles. An embodiment of the present invention will be specifically described.
[1361] System configuration
[1362] Terminal configuration
[1363] 1. Voice recognition means:
[1364] The device is equipped with a microphone that is used to capture voice data from the elderly, which is then analyzed in real time.
[1365] 2. Image Recognition Methods:
[1366] The device is equipped with a camera that is used to capture image data of the elderly person and send it to a server. For example, the device captures the elderly person's facial expressions and body movements.
[1367] Server-side configuration
[1368] 1. Analysis method:
[1369] The server is equipped with a generative AI model that analyzes the voice and image data sent from the device, and this analysis evaluates the health and emotional state of the elderly person.
[1370] 2. Means of notification:
[1371] Based on the analysis results, the server generates necessary advice and reminders for the elderly, which are then sent to the device, which then notifies the elderly by voice or display.
[1372] Hardware and software used
[1373] Hardware:
[1374] Microphone: Used to collect voice data from the elderly.
[1375] Camera: Used to capture image data of the elderly.
[1376] Self-driving vehicles: Used to monitor the health and emotional state of elderly people in real time while they are in the vehicle.
[1377] software:
[1378] speech_recognition library: Used to achieve speech recognition.
[1379] OpenCV library: Used to realize image processing.
[1380] TensorFlow: Used to realize a generative AI model for elderly emotion engine analysis.
[1381] gTTS: Used to generate audio notifications.
[1382] requests: Used to realize data transmission to the server.
[1383] System operation example
[1384] 1. While the elderly person is riding in a self-driving vehicle:
[1385] Microphones and cameras capture the elderly person's voice and image data in real time, which is automatically sent to a server.
[1386] The server uses a generative AI model to analyze audio and image data to assess the health and emotional state of the elderly.
[1387] Based on the assessment results, necessary advice and reminders are generated and sent to the elderly via the device. For example, if a person does not eat enough breakfast, they will receive a voice message saying, "We recommend adding more fruits and vegetables to your next meal."
[1388] Prompt Sentence Examples
[1389] After the elderly person answers questions about health care in the self-driving car, please send the image and voice data to a server and generate a Python program that will provide appropriate advice and reminders.
[1390] As described above, the remote care system of the present invention provides multifaceted support to enable elderly people to use self-driving vehicles with peace of mind.
[1391] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1392] Step 1: Acquire audio data
[1393] The device uses a microphone to capture the elderly person's voice data. The input is the elderly person's voice, which is collected in real time. The device stores the collected voice data either directly or temporarily.
[1394] Step 2: Acquiring image data
[1395] The device uses a camera to acquire image data (facial expressions and body movements) of the elderly. The input is a real-time image of the elderly acquired through the camera, which is captured at high resolution. The device stores the captured image data either directly or temporarily.
[1396] Step 3: Send data to the server
[1397] The terminal sends the collected voice and image data to the server. The input is the data acquired in steps 1 and 2 above, and this is sent to the server via the network. The output is the data sent to the server itself.
[1398] Step 4: Analyzing the audio data
[1399] The server analyzes the received voice data. The input is the voice data sent from the device, which is analyzed using a generative AI model. The content of the elderly person's speech and their emotional state are extracted from the voice data and output as text data.
[1400] Step 5: Analyzing the image data
[1401] The server analyzes the received image data. The input is image data sent from the device, which is analyzed using a generative AI model. The image data is used to evaluate the elderly person's facial expressions, body movements, health condition, etc., and is output as numerical data or evaluation results.
[1402] Step 6: Integrating the analysis results
[1403] The server integrates the results of the voice data analysis and the image data analysis to evaluate the elderly person's comprehensive health and emotional state. The input is the analysis results of steps 4 and 5, which are integrated to comprehensively grasp the elderly person's current condition. The output is the integrated evaluation result.
[1404] Step 7: Generate advice and reminders
[1405] The server generates necessary advice and reminders for the elderly based on the integrated evaluation results. The input is the integrated evaluation results, and appropriate messages are generated based on them. The output is text data of the generated advice and reminders.
[1406] Step 8: Implementing Notification
[1407] The device notifies the elderly of advice and reminders received from the server. The input is text data sent from the server, which is notified visually and audibly. Specifically, messages are displayed on the display and read aloud using gTTS. The output is the advice and reminders notified to the elderly.
[1408] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1409] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1410] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1411] [Third embodiment]
[1412] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1413] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[1414] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1415] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1416] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1417] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1418] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1419] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1420] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1421] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1422] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1423] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1424] This remote care system consists of elderly end users, and a server and terminal (robot) located in a remote location. A specific embodiment of the system is described below. The program code itself is not provided, but the processing flow and each step are described in detail.
[1425] System configuration
[1426] 1. Terminal configuration
[1427] Speech recognition means: The device is equipped with a microphone to collect conversations with the elderly in real time, and speech recognition technology is used to convert the captured speech into text data.
[1428] Image recognition method: The device is equipped with a camera that captures images of the elderly person's daily life, especially when they eat or take medicine. The image data is sent to a server in real time.
[1429] 2. Server-side configuration
[1430] Analysis method: The server uses a generative AI model to analyze the voice and image data sent from the device. This analysis allows for a real-time assessment of the elderly person's health and living conditions.
[1431] Notification method: Based on the analysis results, appropriate advice and reminders are generated and sent to the device, which then communicates this to the elderly via voice.
[1432] Meal management function for the elderly
[1433] Overview: The dietary management function monitors whether the elderly are eating a balanced diet.
[1434] Terminal: Ask the senior, "What did you have for breakfast today?"
[1435] User: "I had bread and milk."
[1436] Device: Converts voice into text data and simultaneously takes pictures of the meal using a camera.
[1437] Server: Analyzes the received text and image data and evaluates the calorie and nutritional balance of the meal.
[1438] Server: Generates advice as needed, such as "Today's breakfast appears to be low in calories. We recommend eating some fruits and vegetables at your next meal."
[1439] Device: Provides advice to the elderly via voice.
[1440] Medication management function
[1441] Overview: Medication management helps seniors take their medications appropriately.
[1442] Server: Detects when it's time to take your medicine and generates a reminder saying, "It's time to take your medicine. Are you ready?"
[1443] Device: Provides voice reminders to seniors.
[1444] User: "Yes, I'll take my medicine."
[1445] Device: Activate the camera and record the scene of taking the medicine.
[1446] Server: Analyzes image data and verifies whether the medication has been taken.
[1447] Server: If the medication is confirmed, generate a confirmation message saying "Your medication was taken correctly" and set the next reminder.
[1448] Terminal: Notify the elderly person of a confirmation message.
[1449] Early detection of dementia
[1450] Overview: The early dementia detection function detects signs of dementia from the everyday conversations of elderly people.
[1451] Device: Initiate everyday conversations such as "Tell me how your day was."
[1452] User: Talk about everyday events.
[1453] Terminal: Converts the voice into text data and sends it to the server.
[1454] Server: Analyzes conversation data and evaluates language patterns and emotional states, for example, detecting vocabulary declines and changes in emotional tone.
[1455] Server: If signs of dementia are detected, it generates an alert such as, "Recently, there has been a decline in vocabulary in conversations."
[1456] Device: Notifies the senior of the alert and, if necessary, notifies caregivers and family members.
[1457] As described above, this remote care system provides comprehensive support for the daily lives of elderly people and makes it easier for caregivers and family members in remote locations to understand the condition of the elderly, thereby improving the quality of life of the elderly and reducing the burden on caregivers.
[1458] The processing flow will be explained below.
[1459] Meal management function for the elderly
[1460] Step 1:
[1461] The device asks the senior, "What did you have for breakfast today?"
[1462] Step 2:
[1463] The user responds, "I had bread and milk."
[1464] Step 3:
[1465] The terminal uses a voice recognition system to convert the user's speech into text data.
[1466] Step 4:
[1467] The device activates the camera and takes pictures of the elderly person eating.
[1468] Step 5:
[1469] The terminal transmits the acquired text data and image data to the server.
[1470] Step 6:
[1471] The server analyzes the received text data and extracts the meal details.
[1472] Step 7:
[1473] The server analyzes the received image data and calculates the calories and nutritional balance of the meal contents.
[1474] Step 8:
[1475] Based on the analysis results, the server generates advice such as, "Today's breakfast appears to be low in calories. We recommend that you eat fruits and vegetables at your next meal."
[1476] Step 9:
[1477] The server sends the generated advice to the terminal.
[1478] Step 10:
[1479] The device will provide advice to the elderly via voice.
[1480] Medication management function
[1481] Step 1:
[1482] The server recognizes pre-set medication times and generates reminders.
[1483] Step 2:
[1484] The device will then send a voice reminder to the elderly person saying, "It's time to take your medicine. Are you ready?"
[1485] Step 3:
[1486] The user responds, "Yes, I'll take my medicine."
[1487] Step 4:
[1488] The terminal uses a voice recognition system to convert the user's response into text data.
[1489] Step 5:
[1490] The device activates the camera and takes a picture of the elderly person taking their medicine.
[1491] Step 6:
[1492] The video data acquired by the terminal is transmitted to the server.
[1493] Step 7:
[1494] The server analyzes the video data and confirms that the medication has been taken.
[1495] Step 8:
[1496] The server generates a confirmation message saying "Your medication was taken correctly" and sets the next reminder.
[1497] Step 9:
[1498] The server sends a confirmation message to the terminal.
[1499] Step 10:
[1500] The device will then send a confirmation message to the elderly person.
[1501] Early detection of dementia
[1502] Step 1:
[1503] The device begins a casual conversation with the elderly person, asking, "Tell me how your day was today."
[1504] Step 2:
[1505] Users talk about everyday events.
[1506] Step 3:
[1507] The terminal uses a voice recognition system to convert the user's speech into text data.
[1508] Step 4:
[1509] The terminal transmits the acquired text data to the server.
[1510] Step 5:
[1511] The server analyzes the text data and evaluates the elderly person's language patterns and emotional state.
[1512] Step 6:
[1513] The server detects signs of dementia based on the analysis results.
[1514] Step 7:
[1515] The server generates an alert such as, "Recently, we have noticed a decline in vocabulary in conversations."
[1516] Step 8:
[1517] The server generates alerts and sends them to the device, which also notifies family members and caregivers.
[1518] Step 9:
[1519] The device will notify the elderly person of the alert.
[1520] These processing steps are designed to ensure that the lives of the elderly are managed effectively and to reduce the burden on caregivers.
[1521] Example 1
[1522] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1523] It is difficult for elderly people to maintain appropriate lifestyle habits and monitor their health. In particular, it is important for elderly people to manage the nutritional balance of their diet, their medication status, and even to detect dementia early, but it is difficult for them to do these things alone. It is also not easy for caregivers and family members in remote locations to understand the condition of elderly people. Conventional systems lack the functionality to comprehensively support the lives of elderly people.
[1524] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1525] In this invention, the server includes a voice recognition means for acquiring voice data of the elderly person, an image recognition means for acquiring image data of the elderly person, an analysis means using a generative AI model to analyze the voice data and image data, a notification means for notifying the elderly person of necessary advice and reminders based on the analysis results, a means for detecting signs of dementia from the elderly person's everyday conversation and generating an alert, a means for reminding the elderly person to take their medicine and confirming that they have taken it, and a means for monitoring the elderly person's diet to evaluate nutritional balance and provide advice. This makes it possible to comprehensively support the elderly's living conditions and enable caregivers and family members in remote locations to easily understand the elderly person's condition.
[1526] The "voice recognition means" is a device or technology for capturing the voice of the elderly person in real time and converting the voice data into text data.
[1527] "Image recognition means" refers to a device or technology for photographing the elderly person's daily life and acquiring image data.
[1528] A "generative AI model" is an artificial intelligence model that analyzes voice and image data and generates appropriate advice and reminders based on the results.
[1529] The "analysis means" is a means of evaluating the health status and living conditions of elderly people using a generative AI model, using the acquired voice and image data.
[1530] The "notification means" is a means of notifying the elderly of advice and reminders generated based on the analysis results via voice or other means.
[1531] The "means for detecting signs of dementia" is a method for analyzing data on everyday conversations of elderly people and detecting early signs of dementia from factors such as a decrease in vocabulary and changes in emotional tone.
[1532] "Reminder measures" are measures that notify elderly people when it is time to take their medicine and help them ensure that they take their medicine.
[1533] A "dietary monitoring method" is a method for monitoring the dietary content of elderly people, evaluating their nutritional balance and calories, and providing necessary advice.
[1534] This invention is a remote care system for monitoring and supporting the living conditions of elderly people. The system consists of an elderly person in a remote location, a server, and a terminal (robot). This system is configured and operates as follows to monitor the elderly person's daily behavior and health status and provide necessary advice and reminders.
[1535] System configuration
[1536] Terminal configuration
[1537] Speech recognition method: The device is equipped with a microphone to collect conversations with the elderly in real time. The captured speech is converted into text data using speech recognition technology (e.g., Google Cloud Speech-to-Text).
[1538] Image recognition method: The device is equipped with a camera that captures images of the elderly person's daily life, especially when they eat or take medicine. The image data is sent to a server in real time.
[1539] Server-side configuration
[1540] Analysis method: The server uses a generative AI model (e.g., OpenAI GPT-4) to analyze the voice and image data sent from the device. This analysis allows for a real-time assessment of the elderly person's health and living conditions.
[1541] Notification method: Based on the analysis results, appropriate advice and reminders are generated and sent to the device, which then communicates this to the elderly via voice.
[1542] Meal management function for the elderly
[1543] Overview: The dietary management function monitors whether the elderly are eating a balanced diet.
[1544] Terminal: Ask the senior, "What did you have for breakfast today?"
[1545] User: "I had bread and milk."
[1546] Device: Converts voice into text data and simultaneously takes pictures of the meal using a camera.
[1547] Server: Analyzes the received text and image data and evaluates the calorie and nutritional balance of the meal.
[1548] Server: Generates advice as needed, such as "Today's breakfast appears to be low in calories. We recommend eating some fruits and vegetables at your next meal."
[1549] Device: Provides advice to the elderly via voice.
[1550] Specific prompt examples:
[1551] "If an elderly person eats bread and milk for breakfast, generate a sentence that evaluates the nutritional balance and advises them to add more fruits and vegetables to their next meal if the calories are low."
[1552] Medication management function
[1553] Overview: Medication management helps seniors take their medications appropriately.
[1554] Server: Detects when it's time to take your medicine and generates a reminder saying, "It's time to take your medicine. Are you ready?"
[1555] Device: Provides voice reminders to seniors.
[1556] User: "Yes, I'll take my medicine."
[1557] Device: Activate the camera and record the scene of taking the medicine.
[1558] Server: Analyzes image data and verifies whether the medication has been taken.
[1559] Server: If the medication is confirmed, generate a confirmation message saying "Your medication was taken correctly" and set the next reminder.
[1560] Terminal: Notify the elderly person of a confirmation message.
[1561] Specific prompt examples:
[1562] "Generate a reminder when it's time for the senior to take their medication, and then create a sentence to confirm if they actually took their medication."
[1563] Early detection of dementia
[1564] Overview: The early dementia detection function detects signs of dementia from the everyday conversations of elderly people.
[1565] Device: Initiate everyday conversations such as "Tell me how your day was."
[1566] User: Talk about everyday events.
[1567] Terminal: Converts the voice into text data and sends it to the server.
[1568] Server: Analyzes conversation data and evaluates language patterns and emotional states, for example, detecting vocabulary declines and changes in emotional tone.
[1569] Server: If signs of dementia are detected, it generates an alert such as, "Recently, there has been a decline in vocabulary in conversations."
[1570] Device: Notifies the senior of the alert and, if necessary, notifies caregivers and family members.
[1571] Specific prompt examples:
[1572] "To detect signs of dementia in everyday conversations of elderly people, create sentences that analyze vocabulary decline and changes in emotional tone."
[1573] As described above, this remote care system provides comprehensive support for the daily lives of elderly people and makes it easier for caregivers and family members in remote locations to understand the condition of the elderly, thereby improving the quality of life of the elderly and reducing the burden on caregivers.
[1574] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1575] Meal management function for the elderly
[1576] Step 1:
[1577] The device asks the senior, "What did you have for breakfast today?"
[1578] Input: Breakfast Question
[1579] Output: Waiting for response from the elderly person
[1580] Step 2:
[1581] The user responds, "I had bread and milk."
[1582] Input: Elderly person's voice response
[1583] Output: Audio data
[1584] Step 3:
[1585] The device converts the voice into text data and simultaneously captures the meal with a camera.
[1586] Input: Audio data, video of elderly people eating
[1587] Output: Text data, image data
[1588] What it does: The microphone captures your voice and converts it into text using a speech recognition engine. The camera captures footage of you eating.
[1589] Step 4:
[1590] The server receives the text and image data and analyzes it using a generative AI model.
[1591] Input: Text data, image data
[1592] Output: Calorie and nutritional balance evaluation results of the meal
[1593] How it works: The server extracts meal items from text data and analyzes the meal contents using an image recognition algorithm. The generative AI model calculates nutritional balance and calories.
[1594] Step 5:
[1595] The server generates advice based on the analysis results.
[1596] Input: Evaluation results of calorie and nutritional balance of meals
[1597] Output: Advice message
[1598] Specific behavior: The server generates advice such as "Today's breakfast is low in calories. We recommend that you eat fruits and vegetables at your next meal."
[1599] Step 6:
[1600] The device will provide advice to the elderly via voice.
[1601] Input: Advice message
[1602] Output: Voice notification to the elderly
[1603] Specific operation: Advice messages are synthesized and conveyed to the elderly.
[1604] Medication management function
[1605] Step 1:
[1606] The server detects when it's time to take your medicine and generates a reminder saying, "It's time to take your medicine. Are you ready?"
[1607] Input: Medication schedule
[1608] Output: Reminder message
[1609] What it does: The server periodically checks the schedule and generates a reminder message when it's time to take your medication.
[1610] Step 2:
[1611] The device will notify the elderly of the reminder via voice.
[1612] Input: Reminder message
[1613] Output: Voice notification to the elderly
[1614] Specific actions: Reminder messages are delivered to elderly people via voice synthesis.
[1615] Step 3:
[1616] The user responds, "Yes, I'll take my medicine."
[1617] Input: Elderly person's voice response
[1618] Output: Response audio data
[1619] Step 4:
[1620] The device activates the camera and captures the scene of the elderly person taking their medicine.
[1621] Input: Video of elderly person taking medication
[1622] Output: Image data of the medication scene
[1623] Specific operation: The camera captures the elderly person taking their medicine and saves the image data on the device.
[1624] Step 5:
[1625] The server analyzes the image data and verifies whether the medication has been taken.
[1626] Input: Image data of a medication scene
[1627] Output: Medication confirmation result
[1628] What it does: The server uses image recognition technology to verify that the medication was taken correctly.
[1629] Step 6:
[1630] The server generates a confirmation message and sets the next reminder.
[1631] Input: Medication confirmation result
[1632] Output: Confirmation message, next reminder setting
[1633] Specific behavior: If the medication is confirmed, the server generates a confirmation message saying "Your medication was taken correctly" and sets the next reminder.
[1634] Step 7:
[1635] The device will then send a confirmation message to the elderly person.
[1636] Input: Confirmation message
[1637] Output: Voice notification to the elderly
[1638] Specific operation: A confirmation message is synthesized and conveyed to the elderly person.
[1639] Early detection of dementia
[1640] Step 1:
[1641] The device will begin a casual conversation such as, "Tell me how your day was today."
[1642] Input: Daily conversation question
[1643] Output: Waiting for response from the elderly person
[1644] Step 2:
[1645] Users talk about everyday events.
[1646] Input: Elderly person's voice
[1647] Output: Audio data
[1648] Step 3:
[1649] The device converts the voice into text data and sends it to the server.
[1650] Input: Audio data
[1651] Output: Text data
[1652] What it does: The microphone captures your voice and the speech recognition engine converts it into text.
[1653] Step 4:
[1654] The server analyzes the conversation data and evaluates language patterns and emotional states.
[1655] Input: Text data
[1656] Output: Evaluation result
[1657] Specific operation: The server uses the generative AI model to detect vocabulary loss, changes in emotional tone, etc.
[1658] Step 5:
[1659] If the server detects signs of dementia, it generates an alert message.
[1660] Input: Evaluation result
[1661] Output: Alert message
[1662] Specific behavior: If signs of dementia are detected, an alert will be generated, such as "Recently, there has been a decline in vocabulary in conversations."
[1663] Step 6:
[1664] The device notifies the elderly person of the alert and, if necessary, notifies caregivers and family members.
[1665] Input: Alert message
[1666] Output: Audio notification to the senior and caregiver / family
[1667] Specific operation: An alert message is synthesized and conveyed to the elderly. At the same time, an alert is sent to caregivers and family members.
[1668] The above are the specific processing steps and their detailed operations. These steps improve the quality of life for the elderly and enable appropriate care to be provided remotely.
[1669] (Application example 1)
[1670] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1671] Remote care systems that support the lives of the elderly are limited in their management functions to the health status of the elderly, and do not address comprehensive home safety. There is a need for a system that comprehensively monitors security aspects such as managing entry and exit within the home and detecting fires, gas leaks, and suspicious individuals. The objective of this invention is to provide a comprehensive support system that not only manages the health of the elderly, but also monitors the safety status of the home.
[1672] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1673] In this invention, the server includes an analysis means, a notification means, and a safety monitoring means, which makes it possible not only to grasp and support the living conditions of the elderly, but also to comprehensively monitor the safety of the home.
[1674] A "remote care system" is a system that monitors and supports the living conditions of elderly people from a remote location.
[1675] The "voice recognition means" is a device that has the function of acquiring voice data from the elderly person and converting it into text data in real time.
[1676] The "image recognition means" is a device that has the function of acquiring image data of an elderly person and analyzing the content of that image data.
[1677] A "generative AI model" is an artificial intelligence model that analyzes acquired voice data and image data and generates the necessary information.
[1678] The "analysis means" is a function that analyzes voice data and image data to evaluate the living conditions of the elderly person.
[1679] "Notification means" is a function that notifies elderly people of advice and reminders based on the analysis results.
[1680] The "safety monitoring means" is a device that has the function of monitoring the safety status of the home, detecting abnormalities, and notifying them.
[1681] "Nutritional balance" is an indicator of whether nutrients are evenly distributed in a particular diet.
[1682] A "calorie" is a unit that indicates the amount of energy ingested through food.
[1683] A "reminder" is a notification that encourages seniors to take a specific action.
[1684] This remote care system is a system for monitoring and supporting the living conditions and home safety conditions of elderly people. A specific embodiment of the system will be described below.
[1685] System configuration
[1686] The server and user terminals (smartphones) work together to monitor the lives of the elderly and the safety of their homes using various sensor devices. This system includes the following hardware and software:
[1687] Hardware
[1688] Smartphone: microphone, camera
[1689] Sensors: Door and window sensors, smoke detectors, gas leak detectors
[1690] software
[1691] Speech recognition API: Used to convert voice data into text data. Typical examples include Google Speech-to-Text and Apple Siri.
[1692] Image recognition API: Used to analyze image data. Representative examples include Google Cloud Vision and Amazon Rekognition.
[1693] Server: Platforms that run generative AI models for data analysis include Google Cloud and AWS (Amazon Web Services).
[1694] System Operation
[1695] The server uses voice recognition, image recognition, and safety monitoring means to notify the user based on the analysis results. Each of these means is described in detail below.
[1696] Voice recognition means
[1697] The microphone on the user's device collects the elderly's voice and converts it into text data in real time using a speech recognition API. For example, if an elderly person says "help me," it is converted into text data and sent to the server.
[1698] Image Recognition Method
[1699] The camera on the user's device captures images of the elderly's living conditions and home situation, and the images are analyzed using an image recognition API. Examples include detecting fires, gas leaks, and suspicious people. It also includes capturing images of elderly people taking medicine, and analyzing the image data on a server.
[1700] Safety monitoring means
[1701] Sensor devices such as door and window sensors, smoke detectors, and gas leak detectors monitor the safety status of the home and notify the server if an abnormality is detected. The server analyzes the abnormality and issues a warning to the user in real time.
[1702] Server analysis method
[1703] The server analyzes the voice and image data using a generative AI model, and based on the analysis results, generates appropriate advice and reminders and sends them to the user's device.
[1704] Notification means
[1705] Based on the analysis results, the system will provide voice reminders and warnings to the elderly, such as specific instructions such as "There is a fire. Please close the doors of each room."
[1706] Specific examples
[1707] In case of fire detection
[1708] 1. The microphone detects the sound of a fire alarm.
[1709] 2. The voice recognition API converts the message "Fire has broken out" into text data.
[1710] 3. Analyze fire data on the server.
[1711] 4. An "emergency fire alert" notification will be sent to your smartphone.
[1712] 5. An audio guide will be given to residents to "close the doors of each room."
[1713] Prompt Sentence Examples
[1714] "Design a program that uses a voice recognition system to detect fire alarm sounds in real time and notify a smartphone of that information. If an abnormal sound is detected, provide voice guidance on first aid measures for the fire."
[1715] This will make it possible to realize a system that comprehensively supports the lives and home safety of the elderly.
[1716] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1717] Step 1:
[1718] The microphone on the user's device captures the elderly person's voice data, which is then sent to a voice recognition API and converted into text data in real time.
[1719] Input: Elderly voice data
[1720] Output: Text data
[1721] Specific operation: When an elderly person says "help me," their voice is picked up by a microphone and converted into text data using a voice recognition API.
[1722] Step 2:
[1723] The camera on the user's device captures image data of the elderly person, which is then sent to an image recognition API for real-time analysis.
[1724] Input: Image data of elderly people
[1725] Output: Analysis data
[1726] Specific operation: The camera captures scenes of elderly people eating or taking medicine, and the image data is analyzed using an image recognition API.
[1727] Step 3:
[1728] The server receives the voice and image data sent from the user's device and analyzes the data using a generative AI model.
[1729] Input: Text data, analysis data
[1730] Output: Analysis results
[1731] Specific operation: The voice data is converted into text and the analysis results of the image data are sent to a server, and the generated AI model analyzes the elderly person's living conditions and health condition.
[1732] Step 4:
[1733] The server generates appropriate advice and reminders based on the analysis results.
[1734] Input: Analysis results
[1735] Output: Advice, reminder
[1736] Specific behavior: For example, if an elderly person's breakfast is analyzed to be insufficient in calories, advice such as "We recommend that you eat fruits and vegetables at your next meal" will be generated.
[1737] Step 5:
[1738] The server transmits the generated advice and reminders to the user terminal using a notification means.
[1739] Input: Advice, Reminder
[1740] Output: Notification data
[1741] Specific operation: The generated advice and reminders are sent to the user's device as push notifications, and the elderly person receives notifications such as "It's time to take their medicine."
[1742] Step 6:
[1743] Sensors on the user's device (doors, windows, smoke, gas leaks) monitor the safety status of the home, and if an abnormality is detected, the data is sent to the server.
[1744] Input: Sensor data
[1745] Output: Anomaly detection data
[1746] Specific operation: If a door or window is opened or closed suspiciously, or if a fire or gas leak is detected, the information is sent to the server.
[1747] Step 7:
[1748] The server analyzes the anomaly detection data, generates necessary warnings, and sends them to the user terminal.
[1749] Input: Anomaly detection data
[1750] Output: Warning notification data
[1751] Specific operation: For example, a warning message such as "A fire has been detected. Please close the doors of each room" is generated and sent to the user terminal.
[1752] Step 8:
[1753] The user terminal notifies the elderly person of the received warning notification data by voice in real time.
[1754] Input: Alert notification data
[1755] Output: Audio notification
[1756] Specific actions: The elderly person will be notified by voice, "A fire has broken out. Please close the doors of each room."
[1757] This will enable comprehensive monitoring of the elderly's living conditions and home safety, and provide appropriate support at the appropriate time.
[1758] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1759] This invention is a remote care system that monitors the living conditions of elderly people and provides appropriate support, and is equipped with an emotion engine that analyzes voice data and image data of the elderly person and recognizes the emotional state of the user. Specific embodiments of the invention will be described below.
[1760] System configuration
[1761] 1. Terminal configuration
[1762] Voice recognition means: The device is equipped with a microphone to collect conversations with the elderly in real time. The collected voice data is converted into text data by a voice recognition system.
[1763] Image recognition means: The device is equipped with a camera that takes pictures of, for example, an elderly person eating or taking medicine. The captured image data is sent to a server in real time.
[1764] 2. Server-side configuration
[1765] Analysis method: The server is equipped with a generative AI model with advanced analytical capabilities that analyzes the voice and image data sent from the device. This analysis allows for a real-time assessment of the elderly person's health and living conditions.
[1766] Emotion engine: The server is also equipped with an emotion engine that recognizes the emotional state of the elderly by analyzing voice data.
[1767] 3. Means of notification
[1768] Advice and reminders: Based on the analysis results and emotional state, advice and reminders are generated for the elderly. This information is sent from the server to the device, which then notifies the elderly by voice.
[1769] Meal management function for the elderly
[1770] Overview: The dietary management function monitors whether the elderly are eating a balanced diet and provides advice as needed.
[1771] Terminal: Ask the senior, "What did you have for breakfast today?"
[1772] User: "I had bread and milk."
[1773] Device: A voice recognition system converts the user's speech into text data. A camera captures the meal.
[1774] Server: Analyzes text and image data to evaluate the calorie and nutritional balance of meal content. Also analyzes the user's emotional state using an emotion engine.
[1775] Server: Generate advice such as, "Your breakfast today seems low in calories. We recommend adding more fruits and vegetables to your next meal."
[1776] Device: Provides advice to the elderly via voice.
[1777] Medication management function
[1778] Abstract: Medication management is a function that helps elderly people take their medications appropriately.
[1779] Server: Detects medication times and generates reminders.
[1780] Device: Reminds seniors, "It's time to take your medicine. Are you ready?"
[1781] User: "Yes, I'll drink it now."
[1782] Device: A voice recognition system converts responses into text data, and a camera captures the process of taking the medicine.
[1783] Server: Analyzes video data and confirms whether the user has taken the medication. An emotion engine also analyzes the user's emotional state.
[1784] Server: Generates a confirmation message saying "Your medication was taken correctly" and sets the next reminder.
[1785] Terminal: Notify the elderly person of a confirmation message.
[1786] Early detection of dementia
[1787] Overview: The early dementia detection function detects signs of dementia from the everyday conversations of elderly people.
[1788] Device: Begins a conversation with "Tell me how your day was."
[1789] User: Talk about everyday events.
[1790] Terminal: The speech recognition system converts the conversation into text data and sends it to the server.
[1791] Server: Analyzes text data and evaluates language patterns and emotional states. An emotion engine also analyzes emotional fluctuations.
[1792] Server: Generate an alert such as, "Recently, we have noticed a decline in vocabulary in conversations."
[1793] Device: Notifies the senior of the alert and, if necessary, notifies caregivers and family members.
[1794] In this way, the remote care system of the present invention supports the lives of elderly people in a variety of ways, and by combining it with an emotion engine, it is possible to provide more accurate advice and appropriate care.
[1795] The processing flow will be explained below.
[1796] Meal management function for the elderly
[1797] Step 1:
[1798] The device asks the senior, "What did you have for breakfast today?"
[1799] Step 2:
[1800] The user responds, "I had bread and milk."
[1801] Step 3:
[1802] The terminal uses a voice recognition system to convert the user's speech into text data.
[1803] Step 4:
[1804] The device activates the camera and takes pictures of the elderly person eating.
[1805] Step 5:
[1806] The terminal transmits the acquired text data and image data to the server.
[1807] Step 6:
[1808] The server analyzes the received text data and extracts the meal details.
[1809] Step 7:
[1810] The server analyzes the received image data and calculates the calories and nutritional balance of the meal contents.
[1811] Step 8:
[1812] The server uses an emotion engine to recognize the user's emotional state from the voice data.
[1813] Step 9:
[1814] Based on the analysis results and emotional state, the server generates advice such as, "Today's breakfast seems low in calories. I recommend eating fruits and vegetables at your next meal."
[1815] Step 10:
[1816] The server sends the generated advice to the terminal.
[1817] Step 11:
[1818] The device will provide advice to the elderly via voice.
[1819] Medication management function
[1820] Step 1:
[1821] The server recognizes pre-set medication times and generates reminders.
[1822] Step 2:
[1823] The device will then send a voice reminder to the elderly person saying, "It's time to take your medicine. Are you ready?"
[1824] Step 3:
[1825] The user responds, "Yes, I'll take my medicine."
[1826] Step 4:
[1827] The terminal uses a voice recognition system to convert the user's response into text data.
[1828] Step 5:
[1829] The device activates the camera and takes a picture of the elderly person taking their medicine.
[1830] Step 6:
[1831] The video data acquired by the terminal is transmitted to the server.
[1832] Step 7:
[1833] The server analyzes the video data and confirms that the medication has been taken.
[1834] Step 8:
[1835] The server uses an emotion engine to analyze the user's emotional state.
[1836] Step 9:
[1837] The server generates a confirmation message saying "Your medication was taken correctly" and sets the next reminder.
[1838] Step 10:
[1839] The server sends a confirmation message to the terminal.
[1840] Step 11:
[1841] The device will then send a confirmation message to the elderly person.
[1842] Early detection of dementia
[1843] Step 1:
[1844] The device begins a casual conversation by asking, "Tell me how your day was today."
[1845] Step 2:
[1846] Users talk about everyday events.
[1847] Step 3:
[1848] The terminal uses a voice recognition system to convert the user's speech into text data.
[1849] Step 4:
[1850] The terminal transmits the acquired text data to the server.
[1851] Step 5:
[1852] The server analyzes the text data and evaluates the elderly person's language patterns.
[1853] Step 6:
[1854] The server uses an emotion engine to analyze the user's emotional state from their speech.
[1855] Step 7:
[1856] Based on the analysis results and emotional state, the server generates an alert such as, "Recently, there has been a decrease in vocabulary in conversations."
[1857] Step 8:
[1858] The server generates an alert and sends it to the device.
[1859] Step 9:
[1860] The device will notify the elderly person of the alert.
[1861] Step 10:
[1862] The server will also notify caregivers and family members of the alerts as needed.
[1863] These processing steps enable the remote care system of the present invention to monitor the user's living conditions with high accuracy and provide appropriate care and support. In addition, the emotion engine enables flexible responses according to the user's emotional state, improving the user's quality of life and reducing the burden on caregivers.
[1864] Example 2
[1865] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1866] For elderly people to maintain independent lifestyles, it is important to properly monitor their daily health and emotional states and provide timely advice and reminders. However, current systems have difficulty effectively analyzing the subject's voice and images and providing appropriate advice and reminders that take their emotional state into account. Furthermore, the lack of personalized support tailored to each individual's health condition and specific lifestyle circumstances makes it difficult for elderly people to maintain appropriate diets and medication intake.
[1867] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1868] In this invention, the server includes an analysis means that uses a generative AI model to analyze voice data and image data and evaluate health conditions and living situations, an emotion engine that recognizes emotional states from voice data, a means that analyzes dietary content and medication status to generate advice and reminders necessary for the elderly, and a notification means that notifies the elderly of the generated advice and reminders by voice. This enables comprehensive analysis using the voices and images of the elderly, and makes it possible to provide appropriate, personalized advice and reminders.
[1869] "Speech recognition means" is a technology that collects speech and converts it into text data.
[1870] "Image recognition means" is a technology that collects images and sends the data to a server.
[1871] A "generative AI model" is an artificial intelligence technology that analyzes voice and image data to assess the health and living conditions of elderly people.
[1872] "Analysis means" refers to the process of analyzing and evaluating audio and image data using a generative AI model.
[1873] "Emotion engine" is a technology that recognizes the user's emotional state from voice data.
[1874] The "notification means" is a technology that notifies the elderly of the generated advice and reminders by voice.
[1875] A "reminder" is a notification that helps seniors remember to take certain actions, such as taking medicine.
[1876] "Text data" refers to character data converted from speech by speech recognition means.
[1877] The present invention is a remote care system that monitors the living conditions of elderly people and provides appropriate support, and is equipped with an emotion engine that analyzes voice data and image data of the elderly person and recognizes the emotional state of the user. Specific embodiments of the present invention will be described below.
[1878] System configuration
[1879] This system consists of a voice recognition means, an image recognition means, an analysis means, an emotion engine, and a notification means.
[1880] Voice recognition means
[1881] The device is equipped with a microphone that collects conversations with elderly people in real time. The collected voice data is converted into text data through a voice recognition system, allowing the user's speech to be captured as text information.
[1882] Image Recognition Method
[1883] The device is equipped with a camera that takes pictures of the elderly person eating and taking their medicine. The captured image data is sent to a server in real time, allowing for a visual record of the elderly person's behavior.
[1884] Analysis means
[1885] The server is equipped with a generative AI model with advanced analytical capabilities that analyzes the voice and image data sent from the device. This analysis makes it possible to evaluate the health and living conditions of elderly people in real time. The generative AI model performs highly accurate analysis of the data, using techniques such as deep learning and natural language processing.
[1886] Emotion Engine
[1887] The server's emotion engine has the ability to recognize the user's emotional state from voice data, allowing it to assess not just their physical state but also their psychological state. Emotion analysis is performed based on the tone of voice and the choice of words used.
[1888] Notification means
[1889] Based on the analysis results and emotional state, the server generates advice and reminders for the elderly, which are then sent to the device, which then notifies the elderly by voice.
[1890] Meal management function
[1891] Overview: The dietary management function monitors whether the elderly are eating a balanced diet and provides advice as needed.
[1892] The device asks the senior, "What did you have for breakfast today?"
[1893] The user responds, "I had bread and milk."
[1894] The device uses a voice recognition system to convert the user's speech into text data and a camera to capture the meal.
[1895] The server analyzes text and image data to evaluate the calorie and nutritional balance of the meal, and also analyzes the user's emotional state using an emotion engine.
[1896] The server generates advice such as, "Today's breakfast seems low in calories. We recommend adding more fruits and vegetables to your next meal."
[1897] The device will provide advice to the elderly via voice.
[1898] Example prompt sentence:
[1899] Prompts for older adults regarding dietary management functions:
[1900] Please explain the analysis process and advice generation steps when an elderly person answers "I had bread and milk" to the question "What did you have for breakfast today?"
[1901] Medication management function
[1902] Abstract: Medication management is a function that helps elderly people take their medications appropriately.
[1903] The server detects when it's time to take the medication and generates a reminder.
[1904] The device will remind the elderly person, "It's time to take your medicine. Are you ready?"
[1905] The user responds, "Yes, I'll drink it now."
[1906] The device uses a voice recognition system to convert responses into text data and a camera to record the patient taking the medicine.
[1907] The server analyzes the video data to confirm whether the user has taken the medicine, and also analyzes the user's emotional state using an emotion engine.
[1908] The server generates a confirmation message saying "Your medication was taken correctly" and sets the next reminder.
[1909] The device will then send a confirmation message to the elderly person.
[1910] Example prompt sentence:
[1911] Prompts regarding medication management functions for seniors:
[1912] Please explain the analysis process and reminder generation steps when an elderly person responds "Yes, I'll take it right away" to the reminder "It's time to take your medicine. Are you ready?"
[1913] Early detection of dementia
[1914] Overview: The early dementia detection function detects signs of dementia from the everyday conversations of elderly people.
[1915] The device begins a casual conversation by asking, "Tell me how your day was today."
[1916] Users talk about everyday events.
[1917] The terminal uses a voice recognition system to convert the conversation into text data and send it to the server.
[1918] The server analyzes the text data, assessing language patterns and emotional states, and also analyzes emotional fluctuations using an emotion engine.
[1919] The server generates an alert such as, "Recently, we have noticed a decline in vocabulary in conversations."
[1920] The device will notify the senior of the alert and, if necessary, their caregiver or family member.
[1921] Example prompt sentence:
[1922] Here are some prompts for early dementia detection:
[1923] Please explain the steps for analyzing and generating alerts after an elderly person talks about their daily life when asked, "Tell me how your day was today."
[1924] In this way, the remote care system of the present invention supports the lives of elderly people in a variety of ways, and by combining it with an emotion engine, it is possible to provide more accurate advice and appropriate care.
[1925] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1926] Meal management function processing steps
[1927] Step 1:
[1928] Voice input and data collection (terminal)
[1929] The device uses a microphone to ask the elderly person, "What did you have for breakfast today?" The elderly person responds, "I had bread and milk." This voice input is collected.
[1930] Input: Voice input for the elderly
[1931] Output: Audio data collected by microphone
[1932] Step 2:
[1933] Speech recognition and data conversion (terminal)
[1934] The device uses a voice recognition system to convert voice data into text data, and also uses a camera to record the elderly person's mealtimes.
[1935] Input: Audio data, video data of meal
[1936] Output: Text data, image data
[1937] Step 3:
[1938] Data transmission (terminal)
[1939] The terminal encrypts the converted text data and the captured image data and transmits them to the server in real time.
[1940] Input: Text data, image data
[1941] Output: Data sent to the server
[1942] Step 4:
[1943] Data analysis (server)
[1944] The server uses the generative AI model to analyze the received text and image data and evaluate the calories and nutritional balance of the meal contents.
[1945] Input: Text data, image data
[1946] Output: Calorie and nutritional balance assessment results
[1947] Step 5:
[1948] Sentiment analysis (server)
[1949] The server's emotion engine recognizes the user's emotional state from the voice data.
[1950] Input: Text data
[1951] Output: User's emotional state
[1952] Step 6:
[1953] Advice Generation (Server)
[1954] The server generates advice based on the meal evaluation and emotional state, such as "We recommend adding more fruits and vegetables to your next meal."
[1955] Input: Calorie and nutritional balance assessment results, user's emotional state
[1956] Output: The generated advice
[1957] Step 7:
[1958] Notifications (Device)
[1959] The device notifies the elderly person of the advice received from the server via voice.
[1960] Input: Advice from the server
[1961] Output: Voice notification to the elderly
[1962] Medication management function processing steps
[1963] Step 1:
[1964] Medication time detection (server)
[1965] The server detects the preset time for taking the medication.
[1966] Input: Pre-set medication time
[1967] Output: Start generating medication reminders
[1968] Step 2:
[1969] Reminder generation (server)
[1970] The server generates a reminder: "It's time to take your medicine. Are you ready?"
[1971] Input: Medication time detection
[1972] Output: The generated reminder
[1973] Step 3:
[1974] Reminder notification (device)
[1975] The device will notify the elderly of reminders via voice.
[1976] Input: Reminder from server
[1977] Output: Voice notification to the elderly
[1978] Step 4:
[1979] User response (terminal)
[1980] The user responds, "Yes, I'll drink it now." The device uses a voice recognition system to convert the response into text data.
[1981] Input: User's voice response
[1982] Output: Text data
[1983] Step 5:
[1984] Medication confirmation (terminal)
[1985] The device uses a camera to record the patient taking the medicine.
[1986] Input: User's medication actions
[1987] Output: Video data
[1988] Step 6:
[1989] Data transmission (terminal)
[1990] The converted text data and video data are sent to the server.
[1991] Input: Text data, video data
[1992] Output: Data sent to the server
[1993] Step 7:
[1994] Data analysis (server)
[1995] The server analyzes the video data to confirm whether the user has taken the medicine, and also analyzes the user's emotional state using an emotion engine.
[1996] Input: Text data, video data
[1997] Output: Medication confirmation result, user's emotional state
[1998] Step 8:
[1999] Confirmation message generation (server)
[2000] The server generates a confirmation message saying "Your medication was taken correctly" and sets the next reminder.
[2001] Input: Medication confirmation result, user's emotional state
[2002] Output: Generated confirmation message, next reminder
[2003] Step 9:
[2004] Confirmation notification (terminal)
[2005] The device will then provide a confirmation message to the elderly person via voice.
[2006] Input: Confirmation message from the server
[2007] Output: Voice notification to the elderly
[2008] Dementia early detection function processing steps
[2009] Step 1:
[2010] Starting a daily conversation (terminal)
[2011] The device speaks to the elderly person, asking, "Tell us how your day was today."
[2012] Input: Pre-configured question prompt
[2013] Output: Start a conversation with the elderly person
[2014] Step 2:
[2015] User response (terminal)
[2016] Users talk about everyday events, and the device converts the conversation into text data using a voice recognition system.
[2017] Input: User's daily conversation
[2018] Output: Text data
[2019] Step 3:
[2020] Data transmission (terminal)
[2021] The converted text data is sent to the server.
[2022] Input: Text data
[2023] Output: Data sent to the server
[2024] Step 4:
[2025] Data analysis (server)
[2026] The server analyzes the text data, assessing language patterns and emotional states, and also analyzes emotional fluctuations using an emotion engine.
[2027] Input: Text data
[2028] Output: Language pattern evaluation results, emotional state evaluation results
[2029] Step 5:
[2030] Alert Generation (Server)
[2031] The server generates an alert such as, "Recently, we have noticed a decline in vocabulary in conversations."
[2032] Input: Language pattern evaluation results, emotional state evaluation results
[2033] Output: The generated alert
[2034] Step 6:
[2035] Alert notification (terminal)
[2036] The device will notify the senior of the alert, and if necessary, a caregiver or family member.
[2037] Input: Alert from the server
[2038] Output: Notification to seniors and caregivers
[2039] (Application example 2)
[2040] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[2041] It is difficult to monitor the health and emotional states of elderly people in real time and provide appropriate advice and reminders when they go about their daily lives, especially when they use self-driving vehicles. Furthermore, conventional remote care systems do not provide support tailored to the elderly's situation in self-driving vehicles, which may reduce the elderly's sense of safety and security. Therefore, there is a need for a remote care system that can monitor the health and emotional states of elderly people in real time and provide appropriate support when they use self-driving vehicles.
[2042] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[2043] In this invention, the server includes a voice recognition means for acquiring voice data of the elderly person, an image recognition means for acquiring image data of the elderly person, an analysis means using a generative AI model to analyze the voice data and image data, a notification means for notifying the elderly person of necessary advice and reminders based on the analysis results, a means for monitoring the elderly person's health and emotional state while in the autonomously driven vehicle, a means for transmitting the data collected by the above means to the server and analyzing it, a means for providing the elderly person with advice regarding their living conditions in the autonomously driven vehicle based on the analysis results, and a means for notifying the elderly person of the advice visually and audibly. This makes it possible to support elderly people in living safely and healthily when using autonomously driven vehicles.
[2044] The "voice recognition means" is a means for acquiring and analyzing voice data of the elderly person.
[2045] The "image recognition means" is a means for acquiring image data of an elderly person and analyzing it.
[2046] A "generative AI model" is an advanced artificial intelligence model that analyzes audio and image data and is used to assess the health and emotional state of an elderly person.
[2047] The "analysis means" is a means for assessing the health and emotional state of an elderly person using audio and image data.
[2048] The "notification means" is a means for visually and audibly notifying the elderly of necessary advice and reminders based on the analysis results.
[2049] An "automated vehicle" is a vehicle that is driven automatically and can be used by seniors.
[2050] "Health status" refers to the physical condition and physical function of an elderly person.
[2051] "Emotional state" refers to the psychological state and emotional movements of elderly people.
[2052] The "means for transmitting data to a server" is a means for transmitting voice data and image data collected from elderly people to a server.
[2053] "Visual and audio notification means" refers to means for conveying advice to the elderly person through a visual display and audio.
[2054] The present invention relates to a remote care system for elderly people to live safe and healthy lives when using autonomous vehicles. An embodiment of the present invention will be specifically described.
[2055] System configuration
[2056] Terminal configuration
[2057] 1. Voice recognition means:
[2058] The device is equipped with a microphone that is used to capture voice data from the elderly, which is then analyzed in real time.
[2059] 2. Image Recognition Methods:
[2060] The device is equipped with a camera that is used to capture image data of the elderly person and send it to a server. For example, the device captures the elderly person's facial expressions and body movements.
[2061] Server-side configuration
[2062] 1. Analysis method:
[2063] The server is equipped with a generative AI model that analyzes the voice and image data sent from the device, and this analysis evaluates the health and emotional state of the elderly person.
[2064] 2. Means of notification:
[2065] Based on the analysis results, the server generates necessary advice and reminders for the elderly, which are then sent to the device, which then notifies the elderly by voice or display.
[2066] Hardware and software used
[2067] Hardware:
[2068] Microphone: Used to collect voice data from the elderly.
[2069] Camera: Used to capture image data of the elderly.
[2070] Self-driving vehicles: Used to monitor the health and emotional state of elderly people in real time while they are in the vehicle.
[2071] software:
[2072] speech_recognition library: Used to achieve speech recognition.
[2073] OpenCV library: Used to realize image processing.
[2074] TensorFlow: Used to realize a generative AI model for elderly emotion engine analysis.
[2075] gTTS: Used to generate audio notifications.
[2076] requests: Used to realize data transmission to the server.
[2077] System operation example
[2078] 1. While the elderly person is riding in a self-driving vehicle:
[2079] Microphones and cameras capture the elderly person's voice and image data in real time, which is automatically sent to a server.
[2080] The server uses a generative AI model to analyze audio and image data to assess the health and emotional state of the elderly.
[2081] Based on the assessment results, necessary advice and reminders are generated and sent to the elderly via the device. For example, if a person does not eat enough breakfast, they will receive a voice message saying, "We recommend adding more fruits and vegetables to your next meal."
[2082] Prompt Sentence Examples
[2083] After the elderly person answers questions about health care in the self-driving car, please send the image and voice data to a server and generate a Python program that will provide appropriate advice and reminders.
[2084] As described above, the remote care system of the present invention provides multifaceted support to enable elderly people to use self-driving vehicles with peace of mind.
[2085] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[2086] Step 1: Acquire audio data
[2087] The device uses a microphone to capture voice data from the elderly. The input is the elderly's voice, which is collected in real time. The device stores the collected voice data either directly or temporarily.
[2088] Step 2: Acquiring image data
[2089] The device uses a camera to acquire image data (facial expressions and body movements) of the elderly. The input is a real-time image of the elderly acquired through the camera, which is captured at high resolution. The device stores the captured image data either directly or temporarily.
[2090] Step 3: Send data to the server
[2091] The terminal sends the collected voice and image data to the server. The input is the data acquired in steps 1 and 2 above, and this is sent to the server via the network. The output is the data itself that was sent to the server.
[2092] Step 4: Analyzing the audio data
[2093] The server analyzes the received voice data. The input is the voice data sent from the device, which is analyzed using a generative AI model. The content of the elderly person's speech and their emotional state are extracted from the voice data and output as text data.
[2094] Step 5: Analyzing the image data
[2095] The server analyzes the received image data. The input is image data sent from the device, which is analyzed using a generative AI model. The image data is used to evaluate the elderly person's facial expressions, body movements, health condition, etc., and is output as numerical data or evaluation results.
[2096] Step 6: Integrating the analysis results
[2097] The server integrates the results of the voice data analysis and the image data analysis to evaluate the elderly person's comprehensive health and emotional state. The input is the analysis results of steps 4 and 5, which are integrated to comprehensively grasp the elderly person's current condition. The output is the integrated evaluation result.
[2098] Step 7: Generate advice and reminders
[2099] The server generates necessary advice and reminders for the elderly based on the integrated evaluation results. The input is the integrated evaluation results, and appropriate messages are generated based on them. The output is text data of the generated advice and reminders.
[2100] Step 8: Implementing Notification
[2101] The device notifies the elderly of advice and reminders received from the server. The input is text data sent from the server, which is notified visually and audibly. Specifically, messages are displayed on the display and read aloud using gTTS. The output is the advice and reminders notified to the elderly.
[2102] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[2103] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2104] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[2105] [Fourth embodiment]
[2106] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[2107] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[2108] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[2109] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[2110] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[2111] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[2112] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[2113] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[2114] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[2115] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[2116] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[2117] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[2118] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2119] This remote care system consists of elderly end users, and a server and terminal (robot) located in a remote location. A specific embodiment of the system is described below. The program code itself is not provided, but the processing flow and each step are described in detail.
[2120] System configuration
[2121] 1. Terminal configuration
[2122] Speech recognition means: The device is equipped with a microphone to collect conversations with the elderly in real time, and speech recognition technology is used to convert the captured speech into text data.
[2123] Image recognition method: The device is equipped with a camera that captures images of the elderly person's daily life, especially when they eat or take medicine. The image data is sent to a server in real time.
[2124] 2. Server-side configuration
[2125] Analysis method: The server uses a generative AI model to analyze the voice and image data sent from the device. This analysis allows for a real-time assessment of the elderly person's health and living conditions.
[2126] Notification method: Based on the analysis results, appropriate advice and reminders are generated and sent to the device, which then communicates this to the elderly via voice.
[2127] Meal management function for the elderly
[2128] Overview: The dietary management function monitors whether the elderly are eating a balanced diet.
[2129] Terminal: Ask the senior, "What did you have for breakfast today?"
[2130] User: "I had bread and milk."
[2131] Device: Converts voice into text data and simultaneously takes pictures of the meal using a camera.
[2132] Server: Analyzes the received text and image data and evaluates the calorie and nutritional balance of the meal.
[2133] Server: Generates advice as needed, such as "Today's breakfast appears to be low in calories. We recommend eating some fruits and vegetables at your next meal."
[2134] Device: Provides advice to the elderly via voice.
[2135] Medication management function
[2136] Overview: Medication management helps seniors take their medications appropriately.
[2137] Server: Detects when it's time to take your medicine and generates a reminder saying, "It's time to take your medicine. Are you ready?"
[2138] Device: Provides voice reminders to seniors.
[2139] User: "Yes, I'll take my medicine."
[2140] Device: Activate the camera and record the scene of taking the medicine.
[2141] Server: Analyzes image data and verifies whether the medication has been taken.
[2142] Server: If the medication is confirmed, generate a confirmation message saying "Your medication was taken correctly" and set the next reminder.
[2143] Terminal: Notify the elderly person of a confirmation message.
[2144] Early detection of dementia
[2145] Overview: The early dementia detection function detects signs of dementia from the everyday conversations of elderly people.
[2146] Device: Initiate everyday conversations such as "Tell me how your day was."
[2147] User: Talk about everyday events.
[2148] Terminal: Converts the voice into text data and sends it to the server.
[2149] Server: Analyzes conversation data and evaluates language patterns and emotional states, for example, detecting vocabulary declines and changes in emotional tone.
[2150] Server: If signs of dementia are detected, it generates an alert such as, "Recently, there has been a decline in vocabulary in conversations."
[2151] Device: Notifies the senior of the alert and, if necessary, notifies caregivers and family members.
[2152] As described above, this remote care system provides comprehensive support for the daily lives of elderly people and makes it easier for caregivers and family members in remote locations to understand the condition of the elderly, thereby improving the quality of life of the elderly and reducing the burden on caregivers.
[2153] The processing flow will be explained below.
[2154] Meal management function for the elderly
[2155] Step 1:
[2156] The device asks the senior, "What did you have for breakfast today?"
[2157] Step 2:
[2158] The user responds, "I had bread and milk."
[2159] Step 3:
[2160] The terminal uses a voice recognition system to convert the user's speech into text data.
[2161] Step 4:
[2162] The device activates the camera and takes pictures of the elderly person eating.
[2163] Step 5:
[2164] The terminal transmits the acquired text data and image data to the server.
[2165] Step 6:
[2166] The server analyzes the received text data and extracts the meal details.
[2167] Step 7:
[2168] The server analyzes the received image data and calculates the calories and nutritional balance of the meal contents.
[2169] Step 8:
[2170] Based on the analysis results, the server generates advice such as, "Today's breakfast appears to be low in calories. We recommend that you eat fruits and vegetables at your next meal."
[2171] Step 9:
[2172] The server sends the generated advice to the terminal.
[2173] Step 10:
[2174] The device will provide advice to the elderly via voice.
[2175] Medication management function
[2176] Step 1:
[2177] The server recognizes pre-set medication times and generates reminders.
[2178] Step 2:
[2179] The device will then send a voice reminder to the elderly person saying, "It's time to take your medicine. Are you ready?"
[2180] Step 3:
[2181] The user responds, "Yes, I'll take my medicine."
[2182] Step 4:
[2183] The terminal uses a voice recognition system to convert the user's response into text data.
[2184] Step 5:
[2185] The device activates the camera and takes a picture of the elderly person taking their medicine.
[2186] Step 6:
[2187] The video data acquired by the terminal is transmitted to the server.
[2188] Step 7:
[2189] The server analyzes the video data and confirms that the medication has been taken.
[2190] Step 8:
[2191] The server generates a confirmation message saying "Your medication was taken correctly" and sets the next reminder.
[2192] Step 9:
[2193] The server sends a confirmation message to the terminal.
[2194] Step 10:
[2195] The device will then send a confirmation message to the elderly person.
[2196] Early detection of dementia
[2197] Step 1:
[2198] The device begins a casual conversation with the elderly person, asking, "Tell me how your day was today."
[2199] Step 2:
[2200] Users talk about everyday events.
[2201] Step 3:
[2202] The terminal uses a voice recognition system to convert the user's speech into text data.
[2203] Step 4:
[2204] The terminal transmits the acquired text data to the server.
[2205] Step 5:
[2206] The server analyzes the text data and evaluates the elderly person's language patterns and emotional state.
[2207] Step 6:
[2208] The server detects signs of dementia based on the analysis results.
[2209] Step 7:
[2210] The server generates an alert such as, "Recently, we have noticed a decline in vocabulary in conversations."
[2211] Step 8:
[2212] The server generates alerts and sends them to the device, which also notifies family members and caregivers.
[2213] Step 9:
[2214] The device will notify the elderly person of the alert.
[2215] These processing steps are designed to ensure that the lives of the elderly are managed effectively and to reduce the burden on caregivers.
[2216] Example 1
[2217] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2218] It is difficult for elderly people to maintain appropriate lifestyle habits and monitor their health. In particular, it is important for elderly people to manage the nutritional balance of their diet, their medication status, and even to detect dementia early, but it is difficult for them to do these things alone. It is also not easy for caregivers and family members in remote locations to understand the condition of elderly people. Conventional systems lack the functionality to comprehensively support the lives of elderly people.
[2219] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[2220] In this invention, the server includes a voice recognition means for acquiring voice data of the elderly person, an image recognition means for acquiring image data of the elderly person, an analysis means using a generative AI model to analyze the voice data and image data, a notification means for notifying the elderly person of necessary advice and reminders based on the analysis results, a means for detecting signs of dementia from the elderly person's everyday conversation and generating an alert, a means for reminding the elderly person to take their medicine and confirming that they have taken it, and a means for monitoring the elderly person's diet to evaluate nutritional balance and provide advice. This makes it possible to comprehensively support the elderly's living conditions and enable caregivers and family members in remote locations to easily understand the elderly person's condition.
[2221] The "voice recognition means" is a device or technology for capturing the voice of the elderly person in real time and converting the voice data into text data.
[2222] "Image recognition means" refers to a device or technology for photographing the elderly person's daily life and acquiring image data.
[2223] A "generative AI model" is an artificial intelligence model that analyzes voice and image data and generates appropriate advice and reminders based on the results.
[2224] The "analysis means" is a means of evaluating the health status and living conditions of elderly people using a generative AI model, using the acquired voice and image data.
[2225] The "notification means" is a means of notifying the elderly of advice and reminders generated based on the analysis results via voice or other means.
[2226] The "means for detecting signs of dementia" is a method for analyzing data on everyday conversations of elderly people and detecting early signs of dementia from factors such as a decrease in vocabulary and changes in emotional tone.
[2227] "Reminder measures" are measures that notify elderly people when it is time to take their medicine and help them ensure that they take their medicine.
[2228] A "dietary monitoring method" is a method for monitoring the dietary content of elderly people, evaluating their nutritional balance and calories, and providing necessary advice.
[2229] This invention is a remote care system for monitoring and supporting the living conditions of elderly people. The system consists of an elderly person in a remote location, a server, and a terminal (robot). This system is configured and operates as follows to monitor the elderly person's daily behavior and health status and provide necessary advice and reminders.
[2230] System configuration
[2231] Terminal configuration
[2232] Speech recognition method: The device is equipped with a microphone to collect conversations with the elderly in real time. The captured speech is converted into text data using speech recognition technology (e.g., Google Cloud Speech-to-Text).
[2233] Image recognition method: The device is equipped with a camera that captures images of the elderly person's daily life, especially when they eat or take medicine. The image data is sent to a server in real time.
[2234] Server-side configuration
[2235] Analysis method: The server uses a generative AI model (e.g., OpenAI GPT-4) to analyze the voice and image data sent from the device. This analysis allows for a real-time assessment of the elderly person's health and living conditions.
[2236] Notification method: Based on the analysis results, appropriate advice and reminders are generated and sent to the device, which then communicates this to the elderly via voice.
[2237] Meal management function for the elderly
[2238] Overview: The dietary management function monitors whether the elderly are eating a balanced diet.
[2239] Terminal: Ask the senior, "What did you have for breakfast today?"
[2240] User: "I had bread and milk."
[2241] Device: Converts voice into text data and simultaneously takes pictures of the meal using a camera.
[2242] Server: Analyzes the received text and image data and evaluates the calorie and nutritional balance of the meal.
[2243] Server: Generates advice as needed, such as "Today's breakfast appears to be low in calories. We recommend eating some fruits and vegetables at your next meal."
[2244] Device: Provides advice to the elderly via voice.
[2245] Specific prompt examples:
[2246] "If an elderly person eats bread and milk for breakfast, generate a sentence that evaluates the nutritional balance and advises them to add more fruits and vegetables to their next meal if the calories are low."
[2247] Medication management function
[2248] Overview: Medication management helps seniors take their medications appropriately.
[2249] Server: Detects when it's time to take your medicine and generates a reminder saying, "It's time to take your medicine. Are you ready?"
[2250] Device: Provides voice reminders to seniors.
[2251] User: "Yes, I'll take my medicine."
[2252] Device: Activate the camera and record the scene of taking the medicine.
[2253] Server: Analyzes image data and verifies whether the medication has been taken.
[2254] Server: If the medication is confirmed, generate a confirmation message saying "Your medication was taken correctly" and set the next reminder.
[2255] Terminal: Notify the elderly person of a confirmation message.
[2256] Specific prompt examples:
[2257] "Generate a reminder when it's time for the senior to take their medication, and then create a sentence to confirm if they actually took their medication."
[2258] Early detection of dementia
[2259] Overview: The early dementia detection function detects signs of dementia from the everyday conversations of elderly people.
[2260] Device: Initiate everyday conversations such as "Tell me how your day was."
[2261] User: Talk about everyday events.
[2262] Terminal: Converts the voice into text data and sends it to the server.
[2263] Server: Analyzes conversation data and evaluates language patterns and emotional states, for example, detecting vocabulary declines and changes in emotional tone.
[2264] Server: If signs of dementia are detected, it generates an alert such as, "Recently, there has been a decline in vocabulary in conversations."
[2265] Device: Notifies the senior of the alert and, if necessary, notifies caregivers and family members.
[2266] Specific prompt examples:
[2267] "To detect signs of dementia in everyday conversations of elderly people, create sentences that analyze vocabulary decline and changes in emotional tone."
[2268] As described above, this remote care system provides comprehensive support for the daily lives of elderly people and makes it easier for caregivers and family members in remote locations to understand the condition of the elderly, thereby improving the quality of life of the elderly and reducing the burden on caregivers.
[2269] The flow of the identification process in the first embodiment will be described with reference to FIG.
[2270] Meal management function for the elderly
[2271] Step 1:
[2272] The device asks the senior, "What did you have for breakfast today?"
[2273] Input: Breakfast Question
[2274] Output: Waiting for response from the elderly person
[2275] Step 2:
[2276] The user responds, "I had bread and milk."
[2277] Input: Elderly person's voice response
[2278] Output: Audio data
[2279] Step 3:
[2280] The device converts the voice into text data and simultaneously captures the meal with a camera.
[2281] Input: Audio data, video of elderly people eating
[2282] Output: Text data, image data
[2283] What it does: The microphone captures your voice and converts it into text using a speech recognition engine. The camera captures footage of you eating.
[2284] Step 4:
[2285] The server receives the text and image data and analyzes it using a generative AI model.
[2286] Input: Text data, image data
[2287] Output: Calorie and nutritional balance evaluation results of the meal
[2288] How it works: The server extracts meal items from text data and analyzes the meal contents using an image recognition algorithm. The generative AI model calculates nutritional balance and calories.
[2289] Step 5:
[2290] The server generates advice based on the analysis results.
[2291] Input: Evaluation results of calorie and nutritional balance of meals
[2292] Output: Advice message
[2293] Specific behavior: The server generates advice such as "Today's breakfast is low in calories. We recommend that you eat fruits and vegetables at your next meal."
[2294] Step 6:
[2295] The device will provide advice to the elderly via voice.
[2296] Input: Advice message
[2297] Output: Voice notification to the elderly
[2298] Specific operation: Advice messages are synthesized and conveyed to the elderly.
[2299] Medication management function
[2300] Step 1:
[2301] The server detects when it's time to take your medicine and generates a reminder saying, "It's time to take your medicine. Are you ready?"
[2302] Input: Medication schedule
[2303] Output: Reminder message
[2304] What it does: The server periodically checks the schedule and generates a reminder message when it's time to take your medication.
[2305] Step 2:
[2306] The device will notify the elderly of the reminder via voice.
[2307] Input: Reminder message
[2308] Output: Voice notification to the elderly
[2309] Specific actions: Reminder messages are delivered to elderly people via voice synthesis.
[2310] Step 3:
[2311] The user responds, "Yes, I'll take my medicine."
[2312] Input: Elderly person's voice response
[2313] Output: Response audio data
[2314] Step 4:
[2315] The device activates the camera and captures the scene of the elderly person taking their medicine.
[2316] Input: Video of elderly person taking medication
[2317] Output: Image data of the medication scene
[2318] Specific operation: The camera captures the elderly person taking their medicine and saves the image data on the device.
[2319] Step 5:
[2320] The server analyzes the image data and verifies whether the medication has been taken.
[2321] Input: Image data of a medication scene
[2322] Output: Medication confirmation result
[2323] What it does: The server uses image recognition technology to verify that the medication was taken correctly.
[2324] Step 6:
[2325] The server generates a confirmation message and sets the next reminder.
[2326] Input: Medication confirmation result
[2327] Output: Confirmation message, next reminder setting
[2328] Specific behavior: If the medication is confirmed, the server generates a confirmation message saying "Your medication was taken correctly" and sets the next reminder.
[2329] Step 7:
[2330] The device will then send a confirmation message to the elderly person.
[2331] Input: Confirmation message
[2332] Output: Voice notification to the elderly
[2333] Specific operation: A confirmation message is synthesized and conveyed to the elderly person.
[2334] Early detection of dementia
[2335] Step 1:
[2336] The device will begin a casual conversation such as, "Tell me how your day was today."
[2337] Input: Daily conversation question
[2338] Output: Waiting for response from the elderly person
[2339] Step 2:
[2340] Users talk about everyday events.
[2341] Input: Elderly person's voice
[2342] Output: Audio data
[2343] Step 3:
[2344] The device converts the voice into text data and sends it to the server.
[2345] Input: Audio data
[2346] Output: Text data
[2347] What it does: The microphone captures your voice and the speech recognition engine converts it into text.
[2348] Step 4:
[2349] The server analyzes the conversation data and evaluates language patterns and emotional states.
[2350] Input: Text data
[2351] Output: Evaluation result
[2352] Specific operation: The server uses the generative AI model to detect vocabulary loss, changes in emotional tone, etc.
[2353] Step 5:
[2354] If the server detects signs of dementia, it generates an alert message.
[2355] Input: Evaluation result
[2356] Output: Alert message
[2357] Specific behavior: If signs of dementia are detected, an alert will be generated, such as "Recently, there has been a decline in vocabulary in conversations."
[2358] Step 6:
[2359] The device notifies the elderly person of the alert and, if necessary, notifies caregivers and family members.
[2360] Input: Alert message
[2361] Output: Audio notification to the senior and caregiver / family
[2362] Specific operation: An alert message is synthesized and conveyed to the elderly. At the same time, an alert is sent to caregivers and family members.
[2363] The above are the specific processing steps and their detailed operations. These steps improve the quality of life for the elderly and enable appropriate care to be provided remotely.
[2364] (Application example 1)
[2365] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2366] Remote care systems that support the lives of the elderly are limited in their management functions to the health status of the elderly, and do not address comprehensive home safety. There is a need for a system that comprehensively monitors security aspects such as managing entry and exit within the home and detecting fires, gas leaks, and suspicious individuals. The objective of this invention is to provide a comprehensive support system that not only manages the health of the elderly, but also monitors the safety status of the home.
[2367] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[2368] In this invention, the server includes an analysis means, a notification means, and a safety monitoring means, which makes it possible not only to grasp and support the living conditions of the elderly, but also to comprehensively monitor the safety of the home.
[2369] A "remote care system" is a system that monitors and supports the living conditions of elderly people from a remote location.
[2370] The "voice recognition means" is a device that has the function of acquiring voice data from the elderly person and converting it into text data in real time.
[2371] The "image recognition means" is a device that has the function of acquiring image data of an elderly person and analyzing the content of that image data.
[2372] A "generative AI model" is an artificial intelligence model that analyzes acquired voice data and image data and generates the necessary information.
[2373] The "analysis means" is a function that analyzes voice data and image data to evaluate the living conditions of the elderly person.
[2374] "Notification means" is a function that notifies elderly people of advice and reminders based on the analysis results.
[2375] The "safety monitoring means" is a device that has the function of monitoring the safety status of the home, detecting abnormalities, and notifying them.
[2376] "Nutritional balance" is an indicator of whether nutrients are evenly distributed in a particular diet.
[2377] A "calorie" is a unit that indicates the amount of energy ingested through food.
[2378] A "reminder" is a notification that encourages seniors to take a specific action.
[2379] This remote care system is a system for monitoring and supporting the living conditions and home safety conditions of elderly people. A specific embodiment of the system will be described below.
[2380] System configuration
[2381] The server and user terminals (smartphones) work together to monitor the lives of the elderly and the safety of their homes using various sensor devices. This system includes the following hardware and software:
[2382] Hardware
[2383] Smartphone: microphone, camera
[2384] Sensors: Door and window sensors, smoke detectors, gas leak detectors
[2385] software
[2386] Speech recognition API: Used to convert voice data into text data. Typical examples include Google Speech-to-Text and Apple Siri.
[2387] Image recognition API: Used to analyze image data. Representative examples include Google Cloud Vision and Amazon Rekognition.
[2388] Server: Platforms that run generative AI models for data analysis include Google Cloud and AWS (Amazon Web Services).
[2389] System Operation
[2390] The server uses voice recognition, image recognition, and safety monitoring means to notify the user based on the analysis results. Each of these means is described in detail below.
[2391] Voice recognition means
[2392] The microphone on the user's device collects the elderly's voice and converts it into text data in real time using a speech recognition API. For example, if an elderly person says "help me," it is converted into text data and sent to the server.
[2393] Image Recognition Method
[2394] The camera on the user's device captures images of the elderly's living conditions and home situation, and the images are analyzed using an image recognition API. Examples include detecting fires, gas leaks, and suspicious people. It also includes capturing images of elderly people taking medicine, and analyzing the image data on a server.
[2395] Safety monitoring means
[2396] Sensor devices such as door and window sensors, smoke detectors, and gas leak detectors monitor the safety status of the home and notify the server if an abnormality is detected. The server analyzes the abnormality and issues a warning to the user in real time.
[2397] Server analysis method
[2398] The server analyzes the voice and image data using a generative AI model, and based on the analysis results, generates appropriate advice and reminders and sends them to the user's device.
[2399] Notification means
[2400] Based on the analysis results, the system will provide voice reminders and warnings to the elderly, such as specific instructions such as "There is a fire. Please close the doors of each room."
[2401] Specific examples
[2402] In case of fire detection
[2403] 1. The microphone detects the sound of a fire alarm.
[2404] 2. The voice recognition API converts the message "Fire has broken out" into text data.
[2405] 3. Analyze fire data on the server.
[2406] 4. An "emergency fire alert" notification will be sent to your smartphone.
[2407] 5. An audio guide will be given to residents to "close the doors of each room."
[2408] Prompt Sentence Examples
[2409] "Design a program that uses a voice recognition system to detect fire alarm sounds in real time and notify a smartphone of that information. If an abnormal sound is detected, provide voice guidance on first aid measures for the fire."
[2410] This will make it possible to realize a system that comprehensively supports the lives and home safety of the elderly.
[2411] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[2412] Step 1:
[2413] The microphone on the user's device captures the elderly person's voice data, which is then sent to a voice recognition API and converted into text data in real time.
[2414] Input: Elderly voice data
[2415] Output: Text data
[2416] Specific operation: When an elderly person says "help me," their voice is picked up by a microphone and converted into text data using a voice recognition API.
[2417] Step 2:
[2418] The camera on the user's device captures image data of the elderly person, which is then sent to an image recognition API for real-time analysis.
[2419] Input: Image data of elderly people
[2420] Output: Analysis data
[2421] Specific operation: The camera captures scenes of elderly people eating or taking medicine, and the image data is analyzed using an image recognition API.
[2422] Step 3:
[2423] The server receives the voice and image data sent from the user's device and analyzes the data using a generative AI model.
[2424] Input: Text data, analysis data
[2425] Output: Analysis results
[2426] Specific operation: The voice data is converted into text and the analysis results of the image data are sent to a server, and the generated AI model analyzes the elderly person's living conditions and health condition.
[2427] Step 4:
[2428] The server generates appropriate advice and reminders based on the analysis results.
[2429] Input: Analysis results
[2430] Output: Advice, reminder
[2431] Specific behavior: For example, if an elderly person's breakfast is analyzed to be insufficient in calories, advice such as "We recommend that you eat fruits and vegetables at your next meal" will be generated.
[2432] Step 5:
[2433] The server transmits the generated advice and reminders to the user terminal using a notification means.
[2434] Input: Advice, Reminder
[2435] Output: Notification data
[2436] Specific operation: The generated advice and reminders are sent to the user's device as push notifications, and the elderly person receives notifications such as "It's time to take their medicine."
[2437] Step 6:
[2438] Sensors on the user's device (doors, windows, smoke, gas leaks) monitor the safety status of the home, and if an abnormality is detected, the data is sent to the server.
[2439] Input: Sensor data
[2440] Output: Anomaly detection data
[2441] Specific operation: If a door or window is opened or closed suspiciously, or if a fire or gas leak is detected, the information is sent to the server.
[2442] Step 7:
[2443] The server analyzes the anomaly detection data, generates necessary warnings, and sends them to the user terminal.
[2444] Input: Anomaly detection data
[2445] Output: Warning notification data
[2446] Specific operation: For example, a warning message such as "A fire has been detected. Please close the doors of each room" is generated and sent to the user terminal.
[2447] Step 8:
[2448] The user terminal notifies the elderly person of the received warning notification data by voice in real time.
[2449] Input: Alert notification data
[2450] Output: Audio notification
[2451] Specific actions: The elderly person will be notified by voice, "A fire has broken out. Please close the doors of each room."
[2452] This will enable comprehensive monitoring of the elderly's living conditions and home safety, and provide appropriate support at the appropriate time.
[2453] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[2454] This invention is a remote care system that monitors the living conditions of elderly people and provides appropriate support, and is equipped with an emotion engine that analyzes voice data and image data of the elderly person and recognizes the emotional state of the user. Specific embodiments of the invention will be described below.
[2455] System configuration
[2456] 1. Terminal configuration
[2457] Voice recognition means: The device is equipped with a microphone to collect conversations with the elderly in real time. The collected voice data is converted into text data by a voice recognition system.
[2458] Image recognition means: The device is equipped with a camera that takes pictures of, for example, an elderly person eating or taking medicine. The captured image data is sent to a server in real time.
[2459] 2. Server-side configuration
[2460] Analysis method: The server is equipped with a generative AI model with advanced analytical capabilities that analyzes the voice and image data sent from the device. This analysis allows for a real-time assessment of the elderly person's health and living conditions.
[2461] Emotion engine: The server is also equipped with an emotion engine that recognizes the emotional state of the elderly by analyzing voice data.
[2462] 3. Means of notification
[2463] Advice and reminders: Based on the analysis results and emotional state, advice and reminders are generated for the elderly. This information is sent from the server to the device, which then notifies the elderly by voice.
[2464] Meal management function for the elderly
[2465] Overview: The dietary management function monitors whether the elderly are eating a balanced diet and provides advice as needed.
[2466] Terminal: Ask the senior, "What did you have for breakfast today?"
[2467] User: "I had bread and milk."
[2468] Device: A voice recognition system converts the user's speech into text data. A camera captures the meal.
[2469] Server: Analyzes text and image data to evaluate the calorie and nutritional balance of meal content. Also analyzes the user's emotional state using an emotion engine.
[2470] Server: Generate advice such as, "Your breakfast today seems low in calories. We recommend adding more fruits and vegetables to your next meal."
[2471] Device: Provides advice to the elderly via voice.
[2472] Medication management function
[2473] Abstract: Medication management is a function that helps elderly people take their medications appropriately.
[2474] Server: Detects medication times and generates reminders.
[2475] Device: Reminds seniors, "It's time to take your medicine. Are you ready?"
[2476] User: "Yes, I'll drink it now."
[2477] Device: A voice recognition system converts responses into text data, and a camera captures the process of taking the medicine.
[2478] Server: Analyzes video data and confirms whether the user has taken the medication. An emotion engine also analyzes the user's emotional state.
[2479] Server: Generates a confirmation message saying "Your medication was taken correctly" and sets the next reminder.
[2480] Terminal: Notify the elderly person of a confirmation message.
[2481] Early detection of dementia
[2482] Overview: The early dementia detection function detects signs of dementia from the everyday conversations of elderly people.
[2483] Device: Begins a conversation with "Tell me how your day was."
[2484] User: Talk about everyday events.
[2485] Terminal: The speech recognition system converts the conversation into text data and sends it to the server.
[2486] Server: Analyzes text data and evaluates language patterns and emotional states. An emotion engine also analyzes emotional fluctuations.
[2487] Server: Generate an alert such as, "Recently, we have noticed a decline in vocabulary in conversations."
[2488] Device: Notifies the senior of the alert and, if necessary, notifies caregivers and family members.
[2489] In this way, the remote care system of the present invention supports the lives of elderly people in a variety of ways, and by combining it with an emotion engine, it is possible to provide more accurate advice and appropriate care.
[2490] The processing flow will be explained below.
[2491] Meal management function for the elderly
[2492] Step 1:
[2493] The device asks the senior, "What did you have for breakfast today?"
[2494] Step 2:
[2495] The user responds, "I had bread and milk."
[2496] Step 3:
[2497] The terminal uses a voice recognition system to convert the user's speech into text data.
[2498] Step 4:
[2499] The device activates the camera and takes pictures of the elderly person eating.
[2500] Step 5:
[2501] The terminal transmits the acquired text data and image data to the server.
[2502] Step 6:
[2503] The server analyzes the received text data and extracts the meal details.
[2504] Step 7:
[2505] The server analyzes the received image data and calculates the calories and nutritional balance of the meal contents.
[2506] Step 8:
[2507] The server uses an emotion engine to recognize the user's emotional state from the voice data.
[2508] Step 9:
[2509] Based on the analysis results and emotional state, the server generates advice such as, "Today's breakfast seems low in calories. I recommend eating fruits and vegetables at your next meal."
[2510] Step 10:
[2511] The server sends the generated advice to the terminal.
[2512] Step 11:
[2513] The device will provide advice to the elderly via voice.
[2514] Medication management function
[2515] Step 1:
[2516] The server recognizes pre-set medication times and generates reminders.
[2517] Step 2:
[2518] The device will then send a voice reminder to the elderly person saying, "It's time to take your medicine. Are you ready?"
[2519] Step 3:
[2520] The user responds, "Yes, I'll take my medicine."
[2521] Step 4:
[2522] The terminal uses a voice recognition system to convert the user's response into text data.
[2523] Step 5:
[2524] The device activates the camera and takes a picture of the elderly person taking their medicine.
[2525] Step 6:
[2526] The video data acquired by the terminal is transmitted to the server.
[2527] Step 7:
[2528] The server analyzes the video data and confirms that the medication has been taken.
[2529] Step 8:
[2530] The server uses an emotion engine to analyze the user's emotional state.
[2531] Step 9:
[2532] The server generates a confirmation message saying "Your medication was taken correctly" and sets the next reminder.
[2533] Step 10:
[2534] The server sends a confirmation message to the terminal.
[2535] Step 11:
[2536] The device will then send a confirmation message to the elderly person.
[2537] Early detection of dementia
[2538] Step 1:
[2539] The device begins a casual conversation by asking, "Tell me how your day was today."
[2540] Step 2:
[2541] Users talk about everyday events.
[2542] Step 3:
[2543] The terminal uses a voice recognition system to convert the user's speech into text data.
[2544] Step 4:
[2545] The terminal transmits the acquired text data to the server.
[2546] Step 5:
[2547] The server analyzes the text data and evaluates the elderly person's language patterns.
[2548] Step 6:
[2549] The server uses an emotion engine to analyze the user's emotional state from their speech.
[2550] Step 7:
[2551] Based on the analysis results and emotional state, the server generates an alert such as, "Recently, there has been a decrease in vocabulary in conversations."
[2552] Step 8:
[2553] The server generates an alert and sends it to the device.
[2554] Step 9:
[2555] The device will notify the elderly person of the alert.
[2556] Step 10:
[2557] The server will also notify caregivers and family members of the alerts as needed.
[2558] These processing steps enable the remote care system of the present invention to monitor the user's living conditions with high accuracy and provide appropriate care and support. In addition, the emotion engine enables flexible responses according to the user's emotional state, improving the user's quality of life and reducing the burden on caregivers.
[2559] Example 2
[2560] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2561] For elderly people to maintain independent lifestyles, it is important to properly monitor their daily health and emotional states and provide timely advice and reminders. However, current systems have difficulty effectively analyzing the subject's voice and images and providing appropriate advice and reminders that take their emotional state into account. Furthermore, the lack of personalized support tailored to each individual's health condition and specific lifestyle circumstances makes it difficult for elderly people to maintain appropriate diets and medication intake.
[2562] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[2563] In this invention, the server includes an analysis means that uses a generative AI model to analyze voice data and image data and evaluate health conditions and living situations, an emotion engine that recognizes emotional states from voice data, a means that analyzes dietary content and medication status to generate advice and reminders necessary for the elderly, and a notification means that notifies the elderly of the generated advice and reminders by voice. This enables comprehensive analysis using the voices and images of the elderly, and makes it possible to provide appropriate, personalized advice and reminders.
[2564] "Speech recognition means" is a technology that collects speech and converts it into text data.
[2565] "Image recognition means" is a technology that collects images and sends the data to a server.
[2566] A "generative AI model" is an artificial intelligence technology that analyzes voice and image data to assess the health and living conditions of elderly people.
[2567] "Analysis means" refers to the process of analyzing and evaluating audio and image data using a generative AI model.
[2568] "Emotion engine" is a technology that recognizes the user's emotional state from voice data.
[2569] The "notification means" is a technology that notifies the elderly of the generated advice and reminders by voice.
[2570] A "reminder" is a notification that helps seniors remember to take certain actions, such as taking medicine.
[2571] "Text data" refers to character data converted from speech by speech recognition means.
[2572] The present invention is a remote care system that monitors the living conditions of elderly people and provides appropriate support, and is equipped with an emotion engine that analyzes voice data and image data of the elderly person and recognizes the emotional state of the user. Specific embodiments of the present invention will be described below.
[2573] System configuration
[2574] This system consists of a voice recognition means, an image recognition means, an analysis means, an emotion engine, and a notification means.
[2575] Voice recognition means
[2576] The device is equipped with a microphone that collects conversations with elderly people in real time. The collected voice data is converted into text data through a voice recognition system, allowing the user's speech to be captured as text information.
[2577] Image Recognition Method
[2578] The device is equipped with a camera that takes pictures of the elderly person eating and taking their medicine. The captured image data is sent to a server in real time, allowing for a visual record of the elderly person's behavior.
[2579] Analysis means
[2580] The server is equipped with a generative AI model with advanced analytical capabilities that analyzes the voice and image data sent from the device. This analysis makes it possible to evaluate the health and living conditions of elderly people in real time. The generative AI model performs highly accurate analysis of the data, using techniques such as deep learning and natural language processing.
[2581] Emotion Engine
[2582] The server's emotion engine has the ability to recognize the user's emotional state from voice data, allowing it to assess not just their physical state but also their psychological state. Emotion analysis is performed based on the tone of voice and the choice of words used.
[2583] Notification means
[2584] Based on the analysis results and emotional state, the server generates advice and reminders for the elderly, which are then sent to the device, which then notifies the elderly by voice.
[2585] Meal management function
[2586] Overview: The dietary management function monitors whether the elderly are eating a balanced diet and provides advice as needed.
[2587] The device asks the senior, "What did you have for breakfast today?"
[2588] The user responds, "I had bread and milk."
[2589] The device uses a voice recognition system to convert the user's speech into text data and a camera to capture the meal.
[2590] The server analyzes text and image data to evaluate the calorie and nutritional balance of the meal, and also analyzes the user's emotional state using an emotion engine.
[2591] The server generates advice such as, "Today's breakfast seems low in calories. We recommend adding more fruits and vegetables to your next meal."
[2592] The device will provide advice to the elderly via voice.
[2593] Example prompt sentence:
[2594] Prompts for older adults regarding dietary management functions:
[2595] Please explain the analysis process and advice generation steps when an elderly person answers "I had bread and milk" to the question "What did you have for breakfast today?"
[2596] Medication management function
[2597] Abstract: Medication management is a function that helps elderly people take their medications appropriately.
[2598] The server detects when it's time to take the medication and generates a reminder.
[2599] The device will remind the elderly person, "It's time to take your medicine. Are you ready?"
[2600] The user responds, "Yes, I'll drink it now."
[2601] The device uses a voice recognition system to convert responses into text data and a camera to record the patient taking the medicine.
[2602] The server analyzes the video data to confirm whether the user has taken the medicine, and also analyzes the user's emotional state using an emotion engine.
[2603] The server generates a confirmation message saying "Your medication was taken correctly" and sets the next reminder.
[2604] The device will then send a confirmation message to the elderly person.
[2605] Example prompt sentence:
[2606] Prompts regarding medication management functions for seniors:
[2607] Please explain the analysis process and reminder generation steps when an elderly person responds "Yes, I'll take it right away" to the reminder "It's time to take your medicine. Are you ready?"
[2608] Early detection of dementia
[2609] Overview: The early dementia detection function detects signs of dementia from the everyday conversations of elderly people.
[2610] The device begins a casual conversation by asking, "Tell me how your day was today."
[2611] Users talk about everyday events.
[2612] The terminal uses a voice recognition system to convert the conversation into text data and send it to the server.
[2613] The server analyzes the text data, assessing language patterns and emotional states, and also analyzes emotional fluctuations using an emotion engine.
[2614] The server generates an alert such as, "Recently, we have noticed a decline in vocabulary in conversations."
[2615] The device will notify the senior of the alert and, if necessary, their caregiver or family member.
[2616] Example prompt sentence:
[2617] Here are some prompts for early dementia detection:
[2618] Please explain the steps for analyzing and generating alerts after an elderly person talks about their daily life when asked, "Tell me how your day was today."
[2619] In this way, the remote care system of the present invention supports the lives of elderly people in a variety of ways, and by combining it with an emotion engine, it is possible to provide more accurate advice and appropriate care.
[2620] The flow of the identification process in the second embodiment will be described with reference to FIG.
[2621] Meal management function processing steps
[2622] Step 1:
[2623] Voice input and data collection (terminal)
[2624] The device uses a microphone to ask the elderly person, "What did you have for breakfast today?" The elderly person responds, "I had bread and milk." This voice input is collected.
[2625] Input: Voice input for the elderly
[2626] Output: Audio data collected by microphone
[2627] Step 2:
[2628] Speech recognition and data conversion (terminal)
[2629] The device uses a voice recognition system to convert voice data into text data, and also uses a camera to record the elderly person's mealtimes.
[2630] Input: Audio data, video data of meal
[2631] Output: Text data, image data
[2632] Step 3:
[2633] Data transmission (terminal)
[2634] The terminal encrypts the converted text data and the captured image data and transmits them to the server in real time.
[2635] Input: Text data, image data
[2636] Output: Data sent to the server
[2637] Step 4:
[2638] Data analysis (server)
[2639] The server uses the generative AI model to analyze the received text and image data and evaluate the calories and nutritional balance of the meal contents.
[2640] Input: Text data, image data
[2641] Output: Calorie and nutritional balance assessment results
[2642] Step 5:
[2643] Sentiment analysis (server)
[2644] The server's emotion engine recognizes the user's emotional state from the voice data.
[2645] Input: Text data
[2646] Output: User's emotional state
[2647] Step 6:
[2648] Advice Generation (Server)
[2649] The server generates advice based on the meal evaluation and emotional state, such as "We recommend adding more fruits and vegetables to your next meal."
[2650] Input: Calorie and nutritional balance assessment results, user's emotional state
[2651] Output: The generated advice
[2652] Step 7:
[2653] Notifications (device)
[2654] The device notifies the elderly person of the advice received from the server via voice.
[2655] Input: Advice from the server
[2656] Output: Voice notification to the elderly
[2657] Medication management function processing steps
[2658] Step 1:
[2659] Medication time detection (server)
[2660] The server detects the preset time for taking the medication.
[2661] Input: Pre-set medication time
[2662] Output: Start generating medication reminders
[2663] Step 2:
[2664] Reminder generation (server)
[2665] The server generates a reminder: "It's time to take your medicine. Are you ready?"
[2666] Input: Medication time detection
[2667] Output: The generated reminder
[2668] Step 3:
[2669] Reminder notification (device)
[2670] The device will notify the elderly of reminders via voice.
[2671] Input: Reminder from server
[2672] Output: Voice notification to the elderly
[2673] Step 4:
[2674] User response (terminal)
[2675] The user responds, "Yes, I'll drink it now." The device uses a voice recognition system to convert the response into text data.
[2676] Input: User's voice response
[2677] Output: Text data
[2678] Step 5:
[2679] Medication confirmation (terminal)
[2680] The device uses a camera to record the patient taking the medicine.
[2681] Input: User's medication actions
[2682] Output: Video data
[2683] Step 6:
[2684] Data transmission (terminal)
[2685] The converted text data and video data are sent to the server.
[2686] Input: Text data, video data
[2687] Output: Data sent to the server
[2688] Step 7:
[2689] Data analysis (server)
[2690] The server analyzes the video data to confirm whether the user has taken the medicine, and also analyzes the user's emotional state using an emotion engine.
[2691] Input: Text data, video data
[2692] Output: Medication confirmation result, user's emotional state
[2693] Step 8:
[2694] Confirmation message generation (server)
[2695] The server generates a confirmation message saying "Your medication was taken correctly" and sets the next reminder.
[2696] Input: Medication confirmation result, user's emotional state
[2697] Output: Generated confirmation message, next reminder
[2698] Step 9:
[2699] Confirmation notification (terminal)
[2700] The device will then provide a confirmation message to the elderly person via voice.
[2701] Input: Confirmation message from the server
[2702] Output: Voice notification to the elderly
[2703] Dementia early detection function processing steps
[2704] Step 1:
[2705] Starting a daily conversation (terminal)
[2706] The device speaks to the elderly person, asking, "Tell us how your day was today."
[2707] Input: Pre-configured question prompt
[2708] Output: Start a conversation with the elderly person
[2709] Step 2:
[2710] User response (terminal)
[2711] Users talk about everyday events, and the device converts the conversation into text data using a voice recognition system.
[2712] Input: User's daily conversation
[2713] Output: Text data
[2714] Step 3:
[2715] Data transmission (terminal)
[2716] The converted text data is sent to the server.
[2717] Input: Text data
[2718] Output: Data sent to the server
[2719] Step 4:
[2720] Data analysis (server)
[2721] The server analyzes the text data, assessing language patterns and emotional states, and also analyzes emotional fluctuations using an emotion engine.
[2722] Input: Text data
[2723] Output: Language pattern evaluation results, emotional state evaluation results
[2724] Step 5:
[2725] Alert Generation (Server)
[2726] The server generates an alert such as, "Recently, we have noticed a decline in vocabulary in conversations."
[2727] Input: Language pattern evaluation results, emotional state evaluation results
[2728] Output: The generated alert
[2729] Step 6:
[2730] Alert notification (terminal)
[2731] The device will notify the senior of the alert, and if necessary, a caregiver or family member.
[2732] Input: Alert from the server
[2733] Output: Notification to seniors and caregivers
[2734] (Application example 2)
[2735] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2736] It is difficult to monitor the health and emotional states of elderly people in real time and provide appropriate advice and reminders when they go about their daily lives, especially when they use self-driving vehicles. Furthermore, conventional remote care systems do not provide support tailored to the elderly's situation in self-driving vehicles, which may reduce the elderly's sense of safety and security. Therefore, there is a need for a remote care system that can monitor the health and emotional states of elderly people in real time and provide appropriate support when they use self-driving vehicles.
[2737] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[2738] In this invention, the server includes a voice recognition means for acquiring voice data of the elderly person, an image recognition means for acquiring image data of the elderly person, an analysis means using a generative AI model to analyze the voice data and image data, a notification means for notifying the elderly person of necessary advice and reminders based on the analysis results, a means for monitoring the elderly person's health and emotional state while in the autonomously driven vehicle, a means for transmitting the data collected by the above means to the server and analyzing it, a means for providing the elderly person with advice regarding their living conditions in the autonomously driven vehicle based on the analysis results, and a means for notifying the elderly person of the advice visually and audibly. This makes it possible to support elderly people in living safely and healthily when using autonomously driven vehicles.
[2739] The "voice recognition means" is a means for acquiring and analyzing voice data of the elderly person.
[2740] The "image recognition means" is a means for acquiring image data of an elderly person and analyzing it.
[2741] A "generative AI model" is an advanced artificial intelligence model that analyzes audio and image data and is used to assess the health and emotional state of an elderly person.
[2742] The "analysis means" is a means for assessing the health and emotional state of an elderly person using audio and image data.
[2743] The "notification means" is a means for visually and audibly notifying the elderly of necessary advice and reminders based on the analysis results.
[2744] An "automated vehicle" is a vehicle that is driven automatically and can be used by seniors.
[2745] "Health status" refers to the physical condition and physical function of an elderly person.
[2746] "Emotional state" refers to the psychological state and emotional movements of elderly people.
[2747] The "means for transmitting data to a server" is a means for transmitting voice data and image data collected from elderly people to a server.
[2748] "Visual and audio notification means" refers to means for conveying advice to the elderly person through a visual display and audio.
[2749] The present invention relates to a remote care system for elderly people to live safe and healthy lives when using autonomous vehicles. An embodiment of the present invention will be specifically described.
[2750] System configuration
[2751] Terminal configuration
[2752] 1. Voice recognition means:
[2753] The device is equipped with a microphone that is used to capture voice data from the elderly, which is then analyzed in real time.
[2754] 2. Image Recognition Methods:
[2755] The device is equipped with a camera that is used to capture image data of the elderly person and send it to a server. For example, the device captures the elderly person's facial expressions and body movements.
[2756] Server-side configuration
[2757] 1. Analysis method:
[2758] The server is equipped with a generative AI model that analyzes the voice and image data sent from the device, and this analysis evaluates the health and emotional state of the elderly person.
[2759] 2. Means of notification:
[2760] Based on the analysis results, the server generates necessary advice and reminders for the elderly, which are then sent to the device, which then notifies the elderly by voice or display.
[2761] Hardware and software used
[2762] Hardware:
[2763] Microphone: Used to collect voice data from the elderly.
[2764] Camera: Used to capture image data of the elderly.
[2765] Self-driving vehicles: Used to monitor the health and emotional state of elderly people in real time while they are in the vehicle.
[2766] software:
[2767] speech_recognition library: Used to achieve speech recognition.
[2768] OpenCV library: Used to realize image processing.
[2769] TensorFlow: Used to realize a generative AI model for elderly emotion engine analysis.
[2770] gTTS: Used to generate audio notifications.
[2771] requests: Used to realize data transmission to the server.
[2772] System operation example
[2773] 1. While the elderly person is riding in a self-driving vehicle:
[2774] Microphones and cameras capture the elderly person's voice and image data in real time, which is automatically sent to a server.
[2775] The server uses a generative AI model to analyze audio and image data to assess the health and emotional state of the elderly.
[2776] Based on the assessment results, necessary advice and reminders are generated and sent to the elderly via the device. For example, if a person does not eat enough breakfast, they will receive a voice message saying, "We recommend adding more fruits and vegetables to your next meal."
[2777] Prompt Sentence Examples
[2778] After the elderly person answers questions about health care in the self-driving car, please send the image and voice data to a server and generate a Python program that will provide appropriate advice and reminders.
[2779] As described above, the remote care system of the present invention provides multifaceted support to enable elderly people to use self-driving vehicles with peace of mind.
[2780] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[2781] Step 1: Acquire audio data
[2782] The device uses a microphone to capture voice data from the elderly. The input is the elderly's voice, which is collected in real time. The device stores the collected voice data either directly or temporarily.
[2783] Step 2: Acquiring image data
[2784] The device uses a camera to acquire image data (facial expressions and body movements) of the elderly. The input is a real-time image of the elderly acquired through the camera, which is captured at high resolution. The device stores the captured image data either directly or temporarily.
[2785] Step 3: Send data to the server
[2786] The terminal sends the collected voice and image data to the server. The input is the data acquired in steps 1 and 2 above, and this is sent to the server via the network. The output is the data itself that was sent to the server.
[2787] Step 4: Analyzing the audio data
[2788] The server analyzes the received voice data. The input is the voice data sent from the device, which is analyzed using a generative AI model. The content of the elderly person's speech and their emotional state are extracted from the voice data and output as text data.
[2789] Step 5: Analyzing the image data
[2790] The server analyzes the received image data. The input is image data sent from the device, which is analyzed using a generative AI model. The image data is used to evaluate the elderly person's facial expressions, body movements, health condition, etc., and is output as numerical data or evaluation results.
[2791] Step 6: Integrating the analysis results
[2792] The server integrates the results of the voice data analysis and the image data analysis to evaluate the elderly person's comprehensive health and emotional state. The input is the analysis results of steps 4 and 5, which are integrated to comprehensively grasp the elderly person's current condition. The output is the integrated evaluation result.
[2793] Step 7: Generate advice and reminders
[2794] The server generates necessary advice and reminders for the elderly based on the integrated evaluation results. The input is the integrated evaluation results, and appropriate messages are generated based on them. The output is text data of the generated advice and reminders.
[2795] Step 8: Implementing Notification
[2796] The device notifies the elderly of advice and reminders received from the server. The input is text data sent from the server, which is notified visually and audibly. Specifically, messages are displayed on the display and read aloud using gTTS. The output is the advice and reminders notified to the elderly.
[2797] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[2798] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2799] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[2800] T...
Claims
1. A remote care system for monitoring and supporting the living conditions of elderly people, a voice recognition means for acquiring voice data of the elderly person; image recognition means for acquiring image data of an elderly person; an analysis means using a generative AI model to analyze the voice data and image data; The system includes a notification means for providing necessary advice and reminders to the elderly based on the analysis results.
2. 2. The system according to claim 1, wherein said analysis means analyzes the dietary contents of the elderly person and evaluates the nutritional balance and calories.
3. 2. The system according to claim 1, wherein the analysis means analyzes the elderly person's medication status and generates a reminder to prevent the elderly person from forgetting to take their medication.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A