system
A robotic system for elderly individuals addresses health monitoring and dementia detection by integrating voice and physical data with AI, facilitating continuous and automatic health assessment and early medical intervention.
Patent Information
- Application Number
- JP2024140486
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-21
- Publication Date
- 2026-03-06
AI Technical Summary
Elderly individuals living alone face challenges in monitoring their health status and connecting with medical institutions, particularly in the early stages of dementia where subjective symptoms are minimal.
A robotic system that conducts preventive health conversations, assesses dementia risk, and integrates voice and physical data using AI to provide feedback and facilitate early medical consultation.
Enables continuous and automatic monitoring of health status, minimizing user burden and promoting early and appropriate medical intervention for elderly individuals.
Smart Images

Figure 2026037461000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In modern society, opportunities for elderly people to live alone are increasing, making health management and early detection of dementia important issues. Conventional systems make it difficult for elderly people to properly monitor their own health status and connect with medical institutions when necessary. In particular, early detection and appropriate care are difficult to achieve because there are few subjective symptoms in the early stages of dementia. [Means for solving the problem]
[0005] The present invention solves the above-mentioned problems by providing a robotic system that conducts preventive health conversations and assesses dementia in the home of elderly people living alone. The system includes a means for recognizing a user, a means for initiating a voice conversation with the user and converting the voice data into text data, a means for analyzing the converted text data and assessing the user's emotions and health status, a means for generating the next question based on the assessment results and presenting it to the user, a means for measuring the user's physical data and collecting the measurement results, a means for integrating the voice and measurement data and assessing dementia risk using AI, a means for providing feedback to the user based on the assessment results, and a means for saving the data and preparing for the next session. This system allows elderly people to monitor their health status on a daily basis and, if necessary, to consult with a medical institution early on.
[0006] "User recognition means" refers to a function that detects the user's face, voice, etc. and identifies them as a specific person.
[0007] "Voice data" refers to data that captures what the user has said to the robot as an acoustic signal.
[0008] "Text data" refers to data obtained by converting voice data into character information.
[0009] A "natural language processing engine" refers to software that analyzes text data and understands its content and emotions.
[0010] "Evaluation means" refers to a function for analyzing text data to infer the user's emotions and health condition.
[0011] "Question generation means" refers to the function of generating the next question to ask based on the analysis results.
[0012] "Physical data" refers to measurement data that indicates the user's physical health status, such as body temperature and blood pressure.
[0013] "Measurement means" refers to equipment and methods for acquiring the user's physical data.
[0014] "AI algorithm" refers to a calculation method that uses artificial intelligence to analyze data and assess dementia risk.
[0015] "Evaluation results" refers to information obtained as a result of analysis or evaluation.
[0016] "Feedback means" refers to a function that returns information or instructions to the user based on the evaluation results.
[0017] "Data Storage" refers to a method or system for storing data collected during a session.
[0018] "Session preparation means" refers to a function that prepares for the next session to proceed smoothly. [Brief explanation of the drawings]
[0019] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8]FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0020] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0021] First, the terms used in the following description will be explained.
[0022] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0023] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0024] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0025] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0027] [First embodiment]
[0028] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0029] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0030] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0031] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0032] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0034] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0035] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0036] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0037] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0038] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0039] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0040] This invention relates to a robotic system for monitoring the health status and diagnosing dementia in the home of elderly people living alone. The system is equipped with functions for user recognition, voice conversation, data analysis, and feedback.
[0041] First, when the device (robot) starts up, it recognizes the user using the built-in camera and voice recognition function. At this point, the robot greets the user by saying, "Good morning, [user name]. How are you today?"
[0042] Next, the user responds to the robot by saying something like, "I feel a bit heavy-headed today." This voice data is converted into text data by a voice recognition module in the device.
[0043] The converted text data is sent to a server and analyzed by a natural language processing engine (NLP). This analysis extracts information to evaluate the user's emotions and health status. For example, an expression such as "my head feels heavy" can be used to infer a decline in health.
[0044] Next, the server generates the next question based on the analysis results and sends it to the device. For example, the question generated might be, "Have you been sleeping well lately?" The device then presents this question to the user.
[0045] At the same time, the device prompts the user to measure their temperature and blood pressure. The user measures their temperature and blood pressure using a thermometer or blood pressure monitor at home and reports the results to the robot. For example, "My temperature is 36.5 degrees and my blood pressure is 120 / 80."
[0046] Health data and text data are integrated on a server, and an AI algorithm is used to assess dementia risk, based on abnormalities in language usage patterns and physical data.
[0047] For example, if the user has become increasingly forgetful in recent conversations or if there are abnormalities in the user's physical data, the server will generate an alert message and send it to the device. The device will then provide feedback such as, "After looking at the results, there are some concerns about your recent physical condition and conversations. We recommend that you consult a doctor."
[0048] All conversation and health data is stored in a database on the server, which allows for preparation for the next session and regular monitoring of the user's condition.
[0049] As a concrete example, the flow of one day is shown below.
[0050] 1. The device starts up, recognizes the user, and greets them with, "Good morning, Tanaka-san. How are you today?"
[0051] 2. The user replies, "Good morning, I'm feeling a bit down today."
[0052] 3. The device converts the voice data into text and sends it to the server.
[0053] 4. The server analyzes the text data, generates a question such as "Have you often felt heavy-headed lately?" and sends it to the device.
[0054] 5. The device asks questions and prompts the user to measure their health data. It prompts the user to "measure their temperature and blood pressure."
[0055] 6. The user reports the measurement results and replies, "Temperature is 36.5°C, blood pressure is 120 / 80."
[0056] 7. The device records this as text data and sends it to the server.
[0057] 8. The server analyzes all data and evaluates the risk of dementia. If the risk is high, an alert is sent to the device.
[0058] 9. The device will provide feedback such as, "After looking at the results, there are some concerns about your recent health condition and conversations. We recommend that you consult a doctor."
[0059] This system will enable daily monitoring of the health status of elderly people and enable early contact with medical institutions if necessary.
[0060] The processing flow will be explained below.
[0061] Step 1:
[0062] The device boots up and uses the built-in camera and microphone to scan the surroundings and identify the user. Once the user is identified, the device greets them with, "Good morning, [username]. How are you today?"
[0063] Step 2:
[0064] The user responds, "Good morning, I feel a bit heavy-headed today." The device collects the voice data and converts it into text data using a voice recognition module.
[0065] Step 3:
[0066] The converted text data is sent to a server, which then analyzes it using a natural language processing engine (NLP).The results of this analysis are used to evaluate the user's emotions and health status.
[0067] Step 4:
[0068] The server generates the next question based on the analysis results, for example, "Have you been sleeping well recently?", and sends this question to the device.
[0069] Step 5:
[0070] The device asks the user, "Have you been sleeping well lately?" While waiting for the user's response, the device instructs the user to measure their temperature and blood pressure. It adds, "Please measure your temperature and blood pressure."
[0071] Step 6:
[0072] The user uses a thermometer or blood pressure monitor to report the measurement results, such as "My body temperature is 36.5 degrees and my blood pressure is 120 / 80." The device's voice recognition module converts this voice back into text data, which is then sent to the server.
[0073] Step 7:
[0074] The server combines the health check data and text data and uses an AI algorithm to assess dementia risk, for example, based on information such as forgetfulness in recent conversations or abnormalities in physical data.
[0075] Step 8:
[0076] If the risk is high based on the assessment results, the server generates an alert message and sends it to the device, such as "There are some concerns about your health. We recommend that you consult a doctor."
[0077] Step 9:
[0078] The device will then provide the user with an alert message, saying, "After looking at the results, we've noticed some concerns about your recent health and conversations. We recommend that you consult a doctor."
[0079] Step 10:
[0080] The server stores all conversation and health data in a database, and periodically prepares data to monitor the user's condition in preparation for the next session.
[0081] Example 1
[0082] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0083] The present invention relates to a robotic system for monitoring the health status and dementia risk of elderly people living alone. Many current health monitoring systems require periodic manual input, which can be a burden for users. Furthermore, conventional systems do not perform detailed analysis of the user's emotions and health status, which can delay appropriate medical treatment when needed. Therefore, the present invention aims to provide a system that can continuously and automatically monitor the user's health status and promptly prompt appropriate medical treatment.
[0084] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0085] In this invention, the server includes means for recognizing a user, means for initiating a voice conversation with the user and converting the voice data into text data, means for analyzing the converted text data and evaluating the user's emotions and health status, means for generating the next question based on the evaluation results and presenting it to the user, means for collecting the user's physical data, means for integrating the voice and physical data and evaluating dementia risk using AI, means for providing feedback to the user based on the evaluation results, means for saving the data and preparing for the next session, and means for image recognition for identifying the user's face and voice recognition for recognizing the user's voice. This makes it possible to continuously and automatically monitor the user's health status in detail while minimizing the burden on the user, and to promote early and appropriate medical treatment.
[0086] "Means for recognizing the user" refers to technology that enables the robot to identify the user's face or specific features using cameras or sensors.
[0087] The "means for initiating a voice conversation and converting voice data into text data" is a voice recognition module for recognizing the voice spoken by the user and converting the voice into text information.
[0088] The "means for analyzing text data and assessing the user's emotions and health condition" refers to an algorithm or program that uses a natural language processing engine to analyze the converted text data and analyze the user's emotions and health condition.
[0089] The "means for generating the next question and presenting it to the user" refers to a device or software that generates a question to further evaluate the user's health condition based on the analysis results and presents it to the user via voice or screen display.
[0090] "Means for collecting user's physical data" refers to a method or device for collecting data on the user's physical condition using measuring instruments such as a thermometer or blood pressure monitor.
[0091] "Means for integrating voice and physical data and assessing dementia risk using AI" refers to a system that integrates and analyzes collected voice data and physical data, and assesses dementia risk using an AI algorithm.
[0092] The "means for providing feedback to the user based on the evaluation results" refers to a method or device that communicates the results of the analysis and evaluation to the user and provides a notice or warning that encourages the user to consult a medical institution if necessary.
[0093] The "means for storing the data and preparing for the next session" refers to a mechanism for recording and storing all data, including voice data and health data, to facilitate the next health monitoring session.
[0094] "Image recognition means" is a technology for analyzing facial features from images captured by a camera and identifying users.
[0095] "Speech recognition means" is a technology for capturing a user's voice and converting it into text data.
[0096] This invention relates to a robotic system for monitoring the health status and dementia risk of elderly people living alone. The system is equipped with functions for user recognition, voice conversation, data analysis, and feedback.
[0097] First, when the device (robot) starts up, it recognizes the user using the built-in camera and voice recognition function. The device greets the user by saying, "Good morning, [user name]. How are you today?" At this time, it uses the camera to identify the face and activates the voice recognition module. As a specific example, the device uses the built-in camera and voice recognition module to identify the user's face and voice.
[0098] Next, if the user responds, "I feel a bit heavy-headed today," this voice data is converted into text data by a voice recognition module. When the terminal converts voice data into text data, it uses a voice recognition module. For example, if the user says, "I'm feeling a bit unwell today," this is recorded as text data.
[0099] The converted text data is sent to a server and analyzed by a natural language processing engine (NLP). This analysis evaluates the user's emotions and health status. For example, the expression "my head feels heavy" can be used to infer a decline in health. As a specific example, the NLP engine analyzes the text "my head feels heavy" and determines that this is a sign of deteriorating health.
[0100] Next, the server generates the next question based on the analysis result and sends it to the terminal. For example, the question "Have you been sleeping well recently?" is generated. The terminal presents this question to the user. An example of a specific prompt sentence in this case is "Have you been sleeping well recently?"
[0101] At the same time, the device instructs the user to measure their temperature and blood pressure. The user uses a thermometer and blood pressure monitor to measure their temperature and report the results to the robot. For example, they might say, "My temperature is 36.5 degrees and my blood pressure is 120 / 80." Once the user reports the measurement results, the voice recognition module converts them back into text data and sends it to the server.
[0102] All data is integrated on a server, and an AI algorithm evaluates the risk of dementia. Risk is assessed based on abnormalities in language usage patterns and physical data. For example, if a person has become increasingly forgetful in recent conversations, this could be assessed as a risk of dementia.
[0103] If the risk is high, the server generates an alert message and sends it to the device, which then provides feedback such as, "After looking at the results, there are some concerns about your recent health condition and conversations. We recommend that you consult a doctor."
[0104] All conversation and health data is stored in a database on the server and ready for the next session, allowing for continuous, automatic, and detailed monitoring of the user's health status, facilitating early and appropriate medical treatment.
[0105] To give an example, here's what a typical day might look like:
[0106] 1. The device starts up, recognizes the user, and greets them with, "Good morning, Tanaka-san. How are you today?"
[0107] 2. The user replies, "Good morning, I'm feeling a bit down today."
[0108] 3. The device converts the voice data into text and sends it to the server.
[0109] 4. The server analyzes the text data, generates a question such as "Have you often felt heavy-headed lately?" and sends it to the device.
[0110] 5. The device asks questions and prompts the user to "take your temperature and blood pressure."
[0111] 6. The user reports the measurement results and replies, "Temperature is 36.5°C, blood pressure is 120 / 80."
[0112] 7. The device records this as text data and sends it to the server.
[0113] 8. The server analyzes all data and evaluates the risk of dementia. If the risk is high, an alert is sent to the device.
[0114] 9. The device will provide feedback such as, "After looking at the results, there are some concerns about your recent health condition and conversations. We recommend that you consult a doctor."
[0115] This system allows for daily monitoring of the health status of elderly people and allows for early linkage with medical institutions.
[0116] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0117] Step 1:
[0118] The device (robot) starts up. Once started, it uses the built-in camera and voice recognition function to recognize the user. Specifically, the camera captures the user's face and identifies the user using a facial recognition algorithm. Once the user is identified, the device greets them with "Good morning, [user name]. How are you today?"
[0119] Input: Start signal, camera image
[0120] Output: User recognition result, greeting message
[0121] Step 2:
[0122] The user responds by saying, "I feel a bit heavy-headed today." This voice data is captured by the device's voice recognition module and converted into text data. Specifically, voice input is captured by a microphone and then converted into text by the voice recognition module.
[0123] Input: User voice
[0124] Output: Text data
[0125] Step 3:
[0126] The device sends the converted text data to the server, where it packages the text data into a standard format such as JSON and securely transmits it to the server through a network module.
[0127] Input: Text data
[0128] Output: Server sent data
[0129] Step 4:
[0130] The server receives the text data and analyzes it using a natural language processing engine (NLP). This analysis evaluates the user's emotions and health status. Specifically, the NLP engine analyzes the text data and extracts keywords and emotions. At this stage, a decline in the user's health is inferred from the user's statement that "my head feels heavy."
[0131] Input: Text data
[0132] Output: Analysis results (emotion and health evaluation)
[0133] Step 5:
[0134] The server generates the next question based on the analysis results and sends it to the device. For example, a question might be generated such as, "Have you been sleeping well lately?" The server then creates an appropriate question based on the analysis results and sends it to the device as a data packet.
[0135] Input: Analysis results
[0136] Output: Next question data
[0137] Step 6:
[0138] The device then presents the generated questions to the user and instructs them to measure their health data. Specifically, it uses a voice synthesis module to prompt the user to "measure their temperature and blood pressure."
[0139] Input: Next question data
[0140] Output: Voice prompts and measurement instructions
[0141] Step 7:
[0142] The user measures their temperature and blood pressure using a thermometer and reports the results to the robot, for example, "My temperature is 36.5 degrees and my blood pressure is 120 / 80." This triggers the voice recognition module to convert the voice data into text.
[0143] Input: Measurement data (body temperature, blood pressure)
[0144] Output: Text data
[0145] Step 8:
[0146] The device sends the recorded health data as text data to the server.
[0147] Input: Text data
[0148] Output: Server sent data
[0149] Step 9:
[0150] The server integrates all the data and uses an AI algorithm to assess the risk of dementia. The AI algorithm analyzes language usage patterns and abnormalities in physical data to assess risk. For example, if the person has become increasingly forgetful in recent conversations, this could be assessed as a risk of dementia.
[0151] Input: Integrated data (voice and health)
[0152] Output: Risk assessment results
[0153] Step 10:
[0154] The server generates an alert message based on the risk assessment results and sends it to the device, which then provides feedback such as, "After looking at the results, there are some concerns about your recent health condition and conversations. We recommend that you consult a doctor."
[0155] Input: Risk assessment results
[0156] Output: Feedback message
[0157] Step 11:
[0158] The server stores all conversation and health data in a database and prepares for the next session, allowing for continuous monitoring of the user's health.
[0159] Input: Voice data, health data
[0160] Output: Save data, prepare for next session
[0161] (Application example 1)
[0162] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0163] It is extremely important to monitor the health status of elderly people living alone at home or in facilities and to detect dementia risk early. However, current systems have difficulty comprehensively and continuously monitoring users' health status and emotions, making early detection and prompt feedback difficult. For this reason, there is a need for a system that provides more comprehensive and accurate monitoring and feedback through devices or smartphone applications installed in facilities that provide healthcare services for the elderly.
[0164] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0165] In this invention, the server includes a means for recognizing the user, a means for initiating a voice conversation with the user and converting the voice data into text data, a means for analyzing the converted text data and evaluating the user's emotions and health status, a means for generating the next question based on the evaluation results and presenting it to the user, and a means for measuring the user's physical data and collecting the measurement results. This enables a system including a means for integrating the voice and measurement data and assessing dementia risk using AI, a means for providing feedback to the user based on the evaluation results, and a means for storing the data and preparing for the next session. This enables more comprehensive and accurate health status monitoring and rapid feedback through terminals and smartphone applications installed in facilities providing services for the elderly.
[0166] "Means for recognizing the user" refers to technology that uses a camera and voice recognition technology to verify the user's personal information.
[0167] The "means for converting voice data into text data" refers to a technique that uses a voice recognition module to convert recorded voice into a corresponding text format.
[0168] "Means for analyzing converted text data and assessing the user's emotions and health status" refers to a technology that utilizes a natural language processing engine and a generative AI model to assess the user's emotions and health status from text data.
[0169] "Means for generating the next question and presenting it to the user" refers to a technology that generates a new question based on the analysis results and presents it to the user in voice or text.
[0170] The "means for collecting user's physical data" refers to a technology for collecting user's physical data such as body temperature and blood pressure using measuring devices such as a thermometer and a blood pressure monitor.
[0171] "Method for integrating voice and measurement data to assess dementia risk using AI" refers to a technology that integrates collected voice data and physical data and uses an AI algorithm to assess dementia risk.
[0172] "Means of providing feedback to users" refers to technology that provides the results of AI evaluation to users via voice or text.
[0173] The "means for saving the data and preparing for the next session" refers to a technique for saving all collected data in a database and using it as the basis for future sessions.
[0174] "Means for conducting voice conversations with the user and collecting health data using a terminal installed in a facility that provides services for the elderly" refers to technology for conducting voice conversations and collecting health data using a device installed within the facility.
[0175] "Means for interacting with the user and managing health data through a smartphone application" refers to technology that uses a smartphone application to exchange information with the user and collect and manage health data.
[0176] "Means of sending data to a cloud environment and analyzing the user's emotions and health status using a natural language processing engine and generative AI model" refers to a technology that sends collected data to the cloud via the internet and analyzes it using a natural language processing engine and generative AI model.
[0177] "Means for providing alerts regarding health status on facility devices and smartphone applications based on evaluation results" refers to technology that provides alerts to users on facility devices and smartphone applications based on the results of AI evaluation.
[0178] This invention is a system for monitoring a user's health status and assessing their dementia risk. The system consists of a terminal installed in a facility, a smartphone application, a cloud server, and multiple software modules for linking these components.
[0179] System Configuration
[0180] 1. Terminal installation location
[0181] In this invention, a robot terminal is installed in a facility that provides services for the elderly. The terminal is equipped with a camera and voice recognition function, and can interact with users. The terminal recognizes the user's voice and facial expressions and identifies personal information.
[0182] 2. Smartphone Applications
[0183] Users can also receive the same services outside the facility using a smartphone application. The application has functions for voice conversations with users, recording health data, and providing consultations. The app synchronizes with the robot terminal and transmits data to a cloud server.
[0184] 3. Cloud environment
[0185] Data analysis and storage are performed in a cloud environment: collected voice and body data are sent to the cloud and processed by a natural language processing engine (e.g., SpaCy or Google® Cloud Natural Language API) and a generative AI model.
[0186] Hardware and Software Details
[0187] Hardware
[0188] Camera: Built into the robot terminal and used to recognize the user's face.
[0189] Microphone: Used to collect voice conversation data.
[0190] Thermometers, blood pressure monitors: Health monitoring devices for measuring users' temperature and blood pressure. These are installed within the facility.
[0191] software
[0192] Speech recognition module: Technology for converting voice data into text data (e.g., Google Cloud Speech-to-Text).
[0193] Natural language processing engines and generative AI models: Analyze users' emotions and health status, and assess dementia risk (e.g., SpaCy, Google Cloud Natural Language API).
[0194] Cloud storage: Infrastructure for storing and managing data.
[0195] Detailed explanation of the process
[0196] 1. User Awareness
[0197] The device uses a camera and voice recognition technology to identify the user. When the user enters the facility, the device greets them with, "Good morning, how are you feeling today?"
[0198] 2. Audio data collection and analysis
[0199] When the user responds, the voice data is converted into text data using a voice recognition module. This data is sent to a cloud server and analyzed using a natural language processing engine. Based on the analysis results, the user's health status and emotions are evaluated.
[0200] 3. Next question generation and presentation
[0201] The next question is generated from the text data by the generative AI model and presented to the user via their device or smartphone app, such as "Have you been sleeping well lately?"
[0202] 4. Collection of health data
[0203] Physical data such as users' temperature and blood pressure will be collected using measuring equipment within the facility, and similar data can also be collected via a smartphone app.
[0204] 5. Data Storage and Feedback
[0205] The collected voice and physical data is stored in cloud storage. Based on the analysis results, feedback is provided to the user. For example, the system may provide feedback such as, "After looking at the results, there are some concerns about your recent physical condition and conversations. Please consult a doctor."
[0206] Specific examples
[0207] 1. Usage scenarios in facilities
[0208] When an elderly person visits the facility, the robot terminal recognizes them and asks, "Good morning, how are you feeling today?" If the elderly person replies, "I have a slight headache," the voice is converted into text data and sent to the cloud.
[0209] 2. Smartphone usage scenarios
[0210] The same elderly person opens the smartphone app at home and is asked the same questions about their health. The collected data is stored in the cloud and used the next time they visit the facility. Based on this data, a generative AI model generates appropriate questions and performs a risk assessment.
[0211] 3. Examples of prompts
[0212] Please analyze the following sentences and assess the user's health and emotions.
[0213] "My head feels a little heavy today."
[0214] This system can continuously monitor the user's health status wherever they are and provide quick feedback, contributing to the early detection of dementia risk.
[0215] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0216] Step 1:
[0217] When a user arrives at a facility, the device uses a camera and voice recognition technology to recognize the user. The camera captures the user's face, and the voice recognition module identifies the user's personal information from the voice data. Specifically, the camera captures the user's image and receives a greeting as voice data. The device outputs a voice message saying, "Good morning, how are you feeling today?"
[0218] Step 2:
[0219] The user responds to the terminal by voice. The terminal captures the voice data and converts it into text data using a voice recognition module. The user's voice input is "I feel a bit heavy-headed today," which is converted into text data. This text data is sent to the server.
[0220] Step 3:
[0221] The server analyzes the received text data and uses a natural language processing engine (e.g., SpaCy or Google Cloud Natural Language API) to extract data from the text data to evaluate the user's emotions and health status. The input is the text data "I feel a bit heavy-headed today," and the output is "The user is feeling unwell."
[0222] Step 4:
[0223] The server uses the generative AI model based on the analysis results to generate the next question. For example, the question "Have you been sleeping well lately?" is generated. This generated question data is sent to the device. Specifically, the text data of the generated question is sent to the device.
[0224] Step 5:
[0225] The device presents the generated question to the user by voice or text. The user then responds to the device. This response data is also converted into text data using a voice recognition module and sent to the server. The specific operation is that the device outputs the question "Have you been sleeping well lately?" by voice.
[0226] Step 6:
[0227] The terminal instructs the user to measure their physical data using a thermometer or blood pressure monitor. The user measures their body temperature and blood pressure using health measurement equipment installed in the facility and reports the results to the terminal. For example, the user reports, "My body temperature is 36.5 degrees and my blood pressure is 120 / 80." This data is converted into text data and sent to the server.
[0228] Step 7:
[0229] The server integrates all the received data and uses an AI algorithm to assess the risk of dementia. The voice data and physical data are analyzed as a single data set to assess the user's risk. The inputs are voice-text data such as "Head feels heavy," body temperature of "36.5 degrees," and blood pressure of "120 / 80," and the output is the "dementia risk assessment result." The specific operation is for the AI algorithm to analyze the data and perform a risk assessment.
[0230] Step 8:
[0231] The server provides feedback to the device and smartphone app based on the evaluation results. For example, it generates an alert message saying, "After looking at the results, there are some concerns about your recent health condition and conversations. Please consult a doctor." and sends it to the device and smartphone app. The user receives the alert message.
[0232] Step 9:
[0233] All data is stored in cloud storage and will be available for the next session. Along with regular data monitoring, it will be possible to continuously track the user's health status. Specifically, the collected voice and physical data will be stored in the cloud and used for future analysis.
[0234] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0235] This invention relates to a robotic system for monitoring the health status and diagnosing dementia in the home of elderly people living alone. The system is equipped with functions for user recognition, voice conversation, data analysis, emotion engine, and feedback.
[0236] First, when the device (robot) starts up, it recognizes the user using the built-in camera and voice recognition function. At this point, the robot greets the user by saying, "Good morning, [user name]. How are you today?"
[0237] The user then responds to the robot by saying something like, "I feel a bit heavy-headed today." This voice data is converted into text data by the device's voice recognition module. At the same time, the device acquires the user's voice and facial expression data and analyzes them using an emotion engine.
[0238] The converted text data and the evaluation data from the emotion engine are sent to a server and analyzed by a natural language processing engine (NLP). This analysis extracts information to evaluate the user's emotions and health status. For example, it can infer a decline in health status from expressions such as "my head feels heavy" and recognize emotions such as "I'm tired."
[0239] The server then generates the next question based on the analysis results and sends it to the device. For example, the question is "Have you been sleeping well lately?", taking into account emotional changes in the user's facial expressions. This allows for more personalized interactions.
[0240] The device will ask the user this question. At the same time, the device will ask the user to take their temperature and blood pressure. It will add, "Please take your temperature and blood pressure."
[0241] The user measures their temperature using a thermometer or blood pressure monitor at home and reports the results to the robot, for example, "My temperature is 36.5 degrees and my blood pressure is 120 / 80."
[0242] Health data, text data, and emotional data are integrated on a server, and an AI algorithm is used to assess dementia risk. The AI assesses risk based on abnormalities in language usage patterns and physical data, while also taking emotional data into account.
[0243] For example, if the user has become increasingly forgetful in recent conversations or has frequently shown emotionally unstable facial expressions, the server will generate an alert message and send it to the device, which will then provide feedback such as, "After looking at the results, there are some concerns about your recent physical condition and emotions. We recommend that you consult a doctor."
[0244] All conversation, health and emotion data is stored in a database on the server, allowing for preparation for the next session and regular monitoring of the user's condition.
[0245] As a concrete example, the flow of one day is shown below.
[0246] 1. The device starts up, recognizes the user, and greets them with, "Good morning, Tanaka-san. How are you today?"
[0247] 2. The user replies, "Good morning, I'm feeling a bit down today."
[0248] 3. The device collects voice and facial expression data, converts the voice into text, and analyzes it using an emotion engine.
[0249] 4. The text data and emotional evaluation data sent to the server are analyzed, and the question "Have you often felt heavy-headed lately?" is generated and sent to the device.
[0250] 5. The device asks questions and prompts the user to measure their health data. It prompts the user to "measure their temperature and blood pressure."
[0251] 6. The user reports the measurement results and replies, "Temperature is 36.5°C, blood pressure is 120 / 80."
[0252] 7. The device records this as text data and emotion data and sends it to the server.
[0253] 8. The server analyzes all data and evaluates the risk of dementia. If the risk is high, an alert is sent to the device.
[0254] 9. The device will provide feedback such as, "After looking at the results, there are some concerns about your recent physical condition and emotions. We recommend that you consult a doctor."
[0255] This system allows for daily monitoring of the health and emotional state of the elderly and allows for early consultation with medical institutions when necessary, thereby improving the peace of mind and quality of life for the elderly and their families.
[0256] The processing flow will be explained below.
[0257] Step 1:
[0258] The device boots up and uses the built-in camera and microphone to scan the surroundings and identify the user. Once the user is identified, the device greets them with, "Good morning, [username]. How are you today?"
[0259] Step 2:
[0260] The user responds, "Good morning, I feel a bit heavy-headed today." The device collects the voice data and converts it into text data using a voice recognition module.
[0261] Step 3:
[0262] The device collects facial expression data along with the captured voice data and analyzes it with an emotion engine, where the user's emotional state is identified as "fatigue."
[0263] Step 4:
[0264] The converted text data and the evaluation data generated by the emotion engine are sent to the server, which then analyzes the received text data using a natural language processing engine (NLP).The analysis results are used to evaluate the user's emotions and health status.
[0265] Step 5:
[0266] The server generates the next question based on the analysis results, for example, "Have you been sleeping well recently?", and sends this question to the device.
[0267] Step 6:
[0268] The device asks the user, "Have you been sleeping well lately?" When asking the question, it also takes into account changes in the user's facial expressions based on their emotions. At the same time, the device instructs the user to measure their temperature and blood pressure. It adds, "Please measure your temperature and blood pressure."
[0269] Step 7:
[0270] The user uses a thermometer or blood pressure monitor to report the measurement results, such as "My body temperature is 36.5 degrees and my blood pressure is 120 / 80." The device's voice recognition module converts this voice back into text data, which is then sent to the server.
[0271] Step 8:
[0272] The server combines the health check data, text data, and emotional data, and uses an AI algorithm to assess dementia risk. For example, if a person has recently shown forgetfulness and emotional instability in their facial expressions, it will determine that they are at high risk.
[0273] Step 9:
[0274] If the risk is high based on the assessment results, the server generates an alert message and sends it to the device, such as "There are some concerns about your health. We recommend that you consult a doctor."
[0275] Step 10:
[0276] The device will then provide the user with an alert message, saying, "After looking at your results, there are some concerns about your recent physical condition and emotions. We suggest you consult a doctor."
[0277] Step 11:
[0278] The server stores all conversation data, health data, and emotion data in a database. It periodically prepares data to monitor the user's condition in preparation for the next session.
[0279] Example 2
[0280] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0281] In modern society, there is a need to routinely monitor the health status and dementia risk of elderly people living alone and detect abnormalities early. However, the means to do so are limited, and there are no systems that use interactive systems to closely observe the health and emotional state of elderly people and provide appropriate feedback. Therefore, there is a need for an efficient and effective method of managing the health of elderly people.
[0282] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for recognizing the user, means for initiating a voice conversation with the user and converting the voice data into text data, means for analyzing the converted text data and evaluating the user's emotions and health status, means for generating the next question based on the evaluation results and presenting it to the user, means for measuring the user's physical data and collecting the measurement results, means for integrating the voice and measurement data and evaluating dementia risk using AI, means for providing feedback to the user based on the evaluation results, and means for saving the data and preparing for the next session. This makes it possible to monitor the health and emotional status of elderly people on a daily basis and to cooperate with medical institutions early if necessary.
[0283] "Means for recognizing a user" refers to the system's ability to identify and recognize the individual user using a camera or voice recognition technology.
[0284] The "means for starting a voice conversation and converting voice data into text data" is a technology for automatically converting voice data acquired through a voice dialogue with a user into text format.
[0285] The "means for analyzing the converted text data and evaluating the user's emotions and health state" is a technology that processes the text data, analyzes its contents, and estimates the user's current emotions and health state.
[0286] The "means for generating the next question and presenting it to the user" is a function for generating and presenting an appropriate next question to the user based on the analysis results.
[0287] "Means for measuring the user's physical data and collecting the measurement results" refers to technology that allows the user to measure physical data such as body temperature and blood pressure and for the system to collect the results.
[0288] "Means for integrating voice and measurement data and assessing dementia risk using AI" refers to a technology that integrates collected voice data and physical data and uses artificial intelligence technology to assess a user's dementia risk.
[0289] The "means for providing feedback to the user based on the evaluation results" is a function for providing the results of the system's evaluation to the user in the form of information or advice in an appropriate format.
[0290] "Means for storing data and preparing for the next session" refers to technology for recording collected and analyzed data and preparing it for use in the next interaction or monitoring.
[0291] The present invention relates to a robotic system for monitoring the health status and dementia risk of elderly people living alone and for detecting abnormalities at an early stage. This system mainly includes a means for recognizing a user, a means for performing voice conversation and data analysis, an emotion analysis engine, a feedback means, and a storage means.
[0292] When the device (robot) starts up, it uses a camera and voice recognition functions to recognize the user. Specifically, it takes a picture of the user's face with the camera and recognizes the user's voice using voice recognition technology. At this time, the device says, "Good morning, [user name]. How are you today?"
[0293] When the user responds, the voice data is converted into text data using the device's voice recognition module. This converted text data and facial expression data captured by the camera are analyzed by the emotion engine. For example, by analyzing a statement such as "I feel a bit heavy-headed today" and the user's facial expression at the time, the user's emotions and health condition can be evaluated.
[0294] The device then sends this data to a server, which incorporates a natural language processing engine (NLP) and performs detailed analysis of the received text data and emotional evaluation data. For example, an expression such as "My head feels heavy" is interpreted as a sign of the user's stress or poor health. Emotional fluctuations are also extracted from facial expression data.
[0295] Based on the analysis results, the server generates the next question. For example, "Have you been sleeping well recently?" and sends it to the device. The system then instructs the user to measure their temperature and blood pressure. The system prompts the user to "measure their temperature and blood pressure."
[0296] The user measures their temperature using a thermometer or blood pressure monitor at home and reports the results to the robot. For example, they might report, "My temperature is 36.5 degrees and my blood pressure is 120 / 80." This health data is collected as text data and sent back to the server.
[0297] The server uses an AI algorithm to comprehensively analyze this data. It evaluates the user's dementia risk by combining voice, emotional, and physical data. For example, if the user has recently shown an increase in forgetfulness and emotional instability in their facial expressions, the server will assess the risk as high and generate an alert message.
[0298] Finally, the device provides the generated feedback to the user, for example, saying, "After looking at the results, there are some concerns about your recent physical condition and emotions. We recommend that you consult a doctor." This makes it possible to monitor the health status of elderly people on a daily basis and to consult with medical institutions early if necessary.
[0299] All conversation, health, and emotion data is stored on the server and prepared for the next session. Through regular monitoring, the health status of the elderly can be continuously managed.
[0300] (Example of a prompt)
[0301] Here are some example prompts for a generative AI model:
[0302] "Design a robotic system to check the health status of elderly people daily. This system should have the ability to recognize the user's face and voice, analyze their health status and emotions, and provide feedback as needed. It should also send the measured health data and analyzed emotional data to a server, enabling regular monitoring."
[0303] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0304] (Program processing steps)
[0305] Step 1:
[0306] The device starts up.
[0307] Input: System power supply
[0308] Output: The terminal is ready
[0309] At startup, the built-in camera and voice recognition function are enabled to prepare for user recognition.
[0310] Step 2:
[0311] The device uses its built-in camera and voice recognition capabilities to identify the user.
[0312] Input: User's face, voice
[0313] Output: Recognized user information
[0314] The camera captures the user's face and uses voice recognition technology to identify the individual from their voice.
[0315] The device greets you with, "Good morning, [username]. How are you today?"
[0316] Step 3:
[0317] The user responds, "I'm feeling a bit heavy-headed today."
[0318] Input: User's voice response
[0319] Output: User's voice data
[0320] The terminal records the user's response.
[0321] Step 4:
[0322] The terminal uses a voice recognition module to convert the user's voice data into text data.
[0323] Input: User's voice data
[0324] Output: Converted text data
[0325] The voice recognition module is activated to convert the voice data into text format.
[0326] Step 5:
[0327] The device collects the user's facial expression data captured by the camera and analyzes it using an emotion engine.
[0328] Input: User's facial expression data
[0329] Output: User's emotion rating data
[0330] A camera captures the user's facial expressions and uses an emotion analysis algorithm to assess their emotional state.
[0331] Step 6:
[0332] The terminal transmits the converted text data and emotion evaluation data to the server.
[0333] Input: Text data, emotion rating data
[0334] Output: Data sent to the server
[0335] The terminal uploads the data to the server via the network.
[0336] Step 7:
[0337] The server uses a natural language processing engine (NLP) to analyze the text data and assess the user's health status.
[0338] Input: Text data, emotion rating data
[0339] Output: User's health status rating
[0340] Using NLP technology, the system analyzes the converted text and, for example, infers a decline in the user's health from an expression like "my head feels heavy." It also takes into account emotional data.
[0341] Step 8:
[0342] The server generates the next question based on the analysis results.
[0343] Input: User's health assessment
[0344] Output: Next question (e.g., "Have you been sleeping well lately?")
[0345] Based on the analysis results, appropriate follow-up questions are automatically generated.
[0346] Step 9:
[0347] The server generates the next question and sends it to the terminal.
[0348] Input: Next question
[0349] Output: Questions sent to the terminal
[0350] The server transmits question data to the terminal via the network.
[0351] Step 10:
[0352] The device asks the user the following question and instructs them to "take your temperature and blood pressure."
[0353] Input: Next question, measurement instructions
[0354] Output: Questions and measurement instructions for the user
[0355] The device will then ask the user the next question via voice and display, and instruct them to measure further health data.
[0356] Step 11:
[0357] The user measures their temperature using a thermometer or blood pressure monitor and reports the results to the robot.
[0358] Input: User measurements
[0359] Output: Reported measurement data (e.g., "Temperature is 36.5°C, Blood pressure is 120 / 80")
[0360] The user follows the instructions to measure their own body temperature and blood pressure and reports the results to the device.
[0361] Step 12:
[0362] The device records the reported health data as text data and sends it to the server.
[0363] Input: Reported measurement data
[0364] Output: Health data sent to the server
[0365] The terminal converts the measurement data into text format and uploads it to a server via the network.
[0366] Step 13:
[0367] The server comprehensively analyzes health data, text data, and emotional data to assess dementia risk.
[0368] Input: Health data, text data, emotion data
[0369] Output: Dementia risk assessment
[0370] An AI algorithm is used to integrate all data and assess dementia risk.
[0371] Step 14:
[0372] If the risk is high, the server generates an alert message and sends it to the terminal.
[0373] Input: Dementia Risk Assessment
[0374] Output: Alert message
[0375] If the risk is determined to be high, the server generates an alert message and sends it to the terminal.
[0376] Step 15:
[0377] The terminal provides feedback to the user.
[0378] Input: Alert message
[0379] Output: Feedback to the user (e.g., "You've noticed some concerns about your physical and emotional state recently. Consult your doctor.")
[0380] The terminal displays an alert message to the user and provides appropriate advice.
[0381] Step 16:
[0382] The server stores all conversation data, health data, and emotion data.
[0383] Input: Conversation data, health data, emotion data
[0384] Output: Saved data
[0385] The server saves all the data in a database ready for the next session.
[0386] (Application example 2)
[0387] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0388] Monitoring the health and emotional state of elderly workers in factories is important, but conventional systems have difficulty grasping the situation and assessing risks in real time, making it difficult to respond effectively. Early response to prevent accidents and illnesses is also often delayed. To solve these problems, a system is needed that can comprehensively analyze workers' health and emotional data and assess risks early.
[0389] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recognizing a user, means for initiating a voice conversation with the user using a voice recognition module and converting the voice data into text data, means including a natural language processing engine for analyzing the converted text data and emotional data and evaluating the user's emotions and health status, means for generating the next question based on the evaluation results and presenting it to the user, means for measuring the user's physical data and collecting the measurement results, means for integrating the voice and measurement data and evaluating health risks using AI, means for providing feedback to the user based on the evaluation results, means for transmitting the health data and emotional data to the server and generating feedback based on the analysis results, and means for storing the data and preparing for the next session. This makes it possible to comprehensively grasp the health and emotional status of workers in real time, evaluate risks early, and provide appropriate feedback and countermeasures.
[0390] "Means for recognizing a user" refers to technology or devices that use cameras or sensors to identify a specific user.
[0391] "Means for initiating a voice conversation and converting voice data into text data" refers to a voice recognition technology or system that converts voice data acquired through a microphone into character string data.
[0392] "Means for analyzing converted text data and assessing a user's emotions and health status" refers to technology or systems that use natural language processing engines and other analytical algorithms to infer and assess a user's emotions and health status from their statements and expressions.
[0393] The "means for generating the next question and presenting it to the user" refers to a technology or system that automatically generates the next appropriate question based on the evaluation results and presents it to the user by voice or display.
[0394] "Means for measuring a user's physical data and collecting the measurement results" refers to a technology or system that uses devices such as a thermometer or blood pressure monitor to measure data related to a user's body and collect the results as data.
[0395] "Means for integrating voice and measurement data and assessing dementia risk using AI" refers to a technology or system that integrates acquired voice data and physical data and uses an AI algorithm to assess a user's dementia risk.
[0396] The "means for providing feedback to the user based on the evaluation results" refers to a technology or system that provides necessary advice or instructions to the user based on the analysis and evaluation results.
[0397] "Means for transmitting health data and emotional data to a server and generating feedback based on the analysis results" refers to a technology or system that transmits collected data to a server and provides appropriate feedback to the user based on the analysis results on the server side.
[0398] "Means for storing the data and preparing for the next session" refers to a technology or system that stores all acquired data in a database and prepares it for the next monitoring or evaluation based on that data.
[0399] The present invention relates to a system for real-time monitoring of the health and emotional states of factory workers and early risk assessment, which is equipped with user recognition, voice conversation, data analysis, emotion engine, and feedback functions.
[0400] Hardware and software used
[0401] Hardware:
[0402] 1. Camera: Used to recognize the face of the worker.
[0403] 2. Microphone: Used for voice recognition.
[0404] 3. Health measurement devices (thermometer, blood pressure monitor, heart rate monitor): Used to obtain physical data of workers.
[0405] software:
[0406] 1. OpenCV: Used to recognize the worker's face from the image data acquired through the camera.
[0407] 2. SpeechRecognition (Python library): Used to convert audio data into text data.
[0408] 3. Natural language processing engine: Analyzes the acquired text and emotion data and uses it to assess the emotions and health status of workers.
[0409] 4. Health Monitoring Library: Provides functions (get_temperature, get_blood_pressure, get_heart_rate) to get body temperature, blood pressure, and heart rate.
[0410] 5. Emotion Recognition Library: Used to analyze emotions from workers' facial expression data or voice data.
[0411] 6. Requests (Python library): Used to send data to and retrieve data from the server.
[0412] Processing flow
[0413] The system first recognizes the user (worker) using a camera and a voice recognition module. For example, when the system starts up, the camera scans the worker's face and greets them verbally, saying, "Good morning, [Worker's name]. How are you feeling?" If the worker replies, "I'm not feeling well today," the voice data is converted into text data by the voice recognition module.
[0414] The converted text data and acquired emotional data (analyzed from facial expressions and voice) are then analyzed using a natural language processing engine, allowing the worker's emotions and health status to be assessed.
[0415] Next, the system asks the worker to measure their body temperature, blood pressure, and heart rate. For example, data such as "body temperature is 37.5 degrees, blood pressure is 130 / 85, and heart rate is 80" is collected from the health measurement device. The collected data is sent to a server, where a health risk assessment is performed.
[0416] Based on the analysis and evaluation results, the system generates the next appropriate question to ask the worker and presents it to them. If the worker's health condition is not good, the system will provide feedback such as "You seem unwell. Please consult a doctor."
[0417] All data is stored in a database and prepared for the next session, allowing for a real-time understanding of the health and emotional state of factory workers, early risk assessment, and appropriate feedback and countermeasures.
[0418] Specific examples
[0419] Example prompt sentence:
[0420] Speech recognition: "I'm not feeling well today"
[0421] Facial Recognition: "Good morning, [Worker's Name]. How are you feeling?"
[0422] Health data collection: "Temperature is 37.5°C, blood pressure is 130 / 85, heart rate is 80."
[0423] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0424] Step 1:
[0425] The terminal starts up and uses the camera to recognize the user's (worker's) face. At this time, the terminal analyzes the image data acquired through the camera using OpenCV and checks whether it matches the facial data registered in advance. The input is the image data acquired from the camera, and the output is the user ID. In concrete terms, the camera scans the worker's face and recognizes matching facial data.
[0426] Step 2:
[0427] The device greets the recognized user by voice and asks about their current health condition. For example, it might ask, "Good morning, [user name]. How are you feeling?" The input is the user ID, and the output is a voice message. Specifically, the device incorporates the user name into the greeting and outputs it by voice.
[0428] Step 3:
[0429] The user responds by saying something like, "I'm not feeling well today." This voice data is acquired through the device's microphone. The input is the user's voice data, and the output is voice data. In concrete terms, the user reports their physical condition by voice into the device.
[0430] Step 4:
[0431] The device uses a speech recognition module to convert the acquired voice data into text data. The input is voice data and the output is text data. Specifically, the device converts the acquired voice data into text data using the SpeechRecognition library.
[0432] Step 5:
[0433] The converted text data and facial expression data are analyzed to evaluate the user's emotions and health status. The device analyzes the text data and facial expression data using a natural language processing engine and an emotion recognition library. The input is the text data and facial expression data, and the output is the evaluation results. Specifically, the device evaluates the emotions and health status based on the analysis results.
[0434] Step 6:
[0435] The device generates the next question based on the evaluation results and presents it to the user. For example, it generates a question such as "Have you been sleeping well recently?" The input is the evaluation results, and the output is the question text. As a specific operation, the generated question is output to the user by voice or display.
[0436] Step 7:
[0437] The device asks the user to measure their body temperature, blood pressure, and heart rate. For example, it instructs the user to "measure their body temperature and blood pressure." The input is the evaluation result, and the output is an instruction message. Specifically, the device gives instructions to measure by voice.
[0438] Step 8:
[0439] A user uses a health measurement device to measure their body temperature, blood pressure, and heart rate, and reports the results to a terminal. For example, they report, "My body temperature is 37.5 degrees, my blood pressure is 130 / 85, and my heart rate is 80." The input is the measurement data, and the output is the reported data. Specifically, the user inputs the measurement data into the terminal.
[0440] Step 9:
[0441] The device sends the collected voice data, health data, and emotion data to the server. The input is the collected data, and the output is the result sent to the server. Specifically, the device uses the Requests library to send data to the server.
[0442] Step 10:
[0443] The server analyzes the received data and evaluates health risks. The input is the received data, and the output is the evaluation result. Specifically, the server uses an AI algorithm to perform risk assessment.
[0444] Step 11:
[0445] The server generates feedback based on the evaluation results and sends it to the device. For example, it generates feedback such as "You seem unwell, please consult a doctor." The input is the evaluation results, and the output is a feedback message. The specific operation is to send the generated feedback to the device.
[0446] Step 12:
[0447] The terminal presents the feedback received from the server to the user. The input is the feedback message, and the output is the result presented to the user. Specifically, the terminal conveys the feedback to the user by voice.
[0448] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0449] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0450] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0451] [Second embodiment]
[0452] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0453] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0454] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0455] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0456] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0457] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0458] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0459] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0460] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0461] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0462] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0463] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0464] This invention relates to a robotic system for monitoring the health status and diagnosing dementia in the home of elderly people living alone. The system is equipped with functions for user recognition, voice conversation, data analysis, and feedback.
[0465] First, when the device (robot) starts up, it recognizes the user using the built-in camera and voice recognition function. At this point, the robot greets the user by saying, "Good morning, [user name]. How are you today?"
[0466] Next, the user responds to the robot by saying something like, "I feel a bit heavy-headed today." This voice data is converted into text data by a voice recognition module in the device.
[0467] The converted text data is sent to a server and analyzed by a natural language processing engine (NLP). This analysis extracts information to evaluate the user's emotions and health status. For example, an expression such as "my head feels heavy" can be used to infer a decline in health.
[0468] Next, the server generates the next question based on the analysis results and sends it to the device. For example, the question generated might be, "Have you been sleeping well lately?" The device then presents this question to the user.
[0469] At the same time, the device prompts the user to measure their temperature and blood pressure. The user measures their temperature and blood pressure using a thermometer or blood pressure monitor at home and reports the results to the robot. For example, "My temperature is 36.5 degrees and my blood pressure is 120 / 80."
[0470] Health data and text data are integrated on a server, and an AI algorithm is used to assess dementia risk, based on abnormalities in language usage patterns and physical data.
[0471] For example, if the user has become increasingly forgetful in recent conversations or if there are abnormalities in the user's physical data, the server will generate an alert message and send it to the device. The device will then provide feedback such as, "After looking at the results, there are some concerns about your recent physical condition and conversations. We recommend that you consult a doctor."
[0472] All conversation and health data is stored in a database on the server, which allows for preparation for the next session and regular monitoring of the user's condition.
[0473] As a concrete example, the flow of one day is shown below.
[0474] 1. The device starts up, recognizes the user, and greets them with, "Good morning, Tanaka-san. How are you today?"
[0475] 2. The user replies, "Good morning, I'm feeling a bit down today."
[0476] 3. The device converts the voice data into text and sends it to the server.
[0477] 4. The server analyzes the text data, generates a question such as "Have you often felt heavy-headed lately?" and sends it to the device.
[0478] 5. The device asks questions and prompts the user to measure their health data. It prompts the user to "measure their temperature and blood pressure."
[0479] 6. The user reports the measurement results and replies, "Temperature is 36.5°C, blood pressure is 120 / 80."
[0480] 7. The device records this as text data and sends it to the server.
[0481] 8. The server analyzes all data and evaluates the risk of dementia. If the risk is high, an alert is sent to the device.
[0482] 9. The device will provide feedback such as, "After looking at the results, there are some concerns about your recent health condition and conversations. We recommend that you consult a doctor."
[0483] This system will enable daily monitoring of the health status of elderly people and enable early contact with medical institutions if necessary.
[0484] The processing flow will be explained below.
[0485] Step 1:
[0486] The device boots up and uses the built-in camera and microphone to scan the surroundings and identify the user. Once the user is identified, the device greets them with, "Good morning, [username]. How are you today?"
[0487] Step 2:
[0488] The user responds, "Good morning, I feel a bit heavy-headed today." The device collects the voice data and converts it into text data using a voice recognition module.
[0489] Step 3:
[0490] The converted text data is sent to a server, which then analyzes it using a natural language processing engine (NLP).The results of this analysis are used to evaluate the user's emotions and health status.
[0491] Step 4:
[0492] The server generates the next question based on the analysis results, for example, "Have you been sleeping well recently?", and sends this question to the device.
[0493] Step 5:
[0494] The device asks the user, "Have you been sleeping well lately?" While waiting for the user's response, the device instructs the user to measure their temperature and blood pressure. It adds, "Please measure your temperature and blood pressure."
[0495] Step 6:
[0496] The user uses a thermometer or blood pressure monitor to report the measurement results, such as "My body temperature is 36.5 degrees and my blood pressure is 120 / 80." The device's voice recognition module converts this voice back into text data, which is then sent to the server.
[0497] Step 7:
[0498] The server combines the health check data and text data and uses an AI algorithm to assess dementia risk, for example, based on information such as forgetfulness in recent conversations or abnormalities in physical data.
[0499] Step 8:
[0500] If the risk is high based on the assessment results, the server generates an alert message and sends it to the device, such as "There are some concerns about your health. We recommend that you consult a doctor."
[0501] Step 9:
[0502] The device will then provide the user with an alert message, saying, "After looking at the results, we've noticed some concerns about your recent health and conversations. We recommend that you consult a doctor."
[0503] Step 10:
[0504] The server stores all conversation and health data in a database, and periodically prepares data to monitor the user's condition in preparation for the next session.
[0505] Example 1
[0506] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0507] The present invention relates to a robotic system for monitoring the health status and dementia risk of elderly people living alone. Many current health monitoring systems require periodic manual input, which can be a burden for users. Furthermore, conventional systems do not perform detailed analysis of the user's emotions and health status, which can delay appropriate medical treatment when needed. Therefore, the present invention aims to provide a system that can continuously and automatically monitor the user's health status and promptly prompt appropriate medical treatment.
[0508] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0509] In this invention, the server includes means for recognizing a user, means for initiating a voice conversation with the user and converting the voice data into text data, means for analyzing the converted text data and evaluating the user's emotions and health status, means for generating the next question based on the evaluation results and presenting it to the user, means for collecting the user's physical data, means for integrating the voice and physical data and evaluating dementia risk using AI, means for providing feedback to the user based on the evaluation results, means for saving the data and preparing for the next session, and means for image recognition for identifying the user's face and voice recognition for recognizing the user's voice. This makes it possible to continuously and automatically monitor the user's health status in detail while minimizing the burden on the user, and to promote early and appropriate medical treatment.
[0510] "Means for recognizing the user" refers to technology that enables the robot to identify the user's face or specific features using cameras or sensors.
[0511] The "means for initiating a voice conversation and converting voice data into text data" is a voice recognition module for recognizing the voice spoken by the user and converting the voice into text information.
[0512] The "means for analyzing text data and assessing the user's emotions and health condition" refers to an algorithm or program that uses a natural language processing engine to analyze the converted text data and analyze the user's emotions and health condition.
[0513] The "means for generating the next question and presenting it to the user" refers to a device or software that generates a question to further evaluate the user's health condition based on the analysis results and presents it to the user via voice or screen display.
[0514] "Means for collecting user's physical data" refers to a method or device for collecting data on the user's physical condition using measuring instruments such as a thermometer or blood pressure monitor.
[0515] "Means for integrating voice and physical data and assessing dementia risk using AI" refers to a system that integrates and analyzes collected voice data and physical data, and assesses dementia risk using an AI algorithm.
[0516] The "means for providing feedback to the user based on the evaluation results" refers to a method or device that communicates the results of the analysis and evaluation to the user and provides a notice or warning that encourages the user to consult a medical institution if necessary.
[0517] The "means for storing the data and preparing for the next session" refers to a mechanism for recording and storing all data, including voice data and health data, to facilitate the next health monitoring session.
[0518] "Image recognition means" is a technology for analyzing facial features from images captured by a camera and identifying users.
[0519] "Speech recognition means" is a technology for capturing a user's voice and converting it into text data.
[0520] This invention relates to a robotic system for monitoring the health status and dementia risk of elderly people living alone. The system is equipped with functions for user recognition, voice conversation, data analysis, and feedback.
[0521] First, when the device (robot) starts up, it recognizes the user using the built-in camera and voice recognition function. The device greets the user by saying, "Good morning, [user name]. How are you today?" At this time, it uses the camera to identify the face and activates the voice recognition module. As a specific example, the device uses the built-in camera and voice recognition module to identify the user's face and voice.
[0522] Next, if the user responds, "I feel a bit heavy-headed today," this voice data is converted into text data by a voice recognition module. When the terminal converts voice data into text data, it uses a voice recognition module. For example, if the user says, "I'm feeling a bit unwell today," this is recorded as text data.
[0523] The converted text data is sent to a server and analyzed by a natural language processing engine (NLP). This analysis evaluates the user's emotions and health status. For example, the expression "my head feels heavy" can be used to infer a decline in health. As a specific example, the NLP engine analyzes the text "my head feels heavy" and determines that this is a sign of deteriorating health.
[0524] Next, the server generates the next question based on the analysis result and sends it to the terminal. For example, the question "Have you been sleeping well recently?" is generated. The terminal presents this question to the user. An example of a specific prompt sentence in this case is "Have you been sleeping well recently?"
[0525] At the same time, the device instructs the user to measure their temperature and blood pressure. The user uses a thermometer and blood pressure monitor to measure their temperature and report the results to the robot. For example, they might say, "My temperature is 36.5 degrees and my blood pressure is 120 / 80." Once the user reports the measurement results, the voice recognition module converts them back into text data and sends it to the server.
[0526] All data is integrated on a server, and an AI algorithm evaluates the risk of dementia. Risk is assessed based on abnormalities in language usage patterns and physical data. For example, if a person has become increasingly forgetful in recent conversations, this could be assessed as a risk of dementia.
[0527] If the risk is high, the server generates an alert message and sends it to the device, which then provides feedback such as, "After looking at the results, there are some concerns about your recent health condition and conversations. We recommend that you consult a doctor."
[0528] All conversation and health data is stored in a database on the server and ready for the next session, allowing for continuous, automatic, and detailed monitoring of the user's health status, facilitating early and appropriate medical treatment.
[0529] To give an example, here's what a typical day might look like:
[0530] 1. The device starts up, recognizes the user, and greets them with, "Good morning, Tanaka-san. How are you today?"
[0531] 2. The user replies, "Good morning, I'm feeling a bit down today."
[0532] 3. The device converts the voice data into text and sends it to the server.
[0533] 4. The server analyzes the text data, generates a question such as "Have you often felt heavy-headed lately?" and sends it to the device.
[0534] 5. The device asks questions and prompts the user to "take your temperature and blood pressure."
[0535] 6. The user reports the measurement results and replies, "Temperature is 36.5°C, blood pressure is 120 / 80."
[0536] 7. The device records this as text data and sends it to the server.
[0537] 8. The server analyzes all data and evaluates the risk of dementia. If the risk is high, an alert is sent to the device.
[0538] 9. The device will provide feedback such as, "After looking at the results, there are some concerns about your recent health condition and conversations. We recommend that you consult a doctor."
[0539] This system allows for daily monitoring of the health status of elderly people and allows for early linkage with medical institutions.
[0540] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0541] Step 1:
[0542] The device (robot) starts up. Once started, it uses the built-in camera and voice recognition function to recognize the user. Specifically, the camera captures the user's face and identifies the user using a facial recognition algorithm. Once the user is identified, the device greets them with "Good morning, [user name]. How are you today?"
[0543] Input: Start signal, camera image
[0544] Output: User recognition result, greeting message
[0545] Step 2:
[0546] The user responds by saying, "I feel a bit heavy-headed today." This voice data is captured by the device's voice recognition module and converted into text data. Specifically, voice input is captured by a microphone and then converted into text by the voice recognition module.
[0547] Input: User voice
[0548] Output: Text data
[0549] Step 3:
[0550] The device sends the converted text data to the server, where it packages the text data into a standard format such as JSON and securely transmits it to the server through a network module.
[0551] Input: Text data
[0552] Output: Server sent data
[0553] Step 4:
[0554] The server receives the text data and analyzes it using a natural language processing engine (NLP). This analysis evaluates the user's emotions and health status. Specifically, the NLP engine analyzes the text data and extracts keywords and emotions. At this stage, a decline in the user's health is inferred from the user's statement that "my head feels heavy."
[0555] Input: Text data
[0556] Output: Analysis results (emotion and health evaluation)
[0557] Step 5:
[0558] The server generates the next question based on the analysis results and sends it to the device. For example, a question might be generated such as, "Have you been sleeping well lately?" The server then creates an appropriate question based on the analysis results and sends it to the device as a data packet.
[0559] Input: Analysis results
[0560] Output: Next question data
[0561] Step 6:
[0562] The device then presents the generated questions to the user and instructs them to measure their health data. Specifically, it uses a voice synthesis module to prompt the user to "measure their temperature and blood pressure."
[0563] Input: Next question data
[0564] Output: Voice prompts and measurement instructions
[0565] Step 7:
[0566] The user measures their temperature and blood pressure using a thermometer and reports the results to the robot, for example, "My temperature is 36.5 degrees and my blood pressure is 120 / 80." This triggers the voice recognition module to convert the voice data into text.
[0567] Input: Measurement data (body temperature, blood pressure)
[0568] Output: Text data
[0569] Step 8:
[0570] The device sends the recorded health data as text data to the server.
[0571] Input: Text data
[0572] Output: Server sent data
[0573] Step 9:
[0574] The server integrates all the data and uses an AI algorithm to assess the risk of dementia. The AI algorithm analyzes language usage patterns and abnormalities in physical data to assess risk. For example, if the person has become increasingly forgetful in recent conversations, this could be assessed as a risk of dementia.
[0575] Input: Integrated data (voice and health)
[0576] Output: Risk assessment results
[0577] Step 10:
[0578] The server generates an alert message based on the risk assessment results and sends it to the device, which then provides feedback such as, "After looking at the results, there are some concerns about your recent health condition and conversations. We recommend that you consult a doctor."
[0579] Input: Risk assessment results
[0580] Output: Feedback message
[0581] Step 11:
[0582] The server stores all conversation and health data in a database and prepares for the next session, allowing for continuous monitoring of the user's health.
[0583] Input: Voice data, health data
[0584] Output: Save data, prepare for next session
[0585] (Application example 1)
[0586] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0587] It is extremely important to monitor the health status of elderly people living alone at home or in facilities and to detect dementia risk early. However, current systems have difficulty comprehensively and continuously monitoring users' health status and emotions, making early detection and prompt feedback difficult. For this reason, there is a need for a system that provides more comprehensive and accurate monitoring and feedback through devices or smartphone applications installed in facilities that provide healthcare services for the elderly.
[0588] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0589] In this invention, the server includes a means for recognizing the user, a means for initiating a voice conversation with the user and converting the voice data into text data, a means for analyzing the converted text data and evaluating the user's emotions and health status, a means for generating the next question based on the evaluation results and presenting it to the user, and a means for measuring the user's physical data and collecting the measurement results. This enables a system including a means for integrating the voice and measurement data and assessing dementia risk using AI, a means for providing feedback to the user based on the evaluation results, and a means for storing the data and preparing for the next session. This enables more comprehensive and accurate health status monitoring and rapid feedback through terminals and smartphone applications installed in facilities providing services for the elderly.
[0590] "Means for recognizing the user" refers to technology that uses a camera and voice recognition technology to verify the user's personal information.
[0591] The "means for converting voice data into text data" refers to a technique that uses a voice recognition module to convert recorded voice into a corresponding text format.
[0592] "Means for analyzing converted text data and assessing the user's emotions and health status" refers to a technology that utilizes a natural language processing engine and a generative AI model to assess the user's emotions and health status from text data.
[0593] "Means for generating the next question and presenting it to the user" refers to a technology that generates a new question based on the analysis results and presents it to the user in voice or text.
[0594] The "means for collecting user's physical data" refers to a technology for collecting user's physical data such as body temperature and blood pressure using measuring devices such as a thermometer and a blood pressure monitor.
[0595] "Method for integrating voice and measurement data to assess dementia risk using AI" refers to a technology that integrates collected voice data and physical data and uses an AI algorithm to assess dementia risk.
[0596] "Means of providing feedback to users" refers to technology that provides the results of AI evaluation to users via voice or text.
[0597] The "means for saving the data and preparing for the next session" refers to a technique for saving all collected data in a database and using it as the basis for future sessions.
[0598] "Means for conducting voice conversations with the user and collecting health data using a terminal installed in a facility that provides services for the elderly" refers to technology for conducting voice conversations and collecting health data using a device installed within the facility.
[0599] "Means for interacting with the user and managing health data through a smartphone application" refers to technology that uses a smartphone application to exchange information with the user and collect and manage health data.
[0600] "Means of sending data to a cloud environment and analyzing the user's emotions and health status using a natural language processing engine and generative AI model" refers to a technology that sends collected data to the cloud via the internet and analyzes it using a natural language processing engine and generative AI model.
[0601] "Means for providing alerts regarding health status on facility devices and smartphone applications based on evaluation results" refers to technology that provides alerts to users on facility devices and smartphone applications based on the results of AI evaluation.
[0602] This invention is a system for monitoring a user's health status and assessing their dementia risk. The system consists of a terminal installed in a facility, a smartphone application, a cloud server, and multiple software modules for linking these components.
[0603] System Configuration
[0604] 1. Terminal installation location
[0605] In this invention, a robot terminal is installed in a facility that provides services for the elderly. The terminal is equipped with a camera and voice recognition function, and can interact with users. The terminal recognizes the user's voice and facial expressions and identifies personal information.
[0606] 2. Smartphone Applications
[0607] Users can also receive the same services outside the facility using a smartphone application. The application has functions for voice conversations with users, recording health data, and providing consultations. The app synchronizes with the robot terminal and transmits data to a cloud server.
[0608] 3. Cloud environment
[0609] Data analysis and storage are performed in a cloud environment: collected voice and body data are sent to the cloud and processed by natural language processing engines (e.g., SpaCy or Google Cloud Natural Language API) and generative AI models.
[0610] Hardware and Software Details
[0611] Hardware
[0612] Camera: Built into the robot terminal and used to recognize the user's face.
[0613] Microphone: Used to collect voice conversation data.
[0614] Thermometers, blood pressure monitors: Health monitoring devices for measuring users' temperature and blood pressure. These are installed within the facility.
[0615] software
[0616] Speech recognition module: Technology for converting voice data into text data (e.g., Google Cloud Speech-to-Text).
[0617] Natural language processing engines and generative AI models: Analyze users' emotions and health status, and assess dementia risk (e.g., SpaCy, Google Cloud Natural Language API).
[0618] Cloud storage: Infrastructure for storing and managing data.
[0619] Detailed explanation of the process
[0620] 1. User Awareness
[0621] The device uses a camera and voice recognition technology to identify the user. When the user enters the facility, the device greets them with, "Good morning, how are you feeling today?"
[0622] 2. Audio data collection and analysis
[0623] When the user responds, the voice data is converted into text data using a voice recognition module. This data is sent to a cloud server and analyzed using a natural language processing engine. Based on the analysis results, the user's health status and emotions are evaluated.
[0624] 3. Next question generation and presentation
[0625] The next question is generated from the text data by the generative AI model and presented to the user via their device or smartphone app, such as "Have you been sleeping well lately?"
[0626] 4. Collection of health data
[0627] Physical data such as users' temperature and blood pressure will be collected using measuring equipment within the facility, and similar data can also be collected via a smartphone app.
[0628] 5. Data Storage and Feedback
[0629] The collected voice and physical data is stored in cloud storage. Based on the analysis results, feedback is provided to the user. For example, the system may provide feedback such as, "After looking at the results, there are some concerns about your recent physical condition and conversations. Please consult a doctor."
[0630] Specific examples
[0631] 1. Usage scenarios in facilities
[0632] When an elderly person visits the facility, the robot terminal recognizes them and asks, "Good morning, how are you feeling today?" If the elderly person replies, "I have a slight headache," the voice is converted into text data and sent to the cloud.
[0633] 2. Smartphone usage scenarios
[0634] The same elderly person opens the smartphone app at home and is asked the same questions about their health. The collected data is stored in the cloud and used the next time they visit the facility. Based on this data, a generative AI model generates appropriate questions and performs a risk assessment.
[0635] 3. Examples of prompts
[0636] Please analyze the following sentences and assess the user's health and emotions.
[0637] "My head feels a little heavy today."
[0638] This system can continuously monitor the user's health status wherever they are and provide quick feedback, contributing to the early detection of dementia risk.
[0639] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0640] Step 1:
[0641] When a user arrives at a facility, the device uses a camera and voice recognition technology to recognize the user. The camera captures the user's face, and the voice recognition module identifies the user's personal information from the voice data. Specifically, the camera captures the user's image and receives a greeting as voice data. The device outputs a voice message saying, "Good morning, how are you feeling today?"
[0642] Step 2:
[0643] The user responds to the terminal by voice. The terminal captures the voice data and converts it into text data using a voice recognition module. The user's voice input is "I feel a bit heavy-headed today," which is converted into text data. This text data is sent to the server.
[0644] Step 3:
[0645] The server analyzes the received text data and uses a natural language processing engine (e.g., SpaCy or Google Cloud Natural Language API) to extract data from the text data to evaluate the user's emotions and health status. The input is the text data "I feel a bit heavy-headed today," and the output is "The user is feeling unwell."
[0646] Step 4:
[0647] The server uses the generative AI model based on the analysis results to generate the next question. For example, the question "Have you been sleeping well lately?" is generated. This generated question data is sent to the device. Specifically, the text data of the generated question is sent to the device.
[0648] Step 5:
[0649] The device presents the generated question to the user by voice or text. The user then responds to the device. This response data is also converted into text data using a voice recognition module and sent to the server. The specific operation is that the device outputs the question "Have you been sleeping well lately?" by voice.
[0650] Step 6:
[0651] The terminal instructs the user to measure their physical data using a thermometer or blood pressure monitor. The user measures their body temperature and blood pressure using health measurement equipment installed in the facility and reports the results to the terminal. For example, the user reports, "My body temperature is 36.5 degrees and my blood pressure is 120 / 80." This data is converted into text data and sent to the server.
[0652] Step 7:
[0653] The server integrates all the received data and uses an AI algorithm to assess the risk of dementia. The voice data and physical data are analyzed as a single data set to assess the user's risk. The inputs are voice-text data such as "Head feels heavy," body temperature of "36.5 degrees," and blood pressure of "120 / 80," and the output is the "dementia risk assessment result." The specific operation is for the AI algorithm to analyze the data and perform a risk assessment.
[0654] Step 8:
[0655] The server provides feedback to the device and smartphone app based on the evaluation results. For example, it generates an alert message saying, "After looking at the results, there are some concerns about your recent health condition and conversations. Please consult a doctor." and sends it to the device and smartphone app. The user receives the alert message.
[0656] Step 9:
[0657] All data is stored in cloud storage and will be available for the next session. Along with regular data monitoring, it will be possible to continuously track the user's health status. Specifically, the collected voice and physical data will be stored in the cloud and used for future analysis.
[0658] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0659] This invention relates to a robotic system for monitoring the health status and diagnosing dementia in the home of elderly people living alone. The system is equipped with functions for user recognition, voice conversation, data analysis, emotion engine, and feedback.
[0660] First, when the device (robot) starts up, it recognizes the user using the built-in camera and voice recognition function. At this point, the robot greets the user by saying, "Good morning, [user name]. How are you today?"
[0661] The user then responds to the robot by saying something like, "I feel a bit heavy-headed today." This voice data is converted into text data by the device's voice recognition module. At the same time, the device acquires the user's voice and facial expression data and analyzes them using an emotion engine.
[0662] The converted text data and the evaluation data from the emotion engine are sent to a server and analyzed by a natural language processing engine (NLP). This analysis extracts information to evaluate the user's emotions and health status. For example, it can infer a decline in health status from expressions such as "my head feels heavy" and recognize emotions such as "I'm tired."
[0663] The server then generates the next question based on the analysis results and sends it to the device. For example, the question is "Have you been sleeping well lately?", taking into account emotional changes in the user's facial expressions. This allows for more personalized interactions.
[0664] The device will ask the user this question. At the same time, the device will ask the user to take their temperature and blood pressure. It will add, "Please take your temperature and blood pressure."
[0665] The user measures their temperature using a thermometer or blood pressure monitor at home and reports the results to the robot, for example, "My temperature is 36.5 degrees and my blood pressure is 120 / 80."
[0666] Health data, text data, and emotional data are integrated on a server, and an AI algorithm is used to assess dementia risk. The AI assesses risk based on abnormalities in language usage patterns and physical data, while also taking emotional data into account.
[0667] For example, if the user has become increasingly forgetful in recent conversations or has frequently shown emotionally unstable facial expressions, the server will generate an alert message and send it to the device, which will then provide feedback such as, "After looking at the results, there are some concerns about your recent physical condition and emotions. We recommend that you consult a doctor."
[0668] All conversation, health and emotion data is stored in a database on the server, allowing for preparation for the next session and regular monitoring of the user's condition.
[0669] As a concrete example, the flow of one day is shown below.
[0670] 1. The device starts up, recognizes the user, and greets them with, "Good morning, Tanaka-san. How are you today?"
[0671] 2. The user replies, "Good morning, I'm feeling a bit down today."
[0672] 3. The device collects voice and facial expression data, converts the voice into text, and analyzes it using an emotion engine.
[0673] 4. The text data and emotional evaluation data sent to the server are analyzed, and the question "Have you often felt heavy-headed lately?" is generated and sent to the device.
[0674] 5. The device asks questions and prompts the user to measure their health data. It prompts the user to "measure their temperature and blood pressure."
[0675] 6. The user reports the measurement results and replies, "Temperature is 36.5°C, blood pressure is 120 / 80."
[0676] 7. The device records this as text data and emotion data and sends it to the server.
[0677] 8. The server analyzes all data and evaluates the risk of dementia. If the risk is high, an alert is sent to the device.
[0678] 9. The device will provide feedback such as, "After looking at the results, there are some concerns about your recent physical condition and emotions. We recommend that you consult a doctor."
[0679] This system allows for daily monitoring of the health and emotional state of the elderly and allows for early consultation with medical institutions when necessary, thereby improving the peace of mind and quality of life for the elderly and their families.
[0680] The processing flow will be explained below.
[0681] Step 1:
[0682] The device boots up and uses the built-in camera and microphone to scan the surroundings and identify the user. Once the user is identified, the device greets them with, "Good morning, [username]. How are you today?"
[0683] Step 2:
[0684] The user responds, "Good morning, I feel a bit heavy-headed today." The device collects the voice data and converts it into text data using a voice recognition module.
[0685] Step 3:
[0686] The device collects facial expression data along with the captured voice data and analyzes it with an emotion engine, where the user's emotional state is identified as "fatigue."
[0687] Step 4:
[0688] The converted text data and the evaluation data generated by the emotion engine are sent to the server, which then analyzes the received text data using a natural language processing engine (NLP).The analysis results are used to evaluate the user's emotions and health status.
[0689] Step 5:
[0690] The server generates the next question based on the analysis results, for example, "Have you been sleeping well recently?", and sends this question to the device.
[0691] Step 6:
[0692] The device asks the user, "Have you been sleeping well lately?" When asking the question, it also takes into account changes in the user's facial expressions based on their emotions. At the same time, the device instructs the user to measure their temperature and blood pressure. It adds, "Please measure your temperature and blood pressure."
[0693] Step 7:
[0694] The user uses a thermometer or blood pressure monitor to report the measurement results, such as "My body temperature is 36.5 degrees and my blood pressure is 120 / 80." The device's voice recognition module converts this voice back into text data, which is then sent to the server.
[0695] Step 8:
[0696] The server combines the health check data, text data, and emotional data, and uses an AI algorithm to assess dementia risk. For example, if a person has recently shown forgetfulness and emotional instability in their facial expressions, it will determine that they are at high risk.
[0697] Step 9:
[0698] If the risk is high based on the assessment results, the server generates an alert message and sends it to the device, such as "There are some concerns about your health. We recommend that you consult a doctor."
[0699] Step 10:
[0700] The device will then provide the user with an alert message, saying, "After looking at your results, there are some concerns about your recent physical condition and emotions. We suggest you consult a doctor."
[0701] Step 11:
[0702] The server stores all conversation data, health data, and emotion data in a database. It periodically prepares data to monitor the user's condition in preparation for the next session.
[0703] Example 2
[0704] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0705] In modern society, there is a need to routinely monitor the health status and dementia risk of elderly people living alone and detect abnormalities early. However, the means to do so are limited, and there are no systems that use interactive systems to closely observe the health and emotional state of elderly people and provide appropriate feedback. Therefore, there is a need for an efficient and effective method of managing the health of elderly people.
[0706] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for recognizing the user, means for initiating a voice conversation with the user and converting the voice data into text data, means for analyzing the converted text data and evaluating the user's emotions and health status, means for generating the next question based on the evaluation results and presenting it to the user, means for measuring the user's physical data and collecting the measurement results, means for integrating the voice and measurement data and evaluating dementia risk using AI, means for providing feedback to the user based on the evaluation results, and means for saving the data and preparing for the next session. This makes it possible to monitor the health and emotional status of elderly people on a daily basis and to cooperate with medical institutions early if necessary.
[0707] "Means for recognizing a user" refers to the system's ability to identify and recognize the individual user using a camera or voice recognition technology.
[0708] The "means for starting a voice conversation and converting voice data into text data" is a technology for automatically converting voice data acquired through a voice dialogue with a user into text format.
[0709] The "means for analyzing the converted text data and evaluating the user's emotions and health state" is a technology that processes the text data, analyzes its contents, and estimates the user's current emotions and health state.
[0710] The "means for generating the next question and presenting it to the user" is a function for generating and presenting an appropriate next question to the user based on the analysis results.
[0711] "Means for measuring the user's physical data and collecting the measurement results" refers to technology that allows the user to measure physical data such as body temperature and blood pressure and for the system to collect the results.
[0712] "Means for integrating voice and measurement data and assessing dementia risk using AI" refers to a technology that integrates collected voice data and physical data and uses artificial intelligence technology to assess a user's dementia risk.
[0713] The "means for providing feedback to the user based on the evaluation results" is a function for providing the results of the system's evaluation to the user in the form of information or advice in an appropriate format.
[0714] "Means for storing data and preparing for the next session" refers to technology for recording collected and analyzed data and preparing it for use in the next interaction or monitoring.
[0715] The present invention relates to a robotic system for monitoring the health status and dementia risk of elderly people living alone and for detecting abnormalities at an early stage. This system mainly includes a means for recognizing a user, a means for performing voice conversation and data analysis, an emotion analysis engine, a feedback means, and a storage means.
[0716] When the device (robot) starts up, it uses a camera and voice recognition functions to recognize the user. Specifically, it takes a picture of the user's face with the camera and recognizes the user's voice using voice recognition technology. At this time, the device says, "Good morning, [user name]. How are you today?"
[0717] When the user responds, the voice data is converted into text data using the device's voice recognition module. This converted text data and facial expression data captured by the camera are analyzed by the emotion engine. For example, by analyzing a statement such as "I feel a bit heavy-headed today" and the user's facial expression at the time, the user's emotions and health condition can be evaluated.
[0718] The device then sends this data to a server, which incorporates a natural language processing engine (NLP) and performs detailed analysis of the received text data and emotional evaluation data. For example, an expression such as "My head feels heavy" is interpreted as a sign of the user's stress or poor health. Emotional fluctuations are also extracted from facial expression data.
[0719] Based on the analysis results, the server generates the next question. For example, "Have you been sleeping well recently?" and sends it to the device. The system then instructs the user to measure their temperature and blood pressure. The system prompts the user to "measure their temperature and blood pressure."
[0720] The user measures their temperature using a thermometer or blood pressure monitor at home and reports the results to the robot. For example, they might report, "My temperature is 36.5 degrees and my blood pressure is 120 / 80." This health data is collected as text data and sent back to the server.
[0721] The server uses an AI algorithm to comprehensively analyze this data. It evaluates the user's dementia risk by combining voice, emotional, and physical data. For example, if the user has recently shown an increase in forgetfulness and emotional instability in their facial expressions, the server will assess the risk as high and generate an alert message.
[0722] Finally, the device provides the generated feedback to the user, for example, saying, "After looking at the results, there are some concerns about your recent physical condition and emotions. We recommend that you consult a doctor." This makes it possible to monitor the health status of elderly people on a daily basis and to consult with medical institutions early if necessary.
[0723] All conversation, health, and emotion data is stored on the server and prepared for the next session. Through regular monitoring, the health status of the elderly can be continuously managed.
[0724] (Example of a prompt)
[0725] Here are some example prompts for a generative AI model:
[0726] "Design a robotic system to check the health status of elderly people daily. This system should have the ability to recognize the user's face and voice, analyze their health status and emotions, and provide feedback as needed. It should also send the measured health data and analyzed emotional data to a server, enabling regular monitoring."
[0727] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0728] (Program processing steps)
[0729] Step 1:
[0730] The device starts up.
[0731] Input: System power supply
[0732] Output: The terminal is ready
[0733] At startup, the built-in camera and voice recognition function are enabled to prepare for user recognition.
[0734] Step 2:
[0735] The device uses its built-in camera and voice recognition capabilities to identify the user.
[0736] Input: User's face, voice
[0737] Output: Recognized user information
[0738] The camera captures the user's face and uses voice recognition technology to identify the individual from their voice.
[0739] The device greets you with, "Good morning, [username]. How are you today?"
[0740] Step 3:
[0741] The user responds, "I'm feeling a bit heavy-headed today."
[0742] Input: User's voice response
[0743] Output: User's voice data
[0744] The terminal records the user's response.
[0745] Step 4:
[0746] The terminal uses a voice recognition module to convert the user's voice data into text data.
[0747] Input: User's voice data
[0748] Output: Converted text data
[0749] The voice recognition module is activated to convert the voice data into text format.
[0750] Step 5:
[0751] The device collects the user's facial expression data captured by the camera and analyzes it using an emotion engine.
[0752] Input: User's facial expression data
[0753] Output: User's emotion rating data
[0754] A camera captures the user's facial expressions and uses an emotion analysis algorithm to assess their emotional state.
[0755] Step 6:
[0756] The terminal transmits the converted text data and emotion evaluation data to the server.
[0757] Input: Text data, emotion rating data
[0758] Output: Data sent to the server
[0759] The terminal uploads the data to the server via the network.
[0760] Step 7:
[0761] The server uses a natural language processing engine (NLP) to analyze the text data and assess the user's health status.
[0762] Input: Text data, emotion rating data
[0763] Output: User's health status rating
[0764] Using NLP technology, the system analyzes the converted text and, for example, infers a decline in the user's health from an expression like "my head feels heavy." It also takes into account emotional data.
[0765] Step 8:
[0766] The server generates the next question based on the analysis results.
[0767] Input: User's health assessment
[0768] Output: Next question (e.g., "Have you been sleeping well lately?")
[0769] Based on the analysis results, appropriate follow-up questions are automatically generated.
[0770] Step 9:
[0771] The server generates the next question and sends it to the terminal.
[0772] Input: Next question
[0773] Output: Questions sent to the terminal
[0774] The server transmits question data to the terminal via the network.
[0775] Step 10:
[0776] The device asks the user the following question and instructs them to "take your temperature and blood pressure."
[0777] Input: Next question, measurement instructions
[0778] Output: Questions and measurement instructions for the user
[0779] The device will then ask the user the next question via voice and display, and instruct them to measure further health data.
[0780] Step 11:
[0781] The user measures their temperature using a thermometer or blood pressure monitor and reports the results to the robot.
[0782] Input: User measurements
[0783] Output: Reported measurement data (e.g., "Temperature is 36.5°C, Blood pressure is 120 / 80")
[0784] The user follows the instructions to measure their own body temperature and blood pressure and reports the results to the device.
[0785] Step 12:
[0786] The device records the reported health data as text data and sends it to the server.
[0787] Input: Reported measurement data
[0788] Output: Health data sent to the server
[0789] The terminal converts the measurement data into text format and uploads it to a server via the network.
[0790] Step 13:
[0791] The server comprehensively analyzes health data, text data, and emotional data to assess dementia risk.
[0792] Input: Health data, text data, emotion data
[0793] Output: Dementia risk assessment
[0794] An AI algorithm is used to integrate all data and assess dementia risk.
[0795] Step 14:
[0796] If the risk is high, the server generates an alert message and sends it to the terminal.
[0797] Input: Dementia Risk Assessment
[0798] Output: Alert message
[0799] If the risk is determined to be high, the server generates an alert message and sends it to the terminal.
[0800] Step 15:
[0801] The terminal provides feedback to the user.
[0802] Input: Alert message
[0803] Output: Feedback to the user (e.g., "You've noticed some concerns about your physical and emotional state recently. Consult your doctor.")
[0804] The terminal displays an alert message to the user and provides appropriate advice.
[0805] Step 16:
[0806] The server stores all conversation data, health data, and emotion data.
[0807] Input: Conversation data, health data, emotion data
[0808] Output: Saved data
[0809] The server saves all the data in a database ready for the next session.
[0810] (Application example 2)
[0811] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0812] Monitoring the health and emotional state of elderly workers in factories is important, but conventional systems have difficulty grasping the situation and assessing risks in real time, making it difficult to respond effectively. Early response to prevent accidents and illnesses is also often delayed. To solve these problems, a system is needed that can comprehensively analyze workers' health and emotional data and assess risks early.
[0813] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recognizing a user, means for initiating a voice conversation with the user using a voice recognition module and converting the voice data into text data, means including a natural language processing engine for analyzing the converted text data and emotional data and evaluating the user's emotions and health status, means for generating the next question based on the evaluation results and presenting it to the user, means for measuring the user's physical data and collecting the measurement results, means for integrating the voice and measurement data and evaluating health risks using AI, means for providing feedback to the user based on the evaluation results, means for transmitting the health data and emotional data to the server and generating feedback based on the analysis results, and means for storing the data and preparing for the next session. This makes it possible to comprehensively grasp the health and emotional status of workers in real time, evaluate risks early, and provide appropriate feedback and countermeasures.
[0814] "Means for recognizing a user" refers to technology or devices that use cameras or sensors to identify a specific user.
[0815] "Means for initiating a voice conversation and converting voice data into text data" refers to a voice recognition technology or system that converts voice data acquired through a microphone into character string data.
[0816] "Means for analyzing converted text data and assessing a user's emotions and health status" refers to technology or systems that use natural language processing engines and other analytical algorithms to infer and assess a user's emotions and health status from their statements and expressions.
[0817] The "means for generating the next question and presenting it to the user" refers to a technology or system that automatically generates the next appropriate question based on the evaluation results and presents it to the user by voice or display.
[0818] "Means for measuring a user's physical data and collecting the measurement results" refers to a technology or system that uses devices such as a thermometer or blood pressure monitor to measure data related to a user's body and collect the results as data.
[0819] "Means for integrating voice and measurement data and assessing dementia risk using AI" refers to a technology or system that integrates acquired voice data and physical data and uses an AI algorithm to assess a user's dementia risk.
[0820] The "means for providing feedback to the user based on the evaluation results" refers to a technology or system that provides necessary advice or instructions to the user based on the analysis and evaluation results.
[0821] "Means for transmitting health data and emotional data to a server and generating feedback based on the analysis results" refers to a technology or system that transmits collected data to a server and provides appropriate feedback to the user based on the analysis results on the server side.
[0822] "Means for storing the data and preparing for the next session" refers to a technology or system that stores all acquired data in a database and prepares it for the next monitoring or evaluation based on that data.
[0823] The present invention relates to a system for real-time monitoring of the health and emotional states of factory workers and early risk assessment, which is equipped with user recognition, voice conversation, data analysis, emotion engine, and feedback functions.
[0824] Hardware and software used
[0825] Hardware:
[0826] 1. Camera: Used to recognize the face of the worker.
[0827] 2. Microphone: Used for voice recognition.
[0828] 3. Health measurement devices (thermometer, blood pressure monitor, heart rate monitor): Used to obtain physical data of workers.
[0829] software:
[0830] 1. OpenCV: Used to recognize the worker's face from the image data acquired through the camera.
[0831] 2. SpeechRecognition (Python library): Used to convert audio data into text data.
[0832] 3. Natural language processing engine: Analyzes the acquired text and emotion data and uses it to assess the emotions and health status of workers.
[0833] 4. Health Monitoring Library: Provides functions (get_temperature, get_blood_pressure, get_heart_rate) to get body temperature, blood pressure, and heart rate.
[0834] 5. Emotion Recognition Library: Used to analyze emotions from workers' facial expression data or voice data.
[0835] 6. Requests (Python library): Used to send data to and retrieve data from the server.
[0836] Processing flow
[0837] The system first recognizes the user (worker) using a camera and a voice recognition module. For example, when the system starts up, the camera scans the worker's face and greets them verbally, saying, "Good morning, [Worker's name]. How are you feeling?" If the worker replies, "I'm not feeling well today," the voice data is converted into text data by the voice recognition module.
[0838] The converted text data and acquired emotional data (analyzed from facial expressions and voice) are then analyzed using a natural language processing engine, allowing the worker's emotions and health status to be assessed.
[0839] Next, the system asks the worker to measure their body temperature, blood pressure, and heart rate. For example, data such as "body temperature is 37.5 degrees, blood pressure is 130 / 85, and heart rate is 80" is collected from the health measurement device. The collected data is sent to a server, where a health risk assessment is performed.
[0840] Based on the analysis and evaluation results, the system generates the next appropriate question to ask the worker and presents it to them. If the worker's health condition is not good, the system will provide feedback such as "You seem unwell. Please consult a doctor."
[0841] All data is stored in a database and prepared for the next session, allowing for a real-time understanding of the health and emotional state of factory workers, early risk assessment, and appropriate feedback and countermeasures.
[0842] Specific examples
[0843] Example prompt sentence:
[0844] Speech recognition: "I'm not feeling well today"
[0845] Facial Recognition: "Good morning, [Worker's Name]. How are you feeling?"
[0846] Health data collection: "Temperature is 37.5°C, blood pressure is 130 / 85, heart rate is 80."
[0847] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0848] Step 1:
[0849] The terminal starts up and uses the camera to recognize the user's (worker's) face. At this time, the terminal analyzes the image data acquired through the camera using OpenCV and checks whether it matches the facial data registered in advance. The input is the image data acquired from the camera, and the output is the user ID. In concrete terms, the camera scans the worker's face and recognizes matching facial data.
[0850] Step 2:
[0851] The device greets the recognized user by voice and asks about their current health condition. For example, it might ask, "Good morning, [user name]. How are you feeling?" The input is the user ID, and the output is a voice message. Specifically, the device incorporates the user name into the greeting and outputs it by voice.
[0852] Step 3:
[0853] The user responds by saying something like, "I'm not feeling well today." This voice data is acquired through the device's microphone. The input is the user's voice data, and the output is voice data. In concrete terms, the user reports their physical condition by voice into the device.
[0854] Step 4:
[0855] The device uses a speech recognition module to convert the acquired voice data into text data. The input is voice data and the output is text data. Specifically, the device converts the acquired voice data into text data using the SpeechRecognition library.
[0856] Step 5:
[0857] The converted text data and facial expression data are analyzed to evaluate the user's emotions and health status. The device analyzes the text data and facial expression data using a natural language processing engine and an emotion recognition library. The input is the text data and facial expression data, and the output is the evaluation results. Specifically, the device evaluates the emotions and health status based on the analysis results.
[0858] Step 6:
[0859] The device generates the next question based on the evaluation results and presents it to the user. For example, it generates a question such as "Have you been sleeping well recently?" The input is the evaluation results, and the output is the question text. As a specific operation, the generated question is output to the user by voice or display.
[0860] Step 7:
[0861] The device asks the user to measure their body temperature, blood pressure, and heart rate. For example, it instructs the user to "measure their body temperature and blood pressure." The input is the evaluation result, and the output is an instruction message. Specifically, the device gives instructions to measure by voice.
[0862] Step 8:
[0863] A user uses a health measurement device to measure their body temperature, blood pressure, and heart rate, and reports the results to a terminal. For example, they report, "My body temperature is 37.5 degrees, my blood pressure is 130 / 85, and my heart rate is 80." The input is the measurement data, and the output is the reported data. Specifically, the user inputs the measurement data into the terminal.
[0864] Step 9:
[0865] The device sends the collected voice data, health data, and emotion data to the server. The input is the collected data, and the output is the result sent to the server. Specifically, the device uses the Requests library to send data to the server.
[0866] Step 10:
[0867] The server analyzes the received data and evaluates health risks. The input is the received data, and the output is the evaluation result. Specifically, the server uses an AI algorithm to perform risk assessment.
[0868] Step 11:
[0869] The server generates feedback based on the evaluation results and sends it to the device. For example, it generates feedback such as "You seem unwell, please consult a doctor." The input is the evaluation results, and the output is a feedback message. The specific operation is to send the generated feedback to the device.
[0870] Step 12:
[0871] The terminal presents the feedback received from the server to the user. The input is the feedback message, and the output is the result presented to the user. Specifically, the terminal conveys the feedback to the user by voice.
[0872] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0873] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0874] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0875] [Third embodiment]
[0876] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0877] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0878] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0879] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0880] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0881] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0882] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0883] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0884] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0885] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0886] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0887] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0888] This invention relates to a robotic system for monitoring the health status and diagnosing dementia in the home of elderly people living alone. The system is equipped with functions for user recognition, voice conversation, data analysis, and feedback.
[0889] First, when the device (robot) starts up, it recognizes the user using the built-in camera and voice recognition function. At this point, the robot greets the user by saying, "Good morning, [user name]. How are you today?"
[0890] Next, the user responds to the robot by saying something like, "I feel a bit heavy-headed today." This voice data is converted into text data by a voice recognition module in the device.
[0891] The converted text data is sent to a server and analyzed by a natural language processing engine (NLP). This analysis extracts information to evaluate the user's emotions and health status. For example, an expression such as "my head feels heavy" can be used to infer a decline in health.
[0892] Next, the server generates the next question based on the analysis results and sends it to the device. For example, the question generated might be, "Have you been sleeping well lately?" The device then presents this question to the user.
[0893] At the same time, the device prompts the user to measure their temperature and blood pressure. The user measures their temperature and blood pressure using a thermometer or blood pressure monitor at home and reports the results to the robot. For example, "My temperature is 36.5 degrees and my blood pressure is 120 / 80."
[0894] Health data and text data are integrated on a server, and an AI algorithm is used to assess dementia risk, based on abnormalities in language usage patterns and physical data.
[0895] For example, if the user has become increasingly forgetful in recent conversations or if there are abnormalities in the user's physical data, the server will generate an alert message and send it to the device. The device will then provide feedback such as, "After looking at the results, there are some concerns about your recent physical condition and conversations. We recommend that you consult a doctor."
[0896] All conversation and health data is stored in a database on the server, which allows for preparation for the next session and regular monitoring of the user's condition.
[0897] As a concrete example, the flow of one day is shown below.
[0898] 1. The device starts up, recognizes the user, and greets them with, "Good morning, Tanaka-san. How are you today?"
[0899] 2. The user replies, "Good morning, I'm feeling a bit down today."
[0900] 3. The device converts the voice data into text and sends it to the server.
[0901] 4. The server analyzes the text data, generates a question such as "Have you often felt heavy-headed lately?" and sends it to the device.
[0902] 5. The device asks questions and prompts the user to measure their health data. It prompts the user to "measure their temperature and blood pressure."
[0903] 6. The user reports the measurement results and replies, "Temperature is 36.5°C, blood pressure is 120 / 80."
[0904] 7. The device records this as text data and sends it to the server.
[0905] 8. The server analyzes all data and evaluates the risk of dementia. If the risk is high, an alert is sent to the device.
[0906] 9. The device will provide feedback such as, "After looking at the results, there are some concerns about your recent health condition and conversations. We recommend that you consult a doctor."
[0907] This system will enable daily monitoring of the health status of elderly people and enable early contact with medical institutions if necessary.
[0908] The processing flow will be explained below.
[0909] Step 1:
[0910] The device boots up and uses the built-in camera and microphone to scan the surroundings and identify the user. Once the user is identified, the device greets them with, "Good morning, [username]. How are you today?"
[0911] Step 2:
[0912] The user responds, "Good morning, I feel a bit heavy-headed today." The device collects the voice data and converts it into text data using a voice recognition module.
[0913] Step 3:
[0914] The converted text data is sent to a server, which then analyzes it using a natural language processing engine (NLP).The results of this analysis are used to evaluate the user's emotions and health status.
[0915] Step 4:
[0916] The server generates the next question based on the analysis results, for example, "Have you been sleeping well recently?", and sends this question to the device.
[0917] Step 5:
[0918] The device asks the user, "Have you been sleeping well lately?" While waiting for the user's response, the device instructs the user to measure their temperature and blood pressure. It adds, "Please measure your temperature and blood pressure."
[0919] Step 6:
[0920] The user uses a thermometer or blood pressure monitor to report the measurement results, such as "My body temperature is 36.5 degrees and my blood pressure is 120 / 80." The device's voice recognition module converts this voice back into text data, which is then sent to the server.
[0921] Step 7:
[0922] The server combines the health check data and text data and uses an AI algorithm to assess dementia risk, for example, based on information such as forgetfulness in recent conversations or abnormalities in physical data.
[0923] Step 8:
[0924] If the risk is high based on the assessment results, the server generates an alert message and sends it to the device, such as "There are some concerns about your health. We recommend that you consult a doctor."
[0925] Step 9:
[0926] The device will then provide the user with an alert message, saying, "After looking at the results, we've noticed some concerns about your recent health and conversations. We recommend that you consult a doctor."
[0927] Step 10:
[0928] The server stores all conversation and health data in a database, and periodically prepares data to monitor the user's condition in preparation for the next session.
[0929] Example 1
[0930] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0931] The present invention relates to a robotic system for monitoring the health status and dementia risk of elderly people living alone. Many current health monitoring systems require periodic manual input, which can be a burden for users. Furthermore, conventional systems do not perform detailed analysis of the user's emotions and health status, which can delay appropriate medical treatment when needed. Therefore, the present invention aims to provide a system that can continuously and automatically monitor the user's health status and promptly prompt appropriate medical treatment.
[0932] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0933] In this invention, the server includes means for recognizing a user, means for initiating a voice conversation with the user and converting the voice data into text data, means for analyzing the converted text data and evaluating the user's emotions and health status, means for generating the next question based on the evaluation results and presenting it to the user, means for collecting the user's physical data, means for integrating the voice and physical data and evaluating dementia risk using AI, means for providing feedback to the user based on the evaluation results, means for saving the data and preparing for the next session, and means for image recognition for identifying the user's face and voice recognition for recognizing the user's voice. This makes it possible to continuously and automatically monitor the user's health status in detail while minimizing the burden on the user, and to promote early and appropriate medical treatment.
[0934] "Means for recognizing the user" refers to technology that enables the robot to identify the user's face or specific features using cameras or sensors.
[0935] The "means for initiating a voice conversation and converting voice data into text data" is a voice recognition module for recognizing the voice spoken by the user and converting the voice into text information.
[0936] The "means for analyzing text data and assessing the user's emotions and health condition" refers to an algorithm or program that uses a natural language processing engine to analyze the converted text data and analyze the user's emotions and health condition.
[0937] The "means for generating the next question and presenting it to the user" refers to a device or software that generates a question to further evaluate the user's health condition based on the analysis results and presents it to the user via voice or screen display.
[0938] "Means for collecting user's physical data" refers to a method or device for collecting data on the user's physical condition using measuring instruments such as a thermometer or blood pressure monitor.
[0939] "Means for integrating voice and physical data and assessing dementia risk using AI" refers to a system that integrates and analyzes collected voice data and physical data, and assesses dementia risk using an AI algorithm.
[0940] The "means for providing feedback to the user based on the evaluation results" refers to a method or device that communicates the results of the analysis and evaluation to the user and provides a notice or warning that encourages the user to consult a medical institution if necessary.
[0941] The "means for storing the data and preparing for the next session" refers to a mechanism for recording and storing all data, including voice data and health data, to facilitate the next health monitoring session.
[0942] "Image recognition means" is a technology for analyzing facial features from images captured by a camera and identifying users.
[0943] "Speech recognition means" is a technology for capturing a user's voice and converting it into text data.
[0944] This invention relates to a robotic system for monitoring the health status and dementia risk of elderly people living alone. The system is equipped with functions for user recognition, voice conversation, data analysis, and feedback.
[0945] First, when the device (robot) starts up, it recognizes the user using the built-in camera and voice recognition function. The device greets the user by saying, "Good morning, [user name]. How are you today?" At this time, it uses the camera to identify the face and activates the voice recognition module. As a specific example, the device uses the built-in camera and voice recognition module to identify the user's face and voice.
[0946] Next, if the user responds, "I feel a bit heavy-headed today," this voice data is converted into text data by a voice recognition module. When the terminal converts voice data into text data, it uses a voice recognition module. For example, if the user says, "I'm feeling a bit unwell today," this is recorded as text data.
[0947] The converted text data is sent to a server and analyzed by a natural language processing engine (NLP). This analysis evaluates the user's emotions and health status. For example, the expression "my head feels heavy" can be used to infer a decline in health. As a specific example, the NLP engine analyzes the text "my head feels heavy" and determines that this is a sign of deteriorating health.
[0948] Next, the server generates the next question based on the analysis result and sends it to the terminal. For example, the question "Have you been sleeping well recently?" is generated. The terminal presents this question to the user. An example of a specific prompt sentence in this case is "Have you been sleeping well recently?"
[0949] At the same time, the device instructs the user to measure their temperature and blood pressure. The user uses a thermometer and blood pressure monitor to measure their temperature and report the results to the robot. For example, they might say, "My temperature is 36.5 degrees and my blood pressure is 120 / 80." Once the user reports the measurement results, the voice recognition module converts them back into text data and sends it to the server.
[0950] All data is integrated on a server, and an AI algorithm evaluates the risk of dementia. Risk is assessed based on abnormalities in language usage patterns and physical data. For example, if a person has become increasingly forgetful in recent conversations, this could be assessed as a risk of dementia.
[0951] If the risk is high, the server generates an alert message and sends it to the device, which then provides feedback such as, "After looking at the results, there are some concerns about your recent health condition and conversations. We recommend that you consult a doctor."
[0952] All conversation and health data is stored in a database on the server and ready for the next session, allowing for continuous, automatic, and detailed monitoring of the user's health status, facilitating early and appropriate medical treatment.
[0953] To give an example, here's what a typical day might look like:
[0954] 1. The device starts up, recognizes the user, and greets them with, "Good morning, Tanaka-san. How are you today?"
[0955] 2. The user replies, "Good morning, I'm feeling a bit down today."
[0956] 3. The device converts the voice data into text and sends it to the server.
[0957] 4. The server analyzes the text data, generates a question such as "Have you often felt heavy-headed lately?" and sends it to the device.
[0958] 5. The device asks questions and prompts the user to "take your temperature and blood pressure."
[0959] 6. The user reports the measurement results and replies, "Temperature is 36.5°C, blood pressure is 120 / 80."
[0960] 7. The device records this as text data and sends it to the server.
[0961] 8. The server analyzes all data and evaluates the risk of dementia. If the risk is high, an alert is sent to the device.
[0962] 9. The device will provide feedback such as, "After looking at the results, there are some concerns about your recent health condition and conversations. We recommend that you consult a doctor."
[0963] This system allows for daily monitoring of the health status of elderly people and allows for early linkage with medical institutions.
[0964] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0965] Step 1:
[0966] The device (robot) starts up. Once started, it uses the built-in camera and voice recognition function to recognize the user. Specifically, the camera captures the user's face and identifies the user using a facial recognition algorithm. Once the user is identified, the device greets them with "Good morning, [user name]. How are you today?"
[0967] Input: Start signal, camera image
[0968] Output: User recognition result, greeting message
[0969] Step 2:
[0970] The user responds by saying, "I feel a bit heavy-headed today." This voice data is captured by the device's voice recognition module and converted into text data. Specifically, voice input is captured by a microphone and then converted into text by the voice recognition module.
[0971] Input: User voice
[0972] Output: Text data
[0973] Step 3:
[0974] The device sends the converted text data to the server, where it packages the text data into a standard format such as JSON and securely transmits it to the server through a network module.
[0975] Input: Text data
[0976] Output: Server sent data
[0977] Step 4:
[0978] The server receives the text data and analyzes it using a natural language processing engine (NLP). This analysis evaluates the user's emotions and health status. Specifically, the NLP engine analyzes the text data and extracts keywords and emotions. At this stage, a decline in the user's health is inferred from the user's statement that "my head feels heavy."
[0979] Input: Text data
[0980] Output: Analysis results (emotion and health evaluation)
[0981] Step 5:
[0982] The server generates the next question based on the analysis results and sends it to the device. For example, a question might be generated such as, "Have you been sleeping well lately?" The server then creates an appropriate question based on the analysis results and sends it to the device as a data packet.
[0983] Input: Analysis results
[0984] Output: Next question data
[0985] Step 6:
[0986] The device then presents the generated questions to the user and instructs them to measure their health data. Specifically, it uses a voice synthesis module to prompt the user to "measure their temperature and blood pressure."
[0987] Input: Next question data
[0988] Output: Voice prompts and measurement instructions
[0989] Step 7:
[0990] The user measures their temperature and blood pressure using a thermometer and reports the results to the robot, for example, "My temperature is 36.5 degrees and my blood pressure is 120 / 80." This triggers the voice recognition module to convert the voice data into text.
[0991] Input: Measurement data (body temperature, blood pressure)
[0992] Output: Text data
[0993] Step 8:
[0994] The device sends the recorded health data as text data to the server.
[0995] Input: Text data
[0996] Output: Server sent data
[0997] Step 9:
[0998] The server integrates all the data and uses an AI algorithm to assess the risk of dementia. The AI algorithm analyzes language usage patterns and abnormalities in physical data to assess risk. For example, if the person has become increasingly forgetful in recent conversations, this could be assessed as a risk of dementia.
[0999] Input: Integrated data (voice and health)
[1000] Output: Risk assessment results
[1001] Step 10:
[1002] The server generates an alert message based on the risk assessment results and sends it to the device, which then provides feedback such as, "After looking at the results, there are some concerns about your recent health condition and conversations. We recommend that you consult a doctor."
[1003] Input: Risk assessment results
[1004] Output: Feedback message
[1005] Step 11:
[1006] The server stores all conversation and health data in a database and prepares for the next session, allowing for continuous monitoring of the user's health.
[1007] Input: Voice data, health data
[1008] Output: Save data, prepare for next session
[1009] (Application example 1)
[1010] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1011] It is extremely important to monitor the health status of elderly people living alone at home or in facilities and to detect dementia risk early. However, current systems have difficulty comprehensively and continuously monitoring users' health status and emotions, making early detection and prompt feedback difficult. For this reason, there is a need for a system that provides more comprehensive and accurate monitoring and feedback through devices or smartphone applications installed in facilities that provide healthcare services for the elderly.
[1012] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1013] In this invention, the server includes a means for recognizing the user, a means for initiating a voice conversation with the user and converting the voice data into text data, a means for analyzing the converted text data and evaluating the user's emotions and health status, a means for generating the next question based on the evaluation results and presenting it to the user, and a means for measuring the user's physical data and collecting the measurement results. This enables a system including a means for integrating the voice and measurement data and assessing dementia risk using AI, a means for providing feedback to the user based on the evaluation results, and a means for storing the data and preparing for the next session. This enables more comprehensive and accurate health status monitoring and rapid feedback through terminals and smartphone applications installed in facilities providing services for the elderly.
[1014] "Means for recognizing the user" refers to technology that uses a camera and voice recognition technology to verify the user's personal information.
[1015] The "means for converting voice data into text data" refers to a technique that uses a voice recognition module to convert recorded voice into a corresponding text format.
[1016] "Means for analyzing converted text data and assessing the user's emotions and health status" refers to a technology that utilizes a natural language processing engine and a generative AI model to assess the user's emotions and health status from text data.
[1017] "Means for generating the next question and presenting it to the user" refers to a technology that generates a new question based on the analysis results and presents it to the user in voice or text.
[1018] The "means for collecting user's physical data" refers to a technology for collecting user's physical data such as body temperature and blood pressure using measuring devices such as a thermometer and a blood pressure monitor.
[1019] "Method for integrating voice and measurement data to assess dementia risk using AI" refers to a technology that integrates collected voice data and physical data and uses an AI algorithm to assess dementia risk.
[1020] "Means of providing feedback to users" refers to technology that provides the results of AI evaluation to users via voice or text.
[1021] The "means for saving the data and preparing for the next session" refers to a technique for saving all collected data in a database and using it as the basis for future sessions.
[1022] "Means for conducting voice conversations with the user and collecting health data using a terminal installed in a facility that provides services for the elderly" refers to technology for conducting voice conversations and collecting health data using a device installed within the facility.
[1023] "Means for interacting with the user and managing health data through a smartphone application" refers to technology that uses a smartphone application to exchange information with the user and collect and manage health data.
[1024] "Means of sending data to a cloud environment and analyzing the user's emotions and health status using a natural language processing engine and generative AI model" refers to a technology that sends collected data to the cloud via the internet and analyzes it using a natural language processing engine and generative AI model.
[1025] "Means for providing alerts regarding health status on facility devices and smartphone applications based on evaluation results" refers to technology that provides alerts to users on facility devices and smartphone applications based on the results of AI evaluation.
[1026] This invention is a system for monitoring a user's health status and assessing their dementia risk. The system consists of a terminal installed in a facility, a smartphone application, a cloud server, and multiple software modules for linking these components.
[1027] System Configuration
[1028] 1. Terminal installation location
[1029] In this invention, a robot terminal is installed in a facility that provides services for the elderly. The terminal is equipped with a camera and voice recognition function, and can interact with users. The terminal recognizes the user's voice and facial expressions and identifies personal information.
[1030] 2. Smartphone Applications
[1031] Users can also receive the same services outside the facility using a smartphone application. The application has functions for voice conversations with users, recording health data, and providing consultations. The app synchronizes with the robot terminal and transmits data to a cloud server.
[1032] 3. Cloud environment
[1033] Data analysis and storage are performed in a cloud environment: collected voice and body data are sent to the cloud and processed by natural language processing engines (e.g., SpaCy or Google Cloud Natural Language API) and generative AI models.
[1034] Hardware and Software Details
[1035] Hardware
[1036] Camera: Built into the robot terminal and used to recognize the user's face.
[1037] Microphone: Used to collect voice conversation data.
[1038] Thermometers, blood pressure monitors: Health monitoring devices for measuring users' temperature and blood pressure. These are installed within the facility.
[1039] software
[1040] Speech recognition module: Technology for converting voice data into text data (e.g., Google Cloud Speech-to-Text).
[1041] Natural language processing engines and generative AI models: Analyze users' emotions and health status, and assess dementia risk (e.g., SpaCy, Google Cloud Natural Language API).
[1042] Cloud storage: Infrastructure for storing and managing data.
[1043] Detailed explanation of the process
[1044] 1. User Awareness
[1045] The device uses a camera and voice recognition technology to identify the user. When the user enters the facility, the device greets them with, "Good morning, how are you feeling today?"
[1046] 2. Audio data collection and analysis
[1047] When the user responds, the voice data is converted into text data using a voice recognition module. This data is sent to a cloud server and analyzed using a natural language processing engine. Based on the analysis results, the user's health status and emotions are evaluated.
[1048] 3. Next question generation and presentation
[1049] The next question is generated from the text data by the generative AI model and presented to the user via their device or smartphone app, such as "Have you been sleeping well lately?"
[1050] 4. Collection of health data
[1051] Physical data such as users' temperature and blood pressure will be collected using measuring equipment within the facility, and similar data can also be collected via a smartphone app.
[1052] 5. Data Storage and Feedback
[1053] The collected voice and physical data is stored in cloud storage. Based on the analysis results, feedback is provided to the user. For example, the system may provide feedback such as, "After looking at the results, there are some concerns about your recent physical condition and conversations. Please consult a doctor."
[1054] Specific examples
[1055] 1. Usage scenarios in facilities
[1056] When an elderly person visits the facility, the robot terminal recognizes them and asks, "Good morning, how are you feeling today?" If the elderly person replies, "I have a slight headache," the voice is converted into text data and sent to the cloud.
[1057] 2. Smartphone usage scenarios
[1058] The same elderly person opens the smartphone app at home and is asked the same questions about their health. The collected data is stored in the cloud and used the next time they visit the facility. Based on this data, a generative AI model generates appropriate questions and performs a risk assessment.
[1059] 3. Examples of prompts
[1060] Please analyze the following sentences and assess the user's health and emotions.
[1061] "My head feels a little heavy today."
[1062] This system can continuously monitor the user's health status wherever they are and provide quick feedback, contributing to the early detection of dementia risk.
[1063] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1064] Step 1:
[1065] When a user arrives at a facility, the device uses a camera and voice recognition technology to recognize the user. The camera captures the user's face, and the voice recognition module identifies the user's personal information from the voice data. Specifically, the camera captures the user's image and receives a greeting as voice data. The device outputs a voice message saying, "Good morning, how are you feeling today?"
[1066] Step 2:
[1067] The user responds to the terminal by voice. The terminal captures the voice data and converts it into text data using a voice recognition module. The user's voice input is "I feel a bit heavy-headed today," which is converted into text data. This text data is sent to the server.
[1068] Step 3:
[1069] The server analyzes the received text data and uses a natural language processing engine (e.g., SpaCy or Google Cloud Natural Language API) to extract data from the text data to evaluate the user's emotions and health status. The input is the text data "I feel a bit heavy-headed today," and the output is "The user is feeling unwell."
[1070] Step 4:
[1071] The server uses the generative AI model based on the analysis results to generate the next question. For example, the question "Have you been sleeping well lately?" is generated. This generated question data is sent to the device. Specifically, the text data of the generated question is sent to the device.
[1072] Step 5:
[1073] The device presents the generated question to the user by voice or text. The user then responds to the device. This response data is also converted into text data using a voice recognition module and sent to the server. The specific operation is that the device outputs the question "Have you been sleeping well lately?" by voice.
[1074] Step 6:
[1075] The terminal instructs the user to measure their physical data using a thermometer or blood pressure monitor. The user measures their body temperature and blood pressure using health measurement equipment installed in the facility and reports the results to the terminal. For example, the user reports, "My body temperature is 36.5 degrees and my blood pressure is 120 / 80." This data is converted into text data and sent to the server.
[1076] Step 7:
[1077] The server integrates all the received data and uses an AI algorithm to assess the risk of dementia. The voice data and physical data are analyzed as a single data set to assess the user's risk. The inputs are voice-text data such as "Head feels heavy," body temperature of "36.5 degrees," and blood pressure of "120 / 80," and the output is the "dementia risk assessment result." The specific operation is for the AI algorithm to analyze the data and perform a risk assessment.
[1078] Step 8:
[1079] The server provides feedback to the device and smartphone app based on the evaluation results. For example, it generates an alert message saying, "After looking at the results, there are some concerns about your recent health condition and conversations. Please consult a doctor." and sends it to the device and smartphone app. The user receives the alert message.
[1080] Step 9:
[1081] All data is stored in cloud storage and will be available for the next session. Along with regular data monitoring, it will be possible to continuously track the user's health status. Specifically, the collected voice and physical data will be stored in the cloud and used for future analysis.
[1082] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1083] This invention relates to a robotic system for monitoring the health status and diagnosing dementia in the home of elderly people living alone. The system is equipped with functions for user recognition, voice conversation, data analysis, emotion engine, and feedback.
[1084] First, when the device (robot) starts up, it recognizes the user using the built-in camera and voice recognition function. At this point, the robot greets the user by saying, "Good morning, [user name]. How are you today?"
[1085] The user then responds to the robot by saying something like, "I feel a bit heavy-headed today." This voice data is converted into text data by the device's voice recognition module. At the same time, the device acquires the user's voice and facial expression data and analyzes them using an emotion engine.
[1086] The converted text data and the evaluation data from the emotion engine are sent to a server and analyzed by a natural language processing engine (NLP). This analysis extracts information to evaluate the user's emotions and health status. For example, it can infer a decline in health status from expressions such as "my head feels heavy" and recognize emotions such as "I'm tired."
[1087] The server then generates the next question based on the analysis results and sends it to the device. For example, the question is "Have you been sleeping well lately?", taking into account emotional changes in the user's facial expressions. This allows for more personalized interactions.
[1088] The device will ask the user this question. At the same time, the device will ask the user to take their temperature and blood pressure. It will add, "Please take your temperature and blood pressure."
[1089] The user measures their temperature using a thermometer or blood pressure monitor at home and reports the results to the robot, for example, "My temperature is 36.5 degrees and my blood pressure is 120 / 80."
[1090] Health data, text data, and emotional data are integrated on a server, and an AI algorithm is used to assess dementia risk. The AI assesses risk based on abnormalities in language usage patterns and physical data, while also taking emotional data into account.
[1091] For example, if the user has become increasingly forgetful in recent conversations or has frequently shown emotionally unstable facial expressions, the server will generate an alert message and send it to the device, which will then provide feedback such as, "After looking at the results, there are some concerns about your recent physical condition and emotions. We recommend that you consult a doctor."
[1092] All conversation, health and emotion data is stored in a database on the server, allowing for preparation for the next session and regular monitoring of the user's condition.
[1093] As a concrete example, the flow of one day is shown below.
[1094] 1. The device starts up, recognizes the user, and greets them with, "Good morning, Tanaka-san. How are you today?"
[1095] 2. The user replies, "Good morning, I'm feeling a bit down today."
[1096] 3. The device collects voice and facial expression data, converts the voice into text, and analyzes it using an emotion engine.
[1097] 4. The text data and emotional evaluation data sent to the server are analyzed, and the question "Have you often felt heavy-headed lately?" is generated and sent to the device.
[1098] 5. The device asks questions and prompts the user to measure their health data. It prompts the user to "measure their temperature and blood pressure."
[1099] 6. The user reports the measurement results and replies, "Temperature is 36.5°C, blood pressure is 120 / 80."
[1100] 7. The device records this as text data and emotion data and sends it to the server.
[1101] 8. The server analyzes all data and evaluates the risk of dementia. If the risk is high, an alert is sent to the device.
[1102] 9. The device will provide feedback such as, "After looking at the results, there are some concerns about your recent physical condition and emotions. We recommend that you consult a doctor."
[1103] This system allows for daily monitoring of the health and emotional state of the elderly and allows for early consultation with medical institutions when necessary, thereby improving the peace of mind and quality of life for the elderly and their families.
[1104] The processing flow will be explained below.
[1105] Step 1:
[1106] The device boots up and uses the built-in camera and microphone to scan the surroundings and identify the user. Once the user is identified, the device greets them with, "Good morning, [username]. How are you today?"
[1107] Step 2:
[1108] The user responds, "Good morning, I feel a bit heavy-headed today." The device collects the voice data and converts it into text data using a voice recognition module.
[1109] Step 3:
[1110] The device collects facial expression data along with the captured voice data and analyzes it with an emotion engine, where the user's emotional state is identified as "fatigue."
[1111] Step 4:
[1112] The converted text data and the evaluation data generated by the emotion engine are sent to the server, which then analyzes the received text data using a natural language processing engine (NLP).The analysis results are used to evaluate the user's emotions and health status.
[1113] Step 5:
[1114] The server generates the next question based on the analysis results, for example, "Have you been sleeping well recently?", and sends this question to the device.
[1115] Step 6:
[1116] The device asks the user, "Have you been sleeping well lately?" When asking the question, it also takes into account changes in the user's facial expressions based on their emotions. At the same time, the device instructs the user to measure their temperature and blood pressure. It adds, "Please measure your temperature and blood pressure."
[1117] Step 7:
[1118] The user uses a thermometer or blood pressure monitor to report the measurement results, such as "My body temperature is 36.5 degrees and my blood pressure is 120 / 80." The device's voice recognition module converts this voice back into text data, which is then sent to the server.
[1119] Step 8:
[1120] The server combines the health check data, text data, and emotional data, and uses an AI algorithm to assess dementia risk. For example, if a person has recently shown forgetfulness and emotional instability in their facial expressions, it will determine that they are at high risk.
[1121] Step 9:
[1122] If the risk is high based on the assessment results, the server generates an alert message and sends it to the device, such as "There are some concerns about your health. We recommend that you consult a doctor."
[1123] Step 10:
[1124] The device will then provide the user with an alert message, saying, "After looking at your results, there are some concerns about your recent physical condition and emotions. We suggest you consult a doctor."
[1125] Step 11:
[1126] The server stores all conversation data, health data, and emotion data in a database. It periodically prepares data to monitor the user's condition in preparation for the next session.
[1127] Example 2
[1128] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1129] In modern society, there is a need to routinely monitor the health status and dementia risk of elderly people living alone and detect abnormalities early. However, the means to do so are limited, and there are no systems that use interactive systems to closely observe the health and emotional state of elderly people and provide appropriate feedback. Therefore, there is a need for an efficient and effective method of managing the health of elderly people.
[1130] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for recognizing the user, means for initiating a voice conversation with the user and converting the voice data into text data, means for analyzing the converted text data and evaluating the user's emotions and health status, means for generating the next question based on the evaluation results and presenting it to the user, means for measuring the user's physical data and collecting the measurement results, means for integrating the voice and measurement data and evaluating dementia risk using AI, means for providing feedback to the user based on the evaluation results, and means for saving the data and preparing for the next session. This makes it possible to monitor the health and emotional status of elderly people on a daily basis and to cooperate with medical institutions early if necessary.
[1131] "Means for recognizing a user" refers to the system's ability to identify and recognize the individual user using a camera or voice recognition technology.
[1132] The "means for starting a voice conversation and converting voice data into text data" is a technology for automatically converting voice data acquired through a voice dialogue with a user into text format.
[1133] The "means for analyzing the converted text data and evaluating the user's emotions and health state" is a technology that processes the text data, analyzes its contents, and estimates the user's current emotions and health state.
[1134] The "means for generating the next question and presenting it to the user" is a function for generating and presenting an appropriate next question to the user based on the analysis results.
[1135] "Means for measuring the user's physical data and collecting the measurement results" refers to technology that allows the user to measure physical data such as body temperature and blood pressure and for the system to collect the results.
[1136] "Means for integrating voice and measurement data and assessing dementia risk using AI" refers to a technology that integrates collected voice data and physical data and uses artificial intelligence technology to assess a user's dementia risk.
[1137] The "means for providing feedback to the user based on the evaluation results" is a function for providing the results of the system's evaluation to the user in the form of information or advice in an appropriate format.
[1138] "Means for storing data and preparing for the next session" refers to technology for recording collected and analyzed data and preparing it for use in the next interaction or monitoring.
[1139] The present invention relates to a robotic system for monitoring the health status and dementia risk of elderly people living alone and for detecting abnormalities at an early stage. This system mainly includes a means for recognizing a user, a means for performing voice conversation and data analysis, an emotion analysis engine, a feedback means, and a storage means.
[1140] When the device (robot) starts up, it uses a camera and voice recognition functions to recognize the user. Specifically, it takes a picture of the user's face with the camera and recognizes the user's voice using voice recognition technology. At this time, the device says, "Good morning, [user name]. How are you today?"
[1141] When the user responds, the voice data is converted into text data using the device's voice recognition module. This converted text data and facial expression data captured by the camera are analyzed by the emotion engine. For example, by analyzing a statement such as "I feel a bit heavy-headed today" and the user's facial expression at the time, the user's emotions and health condition can be evaluated.
[1142] The device then sends this data to a server, which incorporates a natural language processing engine (NLP) and performs detailed analysis of the received text data and emotional evaluation data. For example, an expression such as "My head feels heavy" is interpreted as a sign of the user's stress or poor health. Emotional fluctuations are also extracted from facial expression data.
[1143] Based on the analysis results, the server generates the next question. For example, "Have you been sleeping well recently?" and sends it to the device. The system then instructs the user to measure their temperature and blood pressure. The system prompts the user to "measure their temperature and blood pressure."
[1144] The user measures their temperature using a thermometer or blood pressure monitor at home and reports the results to the robot. For example, they might report, "My temperature is 36.5 degrees and my blood pressure is 120 / 80." This health data is collected as text data and sent back to the server.
[1145] The server uses an AI algorithm to comprehensively analyze this data. It evaluates the user's dementia risk by combining voice, emotional, and physical data. For example, if the user has recently shown an increase in forgetfulness and emotional instability in their facial expressions, the server will assess the risk as high and generate an alert message.
[1146] Finally, the device provides the generated feedback to the user, for example, saying, "After looking at the results, there are some concerns about your recent physical condition and emotions. We recommend that you consult a doctor." This makes it possible to monitor the health status of elderly people on a daily basis and to consult with medical institutions early if necessary.
[1147] All conversation, health, and emotion data is stored on the server and prepared for the next session. Through regular monitoring, the health status of the elderly can be continuously managed.
[1148] (Example of a prompt)
[1149] Here are some example prompts for a generative AI model:
[1150] "Design a robotic system to check the health status of elderly people daily. This system should have the ability to recognize the user's face and voice, analyze their health status and emotions, and provide feedback as needed. It should also send the measured health data and analyzed emotional data to a server, enabling regular monitoring."
[1151] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1152] (Program processing steps)
[1153] Step 1:
[1154] The device starts up.
[1155] Input: System power supply
[1156] Output: The terminal is ready
[1157] At startup, the built-in camera and voice recognition function are enabled to prepare for user recognition.
[1158] Step 2:
[1159] The device uses its built-in camera and voice recognition capabilities to identify the user.
[1160] Input: User's face, voice
[1161] Output: Recognized user information
[1162] The camera captures the user's face and uses voice recognition technology to identify the individual from their voice.
[1163] The device greets you with, "Good morning, [username]. How are you today?"
[1164] Step 3:
[1165] The user responds, "I'm feeling a bit heavy-headed today."
[1166] Input: User's voice response
[1167] Output: User's voice data
[1168] The terminal records the user's response.
[1169] Step 4:
[1170] The terminal uses a voice recognition module to convert the user's voice data into text data.
[1171] Input: User's voice data
[1172] Output: Converted text data
[1173] The voice recognition module is activated to convert the voice data into text format.
[1174] Step 5:
[1175] The device collects the user's facial expression data captured by the camera and analyzes it using an emotion engine.
[1176] Input: User's facial expression data
[1177] Output: User's emotion rating data
[1178] A camera captures the user's facial expressions and uses an emotion analysis algorithm to assess their emotional state.
[1179] Step 6:
[1180] The terminal transmits the converted text data and emotion evaluation data to the server.
[1181] Input: Text data, emotion rating data
[1182] Output: Data sent to the server
[1183] The terminal uploads the data to the server via the network.
[1184] Step 7:
[1185] The server uses a natural language processing engine (NLP) to analyze the text data and assess the user's health status.
[1186] Input: Text data, emotion rating data
[1187] Output: User's health status rating
[1188] Using NLP technology, the system analyzes the converted text and, for example, infers a decline in the user's health from an expression like "my head feels heavy." It also takes into account emotional data.
[1189] Step 8:
[1190] The server generates the next question based on the analysis results.
[1191] Input: User's health assessment
[1192] Output: Next question (e.g., "Have you been sleeping well lately?")
[1193] Based on the analysis results, appropriate follow-up questions are automatically generated.
[1194] Step 9:
[1195] The server generates the next question and sends it to the terminal.
[1196] Input: Next question
[1197] Output: Questions sent to the terminal
[1198] The server transmits question data to the terminal via the network.
[1199] Step 10:
[1200] The device asks the user the following question and instructs them to "take your temperature and blood pressure."
[1201] Input: Next question, measurement instructions
[1202] Output: Questions and measurement instructions for the user
[1203] The device will then ask the user the next question via voice and display, and instruct them to measure further health data.
[1204] Step 11:
[1205] The user measures their temperature using a thermometer or blood pressure monitor and reports the results to the robot.
[1206] Input: User measurements
[1207] Output: Reported measurement data (e.g., "Temperature is 36.5°C, Blood pressure is 120 / 80")
[1208] The user follows the instructions to measure their own body temperature and blood pressure and reports the results to the device.
[1209] Step 12:
[1210] The device records the reported health data as text data and sends it to the server.
[1211] Input: Reported measurement data
[1212] Output: Health data sent to the server
[1213] The terminal converts the measurement data into text format and uploads it to a server via the network.
[1214] Step 13:
[1215] The server comprehensively analyzes health data, text data, and emotional data to assess dementia risk.
[1216] Input: Health data, text data, emotion data
[1217] Output: Dementia risk assessment
[1218] An AI algorithm is used to integrate all data and assess dementia risk.
[1219] Step 14:
[1220] If the risk is high, the server generates an alert message and sends it to the terminal.
[1221] Input: Dementia Risk Assessment
[1222] Output: Alert message
[1223] If the risk is determined to be high, the server generates an alert message and sends it to the terminal.
[1224] Step 15:
[1225] The terminal provides feedback to the user.
[1226] Input: Alert message
[1227] Output: Feedback to the user (e.g., "You've noticed some concerns about your physical and emotional state recently. Consult your doctor.")
[1228] The terminal displays an alert message to the user and provides appropriate advice.
[1229] Step 16:
[1230] The server stores all conversation data, health data, and emotion data.
[1231] Input: Conversation data, health data, emotion data
[1232] Output: Saved data
[1233] The server saves all the data in a database ready for the next session.
[1234] (Application example 2)
[1235] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1236] Monitoring the health and emotional state of elderly workers in factories is important, but conventional systems have difficulty grasping the situation and assessing risks in real time, making it difficult to respond effectively. Early response to prevent accidents and illnesses is also often delayed. To solve these problems, a system is needed that can comprehensively analyze workers' health and emotional data and assess risks early.
[1237] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recognizing a user, means for initiating a voice conversation with the user using a voice recognition module and converting the voice data into text data, means including a natural language processing engine for analyzing the converted text data and emotional data and evaluating the user's emotions and health status, means for generating the next question based on the evaluation results and presenting it to the user, means for measuring the user's physical data and collecting the measurement results, means for integrating the voice and measurement data and evaluating health risks using AI, means for providing feedback to the user based on the evaluation results, means for transmitting the health data and emotional data to the server and generating feedback based on the analysis results, and means for storing the data and preparing for the next session. This makes it possible to comprehensively grasp the health and emotional status of workers in real time, evaluate risks early, and provide appropriate feedback and countermeasures.
[1238] "Means for recognizing a user" refers to technology or devices that use cameras or sensors to identify a specific user.
[1239] "Means for initiating a voice conversation and converting voice data into text data" refers to a voice recognition technology or system that converts voice data acquired through a microphone into character string data.
[1240] "Means for analyzing converted text data and assessing a user's emotions and health status" refers to technology or systems that use natural language processing engines and other analytical algorithms to infer and assess a user's emotions and health status from their statements and expressions.
[1241] The "means for generating the next question and presenting it to the user" refers to a technology or system that automatically generates the next appropriate question based on the evaluation results and presents it to the user by voice or display.
[1242] "Means for measuring a user's physical data and collecting the measurement results" refers to a technology or system that uses devices such as a thermometer or blood pressure monitor to measure data related to a user's body and collect the results as data.
[1243] "Means for integrating voice and measurement data and assessing dementia risk using AI" refers to a technology or system that integrates acquired voice data and physical data and uses an AI algorithm to assess a user's dementia risk.
[1244] The "means for providing feedback to the user based on the evaluation results" refers to a technology or system that provides necessary advice or instructions to the user based on the analysis and evaluation results.
[1245] "Means for transmitting health data and emotional data to a server and generating feedback based on the analysis results" refers to a technology or system that transmits collected data to a server and provides appropriate feedback to the user based on the analysis results on the server side.
[1246] "Means for storing the data and preparing for the next session" refers to a technology or system that stores all acquired data in a database and prepares it for the next monitoring or evaluation based on that data.
[1247] The present invention relates to a system for real-time monitoring of the health and emotional states of factory workers and early risk assessment, which is equipped with user recognition, voice conversation, data analysis, emotion engine, and feedback functions.
[1248] Hardware and software used
[1249] Hardware:
[1250] 1. Camera: Used to recognize the face of the worker.
[1251] 2. Microphone: Used for voice recognition.
[1252] 3. Health measurement devices (thermometer, blood pressure monitor, heart rate monitor): Used to obtain physical data of workers.
[1253] software:
[1254] 1. OpenCV: Used to recognize the worker's face from the image data acquired through the camera.
[1255] 2. SpeechRecognition (Python library): Used to convert audio data into text data.
[1256] 3. Natural language processing engine: Analyzes the acquired text and emotion data and uses it to assess the emotions and health status of workers.
[1257] 4. Health Monitoring Library: Provides functions (get_temperature, get_blood_pressure, get_heart_rate) to get body temperature, blood pressure, and heart rate.
[1258] 5. Emotion Recognition Library: Used to analyze emotions from workers' facial expression data or voice data.
[1259] 6. Requests (Python library): Used to send data to and retrieve data from the server.
[1260] Processing flow
[1261] The system first recognizes the user (worker) using a camera and a voice recognition module. For example, when the system starts up, the camera scans the worker's face and greets them verbally, saying, "Good morning, [Worker's name]. How are you feeling?" If the worker replies, "I'm not feeling well today," the voice data is converted into text data by the voice recognition module.
[1262] The converted text data and acquired emotional data (analyzed from facial expressions and voice) are then analyzed using a natural language processing engine, allowing the worker's emotions and health status to be assessed.
[1263] Next, the system asks the worker to measure their body temperature, blood pressure, and heart rate. For example, data such as "body temperature is 37.5 degrees, blood pressure is 130 / 85, and heart rate is 80" is collected from the health measurement device. The collected data is sent to a server, where a health risk assessment is performed.
[1264] Based on the analysis and evaluation results, the system generates the next appropriate question to ask the worker and presents it to them. If the worker's health condition is not good, the system will provide feedback such as "You seem unwell. Please consult a doctor."
[1265] All data is stored in a database and prepared for the next session, allowing for a real-time understanding of the health and emotional state of factory workers, early risk assessment, and appropriate feedback and countermeasures.
[1266] Specific examples
[1267] Example prompt sentence:
[1268] Speech recognition: "I'm not feeling well today"
[1269] Facial Recognition: "Good morning, [Worker's Name]. How are you feeling?"
[1270] Health data collection: "Temperature is 37.5°C, blood pressure is 130 / 85, heart rate is 80."
[1271] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1272] Step 1:
[1273] The terminal starts up and uses the camera to recognize the user's (worker's) face. At this time, the terminal analyzes the image data acquired through the camera using OpenCV and checks whether it matches the facial data registered in advance. The input is the image data acquired from the camera, and the output is the user ID. In concrete terms, the camera scans the worker's face and recognizes matching facial data.
[1274] Step 2:
[1275] The device greets the recognized user by voice and asks about their current health condition. For example, it might ask, "Good morning, [user name]. How are you feeling?" The input is the user ID, and the output is a voice message. Specifically, the device incorporates the user name into the greeting and outputs it by voice.
[1276] Step 3:
[1277] The user responds by saying something like, "I'm not feeling well today." This voice data is acquired through the device's microphone. The input is the user's voice data, and the output is voice data. In concrete terms, the user reports their physical condition by voice into the device.
[1278] Step 4:
[1279] The device uses a speech recognition module to convert the acquired voice data into text data. The input is voice data and the output is text data. Specifically, the device converts the acquired voice data into text data using the SpeechRecognition library.
[1280] Step 5:
[1281] The converted text data and facial expression data are analyzed to evaluate the user's emotions and health status. The device analyzes the text data and facial expression data using a natural language processing engine and an emotion recognition library. The input is the text data and facial expression data, and the output is the evaluation results. Specifically, the device evaluates the emotions and health status based on the analysis results.
[1282] Step 6:
[1283] The device generates the next question based on the evaluation results and presents it to the user. For example, it generates a question such as "Have you been sleeping well recently?" The input is the evaluation results, and the output is the question text. As a specific operation, the generated question is output to the user by voice or display.
[1284] Step 7:
[1285] The device asks the user to measure their body temperature, blood pressure, and heart rate. For example, it instructs the user to "measure their body temperature and blood pressure." The input is the evaluation result, and the output is an instruction message. Specifically, the device gives instructions to measure by voice.
[1286] Step 8:
[1287] A user uses a health measurement device to measure their body temperature, blood pressure, and heart rate, and reports the results to a terminal. For example, they report, "My body temperature is 37.5 degrees, my blood pressure is 130 / 85, and my heart rate is 80." The input is the measurement data, and the output is the reported data. Specifically, the user inputs the measurement data into the terminal.
[1288] Step 9:
[1289] The device sends the collected voice data, health data, and emotion data to the server. The input is the collected data, and the output is the result sent to the server. Specifically, the device uses the Requests library to send data to the server.
[1290] Step 10:
[1291] The server analyzes the received data and evaluates health risks. The input is the received data, and the output is the evaluation result. Specifically, the server uses an AI algorithm to perform risk assessment.
[1292] Step 11:
[1293] The server generates feedback based on the evaluation results and sends it to the device. For example, it generates feedback such as "You seem unwell, please consult a doctor." The input is the evaluation results, and the output is a feedback message. The specific operation is to send the generated feedback to the device.
[1294] Step 12:
[1295] The terminal presents the feedback received from the server to the user. The input is the feedback message, and the output is the result presented to the user. Specifically, the terminal conveys the feedback to the user by voice.
[1296] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1297] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1298] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1299] [Fourth embodiment]
[1300] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1301] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1302] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1303] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1304] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1305] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1306] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1307] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1308] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1309] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1310] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1311] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1312] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1313] This invention relates to a robotic system for monitoring the health status and diagnosing dementia in the home of elderly people living alone. The system is equipped with functions for user recognition, voice conversation, data analysis, and feedback.
[1314] First, when the device (robot) starts up, it recognizes the user using the built-in camera and voice recognition function. At this point, the robot greets the user by saying, "Good morning, [user name]. How are you today?"
[1315] Next, the user responds to the robot by saying something like, "I feel a bit heavy-headed today." This voice data is converted into text data by a voice recognition module in the device.
[1316] The converted text data is sent to a server and analyzed by a natural language processing engine (NLP). This analysis extracts information to evaluate the user's emotions and health status. For example, an expression such as "my head feels heavy" can be used to infer a decline in health.
[1317] Next, the server generates the next question based on the analysis results and sends it to the device. For example, the question generated might be, "Have you been sleeping well lately?" The device then presents this question to the user.
[1318] At the same time, the device prompts the user to measure their temperature and blood pressure. The user measures their temperature and blood pressure using a thermometer or blood pressure monitor at home and reports the results to the robot. For example, "My temperature is 36.5 degrees and my blood pressure is 120 / 80."
[1319] Health data and text data are integrated on a server, and an AI algorithm is used to assess dementia risk, based on abnormalities in language usage patterns and physical data.
[1320] For example, if the user has become increasingly forgetful in recent conversations or if there are abnormalities in the user's physical data, the server will generate an alert message and send it to the device. The device will then provide feedback such as, "After looking at the results, there are some concerns about your recent physical condition and conversations. We recommend that you consult a doctor."
[1321] All conversation and health data is stored in a database on the server, which allows for preparation for the next session and regular monitoring of the user's condition.
[1322] As a concrete example, the flow of one day is shown below.
[1323] 1. The device starts up, recognizes the user, and greets them with, "Good morning, Tanaka-san. How are you today?"
[1324] 2. The user replies, "Good morning, I'm feeling a bit down today."
[1325] 3. The device converts the voice data into text and sends it to the server.
[1326] 4. The server analyzes the text data, generates a question such as "Have you often felt heavy-headed lately?" and sends it to the device.
[1327] 5. The device asks questions and prompts the user to measure their health data. It prompts the user to "measure their temperature and blood pressure."
[1328] 6. The user reports the measurement results and replies, "Temperature is 36.5°C, blood pressure is 120 / 80."
[1329] 7. The device records this as text data and sends it to the server.
[1330] 8. The server analyzes all data and evaluates the risk of dementia. If the risk is high, an alert is sent to the device.
[1331] 9. The device will provide feedback such as, "After looking at the results, there are some concerns about your recent health condition and conversations. We recommend that you consult a doctor."
[1332] This system will enable daily monitoring of the health status of elderly people and enable early contact with medical institutions if necessary.
[1333] The processing flow will be explained below.
[1334] Step 1:
[1335] The device boots up and uses the built-in camera and microphone to scan the surroundings and identify the user. Once the user is identified, the device greets them with, "Good morning, [username]. How are you today?"
[1336] Step 2:
[1337] The user responds, "Good morning, I feel a bit heavy-headed today." The device collects the voice data and converts it into text data using a voice recognition module.
[1338] Step 3:
[1339] The converted text data is sent to a server, which then analyzes it using a natural language processing engine (NLP).The results of this analysis are used to evaluate the user's emotions and health status.
[1340] Step 4:
[1341] The server generates the next question based on the analysis results, for example, "Have you been sleeping well recently?", and sends this question to the device.
[1342] Step 5:
[1343] The device asks the user, "Have you been sleeping well lately?" While waiting for the user's response, the device instructs the user to measure their temperature and blood pressure. It adds, "Please measure your temperature and blood pressure."
[1344] Step 6:
[1345] The user uses a thermometer or blood pressure monitor to report the measurement results, such as "My body temperature is 36.5 degrees and my blood pressure is 120 / 80." The device's voice recognition module converts this voice back into text data, which is then sent to the server.
[1346] Step 7:
[1347] The server combines the health check data and text data and uses an AI algorithm to assess dementia risk, for example, based on information such as forgetfulness in recent conversations or abnormalities in physical data.
[1348] Step 8:
[1349] If the risk is high based on the assessment results, the server generates an alert message and sends it to the device, such as "There are some concerns about your health. We recommend that you consult a doctor."
[1350] Step 9:
[1351] The device will then provide the user with an alert message, saying, "After looking at the results, we've noticed some concerns about your recent health and conversations. We recommend that you consult a doctor."
[1352] Step 10:
[1353] The server stores all conversation and health data in a database, and periodically prepares data to monitor the user's condition in preparation for the next session.
[1354] Example 1
[1355] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1356] The present invention relates to a robotic system for monitoring the health status and dementia risk of elderly people living alone. Many current health monitoring systems require periodic manual input, which can be a burden for users. Furthermore, conventional systems do not perform detailed analysis of the user's emotions and health status, which can delay appropriate medical treatment when needed. Therefore, the present invention aims to provide a system that can continuously and automatically monitor the user's health status and promptly prompt appropriate medical treatment.
[1357] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1358] In this invention, the server includes means for recognizing a user, means for initiating a voice conversation with the user and converting the voice data into text data, means for analyzing the converted text data and evaluating the user's emotions and health status, means for generating the next question based on the evaluation results and presenting it to the user, means for collecting the user's physical data, means for integrating the voice and physical data and evaluating dementia risk using AI, means for providing feedback to the user based on the evaluation results, means for saving the data and preparing for the next session, and means for image recognition for identifying the user's face and voice recognition for recognizing the user's voice. This makes it possible to continuously and automatically monitor the user's health status in detail while minimizing the burden on the user, and to promote early and appropriate medical treatment.
[1359] "Means for recognizing the user" refers to technology that enables the robot to identify the user's face or specific features using cameras or sensors.
[1360] The "means for initiating a voice conversation and converting voice data into text data" is a voice recognition module for recognizing the voice spoken by the user and converting the voice into text information.
[1361] The "means for analyzing text data and assessing the user's emotions and health condition" refers to an algorithm or program that uses a natural language processing engine to analyze the converted text data and analyze the user's emotions and health condition.
[1362] The "means for generating the next question and presenting it to the user" refers to a device or software that generates a question to further evaluate the user's health condition based on the analysis results and presents it to the user via voice or screen display.
[1363] "Means for collecting user's physical data" refers to a method or device for collecting data on the user's physical condition using measuring instruments such as a thermometer or blood pressure monitor.
[1364] "Means for integrating voice and physical data and assessing dementia risk using AI" refers to a system that integrates and analyzes collected voice data and physical data, and assesses dementia risk using an AI algorithm.
[1365] The "means for providing feedback to the user based on the evaluation results" refers to a method or device that communicates the results of the analysis and evaluation to the user and provides a notice or warning that encourages the user to consult a medical institution if necessary.
[1366] The "means for storing the data and preparing for the next session" refers to a mechanism for recording and storing all data, including voice data and health data, to facilitate the next health monitoring session.
[1367] "Image recognition means" is a technology for analyzing facial features from images captured by a camera and identifying users.
[1368] "Speech recognition means" is a technology for capturing a user's voice and converting it into text data.
[1369] This invention relates to a robotic system for monitoring the health status and dementia risk of elderly people living alone. The system is equipped with functions for user recognition, voice conversation, data analysis, and feedback.
[1370] First, when the device (robot) starts up, it recognizes the user using the built-in camera and voice recognition function. The device greets the user by saying, "Good morning, [user name]. How are you today?" At this time, it uses the camera to identify the face and activates the voice recognition module. As a specific example, the device uses the built-in camera and voice recognition module to identify the user's face and voice.
[1371] Next, if the user responds, "I feel a bit heavy-headed today," this voice data is converted into text data by a voice recognition module. When the terminal converts voice data into text data, it uses a voice recognition module. For example, if the user says, "I'm feeling a bit unwell today," this is recorded as text data.
[1372] The converted text data is sent to a server and analyzed by a natural language processing engine (NLP). This analysis evaluates the user's emotions and health status. For example, the expression "my head feels heavy" can be used to infer a decline in health. As a specific example, the NLP engine analyzes the text "my head feels heavy" and determines that this is a sign of deteriorating health.
[1373] Next, the server generates the next question based on the analysis result and sends it to the terminal. For example, the question "Have you been sleeping well recently?" is generated. The terminal presents this question to the user. An example of a specific prompt sentence in this case is "Have you been sleeping well recently?"
[1374] At the same time, the device instructs the user to measure their temperature and blood pressure. The user uses a thermometer and blood pressure monitor to measure their temperature and report the results to the robot. For example, they might say, "My temperature is 36.5 degrees and my blood pressure is 120 / 80." Once the user reports the measurement results, the voice recognition module converts them back into text data and sends it to the server.
[1375] All data is integrated on a server, and an AI algorithm evaluates the risk of dementia. Risk is assessed based on abnormalities in language usage patterns and physical data. For example, if a person has become increasingly forgetful in recent conversations, this could be assessed as a risk of dementia.
[1376] If the risk is high, the server generates an alert message and sends it to the device, which then provides feedback such as, "After looking at the results, there are some concerns about your recent health condition and conversations. We recommend that you consult a doctor."
[1377] All conversation and health data is stored in a database on the server and ready for the next session, allowing for continuous, automatic, and detailed monitoring of the user's health status, facilitating early and appropriate medical treatment.
[1378] To give an example, here's what a typical day might look like:
[1379] 1. The device starts up, recognizes the user, and greets them with, "Good morning, Tanaka-san. How are you today?"
[1380] 2. The user replies, "Good morning, I'm feeling a bit down today."
[1381] 3. The device converts the voice data into text and sends it to the server.
[1382] 4. The server analyzes the text data, generates a question such as "Have you often felt heavy-headed lately?" and sends it to the device.
[1383] 5. The device asks questions and prompts the user to "take your temperature and blood pressure."
[1384] 6. The user reports the measurement results and replies, "Temperature is 36.5°C, blood pressure is 120 / 80."
[1385] 7. The device records this as text data and sends it to the server.
[1386] 8. The server analyzes all data and evaluates the risk of dementia. If the risk is high, an alert is sent to the device.
[1387] 9. The device will provide feedback such as, "After looking at the results, there are some concerns about your recent health condition and conversations. We recommend that you consult a doctor."
[1388] This system allows for daily monitoring of the health status of elderly people and allows for early linkage with medical institutions.
[1389] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1390] Step 1:
[1391] The device (robot) starts up. Once started, it uses the built-in camera and voice recognition function to recognize the user. Specifically, the camera captures the user's face and identifies the user using a facial recognition algorithm. Once the user is identified, the device greets them with "Good morning, [user name]. How are you today?"
[1392] Input: Start signal, camera image
[1393] Output: User recognition result, greeting message
[1394] Step 2:
[1395] The user responds by saying, "I feel a bit heavy-headed today." This voice data is captured by the device's voice recognition module and converted into text data. Specifically, voice input is captured by a microphone and then converted into text by the voice recognition module.
[1396] Input: User voice
[1397] Output: Text data
[1398] Step 3:
[1399] The device sends the converted text data to the server, where it packages the text data into a standard format such as JSON and securely transmits it to the server through a network module.
[1400] Input: Text data
[1401] Output: Server sent data
[1402] Step 4:
[1403] The server receives the text data and analyzes it using a natural language processing engine (NLP). This analysis evaluates the user's emotions and health status. Specifically, the NLP engine analyzes the text data and extracts keywords and emotions. At this stage, a decline in the user's health is inferred from the user's statement that "my head feels heavy."
[1404] Input: Text data
[1405] Output: Analysis results (emotion and health evaluation)
[1406] Step 5:
[1407] The server generates the next question based on the analysis results and sends it to the device. For example, a question might be generated such as, "Have you been sleeping well lately?" The server then creates an appropriate question based on the analysis results and sends it to the device as a data packet.
[1408] Input: Analysis results
[1409] Output: Next question data
[1410] Step 6:
[1411] The device then presents the generated questions to the user and instructs them to measure their health data. Specifically, it uses a voice synthesis module to prompt the user to "measure their temperature and blood pressure."
[1412] Input: Next question data
[1413] Output: Voice prompts and measurement instructions
[1414] Step 7:
[1415] The user measures their temperature and blood pressure using a thermometer and reports the results to the robot, for example, "My temperature is 36.5 degrees and my blood pressure is 120 / 80." This triggers the voice recognition module to convert the voice data into text.
[1416] Input: Measurement data (body temperature, blood pressure)
[1417] Output: Text data
[1418] Step 8:
[1419] The device sends the recorded health data as text data to the server.
[1420] Input: Text data
[1421] Output: Server sent data
[1422] Step 9:
[1423] The server integrates all the data and uses an AI algorithm to assess the risk of dementia. The AI algorithm analyzes language usage patterns and abnormalities in physical data to assess risk. For example, if the person has become increasingly forgetful in recent conversations, this could be assessed as a risk of dementia.
[1424] Input: Integrated data (voice and health)
[1425] Output: Risk assessment results
[1426] Step 10:
[1427] The server generates an alert message based on the risk assessment results and sends it to the device, which then provides feedback such as, "After looking at the results, there are some concerns about your recent health condition and conversations. We recommend that you consult a doctor."
[1428] Input: Risk assessment results
[1429] Output: Feedback message
[1430] Step 11:
[1431] The server stores all conversation and health data in a database and prepares for the next session, allowing for continuous monitoring of the user's health.
[1432] Input: Voice data, health data
[1433] Output: Save data, prepare for next session
[1434] (Application example 1)
[1435] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1436] It is extremely important to monitor the health status of elderly people living alone at home or in facilities and to detect dementia risk early. However, current systems have difficulty comprehensively and continuously monitoring users' health status and emotions, making early detection and prompt feedback difficult. For this reason, there is a need for a system that provides more comprehensive and accurate monitoring and feedback through devices or smartphone applications installed in facilities that provide healthcare services for the elderly.
[1437] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1438] In this invention, the server includes a means for recognizing the user, a means for initiating a voice conversation with the user and converting the voice data into text data, a means for analyzing the converted text data and evaluating the user's emotions and health status, a means for generating the next question based on the evaluation results and presenting it to the user, and a means for measuring the user's physical data and collecting the measurement results. This enables a system including a means for integrating the voice and measurement data and assessing dementia risk using AI, a means for providing feedback to the user based on the evaluation results, and a means for storing the data and preparing for the next session. This enables more comprehensive and accurate health status monitoring and rapid feedback through terminals and smartphone applications installed in facilities providing services for the elderly.
[1439] "Means for recognizing the user" refers to technology that uses a camera and voice recognition technology to verify the user's personal information.
[1440] The "means for converting voice data into text data" refers to a technique that uses a voice recognition module to convert recorded voice into a corresponding text format.
[1441] "Means for analyzing converted text data and assessing the user's emotions and health status" refers to a technology that utilizes a natural language processing engine and a generative AI model to assess the user's emotions and health status from text data.
[1442] "Means for generating the next question and presenting it to the user" refers to a technology that generates a new question based on the analysis results and presents it to the user in voice or text.
[1443] The "means for collecting user's physical data" refers to a technology for collecting user's physical data such as body temperature and blood pressure using measuring devices such as a thermometer and a blood pressure monitor.
[1444] "Method for integrating voice and measurement data to assess dementia risk using AI" refers to a technology that integrates collected voice data and physical data and uses an AI algorithm to assess dementia risk.
[1445] "Means of providing feedback to users" refers to technology that provides the results of AI evaluation to users via voice or text.
[1446] The "means for saving the data and preparing for the next session" refers to a technique for saving all collected data in a database and using it as the basis for future sessions.
[1447] "Means for conducting voice conversations with the user and collecting health data using a terminal installed in a facility that provides services for the elderly" refers to technology for conducting voice conversations and collecting health data using a device installed within the facility.
[1448] "Means for interacting with the user and managing health data through a smartphone application" refers to technology that uses a smartphone application to exchange information with the user and collect and manage health data.
[1449] "Means of sending data to a cloud environment and analyzing the user's emotions and health status using a natural language processing engine and generative AI model" refers to a technology that sends collected data to the cloud via the internet and analyzes it using a natural language processing engine and generative AI model.
[1450] "Means for providing alerts regarding health status on facility devices and smartphone applications based on evaluation results" refers to technology that provides alerts to users on facility devices and smartphone applications based on the results of AI evaluation.
[1451] This invention is a system for monitoring a user's health status and assessing their dementia risk. The system consists of a terminal installed in a facility, a smartphone application, a cloud server, and multiple software modules for linking these components.
[1452] System Configuration
[1453] 1. Terminal installation location
[1454] In this invention, a robot terminal is installed in a facility that provides services for the elderly. The terminal is equipped with a camera and voice recognition function, and can interact with users. The terminal recognizes the user's voice and facial expressions and identifies personal information.
[1455] 2. Smartphone Applications
[1456] Users can also receive the same services outside the facility using a smartphone application. The application has functions for voice conversations with users, recording health data, and providing consultations. The app synchronizes with the robot terminal and transmits data to a cloud server.
[1457] 3. Cloud environment
[1458] Data analysis and storage are performed in a cloud environment: collected voice and body data are sent to the cloud and processed by natural language processing engines (e.g., SpaCy or Google Cloud Natural Language API) and generative AI models.
[1459] Hardware and Software Details
[1460] Hardware
[1461] Camera: Built into the robot terminal and used to recognize the user's face.
[1462] Microphone: Used to collect voice conversation data.
[1463] Thermometers, blood pressure monitors: Health monitoring devices for measuring users' temperature and blood pressure. These are installed within the facility.
[1464] software
[1465] Speech recognition module: Technology for converting voice data into text data (e.g., Google Cloud Speech-to-Text).
[1466] Natural language processing engines and generative AI models: Analyze users' emotions and health status, and assess dementia risk (e.g., SpaCy, Google Cloud Natural Language API).
[1467] Cloud storage: Infrastructure for storing and managing data.
[1468] Detailed explanation of the process
[1469] 1. User Awareness
[1470] The device uses a camera and voice recognition technology to identify the user. When the user enters the facility, the device greets them with, "Good morning, how are you feeling today?"
[1471] 2. Audio data collection and analysis
[1472] When the user responds, the voice data is converted into text data using a voice recognition module. This data is sent to a cloud server and analyzed using a natural language processing engine. Based on the analysis results, the user's health status and emotions are evaluated.
[1473] 3. Next question generation and presentation
[1474] The next question is generated from the text data by the generative AI model and presented to the user via their device or smartphone app, such as "Have you been sleeping well lately?"
[1475] 4. Collection of health data
[1476] Physical data such as users' temperature and blood pressure will be collected using measuring equipment within the facility, and similar data can also be collected via a smartphone app.
[1477] 5. Data Storage and Feedback
[1478] The collected voice and physical data is stored in cloud storage. Based on the analysis results, feedback is provided to the user. For example, the system may provide feedback such as, "After looking at the results, there are some concerns about your recent physical condition and conversations. Please consult a doctor."
[1479] Specific examples
[1480] 1. Usage scenarios in facilities
[1481] When an elderly person visits the facility, the robot terminal recognizes them and asks, "Good morning, how are you feeling today?" If the elderly person replies, "I have a slight headache," the voice is converted into text data and sent to the cloud.
[1482] 2. Smartphone usage scenarios
[1483] The same elderly person opens the smartphone app at home and is asked the same questions about their health. The collected data is stored in the cloud and used the next time they visit the facility. Based on this data, a generative AI model generates appropriate questions and performs a risk assessment.
[1484] 3. Examples of prompts
[1485] Please analyze the following sentences and assess the user's health and emotions.
[1486] "My head feels a little heavy today."
[1487] This system can continuously monitor the user's health status wherever they are and provide quick feedback, contributing to the early detection of dementia risk.
[1488] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1489] Step 1:
[1490] When a user arrives at a facility, the device uses a camera and voice recognition technology to recognize the user. The camera captures the user's face, and the voice recognition module identifies the user's personal information from the voice data. Specifically, the camera captures the user's image and receives a greeting as voice data. The device outputs a voice message saying, "Good morning, how are you feeling today?"
[1491] Step 2:
[1492] The user responds to the terminal by voice. The terminal captures the voice data and converts it into text data using a voice recognition module. The user's voice input is "I feel a bit heavy-headed today," which is converted into text data. This text data is sent to the server.
[1493] Step 3:
[1494] The server analyzes the received text data and uses a natural language processing engine (e.g., SpaCy or Google Cloud Natural Language API) to extract data from the text data to evaluate the user's emotions and health status. The input is the text data "I feel a bit heavy-headed today," and the output is "The user is feeling unwell."
[1495] Step 4:
[1496] The server uses the generative AI model based on the analysis results to generate the next question. For example, the question "Have you been sleeping well lately?" is generated. This generated question data is sent to the device. Specifically, the text data of the generated question is sent to the device.
[1497] Step 5:
[1498] The device presents the generated question to the user by voice or text. The user then responds to the device. This response data is also converted into text data using a voice recognition module and sent to the server. The specific operation is that the device outputs the question "Have you been sleeping well lately?" by voice.
[1499] Step 6:
[1500] The terminal instructs the user to measure their physical data using a thermometer or blood pressure monitor. The user measures their body temperature and blood pressure using health measurement equipment installed in the facility and reports the results to the terminal. For example, the user reports, "My body temperature is 36.5 degrees and my blood pressure is 120 / 80." This data is converted into text data and sent to the server.
[1501] Step 7:
[1502] The server integrates all the received data and uses an AI algorithm to assess the risk of dementia. The voice data and physical data are analyzed as a single data set to assess the user's risk. The inputs are voice-text data such as "Head feels heavy," body temperature of "36.5 degrees," and blood pressure of "120 / 80," and the output is the "dementia risk assessment result." The specific operation is for the AI algorithm to analyze the data and perform a risk assessment.
[1503] Step 8:
[1504] The server provides feedback to the device and smartphone app based on the evaluation results. For example, it generates an alert message saying, "After looking at the results, there are some concerns about your recent health condition and conversations. Please consult a doctor." and sends it to the device and smartphone app. The user receives the alert message.
[1505] Step 9:
[1506] All data is stored in cloud storage and will be available for the next session. Along with regular data monitoring, it will be possible to continuously track the user's health status. Specifically, the collected voice and physical data will be stored in the cloud and used for future analysis.
[1507] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1508] This invention relates to a robotic system for monitoring the health status and diagnosing dementia in the home of elderly people living alone. The system is equipped with functions for user recognition, voice conversation, data analysis, emotion engine, and feedback.
[1509] First, when the device (robot) starts up, it recognizes the user using the built-in camera and voice recognition function. At this point, the robot greets the user by saying, "Good morning, [user name]. How are you today?"
[1510] The user then responds to the robot by saying something like, "I feel a bit heavy-headed today." This voice data is converted into text data by the device's voice recognition module. At the same time, the device acquires the user's voice and facial expression data and analyzes them using an emotion engine.
[1511] The converted text data and the evaluation data from the emotion engine are sent to a server and analyzed by a natural language processing engine (NLP). This analysis extracts information to evaluate the user's emotions and health status. For example, it can infer a decline in health status from expressions such as "my head feels heavy" and recognize emotions such as "I'm tired."
[1512] The server then generates the next question based on the analysis results and sends it to the device. For example, the question is "Have you been sleeping well lately?", taking into account emotional changes in the user's facial expressions. This allows for more personalized interactions.
[1513] The device will ask the user this question. At the same time, the device will ask the user to take their temperature and blood pressure. It will add, "Please take your temperature and blood pressure."
[1514] The user measures their temperature using a thermometer or blood pressure monitor at home and reports the results to the robot, for example, "My temperature is 36.5 degrees and my blood pressure is 120 / 80."
[1515] Health data, text data, and emotional data are integrated on a server, and an AI algorithm is used to assess dementia risk. The AI assesses risk based on abnormalities in language usage patterns and physical data, while also taking emotional data into account.
[1516] For example, if the user has become increasingly forgetful in recent conversations or has frequently shown emotionally unstable facial expressions, the server will generate an alert message and send it to the device, which will then provide feedback such as, "After looking at the results, there are some concerns about your recent physical condition and emotions. We recommend that you consult a doctor."
[1517] All conversation, health and emotion data is stored in a database on the server, allowing for preparation for the next session and regular monitoring of the user's condition.
[1518] As a concrete example, the flow of one day is shown below.
[1519] 1. The device starts up, recognizes the user, and greets them with, "Good morning, Tanaka-san. How are you today?"
[1520] 2. The user replies, "Good morning, I'm feeling a bit down today."
[1521] 3. The device collects voice and facial expression data, converts the voice into text, and analyzes it using an emotion engine.
[1522] 4. The text data and emotional evaluation data sent to the server are analyzed, and the question "Have you often felt heavy-headed lately?" is generated and sent to the device.
[1523] 5. The device asks questions and prompts the user to measure their health data. It prompts the user to "measure their temperature and blood pressure."
[1524] 6. The user reports the measurement results and replies, "Temperature is 36.5°C, blood pressure is 120 / 80."
[1525] 7. The device records this as text data and emotion data and sends it to the server.
[1526] 8. The server analyzes all data and evaluates the risk of dementia. If the risk is high, an alert is sent to the device.
[1527] 9. The device will provide feedback such as, "After looking at the results, there are some concerns about your recent physical condition and emotions. We recommend that you consult a doctor."
[1528] This system allows for daily monitoring of the health and emotional state of the elderly and allows for early consultation with medical institutions when necessary, thereby improving the peace of mind and quality of life for the elderly and their families.
[1529] The processing flow will be explained below.
[1530] Step 1:
[1531] The device boots up and uses the built-in camera and microphone to scan the surroundings and identify the user. Once the user is identified, the device greets them with, "Good morning, [username]. How are you today?"
[1532] Step 2:
[1533] The user responds, "Good morning, I feel a bit heavy-headed today." The device collects the voice data and converts it into text data using a voice recognition module.
[1534] Step 3:
[1535] The device collects facial expression data along with the captured voice data and analyzes it with an emotion engine, where the user's emotional state is identified as "fatigue."
[1536] Step 4:
[1537] The converted text data and the evaluation data generated by the emotion engine are sent to the server, which then analyzes the received text data using a natural language processing engine (NLP).The analysis results are used to evaluate the user's emotions and health status.
[1538] Step 5:
[1539] The server generates the next question based on the analysis results, for example, "Have you been sleeping well recently?", and sends this question to the device.
[1540] Step 6:
[1541] The device asks the user, "Have you been sleeping well lately?" When asking the question, it also takes into account changes in the user's facial expressions based on their emotions. At the same time, the device instructs the user to measure their temperature and blood pressure. It adds, "Please measure your temperature and blood pressure."
[1542] Step 7:
[1543] The user uses a thermometer or blood pressure monitor to report the measurement results, such as "My body temperature is 36.5 degrees and my blood pressure is 120 / 80." The device's voice recognition module converts this voice back into text data, which is then sent to the server.
[1544] Step 8:
[1545] The server combines the health check data, text data, and emotional data, and uses an AI algorithm to assess dementia risk. For example, if a person has recently shown forgetfulness and emotional instability in their facial expressions, it will determine that they are at high risk.
[1546] Step 9:
[1547] If the risk is high based on the assessment results, the server generates an alert message and sends it to the device, such as "There are some concerns about your health. We recommend that you consult a doctor."
[1548] Step 10:
[1549] The device will then provide the user with an alert message, saying, "After looking at your results, there are some concerns about your recent physical condition and emotions. We suggest you consult a doctor."
[1550] Step 11:
[1551] The server stores all conversation data, health data, and emotion data in a database. It periodically prepares data to monitor the user's condition in preparation for the next session.
[1552] Example 2
[1553] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1554] In modern society, there is a need to routinely monitor the health status and dementia risk of elderly people living alone and detect abnormalities early. However, the means to do so are limited, and there are no systems that use interactive systems to closely observe the health and emotional state of elderly people and provide appropriate feedback. Therefore, there is a need for an efficient and effective method of managing the health of elderly people.
[1555] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for recognizing the user, means for initiating a voice conversation with the user and converting the voice data into text data, means for analyzing the converted text data and evaluating the user's emotions and health status, means for generating the next question based on the evaluation results and presenting it to the user, means for measuring the user's physical data and collecting the measurement results, means for integrating the voice and measurement data and evaluating dementia risk using AI, means for providing feedback to the user based on the evaluation results, and means for saving the data and preparing for the next session. This makes it possible to monitor the health and emotional status of elderly people on a daily basis and to cooperate with medical institutions early if necessary.
[1556] "Means for recognizing a user" refers to the system's ability to identify and recognize the individual user using a camera or voice recognition technology.
[1557] The "means for starting a voice conversation and converting voice data into text data" is a technology for automatically converting voice data acquired through a voice dialogue with a user into text format.
[1558] The "means for analyzing the converted text data and evaluating the user's emotions and health state" is a technology that processes the text data, analyzes its contents, and estimates the user's current emotions and health state.
[1559] The "means for generating the next question and presenting it to the user" is a function for generating and presenting an appropriate next question to the user based on the analysis results.
[1560] "Means for measuring the user's physical data and collecting the measurement results" refers to technology that allows the user to measure physical data such as body temperature and blood pressure and for the system to collect the results.
[1561] "Means for integrating voice and measurement data and assessing dementia risk using AI" refers to a technology that integrates collected voice data and physical data and uses artificial intelligence technology to assess a user's dementia risk.
[1562] The "means for providing feedback to the user based on the evaluation results" is a function for providing the results of the system's evaluation to the user in the form of information or advice in an appropriate format.
[1563] "Means for storing data and preparing for the next session" refers to technology for recording collected and analyzed data and preparing it for use in the next interaction or monitoring.
[1564] The present invention relates to a robotic system for monitoring the health status and dementia risk of elderly people living alone and for detecting abnormalities at an early stage. This system mainly includes a means for recognizing a user, a means for performing voice conversation and data analysis, an emotion analysis engine, a feedback means, and a storage means.
[1565] When the device (robot) starts up, it uses a camera and voice recognition functions to recognize the user. Specifically, it takes a picture of the user's face with the camera and recognizes the user's voice using voice recognition technology. At this time, the device says, "Good morning, [user name]. How are you today?"
[1566] When the user responds, the voice data is converted into text data using the device's voice recognition module. This converted text data and facial expression data captured by the camera are analyzed by the emotion engine. For example, by analyzing a statement such as "I feel a bit heavy-headed today" and the user's facial expression at the time, the user's emotions and health condition can be evaluated.
[1567] The device then sends this data to a server, which incorporates a natural language processing engine (NLP) and performs detailed analysis of the received text data and emotional evaluation data. For example, an expression such as "My head feels heavy" is interpreted as a sign of the user's stress or poor health. Emotional fluctuations are also extracted from facial expression data.
[1568] Based on the analysis results, the server generates the next question. For example, "Have you been sleeping well recently?" and sends it to the device. The system then instructs the user to measure their temperature and blood pressure. The system prompts the user to "measure their temperature and blood pressure."
[1569] The user measures their temperature using a thermometer or blood pressure monitor at home and reports the results to the robot. For example, they might report, "My temperature is 36.5 degrees and my blood pressure is 120 / 80." This health data is collected as text data and sent back to the server.
[1570] The server uses an AI algorithm to comprehensively analyze this data. It evaluates the user's dementia risk by combining voice, emotional, and physical data. For example, if the user has recently shown an increase in forgetfulness and emotional instability in their facial expressions, the server will assess the risk as high and generate an alert message.
[1571] Finally, the device provides the generated feedback to the user, for example, saying, "After looking at the results, there are some concerns about your recent physical condition and emotions. We recommend that you consult a doctor." This makes it possible to monitor the health status of elderly people on a daily basis and to consult with medical institutions early if necessary.
[1572] All conversation, health, and emotion data is stored on the server and prepared for the next session. Through regular monitoring, the health status of the elderly can be continuously managed.
[1573] (Example of a prompt)
[1574] Here are some example prompts for a generative AI model:
[1575] "Design a robotic system to check the health status of elderly people daily. This system should have the ability to recognize the user's face and voice, analyze their health status and emotions, and provide feedback as needed. It should also send the measured health data and analyzed emotional data to a server, enabling regular monitoring."
[1576] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1577] (Program processing steps)
[1578] Step 1:
[1579] The device starts up.
[1580] Input: System power supply
[1581] Output: The terminal is ready
[1582] At startup, the built-in camera and voice recognition function are enabled to prepare for user recognition.
[1583] Step 2:
[1584] The device uses its built-in camera and voice recognition capabilities to identify the user.
[1585] Input: User's face, voice
[1586] Output: Recognized user information
[1587] The camera captures the user's face and uses voice recognition technology to identify the individual from their voice.
[1588] The device greets you with, "Good morning, [username]. How are you today?"
[1589] Step 3:
[1590] The user responds, "I'm feeling a bit heavy-headed today."
[1591] Input: User's voice response
[1592] Output: User's voice data
[1593] The terminal records the user's response.
[1594] Step 4:
[1595] The terminal uses a voice recognition module to convert the user's voice data into text data.
[1596] Input: User's voice data
[1597] Output: Converted text data
[1598] The voice recognition module is activated to convert the voice data into text format.
[1599] Step 5:
[1600] The device collects the user's facial expression data captured by the camera and analyzes it using an emotion engine.
[1601] Input: User's facial expression data
[1602] Output: User's emotion rating data
[1603] A camera captures the user's facial expressions and uses an emotion analysis algorithm to assess their emotional state.
[1604] Step 6:
[1605] The terminal transmits the converted text data and emotion evaluation data to the server.
[1606] Input: Text data, emotion rating data
[1607] Output: Data sent to the server
[1608] The terminal uploads the data to the server via the network.
[1609] Step 7:
[1610] The server uses a natural language processing engine (NLP) to analyze the text data and assess the user's health status.
[1611] Input: Text data, emotion rating data
[1612] Output: User's health status rating
[1613] Using NLP technology, the system analyzes the converted text and, for example, infers a decline in the user's health from an expression like "my head feels heavy." It also takes into account emotional data.
[1614] Step 8:
[1615] The server generates the next question based on the analysis results.
[1616] Input: User's health assessment
[1617] Output: Next question (e.g., "Have you been sleeping well lately?")
[1618] Based on the analysis results, appropriate follow-up questions are automatically generated.
[1619] Step 9:
[1620] The server generates the next question and sends it to the terminal.
[1621] Input: Next question
[1622] Output: Questions sent to the terminal
[1623] The server transmits question data to the terminal via the network.
[1624] Step 10:
[1625] The device asks the user the following question and instructs them to "take your temperature and blood pressure."
[1626] Input: Next question, measurement instructions
[1627] Output: Questions and measurement instructions for the user
[1628] The device will then ask the user the next question via voice and display, and instruct them to measure further health data.
[1629] Step 11:
[1630] The user measures their temperature using a thermometer or blood pressure monitor and reports the results to the robot.
[1631] Input: User measurements
[1632] Output: Reported measurement data (e.g., "Temperature is 36.5°C, Blood pressure is 120 / 80")
[1633] The user follows the instructions to measure their own body temperature and blood pressure and reports the results to the device.
[1634] Step 12:
[1635] The device records the reported health data as text data and sends it to the server.
[1636] Input: Reported measurement data
[1637] Output: Health data sent to the server
[1638] The terminal converts the measurement data into text format and uploads it to a server via the network.
[1639] Step 13:
[1640] The server comprehensively analyzes health data, text data, and emotional data to assess dementia risk.
[1641] Input: Health data, text data, emotion data
[1642] Output: Dementia risk assessment
[1643] An AI algorithm is used to integrate all data and assess dementia risk.
[1644] Step 14:
[1645] If the risk is high, the server generates an alert message and sends it to the terminal.
[1646] Input: Dementia Risk Assessment
[1647] Output: Alert message
[1648] If the risk is determined to be high, the server generates an alert message and sends it to the terminal.
[1649] Step 15:
[1650] The terminal provides feedback to the user.
[1651] Input: Alert message
[1652] Output: Feedback to the user (e.g., "You've noticed some concerns about your physical and emotional state recently. Consult your doctor.")
[1653] The terminal displays an alert message to the user and provides appropriate advice.
[1654] Step 16:
[1655] The server stores all conversation data, health data, and emotion data.
[1656] Input: Conversation data, health data, emotion data
[1657] Output: Saved data
[1658] The server saves all the data in a database ready for the next session.
[1659] (Application example 2)
[1660] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1661] Monitoring the health and emotional state of elderly workers in factories is important, but conventional systems have difficulty grasping the situation and assessing risks in real time, making it difficult to respond effectively. Early response to prevent accidents and illnesses is also often delayed. To solve these problems, a system is needed that can comprehensively analyze workers' health and emotional data and assess risks early.
[1662] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recognizing a user, means for initiating a voice conversation with the user using a voice recognition module and converting the voice data into text data, means including a natural language processing engine for analyzing the converted text data and emotional data and evaluating the user's emotions and health status, means for generating the next question based on the evaluation results and presenting it to the user, means for measuring the user's physical data and collecting the measurement results, means for integrating the voice and measurement data and evaluating health risks using AI, means for providing feedback to the user based on the evaluation results, means for transmitting the health data and emotional data to the server and generating feedback based on the analysis results, and means for storing the data and preparing for the next session. This makes it possible to comprehensively grasp the health and emotional status of workers in real time, evaluate risks early, and provide appropriate feedback and countermeasures.
[1663] "Means for recognizing a user" refers to technology or devices that use cameras or sensors to identify a specific user.
[1664] "Means for initiating a voice conversation and converting voice data into text data" refers to a voice recognition technology or system that converts voice data acquired through a microphone into character string data.
[1665] "Means for analyzing converted text data and assessing a user's emotions and health status" refers to technology or systems that use natural language processing engines and other analytical algorithms to infer and assess a user's emotions and health status from their statements and expressions.
[1666] The "means for generating the next question and presenting it to the user" refers to a technology or system that automatically generates the next appropriate question based on the evaluation results and presents it to the user by voice or display.
[1667] "Means for measuring a user's physical data and collecting the measurement results" refers to a technology or system that uses devices such as a thermometer or blood pressure monitor to measure data related to a user's body and collect the results as data.
[1668] "Means for integrating voice and measurement data and assessing dementia risk using AI" refers to a technology or system that integrates acquired voice data and physical data and uses an AI algorithm to assess a user's dementia risk.
[1669] The "means for providing feedback to the user based on the evaluation results" refers to a technology or system that provides necessary advice or instructions to the user based on the analysis and evaluation results.
[1670] "Means for transmitting health data and emotional data to a server and generating feedback based on the analysis results" refers to a technology or system that transmits collected data to a server and provides appropriate feedback to the user based on the analysis results on the server side.
[1671] "Means for storing the data and preparing for the next session" refers to a technology or system that stores all acquired data in a database and prepares it for the next monitoring or evaluation based on that data.
[1672] The present invention relates to a system for real-time monitoring of the health and emotional states of factory workers and early risk assessment, which is equipped with user recognition, voice conversation, data analysis, emotion engine, and feedback functions.
[1673] Hardware and software used
[1674] Hardware:
[1675] 1. Camera: Used to recognize the face of the worker.
[1676] 2. Microphone: Used for voice recognition.
[1677] 3. Health measurement devices (thermometer, blood pressure monitor, heart rate monitor): Used to obtain physical data of workers.
[1678] software:
[1679] 1. OpenCV: Used to recognize the worker's face from the image data acquired through the camera.
[1680] 2. SpeechRecognition (Python library): Used to convert audio data into text data.
[1681] 3. Natural language processing engine: Analyzes the acquired text and emotion data and uses it to assess the emotions and health status of workers.
[1682] 4. Health Monitoring Library: Provides functions (get_temperature, get_blood_pressure, get_heart_rate) to get body temperature, blood pressure, and heart rate.
[1683] 5. Emotion Recognition Library: Used to analyze emotions from workers' facial expression data or voice data.
[1684] 6. Requests (Python library): Used to send data to and retrieve data from the server.
[1685] Processing flow
[1686] The system first recognizes the user (worker) using a camera and a voice recognition module. For example, when the system starts up, the camera scans the worker's face and greets them verbally, saying, "Good morning, [Worker's name]. How are you feeling?" If the worker replies, "I'm not feeling well today," the voice data is converted into text data by the voice recognition module.
[1687] The converted text data and acquired emotional data (analyzed from facial expressions and voice) are then analyzed using a natural language processing engine, allowing the worker's emotions and health status to be assessed.
[1688] Next, the system asks the worker to measure their body temperature, blood pressure, and heart rate. For example, data such as "body temperature is 37.5 degrees, blood pressure is 130 / 85, and heart rate is 80" is collected from the health measurement device. The collected data is sent to a server, where a health risk assessment is performed.
[1689] Based on the analysis and evaluation results, the system generates the next appropriate question to ask the worker and presents it to them. If the worker's health condition is not good, the system will provide feedback such as "You seem unwell. Please consult a doctor."
[1690] All data is stored in a database and prepared for the next session, allowing for a real-time understanding of the health and emotional state of factory workers, early risk assessment, and appropriate feedback and countermeasures.
[1691] Specific examples
[1692] Example prompt sentence:
[1693] Speech recognition: "I'm not feeling well today"
[1694] Facial Recognition: "Good morning, [Worker's Name]. How are you feeling?"
[1695] Health data collection: "Temperature is 37.5°C, blood pressure is 130 / 85, heart rate is 80."
[1696] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1697] Step 1:
[1698] The terminal starts up and uses the camera to recognize the user's (worker's) face. At this time, the terminal analyzes the image data acquired through the camera using OpenCV and checks whether it matches the facial data registered in advance. The input is the image data acquired from the camera, and the output is the user ID. In concrete terms, the camera scans the worker's face and recognizes matching facial data.
[1699] Step 2:
[1700] The device greets the recognized user by voice and asks about their current health condition. For example, it might ask, "Good morning, [user name]. How are you feeling?" The input is the user ID, and the output is a voice message. Specifically, the device incorporates the user name into the greeting and outputs it by voice.
[1701] Step 3:
[1702] The user responds by saying something like, "I'm not feeling well today." This voice data is acquired through the device's microphone. The input is the user's voice data, and the output is voice data. In concrete terms, the user reports their physical condition by voice into the device.
[1703] Step 4:
[1704] The device uses a speech recognition module to convert the acquired voice data into text data. The input is voice data and the output is text data. Specifically, the device converts the acquired voice data into text data using the SpeechRecognition library.
[1705] Step 5:
[1706] The converted text data and facial expression data are analyzed to evaluate the user's emotions and health status. The device analyzes the text data and facial expression data using a natural language processing engine and an emotion recognition library. The input is the text data and facial expression data, and the output is the evaluation results. Specifically, the device evaluates the emotions and health status based on the analysis results.
[1707] Step 6:
[1708] The device generates the next question based on the evaluation results and presents it to the user. For example, it generates a question such as "Have you been sleeping well recently?" The input is the evaluation results, and the output is the question text. As a specific operation, the generated question is output to the user by voice or display.
[1709] Step 7:
[1710] The device asks the user to measure their body temperature, blood pressure, and heart rate. For example, it instructs the user to "measure their body temperature and blood pressure." The input is the evaluation result, and the output is an instruction message. Specifically, the device gives instructions to measure by voice.
[1711] Step 8:
[1712] A user uses a health measurement device to measure their body temperature, blood pressure, and heart rate, and reports the results to a terminal. For example, they report, "My body temperature is 37.5 degrees, my blood pressure is 130 / 85, and my heart rate is 80." The input is the measurement data, and the output is the reported data. Specifically, the user inputs the measurement data into the terminal.
[1713] Step 9:
[1714] The device sends the collected voice data, health data, and emotion data to the server. The input is the collected data, and the output is the result sent to the server. Specifically, the device uses the Requests library to send data to the server.
[1715] Step 10:
[1716] The server analyzes the received data and evaluates health risks. The input is the received data, and the output is the evaluation result. Specifically, the server uses an AI algorithm to perform risk assessment.
[1717] Step 11:
[1718] The server generates feedback based on the evaluation results and sends it to the device. For example, it generates feedback such as "You seem unwell, please consult a doctor." The input is the evaluation results, and the output is a feedback message. The specific operation is to send the generated feedback to the device.
[1719] Step 12:
[1720] The terminal presents the feedback received from the server to the user. The input is the feedback message, and the output is the result presented to the user. Specifically, the terminal conveys the feedback to the user by voice.
[1721] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1722] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1723] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1724] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1725] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1726] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1727] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1728] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1729] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1730] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1731] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1732] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1733] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1734] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1735] It is not necessary to store all of the sp...
Claims
1. a means for recognizing a user; means for initiating a voice conversation with a user and converting voice data into text data; A means for analyzing the converted text data and evaluating the user's emotions and health status; means for generating a next question based on the evaluation result and presenting the next question to the user; means for measuring a user's physical data and collecting the measurement results; A method for integrating voice and measurement data to assess dementia risk using AI, a means for providing feedback to the user based on the evaluation results; The system includes means for storing said data and preparing it for the next session.
2. 2. The system of claim 1, wherein the means for converting voice data to text data uses a voice recognition module.
3. The system of claim 1 , wherein the means for assessing the user's emotions and health status includes a natural language processing engine.
4. The system of claim 1, wherein the means for assessing the risk of dementia uses an AI algorithm.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A