System

A system for early detection of frailty in elderly adults by analyzing daily life data with multimodal AI and generating personalized care suggestions addresses the inadequacies of conventional methods, enabling timely and effective health management.

JP2026028163APending Publication Date: 2026-02-19SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024130461
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-06
Publication Date
2026-02-19

AI Technical Summary

Technical Problem

Conventional methods for detecting early decline in physical, psychological, and social functions in elderly people and adults are inadequate, as regular medical checkups and self-checks fail to provide sufficient measures, and family members and caregivers lack specialized knowledge for appropriate care.

Method used

A system that collects video and audio data from daily life, analyzes it using multimodal AI to assess frailty risk, generates care suggestions, and provides timely notifications and responses to user queries, while periodically comparing new data with past data for continuous health monitoring.

Benefits of technology

Enables early detection of frailty risk and provides appropriate care by continuously analyzing daily life data, generating personalized care methods, and responding to user questions, thus effectively managing health at home.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026028163000001_ABST
    Figure 2026028163000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for collecting video and audio data of daily life; means for transmitting the collected video and audio data to a server; means for analyzing the received video and audio data and evaluating a risk of deterioration in physical, psychological, and social functions of a user; means for generating an appropriate care method and lifestyle improvement proposal for the user based on an evaluation result; and means for notifying the user of the generated care proposal.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] There is a challenge in early detection of the decline in physical, psychological, and social function (frailty) experienced in the daily lives of elderly people and adults who need prevention, and in providing appropriate care. Conventional regular medical checkups at medical institutions and self-checks are difficult to provide sufficient measures, and early detection and prevention of frailty are not always achieved. Furthermore, there is also the challenge for family members and caregivers, who lack specialized knowledge and find it difficult to provide appropriate care. [Means for solving the problem]

[0005] The present invention is a system including a means for collecting video and audio data of daily life, a means for transmitting the collected video and audio data to a server, a means for analyzing the received video and audio data and assessing the user's risk of decline in physical, psychological, and social functions, a means for generating appropriate care methods and lifestyle improvement suggestions for the user based on the assessment results, and a means for notifying the user of the generated care suggestions.The system also includes a means for receiving questions and requests for advice from the user and transmitting the information to a server, a means for providing the user with the answers generated by the server, and a means for periodically collecting new video and audio data and comparing it with past data, thereby enabling early detection of a user's risk of frailty and providing appropriate care.

[0006] "Video and audio data" refers to video and audio recordings collected in the user's daily life.

[0007] "Means of collection" refers to devices and software that collect video and audio data using cameras and microphones, etc.

[0008] "Transmission means" refers to devices and protocols for transmitting collected video and audio data to a server via communication means such as the Internet.

[0009] "Server" refers to a computer system for receiving and analyzing collected data.

[0010] "Means for analyzing" refers to algorithms or software for analyzing received video and audio data.

[0011] "Risk of decline" refers to the risk associated with a decline in the user's physical, psychological, and social functioning.

[0012] "Means of evaluation" refers to algorithms or systems that evaluate a user's frailty risk as a numerical value or evaluation value based on the analysis results.

[0013] "Care methods" refer to specific behaviors and lifestyle improvement measures suggested in response to the user's frailty risk.

[0014] "Improvement suggestions" refer to specific advice and instructions regarding care methods and lifestyle improvements.

[0015] "Means for notifying" refers to a device or software that notifies the user of the generated care suggestions.

[0016] "Means for accepting questions or requests for advice" refers to the interface or system through which users can ask questions or seek further advice.

[0017] The "means for providing an answer" is the means by which the server communicates the answer generated by the server to the user.

[0018] "Comparative measures" refer to algorithms or systems that compare new data collected with past data and analyze changes. [Brief explanation of the drawings]

[0019] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0020] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0021] First, the terms used in the following description will be explained.

[0022] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0023] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0024] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0025] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0027] [First embodiment]

[0028] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0029] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0030] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0031] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0032] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0034] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0035] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0036] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0037] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0038] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0039] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0040] This invention relates to a system for early detection of risk of decline in physical, psychological, and social functions (frailty) in a user's daily life and providing appropriate care. An embodiment of this system will be described in detail below.

[0041] overview

[0042] The system uses a device to collect video and audio data from the user's daily life and transmit it to a server. The server analyzes the received data and assesses the user's risk of frailty. Based on the results, it generates appropriate care methods and lifestyle improvement suggestions and notifies the user. It also responds to questions and requests for advice from the user and periodically compares new data with past data.

[0043] Data collection and transmission

[0044] The device uses a camera and microphone to collect video and audio data of the user's daily life. For example, the device is installed in the user's living space and automatically records walking, eating, conversations, and other activities that occur there.

[0045] The collected data is encrypted to protect privacy and sent to a server via the Internet. This transmission is carried out periodically (e.g., every night at midnight), and the data for the day is aggregated on the server.

[0046] Data analysis and evaluation

[0047] The server uses a multimodal AI model to analyze the received video and audio data. Specifically, it analyzes the user's movements, walking speed, and posture from the video data, and the tone of voice, speech tempo, and frequency of conversation from the audio data.

[0048] Based on the analysis results, the user's risk of decline in physical, psychological, and social functions is assessed. For example, if a decrease in walking speed or frequency of conversation is observed, it is determined that the risk of frailty is increasing.

[0049] Care suggestion generation and notification

[0050] Based on the analysis results, the server generates appropriate care methods and lifestyle improvement suggestions for the user, including daily exercise, a balanced diet, and ways to socialize.

[0051] The generated care suggestions are then sent back to the device and notified to the user. For example, specific instructions such as "Today's recommended exercise: 30 minutes of walking" are displayed to the user through the device.

[0052] User interaction

[0053] The device also has the ability to accept questions and requests for advice from users. When a user inputs a question through the device's interface, the information is sent to the server. The server uses generative AI to generate an appropriate answer to the question and returns it to the device. For example, in response to the question, "Would you like to change your exercise plan for today?", the device might provide a response such as, "Today's recommended exercise is 30 minutes of walking, but if the weather is bad, we recommend stretching indoors."

[0054] Data collection and comparison

[0055] The device also has the ability to periodically collect new video and audio data and compare it with past data, allowing it to track changes in the user's health and update care recommendations as needed. The device collects data daily, and the server uses this data to continuously assess risk.

[0056] As described above, this system detects users' risk of frailty early and provides appropriate care, enabling elderly people and adults who need preventative care to effectively manage their health at home.

[0057] The processing flow will be explained below.

[0058] Step 1:

[0059] The device collects video and audio data from the user's daily life through a camera and microphone, for example recording the user walking around the living room or talking while watching TV.

[0060] Step 2:

[0061] The device encrypts the collected video and audio data and transmits it securely to a server, taking privacy into consideration. The data is transmitted on a scheduled basis, such as at night or during periods of low load.

[0062] Step 3:

[0063] The server analyzes the received video and audio data, where a multimodal AI model is activated to extract the user's walking speed, posture, and movement patterns from the video data, and evaluate the tone of voice, speech tempo, and frequency of conversation from the audio data.

[0064] Step 4:

[0065] The server evaluates the risk of decline in the user's physical, psychological, and social functions based on the analysis results, for example, by comparing them with past data to see if their walking speed or frequency of conversation has decreased.

[0066] Step 5:

[0067] Based on the assessment results, the server generates appropriate care methods and lifestyle improvement suggestions, such as recommendations for appropriate exercise, a balanced diet, and ways to increase social interaction.

[0068] Step 6:

[0069] The server then sends the generated care recommendations back to the device, which then notifies the user. For example, a push notification will be displayed saying, "Today's recommended exercise is 30 minutes of walking."

[0070] Step 7:

[0071] Users can ask questions or ask for additional advice through the device, for example, "What is a good exercise for a rainy day?"

[0072] Step 8:

[0073] The device sends the user's question to the server, which then uses generative AI to generate an answer to the question. The answer is then sent to the device, which displays it to the user. For example, the device might return an answer such as "We recommend stretching indoors."

[0074] Step 9:

[0075] The device continues to periodically collect new video and audio data. The server compares this new data with past data and continually updates the user's risk assessment. For example, it analyzes the data collected daily to track changes in health status.

[0076] In this way, the system continuously collects and analyzes data throughout the user's daily life and provides appropriate care, making it possible to detect the user's risk of frailty early and respond appropriately.

[0077] Example 1

[0078] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0079] There is a lack of means to detect early risk of decline in physical, psychological, and social function (frailty) at home for elderly people and adults who need preventative care, and to provide appropriate care. Conventional systems often do not adequately collect and analyze data closely related to users' daily lives, resulting in delayed risk assessment. There is also a need for a system that can respond quickly and appropriately to user questions.

[0080] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0081] In this invention, the server includes a means for encrypting video and audio data of daily life and transmitting it to the server, a means for analyzing the received video and audio data using a multimodal AI model to assess the user's risk of decline in physical, psychological, and social function, and a means for generating appropriate care methods and lifestyle improvement suggestions for the user based on the assessment results. This allows users to easily collect data from their daily lives and have it analyzed by the server, enabling early detection of frailty risk and appropriate care. Furthermore, answer generation using the generative AI model enables prompt and appropriate responses to user questions.

[0082] "Video and audio data of everyday life" refers to video and audio data that records the actions and conversations of users in their daily lives.

[0083] "Encryption" is a technology that converts the contents of data into a format that cannot be understood by third parties, and is used particularly to protect privacy and improve security.

[0084] A "server" is a computer system that receives data from client terminals via a network and analyzes and processes the data.

[0085] A "multimodal AI model" is an artificial intelligence model that can comprehensively analyze multiple types of data, such as video data and audio data.

[0086] "Frailty risk" is an indicator of the risk of decline in physical, psychological, and social functioning, and is particularly important when assessing the health status of elderly people and adults who require prevention.

[0087] "Care Suggestions" are advice including specific exercises, balanced meals, and social interaction methods aimed at maintaining the user's health and restoring function.

[0088] A "generative AI model" is an artificial intelligence model that automatically generates appropriate answers or suggestions based on input data.

[0089] "Comparison" is the act of matching new data with past data to detect changes over time.

[0090] "Notification" refers to the act of conveying important information to the user, and in particular, care suggestions and responses from the system.

[0091] This invention relates to a system that detects early risk of decline in physical, psychological, and social function (frailty) in a user's daily life and provides appropriate care. This system can execute a series of processes, from collecting and analyzing video and audio data to generating and notifying care suggestions.

[0092] Data collection and transmission

[0093] The device uses a camera and microphone to collect video and audio data of the user's daily life. The device is equipped with a Python program that uses the OpenCV library to capture video and the PyAudio library to record audio. For example, the device can be installed in the user's living space and automatically record walking, eating, conversations, and other activities.

[0094] The collected data is encrypted using encryption technology such as AES. The encrypted data is periodically sent to a server via the Internet. Data transmission is managed by a regularly executed script such as a Cron job. Specifically, the collected data is encrypted every night at midnight and sent using the HTTPS protocol.

[0095] Data analysis

[0096] The server uses a multimodal AI model to analyze the received data. The multimodal AI model includes a model that comprehensively analyzes the features of video data and audio data (e.g., a model trained using PyTorch). Specifically, the server analyzes the user's walking speed and posture from the video data, and the tone of voice, speech tempo, and frequency of conversation from the audio data.

[0097] Risk assessment and care recommendation generation

[0098] The server evaluates the user's risk of frailty based on the analysis results. If a decrease in walking speed or frequency of conversation is observed, it is determined that the risk of frailty is increasing. Based on the evaluation results, the server generates appropriate care methods and lifestyle improvement suggestions. These suggestions include daily exercise, a balanced diet, and ways to socialize. The generated care suggestions are formatted into text.

[0099] Care Suggestion Notification

[0100] The generated care suggestions are sent back to the device and notified to the user. The suggestions are displayed as push notifications or alerts. For example, a message such as "Today's recommended exercise: 30 minutes of walking" is displayed on the user's device.

[0101] Responding to user questions

[0102] The device has the ability to accept questions and requests for advice from users. When a user inputs a question through the device's interface, the information is sent to the server. The server uses a generative AI model to generate an appropriate answer to the question and returns it to the device. For example, in response to the question, "Would you like to change your exercise plan for today?", the device might provide an answer such as, "Today's recommended exercise is 30 minutes of walking, but if the weather is bad, we recommend stretching indoors."

[0103] Data collection and comparison

[0104] The server periodically collects new data and compares it with past data. This allows the server to track changes in the user's health and update care recommendations as needed. For example, if walking speed data from the past three months is compared and a clear decline is detected, increased exercise is suggested. In a specific scenario, the device collects data daily, and the server continuously performs risk assessments based on this data.

[0105] Prompt Sentence Examples

[0106] "Analyze walking speed data from the past three months and generate appropriate care suggestions if significant decline is observed."

[0107] "If the frequency of conversations during meals decreases, please provide psychological support suggestions accordingly."

[0108] As described above, the purpose of this system is to detect early risks of decline in the user's physical, psychological, and social functions in their daily lives and to provide appropriate care.

[0109] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0110] Step 1: Data collection

[0111] The device collects video and audio data from the user's daily life. The input is video and audio data from the camera and microphone. The device uses a Python program to capture video using the OpenCV library and record audio using the PyAudio library. The output is the collected raw data. Specifically, the device activates the camera and microphone every morning at 6:00 and begins recording movements and conversations in the living room and kitchen.

[0112] Step 2: Send data

[0113] The terminal encrypts the collected video and audio data using AES and sends it to the server. The input is the raw data collected in step 1. Data transmission is managed by a regularly executed script such as a Cron job. The output is the encrypted data sent to the server. Specifically, the collected data is encrypted every night at midnight and transmitted using the HTTPS protocol.

[0114] Step 3: Data analysis

[0115] The server analyzes the received data using a multimodal AI model. The input is encrypted video and audio data. The server decrypts the data and extracts features from the video data using OpenCV and the audio data using Librosa. The output is the analysis results. Specifically, the server decompresses the data, inputs it into the AI ​​model for analysis, and extracts features such as walking speed, tone of voice, and frequency of conversation.

[0116] Step 4: Risk assessment

[0117] The server evaluates the user's risk of frailty based on the analysis results. The input is the analysis results obtained in step 3. The server calculates a risk score based on this data and determines the risk level. The output is the risk assessment result. Specifically, the server evaluates a decrease in walking speed and frequency of conversation, and calculates a risk level of "low," "medium," or "high."

[0118] Step 5: Generate a care proposal

[0119] The server generates appropriate care suggestions based on the risk assessment results. The input is a database of risk assessment results and corresponding care suggestions. The server uses the generative AI model to generate suggestions such as exercise, diet, and social interaction. The output is specific care suggestions. Specific operations include the server formatting "daily exercises" in text format based on the AI ​​model.

[0120] Step 6: User Notification

[0121] The server sends the generated care suggestions to the device and notifies the user. The input is the text data of the care suggestions. The server sends the suggestions to the device, and the device displays a notification to the user. The output is the state in which the notification is displayed to the user. Specifically, the server sends content such as "Today's recommended exercise: 30 minutes of walking" to the device, and the device displays it as a push notification.

[0122] Step 7: Respond to user questions

[0123] The device accepts questions and requests for advice from the user. The input is the question text from the user. The device sends this information to the server, which uses a generative AI model to generate an answer. The output is the generated answer. Specifically, the user enters the question, "Would you like to change your exercise routine today?" and the server provides the answer, "If the weather is bad, we recommend stretching indoors."

[0124] Step 8: Data comparison and update

[0125] The server periodically collects new data and compares it with past data. The input is new video and audio data and previously collected data. The server compares these and detects changes in health status. The output is updated care recommendations. Specifically, the server compares walking speed data from the past three months and generates exercise enhancements or new care recommendations as needed, and notifies the user.

[0126] (Application example 1)

[0127] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0128] This invention relates to a system that detects the risk of decline in physical, psychological, and social functions in daily life (frailty) at an early stage and provides appropriate care. The purpose of this invention is to solve the problem of the difficulty of monitoring the health status of customers in real time in a physical store and providing timely care and lifestyle improvement suggestions. It is also required to respond quickly to customers' questions and requests for advice.

[0129] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0130] In this invention, the server includes means for collecting video and audio data of daily life, means for transmitting the collected video and audio data to the server, means for analyzing the received video and audio data and assessing the risk of decline in the user's physical, psychological, and social functions, means for generating appropriate care methods and lifestyle improvement suggestions for the user based on the assessment results, means for notifying the user of the generated care suggestions, and means for monitoring the physical and psychological functions of customers in a physical store in real time using a smart device and assessing health risks. This makes it possible to assess the health status of customers in real time and provide appropriate care and lifestyle improvement suggestions as needed.

[0131] "Video and audio data from everyday life" refers to video footage and audio recordings taken in the user's living environment.

[0132] "Server" refers to a computer system that receives, analyzes, and evaluates collected video and audio data.

[0133] "Means for assessing the risk of decline in physical, psychological, and social function" refers to algorithms and programs that analyze data indicating the user's health status and assess the risk of frailty.

[0134] The "means for generating suggestions for improving care methods and lifestyle habits" refers to a means for suggesting to the user care and lifestyle habits improvement methods that should be performed based on the evaluation results.

[0135] "Means for notifying the user of the generated care suggestions" refers to a device or interface for notifying the user of the generated care suggestions in an easy-to-understand manner.

[0136] "Means of monitoring the physical and psychological functions of customers in real time using smart devices in physical stores" refers to smart devices such as cameras and microphones used to observe and record the health status of customers in physical stores, as well as systems for analyzing that data.

[0137] This invention is a system that monitors the health status of customers in a physical store in real time and provides appropriate care and lifestyle improvement suggestions. This invention is realized by collecting video and audio data of daily life using smart devices and sending it to a server.

[0138] System configuration

[0139] Hardware and Software Used

[0140] 1. Smart Devices:

[0141] Using cameras and microphones (e.g., smart glasses), video and audio recordings of customers' everyday activities are collected.

[0142] 2. Server:

[0143] Uses Python, OpenCV, PyAudio, and the Requests library to receive, store, and analyze data.

[0144] A generative AI model is used to analyze the collected data and generate appropriate care methods and improvement suggestions.

[0145] Data collection and transmission

[0146] 1. Data Collection:

[0147] The smart glasses are equipped with a camera and microphone that collects video and audio data of customers, specifically monitoring their walking speed, posture, tone of voice, and frequency of language use.

[0148] 2. Data transmission:

[0149] The collected data is encrypted to protect privacy and sent to a local server at the physical store or to a cloud server.

[0150] Data analysis and evaluation

[0151] The server uses the generative AI model to analyze the received video and audio data. Specifically, it performs the following data analysis:

[0152] 1. Video Data Analysis:

[0153] Analyzes user movements, walking speed, and posture.

[0154] 2. Audio data analysis:

[0155] Analyze tone of voice and frequency of language usage.

[0156] Based on the analysis results, the server evaluates the user's risk of declining physical and psychological function. If the evaluation finds a decrease in walking speed or frequency of conversation, it determines that the risk of frailty is increasing.

[0157] Care suggestion generation and notification

[0158] Based on the analysis results, the server generates appropriate care methods and lifestyle improvement suggestions for the customer, including:

[0159] 1. Exercise Suggestions:

[0160] Generate specific instructions such as "Today's recommended exercise: 30 minutes of walking."

[0161] 2. Dietary suggestions:

[0162] They offer advice such as, "Try to eat a balanced diet."

[0163] The generated care suggestions are communicated to the customer via the smart glasses display.

[0164] User interaction features

[0165] Customers can input questions or requests for advice through the smart glasses interface. This information is sent to the server, where the AI ​​generates an appropriate answer and sends it back to the smart glasses.

[0166] "What exercises would you recommend today?"

[0167] What measures should be taken if walking speed slows down?

[0168] "How would you assess your health status based on today's walking data?"

[0169] Answers to prompts such as:

[0170] Regularly collect and compare data

[0171] New data is collected periodically and compared with past data to track changes in the health status of customers. To achieve this, the smart glasses collect data daily, and the server continuously performs risk assessments based on this data.

[0172] The above is a specific embodiment of the present invention, which makes it possible to evaluate the health status of customers in real time within a physical store and provide appropriate care and lifestyle improvement suggestions.

[0173] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0174] Step 1:

[0175] The device collects video and audio data from the user's daily life.

[0176] Input: Video and audio data captured by the camera and microphone on the smart glasses.

[0177] Data processing and calculation: Use OpenCV to capture video and PyAudio to record audio.

[0178] Output: Collected video data (e.g., JPEG images) and audio data (e.g., WAV files).

[0179] Step 2:

[0180] The terminal transmits the collected video and audio data to the server.

[0181] Input: Video and audio data collected in step 1.

[0182] Data processing and calculation: The data is encrypted to protect privacy and sent to the server using an HTTP POST request.

[0183] Output: Encrypted video and audio data is sent to the server.

[0184] Step 3:

[0185] The server analyzes the received video and audio data and assesses the user's risk of decline in physical and psychological function.

[0186] Input: Encrypted video and audio data sent from the device.

[0187] Data processing and calculation: Generative AI models are used to analyze factors such as walking speed, posture, tone of voice, and frequency of language use.

[0188] Output: Results of the user's frailty risk assessment.

[0189] Step 4:

[0190] Based on the evaluation results, the server generates appropriate care methods and lifestyle improvement suggestions for the user.

[0191] Input: Frailty risk assessment results obtained in Step 3.

[0192] Data processing and calculation: Based on the assessment results, care recommendations such as specific exercises, dietary improvements, and social interaction methods are generated.

[0193] Output: Generated care suggestions (e.g., "Today's recommended exercise: 30 minutes walking").

[0194] Step 5:

[0195] The terminal notifies the user of the generated care suggestions.

[0196] Input: Care suggestions generated in step 4.

[0197] Data processing and calculation: Care suggestions are displayed on the smart glasses display.

[0198] Output: Notification to the user (e.g. "Recommended exercise for today: 30 minutes walking" appears on the display).

[0199] Step 6:

[0200] The user inputs questions or requests for advice through the device interface.

[0201] Input: Text data of a user question or request (e.g., "What exercise would you recommend for me today?").

[0202] Data processing and calculation: The terminal sends this information to the server.

[0203] Output: The user's question is sent to the server.

[0204] Step 7:

[0205] The server uses a generative AI model to generate appropriate answers to the questions and returns them to the device.

[0206] Input: The user's question submitted in step 6.

[0207] Data processing and computation: Generative AI models are used to generate answers to questions.

[0208] Output: The generated answer (e.g., "Today's recommended exercise is 30 minutes of walking, but if the weather is bad, indoor stretching is recommended").

[0209] Step 8:

[0210] The terminal provides the answer from the server to the user.

[0211] Input: The answer from the server generated in step 7.

[0212] Data processing and calculation: The answer is displayed on the smart glasses display.

[0213] Output: Notify the user of the answer (e.g., "Today's recommended exercise is 30 minutes of walking, but if the weather is bad, we recommend stretching indoors.")

[0214] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0215] This invention relates to a system that detects early risk of decline in physical, psychological, and social functions (frailty) in a user's daily life and provides appropriate care, and further combines it with an emotion engine that recognizes the user's emotions. An embodiment of this system will be described in detail below.

[0216] overview

[0217] The system uses a device to collect video and audio data from the user's daily life and send it to a server. The server not only analyzes the received data and assesses the user's risk of frailty, but also recognizes the user's emotions using an emotion engine. Based on the results, it generates appropriate care methods and lifestyle improvement suggestions and notifies the user. It also responds to questions and requests for advice from the user and regularly compares new data with past data.

[0218] Data collection and transmission

[0219] The device uses a camera and microphone to collect video and audio data from the user's daily life, such as when they walk around the living room or talk while watching TV. The emotion engine then analyzes the user's facial expressions and tone of voice in real time to recognize their emotions.

[0220] The collected data is encrypted to protect privacy and sent to a server via the Internet. This transmission is carried out periodically (e.g., every night at midnight), and the data for the day is aggregated on the server.

[0221] Data analysis and evaluation

[0222] The server analyzes the received video and audio data. A multimodal AI model is activated to extract the user's movements, walking speed, and posture from the video data, and evaluate the tone of voice, speech tempo, and frequency of conversation from the audio data. In addition, an emotion engine analyzes the user's emotions from changes in facial expressions and tone of voice. This also allows the collection of user emotional data.

[0223] Based on the analysis results, the system assesses the user's risk of decline in physical, psychological, and social function. For example, it compares data from the past to see if their walking speed or frequency of conversation has decreased. Emotional data is also taken into account, and if stress or depression is observed, it determines that the risk is high.

[0224] Care suggestion generation and notification

[0225] Based on the analysis results, the server generates appropriate care methods and lifestyle improvement suggestions for the user, including suggestions for appropriate exercise, a balanced diet, and ways to increase social interactions. Emotional data is also taken into account, and suggestions are made based on the user's emotional state.

[0226] The generated care suggestions are then sent back to the device, which then notifies the user. For example, a push notification will be displayed saying, "Today's recommended exercise is 30 minutes of walking. We also recommend yoga for relaxation."

[0227] User interaction

[0228] The device has the ability to accept questions and requests for advice from users. When a user inputs a question through the device's interface, the information is sent to the server. The server uses generative AI to generate an appropriate answer to the question and returns it to the device. For example, in response to the question, "Would you like to change your exercise plan for today?", the device may provide an answer such as, "Today's recommended exercise is 30 minutes of walking, but if the weather is bad, we recommend stretching indoors."

[0229] Data collection and comparison

[0230] The device continues to periodically collect new video and audio data. The server compares this new data with past data and continually updates the user's risk assessment. For example, it analyzes the data collected daily to track changes in health status or emotions.

[0231] In this way, this system continuously collects and analyzes data from the user's daily life and provides appropriate care, making it possible to detect the user's risk of frailty early and respond appropriately.Furthermore, by combining it with an emotion engine, it is possible to realize individual responses that take into account the user's emotional state.

[0232] The processing flow will be explained below.

[0233] Step 1:

[0234] The device collects video and audio data from the user's daily life through a camera and microphone, for example, recording the user walking around the living room or talking while watching TV. The emotion engine analyzes the user's facial expressions and tone of voice in real time to recognize their emotions.

[0235] Step 2:

[0236] The device encrypts the collected video and audio data and transmits it securely to a server, taking privacy into consideration. The data is transmitted on a scheduled basis, such as at night or during periods of low load.

[0237] Step 3:

[0238] The server analyzes the received video and audio data. Here, a multimodal AI model is activated, extracting the user's walking speed, posture, and movement patterns from the video data, and evaluating the tone of voice, speech tempo, and frequency of conversation from the audio data. Furthermore, an emotion engine analyzes the user's emotions from changes in facial expressions and tone of voice.

[0239] Step 4:

[0240] The server evaluates the user's risk of decline in physical, psychological, and social function based on the analysis results. For example, it compares the results with past data to see if their walking speed has slowed or if they are talking less frequently. Emotional data is also taken into account, and if stress or depression is observed, it determines that the risk is high.

[0241] Step 5:

[0242] The server generates appropriate care methods and lifestyle improvement suggestions based on the evaluation results. These include suggestions for appropriate exercise, a balanced diet, and ways to increase social interaction. It also takes into account emotional data, suggesting relaxation techniques if the user is feeling stressed, for example.

[0243] Step 6:

[0244] The server then sends the generated care recommendations back to the device, which then notifies the user. For example, a push notification will be displayed saying, "Today's recommended exercise is 30 minutes of walking. We also recommend yoga for relaxation."

[0245] Step 7:

[0246] Users can ask questions or ask for additional advice through the device, for example, "What is a good exercise for a rainy day?"

[0247] Step 8:

[0248] The device sends a question from the user to the server, which then uses generative AI to generate an answer to the question. The answer is then sent to the device, which then displays it to the user. For example, the device might provide an answer such as "We recommend stretching indoors."

[0249] Step 9:

[0250] The device will continue to periodically collect new video and audio data. The server will compare this new data with past data and continually update the user's risk assessment. For example, it will analyze the data collected daily to track changes in health status or emotions.

[0251] In this way, this system continuously collects and analyzes data from the user's daily life and provides appropriate care, making it possible to detect the user's risk of frailty early and respond appropriately.Furthermore, by combining it with an emotion engine, it is possible to realize individual responses that take into account the user's emotional state.

[0252] Example 2

[0253] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0254] In modern society, the aging population is increasing the risk of decline in physical, psychological, and social functions (frailty). Early detection of this risk and the provision of appropriate care and lifestyle improvement suggestions are required, but currently there is a lack of systems that can effectively do this. Individualized care that takes into account the user's emotional state is also required, but systems that can address this are limited. To solve this problem, a system is needed that can continuously collect and analyze data from the user's daily life and provide appropriate care.

[0255] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0256] In this invention, the server includes a means for collecting video and audio data of daily activities, a means for comparing the analysis results with past data and assessing the user's risk of decline in physical, psychological, and social functions, and a means for integrating emotional data into the assessment results. This enables early detection of the user's risk of frailty and provides appropriate care and lifestyle improvement suggestions in real time. Furthermore, by taking the user's emotional state into consideration, personalized suggestions can be made, improving the user's quality of life.

[0257] "Daily activities" refers to the various actions and tasks that users perform on a daily basis, including, in particular, movement, conversation, and activities within the home.

[0258] "Video and audio data" means video and audio recordings of your movements and speech, captured through cameras, microphones, and other sensors.

[0259] "Collection means" refers to a combination of devices and software for acquiring data, including, for example, a camera, microphone, data collection application, etc.

[0260] A "computer" is a system consisting of hardware and software for processing and analyzing data, and primarily refers to a server, but also includes cloud computing resources.

[0261] "Encryption" refers to the process of transforming data using a specific algorithm to protect its security. Encrypted data can only be read with a specific key.

[0262] "Decryption" refers to the process of restoring encrypted data to its original state, also known as decoding.

[0263] "Means for analysis" refers to devices and software that use specific algorithms and models to extract meaning from and evaluate acquired data.

[0264] "Extracting movement, posture, and walking speed" refers to the process of recognizing and quantifying the user's physical activity patterns, posture, and walking speed from video data.

[0265] "Evaluating tone, tempo, and frequency of speech" refers to the process of extracting characteristics of a user's speaking style from audio data and analyzing those changes.

[0266] "Emotional data" refers to information about a user's emotional state that can be inferred from their facial expressions and tone of voice, and can include stress, joy, sadness, etc.

[0267] "Means for integrating into evaluation results" refers to devices and software that integrate analyzed emotional data with physical and psychological evaluation results to comprehensively evaluate the user's condition.

[0268] "Generating suggestions for improving care methods and lifestyle habits" refers to the process of recommending specific actions and habits for maintaining health and improving quality of life based on the user's evaluation results.

[0269] "Means of notification" refers to the means by which the system communicates care suggestions generated by the system to the user, including, for example, push notifications, emails, in-app messages, etc.

[0270] MODE FOR CARRYING OUT THE INVENTION

[0271] This invention relates to a system that detects early risk of decline in physical, psychological, and social functions (frailty) in the user's daily life and provides appropriate care. Furthermore, this system incorporates an emotion engine that recognizes the user's emotions.

[0272] overview

[0273] The system uses a device that collects video and audio data from the user's daily life and sends it to a server. The server not only analyzes the received data and assesses the user's risk of frailty, but also recognizes the user's emotions using an emotion engine. Based on the assessment results, the system generates appropriate care methods and lifestyle improvement suggestions and notifies the user. The system also responds to questions and requests for advice from the user and periodically compares new data with past data.

[0274] Data collection and transmission

[0275] The device uses a camera and microphone to collect video and audio data from the user's daily life. For example, it records the user walking around the living room or talking while watching TV. In addition, an emotion engine analyzes the user's facial expressions and tone of voice in real time to recognize emotions. The collected data is encrypted to protect privacy and sent to a server via the Internet. This transmission occurs periodically (for example, every night at midnight), and the data for the day is aggregated on the server.

[0276] Data analysis and evaluation

[0277] The server analyzes the received video and audio data. A multimodal AI model is activated to extract the user's movements, walking speed, and posture from the video data, and evaluate the tone of voice, speech tempo, and frequency of conversation from the audio data. Furthermore, an emotion engine analyzes emotions from the user's facial expressions and tone of voice, thereby collecting emotional data. Based on the analysis results, the server evaluates the risk of decline in the user's physical, psychological, and social functions. For example, by comparing with past data, it checks whether walking speed has slowed or frequency of conversation has decreased. Emotional data is also taken into consideration, and if stress or depression is observed, it determines that there is a high risk.

[0278] Care suggestion generation and notification

[0279] Based on the analysis results, the server generates appropriate care methods and lifestyle improvement suggestions for the user. These include suggestions for appropriate exercise, a balanced diet, and ways to increase social interaction. Emotional data is also taken into consideration, and suggestions are made based on the user's emotional state. The generated care suggestions are sent back to the device, which then notifies the user. For example, a push notification may be displayed saying, "Today's recommended exercise is 30 minutes of walking. We also recommend yoga for relaxation."

[0280] User interaction

[0281] The device has the ability to accept questions and requests for advice from users. When a user inputs a question through the device's interface, the information is sent to the server. The server uses generative AI to generate an appropriate answer to the question and returns it to the device. For example, in response to the question, "Would you like to change your exercise plan for today?", the device may provide an answer such as, "Today's recommended exercise is 30 minutes of walking, but if the weather is bad, we recommend stretching indoors."

[0282] Data collection and comparison

[0283] The device periodically collects new video and audio data. The server compares this new data with past data and continuously updates the user's risk assessment. For example, it analyzes data collected daily to track changes in health and emotional state. Through these steps, the system continuously collects and analyzes data throughout the user's daily life, enabling it to provide appropriate care. It also features a dialogue function using a generative AI model, providing information tailored to the user's needs.

[0284] Specific examples

[0285] Specific examples are shown below.

[0286] Example 1: Video footage of a user walking around the living room and audio data of the user talking are collected. The server analyzes these and compares them with past data to determine that the user's walking speed and frequency of conversation have decreased, indicating an increased risk of frailty.

[0287] Example 2: A user uses a smartphone to input a question. In response to the question, "Would you like to change your exercise routine today?", the server uses weather data to generate a response such as, "If the weather is bad, we recommend stretching indoors."

[0288] Prompt Sentence Examples

[0289] "Please tell me the process flow of a system that collects data on walking and talking in a user's living room. I also heard that the system can recognize the user's emotions. Can you give me a concrete example?"

[0290] "Describe the process by which the server generates an appropriate answer when a user types a question into the terminal. Include examples of questions users might ask about today's exercise."

[0291] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0292] Step 1: Collect data

[0293] 1. The device uses a built-in camera and microphone to collect video and audio data from the user's daily life, and the input is real-world data of the user's actions and conversations.

[0294] 2. The device temporarily stores the collected data in its internal memory. The output is raw video and audio data stored in its internal memory. Specifically, it records the user walking around the living room, watching TV, and talking.

[0295] Step 2: Sentiment Analysis

[0296] 1. The device sends the video and audio data collected in real time to the emotion engine. The input is the video and audio data collected in step 1.

[0297] 2. The emotion engine analyzes the user's facial expressions and tone of voice to recognize their current emotion (happiness, sadness, anger, stress, etc.). The output is data indicating the user's emotion.

[0298] Step 3: Encrypt the data

[0299] 1. The device encrypts the collected video and audio data and emotion analysis results. The inputs are raw video, audio, and emotion data.

[0300] 2. Encryption uses technologies such as AES (Advanced Encryption Standard), and the output is encrypted video, audio, and emotion data.

[0301] Step 4: Sending data

[0302] 1. The terminal sends encrypted data to the server via the Internet. This transmission occurs at midnight every day. The input is the encrypted data.

[0303] 2. The terminal confirms that the transmission was successful. The output is the data that was successfully sent to the server.

[0304] Step 5: Decrypting the received data

[0305] 1. The server decrypts the received encrypted data. The input is the encrypted video, audio, and emotion data.

[0306] 2. The data is restored using the AES decryption process. The output is the original video, audio, and emotion data.

[0307] Step 6: Analyze video and audio data

[0308] 1. The server runs a multimodal AI model to extract user movement, walking speed, and posture from the video data. The input is the decoded video data.

[0309] 2. Perform motion recognition using a deep learning model. The output is extracted user motion, walking speed, and posture data.

[0310] Step 7: Evaluate the audio data

[0311] 1. The server analyzes the audio data and evaluates the tone of voice, speech tempo, and frequency of speech. The input is the decoded audio data.

[0312] 2. Natural language processing (NLP) techniques are used to extract information from speech, and the output is analysis data on tone, tempo, and frequency of speech.

[0313] Step 8: Integrating Emotional Data

[0314] 1. The server integrates the received emotion analysis results with other analysis data. The inputs are emotion analysis data, motion data, and voice data.

[0315] 2. Graph weekly reports and sentiment changes. The output is a consolidated user overall evaluation data.

[0316] Step 9: Risk Assessment

[0317] 1. The server compares the analysis results with past data and evaluates the risk of decline in the user's physical, psychological, and social functions. The input is the integrated assessment data and past data.

[0318] 2. For example, identify slower walking speed or decreased frequency of conversation. The output is the user's risk assessment.

[0319] Step 10: Generate a care recommendation

[0320] 1. The server generates appropriate care methods and lifestyle improvement suggestions for the user based on the assessment results. The input is the risk assessment results.

[0321] 2. The specific suggestion is, for example, "Today's recommended exercise is 30 minutes of walking." The output is the generated care suggestion.

[0322] Step 11: Notification of Care Suggestions

[0323] 1. The server sends the generated care suggestion to the terminal. The input is the generated care suggestion.

[0324] 2. The device notifies the user of the received care suggestion. For example, a push notification might say, "Today's recommended exercise is 30 minutes of walking." The output is the notified care suggestion.

[0325] Step 12: Accepting user questions

[0326] 1. The user enters a question through the device interface, for example, "Are there any changes to today's exercise?"

[0327] 2. The terminal sends the entered question to the server. The input is the user's question. The output is the question sent to the server.

[0328] Step 13: Generate answers to questions

[0329] 1. The server uses a generative AI to generate an appropriate answer. The input is the submitted user question.

[0330] 2. Generate an answer, for example, "Today's recommended exercise is 30 minutes of walking, but if the weather is bad, we recommend stretching indoors." The output is the generated answer.

[0331] Step 14: Notification of Response

[0332] 1. The server sends the generated answer to the terminal. The input is the generated answer.

[0333] 2. The terminal displays the received answer to the user. The output is the notified answer.

[0334] Step 15: Continuous data collection and comparison

[0335] 1. The device periodically collects new video and audio data. The input is data from the user's daily activities.

[0336] 2. The server continuously compares the historical data with the new data and updates the user's risk assessment. The output is an updated risk assessment result.

[0337] (Application example 2)

[0338] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0339] Conventional early detection systems for frailty (risk of decline in physical, psychological, and social function) assess risk by collecting and analyzing video and audio data from users' daily lives, but do not take the user's emotional state into consideration. This makes it difficult to provide more precise care recommendations or lifestyle improvements, and there remains the issue of strengthening security, particularly when it comes to electronic payments.

[0340] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting video and audio data of daily life, means for transmitting the collected video and audio data to the server, means for analyzing the received video and audio data and assessing the risk of decline in the user's physical, psychological, and social functions, means for generating appropriate care methods and lifestyle improvement suggestions for the user based on the assessment results, means for notifying the user of the generated care suggestions, means including an emotion engine that recognizes the user's emotions in real time, and means for making security enhancement and customization suggestions for electronic payments based on the user's emotional state. This makes it possible to enhance security and customize suggestions based on the user's emotional state.

[0341] "Video and audio data of everyday life" refers to video and audio data that includes records of a user's daily activities and conversations.

[0342] "Means of collection" refers to means of obtaining video and audio data using devices such as cameras and microphones.

[0343] The "means for transmitting to the server" is a means for sending collected data to the server via a communication network such as the Internet.

[0344] "Means for analyzing and assessing the risk of decline in the user's physical, psychological, and social functions" refers to a data analysis system and algorithms that use the received data to assess the user's health status from multiple angles.

[0345] The "means for generating appropriate care methods and lifestyle improvement suggestions" is a system that automatically suggests optimal care and improvement measures to users based on the evaluation results.

[0346] "Means of notification" refers to the means for notifying users of the generated care suggestions and improvement measures, and includes push notifications, emails, etc.

[0347] The "real-time emotion recognition engine" is an algorithm and system that analyzes a user's facial expressions and tone of voice to identify their emotional state at that moment.

[0348] The "means for enhancing security and providing customized suggestions during electronic payments" is a system that provides security measures and customized suggestions according to the user's psychological state based on the results of the emotion engine.

[0349] This invention is a system that collects video and audio data from users' daily lives, detects the risk of decline in physical, psychological, and social function (frailty) at an early stage, and provides appropriate care. In addition, by combining it with an emotion engine, it can enhance security and provide customized suggestions during electronic payments based on the user's emotional state.

[0350] Hardware and software used

[0351] The following hardware and software are used to implement this system.

[0352] Hardware:

[0353] Camera and microphone: Collects video and audio data from your daily life.

[0354] Smartphone or PC: Used as a device to collect data and send it to the server.

[0355] software:

[0356] OpenCV: Used for video data processing and face detection.

[0357] Keras: A neural network library for running emotion recognition models.

[0358] SpeechRecognition: A library for converting speech data into text.

[0359] Transformers (Hugging Face): A library for recognizing emotions from speech using the wav2vec2 model.

[0360] Data collection and analysis

[0361] 1. The device uses a camera and microphone to collect video and audio data from everyday life, for example recording you walking around a room or having a conversation with someone.

[0362] 2. The device encrypts the collected video and audio data and sends it to a server via the Internet.

[0363] 3. The server uses OpenCV to detect faces from video data and Keras models to analyze emotions in real time. It also converts audio data to text using SpeechRecognition and uses wav2vec2 models to recognize emotions in the audio.

[0364] Assessment of frailty risk and emotional state

[0365] 1. The server analyzes the received video and audio data to assess the risk of decline in the user's physical, psychological, and social functions. For example, it analyzes the user's walking speed and frequency of conversation and compares them with past data.

[0366] 2. Based on the analysis results, the server generates appropriate care methods and lifestyle improvement suggestions for the user. For example, a suggestion might be, "Today's recommended exercise is 30 minutes of walking. We also recommend yoga for relaxation."

[0367] 3. Based on the results of emotion recognition, the system can also suggest enhanced security measures and customized special offers during electronic payments, for example, requesting additional authentication measures if the user is stressed.

[0368] Notifications and Interactivity

[0369] 1. The generated care suggestions and customization suggestions are sent to the device and notified to the user, for example, by a push notification on a smartphone.

[0370] 2. The device accepts questions and requests for advice from users. The user's questions are sent to the server, which uses a generative AI model to generate appropriate answers and return them to the device.

[0371] Examples of specific examples and prompts

[0372] For example, if a user has their camera and microphone enabled when making an electronic payment, the system will generate the following prompt to perform emotion recognition:

[0373] Example prompt sentence:

[0374] Analyze users' emotions in real time and assess their stress levels.

[0375] This makes it possible to enhance security and provide customized suggestions that take into account the user's emotional state.

[0376] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0377] Step 1:

[0378] The device uses a camera and microphone to collect video and audio data from the user's daily life, for example, recording the user walking around a room or talking to someone. The input of this step is the camera and microphone, and the output is the collected video and audio data.

[0379] Step 2:

[0380] The terminal encrypts the collected video and audio data and transmits it to a server via the Internet. The input of this step is the collected video and audio data, and the output is an encrypted data stream.

[0381] Step 3:

[0382] The server processes the received video data using the OpenCV library to detect faces. Once the face region is detected, it analyzes the emotion in real time using a Keras model. The input of this step is the encrypted video data, and the output is the user's face region and emotional state (e.g., "happy," "sad," "anger," etc.).

[0383] Step 4:

[0384] The server converts the received audio data into text using the SpeechRecognition library and recognizes emotions from the audio using the wav2vec2 model. The input of this step is the encrypted audio data, and the output is the emotional state extracted from the audio.

[0385] Step 5:

[0386] The server integrates the results of the video and audio data analysis to assess the user's risk of decline in physical, psychological, and social function. For example, it analyzes the user's walking speed and frequency of conversation and compares them with past data. The inputs to this step are the analyzed video and audio data and past data, and the output is the user's risk assessment result.

[0387] Step 6:

[0388] Based on the analysis results, the server generates appropriate care methods and lifestyle improvement suggestions. For example, a suggestion might be generated such as, "Today's recommended exercise is 30 minutes of walking. We also recommend yoga for relaxation." The input for this step is the risk assessment result, and the output is the generated care suggestions.

[0389] Step 7:

[0390] The server sends the generated care suggestion to the device. The device notifies the user of the received suggestion, for example, by displaying a push notification on a smartphone. The input of this step is the generated care suggestion, and the output is a notification to the user.

[0391] Step 8:

[0392] Based on the emotion recognition results, the server can then customize the electronic payment process by suggesting enhanced security measures or special offers. For example, if the user is in a stressed state, it can request additional authentication measures. The input to this step is the emotion recognition results, and the output is the enhanced security measures and customized offers.

[0393] Step 9:

[0394] The device accepts questions and requests for advice from the user and sends that information to the server. For example, the user might ask, "Are there any changes to today's exercise?" The input to this step is the user's question or request, and the output is the data sent to the server.

[0395] Step 10:

[0396] The server uses the generative AI model to generate an appropriate answer to the user's question and return it to the device. For example, a response might be, "Today's recommended exercise is 30 minutes of walking, but if the weather is bad, we recommend indoor stretching." The input for this step is the user's question or request, and the output is the generated answer.

[0397] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0398] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0399] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0400] [Second embodiment]

[0401] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0402] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0403] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0404] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0405] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0406] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0407] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0408] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0409] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0410] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0411] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0412] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0413] This invention relates to a system for early detection of risk of decline in physical, psychological, and social functions (frailty) in a user's daily life and providing appropriate care. An embodiment of this system will be described in detail below.

[0414] overview

[0415] The system uses a device to collect video and audio data from the user's daily life and transmit it to a server. The server analyzes the received data and assesses the user's risk of frailty. Based on the results, it generates appropriate care methods and lifestyle improvement suggestions and notifies the user. It also responds to questions and requests for advice from the user and periodically compares new data with past data.

[0416] Data collection and transmission

[0417] The device uses a camera and microphone to collect video and audio data of the user's daily life. For example, the device is installed in the user's living space and automatically records walking, eating, conversations, and other activities that occur there.

[0418] The collected data is encrypted to protect privacy and sent to a server via the Internet. This transmission is carried out periodically (e.g., every night at midnight), and the data for the day is aggregated on the server.

[0419] Data analysis and evaluation

[0420] The server uses a multimodal AI model to analyze the received video and audio data. Specifically, it analyzes the user's movements, walking speed, and posture from the video data, and the tone of voice, speech tempo, and frequency of conversation from the audio data.

[0421] Based on the analysis results, the user's risk of decline in physical, psychological, and social functions is assessed. For example, if a decrease in walking speed or frequency of conversation is observed, it is determined that the risk of frailty is increasing.

[0422] Care suggestion generation and notification

[0423] Based on the analysis results, the server generates appropriate care methods and lifestyle improvement suggestions for the user, including daily exercise, a balanced diet, and ways to socialize.

[0424] The generated care suggestions are then sent back to the device and notified to the user. For example, specific instructions such as "Today's recommended exercise: 30 minutes of walking" are displayed to the user through the device.

[0425] User interaction

[0426] The device also has the ability to accept questions and requests for advice from users. When a user inputs a question through the device's interface, the information is sent to the server. The server uses generative AI to generate an appropriate answer to the question and returns it to the device. For example, in response to the question, "Would you like to change your exercise plan for today?", the device might provide a response such as, "Today's recommended exercise is 30 minutes of walking, but if the weather is bad, we recommend stretching indoors."

[0427] Data collection and comparison

[0428] The device also has the ability to periodically collect new video and audio data and compare it with past data, allowing it to track changes in the user's health and update care recommendations as needed. The device collects data daily, and the server uses this data to continuously assess risk.

[0429] As described above, this system detects users' risk of frailty early and provides appropriate care, enabling elderly people and adults who need preventative care to effectively manage their health at home.

[0430] The processing flow will be explained below.

[0431] Step 1:

[0432] The device collects video and audio data from the user's daily life through a camera and microphone, for example recording the user walking around the living room or talking while watching TV.

[0433] Step 2:

[0434] The device encrypts the collected video and audio data and transmits it securely to a server, taking privacy into consideration. The data is transmitted on a scheduled basis, such as at night or during periods of low load.

[0435] Step 3:

[0436] The server analyzes the received video and audio data, where a multimodal AI model is activated to extract the user's walking speed, posture, and movement patterns from the video data, and evaluate the tone of voice, speech tempo, and frequency of conversation from the audio data.

[0437] Step 4:

[0438] The server evaluates the risk of decline in the user's physical, psychological, and social functions based on the analysis results, for example, by comparing them with past data to see if their walking speed or frequency of conversation has decreased.

[0439] Step 5:

[0440] Based on the assessment results, the server generates appropriate care methods and lifestyle improvement suggestions, such as recommendations for appropriate exercise, a balanced diet, and ways to increase social interaction.

[0441] Step 6:

[0442] The server then sends the generated care recommendations back to the device, which then notifies the user. For example, a push notification will be displayed saying, "Today's recommended exercise is 30 minutes of walking."

[0443] Step 7:

[0444] Users can ask questions or ask for additional advice through the device, for example, "What is a good exercise for a rainy day?"

[0445] Step 8:

[0446] The device sends the user's question to the server, which then uses generative AI to generate an answer to the question. The answer is then sent to the device, which displays it to the user. For example, the device might return an answer such as "We recommend stretching indoors."

[0447] Step 9:

[0448] The device continues to periodically collect new video and audio data. The server compares this new data with past data and continually updates the user's risk assessment. For example, it analyzes the data collected daily to track changes in health status.

[0449] In this way, the system continuously collects and analyzes data throughout the user's daily life and provides appropriate care, making it possible to detect the user's risk of frailty early and respond appropriately.

[0450] Example 1

[0451] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0452] There is a lack of means to detect early risk of decline in physical, psychological, and social function (frailty) at home for elderly people and adults who need preventative care, and to provide appropriate care. Conventional systems often do not adequately collect and analyze data closely related to users' daily lives, resulting in delayed risk assessment. There is also a need for a system that can respond quickly and appropriately to user questions.

[0453] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0454] In this invention, the server includes a means for encrypting video and audio data of daily life and transmitting it to the server, a means for analyzing the received video and audio data using a multimodal AI model to assess the user's risk of decline in physical, psychological, and social function, and a means for generating appropriate care methods and lifestyle improvement suggestions for the user based on the assessment results. This allows users to easily collect data from their daily lives and have it analyzed by the server, enabling early detection of frailty risk and appropriate care. Furthermore, answer generation using the generative AI model enables prompt and appropriate responses to user questions.

[0455] "Video and audio data of everyday life" refers to video and audio data that records the actions and conversations of users in their daily lives.

[0456] "Encryption" is a technology that converts the contents of data into a format that cannot be understood by third parties, and is used particularly to protect privacy and improve security.

[0457] A "server" is a computer system that receives data from client terminals via a network and analyzes and processes the data.

[0458] A "multimodal AI model" is an artificial intelligence model that can comprehensively analyze multiple types of data, such as video data and audio data.

[0459] "Frailty risk" is an indicator of the risk of decline in physical, psychological, and social functioning, and is particularly important when assessing the health status of elderly people and adults who require prevention.

[0460] "Care Suggestions" are advice including specific exercises, balanced meals, and social interaction methods aimed at maintaining the user's health and restoring function.

[0461] A "generative AI model" is an artificial intelligence model that automatically generates appropriate answers or suggestions based on input data.

[0462] "Comparison" is the act of matching new data with past data to detect changes over time.

[0463] "Notification" refers to the act of conveying important information to the user, and in particular, care suggestions and responses from the system.

[0464] This invention relates to a system that detects early risk of decline in physical, psychological, and social function (frailty) in a user's daily life and provides appropriate care. This system can execute a series of processes, from collecting and analyzing video and audio data to generating and notifying care suggestions.

[0465] Data collection and transmission

[0466] The device uses a camera and microphone to collect video and audio data of the user's daily life. The device is equipped with a Python program that uses the OpenCV library to capture video and the PyAudio library to record audio. For example, the device can be installed in the user's living space and automatically record walking, eating, conversations, and other activities.

[0467] The collected data is encrypted using encryption technology such as AES. The encrypted data is periodically sent to a server via the Internet. Data transmission is managed by a regularly executed script such as a Cron job. Specifically, the collected data is encrypted every night at midnight and sent using the HTTPS protocol.

[0468] Data analysis

[0469] The server uses a multimodal AI model to analyze the received data. The multimodal AI model includes a model that comprehensively analyzes the features of video data and audio data (e.g., a model trained using PyTorch). Specifically, the server analyzes the user's walking speed and posture from the video data, and the tone of voice, speech tempo, and frequency of conversation from the audio data.

[0470] Risk assessment and care recommendation generation

[0471] The server evaluates the user's risk of frailty based on the analysis results. If a decrease in walking speed or frequency of conversation is observed, it is determined that the risk of frailty is increasing. Based on the evaluation results, the server generates appropriate care methods and lifestyle improvement suggestions. These suggestions include daily exercise, a balanced diet, and ways to socialize. The generated care suggestions are formatted into text.

[0472] Care Suggestion Notification

[0473] The generated care suggestions are sent back to the device and notified to the user. The suggestions are displayed as push notifications or alerts. For example, a message such as "Today's recommended exercise: 30 minutes of walking" is displayed on the user's device.

[0474] Responding to user questions

[0475] The device has the ability to accept questions and requests for advice from users. When a user inputs a question through the device's interface, the information is sent to the server. The server uses a generative AI model to generate an appropriate answer to the question and returns it to the device. For example, in response to the question, "Would you like to change your exercise plan for today?", the device might provide an answer such as, "Today's recommended exercise is 30 minutes of walking, but if the weather is bad, we recommend stretching indoors."

[0476] Data collection and comparison

[0477] The server periodically collects new data and compares it with past data. This allows the server to track changes in the user's health and update care recommendations as needed. For example, if walking speed data from the past three months is compared and a clear decline is detected, increased exercise is suggested. In a specific scenario, the device collects data daily, and the server continuously performs risk assessments based on this data.

[0478] Prompt Sentence Examples

[0479] "Analyze walking speed data from the past three months and generate appropriate care suggestions if significant decline is observed."

[0480] "If the frequency of conversations during meals decreases, please provide appropriate psychological support."

[0481] As described above, the purpose of this system is to detect early risks of decline in the user's physical, psychological, and social functions in their daily lives and to provide appropriate care.

[0482] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0483] Step 1: Data collection

[0484] The device collects video and audio data from the user's daily life. The input is video and audio data from the camera and microphone. The device uses a Python program to capture video using the OpenCV library and record audio using the PyAudio library. The output is the collected raw data. Specifically, the device activates the camera and microphone every morning at 6:00 and begins recording movements and conversations in the living room and kitchen.

[0485] Step 2: Send data

[0486] The terminal encrypts the collected video and audio data using AES and sends it to the server. The input is the raw data collected in step 1. Data transmission is managed by a regularly executed script such as a Cron job. The output is the encrypted data sent to the server. Specifically, the collected data is encrypted every night at midnight and transmitted using the HTTPS protocol.

[0487] Step 3: Data analysis

[0488] The server analyzes the received data using a multimodal AI model. The input is encrypted video and audio data. The server decrypts the data and extracts features from the video data using OpenCV and the audio data using Librosa. The output is the analysis results. Specifically, the server decompresses the data, inputs it into the AI ​​model for analysis, and extracts features such as walking speed, tone of voice, and frequency of conversation.

[0489] Step 4: Risk assessment

[0490] The server evaluates the user's risk of frailty based on the analysis results. The input is the analysis results obtained in step 3. The server calculates a risk score based on this data and determines the risk level. The output is the risk assessment result. Specifically, the server evaluates a decrease in walking speed and frequency of conversation, and calculates a risk level of "low," "medium," or "high."

[0491] Step 5: Generate a care proposal

[0492] The server generates appropriate care suggestions based on the risk assessment results. The input is a database of risk assessment results and corresponding care suggestions. The server uses the generative AI model to generate suggestions such as exercise, diet, and social interaction. The output is specific care suggestions. Specific operations include the server formatting "daily exercises" in text format based on the AI ​​model.

[0493] Step 6: User Notification

[0494] The server sends the generated care suggestions to the device and notifies the user. The input is the text data of the care suggestions. The server sends the suggestions to the device, and the device displays a notification to the user. The output is the state in which the notification is displayed to the user. Specifically, the server sends content such as "Today's recommended exercise: 30 minutes of walking" to the device, and the device displays it as a push notification.

[0495] Step 7: Respond to user questions

[0496] The device accepts questions and requests for advice from the user. The input is the question text from the user. The device sends this information to the server, which uses a generative AI model to generate an answer. The output is the generated answer. Specifically, the user enters the question, "Would you like to change your exercise routine today?" and the server provides the answer, "If the weather is bad, we recommend stretching indoors."

[0497] Step 8: Data comparison and update

[0498] The server periodically collects new data and compares it with past data. The input is new video and audio data and previously collected data. The server compares these and detects changes in health status. The output is updated care recommendations. Specifically, the server compares walking speed data from the past three months and generates exercise enhancements or new care recommendations as needed, and notifies the user.

[0499] (Application example 1)

[0500] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0501] This invention relates to a system that detects the risk of decline in physical, psychological, and social functions in daily life (frailty) at an early stage and provides appropriate care. The purpose of this invention is to solve the problem of the difficulty of monitoring the health status of customers in real time in a physical store and providing timely care and lifestyle improvement suggestions. It is also required to respond quickly to customers' questions and requests for advice.

[0502] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0503] In this invention, the server includes means for collecting video and audio data of daily life, means for transmitting the collected video and audio data to the server, means for analyzing the received video and audio data and assessing the risk of decline in the user's physical, psychological, and social functions, means for generating appropriate care methods and lifestyle improvement suggestions for the user based on the assessment results, means for notifying the user of the generated care suggestions, and means for monitoring the physical and psychological functions of customers in a physical store in real time using a smart device and assessing health risks. This makes it possible to assess the health status of customers in real time and provide appropriate care and lifestyle improvement suggestions as needed.

[0504] "Video and audio data from everyday life" refers to video footage and audio recordings taken in the user's living environment.

[0505] "Server" refers to a computer system that receives, analyzes, and evaluates collected video and audio data.

[0506] "Means for assessing the risk of decline in physical, psychological, and social functioning" refers to algorithms and programs that analyze data indicating the user's health status and assess the risk of frailty.

[0507] The "means for generating suggestions for improving care methods and lifestyle habits" refers to a means for suggesting to the user care and lifestyle habits improvement methods that should be performed based on the evaluation results.

[0508] "Means for notifying the user of the generated care suggestions" refers to a device or interface for notifying the user of the generated care suggestions in an easy-to-understand manner.

[0509] "Means of monitoring the physical and psychological functions of customers in real time using smart devices in physical stores" refers to smart devices such as cameras and microphones used to observe and record the health status of customers in physical stores, as well as systems for analyzing that data.

[0510] This invention is a system that monitors the health status of customers in a physical store in real time and provides appropriate care and lifestyle improvement suggestions. This invention is realized by collecting video and audio data of daily life using smart devices and sending it to a server.

[0511] System configuration

[0512] Hardware and Software Used

[0513] 1. Smart Devices:

[0514] Using cameras and microphones (e.g., smart glasses), video and audio recordings of customers' everyday activities are collected.

[0515] 2. Server:

[0516] Uses Python, OpenCV, PyAudio, and the Requests library to receive, store, and analyze data.

[0517] A generative AI model is used to analyze the collected data and generate appropriate care methods and improvement suggestions.

[0518] Data collection and transmission

[0519] 1. Data Collection:

[0520] The smart glasses are equipped with a camera and microphone that collects video and audio data of customers, specifically monitoring their walking speed, posture, tone of voice, and frequency of language use.

[0521] 2. Data transmission:

[0522] The collected data is encrypted to protect privacy and sent to a local server at the physical store or to a cloud server.

[0523] Data analysis and evaluation

[0524] The server uses the generative AI model to analyze the received video and audio data. Specifically, it performs the following data analysis:

[0525] 1. Video Data Analysis:

[0526] Analyzes user movements, walking speed, and posture.

[0527] 2. Audio data analysis:

[0528] Analyze tone of voice and frequency of language usage.

[0529] Based on the analysis results, the server evaluates the user's risk of declining physical and psychological function. If the evaluation finds a decrease in walking speed or frequency of conversation, it determines that the risk of frailty is increasing.

[0530] Care suggestion generation and notification

[0531] Based on the analysis results, the server generates appropriate care methods and lifestyle improvement suggestions for the customer, including:

[0532] 1. Exercise Suggestions:

[0533] Generate specific instructions such as "Today's recommended exercise: 30 minutes of walking."

[0534] 2. Dietary suggestions:

[0535] They offer advice such as, "Try to eat a balanced diet."

[0536] The generated care suggestions are communicated to the customer via the smart glasses display.

[0537] User interaction features

[0538] Customers can input questions or requests for advice through the smart glasses interface. This information is sent to the server, where the AI ​​generates an appropriate answer and sends it back to the smart glasses.

[0539] "What exercises would you recommend today?"

[0540] What measures should be taken if walking speed slows down?

[0541] "How would you assess your health status based on today's walking data?"

[0542] Answers to prompts such as:

[0543] Regularly collect and compare data

[0544] New data is collected periodically and compared with past data to track changes in the health status of customers. To achieve this, the smart glasses collect data daily, and the server continuously performs risk assessments based on this data.

[0545] The above is a specific embodiment of the present invention, which makes it possible to evaluate the health status of customers in real time within a physical store and provide appropriate care and lifestyle improvement suggestions.

[0546] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0547] Step 1:

[0548] The device collects video and audio data from the user's daily life.

[0549] Input: Video and audio data captured by the camera and microphone on the smart glasses.

[0550] Data processing and calculation: Use OpenCV to capture video and PyAudio to record audio.

[0551] Output: Collected video data (e.g., JPEG images) and audio data (e.g., WAV files).

[0552] Step 2:

[0553] The terminal transmits the collected video and audio data to the server.

[0554] Input: Video and audio data collected in step 1.

[0555] Data processing and calculation: The data is encrypted to protect privacy and sent to the server using an HTTP POST request.

[0556] Output: Encrypted video and audio data is sent to the server.

[0557] Step 3:

[0558] The server analyzes the received video and audio data and assesses the user's risk of decline in physical and psychological function.

[0559] Input: Encrypted video and audio data sent from the device.

[0560] Data processing and calculation: Generative AI models are used to analyze factors such as walking speed, posture, tone of voice, and frequency of language use.

[0561] Output: Results of the user's frailty risk assessment.

[0562] Step 4:

[0563] Based on the evaluation results, the server generates appropriate care methods and lifestyle improvement suggestions for the user.

[0564] Input: Frailty risk assessment results obtained in Step 3.

[0565] Data processing and calculation: Based on the assessment results, care recommendations such as specific exercises, dietary improvements, and social interaction methods are generated.

[0566] Output: Generated care suggestions (e.g., "Today's recommended exercise: 30 minutes walking").

[0567] Step 5:

[0568] The terminal notifies the user of the generated care suggestions.

[0569] Input: Care suggestions generated in step 4.

[0570] Data processing and calculation: Care suggestions are displayed on the smart glasses display.

[0571] Output: Notification to the user (e.g. "Recommended exercise for today: 30 minutes walking" appears on the display).

[0572] Step 6:

[0573] The user inputs questions or requests for advice through the device interface.

[0574] Input: Text data of a user question or request (e.g., "What exercise would you recommend for me today?").

[0575] Data processing and calculation: The terminal sends this information to the server.

[0576] Output: The user's question is sent to the server.

[0577] Step 7:

[0578] The server uses a generative AI model to generate appropriate answers to the questions and returns them to the device.

[0579] Input: The user's question submitted in step 6.

[0580] Data processing and computation: Generative AI models are used to generate answers to questions.

[0581] Output: The generated answer (e.g., "Today's recommended exercise is 30 minutes of walking, but if the weather is bad, indoor stretching is recommended").

[0582] Step 8:

[0583] The terminal provides the answer from the server to the user.

[0584] Input: The answer from the server generated in step 7.

[0585] Data processing and calculation: The answer is displayed on the smart glasses display.

[0586] Output: Notify the user of the answer (e.g., "Today's recommended exercise is 30 minutes of walking, but if the weather is bad, we recommend stretching indoors.")

[0587] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0588] This invention relates to a system that detects early risk of decline in physical, psychological, and social functions (frailty) in a user's daily life and provides appropriate care, and further combines it with an emotion engine that recognizes the user's emotions. An embodiment of this system will be described in detail below.

[0589] overview

[0590] The system uses a device to collect video and audio data from the user's daily life and send it to a server. The server not only analyzes the received data and assesses the user's risk of frailty, but also recognizes the user's emotions using an emotion engine. Based on the results, it generates appropriate care methods and lifestyle improvement suggestions and notifies the user. It also responds to questions and requests for advice from the user and regularly compares new data with past data.

[0591] Data collection and transmission

[0592] The device uses a camera and microphone to collect video and audio data from the user's daily life, such as when they walk around the living room or talk while watching TV. The emotion engine then analyzes the user's facial expressions and tone of voice in real time to recognize their emotions.

[0593] The collected data is encrypted to protect privacy and sent to a server via the Internet. This transmission is carried out periodically (e.g., every night at midnight), and the data for the day is aggregated on the server.

[0594] Data analysis and evaluation

[0595] The server analyzes the received video and audio data. A multimodal AI model is activated to extract the user's movements, walking speed, and posture from the video data, and evaluate the tone of voice, speech tempo, and frequency of conversation from the audio data. In addition, an emotion engine analyzes the user's emotions from changes in facial expressions and tone of voice. This also allows the collection of user emotional data.

[0596] Based on the analysis results, the system assesses the user's risk of decline in physical, psychological, and social functioning. For example, it compares data from the past to see if their walking speed or frequency of conversation has decreased. Emotional data is also taken into consideration, and if stress or depression is observed, it determines that the risk is high.

[0597] Care suggestion generation and notification

[0598] Based on the analysis results, the server generates appropriate care methods and lifestyle improvement suggestions for the user, including suggestions for appropriate exercise, a balanced diet, and ways to increase social interactions. Emotional data is also taken into account, and suggestions are made based on the user's emotional state.

[0599] The generated care suggestions are then sent back to the device, which then notifies the user. For example, a push notification will be displayed saying, "Today's recommended exercise is 30 minutes of walking. We also recommend yoga for relaxation."

[0600] User interaction

[0601] The device has the ability to accept questions and requests for advice from users. When a user inputs a question through the device's interface, the information is sent to the server. The server uses generative AI to generate an appropriate answer to the question and returns it to the device. For example, in response to the question, "Would you like to change your exercise plan for today?", the device may provide an answer such as, "Today's recommended exercise is 30 minutes of walking, but if the weather is bad, we recommend stretching indoors."

[0602] Data collection and comparison

[0603] The device will continue to periodically collect new video and audio data. The server will compare this new data with past data and continually update the user's risk assessment. For example, it will analyze the data collected daily to track changes in health status and emotional state.

[0604] In this way, this system continuously collects and analyzes data from the user's daily life and provides appropriate care, making it possible to detect the user's risk of frailty early and respond appropriately.Furthermore, by combining it with an emotion engine, it is possible to realize individual responses that take into account the user's emotional state.

[0605] The processing flow will be explained below.

[0606] Step 1:

[0607] The device collects video and audio data from the user's daily life through a camera and microphone, for example, recording the user walking around the living room or talking while watching TV. The emotion engine analyzes the user's facial expressions and tone of voice in real time to recognize their emotions.

[0608] Step 2:

[0609] The device encrypts the collected video and audio data and transmits it securely to a server, taking privacy into consideration. The data is transmitted on a scheduled basis, such as at night or during periods of low load.

[0610] Step 3:

[0611] The server analyzes the received video and audio data. Here, a multimodal AI model is activated, extracting the user's walking speed, posture, and movement patterns from the video data, and evaluating the tone of voice, speech tempo, and frequency of conversation from the audio data. Furthermore, an emotion engine analyzes the user's emotions from changes in facial expressions and tone of voice.

[0612] Step 4:

[0613] The server evaluates the user's risk of decline in physical, psychological, and social function based on the analysis results. For example, it compares the results with past data to see if their walking speed has slowed or if they are talking less frequently. Emotional data is also taken into account, and if stress or depression is observed, it determines that the risk is high.

[0614] Step 5:

[0615] The server generates appropriate care methods and lifestyle improvement suggestions based on the evaluation results. These include suggestions for appropriate exercise, a balanced diet, and ways to increase social interaction. It also takes into account emotional data, suggesting relaxation techniques if the user is feeling stressed, for example.

[0616] Step 6:

[0617] The server then sends the generated care recommendations back to the device, which then notifies the user. For example, a push notification will be displayed saying, "Today's recommended exercise is 30 minutes of walking. We also recommend yoga for relaxation."

[0618] Step 7:

[0619] Users can ask questions or ask for additional advice through the device, for example, "What is a good exercise for a rainy day?"

[0620] Step 8:

[0621] The device sends a question from the user to the server, which then uses generative AI to generate an answer to the question. The answer is then sent to the device, which then displays it to the user. For example, the device might provide an answer such as "We recommend stretching indoors."

[0622] Step 9:

[0623] The device will continue to periodically collect new video and audio data. The server will compare this new data with past data and continually update the user's risk assessment. For example, it will analyze the data collected daily to track changes in health status or emotions.

[0624] In this way, this system continuously collects and analyzes data from the user's daily life and provides appropriate care, making it possible to detect the user's risk of frailty early and respond appropriately.Furthermore, by combining it with an emotion engine, it is possible to realize individual responses that take into account the user's emotional state.

[0625] Example 2

[0626] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0627] In modern society, the aging population is increasing the risk of decline in physical, psychological, and social functions (frailty). Early detection of this risk and the provision of appropriate care and lifestyle improvement suggestions are required, but currently there is a lack of systems that can effectively do this. Individualized care that takes into account the user's emotional state is also required, but systems that can address this are limited. To solve this problem, a system is needed that can continuously collect and analyze data from the user's daily life and provide appropriate care.

[0628] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0629] In this invention, the server includes a means for collecting video and audio data of daily activities, a means for comparing the analysis results with past data and assessing the user's risk of decline in physical, psychological, and social functions, and a means for integrating emotional data into the assessment results. This enables early detection of the user's risk of frailty and provides appropriate care and lifestyle improvement suggestions in real time. Furthermore, by taking the user's emotional state into consideration, personalized suggestions can be made, improving the user's quality of life.

[0630] "Daily activities" refers to the various actions and tasks that users perform on a daily basis, including, in particular, movement, conversation, and activities within the home.

[0631] "Video and audio data" means video and audio recordings of your movements and speech, captured through cameras, microphones, and other sensors.

[0632] "Collection means" refers to a combination of devices and software for acquiring data, including, for example, a camera, microphone, data collection application, etc.

[0633] A "computer" is a system consisting of hardware and software for processing and analyzing data, and primarily refers to a server, but also includes cloud computing resources.

[0634] "Encryption" refers to the process of transforming data using a specific algorithm to protect its security. Encrypted data can only be read with a specific key.

[0635] "Decryption" refers to the process of restoring encrypted data to its original state, also known as decoding.

[0636] "Means for analysis" refers to devices and software that use specific algorithms and models to extract meaning from and evaluate acquired data.

[0637] "Extracting movement, posture, and walking speed" refers to the process of recognizing and quantifying the user's physical activity patterns, posture, and walking speed from video data.

[0638] "Evaluating tone, tempo, and frequency of speech" refers to the process of extracting characteristics of a user's speaking style from audio data and analyzing those changes.

[0639] "Emotional data" refers to information about a user's emotional state that can be inferred from their facial expressions and tone of voice, and can include stress, joy, sadness, etc.

[0640] "Means for integrating into evaluation results" refers to devices and software that integrate analyzed emotional data with physical and psychological evaluation results to comprehensively evaluate the user's condition.

[0641] "Generating suggestions for improving care methods and lifestyle habits" refers to the process of recommending specific actions and habits for maintaining health and improving quality of life based on the user's evaluation results.

[0642] "Means of notification" refers to the means by which the system communicates care suggestions generated by the system to the user, including, for example, push notifications, emails, in-app messages, etc.

[0643] MODE FOR CARRYING OUT THE INVENTION

[0644] This invention relates to a system that detects early risk of decline in physical, psychological, and social functions (frailty) in the user's daily life and provides appropriate care. Furthermore, this system incorporates an emotion engine that recognizes the user's emotions.

[0645] overview

[0646] The system uses a device that collects video and audio data from the user's daily life and sends it to a server. The server not only analyzes the received data and assesses the user's risk of frailty, but also recognizes the user's emotions using an emotion engine. Based on the assessment results, the system generates appropriate care methods and lifestyle improvement suggestions and notifies the user. The system also responds to questions and requests for advice from the user and periodically compares new data with past data.

[0647] Data collection and transmission

[0648] The device uses a camera and microphone to collect video and audio data from the user's daily life. For example, it records the user walking around the living room or talking while watching TV. In addition, an emotion engine analyzes the user's facial expressions and tone of voice in real time to recognize emotions. The collected data is encrypted to protect privacy and sent to a server via the Internet. This transmission occurs periodically (for example, every night at midnight), and the data for the day is aggregated on the server.

[0649] Data analysis and evaluation

[0650] The server analyzes the received video and audio data. A multimodal AI model is activated to extract the user's movements, walking speed, and posture from the video data, and evaluate the tone of voice, speech tempo, and frequency of conversation from the audio data. Furthermore, an emotion engine analyzes emotions from the user's facial expressions and tone of voice, thereby collecting emotional data. Based on the analysis results, the server evaluates the risk of decline in the user's physical, psychological, and social functions. For example, by comparing with past data, it checks whether walking speed has slowed or frequency of conversation has decreased. Emotional data is also taken into consideration, and if stress or depression is observed, it determines that there is a high risk.

[0651] Care suggestion generation and notification

[0652] Based on the analysis results, the server generates appropriate care methods and lifestyle improvement suggestions for the user. These include suggestions for appropriate exercise, a balanced diet, and ways to increase social interaction. Emotional data is also taken into consideration, and suggestions are made based on the user's emotional state. The generated care suggestions are sent back to the device, which then notifies the user. For example, a push notification may be displayed saying, "Today's recommended exercise is 30 minutes of walking. We also recommend yoga for relaxation."

[0653] User interaction

[0654] The device has the ability to accept questions and requests for advice from users. When a user inputs a question through the device's interface, the information is sent to the server. The server uses generative AI to generate an appropriate answer to the question and returns it to the device. For example, in response to the question, "Would you like to change your exercise plan for today?", the device may provide an answer such as, "Today's recommended exercise is 30 minutes of walking, but if the weather is bad, we recommend stretching indoors."

[0655] Data collection and comparison

[0656] The device periodically collects new video and audio data. The server compares this new data with past data and continuously updates the user's risk assessment. For example, it analyzes data collected daily to track changes in health and emotional state. Through these steps, the system continuously collects and analyzes data throughout the user's daily life, enabling it to provide appropriate care. It also features a dialogue function using a generative AI model, providing information tailored to the user's needs.

[0657] Specific examples

[0658] Specific examples are shown below.

[0659] Example 1: Video footage of a user walking around the living room and audio data of the user talking are collected. The server analyzes these and compares them with past data to determine that the user's walking speed and frequency of conversation have decreased, indicating an increased risk of frailty.

[0660] Example 2: A user uses a smartphone to input a question. In response to the question, "Would you like to change your exercise routine today?", the server uses weather data to generate a response such as, "If the weather is bad, we recommend stretching indoors."

[0661] Prompt Sentence Examples

[0662] "Please tell me the process flow of a system that collects data on walking and talking in a user's living room. I also heard that the system can recognize the user's emotions. Can you give me a concrete example?"

[0663] "Describe the process by which the server generates an appropriate answer when a user types a question into the terminal. Include examples of questions users might ask about today's exercise."

[0664] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0665] Step 1: Collect data

[0666] 1. The device uses a built-in camera and microphone to collect video and audio data from the user's daily life, and the input is real-world data of the user's actions and conversations.

[0667] 2. The device temporarily stores the collected data in its internal memory. The output is raw video and audio data stored in its internal memory. Specifically, it records the user walking around the living room, watching TV, and talking.

[0668] Step 2: Sentiment Analysis

[0669] 1. The device sends the video and audio data collected in real time to the emotion engine. The input is the video and audio data collected in step 1.

[0670] 2. The emotion engine analyzes the user's facial expressions and tone of voice to recognize their current emotion (happiness, sadness, anger, stress, etc.). The output is data indicating the user's emotion.

[0671] Step 3: Encrypt the data

[0672] 1. The device encrypts the collected video and audio data and emotion analysis results. The inputs are raw video, audio, and emotion data.

[0673] 2. Encryption uses technologies such as AES (Advanced Encryption Standard), and the output is encrypted video, audio, and emotion data.

[0674] Step 4: Sending data

[0675] 1. The terminal sends encrypted data to the server via the Internet. This transmission occurs at midnight every day. The input is the encrypted data.

[0676] 2. The terminal confirms that the transmission was successful. The output is the data that was successfully sent to the server.

[0677] Step 5: Decrypting the received data

[0678] 1. The server decrypts the received encrypted data. The input is the encrypted video, audio, and emotion data.

[0679] 2. The data is restored using the AES decryption process. The output is the original video, audio, and emotion data.

[0680] Step 6: Analyze video and audio data

[0681] 1. The server runs a multimodal AI model to extract user movement, walking speed, and posture from the video data. The input is the decoded video data.

[0682] 2. Perform motion recognition using a deep learning model. The output is extracted user motion, walking speed, and posture data.

[0683] Step 7: Evaluate the audio data

[0684] 1. The server analyzes the audio data and evaluates the tone of voice, speech tempo, and frequency of speech. The input is the decoded audio data.

[0685] 2. Natural language processing (NLP) techniques are used to extract information from speech, and the output is analysis data on tone, tempo, and frequency of speech.

[0686] Step 8: Integrating Emotional Data

[0687] 1. The server integrates the received emotion analysis results with other analysis data. The inputs are emotion analysis data, motion data, and voice data.

[0688] 2. Graph weekly reports and sentiment changes. The output is a consolidated user overall evaluation data.

[0689] Step 9: Risk Assessment

[0690] 1. The server compares the analysis results with past data and evaluates the user's risk of decline in physical, psychological, and social function. The input is the integrated assessment data and past data.

[0691] 2. For example, identify slower walking speed or decreased frequency of conversation. The output is the user's risk assessment.

[0692] Step 10: Generate a care recommendation

[0693] 1. The server generates appropriate care methods and lifestyle improvement suggestions for the user based on the assessment results. The input is the risk assessment results.

[0694] 2. The specific suggestion is, for example, "Today's recommended exercise is 30 minutes of walking." The output is the generated care suggestion.

[0695] Step 11: Notification of Care Suggestion

[0696] 1. The server sends the generated care suggestion to the terminal. The input is the generated care suggestion.

[0697] 2. The device notifies the user of the received care suggestion. For example, a push notification might say, "Today's recommended exercise is 30 minutes of walking." The output is the notified care suggestion.

[0698] Step 12: Accepting user questions

[0699] 1. The user enters a question through the device interface, for example, "Are there any changes to today's exercise?"

[0700] 2. The terminal sends the entered question to the server. The input is the user's question. The output is the question sent to the server.

[0701] Step 13: Generate answers to questions

[0702] 1. The server uses a generative AI to generate an appropriate answer. The input is the submitted user question.

[0703] 2. Generate an answer, for example, "Today's recommended exercise is 30 minutes of walking, but if the weather is bad, we recommend stretching indoors." The output is the generated answer.

[0704] Step 14: Notification of Response

[0705] 1. The server sends the generated answer to the terminal. The input is the generated answer.

[0706] 2. The terminal displays the received answer to the user. The output is the notified answer.

[0707] Step 15: Continuous data collection and comparison

[0708] 1. The device periodically collects new video and audio data. The input is data from the user's daily activities.

[0709] 2. The server continuously compares the historical data with the new data and updates the user's risk assessment. The output is an updated risk assessment result.

[0710] (Application example 2)

[0711] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0712] Conventional early detection systems for frailty (risk of decline in physical, psychological, and social function) assess risk by collecting and analyzing video and audio data from users' daily lives, but do not take the user's emotional state into consideration. This makes it difficult to provide more precise care recommendations or lifestyle improvements, and there remains the issue of strengthening security, particularly when it comes to electronic payments.

[0713] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting video and audio data of daily life, means for transmitting the collected video and audio data to the server, means for analyzing the received video and audio data and assessing the risk of decline in the user's physical, psychological, and social functions, means for generating appropriate care methods and lifestyle improvement suggestions for the user based on the assessment results, means for notifying the user of the generated care suggestions, means including an emotion engine that recognizes the user's emotions in real time, and means for making security enhancement and customization suggestions for electronic payments based on the user's emotional state. This makes it possible to enhance security and make customization suggestions that take the user's emotional state into consideration.

[0714] "Video and audio data of everyday life" refers to video and audio data that includes records of a user's daily activities and conversations.

[0715] "Means of collection" refers to means of obtaining video and audio data using devices such as cameras and microphones.

[0716] The "means for transmitting to the server" is a means for sending collected data to the server via a communication network such as the Internet.

[0717] "Means for analyzing and assessing the risk of decline in the user's physical, psychological, and social functions" refers to a data analysis system and algorithms that use the received data to assess the user's health status from multiple angles.

[0718] The "means for generating appropriate care methods and lifestyle improvement suggestions" is a system that automatically suggests optimal care and improvement measures to users based on the evaluation results.

[0719] "Means of notification" refers to the means for notifying users of the generated care suggestions and improvement measures, and includes push notifications, emails, etc.

[0720] The "real-time emotion recognition engine" is an algorithm and system that analyzes a user's facial expressions and tone of voice to identify their emotional state at that moment.

[0721] The "means for enhancing security and providing customized suggestions during electronic payments" is a system that provides security measures and customized suggestions according to the user's psychological state based on the results of the emotion engine.

[0722] This invention is a system that collects video and audio data from users' daily lives, detects the risk of decline in physical, psychological, and social function (frailty) at an early stage, and provides appropriate care. In addition, by combining it with an emotion engine, it can enhance security and provide customized suggestions during electronic payments based on the user's emotional state.

[0723] Hardware and software used

[0724] The following hardware and software are used to implement this system.

[0725] Hardware:

[0726] Camera and microphone: Collects video and audio data from your daily life.

[0727] Smartphone or PC: Used as a device to collect data and send it to the server.

[0728] software:

[0729] OpenCV: Used for video data processing and face detection.

[0730] Keras: A neural network library for running emotion recognition models.

[0731] SpeechRecognition: A library for converting speech data into text.

[0732] Transformers (Hugging Face): A library for recognizing emotions from speech using the wav2vec2 model.

[0733] Data collection and analysis

[0734] 1. The device uses a camera and microphone to collect video and audio data from everyday life, for example recording you walking around a room or having a conversation with someone.

[0735] 2. The device encrypts the collected video and audio data and sends it to a server via the Internet.

[0736] 3. The server uses OpenCV to detect faces from video data and Keras models to analyze emotions in real time. It also converts audio data to text using SpeechRecognition and uses wav2vec2 models to recognize emotions in the audio.

[0737] Assessment of frailty risk and emotional state

[0738] 1. The server analyzes the received video and audio data to assess the risk of decline in the user's physical, psychological, and social functions. For example, it analyzes the user's walking speed and frequency of conversation and compares them with past data.

[0739] 2. Based on the analysis results, the server generates appropriate care methods and lifestyle improvement suggestions for the user. For example, a suggestion might be, "Today's recommended exercise is 30 minutes of walking. We also recommend yoga for relaxation."

[0740] 3. Based on the results of emotion recognition, the system can also suggest enhanced security measures and customized special offers during electronic payments, for example, requesting additional authentication measures if the user is stressed.

[0741] Notifications and Interactivity

[0742] 1. The generated care suggestions and customization suggestions are sent to the device and notified to the user, for example, by a push notification on a smartphone.

[0743] 2. The device accepts questions and requests for advice from users. The user's questions are sent to the server, which uses a generative AI model to generate appropriate answers and return them to the device.

[0744] Examples of concrete examples and prompts

[0745] For example, if a user has their camera and microphone enabled when making an electronic payment, the system will generate the following prompt to perform emotion recognition:

[0746] Example prompt sentence:

[0747] Analyze users' emotions in real time and assess their stress levels.

[0748] This makes it possible to enhance security and provide customized suggestions that take into account the user's emotional state.

[0749] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0750] Step 1:

[0751] The device uses a camera and microphone to collect video and audio data from the user's daily life, for example, recording the user walking around a room or talking to someone. The input of this step is the camera and microphone, and the output is the collected video and audio data.

[0752] Step 2:

[0753] The terminal encrypts the collected video and audio data and transmits it to a server via the Internet. The input of this step is the collected video and audio data, and the output is an encrypted data stream.

[0754] Step 3:

[0755] The server processes the received video data using the OpenCV library to detect faces. Once the face region is detected, it analyzes the emotion in real time using a Keras model. The input of this step is the encrypted video data, and the output is the user's face region and emotional state (e.g., "happy," "sad," "anger," etc.).

[0756] Step 4:

[0757] The server converts the received audio data into text using the SpeechRecognition library and recognizes emotions from the audio using the wav2vec2 model. The input of this step is the encrypted audio data, and the output is the emotional state extracted from the audio.

[0758] Step 5:

[0759] The server integrates the results of the video and audio data analysis to assess the user's risk of decline in physical, psychological, and social function. For example, it analyzes the user's walking speed and frequency of conversation and compares them with past data. The inputs to this step are the analyzed video and audio data and past data, and the output is the user's risk assessment result.

[0760] Step 6:

[0761] Based on the analysis results, the server generates appropriate care methods and lifestyle improvement suggestions. For example, a suggestion might be generated such as, "Today's recommended exercise is 30 minutes of walking. We also recommend yoga for relaxation." The input for this step is the risk assessment result, and the output is the generated care suggestions.

[0762] Step 7:

[0763] The server sends the generated care suggestion to the device. The device notifies the user of the received suggestion, for example, by displaying a push notification on a smartphone. The input of this step is the generated care suggestion, and the output is a notification to the user.

[0764] Step 8:

[0765] Based on the emotion recognition results, the server can then customize the electronic payment process by suggesting enhanced security measures or special offers. For example, if the user is in a stressed state, it can request additional authentication measures. The input to this step is the emotion recognition results, and the output is the enhanced security measures and customized offers.

[0766] Step 9:

[0767] The device accepts questions and requests for advice from the user and sends that information to the server. For example, the user might ask, "Are there any changes to today's exercise?" The input to this step is the user's question or request, and the output is the data sent to the server.

[0768] Step 10:

[0769] The server uses the generative AI model to generate an appropriate answer to the user's question and return it to the device. For example, a response might be, "Today's recommended exercise is 30 minutes of walking, but if the weather is bad, we recommend indoor stretching." The input for this step is the user's question or request, and the output is the generated answer.

[0770] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0771] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0772] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0773] [Third embodiment]

[0774] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0775] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0776] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0777] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0778] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0779] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0780] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0781] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0782] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0783] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0784] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0785] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0786] This invention relates to a system for early detection of risk of decline in physical, psychological, and social functions (frailty) in a user's daily life and providing appropriate care. An embodiment of this system will be described in detail below.

[0787] overview

[0788] The system uses a device to collect video and audio data from the user's daily life and transmit it to a server. The server analyzes the received data and assesses the user's risk of frailty. Based on the results, it generates appropriate care methods and lifestyle improvement suggestions and notifies the user. It also responds to questions and requests for advice from the user and periodically compares new data with past data.

[0789] Data collection and transmission

[0790] The device uses a camera and microphone to collect video and audio data of the user's daily life. For example, the device is installed in the user's living space and automatically records walking, eating, conversations, and other activities that occur there.

[0791] The collected data is encrypted to protect privacy and sent to a server via the Internet. This transmission is carried out periodically (e.g., every night at midnight), and the data for the day is aggregated on the server.

[0792] Data analysis and evaluation

[0793] The server uses a multimodal AI model to analyze the received video and audio data. Specifically, it analyzes the user's movements, walking speed, and posture from the video data, and the tone of voice, speech tempo, and frequency of conversation from the audio data.

[0794] Based on the analysis results, the user's risk of decline in physical, psychological, and social functions is assessed. For example, if a decrease in walking speed or frequency of conversation is observed, it is determined that the risk of frailty is increasing.

[0795] Care suggestion generation and notification

[0796] Based on the analysis results, the server generates appropriate care methods and lifestyle improvement suggestions for the user, including daily exercise, a balanced diet, and ways to socialize.

[0797] The generated care suggestions are then sent back to the device and notified to the user. For example, specific instructions such as "Today's recommended exercise: 30 minutes of walking" are displayed to the user through the device.

[0798] User interaction

[0799] The device also has the ability to accept questions and requests for advice from users. When a user inputs a question through the device's interface, the information is sent to the server. The server uses generative AI to generate an appropriate answer to the question and returns it to the device. For example, in response to the question, "Would you like to change your exercise plan for today?", the device might provide a response such as, "Today's recommended exercise is 30 minutes of walking, but if the weather is bad, we recommend stretching indoors."

[0800] Data collection and comparison

[0801] The device also has the ability to periodically collect new video and audio data and compare it with past data, allowing it to track changes in the user's health and update care recommendations as needed. The device collects data daily, and the server uses this data to continuously assess risk.

[0802] As described above, this system detects users' risk of frailty early and provides appropriate care, enabling elderly people and adults who need preventative care to effectively manage their health at home.

[0803] The processing flow will be explained below.

[0804] Step 1:

[0805] The device collects video and audio data from the user's daily life through a camera and microphone, for example recording the user walking around the living room or talking while watching TV.

[0806] Step 2:

[0807] The device encrypts the collected video and audio data and transmits it securely to a server, taking privacy into consideration. The data is transmitted on a scheduled basis, such as at night or during periods of low load.

[0808] Step 3:

[0809] The server analyzes the received video and audio data, where a multimodal AI model is activated to extract the user's walking speed, posture, and movement patterns from the video data, and evaluate the tone of voice, speech tempo, and frequency of conversation from the audio data.

[0810] Step 4:

[0811] The server evaluates the risk of decline in the user's physical, psychological, and social functions based on the analysis results, for example, by comparing them with past data to see if their walking speed or frequency of conversation has decreased.

[0812] Step 5:

[0813] Based on the assessment results, the server generates appropriate care methods and lifestyle improvement suggestions, such as recommendations for appropriate exercise, a balanced diet, and ways to increase social interaction.

[0814] Step 6:

[0815] The server then sends the generated care recommendations back to the device, which then notifies the user. For example, a push notification will be displayed saying, "Today's recommended exercise is 30 minutes of walking."

[0816] Step 7:

[0817] Users can ask questions or ask for additional advice through the device, for example, "What is a good exercise for a rainy day?"

[0818] Step 8:

[0819] The device sends the user's question to the server, which then uses generative AI to generate an answer to the question. The answer is then sent to the device, which displays it to the user. For example, the device might return an answer such as "We recommend stretching indoors."

[0820] Step 9:

[0821] The device continues to periodically collect new video and audio data. The server compares this new data with past data and continually updates the user's risk assessment. For example, it analyzes the data collected daily to track changes in health status.

[0822] In this way, the system continuously collects and analyzes data throughout the user's daily life and provides appropriate care, making it possible to detect the user's risk of frailty early and respond appropriately.

[0823] Example 1

[0824] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0825] There is a lack of means to detect early risk of decline in physical, psychological, and social function (frailty) at home for elderly people and adults who need preventative care, and to provide appropriate care. Conventional systems often do not adequately collect and analyze data closely related to users' daily lives, resulting in delayed risk assessment. There is also a need for a system that can respond quickly and appropriately to user questions.

[0826] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0827] In this invention, the server includes a means for encrypting video and audio data of daily life and transmitting it to the server, a means for analyzing the received video and audio data using a multimodal AI model to assess the user's risk of decline in physical, psychological, and social function, and a means for generating appropriate care methods and lifestyle improvement suggestions for the user based on the assessment results. This allows users to easily collect data from their daily lives and have it analyzed by the server, enabling early detection of frailty risk and appropriate care. Furthermore, answer generation using the generative AI model enables prompt and appropriate responses to user questions.

[0828] "Video and audio data of everyday life" refers to video and audio data that records the actions and conversations of users in their daily lives.

[0829] "Encryption" is a technology that converts the contents of data into a format that cannot be understood by third parties, and is used particularly to protect privacy and improve security.

[0830] A "server" is a computer system that receives data from client terminals via a network and analyzes and processes the data.

[0831] A "multimodal AI model" is an artificial intelligence model that can comprehensively analyze multiple types of data, such as video data and audio data.

[0832] "Frailty risk" is an indicator of the risk of decline in physical, psychological, and social functioning, and is particularly important when assessing the health status of elderly people and adults who require prevention.

[0833] "Care Suggestions" are advice including specific exercises, balanced meals, and social interaction methods aimed at maintaining the user's health and restoring function.

[0834] A "generative AI model" is an artificial intelligence model that automatically generates appropriate answers or suggestions based on input data.

[0835] "Comparison" is the act of matching new data with past data to detect changes over time.

[0836] "Notification" refers to the act of conveying important information to the user, and in particular, care suggestions and responses from the system.

[0837] This invention relates to a system that detects early risk of decline in physical, psychological, and social function (frailty) in a user's daily life and provides appropriate care. This system can execute a series of processes, from collecting and analyzing video and audio data to generating and notifying care suggestions.

[0838] Data collection and transmission

[0839] The device uses a camera and microphone to collect video and audio data of the user's daily life. The device is equipped with a Python program that uses the OpenCV library to capture video and the PyAudio library to record audio. For example, the device can be installed in the user's living space and automatically record walking, eating, conversations, and other activities.

[0840] The collected data is encrypted using encryption technology such as AES. The encrypted data is periodically sent to a server via the Internet. Data transmission is managed by a regularly executed script such as a Cron job. Specifically, the collected data is encrypted every night at midnight and sent using the HTTPS protocol.

[0841] Data analysis

[0842] The server uses a multimodal AI model to analyze the received data. The multimodal AI model includes a model that comprehensively analyzes the features of video data and audio data (e.g., a model trained using PyTorch). Specifically, the server analyzes the user's walking speed and posture from the video data, and the tone of voice, speech tempo, and frequency of conversation from the audio data.

[0843] Risk assessment and care recommendation generation

[0844] The server evaluates the user's risk of frailty based on the analysis results. If a decrease in walking speed or frequency of conversation is observed, it is determined that the risk of frailty is increasing. Based on the evaluation results, the server generates appropriate care methods and lifestyle improvement suggestions. These suggestions include daily exercise, a balanced diet, and ways to socialize. The generated care suggestions are formatted into text.

[0845] Care Suggestion Notification

[0846] The generated care suggestions are sent back to the device and notified to the user. The suggestions are displayed as push notifications or alerts. For example, a message such as "Today's recommended exercise: 30 minutes of walking" is displayed on the user's device.

[0847] Responding to user questions

[0848] The device has the ability to accept questions and requests for advice from users. When a user inputs a question through the device's interface, the information is sent to the server. The server uses a generative AI model to generate an appropriate answer to the question and returns it to the device. For example, in response to the question, "Would you like to change your exercise plan for today?", the device might provide an answer such as, "Today's recommended exercise is 30 minutes of walking, but if the weather is bad, we recommend stretching indoors."

[0849] Data collection and comparison

[0850] The server periodically collects new data and compares it with past data. This allows the server to track changes in the user's health and update care recommendations as needed. For example, if walking speed data from the past three months is compared and a clear decline is detected, increased exercise is suggested. In a specific scenario, the device collects data daily, and the server continuously performs risk assessments based on this data.

[0851] Prompt Sentence Examples

[0852] "Analyze walking speed data from the past three months and generate appropriate care suggestions if significant decline is observed."

[0853] "If the frequency of conversations during meals decreases, please provide appropriate psychological support."

[0854] As described above, the purpose of this system is to detect early risks of decline in the user's physical, psychological, and social functions in their daily lives and to provide appropriate care.

[0855] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0856] Step 1: Data collection

[0857] The device collects video and audio data from the user's daily life. The input is video and audio data from the camera and microphone. The device uses a Python program to capture video using the OpenCV library and record audio using the PyAudio library. The output is the collected raw data. Specifically, the device activates the camera and microphone every morning at 6:00 and begins recording movements and conversations in the living room and kitchen.

[0858] Step 2: Send data

[0859] The terminal encrypts the collected video and audio data using AES and sends it to the server. The input is the raw data collected in step 1. Data transmission is managed by a regularly executed script such as a Cron job. The output is the encrypted data sent to the server. Specifically, the collected data is encrypted every night at midnight and transmitted using the HTTPS protocol.

[0860] Step 3: Data analysis

[0861] The server analyzes the received data using a multimodal AI model. The input is encrypted video and audio data. The server decrypts the data and extracts features from the video data using OpenCV and the audio data using Librosa. The output is the analysis results. Specifically, the server decompresses the data, inputs it into the AI ​​model for analysis, and extracts features such as walking speed, tone of voice, and frequency of conversation.

[0862] Step 4: Risk assessment

[0863] The server evaluates the user's risk of frailty based on the analysis results. The input is the analysis results obtained in step 3. The server calculates a risk score based on this data and determines the risk level. The output is the risk assessment result. Specifically, the server evaluates a decrease in walking speed and frequency of conversation, and calculates a risk level of "low," "medium," or "high."

[0864] Step 5: Generate a care proposal

[0865] The server generates appropriate care suggestions based on the risk assessment results. The input is a database of risk assessment results and corresponding care suggestions. The server uses the generative AI model to generate suggestions such as exercise, diet, and social interaction. The output is specific care suggestions. Specific operations include the server formatting "daily exercises" in text format based on the AI ​​model.

[0866] Step 6: User Notification

[0867] The server sends the generated care suggestions to the device and notifies the user. The input is the text data of the care suggestions. The server sends the suggestions to the device, and the device displays a notification to the user. The output is the state in which the notification is displayed to the user. Specifically, the server sends content such as "Today's recommended exercise: 30 minutes of walking" to the device, and the device displays it as a push notification.

[0868] Step 7: Respond to user questions

[0869] The device accepts questions and requests for advice from the user. The input is the question text from the user. The device sends this information to the server, which uses a generative AI model to generate an answer. The output is the generated answer. Specifically, the user enters the question, "Would you like to change your exercise routine today?" and the server provides the answer, "If the weather is bad, we recommend stretching indoors."

[0870] Step 8: Data comparison and update

[0871] The server periodically collects new data and compares it with past data. The input is new video and audio data and previously collected data. The server compares these and detects changes in health status. The output is updated care recommendations. Specifically, the server compares walking speed data from the past three months and generates exercise enhancements or new care recommendations as needed, and notifies the user.

[0872] (Application example 1)

[0873] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0874] This invention relates to a system that detects the risk of decline in physical, psychological, and social functions in daily life (frailty) at an early stage and provides appropriate care. The purpose of this invention is to solve the problem of the difficulty of monitoring the health status of customers in real time in a physical store and providing timely care and lifestyle improvement suggestions. It is also required to respond quickly to customers' questions and requests for advice.

[0875] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0876] In this invention, the server includes means for collecting video and audio data of daily life, means for transmitting the collected video and audio data to the server, means for analyzing the received video and audio data and assessing the risk of decline in the user's physical, psychological, and social functions, means for generating appropriate care methods and lifestyle improvement suggestions for the user based on the assessment results, means for notifying the user of the generated care suggestions, and means for monitoring the physical and psychological functions of customers in a physical store in real time using a smart device and assessing health risks. This makes it possible to assess the health status of customers in real time and provide appropriate care and lifestyle improvement suggestions as needed.

[0877] "Video and audio data from everyday life" refers to video footage and audio recordings taken in the user's living environment.

[0878] "Server" refers to a computer system that receives, analyzes, and evaluates collected video and audio data.

[0879] "Means for assessing the risk of decline in physical, psychological, and social functioning" refers to algorithms and programs that analyze data indicating the user's health status and assess the risk of frailty.

[0880] The "means for generating suggestions for improving care methods and lifestyle habits" refers to a means for suggesting to the user care and lifestyle habits improvement methods that should be performed based on the evaluation results.

[0881] "Means for notifying the user of the generated care suggestions" refers to a device or interface for notifying the user of the generated care suggestions in an easy-to-understand manner.

[0882] "Means of monitoring the physical and psychological functions of customers in real time using smart devices in physical stores" refers to smart devices such as cameras and microphones used to observe and record the health status of customers in physical stores, as well as systems for analyzing that data.

[0883] This invention is a system that monitors the health status of customers in a physical store in real time and provides appropriate care and lifestyle improvement suggestions. This invention is realized by collecting video and audio data of daily life using smart devices and sending it to a server.

[0884] System configuration

[0885] Hardware and Software Used

[0886] 1. Smart Devices:

[0887] Using cameras and microphones (e.g., smart glasses), video and audio recordings of customers' everyday activities are collected.

[0888] 2. Server:

[0889] Uses Python, OpenCV, PyAudio, and the Requests library to receive, store, and analyze data.

[0890] A generative AI model is used to analyze the collected data and generate appropriate care methods and improvement suggestions.

[0891] Data collection and transmission

[0892] 1. Data Collection:

[0893] The smart glasses are equipped with a camera and microphone that collects video and audio data of customers, specifically monitoring their walking speed, posture, tone of voice, and frequency of language use.

[0894] 2. Data transmission:

[0895] The collected data is encrypted to protect privacy and sent to a local server at the physical store or to a cloud server.

[0896] Data analysis and evaluation

[0897] The server uses the generative AI model to analyze the received video and audio data. Specifically, it performs the following data analysis:

[0898] 1. Video Data Analysis:

[0899] Analyzes user movements, walking speed, and posture.

[0900] 2. Audio data analysis:

[0901] Analyze tone of voice and frequency of language usage.

[0902] Based on the analysis results, the server evaluates the user's risk of declining physical and psychological function. If the evaluation finds a decrease in walking speed or frequency of conversation, it determines that the risk of frailty is increasing.

[0903] Care suggestion generation and notification

[0904] Based on the analysis results, the server generates appropriate care methods and lifestyle improvement suggestions for the customer, including:

[0905] 1. Exercise Suggestions:

[0906] Generate specific instructions such as "Today's recommended exercise: 30 minutes of walking."

[0907] 2. Dietary suggestions:

[0908] They offer advice such as, "Try to eat a balanced diet."

[0909] The generated care suggestions are communicated to the customer via the smart glasses display.

[0910] User interaction features

[0911] Customers can input questions or requests for advice through the smart glasses interface. This information is sent to the server, where the AI ​​generates an appropriate answer and sends it back to the smart glasses.

[0912] "What exercises would you recommend today?"

[0913] What measures should be taken if walking speed slows down?

[0914] "How would you assess your health status based on today's walking data?"

[0915] Answers to prompts such as:

[0916] Regularly collect and compare data

[0917] New data is collected periodically and compared with past data to track changes in the health status of customers. To achieve this, the smart glasses collect data daily, and the server continuously performs risk assessments based on this data.

[0918] The above is a specific embodiment of the present invention, which makes it possible to evaluate the health status of customers in real time within a physical store and provide appropriate care and lifestyle improvement suggestions.

[0919] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0920] Step 1:

[0921] The device collects video and audio data from the user's daily life.

[0922] Input: Video and audio data captured by the camera and microphone on the smart glasses.

[0923] Data processing and calculation: Use OpenCV to capture video and PyAudio to record audio.

[0924] Output: Collected video data (e.g., JPEG images) and audio data (e.g., WAV files).

[0925] Step 2:

[0926] The terminal transmits the collected video and audio data to the server.

[0927] Input: Video and audio data collected in step 1.

[0928] Data processing and calculation: The data is encrypted to protect privacy and sent to the server using an HTTP POST request.

[0929] Output: Encrypted video and audio data is sent to the server.

[0930] Step 3:

[0931] The server analyzes the received video and audio data and assesses the user's risk of decline in physical and psychological function.

[0932] Input: Encrypted video and audio data sent from the device.

[0933] Data processing and calculation: Generative AI models are used to analyze factors such as walking speed, posture, tone of voice, and frequency of language use.

[0934] Output: Results of the user's frailty risk assessment.

[0935] Step 4:

[0936] Based on the evaluation results, the server generates appropriate care methods and lifestyle improvement suggestions for the user.

[0937] Input: Frailty risk assessment results obtained in Step 3.

[0938] Data processing and calculation: Based on the assessment results, care recommendations such as specific exercises, dietary improvements, and social interaction methods are generated.

[0939] Output: Generated care suggestions (e.g., "Today's recommended exercise: 30 minutes walking").

[0940] Step 5:

[0941] The terminal notifies the user of the generated care suggestions.

[0942] Input: Care suggestions generated in step 4.

[0943] Data processing and calculation: Care suggestions are displayed on the smart glasses display.

[0944] Output: Notification to the user (e.g. "Recommended exercise for today: 30 minutes walking" appears on the display).

[0945] Step 6:

[0946] The user inputs questions or requests for advice through the device interface.

[0947] Input: Text data of a user question or request (e.g., "What exercise would you recommend for me today?").

[0948] Data processing and calculation: The terminal sends this information to the server.

[0949] Output: The user's question is sent to the server.

[0950] Step 7:

[0951] The server uses a generative AI model to generate appropriate answers to the questions and returns them to the device.

[0952] Input: The user's question submitted in step 6.

[0953] Data processing and computation: Generative AI models are used to generate answers to questions.

[0954] Output: The generated answer (e.g., "Today's recommended exercise is 30 minutes of walking, but if the weather is bad, indoor stretching is recommended").

[0955] Step 8:

[0956] The terminal provides the answer from the server to the user.

[0957] Input: The answer from the server generated in step 7.

[0958] Data processing and calculation: The answer is displayed on the smart glasses display.

[0959] Output: Notify the user of the answer (e.g., "Today's recommended exercise is 30 minutes of walking, but if the weather is bad, we recommend stretching indoors.")

[0960] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0961] This invention relates to a system that detects early risk of decline in physical, psychological, and social functions (frailty) in a user's daily life and provides appropriate care, and further combines it with an emotion engine that recognizes the user's emotions. An embodiment of this system will be described in detail below.

[0962] overview

[0963] The system uses a device to collect video and audio data from the user's daily life and send it to a server. The server not only analyzes the received data and assesses the user's risk of frailty, but also recognizes the user's emotions using an emotion engine. Based on the results, it generates appropriate care methods and lifestyle improvement suggestions and notifies the user. It also responds to questions and requests for advice from the user and regularly compares new data with past data.

[0964] Data collection and transmission

[0965] The device uses a camera and microphone to collect video and audio data from the user's daily life, such as when they walk around the living room or talk while watching TV. The emotion engine then analyzes the user's facial expressions and tone of voice in real time to recognize their emotions.

[0966] The collected data is encrypted to protect privacy and sent to a server via the Internet. This transmission is carried out periodically (e.g., every night at midnight), and the data for the day is aggregated on the server.

[0967] Data analysis and evaluation

[0968] The server analyzes the received video and audio data. A multimodal AI model is activated to extract the user's movements, walking speed, and posture from the video data, and evaluate the tone of voice, speech tempo, and frequency of conversation from the audio data. In addition, an emotion engine analyzes the user's emotions from changes in facial expressions and tone of voice. This also allows the collection of user emotional data.

[0969] Based on the analysis results, the system assesses the user's risk of decline in physical, psychological, and social functioning. For example, it compares data from the past to see if their walking speed or frequency of conversation has decreased. Emotional data is also taken into consideration, and if stress or depression is observed, it determines that the risk is high.

[0970] Care suggestion generation and notification

[0971] Based on the analysis results, the server generates appropriate care methods and lifestyle improvement suggestions for the user, including suggestions for appropriate exercise, a balanced diet, and ways to increase social interactions. Emotional data is also taken into account, and suggestions are made based on the user's emotional state.

[0972] The generated care suggestions are then sent back to the device, which then notifies the user. For example, a push notification will be displayed saying, "Today's recommended exercise is 30 minutes of walking. We also recommend yoga for relaxation."

[0973] User interaction

[0974] The device has the ability to accept questions and requests for advice from users. When a user inputs a question through the device's interface, the information is sent to the server. The server uses generative AI to generate an appropriate answer to the question and returns it to the device. For example, in response to the question, "Would you like to change your exercise plan for today?", the device may provide an answer such as, "Today's recommended exercise is 30 minutes of walking, but if the weather is bad, we recommend stretching indoors."

[0975] Data collection and comparison

[0976] The device will continue to periodically collect new video and audio data. The server will compare this new data with past data and continually update the user's risk assessment. For example, it will analyze the data collected daily to track changes in health status and emotional state.

[0977] In this way, this system continuously collects and analyzes data from the user's daily life and provides appropriate care, making it possible to detect the user's risk of frailty early and respond appropriately.Furthermore, by combining it with an emotion engine, it is possible to realize individual responses that take into account the user's emotional state.

[0978] The processing flow will be explained below.

[0979] Step 1:

[0980] The device collects video and audio data from the user's daily life through a camera and microphone, for example, recording the user walking around the living room or talking while watching TV. The emotion engine analyzes the user's facial expressions and tone of voice in real time to recognize their emotions.

[0981] Step 2:

[0982] The device encrypts the collected video and audio data and transmits it securely to a server, taking privacy into consideration. The data is transmitted on a scheduled basis, such as at night or during periods of low load.

[0983] Step 3:

[0984] The server analyzes the received video and audio data. Here, a multimodal AI model is activated, extracting the user's walking speed, posture, and movement patterns from the video data, and evaluating the tone of voice, speech tempo, and frequency of conversation from the audio data. Furthermore, an emotion engine analyzes the user's emotions from changes in facial expressions and tone of voice.

[0985] Step 4:

[0986] The server evaluates the user's risk of decline in physical, psychological, and social function based on the analysis results. For example, it compares the results with past data to see if their walking speed has slowed or if they are talking less frequently. Emotional data is also taken into account, and if stress or depression is observed, it determines that the risk is high.

[0987] Step 5:

[0988] The server generates appropriate care methods and lifestyle improvement suggestions based on the evaluation results. These include suggestions for appropriate exercise, a balanced diet, and ways to increase social interaction. It also takes into account emotional data, suggesting relaxation techniques if the user is feeling stressed, for example.

[0989] Step 6:

[0990] The server then sends the generated care recommendations back to the device, which then notifies the user. For example, a push notification will be displayed saying, "Today's recommended exercise is 30 minutes of walking. We also recommend yoga for relaxation."

[0991] Step 7:

[0992] Users can ask questions or ask for additional advice through the device, for example, "What is a good exercise for a rainy day?"

[0993] Step 8:

[0994] The device sends a question from the user to the server, which then uses generative AI to generate an answer to the question. The answer is then sent to the device, which then displays it to the user. For example, the device might provide an answer such as "We recommend stretching indoors."

[0995] Step 9:

[0996] The device will continue to periodically collect new video and audio data. The server will compare this new data with past data and continually update the user's risk assessment. For example, it will analyze the data collected daily to track changes in health status or emotions.

[0997] In this way, this system continuously collects and analyzes data from the user's daily life and provides appropriate care, making it possible to detect the user's risk of frailty early and respond appropriately.Furthermore, by combining it with an emotion engine, it is possible to realize individual responses that take into account the user's emotional state.

[0998] Example 2

[0999] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1000] In modern society, the aging population is increasing the risk of decline in physical, psychological, and social functions (frailty). Early detection of this risk and the provision of appropriate care and lifestyle improvement suggestions are required, but currently there is a lack of systems that can effectively do this. Individualized care that takes into account the user's emotional state is also required, but systems that can address this are limited. To solve this problem, a system is needed that can continuously collect and analyze data from the user's daily life and provide appropriate care.

[1001] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1002] In this invention, the server includes a means for collecting video and audio data of daily activities, a means for comparing the analysis results with past data and assessing the user's risk of decline in physical, psychological, and social functions, and a means for integrating emotional data into the assessment results. This enables early detection of the user's risk of frailty and provides appropriate care and lifestyle improvement suggestions in real time. Furthermore, by taking the user's emotional state into consideration, personalized suggestions can be made, improving the user's quality of life.

[1003] "Daily activities" refers to the various actions and tasks that users perform on a daily basis, including, in particular, movement, conversation, and activities within the home.

[1004] "Video and audio data" means video and audio recordings of your movements and speech, captured through cameras, microphones, and other sensors.

[1005] "Collection means" refers to a combination of devices and software for acquiring data, including, for example, a camera, microphone, data collection application, etc.

[1006] A "computer" is a system consisting of hardware and software for processing and analyzing data, and primarily refers to a server, but also includes cloud computing resources.

[1007] "Encryption" refers to the process of transforming data using a specific algorithm to protect its security. Encrypted data can only be read with a specific key.

[1008] "Decryption" refers to the process of restoring encrypted data to its original state, also known as decoding.

[1009] "Means for analysis" refers to devices and software that use specific algorithms and models to extract meaning from and evaluate acquired data.

[1010] "Extracting movement, posture, and walking speed" refers to the process of recognizing and quantifying the user's physical activity patterns, posture, and walking speed from video data.

[1011] "Evaluating tone, tempo, and frequency of speech" refers to the process of extracting characteristics of a user's speaking style from audio data and analyzing those changes.

[1012] "Emotional data" refers to information about a user's emotional state that can be inferred from their facial expressions and tone of voice, and can include stress, joy, sadness, etc.

[1013] "Means for integrating into evaluation results" refers to devices and software that integrate analyzed emotional data with physical and psychological evaluation results to comprehensively evaluate the user's condition.

[1014] "Generating suggestions for improving care methods and lifestyle habits" refers to the process of recommending specific actions and habits for maintaining health and improving quality of life based on the user's evaluation results.

[1015] "Means of notification" refers to the means by which the system communicates care suggestions generated by the system to the user, including, for example, push notifications, emails, in-app messages, etc.

[1016] MODE FOR CARRYING OUT THE INVENTION

[1017] This invention relates to a system that detects early risk of decline in physical, psychological, and social functions (frailty) in the user's daily life and provides appropriate care. Furthermore, this system incorporates an emotion engine that recognizes the user's emotions.

[1018] overview

[1019] The system uses a device that collects video and audio data from the user's daily life and sends it to a server. The server not only analyzes the received data and assesses the user's risk of frailty, but also recognizes the user's emotions using an emotion engine. Based on the assessment results, the system generates appropriate care methods and lifestyle improvement suggestions and notifies the user. The system also responds to questions and requests for advice from the user and periodically compares new data with past data.

[1020] Data collection and transmission

[1021] The device uses a camera and microphone to collect video and audio data from the user's daily life. For example, it records the user walking around the living room or talking while watching TV. In addition, an emotion engine analyzes the user's facial expressions and tone of voice in real time to recognize emotions. The collected data is encrypted to protect privacy and sent to a server via the Internet. This transmission occurs periodically (for example, every night at midnight), and the data for the day is aggregated on the server.

[1022] Data analysis and evaluation

[1023] The server analyzes the received video and audio data. A multimodal AI model is activated to extract the user's movements, walking speed, and posture from the video data, and evaluate the tone of voice, speech tempo, and frequency of conversation from the audio data. Furthermore, an emotion engine analyzes emotions from the user's facial expressions and tone of voice, thereby collecting emotional data. Based on the analysis results, the server evaluates the risk of decline in the user's physical, psychological, and social functions. For example, by comparing with past data, it checks whether walking speed has slowed or frequency of conversation has decreased. Emotional data is also taken into consideration, and if stress or depression is observed, it determines that there is a high risk.

[1024] Care suggestion generation and notification

[1025] Based on the analysis results, the server generates appropriate care methods and lifestyle improvement suggestions for the user. These include suggestions for appropriate exercise, a balanced diet, and ways to increase social interaction. Emotional data is also taken into consideration, and suggestions are made based on the user's emotional state. The generated care suggestions are sent back to the device, which then notifies the user. For example, a push notification may be displayed saying, "Today's recommended exercise is 30 minutes of walking. We also recommend yoga for relaxation."

[1026] User interaction

[1027] The device has the ability to accept questions and requests for advice from users. When a user inputs a question through the device's interface, the information is sent to the server. The server uses generative AI to generate an appropriate answer to the question and returns it to the device. For example, in response to the question, "Would you like to change your exercise plan for today?", the device may provide an answer such as, "Today's recommended exercise is 30 minutes of walking, but if the weather is bad, we recommend stretching indoors."

[1028] Data collection and comparison

[1029] The device periodically collects new video and audio data. The server compares this new data with past data and continuously updates the user's risk assessment. For example, it analyzes data collected daily to track changes in health and emotional state. Through these steps, the system continuously collects and analyzes data throughout the user's daily life, enabling it to provide appropriate care. It also features a dialogue function using a generative AI model, providing information tailored to the user's needs.

[1030] Specific examples

[1031] Specific examples are shown below.

[1032] Example 1: Video footage of a user walking around the living room and audio data of the user talking are collected. The server analyzes these and compares them with past data to determine that the user's walking speed and frequency of conversation have decreased, indicating an increased risk of frailty.

[1033] Example 2: A user uses a smartphone to input a question. In response to the question, "Would you like to change your exercise routine today?", the server uses weather data to generate a response such as, "If the weather is bad, we recommend stretching indoors."

[1034] Prompt Sentence Examples

[1035] "Please tell me the process flow of a system that collects data on walking and talking in a user's living room. I also heard that the system can recognize the user's emotions. Can you give me a concrete example?"

[1036] "Describe the process by which the server generates an appropriate answer when a user types a question into the terminal. Include examples of questions users might ask about today's exercise."

[1037] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1038] Step 1: Collect data

[1039] 1. The device uses a built-in camera and microphone to collect video and audio data from the user's daily life, and the input is real-world data of the user's actions and conversations.

[1040] 2. The device temporarily stores the collected data in its internal memory. The output is raw video and audio data stored in its internal memory. Specifically, it records the user walking around the living room, watching TV, and talking.

[1041] Step 2: Sentiment Analysis

[1042] 1. The device sends the video and audio data collected in real time to the emotion engine. The input is the video and audio data collected in step 1.

[1043] 2. The emotion engine analyzes the user's facial expressions and tone of voice to recognize their current emotion (happiness, sadness, anger, stress, etc.). The output is data indicating the user's emotion.

[1044] Step 3: Encrypt the data

[1045] 1. The device encrypts the collected video and audio data and emotion analysis results. The inputs are raw video, audio, and emotion data.

[1046] 2. Encryption uses technologies such as AES (Advanced Encryption Standard), and the output is encrypted video, audio, and emotion data.

[1047] Step 4: Sending data

[1048] 1. The terminal sends encrypted data to the server via the Internet. This transmission occurs at midnight every day. The input is the encrypted data.

[1049] 2. The terminal confirms that the transmission was successful. The output is the data that was successfully sent to the server.

[1050] Step 5: Decrypting the received data

[1051] 1. The server decrypts the received encrypted data. The input is the encrypted video, audio, and emotion data.

[1052] 2. The data is restored using the AES decryption process. The output is the original video, audio, and emotion data.

[1053] Step 6: Analyze video and audio data

[1054] 1. The server runs a multimodal AI model to extract user movement, walking speed, and posture from the video data. The input is the decoded video data.

[1055] 2. Perform motion recognition using a deep learning model. The output is extracted user motion, walking speed, and posture data.

[1056] Step 7: Evaluate the audio data

[1057] 1. The server analyzes the audio data and evaluates the tone of voice, speech tempo, and frequency of speech. The input is the decoded audio data.

[1058] 2. Natural language processing (NLP) techniques are used to extract information from speech, and the output is analysis data on tone, tempo, and frequency of speech.

[1059] Step 8: Integrating Emotional Data

[1060] 1. The server integrates the received emotion analysis results with other analysis data. The inputs are emotion analysis data, motion data, and voice data.

[1061] 2. Graph weekly reports and sentiment changes. The output is a consolidated user overall evaluation data.

[1062] Step 9: Risk Assessment

[1063] 1. The server compares the analysis results with past data and evaluates the user's risk of decline in physical, psychological, and social function. The input is the integrated assessment data and past data.

[1064] 2. For example, identify slower walking speed or decreased frequency of conversation. The output is the user's risk assessment.

[1065] Step 10: Generate a care recommendation

[1066] 1. The server generates appropriate care methods and lifestyle improvement suggestions for the user based on the assessment results. The input is the risk assessment results.

[1067] 2. The specific suggestion is, for example, "Today's recommended exercise is 30 minutes of walking." The output is the generated care suggestion.

[1068] Step 11: Notification of Care Suggestion

[1069] 1. The server sends the generated care suggestion to the terminal. The input is the generated care suggestion.

[1070] 2. The device notifies the user of the received care suggestion. For example, a push notification might say, "Today's recommended exercise is 30 minutes of walking." The output is the notified care suggestion.

[1071] Step 12: Accepting user questions

[1072] 1. The user enters a question through the device interface, for example, "Are there any changes to today's exercise?"

[1073] 2. The terminal sends the entered question to the server. The input is the user's question. The output is the question sent to the server.

[1074] Step 13: Generate answers to questions

[1075] 1. The server uses a generative AI to generate an appropriate answer. The input is the submitted user question.

[1076] 2. Generate an answer, for example, "Today's recommended exercise is 30 minutes of walking, but if the weather is bad, we recommend stretching indoors." The output is the generated answer.

[1077] Step 14: Notification of Response

[1078] 1. The server sends the generated answer to the terminal. The input is the generated answer.

[1079] 2. The terminal displays the received answer to the user. The output is the notified answer.

[1080] Step 15: Continuous data collection and comparison

[1081] 1. The device periodically collects new video and audio data. The input is data from the user's daily activities.

[1082] 2. The server continuously compares the historical data with the new data and updates the user's risk assessment. The output is an updated risk assessment result.

[1083] (Application example 2)

[1084] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1085] Conventional early detection systems for frailty (risk of decline in physical, psychological, and social function) assess risk by collecting and analyzing video and audio data from users' daily lives, but do not take the user's emotional state into consideration. This makes it difficult to provide more precise care recommendations or lifestyle improvements, and there remains the issue of strengthening security, particularly when it comes to electronic payments.

[1086] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting video and audio data of daily life, means for transmitting the collected video and audio data to the server, means for analyzing the received video and audio data and assessing the risk of decline in the user's physical, psychological, and social functions, means for generating appropriate care methods and lifestyle improvement suggestions for the user based on the assessment results, means for notifying the user of the generated care suggestions, means including an emotion engine that recognizes the user's emotions in real time, and means for making security enhancement and customization suggestions for electronic payments based on the user's emotional state. This makes it possible to enhance security and make customization suggestions that take the user's emotional state into consideration.

[1087] "Video and audio data of everyday life" refers to video and audio data that includes records of a user's daily activities and conversations.

[1088] "Means of collection" refers to means of obtaining video and audio data using devices such as cameras and microphones.

[1089] The "means for transmitting to the server" is a means for sending collected data to the server via a communication network such as the Internet.

[1090] "Means for analyzing and assessing the risk of decline in the user's physical, psychological, and social functions" refers to a data analysis system and algorithms that use the received data to assess the user's health status from multiple angles.

[1091] The "means for generating appropriate care methods and lifestyle improvement suggestions" is a system that automatically suggests optimal care and improvement measures to users based on the evaluation results.

[1092] "Means of notification" refers to the means for notifying users of the generated care suggestions and improvement measures, and includes push notifications, emails, etc.

[1093] The "real-time emotion recognition engine" is an algorithm and system that analyzes a user's facial expressions and tone of voice to identify their emotional state at that moment.

[1094] The "means for enhancing security and providing customized suggestions during electronic payments" is a system that provides security measures and customized suggestions according to the user's psychological state based on the results of the emotion engine.

[1095] This invention is a system that collects video and audio data from users' daily lives, detects the risk of decline in physical, psychological, and social function (frailty) at an early stage, and provides appropriate care. In addition, by combining it with an emotion engine, it can enhance security and provide customized suggestions during electronic payments based on the user's emotional state.

[1096] Hardware and software used

[1097] The following hardware and software are used to implement this system.

[1098] Hardware:

[1099] Camera and microphone: Collects video and audio data from your daily life.

[1100] Smartphone or PC: Used as a device to collect data and send it to the server.

[1101] software:

[1102] OpenCV: Used for video data processing and face detection.

[1103] Keras: A neural network library for running emotion recognition models.

[1104] SpeechRecognition: A library for converting speech data into text.

[1105] Transformers (Hugging Face): A library for recognizing emotions from speech using the wav2vec2 model.

[1106] Data collection and analysis

[1107] 1. The device uses a camera and microphone to collect video and audio data from everyday life, for example recording you walking around a room or having a conversation with someone.

[1108] 2. The device encrypts the collected video and audio data and sends it to a server via the Internet.

[1109] 3. The server uses OpenCV to detect faces from video data and Keras models to analyze emotions in real time. It also converts audio data to text using SpeechRecognition and uses wav2vec2 models to recognize emotions in the audio.

[1110] Assessment of frailty risk and emotional state

[1111] 1. The server analyzes the received video and audio data to assess the risk of decline in the user's physical, psychological, and social functions. For example, it analyzes the user's walking speed and frequency of conversation and compares them with past data.

[1112] 2. Based on the analysis results, the server generates appropriate care methods and lifestyle improvement suggestions for the user. For example, a suggestion might be, "Today's recommended exercise is 30 minutes of walking. We also recommend yoga for relaxation."

[1113] 3. Based on the results of emotion recognition, the system can also suggest enhanced security measures and customized special offers during electronic payments, for example, requesting additional authentication measures if the user is stressed.

[1114] Notifications and Interactivity

[1115] 1. The generated care suggestions and customization suggestions are sent to the device and notified to the user, for example, by a push notification on a smartphone.

[1116] 2. The device accepts questions and requests for advice from users. The user's questions are sent to the server, which uses a generative AI model to generate appropriate answers and return them to the device.

[1117] Examples of concrete examples and prompts

[1118] For example, if a user has their camera and microphone enabled when making an electronic payment, the system will generate the following prompt to perform emotion recognition:

[1119] Example prompt sentence:

[1120] Analyze users' emotions in real time and assess their stress levels.

[1121] This makes it possible to enhance security and provide customized suggestions that take into account the user's emotional state.

[1122] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1123] Step 1:

[1124] The device uses a camera and microphone to collect video and audio data from the user's daily life, for example, recording the user walking around a room or talking to someone. The input of this step is the camera and microphone, and the output is the collected video and audio data.

[1125] Step 2:

[1126] The terminal encrypts the collected video and audio data and transmits it to a server via the Internet. The input of this step is the collected video and audio data, and the output is an encrypted data stream.

[1127] Step 3:

[1128] The server processes the received video data using the OpenCV library to detect faces. Once the face region is detected, it analyzes the emotion in real time using a Keras model. The input of this step is the encrypted video data, and the output is the user's face region and emotional state (e.g., "happy," "sad," "anger," etc.).

[1129] Step 4:

[1130] The server converts the received audio data into text using the SpeechRecognition library and recognizes emotions from the audio using the wav2vec2 model. The input of this step is the encrypted audio data, and the output is the emotional state extracted from the audio.

[1131] Step 5:

[1132] The server integrates the results of the video and audio data analysis to assess the user's risk of decline in physical, psychological, and social function. For example, it analyzes the user's walking speed and frequency of conversation and compares them with past data. The inputs to this step are the analyzed video and audio data and past data, and the output is the user's risk assessment result.

[1133] Step 6:

[1134] Based on the analysis results, the server generates appropriate care methods and lifestyle improvement suggestions. For example, a suggestion might be generated such as, "Today's recommended exercise is 30 minutes of walking. We also recommend yoga for relaxation." The input for this step is the risk assessment result, and the output is the generated care suggestions.

[1135] Step 7:

[1136] The server sends the generated care suggestion to the device. The device notifies the user of the received suggestion, for example, by displaying a push notification on a smartphone. The input of this step is the generated care suggestion, and the output is a notification to the user.

[1137] Step 8:

[1138] Based on the emotion recognition results, the server can then customize the electronic payment process by suggesting enhanced security measures or special offers. For example, if the user is in a stressed state, it can request additional authentication measures. The input to this step is the emotion recognition results, and the output is the enhanced security measures and customized offers.

[1139] Step 9:

[1140] The device accepts questions and requests for advice from the user and sends that information to the server. For example, the user might ask, "Are there any changes to today's exercise?" The input to this step is the user's question or request, and the output is the data sent to the server.

[1141] Step 10:

[1142] The server uses the generative AI model to generate an appropriate answer to the user's question and return it to the device. For example, a response might be, "Today's recommended exercise is 30 minutes of walking, but if the weather is bad, we recommend indoor stretching." The input for this step is the user's question or request, and the output is the generated answer.

[1143] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1144] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1145] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1146] [Fourth embodiment]

[1147] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1148] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1149] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1150] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1151] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1152] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1153] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1154] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1155] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1156] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1157] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1158] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1159] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1160] This invention relates to a system for early detection of risk of decline in physical, psychological, and social functions (frailty) in a user's daily life and providing appropriate care. An embodiment of this system will be described in detail below.

[1161] overview

[1162] The system uses a device to collect video and audio data from the user's daily life and transmit it to a server. The server analyzes the received data and assesses the user's risk of frailty. Based on the results, it generates appropriate care methods and lifestyle improvement suggestions and notifies the user. It also responds to questions and requests for advice from the user and periodically compares new data with past data.

[1163] Data collection and transmission

[1164] The device uses a camera and microphone to collect video and audio data of the user's daily life. For example, the device is installed in the user's living space and automatically records walking, eating, conversations, and other activities that occur there.

[1165] The collected data is encrypted to protect privacy and sent to a server via the Internet. This transmission is carried out periodically (e.g., every night at midnight), and the data for the day is aggregated on the server.

[1166] Data analysis and evaluation

[1167] The server uses a multimodal AI model to analyze the received video and audio data. Specifically, it analyzes the user's movements, walking speed, and posture from the video data, and the tone of voice, speech tempo, and frequency of conversation from the audio data.

[1168] Based on the analysis results, the user's risk of decline in physical, psychological, and social functions is assessed. For example, if a decrease in walking speed or frequency of conversation is observed, it is determined that the risk of frailty is increasing.

[1169] Care suggestion generation and notification

[1170] Based on the analysis results, the server generates appropriate care methods and lifestyle improvement suggestions for the user, including daily exercise, a balanced diet, and ways to socialize.

[1171] The generated care suggestions are then sent back to the device and notified to the user. For example, specific instructions such as "Today's recommended exercise: 30 minutes of walking" are displayed to the user through the device.

[1172] User interaction

[1173] The device also has the ability to accept questions and requests for advice from users. When a user inputs a question through the device's interface, the information is sent to the server. The server uses generative AI to generate an appropriate answer to the question and returns it to the device. For example, in response to the question, "Would you like to change your exercise plan for today?", the device might provide a response such as, "Today's recommended exercise is 30 minutes of walking, but if the weather is bad, we recommend stretching indoors."

[1174] Data collection and comparison

[1175] The device also has the ability to periodically collect new video and audio data and compare it with past data, allowing it to track changes in the user's health and update care recommendations as needed. The device collects data daily, and the server uses this data to continuously assess risk.

[1176] As described above, this system detects users' risk of frailty early and provides appropriate care, enabling elderly people and adults who need preventative care to effectively manage their health at home.

[1177] The processing flow will be explained below.

[1178] Step 1:

[1179] The device collects video and audio data from the user's daily life through a camera and microphone, for example recording the user walking around the living room or talking while watching TV.

[1180] Step 2:

[1181] The device encrypts the collected video and audio data and transmits it securely to a server, taking privacy into consideration. The data is transmitted on a scheduled basis, such as at night or during periods of low load.

[1182] Step 3:

[1183] The server analyzes the received video and audio data, where a multimodal AI model is activated to extract the user's walking speed, posture, and movement patterns from the video data, and evaluate the tone of voice, speech tempo, and frequency of conversation from the audio data.

[1184] Step 4:

[1185] The server evaluates the risk of decline in the user's physical, psychological, and social functions based on the analysis results, for example, by comparing them with past data to see if their walking speed or frequency of conversation has decreased.

[1186] Step 5:

[1187] Based on the assessment results, the server generates appropriate care methods and lifestyle improvement suggestions, such as recommendations for appropriate exercise, a balanced diet, and ways to increase social interaction.

[1188] Step 6:

[1189] The server then sends the generated care recommendations back to the device, which then notifies the user. For example, a push notification will be displayed saying, "Today's recommended exercise is 30 minutes of walking."

[1190] Step 7:

[1191] Users can ask questions or ask for additional advice through the device, for example, "What is a good exercise for a rainy day?"

[1192] Step 8:

[1193] The device sends the user's question to the server, which then uses generative AI to generate an answer to the question. The answer is then sent to the device, which displays it to the user. For example, the device might return an answer such as "We recommend stretching indoors."

[1194] Step 9:

[1195] The device continues to periodically collect new video and audio data. The server compares this new data with past data and continually updates the user's risk assessment. For example, it analyzes the data collected daily to track changes in health status.

[1196] In this way, the system continuously collects and analyzes data throughout the user's daily life and provides appropriate care, making it possible to detect the user's risk of frailty early and respond appropriately.

[1197] Example 1

[1198] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1199] Elderly people and adults who need preventative care lack the means to detect early risk of decline in physical, psychological, and social function (frailty) at home and receive appropriate care. Conventional systems often do not adequately collect and analyze data closely related to users' daily lives, resulting in delayed risk assessment. There is also a need for a system that can respond quickly and appropriately to user questions.

[1200] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1201] In this invention, the server includes a means for encrypting video and audio data of daily life and transmitting it to the server, a means for analyzing the received video and audio data using a multimodal AI model to assess the user's risk of decline in physical, psychological, and social function, and a means for generating appropriate care methods and lifestyle improvement suggestions for the user based on the assessment results. This allows users to easily collect data from their daily lives and have it analyzed by the server, enabling early detection of frailty risk and appropriate care. Furthermore, answer generation using the generative AI model enables prompt and appropriate responses to user questions.

[1202] "Video and audio data of everyday life" refers to video and audio data that records the actions and conversations of users in their daily lives.

[1203] "Encryption" is a technology that converts the contents of data into a format that cannot be understood by third parties, and is used particularly to protect privacy and improve security.

[1204] A "server" is a computer system that receives data from client terminals via a network and analyzes and processes the data.

[1205] A "multimodal AI model" is an artificial intelligence model that can comprehensively analyze multiple types of data, such as video data and audio data.

[1206] "Frailty risk" is an indicator of the risk of decline in physical, psychological, and social functioning, and is particularly important when assessing the health status of elderly people and adults who require prevention.

[1207] "Care Suggestions" are advice including specific exercises, balanced meals, and social interaction methods aimed at maintaining the user's health and restoring function.

[1208] A "generative AI model" is an artificial intelligence model that automatically generates appropriate answers or suggestions based on input data.

[1209] "Comparison" is the act of matching new data with past data to detect changes over time.

[1210] "Notification" refers to the act of conveying important information to the user, and in particular, care suggestions and responses from the system.

[1211] This invention relates to a system that detects early risk of decline in physical, psychological, and social function (frailty) in a user's daily life and provides appropriate care. This system can execute a series of processes, from collecting and analyzing video and audio data to generating and notifying care suggestions.

[1212] Data collection and transmission

[1213] The device uses a camera and microphone to collect video and audio data of the user's daily life. The device is equipped with a Python program that uses the OpenCV library to capture video and the PyAudio library to record audio. For example, the device can be installed in the user's living space and automatically record walking, eating, conversations, and other activities.

[1214] The collected data is encrypted using encryption technology such as AES. The encrypted data is periodically sent to a server via the Internet. Data transmission is managed by a regularly executed script such as a Cron job. Specifically, the collected data is encrypted every night at midnight and sent using the HTTPS protocol.

[1215] Data analysis

[1216] The server uses a multimodal AI model to analyze the received data. The multimodal AI model includes a model that comprehensively analyzes the features of video data and audio data (e.g., a model trained using PyTorch). Specifically, the server analyzes the user's walking speed and posture from the video data, and the tone of voice, speech tempo, and frequency of conversation from the audio data.

[1217] Risk assessment and care recommendation generation

[1218] The server evaluates the user's risk of frailty based on the analysis results. If a decrease in walking speed or frequency of conversation is observed, it is determined that the risk of frailty is increasing. Based on the evaluation results, the server generates appropriate care methods and lifestyle improvement suggestions. These suggestions include daily exercise, a balanced diet, and ways to socialize. The generated care suggestions are formatted into text.

[1219] Care Suggestion Notification

[1220] The generated care suggestions are sent back to the device and notified to the user. The suggestions are displayed as push notifications or alerts. For example, a message such as "Today's recommended exercise: 30 minutes of walking" is displayed on the user's device.

[1221] Responding to user questions

[1222] The device has the ability to accept questions and requests for advice from users. When a user inputs a question through the device's interface, the information is sent to the server. The server uses a generative AI model to generate an appropriate answer to the question and returns it to the device. For example, in response to the question, "Would you like to change your exercise plan for today?", the device might provide an answer such as, "Today's recommended exercise is 30 minutes of walking, but if the weather is bad, we recommend stretching indoors."

[1223] Data collection and comparison

[1224] The server periodically collects new data and compares it with past data. This allows the server to track changes in the user's health and update care recommendations as needed. For example, if walking speed data from the past three months is compared and a clear decline is detected, increased exercise is suggested. In a specific scenario, the device collects data daily, and the server continuously performs risk assessments based on this data.

[1225] Prompt Sentence Examples

[1226] "Analyze walking speed data from the past three months and generate appropriate care suggestions if significant decline is observed."

[1227] "If the frequency of conversations during meals decreases, please provide psychological support suggestions accordingly."

[1228] As described above, the purpose of this system is to detect early risks of decline in the user's physical, psychological, and social functions in their daily lives and to provide appropriate care.

[1229] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1230] Step 1: Data collection

[1231] The device collects video and audio data from the user's daily life. The input is video and audio data from the camera and microphone. The device uses a Python program to capture video using the OpenCV library and record audio using the PyAudio library. The output is the collected raw data. Specifically, the device activates the camera and microphone every morning at 6:00 and begins recording movements and conversations in the living room and kitchen.

[1232] Step 2: Send data

[1233] The terminal encrypts the collected video and audio data using AES and sends it to the server. The input is the raw data collected in step 1. Data transmission is managed by a regularly executed script such as a Cron job. The output is the encrypted data sent to the server. Specifically, the collected data is encrypted every night at midnight and transmitted using the HTTPS protocol.

[1234] Step 3: Data analysis

[1235] The server analyzes the received data using a multimodal AI model. The input is encrypted video and audio data. The server decrypts the data and extracts features from the video data using OpenCV and the audio data using Librosa. The output is the analysis results. Specifically, the server decompresses the data, inputs it into the AI ​​model for analysis, and extracts features such as walking speed, tone of voice, and frequency of conversation.

[1236] Step 4: Risk assessment

[1237] The server evaluates the user's risk of frailty based on the analysis results. The input is the analysis results obtained in step 3. The server calculates a risk score based on this data and determines the risk level. The output is the risk assessment result. Specifically, the server evaluates a decrease in walking speed and frequency of conversation, and calculates a risk level of "low," "medium," or "high."

[1238] Step 5: Generate a care proposal

[1239] The server generates appropriate care suggestions based on the risk assessment results. The input is a database of risk assessment results and corresponding care suggestions. The server uses the generative AI model to generate suggestions such as exercise, diet, and social interaction. The output is specific care suggestions. Specific operations include the server formatting "daily exercises" in text format based on the AI ​​model.

[1240] Step 6: User Notification

[1241] The server sends the generated care suggestions to the device and notifies the user. The input is the text data of the care suggestions. The server sends the suggestions to the device, and the device displays a notification to the user. The output is the state in which the notification is displayed to the user. Specifically, the server sends content such as "Today's recommended exercise: 30 minutes of walking" to the device, and the device displays it as a push notification.

[1242] Step 7: Respond to user questions

[1243] The device accepts questions and requests for advice from the user. The input is the question text from the user. The device sends this information to the server, which uses a generative AI model to generate an answer. The output is the generated answer. Specifically, the user enters the question, "Would you like to change your exercise routine today?" and the server provides the answer, "If the weather is bad, we recommend stretching indoors."

[1244] Step 8: Data comparison and update

[1245] The server periodically collects new data and compares it with past data. The input is new video and audio data and previously collected data. The server compares these and detects changes in health status. The output is updated care recommendations. Specifically, the server compares walking speed data from the past three months and generates exercise enhancements or new care recommendations as needed, and notifies the user.

[1246] (Application example 1)

[1247] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1248] This invention relates to a system that detects the risk of decline in physical, psychological, and social functions in daily life (frailty) at an early stage and provides appropriate care. The purpose of this invention is to solve the problem of the difficulty of monitoring the health status of customers in real time in a physical store and providing timely care and lifestyle improvement suggestions. It is also required to respond quickly to customers' questions and requests for advice.

[1249] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1250] In this invention, the server includes means for collecting video and audio data of daily life, means for transmitting the collected video and audio data to the server, means for analyzing the received video and audio data and assessing the risk of decline in the user's physical, psychological, and social functions, means for generating appropriate care methods and lifestyle improvement suggestions for the user based on the assessment results, means for notifying the user of the generated care suggestions, and means for monitoring the physical and psychological functions of customers in a physical store in real time using a smart device and assessing health risks. This makes it possible to assess the health status of customers in real time and provide appropriate care and lifestyle improvement suggestions as needed.

[1251] "Video and audio data from everyday life" refers to video footage and audio recordings taken in the user's living environment.

[1252] "Server" refers to a computer system that receives, analyzes, and evaluates collected video and audio data.

[1253] "Means for assessing the risk of decline in physical, psychological, and social function" refers to algorithms and programs that analyze data indicating the user's health status and assess the risk of frailty.

[1254] The "means for generating suggestions for improving care methods and lifestyle habits" refers to a means for suggesting to the user care and lifestyle habits improvement methods that should be performed based on the evaluation results.

[1255] "Means for notifying the user of the generated care suggestions" refers to a device or interface for notifying the user of the generated care suggestions in an easy-to-understand manner.

[1256] "Means of monitoring the physical and psychological functions of customers in real time using smart devices in physical stores" refers to smart devices such as cameras and microphones used to observe and record the health status of customers in physical stores, as well as systems for analyzing that data.

[1257] This invention is a system that monitors the health status of customers in a physical store in real time and provides appropriate care and lifestyle improvement suggestions. This invention is realized by collecting video and audio data of daily life using smart devices and sending it to a server.

[1258] System configuration

[1259] Hardware and Software Used

[1260] 1. Smart Devices:

[1261] Using cameras and microphones (e.g., smart glasses), video and audio recordings of customers' everyday activities are collected.

[1262] 2. Server:

[1263] Uses Python, OpenCV, PyAudio, and the Requests library to receive, store, and analyze data.

[1264] A generative AI model is used to analyze the collected data and generate appropriate care methods and improvement suggestions.

[1265] Data collection and transmission

[1266] 1. Data Collection:

[1267] The smart glasses are equipped with a camera and microphone that collects video and audio data of customers, specifically monitoring their walking speed, posture, tone of voice, and frequency of language use.

[1268] 2. Data transmission:

[1269] The collected data is encrypted to protect privacy and sent to a local server at the physical store or to a cloud server.

[1270] Data analysis and evaluation

[1271] The server uses the generative AI model to analyze the received video and audio data. Specifically, it performs the following data analysis:

[1272] 1. Video Data Analysis:

[1273] Analyzes user movements, walking speed, and posture.

[1274] 2. Audio data analysis:

[1275] Analyze tone of voice and frequency of language usage.

[1276] Based on the analysis results, the server evaluates the user's risk of declining physical and psychological function. If the evaluation finds a decrease in walking speed or frequency of conversation, it determines that the risk of frailty is increasing.

[1277] Care suggestion generation and notification

[1278] Based on the analysis results, the server generates appropriate care methods and lifestyle improvement suggestions for the customer, including:

[1279] 1. Exercise Suggestions:

[1280] Generate specific instructions such as "Today's recommended exercise: 30 minutes of walking."

[1281] 2. Dietary suggestions:

[1282] They offer advice such as, "Try to eat a balanced diet."

[1283] The generated care suggestions are communicated to the customer via the smart glasses display.

[1284] User interaction features

[1285] Customers can input questions or requests for advice through the smart glasses interface. This information is sent to the server, where the AI ​​generates an appropriate answer and sends it back to the smart glasses.

[1286] "What exercises would you recommend today?"

[1287] What measures should be taken if walking speed slows down?

[1288] "How would you assess your health status based on today's walking data?"

[1289] Answers to prompts such as:

[1290] Regularly collect and compare data

[1291] New data is collected periodically and compared with past data to track changes in the health status of customers. To achieve this, the smart glasses collect data daily, and the server continuously performs risk assessments based on this data.

[1292] The above is a specific embodiment of the present invention, which makes it possible to evaluate the health status of customers in real time within a physical store and provide appropriate care and lifestyle improvement suggestions.

[1293] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1294] Step 1:

[1295] The device collects video and audio data from the user's daily life.

[1296] Input: Video and audio data captured by the camera and microphone on the smart glasses.

[1297] Data processing and calculation: Use OpenCV to capture video and PyAudio to record audio.

[1298] Output: Collected video data (e.g., JPEG images) and audio data (e.g., WAV files).

[1299] Step 2:

[1300] The terminal transmits the collected video and audio data to the server.

[1301] Input: Video and audio data collected in step 1.

[1302] Data processing and calculation: The data is encrypted to protect privacy and sent to the server using an HTTP POST request.

[1303] Output: Encrypted video and audio data is sent to the server.

[1304] Step 3:

[1305] The server analyzes the received video and audio data and assesses the user's risk of decline in physical and psychological function.

[1306] Input: Encrypted video and audio data sent from the device.

[1307] Data processing and calculation: Generative AI models are used to analyze factors such as walking speed, posture, tone of voice, and frequency of language use.

[1308] Output: Results of the user's frailty risk assessment.

[1309] Step 4:

[1310] Based on the evaluation results, the server generates appropriate care methods and lifestyle improvement suggestions for the user.

[1311] Input: Frailty risk assessment results obtained in Step 3.

[1312] Data processing and calculation: Based on the assessment results, care recommendations such as specific exercises, dietary improvements, and social interaction methods are generated.

[1313] Output: Generated care suggestions (e.g., "Today's recommended exercise: 30 minutes walking").

[1314] Step 5:

[1315] The terminal notifies the user of the generated care suggestions.

[1316] Input: Care suggestions generated in step 4.

[1317] Data processing and calculation: Care suggestions are displayed on the smart glasses display.

[1318] Output: Notification to the user (e.g. "Recommended exercise for today: 30 minutes walking" appears on the display).

[1319] Step 6:

[1320] The user inputs questions or requests for advice through the device interface.

[1321] Input: Text data of a user question or request (e.g., "What exercise would you recommend for me today?").

[1322] Data processing and calculation: The terminal sends this information to the server.

[1323] Output: The user's question is sent to the server.

[1324] Step 7:

[1325] The server uses a generative AI model to generate appropriate answers to the questions and returns them to the device.

[1326] Input: The user's question submitted in step 6.

[1327] Data processing and computation: Generative AI models are used to generate answers to questions.

[1328] Output: The generated answer (e.g., "Today's recommended exercise is 30 minutes of walking, but if the weather is bad, indoor stretching is recommended").

[1329] Step 8:

[1330] The terminal provides the answer from the server to the user.

[1331] Input: The answer from the server generated in step 7.

[1332] Data processing and calculation: The answer is displayed on the smart glasses display.

[1333] Output: Notify the user of the answer (e.g., "Today's recommended exercise is 30 minutes of walking, but if the weather is bad, we recommend stretching indoors.")

[1334] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1335] This invention relates to a system that detects early risk of decline in physical, psychological, and social functions (frailty) in a user's daily life and provides appropriate care, and further combines it with an emotion engine that recognizes the user's emotions. An embodiment of this system will be described in detail below.

[1336] overview

[1337] The system uses a device to collect video and audio data from the user's daily life and send it to a server. The server not only analyzes the received data and assesses the user's risk of frailty, but also recognizes the user's emotions using an emotion engine. Based on the results, it generates appropriate care methods and lifestyle improvement suggestions and notifies the user. It also responds to questions and requests for advice from the user and regularly compares new data with past data.

[1338] Data collection and transmission

[1339] The device uses a camera and microphone to collect video and audio data from the user's daily life, such as when they walk around the living room or talk while watching TV. The emotion engine then analyzes the user's facial expressions and tone of voice in real time to recognize their emotions.

[1340] The collected data is encrypted to protect privacy and sent to a server via the Internet. This transmission is carried out periodically (e.g., every night at midnight), and the data for the day is aggregated on the server.

[1341] Data analysis and evaluation

[1342] The server analyzes the received video and audio data. A multimodal AI model is activated to extract the user's movements, walking speed, and posture from the video data, and evaluate the tone of voice, speech tempo, and frequency of conversation from the audio data. In addition, an emotion engine analyzes the user's emotions from changes in facial expressions and tone of voice. This also allows the collection of user emotional data.

[1343] Based on the analysis results, the system assesses the user's risk of decline in physical, psychological, and social functioning. For example, it compares data from the past to see if their walking speed or frequency of conversation has decreased. Emotional data is also taken into consideration, and if stress or depression is observed, it determines that the risk is high.

[1344] Care suggestion generation and notification

[1345] Based on the analysis results, the server generates appropriate care methods and lifestyle improvement suggestions for the user, including suggestions for appropriate exercise, a balanced diet, and ways to increase social interactions. Emotional data is also taken into account, and suggestions are made based on the user's emotional state.

[1346] The generated care suggestions are then sent back to the device, which then notifies the user. For example, a push notification will be displayed saying, "Today's recommended exercise is 30 minutes of walking. We also recommend yoga for relaxation."

[1347] User interaction

[1348] The device has the ability to accept questions and requests for advice from users. When a user inputs a question through the device's interface, the information is sent to the server. The server uses generative AI to generate an appropriate answer to the question and returns it to the device. For example, in response to the question, "Would you like to change your exercise plan for today?", the device may provide an answer such as, "Today's recommended exercise is 30 minutes of walking, but if the weather is bad, we recommend stretching indoors."

[1349] Data collection and comparison

[1350] The device continues to periodically collect new video and audio data. The server compares this new data with past data and continually updates the user's risk assessment. For example, it analyzes the data collected daily to track changes in health status or emotions.

[1351] In this way, this system continuously collects and analyzes data from the user's daily life and provides appropriate care, making it possible to detect the user's risk of frailty early and respond appropriately.Furthermore, by combining it with an emotion engine, it is possible to realize individual responses that take into account the user's emotional state.

[1352] The processing flow will be explained below.

[1353] Step 1:

[1354] The device collects video and audio data from the user's daily life through a camera and microphone, for example, recording the user walking around the living room or talking while watching TV. The emotion engine analyzes the user's facial expressions and tone of voice in real time to recognize their emotions.

[1355] Step 2:

[1356] The device encrypts the collected video and audio data and transmits it securely to a server, taking privacy into consideration. The data is transmitted on a scheduled basis, such as at night or during periods of low load.

[1357] Step 3:

[1358] The server analyzes the received video and audio data. Here, a multimodal AI model is activated, extracting the user's walking speed, posture, and movement patterns from the video data, and evaluating the tone of voice, speech tempo, and frequency of conversation from the audio data. Furthermore, an emotion engine analyzes the user's emotions from changes in facial expressions and tone of voice.

[1359] Step 4:

[1360] The server evaluates the user's risk of decline in physical, psychological, and social function based on the analysis results. For example, it compares the results with past data to see if their walking speed has slowed or if they are talking less frequently. Emotional data is also taken into account, and if stress or depression is observed, it determines that the risk is high.

[1361] Step 5:

[1362] The server generates appropriate care methods and lifestyle improvement suggestions based on the evaluation results. These include suggestions for appropriate exercise, a balanced diet, and ways to increase social interaction. It also takes into account emotional data, suggesting relaxation techniques if the user is feeling stressed, for example.

[1363] Step 6:

[1364] The server then sends the generated care recommendations back to the device, which then notifies the user. For example, a push notification will be displayed saying, "Today's recommended exercise is 30 minutes of walking. We also recommend yoga for relaxation."

[1365] Step 7:

[1366] Users can ask questions or ask for additional advice through the device, for example, "What is a good exercise for a rainy day?"

[1367] Step 8:

[1368] The device sends a question from the user to the server, which then uses generative AI to generate an answer to the question. The answer is then sent to the device, which then displays it to the user. For example, the device might provide an answer such as "We recommend stretching indoors."

[1369] Step 9:

[1370] The device will continue to periodically collect new video and audio data. The server will compare this new data with past data and continually update the user's risk assessment. For example, it will analyze the data collected daily to track changes in health status or emotions.

[1371] In this way, this system continuously collects and analyzes data from the user's daily life and provides appropriate care, making it possible to detect the user's risk of frailty early and respond appropriately.Furthermore, by combining it with an emotion engine, it is possible to realize individual responses that take into account the user's emotional state.

[1372] Example 2

[1373] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1374] In modern society, the aging population is increasing the risk of decline in physical, psychological, and social functions (frailty). Early detection of this risk and the provision of appropriate care and lifestyle improvement suggestions are required, but currently there is a lack of systems that can effectively do this. Individualized care that takes into account the user's emotional state is also required, but systems that can address this are limited. To solve this problem, a system that can continuously collect and analyze data from the user's daily life and provide appropriate care is needed.

[1375] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1376] In this invention, the server includes a means for collecting video and audio data of daily activities, a means for comparing the analysis results with past data and assessing the user's risk of decline in physical, psychological, and social functions, and a means for integrating emotional data into the assessment results. This enables early detection of the user's risk of frailty and provides appropriate care and lifestyle improvement suggestions in real time. Furthermore, by taking the user's emotional state into consideration, personalized suggestions can be made, improving the user's quality of life.

[1377] "Daily activities" refers to the various actions and tasks that users perform on a daily basis, including, in particular, movement, conversation, and activities within the home.

[1378] "Video and audio data" means video and audio recordings of your movements and speech, captured through cameras, microphones, and other sensors.

[1379] "Collection means" refers to a combination of devices and software for acquiring data, including, for example, a camera, microphone, data collection application, etc.

[1380] A "computer" is a system consisting of hardware and software for processing and analyzing data, and primarily refers to a server, but also includes cloud computing resources.

[1381] "Encryption" refers to the process of transforming data using a specific algorithm to protect its security. Encrypted data can only be read with a specific key.

[1382] "Decryption" refers to the process of restoring encrypted data to its original state, also known as decoding.

[1383] "Means for analysis" refers to devices and software that use specific algorithms and models to extract meaning from and evaluate acquired data.

[1384] "Extracting movement, posture, and walking speed" refers to the process of recognizing and quantifying the user's physical activity patterns, posture, and walking speed from video data.

[1385] "Evaluating tone, tempo, and frequency of speech" refers to the process of extracting characteristics of a user's speaking style from audio data and analyzing those changes.

[1386] "Emotional data" refers to information about a user's emotional state that can be inferred from their facial expressions and tone of voice, and can include stress, joy, sadness, etc.

[1387] "Means for integrating into evaluation results" refers to devices and software that integrate analyzed emotional data with physical and psychological evaluation results to comprehensively evaluate the user's condition.

[1388] "Generating suggestions for improving care methods and lifestyle habits" refers to the process of recommending specific actions and habits for maintaining health and improving quality of life based on the user's evaluation results.

[1389] "Means of notification" refers to the means by which the system communicates care suggestions generated by the system to the user, including, for example, push notifications, emails, in-app messages, etc.

[1390] MODE FOR CARRYING OUT THE INVENTION

[1391] This invention relates to a system that detects early risk of decline in physical, psychological, and social functions (frailty) in the user's daily life and provides appropriate care. Furthermore, this system incorporates an emotion engine that recognizes the user's emotions.

[1392] overview

[1393] The system uses a device that collects video and audio data from the user's daily life and sends it to a server. The server not only analyzes the received data and assesses the user's risk of frailty, but also recognizes the user's emotions using an emotion engine. Based on the assessment results, the system generates appropriate care methods and lifestyle improvement suggestions and notifies the user. The system also responds to questions and requests for advice from the user and periodically compares new data with past data.

[1394] Data collection and transmission

[1395] The device uses a camera and microphone to collect video and audio data from the user's daily life. For example, it records the user walking around the living room or talking while watching TV. In addition, an emotion engine analyzes the user's facial expressions and tone of voice in real time to recognize emotions. The collected data is encrypted to protect privacy and sent to a server via the Internet. This transmission occurs periodically (for example, every night at midnight), and the data for the day is aggregated on the server.

[1396] Data analysis and evaluation

[1397] The server analyzes the received video and audio data. A multimodal AI model is activated to extract the user's movements, walking speed, and posture from the video data, and evaluate the tone of voice, speech tempo, and frequency of conversation from the audio data. Furthermore, an emotion engine analyzes emotions from the user's facial expressions and tone of voice, thereby collecting emotional data. Based on the analysis results, the server evaluates the risk of decline in the user's physical, psychological, and social functions. For example, by comparing with past data, it checks whether walking speed has slowed or frequency of conversation has decreased. Emotional data is also taken into consideration, and if stress or depression is observed, it determines that there is a high risk.

[1398] Care suggestion generation and notification

[1399] Based on the analysis results, the server generates appropriate care methods and lifestyle improvement suggestions for the user. These include suggestions for appropriate exercise, a balanced diet, and ways to increase social interaction. Emotional data is also taken into consideration, and suggestions are made based on the user's emotional state. The generated care suggestions are sent back to the device, which then notifies the user. For example, a push notification may be displayed saying, "Today's recommended exercise is 30 minutes of walking. We also recommend yoga for relaxation."

[1400] User interaction

[1401] The device has the ability to accept questions and requests for advice from users. When a user inputs a question through the device's interface, the information is sent to the server. The server uses generative AI to generate an appropriate answer to the question and returns it to the device. For example, in response to the question, "Would you like to change your exercise plan for today?", the device may provide an answer such as, "Today's recommended exercise is 30 minutes of walking, but if the weather is bad, we recommend stretching indoors."

[1402] Data collection and comparison

[1403] The device periodically collects new video and audio data. The server compares this new data with past data and continuously updates the user's risk assessment. For example, it analyzes data collected daily to track changes in health and emotional state. Through these steps, the system continuously collects and analyzes data throughout the user's daily life, enabling it to provide appropriate care. It also features a dialogue function using a generative AI model, providing information tailored to the user's needs.

[1404] Specific examples

[1405] Specific examples are shown below.

[1406] Example 1: Video footage of a user walking around the living room and audio data of the user talking are collected. The server analyzes these and compares them with past data to determine that the user's walking speed and frequency of conversation have decreased, indicating an increased risk of frailty.

[1407] Example 2: A user uses a smartphone to input a question. In response to the question, "Would you like to change your exercise routine today?", the server uses weather data to generate a response such as, "If the weather is bad, we recommend stretching indoors."

[1408] Prompt Sentence Examples

[1409] "Please tell me the process flow of the system that collects data on walking and talking in the user's living room. I also heard that the system can recognize the user's emotions. Can you give me a concrete example?"

[1410] "Describe the process by which the server generates an appropriate answer when a user types a question into the terminal. Include examples of prompts users might ask about today's exercise."

[1411] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1412] Step 1: Collect data

[1413] 1. The device uses a built-in camera and microphone to collect video and audio data from the user's daily life, and the input is real-world data of the user's actions and conversations.

[1414] 2. The device temporarily stores the collected data in its internal memory. The output is raw video and audio data stored in its internal memory. Specifically, it records the user walking around the living room, watching TV, and talking.

[1415] Step 2: Sentiment Analysis

[1416] 1. The device sends the video and audio data collected in real time to the emotion engine. The input is the video and audio data collected in step 1.

[1417] 2. The emotion engine analyzes the user's facial expressions and tone of voice to recognize their current emotion (happiness, sadness, anger, stress, etc.). The output is data that indicates the user's emotion.

[1418] Step 3: Encrypt the data

[1419] 1. The device encrypts the collected video and audio data and emotion analysis results. The inputs are raw video, audio, and emotion data.

[1420] 2. Encryption uses technologies such as AES (Advanced Encryption Standard), and the output is encrypted video, audio, and emotion data.

[1421] Step 4: Sending data

[1422] 1. The terminal sends encrypted data to the server via the Internet. This transmission occurs at midnight every day. The input is the encrypted data.

[1423] 2. The terminal confirms that the transmission was successful. The output is the data that was successfully sent to the server.

[1424] Step 5: Decrypting the received data

[1425] 1. The server decrypts the received encrypted data. The input is the encrypted video, audio, and emotion data.

[1426] 2. The data is restored using the AES decryption process. The output is the original video, audio, and emotion data.

[1427] Step 6: Analyze video and audio data

[1428] 1. The server runs a multimodal AI model to extract user movement, walking speed, and posture from the video data. The input is the decoded video data.

[1429] 2. Perform motion recognition using a deep learning model. The output is extracted user motion, walking speed, and posture data.

[1430] Step 7: Evaluate the audio data

[1431] 1. The server analyzes the audio data and evaluates the tone of voice, speech tempo, and frequency of speech. The input is the decoded audio data.

[1432] 2. Natural language processing (NLP) techniques are used to extract information from speech, and the output is analysis data on tone, tempo, and frequency of speech.

[1433] Step 8: Integrating Emotional Data

[1434] 1. The server integrates the received emotion analysis results with other analysis data. The inputs are emotion analysis data, motion data, and voice data.

[1435] 2. Graph weekly reports and sentiment changes. The output is a consolidated user overall evaluation data.

[1436] Step 9: Risk Assessment

[1437] 1. The server compares the analysis results with past data and evaluates the risk of decline in the user's physical, psychological, and social functions. The input is the integrated assessment data and past data.

[1438] 2. For example, identify slower walking speed or decreased frequency of conversation. The output is the user's risk assessment.

[1439] Step 10: Generate a care recommendation

[1440] 1. The server generates appropriate care methods and lifestyle improvement suggestions for the user based on the assessment results. The input is the risk assessment results.

[1441] 2. The specific suggestion is, for example, "Today's recommended exercise is 30 minutes of walking." The output is the generated care suggestion.

[1442] Step 11: Notification of Care Suggestions

[1443] 1. The server sends the generated care suggestion to the terminal. The input is the generated care suggestion.

[1444] 2. The device notifies the user of the received care suggestion. For example, a push notification might say, "Today's recommended exercise is 30 minutes of walking." The output is the notified care suggestion.

[1445] Step 12: Accepting user questions

[1446] 1. The user enters a question through the device interface, for example, "Are there any changes to today's exercise?"

[1447] 2. The terminal sends the entered question to the server. The input is the user's question. The output is the question sent to the server.

[1448] Step 13: Generate answers to questions

[1449] 1. The server uses a generative AI to generate an appropriate answer. The input is the submitted user question.

[1450] 2. Generate an answer, for example, "Today's recommended exercise is 30 minutes of walking, but if the weather is bad, we recommend stretching indoors." The output is the generated answer.

[1451] Step 14: Notification of Response

[1452] 1. The server sends the generated answer to the terminal. The input is the generated answer.

[1453] 2. The terminal displays the received answer to the user. The output is the notified answer.

[1454] Step 15: Continuous data collection and comparison

[1455] 1. The device periodically collects new video and audio data. The input is data from the user's daily activities.

[1456] 2. The server continuously compares the historical data with the new data and updates the user's risk assessment. The output is an updated risk assessment result.

[1457] (Application example 2)

[1458] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1459] Conventional early detection systems for frailty (risk of decline in physical, psychological, and social function) assess risk by collecting and analyzing video and audio data from users' daily lives, but do not take the user's emotional state into consideration. This makes it difficult to provide more precise care recommendations or lifestyle improvements, and there remains the issue of strengthening security, particularly when it comes to electronic payments.

[1460] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting video and audio data of daily life, means for transmitting the collected video and audio data to the server, means for analyzing the received video and audio data and assessing the risk of decline in the user's physical, psychological, and social functions, means for generating appropriate care methods and lifestyle improvement suggestions for the user based on the assessment results, means for notifying the user of the generated care suggestions, means including an emotion engine that recognizes the user's emotions in real time, and means for making security enhancement and customization suggestions for electronic payments based on the user's emotional state. This makes it possible to enhance security and make customization suggestions that take the user's emotional state into consideration.

[1461] "Video and audio data of everyday life" refers to video and audio data that includes records of a user's daily activities and conversations.

[1462] "Means of collection" refers to means of obtaining video and audio data using devices such as cameras and microphones.

[1463] The "means for transmitting to the server" is a means for sending collected data to the server via a communication network such as the Internet.

[1464] "Means for analyzing and assessing the risk of decline in the user's physical, psychological, and social functions" refers to a data analysis system and algorithms that use the received data to assess the user's health status from multiple angles.

[1465] The "means for generating appropriate care methods and lifestyle improvement suggestions" is a system that automatically suggests optimal care and improvement measures to users based on the evaluation results.

[1466] "Means of notification" refers to the means for notifying users of the generated care suggestions and improvement measures, and includes push notifications, emails, etc.

[1467] The "real-time emotion recognition engine" is an algorithm and system that analyzes a user's facial expressions and tone of voice to identify their emotional state at that moment.

[1468] The "means for enhancing security and providing customized suggestions during electronic payments" is a system that provides security measures and customized suggestions according to the user's psychological state based on the results of the emotion engine.

[1469] This invention is a system that collects video and audio data from users' daily lives, detects the risk of decline in physical, psychological, and social function (frailty) at an early stage, and provides appropriate care. In addition, by combining it with an emotion engine, it can enhance security and provide customized suggestions during electronic payments based on the user's emotional state.

[1470] Hardware and software used

[1471] The following hardware and software are used to implement this system.

[1472] Hardware:

[1473] Camera and microphone: Collects video and audio data from your daily life.

[1474] Smartphone or PC: Used as a device to collect data and send it to the server.

[1475] software:

[1476] OpenCV: Used for video data processing and face detection.

[1477] Keras: A neural network library for running emotion recognition models.

[1478] SpeechRecognition: A library for converting speech data into text.

[1479] Transformers (Hugging Face): A library for recognizing emotions from speech using the wav2vec2 model.

[1480] Data collection and analysis

[1481] 1. The device uses a camera and microphone to collect video and audio data from everyday life, for example recording you walking around a room or having a conversation with someone.

[1482] 2. The device encrypts the collected video and audio data and sends it to a server via the Internet.

[1483] 3. The server uses OpenCV to detect faces from video data and Keras models to analyze emotions in real time. It also converts audio data to text using SpeechRecognition and uses wav2vec2 models to recognize emotions in the audio.

[1484] Assessment of frailty risk and emotional state

[1485] 1. The server analyzes the received video and audio data to assess the risk of decline in the user's physical, psychological, and social functions. For example, it analyzes the user's walking speed and frequency of conversation and compares them with past data.

[1486] 2. Based on the analysis results, the server generates appropriate care methods and lifestyle improvement suggestions for the user. For example, a suggestion might be, "Today's recommended exercise is 30 minutes of walking. We also recommend yoga for relaxation."

[1487] 3. Based on the results of emotion recognition, the system can also suggest enhanced security measures and customized special offers during electronic payments, for example, requesting additional authentication measures if the user is stressed.

[1488] Notifications and Interactivity

[1489] 1. The generated care suggestions and customization suggestions are sent to the device and notified to the user, for example, by a push notification on a smartphone.

[1490] 2. The device accepts questions and requests for advice from users. The user's questions are sent to the server, which uses a generative AI model to generate appropriate answers and return them to the device.

[1491] Examples of concrete examples and prompts

[1492] For example, if a user has their camera and microphone enabled when making an electronic payment, the system will generate the following prompt to perform emotion recognition:

[1493] Example prompt sentence:

[1494] Analyze users' emotions in real time and assess their stress levels.

[1495] This makes it possible to enhance security and provide customized suggestions that take into account the user's emotional state.

[1496] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1497] Step 1:

[1498] The device uses a camera and microphone to collect video and audio data from the user's daily life, for example, recording the user walking around a room or talking to someone. The input of this step is the camera and microphone, and the output is the collected video and audio data.

[1499] Step 2:

[1500] The terminal encrypts the collected video and audio data and transmits it to a server via the Internet. The input of this step is the collected video and audio data, and the output is an encrypted data stream.

[1501] Step 3:

[1502] The server processes the received video data using the OpenCV library to detect faces. Once the face region is detected, it analyzes the emotion in real time using a Keras model. The input of this step is the encrypted video data, and the output is the user's face region and emotional state (e.g., "happy," "sad," "anger," etc.).

[1503] Step 4:

[1504] The server converts the received audio data into text using the SpeechRecognition library and recognizes emotions from the audio using the wav2vec2 model. The input of this step is the encrypted audio data, and the output is the emotional state extracted from the audio.

[1505] Step 5:

[1506] The server integrates the results of the video and audio data analysis to assess the user's risk of decline in physical, psychological, and social function. For example, it analyzes the user's walking speed and frequency of conversation and compares them with past data. The inputs to this step are the analyzed video and audio data and past data, and the output is the user's risk assessment result.

[1507] Step 6:

[1508] Based on the analysis results, the server generates appropriate care methods and lifestyle improvement suggestions. For example, a suggestion might be generated such as, "Today's recommended exercise is 30 minutes of walking. We also recommend yoga for relaxation." The input for this step is the risk assessment result, and the output is the generated care suggestions.

[1509] Step 7:

[1510] The server sends the generated care suggestion to the device. The device notifies the user of the received suggestion, for example, by displaying a push notification on a smartphone. The input of this step is the generated care suggestion, and the output is a notification to the user.

[1511] Step 8:

[1512] Based on the emotion recognition results, the server can then customize the electronic payment process by suggesting enhanced security measures or special offers. For example, if the user is in a stressed state, it can request additional authentication measures. The input to this step is the emotion recognition results, and the output is the enhanced security measures and customized offers.

[1513] Step 9:

[1514] The device accepts questions and requests for advice from the user and sends that information to the server. For example, the user might ask, "Are there any changes to today's exercise?" The input to this step is the user's question or request, and the output is the data sent to the server.

[1515] Step 10:

[1516] The server uses the generative AI model to generate an appropriate answer to the user's question and return it to the device. For example, a response might be, "Today's recommended exercise is 30 minutes of walking, but if the weather is bad, we recommend indoor stretching." The input for this step is the user's question or request, and the output is the generated answer.

[1517] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1518] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1519] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1520] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1521] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1522] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1523] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1524] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1525] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1526] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1527] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1528] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1529] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1530] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1531] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1532] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1533] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1534] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1535] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1536] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1537] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1538] The following is further disclosed regarding the above embodiment.

[1539] (Claim 1)

[1540] A means for collecting video and audio data of daily life;

[1541] means for transmitting the collected video and audio data to a server;

[1542] A means for analyzing the received video and audio data and assessing the risk of decline in the user's physical, psychological, and social functions;

[1543] A means for generating appropriate care methods and lifestyle improvement suggestions for users based on the evaluation results;

[1544] means for notifying the user of the generated care suggestions;

[1545] A system including:

[1546] (Claim 2)

[1547] A means for accepting questions and requests for advice from users and transmitting that information to a server;

[1548] 10. The system of claim 1, further comprising: means for providing the server-generated answer to the user.

[1549] (Claim 3)

[1550] 10. The system of claim 1, further comprising means for periodically collecting new video and audio data and comparing it with past data.

[1551] "Example 1"

[1552] (Claim 1)

[1553] A means for collecting video and audio data of daily life;

[1554] A means for encrypting the collected video and audio data and transmitting it to a server;

[1555] A means of analyzing the received video and audio data using a multimodal AI model to assess the risk of decline in the user's physical, psychological, and social functions;

[1556] A means for generating appropriate care methods and lifestyle improvement suggestions for users based on the evaluation results;

[1557] means for notifying the user of the generated care suggestions;

[1558] A system including:

[1559] (Claim 2)

[1560] A means for accepting questions and requests for advice from users and transmitting that information to a server;

[1561] 10. The system of claim 1, further comprising means for the server to provide the user with an answer generated using the generative AI model.

[1562] (Claim 3)

[1563] 10. The system of claim 1, further comprising means for periodically collecting new video and audio data and comparing it with past data.

[1564] "Application Example 1"

[1565] (Claim 1)

[1566] A means for collecting video and audio data of daily life;

[1567] means for transmitting the collected video and audio data to a server;

[1568] A means for analyzing the received video and audio data and assessing the risk of decline in the user's physical, psychological, and social functions;

[1569] A means for generating appropriate care methods and lifestyle improvement suggestions for users based on the evaluation results;

[1570] means for notifying the user of the generated care suggestions;

[1571] A method to monitor the physical and psychological functions of customers in real time using smart devices in physical stores and assess their health risks.

[1572] A system including:

[1573] (Claim 2)

[1574] A means for accepting questions and requests for advice from users and transmitting that information to a server;

[1575] 10. The system of claim 1, further comprising: means for providing the server-generated answer to the user.

[1576] (Claim 3)

[1577] 10. The system of claim 1, further comprising means for periodically collecting new video and audio data and comparing it with past data.

[1578] "Example 2: Combining Emotion Engines"

[1579] (Claim 1)

[1580] means for collecting video and audio data of daily activities;

[1581] means for transmitting the collected video and audio data to a computer;

[1582] A means for decrypting the encrypted data, extracting the user's movements, posture, and walking speed from the video data, and evaluating the tone, tempo, and frequency of conversation from the audio data;

[1583] A means of collecting emotional data by analyzing the user's facial expressions and tone of voice,

[1584] a means for integrating the collected emotion data into the evaluation results;

[1585] A means of comparing the analysis results with past data to assess the risk of decline in the user's physical, psychological, and social functions;

[1586] A means for generating appropriate care methods and lifestyle improvement suggestions for users based on the evaluation results;

[1587] means for notifying the user of the generated care suggestions;

[1588] A system including:

[1589] (Claim 2)

[1590] means for receiving questions and requests for advice from users and transmitting that information to the computer;

[1591] 10. The system of claim 1, further comprising means for providing the computer-generated answer to the user.

[1592] (Claim 3)

[1593] 10. The system of claim 1, further comprising means for periodically collecting new video and audio data and comparing it with past data.

[1594] "Application example 2 when combining emotion engines"

[1595] (Claim 1)

[1596] A means for collecting video and audio data of daily life;

[1597] means for transmitting the collected video and audio data to a server;

[1598] A means for analyzing the received video and audio data and assessing the risk of decline in the user's physical, psychological, and social functions;

[1599] A means for generating appropriate care methods and lifestyle improvement suggestions for users based on the evaluation results;

[1600] means for notifying the user of the generated care suggestions;

[1601] means including an emotion engine for recognizing user emotions in real time;

[1602] A means for enhancing security and providing customized suggestions during electronic payments based on the user's emotional state;

[1603] A system including:

[1604] (Claim 2)

[1605] A means for accepting questions and requests for advice from users and transmitting that information to a server;

[1606] 10. The system of claim 1, further comprising means for the server to provide the user with an answer generated using the generative AI model.

[1607] (Claim 3)

[1608] 10. The system of claim 1, further comprising means for periodically collecting new video and audio data and comparing it with past data. [Explanation of symbols]

[1609] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for collecting video and audio data of daily life; means for transmitting the collected video and audio data to a server; A means for analyzing the received video and audio data to assess the risk of decline in the user's physical, psychological, and social functions; A means for generating appropriate care methods and lifestyle improvement suggestions for users based on the evaluation results; means for notifying the user of the generated care suggestions; A system including:

2. A means for accepting questions and requests for advice from users and transmitting that information to a server; 10. The system of claim 1, further comprising means for providing the server-generated answer to the user.

3. 10. The system of claim 1, further comprising means for periodically collecting new video and audio data and comparing it with past data.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A