system

A system for monitoring elderly individuals' health and emotional states through data collection and analysis addresses the challenges of manpower shortages and privacy concerns, enabling early detection and automated responses to abnormalities.

JP2026070907APending Publication Date: 2026-04-28SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-16
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Conventional care methods for elderly individuals living alone face challenges such as a shortage of manpower and ensuring privacy, making it difficult to monitor their living conditions and detect health issues like dementia effectively.

Method used

A system that collects images and audio data, analyzes lifestyle patterns, and uses a generative model to determine dementia risk, providing audio warnings and notifying external parties, all without human intervention.

Benefits of technology

Enables comprehensive monitoring of elderly individuals' health and emotional states, facilitating early detection of abnormalities and reducing the burden of care by automating responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026070907000001_ABST
    Figure 2026070907000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A data collection means for collecting images and audio, An analysis means for creating a lifestyle pattern model by analyzing collected image and audio data and detecting anomalies, A notification system that alerts the user via voice when an anomaly is detected, A means of communication to transmit details of the anomaly to external parties, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] With the increase in the number of elderly people living alone at a distance, it is important to appropriately monitor the living conditions of the elderly and detect the progress of their health conditions and dementia at an early stage. However, conventional care methods have problems such as a shortage of manpower and ensuring the privacy of the elderly, and there is a need for a method to effectively solve these problems.

Means for Solving the Problems

[0005] This invention provides a system for detecting abnormalities by collecting images and audio, analyzing this data, and modeling the lifestyle patterns of elderly individuals. Furthermore, it supports the elderly's peace of mind by providing an audio warning to the user upon detection of an abnormality and notifying external parties of the details of the abnormality. The system also includes means for determining the risk of dementia from audio data using a generative model and monitoring health status using biometric information acquired by a wearable device. This enables a comprehensive understanding of the elderly's condition without relying on human intervention, facilitating early detection and reducing the burden of care.

[0006] "Data acquisition means" refers to devices and equipment used to acquire images and audio.

[0007] "Analysis methods" refer to processes and technologies used to model lifestyle patterns based on collected data and detect anomalies.

[0008] "Notification means" refers to technologies and devices used to communicate detected anomalies to the user via voice.

[0009] "Communication means" refers to circuits, devices, or software used to transmit detailed information about an anomaly to external parties.

[0010] A "generative model" refers to an algorithm or artificial intelligence technology that analyzes voice data to determine the risk level of dementia.

[0011] A "wearable device," literally translated, is a device that can be attached to the body, and refers to an electronic device that is attached to the human body to acquire biometric information.

[0012] "Biometric information" refers to data about health status that can be collected from the human body, such as heart rate and step count. [Brief explanation of the drawing]

[0013] [Figure 1]This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]

[0014] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0015] First, the terms used in the following description will be explained.

[0016] In the following embodiments, a tagged processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0017] In the following embodiments, a tagged RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0018] In the following embodiments, a tagged storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0019] In the following embodiments, a tagged communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), etc.

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0021] [First Embodiment]

[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0034] This invention relates to a system used in the living environment of elderly people, which monitors their lifestyle patterns using various sensors and devices and detects abnormalities. Specifically, this system consists of a terminal installed in the living space and a server that processes data in the background.

[0035] terminal

[0036] The device is placed in the user's living environment and is equipped with a camera and microphone to collect the user's actions and conversations in daily life. The device also works in conjunction with wearable devices to acquire biometric information (e.g., steps taken, heart rate, etc.). The collected data is transmitted to a server via a secure communication protocol.

[0037] server

[0038] The server analyzes the received image and audio data to model the typical daily routines of elderly individuals. The modeled data is compared with historical data to monitor for any anomalies. If an anomaly is detected, immediate action is taken, and the user is alerted via voice.

[0039] Next, the analyzed audio data is evaluated using a generative model to automatically determine the risk of dementia. This evaluation result is then used as valuable information for the user and their family and friends.

[0040] User

[0041] When an anomaly is detected, the user will receive appropriate verbal notifications from their device. For example, a message such as, "You woke up at [time], is everything alright?" The server will also send details of the anomaly to family members and care managers via email or app notifications.

[0042] This system eliminates the need for specialized care staff and automatically monitors the current situation of elderly individuals, prompting them to take necessary actions. For example, even during times when care services are unavailable, it can detect signs of nighttime wandering and respond quickly. Furthermore, if the risk of dementia is high, early diagnosis and treatment are encouraged, leading to more efficient risk management.

[0043] The following describes the processing flow.

[0044] Step 1:

[0045] The device acquires images and audio data from the living space through its camera and microphone. It also collects biometric information such as heart rate and steps taken from wearable devices.

[0046] Step 2:

[0047] The terminal converts the acquired data into a predetermined format and sends it to the server using a secure communication protocol.

[0048] Step 3:

[0049] The server analyzes the received image data using an image recognition algorithm and automatically classifies the user's behavior. This allows for the modeling of daily activity patterns.

[0050] Step 4:

[0051] The server converts the audio data into text using speech recognition technology, analyzes the conversation content, and assesses the risk of dementia.

[0052] Step 5:

[0053] The server models normal lifestyle patterns based on collected lifestyle data and detects anomalies by comparing them with the latest data in real time.

[0054] Step 6:

[0055] When an anomaly is detected, the server determines the appropriate course of action and generates an appropriate message.

[0056] Step 7:

[0057] The terminal provides warnings and advice to the user via voice messages, based on instructions from the server.

[0058] Step 8:

[0059] The server compiles the details of the detected anomalies and sends notifications to family members and care managers.

[0060] Step 9:

[0061] The server generates periodic reports at regular intervals and sends comprehensive analysis results regarding the user's lifestyle and health status to relevant parties.

[0062] (Example 1)

[0063] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0064] In the living environment of the elderly, it is important to monitor their health while maintaining their independence. However, it is not realistic to have professional care staff present at all times, so a system that can quickly detect and respond to abnormalities is needed. Furthermore, it is important to detect and address cognitive decline early. To solve these problems, a system is needed that can acquire and analyze data in an efficient and reliable manner.

[0065] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0066] In this invention, the server includes information acquisition means for acquiring behavior and conversation; analysis means for analyzing the acquired behavior and conversation data to generate a lifestyle model and detect anomalies; and notification means for providing voice warnings to the user when an anomaly is detected. This enables rapid detection and response to anomalies in the living environment of elderly people.

[0067] "Information acquisition means" refers to devices or technologies for sensing the behavior and conversations of elderly people and acquiring them as data.

[0068] "Analysis method" refers to the process of analyzing the lifestyle of elderly people based on acquired data and detecting deviations from normal patterns.

[0069] A "notification mechanism" is a system element that has the function of communicating a warning to the user via voice when an anomaly is detected.

[0070] "Information transmission means" refers to a mechanism for communicating details of detected anomalies to other relevant parties, and this is done via email, app notifications, etc.

[0071] A "generative model" is a computational model used to generate assessments of cognitive function in older adults from data.

[0072] A "portable device" is a device used to acquire a user's biometric information in real time, and it generally takes the form of something that can be worn on the body.

[0073] "Biometric information" refers to data that indicates a user's health status, including physical indicators such as heart rate and step count.

[0074] This invention is a system for rapidly detecting and responding to abnormalities in the living environment of elderly people. This system consists of a terminal for acquiring information, a server for analyzing the data, and a user interface for providing notifications.

[0075] The terminals are installed in the living environments of elderly people and are equipped with cameras and microphones to capture their actions and conversations. Additionally, a portable device worn by the user is used to acquire biometric information such as heart rate and step count in real time. This information is transmitted to a server via a secure protocol.

[0076] Specifically, the device can detect user movement and falls, for example, using a fixed camera in the room. If a user falls, the device immediately notifies the server so that appropriate action can be taken.

[0077] The server analyzes data transmitted from the terminal and models the lifestyle of the elderly. Based on the acquired video and audio data, a generative AI model performs predictive analysis and detects anomalies. Furthermore, it evaluates the risk of cognitive function from the audio data and saves the results. This evaluation is based on pre-configured prompts.

[0078] A concrete example of a prompt message is, "Based on the user's latest voice and biometric data, evaluate whether any anomalies have been detected and report the results of the cognitive function risk analysis." This generates an appropriate alert that responds to the situation.

[0079] The user is in a position to receive information provided by the system. If an anomaly is detected, an audio notification will be sent from the device. For example, a specific alert such as, "A value higher than your normal resting heart rate has been detected. We recommend that you take a break," will be conveyed to the user. Furthermore, the anomaly information is also sent to external parties and communicated to family members and care managers via email or application notifications.

[0080] This system enables quick and effective responses to ensure the safety of elderly individuals, even in the absence of specialized care staff. By utilizing biometric information, it allows for detailed management tailored to each individual's health condition.

[0081] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0082] Step 1:

[0083] The device uses a camera and microphone installed in the elderly person's living space to capture their actions and conversations. The input is real-time video and audio data. Specifically, the device continuously captures video frames and records audio signals. This data is used as foundational data for anomaly detection. The output is the captured video and audio data.

[0084] Step 2:

[0085] The terminal acquires biometric data from a portable device worn by the user. The input is biometric information such as heart rate and step count, which are updated in real time. Specifically, the terminal periodically receives data from the portable device using communication functions such as Bluetooth or Wi-Fi. The output is the collected biometric data.

[0086] Step 3:

[0087] The terminal transmits collected video, audio, and biometric data to the server via a secure protocol. The input is all data acquired by the terminal. Specifically, the terminal compresses and encrypts the data before transmitting it over the network. The output is the data transferred to the server.

[0088] Step 4:

[0089] The server receives the transmitted data and begins analysis. The input consists of video, audio, and biometric data sent from the terminal. The server first analyzes the video data to extract behavioral features. Next, it converts the audio data into text and analyzes the content of the conversation using natural language processing techniques. Furthermore, it analyzes the biometric data to detect deviations from normal patterns. The output is the analysis result.

[0090] Step 5:

[0091] The server uses a generative AI model to assess cognitive function risk from voice data. The input is pre-analyzed voice data. Specifically, the server inputs prompt sentences into the generative AI model and retrieves the corresponding risk analysis results. The output is the cognitive function risk assessment result.

[0092] Step 6:

[0093] The server detects anomalies based on the analysis results and notifies the user as necessary. The input consists of the analysis results and risk assessment results generated by the server. Specifically, the server determines the anomaly, generates a warning message using speech synthesis technology, and sends it to the terminal. The output is an audio warning to the user.

[0094] Step 7:

[0095] The server sends detailed information about the anomaly to external parties. The input is all information about the detected anomaly. Specifically, the server notifies parties via email or application. The output is the alert notification sent to the parties.

[0096] (Application Example 1)

[0097] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0098] The goal is to solve the problem of difficulty in quickly detecting abnormal movements or inappropriate postures while ensuring worker safety in work environments such as factories.

[0099] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0100] In this invention, the server includes a data acquisition device, a processing device that analyzes the acquired data to generate behavioral patterns and detect anomalies, a notification device that issues an audible warning to the user when an anomaly is detected, a communication device that transmits details of the anomaly to external parties, a device that learns the worker's movements, and a device that monitors the movements in the work environment. This improves the safety of workers in the work environment and enables rapid response to anomalies.

[0101] A "data acquisition device" is a device that collects data such as images and sounds from the surrounding environment.

[0102] A "processing device" is a device that analyzes acquired data, generates behavioral patterns, and has the function of detecting anomalies.

[0103] A "notification device" is a device that alerts the user via voice when an abnormality is detected.

[0104] A "communication device" is a device used to transmit details of detected anomalies to external parties.

[0105] A "device that learns worker movements" is a device that recognizes the movements of workers in a work environment such as a factory and learns those patterns.

[0106] A "device for monitoring actions" is a device for monitoring the actions of workers in the work environment in real time.

[0107] The system of this invention includes a device for ensuring and improving worker safety in a work environment such as a factory. The server, as a data acquisition device, collects image and audio data from the work environment using a camera and microphone. The processing device analyzes the collected data using image processing libraries such as OpenCV and machine learning frameworks such as TENSORFLOW® to generate worker behavior patterns and detects anomalies by comparing these patterns with actual behavior.

[0108] If an anomaly is detected, the notification device will issue an audio warning to the worker. For example, a message such as, "That posture is dangerous. Take a break," will be issued. Furthermore, details of the anomaly are sent to an external administrator via a communication device, allowing the administrator to respond quickly.

[0109] Furthermore, the device that learns the worker's movements analyzes past data to learn work patterns and promote appropriate work habits. This system, as a movement monitoring device, performs real-time work monitoring and improves safety on site.

[0110] The following is an example of a prompt message for a specific generative AI model.

[0111] "Use image data to analyze worker behavior patterns and design a machine learning model to detect anomalies. Propose how to issue warnings and notify managers when an anomaly is detected."

[0112] This system makes it possible to improve work efficiency and safety within the factory.

[0113] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0114] Step 1:

[0115] The terminal uses a camera and microphone to acquire image and audio data from the factory work environment. This input data includes worker actions and conversations. This data is transmitted to the server in real time.

[0116] Step 2:

[0117] The server processes the received image data using OpenCV to analyze the worker's posture and movements. Specifically, it detects human movement within the image and compares it to normal movement patterns. This process makes it possible to detect abnormal movements based on data from normal operation.

[0118] Step 3:

[0119] The server analyzes audio data using TensorFlow, converting the workers' conversations into text data. Then, it performs feature extraction and pattern analysis. This process enables the detection of abnormal conversation content and unusual sounds.

[0120] Step 4:

[0121] If the server's processing unit detects an anomaly based on the analysis results, it will issue an audible warning to the worker via a notification device. For example, if a worker is in a dangerous position, it will warn, "This is a dangerous position. Please be careful." This output plays a role in enhancing worker safety.

[0122] Step 5:

[0123] The server transmits details of detected anomalies to an external administrator via a communication device. This allows the administrator to take prompt action. This includes a timestamp of the anomaly and environmental data.

[0124] Step 6:

[0125] The server continuously learns work patterns and updates this information using a generative AI model. An example prompt might be, "Use image data to analyze worker behavior patterns and design a machine learning model to detect anomalies." This process ensures optimal anomaly detection at all times.

[0126] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0127] This invention relates to a system for use in the living environments of elderly people, which utilizes various sensors and devices to monitor and analyze lifestyle patterns and emotional states, and to provide appropriate responses. In particular, this system has the function of collecting image and audio data and simultaneously analyzing abnormal lifestyle patterns and emotional changes.

[0128] terminal

[0129] The terminal is installed in the user's living space and has the function of acquiring image and audio data through a camera and microphone. Furthermore, it can communicate with wearable devices and acquire biometric information. The acquired data is sequentially sent to a server for analysis in real time.

[0130] server

[0131] The server first analyzes the received data based on a lifestyle pattern model. This allows it to detect abnormalities if the data deviates from the normal lifestyle pattern. It also uses a generative model to analyze voice data and assess the risk of dementia.

[0132] Furthermore, this invention uses an emotion engine to recognize emotions from the user's voice and facial expression data. For example, if the user is showing signs of stress or anxiety, this engine captures that state and uses it as information for making appropriate decisions. The server analyzes the changes in emotion, generates a customized message based on that analysis, and provides it to the user.

[0133] User

[0134] Users receive notifications via voice messages, and advice based on abnormalities and emotional changes is provided through their devices. For example, if a user is clearly feeling anxious, the system can say, "Are you okay? We'll send help if you need it."

[0135] If emotional changes are deemed persistent or abnormal, the server notifies family members or care managers of the details via email or a dedicated app. This allows for prompt action and enhanced mental and physical health support.

[0136] This system allows for the automatic monitoring of elderly individuals' emotional states and living conditions, and the provision of appropriate support, without the need for constant supervision by specialized care staff. This enables more comprehensive and precise monitoring compared to conventional methods.

[0137] The following describes the processing flow.

[0138] Step 1:

[0139] The device uses cameras and microphones installed within the living space to collect the user's image and voice data in real time. It also acquires biometric information from wearable devices.

[0140] Step 2:

[0141] The device transmits collected images, audio, and biometric information to the server via a predetermined communication protocol. The data is encrypted during transmission, ensuring security.

[0142] Step 3:

[0143] The server converts the received audio data into text using speech recognition technology and analyzes the user's conversation. The analyzed data is then used with a generative model to assess the risk of dementia.

[0144] Step 4:

[0145] The server analyzes image data to identify the user's behavior. Based on the analysis results, it detects an anomaly if the user's lifestyle pattern deviates from the norm.

[0146] Step 5:

[0147] The server uses an emotion engine to analyze voice and facial expression data and evaluate the user's emotional state. It recognizes changes in emotion and signs of anxiety and stress.

[0148] Step 6:

[0149] When the server detects anomalies or changes in the user's emotional state, it generates a customized message. This message includes advice and warnings based on both the user's life and emotions.

[0150] Step 7:

[0151] The device, based on instructions from the server, delivers voice messages to the user through its speaker. For example, it might convey information about unusual circumstances in daily life or emotional anxieties.

[0152] Step 8:

[0153] The server notifies family members or care managers via email or app of details about any anomalies or emotional changes. This allows for a quick response if necessary.

[0154] Step 9:

[0155] The server periodically aggregates the analysis results and generates a report. This report provides important information about users' lifestyle patterns and emotional states and is sent to relevant parties on a regular basis.

[0156] (Example 2)

[0157] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0158] There is a growing need to monitor the living environment of the elderly and appropriately assess their emotions and health status in order to provide prompt and accurate support. However, achieving this requires a system that integrates data from multiple sources and analyzes it efficiently. Conventional methods have made it difficult to quickly capture changes in behavior and emotions, and appropriate responses have sometimes been delayed even when abnormalities occur. This invention aims to solve these problems.

[0159] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0160] In this invention, the server includes information gathering means for collecting images and sounds, analysis means for analyzing the collected image and sound information to create behavioral patterns and detect anomalies, and emotion recognition means for analyzing emotion information to evaluate the user's emotional state. This makes it possible to quickly detect anomalies and changes in emotions in the user's living environment and provide appropriate responses.

[0161] "Information gathering means" refers to devices and technologies for acquiring image and audio information, and has the function of collecting data in the user's living environment.

[0162] "Analysis means" refers to processes and technologies for analyzing collected image and audio information to identify user behavior patterns and anomalies.

[0163] "Emotion recognition means" refers to technologies and algorithms that evaluate a user's emotional state based on voice and facial expression data.

[0164] A "guidance system" is a function that notifies the user of detected anomalies or emotional states and directly conveys messages through voice or other means.

[0165] "Means of communication" refers to communication technologies and infrastructure used to transmit information about abnormal or important emotional states to external parties.

[0166] "Generative technology" refers to methods for performing specific risks and assessments using data generation models, which are used in the analysis of audio data and other similar materials.

[0167] A "wearable device" refers to a device that is attached to a user's body to acquire biometric data.

[0168] This invention is a system for monitoring the living environment of elderly people and evaluating abnormalities and emotional states. The main components necessary for carrying out the invention include a terminal for collecting information, a server for analyzing the data, and means for notifying the user.

[0169] The terminal is installed in the living space and uses a camera and microphone to collect the user's image and voice data in real time. This terminal can connect with wearable devices as needed and transmit information, including biometric data, to a server. For example, if the user wears a smartwatch, additional data such as heart rate and body temperature can be collected.

[0170] The server aggregates and analyzes data sent from the terminal. For image and audio data, advanced analytical algorithms are used to identify behavioral patterns and detect anomalies. Furthermore, a generative AI model is used to analyze audio data and provide prompts to assess the user's risk of cognitive decline. For example, by inputting "Assess the likelihood of cognitive decline based on the user's voice characteristics," the model provides an output.

[0171] Furthermore, the server utilizes an emotion engine to evaluate the user's emotional state. Based on this, if emotional changes such as stress or anxiety are identified, corrective measures are taken.

[0172] Users receive feedback on their health and emotional state through notifications provided by the system. For example, if a detected anomaly is related to everyday stress, the user may receive a voice notification such as, "You need to relax. Take a break," to facilitate appropriate support.

[0173] This system will create an environment where elderly people can live more independently and with greater peace of mind, and will also enable efficient and effective care support through the rapid provision of information to external stakeholders.

[0174] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0175] Step 1:

[0176] The device uses a camera and microphone to collect user image and audio data in real time. Specifically, the camera captures multiple frames of images per second, and the microphone continuously records ambient sounds. The device compresses this data and sends it to the server. The input is images and audio from the living space, and the output is a compressed data file.

[0177] Step 2:

[0178] The terminal acquires the user's biometric data by communicating with a wearable device. This data includes information such as heart rate and body temperature. The terminal receives data from the device via wireless communication such as Bluetooth and transmits it to the server. The input is biometric data from the wearable device, and the output is the data transmitted to the server.

[0179] Step 3:

[0180] The server receives image and audio data transmitted from the terminal and analyzes the behavioral patterns. The server compares this data with historical data and uses a statistical model to identify deviations from normal patterns. The input is image and audio data files, and the output is information about the detected anomalies.

[0181] Step 4:

[0182] The server analyzes voice data using a generative AI model to assess the risk of cognitive decline. It inputs a prompt sentence into the model and calculates a risk score. The input consists of voice data and the prompt sentence "Assess the likelihood of cognitive decline based on the user's voice characteristics," and the output is the generated risk assessment score.

[0183] Step 5:

[0184] The server uses an emotion engine to identify emotional states from voice and facial expression data. Data analysis identifies emotions such as stress or anxiety and determines how to respond to the user. The input is voice and facial expression data, and the output is the evaluation result of the emotional state.

[0185] Step 6:

[0186] The user receives an audio notification from their device. A response message generated by the server is sent to the device and provided to the user as an audio message. Specifically, advice such as "You need to relax. Please take a break." is delivered via audio. The input is the response message from the server, and the output is the audio notification to the user.

[0187] Step 7:

[0188] The server will send information to external parties if emotions or anomalies persist. Details of the situation and recommended actions will be communicated to relevant parties via email or a dedicated app. Input is the detected anomaly or emotional state, and output is the notification of information to external parties.

[0189] (Application Example 2)

[0190] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0191] In environments where elderly people live alone, there is a growing need for effective monitoring systems that can quickly and accurately detect everyday abnormalities and emotional changes, and take appropriate measures. Conventional methods have challenges in that it is difficult to quickly detect abnormalities and there is insufficient notification to family members or caregivers.

[0192] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes an information collection means for collecting images and sounds, a data analysis means for analyzing the collected image and sound information to create a lifestyle pattern model and detect abnormalities, and a notification means for conveying a voice warning to the user when an abnormality is detected. This makes it possible to immediately detect abnormalities in daily life and changes in emotions in the living space of elderly people and to quickly and accurately notify relevant parties.

[0193] "Information gathering means" refers to a device or system for acquiring image and audio information from an object.

[0194] "Data analysis means" refers to a device or method for processing collected image and audio information, creating a lifestyle pattern model, and detecting anomalies within it.

[0195] "Notification means" refers to a device or system for conveying warnings or notifications to users via voice.

[0196] "Information transmission means" refers to a device or system for transmitting collected and analyzed information to external parties.

[0197] "Computation processing means" refers to a device or system that operates on a cloud infrastructure and performs real-time analysis of collected information.

[0198] "Information distribution means" refers to a device or system for notifying external parties based on analysis results.

[0199] An "information processing system" is a device or method for processing information and performing specific analyses or evaluations.

[0200] An "information management system" is a device or system for managing acquired information and monitoring the status of users.

[0201] The system implementing this invention includes a terminal installed in the user's living space and a server operating on a cloud infrastructure. This section provides a detailed explanation of information gathering, information analysis, and notification.

[0202] The device is equipped with a camera and microphone, and collects image and audio information on a daily basis. This allows for the rapid detection of any anomalies occurring in the user's living space. The device transmits the collected data to a server via the internet.

[0203] The servers are located on a cloud infrastructure and analyze the collected data in real time. TensorFlow and other machine learning frameworks are used for data analysis. This analysis method creates a lifestyle pattern model and detects anomalies when users exhibit behavior that deviates from their daily routine. In addition, a generative AI model is used to evaluate emotional states from voice information and detect specific emotional changes.

[0204] Based on the analysis results, the server uses speech synthesis technology to deliver messages to users via their devices, providing voice warnings and advice. Furthermore, information regarding anomalies and emotional changes is sent to external parties via email or SMS using the Twilio API.

[0205] For example, if a user fails to take their usual afternoon walk, the server will generate a message saying, "It appears you haven't completed your morning exercise. If you are feeling unwell, we recommend taking a rest," and notify the user. It will also send a message to family members saying, "The user hasn't done their morning exercise yet. Please check on their health."

[0206] Examples of prompt statements for a generative AI model are as follows:

[0207] "Please create an algorithm that analyzes lifestyle pattern data and emotional voice data of elderly individuals to detect abnormalities."

[0208] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0209] Step 1:

[0210] The device collects image and audio information of the user in their living space using a camera and microphone. The input is raw data obtained from the camera and microphone, and the output is this raw data converted directly into a digital format. Specifically, it records images captured by the camera and audio recorded by the microphone in real time.

[0211] Step 2:

[0212] The device transmits collected image and audio data to a server via the internet. The input is processed digital data, and the output is data upload to a cloud server. Specifically, it reliably transmits data to the cloud using a communication module.

[0213] Step 3:

[0214] The server analyzes received images and audio and detects anomalies based on a lifestyle pattern model. The input is digital data sent from the terminal, and the output is the result of detecting anomalies in the lifestyle pattern. Specifically, it uses TensorFlow to perform pattern analysis with a pre-trained model and identify anomalies.

[0215] Step 4:

[0216] The server utilizes a generative AI model to analyze the user's emotional state from voice data. The input is voice data, and the output is the evaluation result of the emotional state. Specifically, it identifies the user's emotions through the voice data and performs voice analysis to capture emotional changes such as stress and anxiety.

[0217] Step 5:

[0218] The server uses speech synthesis technology to generate messages to be delivered to the user via the terminal, based on the results of anomaly detection and sentiment analysis. The input is the analysis results, and the output is the voice message to the user. Specifically, it generates appropriate advice and warning messages and delivers them to the user via the speaker.

[0219] Step 6:

[0220] The server uses the Twilio API to notify external stakeholders of anomalies and emotional changes obtained through analysis. Input is detailed information about anomalies and emotional changes, and output is the result of sending emails or SMS messages to stakeholders. The specific operation involves generating notification content and delivering the information quickly and reliably to stakeholders.

[0221] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0222] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0223] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0224] [Second Embodiment]

[0225] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0226] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0227] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0228] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0229] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0230] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0231] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0232] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0233] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0234] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0235] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0236] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0237] This invention relates to a system used in the living environment of elderly people, which monitors their lifestyle patterns using various sensors and devices and detects abnormalities. Specifically, this system consists of a terminal installed in the living space and a server that processes data in the background.

[0238] terminal

[0239] The device is placed in the user's living environment and is equipped with a camera and microphone to collect the user's actions and conversations in daily life. The device also works in conjunction with wearable devices to acquire biometric information (e.g., steps taken, heart rate, etc.). The collected data is transmitted to a server via a secure communication protocol.

[0240] server

[0241] The server analyzes the received image and audio data to model the typical daily routines of elderly individuals. The modeled data is compared with historical data to monitor for any anomalies. If an anomaly is detected, immediate action is taken, and the user is alerted via voice.

[0242] Next, the analyzed audio data is evaluated using a generative model to automatically determine the risk of dementia. This evaluation result is then used as valuable information for the user and their family and friends.

[0243] User

[0244] When an anomaly is detected, the user will receive appropriate verbal notifications from their device. For example, a message such as, "You woke up at [time], is everything alright?" The server will also send details of the anomaly to family members and care managers via email or app notifications.

[0245] This system eliminates the need for specialized care staff and automatically monitors the current situation of elderly individuals, prompting them to take necessary actions. For example, even during times when care services are unavailable, it can detect signs of nighttime wandering and respond quickly. Furthermore, if the risk of dementia is high, early diagnosis and treatment are encouraged, leading to more efficient risk management.

[0246] The following describes the processing flow.

[0247] Step 1:

[0248] The device acquires images and audio data from the living space through its camera and microphone. It also collects biometric information such as heart rate and steps taken from wearable devices.

[0249] Step 2:

[0250] The terminal converts the acquired data into a predetermined format and sends it to the server using a secure communication protocol.

[0251] Step 3:

[0252] The server analyzes the received image data using an image recognition algorithm and automatically classifies the user's behavior. This allows for the modeling of daily activity patterns.

[0253] Step 4:

[0254] The server converts the audio data into text using speech recognition technology, analyzes the conversation content, and assesses the risk of dementia.

[0255] Step 5:

[0256] The server models normal lifestyle patterns based on collected lifestyle data and detects anomalies by comparing them with the latest data in real time.

[0257] Step 6:

[0258] When an anomaly is detected, the server determines the appropriate course of action and generates an appropriate message.

[0259] Step 7:

[0260] The terminal provides warnings and advice to the user via voice messages, based on instructions from the server.

[0261] Step 8:

[0262] The server compiles the details of the detected anomalies and sends notifications to family members and care managers.

[0263] Step 9:

[0264] The server generates periodic reports at regular intervals and sends comprehensive analysis results regarding the user's lifestyle and health status to relevant parties.

[0265] (Example 1)

[0266] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0267] In the living environment of the elderly, it is important to monitor their health while maintaining their independence. However, it is not realistic to have professional care staff present at all times, so a system that can quickly detect and respond to abnormalities is needed. Furthermore, it is important to detect and address cognitive decline early. To solve these problems, a system is needed that can acquire and analyze data in an efficient and reliable manner.

[0268] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0269] In this invention, the server includes information acquisition means for acquiring behavior and conversation; analysis means for analyzing the acquired behavior and conversation data to generate a lifestyle model and detect anomalies; and notification means for providing voice warnings to the user when an anomaly is detected. This enables rapid detection and response to anomalies in the living environment of elderly people.

[0270] "Information acquisition means" refers to devices or technologies for sensing the behavior and conversations of elderly people and acquiring them as data.

[0271] "Analysis method" refers to the process of analyzing the lifestyle of elderly people based on acquired data and detecting deviations from normal patterns.

[0272] A "notification mechanism" is a system element that has the function of communicating a warning to the user via voice when an anomaly is detected.

[0273] "Information transmission means" refers to a mechanism for communicating details of detected anomalies to other relevant parties, and this is done via email, app notifications, etc.

[0274] A "generative model" is a computational model used to generate assessments of cognitive function in older adults from data.

[0275] A "portable device" is a device used to acquire a user's biometric information in real time, and it generally takes the form of something that can be worn on the body.

[0276] "Biometric information" refers to data that indicates a user's health status, including physical indicators such as heart rate and step count.

[0277] This invention is a system for rapidly detecting and responding to abnormalities in the living environment of elderly people. This system consists of a terminal for acquiring information, a server for analyzing the data, and a user interface for providing notifications.

[0278] The terminals are installed in the living environments of elderly people and are equipped with cameras and microphones to capture their actions and conversations. Additionally, a portable device worn by the user is used to acquire biometric information such as heart rate and step count in real time. This information is transmitted to a server via a secure protocol.

[0279] Specifically, the device can detect user movement and falls, for example, using a fixed camera in the room. If a user falls, the device immediately notifies the server so that appropriate action can be taken.

[0280] The server is responsible for analyzing the data sent from the terminal and modeling the lifestyle of the elderly. Based on the acquired video data and audio data, the generated AI model performs predictive analysis and detects abnormalities. Furthermore, it evaluates the risk of cognitive function from the audio data and saves the results. This evaluation is based on pre-set prompt sentences.

[0281] As a specific example of the prompt sentence, there is "Please evaluate whether an abnormality has been detected based on the user's latest audio and biometric information data and report the risk analysis results of cognitive function." Thereby, appropriate alerts corresponding to the situation are generated.

[0282] The user is in a position to receive the information provided by the system. When an abnormality is detected, an audio notification is sent from the terminal. For example, specific alerts such as "A value higher than the normal resting heart rate has been detected. It is recommended that you take a rest." are conveyed to the user. Furthermore, the abnormality information is also sent to external relevant parties and transmitted to family members and care managers via email or application notifications.

[0283] With this system, even without professional care staff, a quick and effective response can be made to ensure the safety of the elderly. By utilizing biometric information, detailed management based on individual health conditions can be realized.

[0284] The flow of the specific process in Example 1 will be described using FIG. 11.

[0285] Step 1:

[0286] The terminal uses cameras and microphones installed in the living space of the elderly to acquire actions and conversations. The input is video and audio data acquired in real time. As specific operations, the terminal continuously captures video frames and records audio signals. These data are used as basic data for abnormality detection. The output is the acquired video and audio data.

[0287] Step 2:

[0288] The terminal acquires biometric data from a portable device worn by the user. The input is biometric information such as heart rate and step count, which are updated in real time. Specifically, the terminal periodically receives data from the portable device using communication functions such as Bluetooth or Wi-Fi. The output is the collected biometric data.

[0289] Step 3:

[0290] The terminal transmits collected video, audio, and biometric data to the server via a secure protocol. The input is all data acquired by the terminal. Specifically, the terminal compresses and encrypts the data before transmitting it over the network. The output is the data transferred to the server.

[0291] Step 4:

[0292] The server receives the transmitted data and begins analysis. The input consists of video, audio, and biometric data sent from the terminal. The server first analyzes the video data to extract behavioral features. Next, it converts the audio data into text and analyzes the content of the conversation using natural language processing techniques. Furthermore, it analyzes the biometric data to detect deviations from normal patterns. The output is the analysis result.

[0293] Step 5:

[0294] The server uses a generative AI model to assess cognitive function risk from voice data. The input is pre-analyzed voice data. Specifically, the server inputs prompt sentences into the generative AI model and retrieves the corresponding risk analysis results. The output is the cognitive function risk assessment result.

[0295] Step 6:

[0296] The server detects anomalies based on the analysis results and notifies the user as necessary. The input consists of the analysis results and risk assessment results generated by the server. Specifically, the server determines the anomaly, generates a warning message using speech synthesis technology, and sends it to the terminal. The output is an audio warning to the user.

[0297] Step 7:

[0298] The server sends detailed information about the anomaly to external parties. The input is all information about the detected anomaly. Specifically, the server notifies parties via email or application. The output is the alert notification sent to the parties.

[0299] (Application Example 1)

[0300] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0301] The goal is to solve the problem of difficulty in quickly detecting abnormal movements or inappropriate postures while ensuring worker safety in work environments such as factories.

[0302] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0303] In this invention, the server includes a data acquisition device, a processing device that analyzes the acquired data to generate behavioral patterns and detect anomalies, a notification device that issues an audible warning to the user when an anomaly is detected, a communication device that transmits details of the anomaly to external parties, a device that learns the worker's movements, and a device that monitors the movements in the work environment. This improves the safety of workers in the work environment and enables rapid response to anomalies.

[0304] A "data acquisition device" is a device that collects data such as images and sounds from the surrounding environment.

[0305] The "processing device" is a device that analyzes the acquired data, generates an action pattern, and has the function of detecting abnormalities.

[0306] The "notification device" is a device for warning the user of the information by voice when an abnormality is detected.

[0307] The "communication device" is a device for transmitting the details of the detected abnormality to external parties concerned.

[0308] The "device for learning the operator's actions" is a device that recognizes the actions of an operator in a working environment such as a factory and learns the patterns thereof.

[0309] The "device for monitoring actions" is a device for monitoring the actions of an operator in a working environment in real time.

[0310] The system of this invention includes devices for ensuring the safety of an operator and improving safety in a working environment such as a factory. The server, as a device for acquiring data, collects image and audio data from the working environment using a camera and a microphone. Further, as a processing device, the collected data is analyzed using an image processing library such as OpenCV or a machine learning framework such as TensorFlow, an action pattern of the operator is generated, and an abnormality is detected by comparing the pattern with the actual action.

[0311] When an abnormality is detected, the notification device gives a voice warning to the operator. For example, a message such as "That posture is dangerous. Take a break" is issued. Further, the details of the abnormality are transmitted to an external administrator via the communication device so that the administrator can respond promptly.

[0312] Furthermore, the device that learns the worker's movements analyzes past data to learn work patterns and promote appropriate work habits. This system, as a movement monitoring device, performs real-time work monitoring and improves safety on site.

[0313] The following is an example of a prompt message for a specific generative AI model.

[0314] "Use image data to analyze worker behavior patterns and design a machine learning model to detect anomalies. Propose how to issue warnings and notify managers when an anomaly is detected."

[0315] This system makes it possible to improve work efficiency and safety within the factory.

[0316] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0317] Step 1:

[0318] The terminal uses a camera and microphone to acquire image and audio data from the factory work environment. This input data includes worker actions and conversations. This data is transmitted to the server in real time.

[0319] Step 2:

[0320] The server processes the received image data using OpenCV to analyze the worker's posture and movements. Specifically, it detects human movement within the image and compares it to normal movement patterns. This process makes it possible to detect abnormal movements based on data from normal operation.

[0321] Step 3:

[0322] The server analyzes audio data using TensorFlow, converting the workers' conversations into text data. Then, it performs feature extraction and pattern analysis. This process enables the detection of abnormal conversation content and unusual sounds.

[0323] Step 4:

[0324] If the server's processing unit detects an anomaly based on the analysis results, it will issue an audible warning to the worker via a notification device. For example, if a worker is in a dangerous position, it will warn, "This is a dangerous position. Please be careful." This output plays a role in enhancing worker safety.

[0325] Step 5:

[0326] The server transmits details of detected anomalies to an external administrator via a communication device. This allows the administrator to take prompt action. This includes a timestamp of the anomaly and environmental data.

[0327] Step 6:

[0328] The server continuously learns work patterns and updates this information using a generative AI model. An example prompt might be, "Use image data to analyze worker behavior patterns and design a machine learning model to detect anomalies." This process ensures optimal anomaly detection at all times.

[0329] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0330] This invention relates to a system for use in the living environments of elderly people, which utilizes various sensors and devices to monitor and analyze lifestyle patterns and emotional states, and to provide appropriate responses. In particular, this system has the function of collecting image and audio data and simultaneously analyzing abnormal lifestyle patterns and emotional changes.

[0331] terminal

[0332] The terminal is installed in the user's living space and has the function of acquiring image and audio data through a camera and microphone. Furthermore, it can communicate with wearable devices and acquire biometric information. The acquired data is sequentially sent to a server for analysis in real time.

[0333] server

[0334] The server first analyzes the received data based on a lifestyle pattern model. This allows it to detect abnormalities if the data deviates from the normal lifestyle pattern. It also uses a generative model to analyze voice data and assess the risk of dementia.

[0335] Furthermore, this invention uses an emotion engine to recognize emotions from the user's voice and facial expression data. For example, if the user is showing signs of stress or anxiety, this engine captures that state and uses it as information for making appropriate decisions. The server analyzes the changes in emotion, generates a customized message based on that analysis, and provides it to the user.

[0336] User

[0337] Users receive notifications via voice messages, and advice based on abnormalities and emotional changes is provided through their devices. For example, if a user is clearly feeling anxious, the system can say, "Are you okay? We'll send help if you need it."

[0338] If emotional changes are deemed persistent or abnormal, the server notifies family members or care managers of the details via email or a dedicated app. This allows for prompt action and enhanced mental and physical health support.

[0339] This system allows for the automatic monitoring of elderly individuals' emotional states and living conditions, and the provision of appropriate support, without the need for constant supervision by specialized care staff. This enables more comprehensive and precise monitoring compared to conventional methods.

[0340] The following describes the processing flow.

[0341] Step 1:

[0342] The device uses cameras and microphones installed within the living space to collect the user's image and voice data in real time. It also acquires biometric information from wearable devices.

[0343] Step 2:

[0344] The device transmits collected images, audio, and biometric information to the server via a predetermined communication protocol. The data is encrypted during transmission, ensuring security.

[0345] Step 3:

[0346] The server converts the received audio data into text using speech recognition technology and analyzes the user's conversation. The analyzed data is then used with a generative model to assess the risk of dementia.

[0347] Step 4:

[0348] The server analyzes image data to identify the user's behavior. Based on the analysis results, it detects an anomaly if the user's lifestyle pattern deviates from the norm.

[0349] Step 5:

[0350] The server uses an emotion engine to analyze voice and facial expression data and evaluate the user's emotional state. It recognizes changes in emotion and signs of anxiety and stress.

[0351] Step 6:

[0352] When the server detects anomalies or changes in the user's emotional state, it generates a customized message. This message includes advice and warnings based on both the user's life and emotions.

[0353] Step 7:

[0354] The device, based on instructions from the server, delivers voice messages to the user through its speaker. For example, it might convey information about unusual circumstances in daily life or emotional anxieties.

[0355] Step 8:

[0356] The server notifies family members or care managers via email or app of details about any anomalies or emotional changes. This allows for a quick response if necessary.

[0357] Step 9:

[0358] The server periodically aggregates the analysis results and generates a report. This report provides important information about users' lifestyle patterns and emotional states and is sent to relevant parties on a regular basis.

[0359] (Example 2)

[0360] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0361] There is a growing need to monitor the living environment of the elderly and appropriately assess their emotions and health status in order to provide prompt and accurate support. However, achieving this requires a system that integrates data from multiple sources and analyzes it efficiently. Conventional methods have made it difficult to quickly capture changes in behavior and emotions, and appropriate responses have sometimes been delayed even when abnormalities occur. This invention aims to solve these problems.

[0362] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0363] In this invention, the server includes information gathering means for collecting images and sounds, analysis means for analyzing the collected image and sound information to create behavioral patterns and detect anomalies, and emotion recognition means for analyzing emotion information to evaluate the user's emotional state. This makes it possible to quickly detect anomalies and changes in emotions in the user's living environment and provide appropriate responses.

[0364] "Information gathering means" refers to devices and technologies for acquiring image and audio information, and has the function of collecting data in the user's living environment.

[0365] "Analysis means" refers to processes and technologies for analyzing collected image and audio information to identify user behavior patterns and anomalies.

[0366] "Emotion recognition means" refers to technologies and algorithms that evaluate a user's emotional state based on voice and facial expression data.

[0367] A "guidance system" is a function that notifies the user of detected anomalies or emotional states and directly conveys messages through voice or other means.

[0368] "Means of communication" refers to communication technologies and infrastructure used to transmit information about abnormal or important emotional states to external parties.

[0369] "Generative technology" refers to methods for performing specific risks and assessments using data generation models, which are used in the analysis of audio data and other similar materials.

[0370] A "wearable device" refers to a device that is attached to a user's body to acquire biometric data.

[0371] This invention is a system for monitoring the living environment of elderly people and evaluating abnormalities and emotional states. The main components necessary for carrying out the invention include a terminal for collecting information, a server for analyzing the data, and means for notifying the user.

[0372] The terminal is installed in the living space and uses a camera and microphone to collect the user's image and voice data in real time. This terminal can connect with wearable devices as needed and transmit information, including biometric data, to a server. For example, if the user wears a smartwatch, additional data such as heart rate and body temperature can be collected.

[0373] The server aggregates and analyzes data sent from the terminal. For image and audio data, advanced analytical algorithms are used to identify behavioral patterns and detect anomalies. Furthermore, a generative AI model is used to analyze audio data and provide prompts to assess the user's risk of cognitive decline. For example, by inputting "Assess the likelihood of cognitive decline based on the user's voice characteristics," the model provides an output.

[0374] Furthermore, the server utilizes an emotion engine to evaluate the user's emotional state. Based on this, if emotional changes such as stress or anxiety are identified, corrective measures are taken.

[0375] Users receive feedback on their health and emotional state through notifications provided by the system. For example, if a detected anomaly is related to everyday stress, the user may receive a voice notification such as, "You need to relax. Take a break," to facilitate appropriate support.

[0376] This system will create an environment where elderly people can live more independently and with greater peace of mind, and will also enable efficient and effective care support through the rapid provision of information to external stakeholders.

[0377] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0378] Step 1:

[0379] The device uses a camera and microphone to collect user image and audio data in real time. Specifically, the camera captures multiple frames of images per second, and the microphone continuously records ambient sounds. The device compresses this data and sends it to the server. The input is images and audio from the living space, and the output is a compressed data file.

[0380] Step 2:

[0381] The terminal acquires the user's biometric data by communicating with a wearable device. This data includes information such as heart rate and body temperature. The terminal receives data from the device via wireless communication such as Bluetooth and transmits it to the server. The input is biometric data from the wearable device, and the output is the data transmitted to the server.

[0382] Step 3:

[0383] The server receives image and audio data transmitted from the terminal and analyzes the behavioral patterns. The server compares this data with historical data and uses a statistical model to identify deviations from normal patterns. The input is image and audio data files, and the output is information about the detected anomalies.

[0384] Step 4:

[0385] The server analyzes voice data using a generative AI model to assess the risk of cognitive decline. It inputs a prompt sentence into the model and calculates a risk score. The input consists of voice data and the prompt sentence "Assess the likelihood of cognitive decline based on the user's voice characteristics," and the output is the generated risk assessment score.

[0386] Step 5:

[0387] The server uses an emotion engine to identify emotional states from voice and facial expression data. Data analysis identifies emotions such as stress or anxiety and determines how to respond to the user. The input is voice and facial expression data, and the output is the evaluation result of the emotional state.

[0388] Step 6:

[0389] The user receives an audio notification from their device. A response message generated by the server is sent to the device and provided to the user as an audio message. Specifically, advice such as "You need to relax. Please take a break." is delivered via audio. The input is the response message from the server, and the output is the audio notification to the user.

[0390] Step 7:

[0391] The server will send information to external parties if emotions or anomalies persist. Details of the situation and recommended actions will be communicated to relevant parties via email or a dedicated app. Input is the detected anomaly or emotional state, and output is the notification of information to external parties.

[0392] (Application Example 2)

[0393] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0394] In environments where elderly people live alone, there is a growing need for effective monitoring systems that can quickly and accurately detect everyday abnormalities and emotional changes, and take appropriate measures. Conventional methods have challenges in that it is difficult to quickly detect abnormalities and there is insufficient notification to family members or caregivers.

[0395] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes an information collection means for collecting images and sounds, a data analysis means for analyzing the collected image and sound information to create a lifestyle pattern model and detect abnormalities, and a notification means for conveying a voice warning to the user when an abnormality is detected. This makes it possible to immediately detect abnormalities in daily life and changes in emotions in the living space of elderly people and to quickly and accurately notify relevant parties.

[0396] "Information gathering means" refers to a device or system for acquiring image and audio information from an object.

[0397] "Data analysis means" refers to a device or method for processing collected image and audio information, creating a lifestyle pattern model, and detecting anomalies within it.

[0398] "Notification means" refers to a device or system for conveying warnings or notifications to users via voice.

[0399] "Information transmission means" refers to a device or system for transmitting collected and analyzed information to external parties.

[0400] "Computation processing means" refers to a device or system that operates on a cloud infrastructure and performs real-time analysis of collected information.

[0401] "Information distribution means" refers to a device or system for notifying external parties based on analysis results.

[0402] An "information processing system" is a device or method for processing information and performing specific analyses or evaluations.

[0403] An "information management system" is a device or system for managing acquired information and monitoring the status of users.

[0404] The system implementing this invention includes a terminal installed in the user's living space and a server operating on a cloud infrastructure. This section provides a detailed explanation of information gathering, information analysis, and notification.

[0405] The device is equipped with a camera and microphone, and collects image and audio information on a daily basis. This allows for the rapid detection of any anomalies occurring in the user's living space. The device transmits the collected data to a server via the internet.

[0406] The servers are located on a cloud infrastructure and analyze the collected data in real time. TensorFlow and other machine learning frameworks are used for data analysis. This analysis method creates a lifestyle pattern model and detects anomalies when users exhibit behavior that deviates from their daily routine. In addition, a generative AI model is used to evaluate emotional states from voice information and detect specific emotional changes.

[0407] Based on the analysis results, the server uses speech synthesis technology to deliver messages to users via their devices, providing voice warnings and advice. Furthermore, information regarding anomalies and emotional changes is sent to external parties via email or SMS using the Twilio API.

[0408] For example, if a user fails to take their usual afternoon walk, the server will generate a message saying, "It appears you haven't completed your morning exercise. If you are feeling unwell, we recommend taking a rest," and notify the user. It will also send a message to family members saying, "The user hasn't done their morning exercise yet. Please check on their health."

[0409] Examples of prompt statements for a generative AI model are as follows:

[0410] "Please create an algorithm that analyzes lifestyle pattern data and emotional voice data of elderly individuals to detect abnormalities."

[0411] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0412] Step 1:

[0413] The device collects image and audio information of the user in their living space using a camera and microphone. The input is raw data obtained from the camera and microphone, and the output is this raw data converted directly into a digital format. Specifically, it records images captured by the camera and audio recorded by the microphone in real time.

[0414] Step 2:

[0415] The device transmits collected image and audio data to a server via the internet. The input is processed digital data, and the output is data upload to a cloud server. Specifically, it reliably transmits data to the cloud using a communication module.

[0416] Step 3:

[0417] The server analyzes received images and audio and detects anomalies based on a lifestyle pattern model. The input is digital data sent from the terminal, and the output is the result of detecting anomalies in the lifestyle pattern. Specifically, it uses TensorFlow to perform pattern analysis with a pre-trained model and identify anomalies.

[0418] Step 4:

[0419] The server utilizes a generative AI model to analyze the user's emotional state from voice data. The input is voice data, and the output is the evaluation result of the emotional state. Specifically, it identifies the user's emotions through the voice data and performs voice analysis to capture emotional changes such as stress and anxiety.

[0420] Step 5:

[0421] The server uses speech synthesis technology to generate messages to be delivered to the user via the terminal, based on the results of anomaly detection and sentiment analysis. The input is the analysis results, and the output is the voice message to the user. Specifically, it generates appropriate advice and warning messages and delivers them to the user via the speaker.

[0422] Step 6:

[0423] The server uses the Twilio API to notify external stakeholders of anomalies and emotional changes obtained through analysis. Input is detailed information about anomalies and emotional changes, and output is the result of sending emails or SMS messages to stakeholders. The specific operation involves generating notification content and delivering the information quickly and reliably to stakeholders.

[0424] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0425] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0426] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0427] [Third Embodiment]

[0428] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0429] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0430] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0431] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0432] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0433] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0434] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0435] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0436] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0437] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0438] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0439] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0440] This invention relates to a system used in the living environment of elderly people, which monitors their lifestyle patterns using various sensors and devices and detects abnormalities. Specifically, this system consists of a terminal installed in the living space and a server that processes data in the background.

[0441] terminal

[0442] The device is placed in the user's living environment and is equipped with a camera and microphone to collect the user's actions and conversations in daily life. The device also works in conjunction with wearable devices to acquire biometric information (e.g., steps taken, heart rate, etc.). The collected data is transmitted to a server via a secure communication protocol.

[0443] server

[0444] The server analyzes the received image and audio data to model the typical daily routines of elderly individuals. The modeled data is compared with historical data to monitor for any anomalies. If an anomaly is detected, immediate action is taken, and the user is alerted via voice.

[0445] Next, the analyzed audio data is evaluated using a generative model to automatically determine the risk of dementia. This evaluation result is then used as valuable information for the user and their family and friends.

[0446] User

[0447] When an anomaly is detected, the user will receive appropriate verbal notifications from their device. For example, a message such as, "You woke up at [time], is everything alright?" The server will also send details of the anomaly to family members and care managers via email or app notifications.

[0448] This system eliminates the need for specialized care staff and automatically monitors the current situation of elderly individuals, prompting them to take necessary actions. For example, even during times when care services are unavailable, it can detect signs of nighttime wandering and respond quickly. Furthermore, if the risk of dementia is high, early diagnosis and treatment are encouraged, leading to more efficient risk management.

[0449] The following describes the processing flow.

[0450] Step 1:

[0451] The device acquires images and audio data from the living space through its camera and microphone. It also collects biometric information such as heart rate and steps taken from wearable devices.

[0452] Step 2:

[0453] The terminal converts the acquired data into a predetermined format and sends it to the server using a secure communication protocol.

[0454] Step 3:

[0455] The server analyzes the received image data using an image recognition algorithm and automatically classifies the user's behavior. This allows for the modeling of daily activity patterns.

[0456] Step 4:

[0457] The server converts the audio data into text using speech recognition technology, analyzes the conversation content, and assesses the risk of dementia.

[0458] Step 5:

[0459] The server models normal lifestyle patterns based on collected lifestyle data and detects anomalies by comparing them with the latest data in real time.

[0460] Step 6:

[0461] When an anomaly is detected, the server determines the appropriate course of action and generates an appropriate message.

[0462] Step 7:

[0463] The terminal provides warnings and advice to the user via voice messages, based on instructions from the server.

[0464] Step 8:

[0465] The server compiles the details of the detected anomalies and sends notifications to family members and care managers.

[0466] Step 9:

[0467] The server generates periodic reports at regular intervals and sends comprehensive analysis results regarding the user's lifestyle and health status to relevant parties.

[0468] (Example 1)

[0469] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0470] In the living environment of the elderly, it is important to monitor their health while maintaining their independence. However, it is not realistic to have professional care staff present at all times, so a system that can quickly detect and respond to abnormalities is needed. Furthermore, it is important to detect and address cognitive decline early. To solve these problems, a system is needed that can acquire and analyze data in an efficient and reliable manner.

[0471] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0472] In this invention, the server includes information acquisition means for acquiring behavior and conversation; analysis means for analyzing the acquired behavior and conversation data to generate a lifestyle model and detect anomalies; and notification means for providing voice warnings to the user when an anomaly is detected. This enables rapid detection and response to anomalies in the living environment of elderly people.

[0473] "Information acquisition means" refers to devices or technologies for sensing the behavior and conversations of elderly people and acquiring them as data.

[0474] "Analysis method" refers to the process of analyzing the lifestyle of elderly people based on acquired data and detecting deviations from normal patterns.

[0475] A "notification mechanism" is a system element that has the function of communicating a warning to the user via voice when an anomaly is detected.

[0476] "Information transmission means" refers to a mechanism for communicating details of detected anomalies to other relevant parties, and this is done via email, app notifications, etc.

[0477] A "generative model" is a computational model used to generate assessments of cognitive function in older adults from data.

[0478] A "portable device" is a device used to acquire a user's biometric information in real time, and it generally takes the form of something that can be worn on the body.

[0479] "Biometric information" refers to data that indicates a user's health status, including physical indicators such as heart rate and step count.

[0480] This invention is a system for rapidly detecting and responding to abnormalities in the living environment of elderly people. This system consists of a terminal for acquiring information, a server for analyzing the data, and a user interface for providing notifications.

[0481] The terminals are installed in the living environments of elderly people and are equipped with cameras and microphones to capture their actions and conversations. Additionally, a portable device worn by the user is used to acquire biometric information such as heart rate and step count in real time. This information is transmitted to a server via a secure protocol.

[0482] Specifically, the device can detect user movement and falls, for example, using a fixed camera in the room. If a user falls, the device immediately notifies the server so that appropriate action can be taken.

[0483] The server analyzes data transmitted from the terminal and models the lifestyle of the elderly. Based on the acquired video and audio data, a generative AI model performs predictive analysis and detects anomalies. Furthermore, it evaluates the risk of cognitive function from the audio data and saves the results. This evaluation is based on pre-configured prompts.

[0484] A concrete example of a prompt message is, "Based on the user's latest voice and biometric data, evaluate whether any anomalies have been detected and report the results of the cognitive function risk analysis." This generates an appropriate alert that responds to the situation.

[0485] The user is in a position to receive information provided by the system. If an anomaly is detected, an audio notification will be sent from the device. For example, a specific alert such as, "A value higher than your normal resting heart rate has been detected. We recommend that you take a break," will be conveyed to the user. Furthermore, the anomaly information is also sent to external parties and communicated to family members and care managers via email or application notifications.

[0486] This system enables quick and effective responses to ensure the safety of elderly individuals, even in the absence of specialized care staff. By utilizing biometric information, it allows for detailed management tailored to each individual's health condition.

[0487] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0488] Step 1:

[0489] The device uses a camera and microphone installed in the elderly person's living space to capture their actions and conversations. The input is real-time video and audio data. Specifically, the device continuously captures video frames and records audio signals. This data is used as foundational data for anomaly detection. The output is the captured video and audio data.

[0490] Step 2:

[0491] The terminal acquires biometric data from a portable device worn by the user. The input is biometric information such as heart rate and step count, which are updated in real time. Specifically, the terminal periodically receives data from the portable device using communication functions such as Bluetooth or Wi-Fi. The output is the collected biometric data.

[0492] Step 3:

[0493] The terminal transmits collected video, audio, and biometric data to the server via a secure protocol. The input is all data acquired by the terminal. Specifically, the terminal compresses and encrypts the data before transmitting it over the network. The output is the data transferred to the server.

[0494] Step 4:

[0495] The server receives the transmitted data and begins analysis. The input consists of video, audio, and biometric data sent from the terminal. The server first analyzes the video data to extract behavioral features. Next, it converts the audio data into text and analyzes the content of the conversation using natural language processing techniques. Furthermore, it analyzes the biometric data to detect deviations from normal patterns. The output is the analysis result.

[0496] Step 5:

[0497] The server uses a generative AI model to assess cognitive function risk from voice data. The input is pre-analyzed voice data. Specifically, the server inputs prompt sentences into the generative AI model and retrieves the corresponding risk analysis results. The output is the cognitive function risk assessment result.

[0498] Step 6:

[0499] The server detects anomalies based on the analysis results and notifies the user as necessary. The input consists of the analysis results and risk assessment results generated by the server. Specifically, the server determines the anomaly, generates a warning message using speech synthesis technology, and sends it to the terminal. The output is an audio warning to the user.

[0500] Step 7:

[0501] The server sends detailed information about the anomaly to external parties. The input is all information about the detected anomaly. Specifically, the server notifies parties via email or application. The output is the alert notification sent to the parties.

[0502] (Application Example 1)

[0503] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0504] The goal is to solve the problem of difficulty in quickly detecting abnormal movements or inappropriate postures while ensuring worker safety in work environments such as factories.

[0505] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0506] In this invention, the server includes a data acquisition device, a processing device that analyzes the acquired data to generate behavioral patterns and detect anomalies, a notification device that issues an audible warning to the user when an anomaly is detected, a communication device that transmits details of the anomaly to external parties, a device that learns the worker's movements, and a device that monitors the movements in the work environment. This improves the safety of workers in the work environment and enables rapid response to anomalies.

[0507] A "data acquisition device" is a device that collects data such as images and sounds from the surrounding environment.

[0508] A "processing device" is a device that analyzes acquired data, generates behavioral patterns, and has the function of detecting anomalies.

[0509] A "notification device" is a device that alerts the user via voice when an abnormality is detected.

[0510] A "communication device" is a device used to transmit details of detected anomalies to external parties.

[0511] A "device that learns worker movements" is a device that recognizes the movements of workers in a work environment such as a factory and learns those patterns.

[0512] A "device for monitoring actions" is a device for monitoring the actions of workers in the work environment in real time.

[0513] The system of this invention includes a device for ensuring and improving worker safety in a work environment such as a factory. The server, as a data acquisition device, collects image and audio data from the work environment using a camera and microphone. The processing device analyzes the collected data using image processing libraries such as OpenCV and machine learning frameworks such as TensorFlow to generate worker behavior patterns and detects anomalies by comparing these patterns with actual behavior.

[0514] If an anomaly is detected, the notification device will issue an audio warning to the worker. For example, a message such as, "That posture is dangerous. Take a break," will be issued. Furthermore, details of the anomaly are sent to an external administrator via a communication device, allowing the administrator to respond quickly.

[0515] Furthermore, the device that learns the worker's movements analyzes past data to learn work patterns and promote appropriate work habits. This system, as a movement monitoring device, performs real-time work monitoring and improves safety on site.

[0516] The following is an example of a prompt message for a specific generative AI model.

[0517] "Use image data to analyze worker behavior patterns and design a machine learning model to detect anomalies. Propose how to issue warnings and notify managers when an anomaly is detected."

[0518] This system makes it possible to improve work efficiency and safety within the factory.

[0519] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0520] Step 1:

[0521] The terminal uses a camera and microphone to acquire image and audio data from the factory work environment. This input data includes worker actions and conversations. This data is transmitted to the server in real time.

[0522] Step 2:

[0523] The server processes the received image data using OpenCV to analyze the worker's posture and movements. Specifically, it detects human movement within the image and compares it to normal movement patterns. This process makes it possible to detect abnormal movements based on data from normal operation.

[0524] Step 3:

[0525] The server analyzes audio data using TensorFlow, converting the workers' conversations into text data. Then, it performs feature extraction and pattern analysis. This process enables the detection of abnormal conversation content and unusual sounds.

[0526] Step 4:

[0527] If the server's processing unit detects an anomaly based on the analysis results, it will issue an audible warning to the worker via a notification device. For example, if a worker is in a dangerous position, it will warn, "This is a dangerous position. Please be careful." This output plays a role in enhancing worker safety.

[0528] Step 5:

[0529] The server transmits details of detected anomalies to an external administrator via a communication device. This allows the administrator to take prompt action. This includes a timestamp of the anomaly and environmental data.

[0530] Step 6:

[0531] The server continuously learns work patterns and updates this information using a generative AI model. An example prompt might be, "Use image data to analyze worker behavior patterns and design a machine learning model to detect anomalies." This process ensures optimal anomaly detection at all times.

[0532] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0533] This invention relates to a system for use in the living environments of elderly people, which utilizes various sensors and devices to monitor and analyze lifestyle patterns and emotional states, and to provide appropriate responses. In particular, this system has the function of collecting image and audio data and simultaneously analyzing abnormal lifestyle patterns and emotional changes.

[0534] terminal

[0535] The terminal is installed in the user's living space and has the function of acquiring image and audio data through a camera and microphone. Furthermore, it can communicate with wearable devices and acquire biometric information. The acquired data is sequentially sent to a server for analysis in real time.

[0536] server

[0537] The server first analyzes the received data based on a lifestyle pattern model. This allows it to detect abnormalities if the data deviates from the normal lifestyle pattern. It also uses a generative model to analyze voice data and assess the risk of dementia.

[0538] Furthermore, this invention uses an emotion engine to recognize emotions from the user's voice and facial expression data. For example, if the user is showing signs of stress or anxiety, this engine captures that state and uses it as information for making appropriate decisions. The server analyzes the changes in emotion, generates a customized message based on that analysis, and provides it to the user.

[0539] User

[0540] Users receive notifications via voice messages, and advice based on abnormalities and emotional changes is provided through their devices. For example, if a user is clearly feeling anxious, the system can say, "Are you okay? We'll send help if you need it."

[0541] If emotional changes are deemed persistent or abnormal, the server notifies family members or care managers of the details via email or a dedicated app. This allows for prompt action and enhanced mental and physical health support.

[0542] This system allows for the automatic monitoring of elderly individuals' emotional states and living conditions, and the provision of appropriate support, without the need for constant supervision by specialized care staff. This enables more comprehensive and precise monitoring compared to conventional methods.

[0543] The following describes the processing flow.

[0544] Step 1:

[0545] The device uses cameras and microphones installed within the living space to collect the user's image and voice data in real time. It also acquires biometric information from wearable devices.

[0546] Step 2:

[0547] The device transmits collected images, audio, and biometric information to the server via a predetermined communication protocol. The data is encrypted during transmission, ensuring security.

[0548] Step 3:

[0549] The server converts the received audio data into text using speech recognition technology and analyzes the user's conversation. The analyzed data is then used with a generative model to assess the risk of dementia.

[0550] Step 4:

[0551] The server analyzes image data to identify the user's behavior. Based on the analysis results, it detects an anomaly if the user's lifestyle pattern deviates from the norm.

[0552] Step 5:

[0553] The server uses an emotion engine to analyze voice and facial expression data and evaluate the user's emotional state. It recognizes changes in emotion and signs of anxiety and stress.

[0554] Step 6:

[0555] When the server detects anomalies or changes in the user's emotional state, it generates a customized message. This message includes advice and warnings based on both the user's life and emotions.

[0556] Step 7:

[0557] The device, based on instructions from the server, delivers voice messages to the user through its speaker. For example, it might convey information about unusual circumstances in daily life or emotional anxieties.

[0558] Step 8:

[0559] The server notifies family members or care managers via email or app of details about any anomalies or emotional changes. This allows for a quick response if necessary.

[0560] Step 9:

[0561] The server periodically aggregates the analysis results and generates a report. This report provides important information about users' lifestyle patterns and emotional states and is sent to relevant parties on a regular basis.

[0562] (Example 2)

[0563] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0564] There is a growing need to monitor the living environment of the elderly and appropriately assess their emotions and health status in order to provide prompt and accurate support. However, achieving this requires a system that integrates data from multiple sources and analyzes it efficiently. Conventional methods have made it difficult to quickly capture changes in behavior and emotions, and appropriate responses have sometimes been delayed even when abnormalities occur. This invention aims to solve these problems.

[0565] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0566] In this invention, the server includes information gathering means for collecting images and sounds, analysis means for analyzing the collected image and sound information to create behavioral patterns and detect anomalies, and emotion recognition means for analyzing emotion information to evaluate the user's emotional state. This makes it possible to quickly detect anomalies and changes in emotions in the user's living environment and provide appropriate responses.

[0567] "Information gathering means" refers to devices and technologies for acquiring image and audio information, and has the function of collecting data in the user's living environment.

[0568] "Analysis means" refers to processes and technologies for analyzing collected image and audio information to identify user behavior patterns and anomalies.

[0569] "Emotion recognition means" refers to technologies and algorithms that evaluate a user's emotional state based on voice and facial expression data.

[0570] A "guidance system" is a function that notifies the user of detected anomalies or emotional states and directly conveys messages through voice or other means.

[0571] "Means of communication" refers to communication technologies and infrastructure used to transmit information about abnormal or important emotional states to external parties.

[0572] "Generative technology" refers to methods for performing specific risks and assessments using data generation models, which are used in the analysis of audio data and other similar materials.

[0573] A "wearable device" refers to a device that is attached to a user's body to acquire biometric data.

[0574] This invention is a system for monitoring the living environment of elderly people and evaluating abnormalities and emotional states. The main components necessary for carrying out the invention include a terminal for collecting information, a server for analyzing the data, and means for notifying the user.

[0575] The terminal is installed in the living space and uses a camera and microphone to collect the user's image and voice data in real time. This terminal can connect with wearable devices as needed and transmit information, including biometric data, to a server. For example, if the user wears a smartwatch, additional data such as heart rate and body temperature can be collected.

[0576] The server aggregates and analyzes data sent from the terminal. For image and audio data, advanced analytical algorithms are used to identify behavioral patterns and detect anomalies. Furthermore, a generative AI model is used to analyze audio data and provide prompts to assess the user's risk of cognitive decline. For example, by inputting "Assess the likelihood of cognitive decline based on the user's voice characteristics," the model provides an output.

[0577] Furthermore, the server utilizes an emotion engine to evaluate the user's emotional state. Based on this, if emotional changes such as stress or anxiety are identified, corrective measures are taken.

[0578] Users receive feedback on their health and emotional state through notifications provided by the system. For example, if a detected anomaly is related to everyday stress, the user may receive a voice notification such as, "You need to relax. Take a break," to facilitate appropriate support.

[0579] This system will create an environment where elderly people can live more independently and with greater peace of mind, and will also enable efficient and effective care support through the rapid provision of information to external stakeholders.

[0580] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0581] Step 1:

[0582] The device uses a camera and microphone to collect user image and audio data in real time. Specifically, the camera captures multiple frames of images per second, and the microphone continuously records ambient sounds. The device compresses this data and sends it to the server. The input is images and audio from the living space, and the output is a compressed data file.

[0583] Step 2:

[0584] The terminal acquires the user's biometric data by communicating with a wearable device. This data includes information such as heart rate and body temperature. The terminal receives data from the device via wireless communication such as Bluetooth and transmits it to the server. The input is biometric data from the wearable device, and the output is the data transmitted to the server.

[0585] Step 3:

[0586] The server receives image and audio data transmitted from the terminal and analyzes the behavioral patterns. The server compares this data with historical data and uses a statistical model to identify deviations from normal patterns. The input is image and audio data files, and the output is information about the detected anomalies.

[0587] Step 4:

[0588] The server analyzes voice data using a generative AI model to assess the risk of cognitive decline. It inputs a prompt sentence into the model and calculates a risk score. The input consists of voice data and the prompt sentence "Assess the likelihood of cognitive decline based on the user's voice characteristics," and the output is the generated risk assessment score.

[0589] Step 5:

[0590] The server uses an emotion engine to identify emotional states from voice and facial expression data. Data analysis identifies emotions such as stress or anxiety and determines how to respond to the user. The input is voice and facial expression data, and the output is the evaluation result of the emotional state.

[0591] Step 6:

[0592] The user receives an audio notification from their device. A response message generated by the server is sent to the device and provided to the user as an audio message. Specifically, advice such as "You need to relax. Please take a break." is delivered via audio. The input is the response message from the server, and the output is the audio notification to the user.

[0593] Step 7:

[0594] The server will send information to external parties if emotions or anomalies persist. Details of the situation and recommended actions will be communicated to relevant parties via email or a dedicated app. Input is the detected anomaly or emotional state, and output is the notification of information to external parties.

[0595] (Application Example 2)

[0596] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0597] In environments where elderly people live alone, there is a growing need for effective monitoring systems that can quickly and accurately detect everyday abnormalities and emotional changes, and take appropriate measures. Conventional methods have challenges in that it is difficult to quickly detect abnormalities and there is insufficient notification to family members or caregivers.

[0598] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes an information collection means for collecting images and sounds, a data analysis means for analyzing the collected image and sound information to create a lifestyle pattern model and detect abnormalities, and a notification means for conveying a voice warning to the user when an abnormality is detected. This makes it possible to immediately detect abnormalities in daily life and changes in emotions in the living space of elderly people and to quickly and accurately notify relevant parties.

[0599] "Information gathering means" refers to a device or system for acquiring image and audio information from an object.

[0600] "Data analysis means" refers to a device or method for processing collected image and audio information, creating a lifestyle pattern model, and detecting anomalies within it.

[0601] "Notification means" refers to a device or system for conveying warnings or notifications to users via voice.

[0602] "Information transmission means" refers to a device or system for transmitting collected and analyzed information to external parties.

[0603] "Computation processing means" refers to a device or system that operates on a cloud infrastructure and performs real-time analysis of collected information.

[0604] "Information distribution means" refers to a device or system for notifying external parties based on analysis results.

[0605] An "information processing system" is a device or method for processing information and performing specific analyses or evaluations.

[0606] An "information management system" is a device or system for managing acquired information and monitoring the status of users.

[0607] The system implementing this invention includes a terminal installed in the user's living space and a server operating on a cloud infrastructure. This section provides a detailed explanation of information gathering, information analysis, and notification.

[0608] The device is equipped with a camera and microphone, and collects image and audio information on a daily basis. This allows for the rapid detection of any anomalies occurring in the user's living space. The device transmits the collected data to a server via the internet.

[0609] The servers are located on a cloud infrastructure and analyze the collected data in real time. TensorFlow and other machine learning frameworks are used for data analysis. This analysis method creates a lifestyle pattern model and detects anomalies when users exhibit behavior that deviates from their daily routine. In addition, a generative AI model is used to evaluate emotional states from voice information and detect specific emotional changes.

[0610] Based on the analysis results, the server uses speech synthesis technology to deliver messages to users via their devices, providing voice warnings and advice. Furthermore, information regarding anomalies and emotional changes is sent to external parties via email or SMS using the Twilio API.

[0611] For example, if a user fails to take their usual afternoon walk, the server will generate a message saying, "It appears you haven't completed your morning exercise. If you are feeling unwell, we recommend taking a rest," and notify the user. It will also send a message to family members saying, "The user hasn't done their morning exercise yet. Please check on their health."

[0612] Examples of prompt statements for a generative AI model are as follows:

[0613] "Please create an algorithm that analyzes lifestyle pattern data and emotional voice data of elderly individuals to detect abnormalities."

[0614] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0615] Step 1:

[0616] The device collects image and audio information of the user in their living space using a camera and microphone. The input is raw data obtained from the camera and microphone, and the output is this raw data converted directly into a digital format. Specifically, it records images captured by the camera and audio recorded by the microphone in real time.

[0617] Step 2:

[0618] The device transmits collected image and audio data to a server via the internet. The input is processed digital data, and the output is data upload to a cloud server. Specifically, it reliably transmits data to the cloud using a communication module.

[0619] Step 3:

[0620] The server analyzes received images and audio and detects anomalies based on a lifestyle pattern model. The input is digital data sent from the terminal, and the output is the result of detecting anomalies in the lifestyle pattern. Specifically, it uses TensorFlow to perform pattern analysis with a pre-trained model and identify anomalies.

[0621] Step 4:

[0622] The server utilizes a generative AI model to analyze the user's emotional state from voice data. The input is voice data, and the output is the evaluation result of the emotional state. Specifically, it identifies the user's emotions through the voice data and performs voice analysis to capture emotional changes such as stress and anxiety.

[0623] Step 5:

[0624] The server uses speech synthesis technology to generate messages to be delivered to the user via the terminal, based on the results of anomaly detection and sentiment analysis. The input is the analysis results, and the output is the voice message to the user. Specifically, it generates appropriate advice and warning messages and delivers them to the user via the speaker.

[0625] Step 6:

[0626] The server uses the Twilio API to notify external stakeholders of anomalies and emotional changes obtained through analysis. Input is detailed information about anomalies and emotional changes, and output is the result of sending emails or SMS messages to stakeholders. The specific operation involves generating notification content and delivering the information quickly and reliably to stakeholders.

[0627] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0628] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0629] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0630] [Fourth Embodiment]

[0631] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0632] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0633] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0634] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0635] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0636] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0637] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0638] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0639] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0640] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0641] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0642] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0643] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0644] This invention relates to a system used in the living environment of elderly people, which monitors their lifestyle patterns using various sensors and devices and detects abnormalities. Specifically, this system consists of a terminal installed in the living space and a server that processes data in the background.

[0645] terminal

[0646] The device is placed in the user's living environment and is equipped with a camera and microphone to collect the user's actions and conversations in daily life. The device also works in conjunction with wearable devices to acquire biometric information (e.g., steps taken, heart rate, etc.). The collected data is transmitted to a server via a secure communication protocol.

[0647] server

[0648] The server analyzes the received image and audio data to model the typical daily routines of elderly individuals. The modeled data is compared with historical data to monitor for any anomalies. If an anomaly is detected, immediate action is taken, and the user is alerted via voice.

[0649] Next, the analyzed audio data is evaluated using a generative model to automatically determine the risk of dementia. This evaluation result is then used as valuable information for the user and their family and friends.

[0650] User

[0651] When an anomaly is detected, the user will receive appropriate verbal notifications from their device. For example, a message such as, "You woke up at [time], is everything alright?" The server will also send details of the anomaly to family members and care managers via email or app notifications.

[0652] This system eliminates the need for specialized care staff and automatically monitors the current situation of elderly individuals, prompting them to take necessary actions. For example, even during times when care services are unavailable, it can detect signs of nighttime wandering and respond quickly. Furthermore, if the risk of dementia is high, early diagnosis and treatment are encouraged, leading to more efficient risk management.

[0653] The following describes the processing flow.

[0654] Step 1:

[0655] The device acquires images and audio data from the living space through its camera and microphone. It also collects biometric information such as heart rate and steps taken from wearable devices.

[0656] Step 2:

[0657] The terminal converts the acquired data into a predetermined format and sends it to the server using a secure communication protocol.

[0658] Step 3:

[0659] The server analyzes the received image data using an image recognition algorithm and automatically classifies the user's behavior. This allows for the modeling of daily activity patterns.

[0660] Step 4:

[0661] The server converts the audio data into text using speech recognition technology, analyzes the conversation content, and assesses the risk of dementia.

[0662] Step 5:

[0663] The server models normal lifestyle patterns based on collected lifestyle data and detects anomalies by comparing them with the latest data in real time.

[0664] Step 6:

[0665] When an anomaly is detected, the server determines the appropriate course of action and generates an appropriate message.

[0666] Step 7:

[0667] The terminal provides warnings and advice to the user via voice messages, based on instructions from the server.

[0668] Step 8:

[0669] The server compiles the details of the detected anomalies and sends notifications to family members and care managers.

[0670] Step 9:

[0671] The server generates periodic reports at regular intervals and sends comprehensive analysis results regarding the user's lifestyle and health status to relevant parties.

[0672] (Example 1)

[0673] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0674] In the living environment of the elderly, it is important to monitor their health while maintaining their independence. However, it is not realistic to have professional care staff present at all times, so a system that can quickly detect and respond to abnormalities is needed. Furthermore, it is important to detect and address cognitive decline early. To solve these problems, a system is needed that can acquire and analyze data in an efficient and reliable manner.

[0675] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0676] In this invention, the server includes information acquisition means for acquiring behavior and conversation; analysis means for analyzing the acquired behavior and conversation data to generate a lifestyle model and detect anomalies; and notification means for providing voice warnings to the user when an anomaly is detected. This enables rapid detection and response to anomalies in the living environment of elderly people.

[0677] "Information acquisition means" refers to devices or technologies for sensing the behavior and conversations of elderly people and acquiring them as data.

[0678] "Analysis method" refers to the process of analyzing the lifestyle of elderly people based on acquired data and detecting deviations from normal patterns.

[0679] A "notification mechanism" is a system element that has the function of communicating a warning to the user via voice when an anomaly is detected.

[0680] "Information transmission means" refers to a mechanism for communicating details of detected anomalies to other relevant parties, and this is done via email, app notifications, etc.

[0681] A "generative model" is a computational model used to generate assessments of cognitive function in older adults from data.

[0682] A "portable device" is a device used to acquire a user's biometric information in real time, and it generally takes the form of something that can be worn on the body.

[0683] "Biometric information" refers to data that indicates a user's health status, including physical indicators such as heart rate and step count.

[0684] This invention is a system for rapidly detecting and responding to abnormalities in the living environment of elderly people. This system consists of a terminal for acquiring information, a server for analyzing the data, and a user interface for providing notifications.

[0685] The terminals are installed in the living environments of elderly people and are equipped with cameras and microphones to capture their actions and conversations. Additionally, a portable device worn by the user is used to acquire biometric information such as heart rate and step count in real time. This information is transmitted to a server via a secure protocol.

[0686] Specifically, the device can detect user movement and falls, for example, using a fixed camera in the room. If a user falls, the device immediately notifies the server so that appropriate action can be taken.

[0687] The server analyzes data transmitted from the terminal and models the lifestyle of the elderly. Based on the acquired video and audio data, a generative AI model performs predictive analysis and detects anomalies. Furthermore, it evaluates the risk of cognitive function from the audio data and saves the results. This evaluation is based on pre-configured prompts.

[0688] A concrete example of a prompt message is, "Based on the user's latest voice and biometric data, evaluate whether any anomalies have been detected and report the results of the cognitive function risk analysis." This generates an appropriate alert that responds to the situation.

[0689] The user is in a position to receive information provided by the system. If an anomaly is detected, an audio notification will be sent from the device. For example, a specific alert such as, "A value higher than your normal resting heart rate has been detected. We recommend that you take a break," will be conveyed to the user. Furthermore, the anomaly information is also sent to external parties and communicated to family members and care managers via email or application notifications.

[0690] This system enables quick and effective responses to ensure the safety of elderly individuals, even in the absence of specialized care staff. By utilizing biometric information, it allows for detailed management tailored to each individual's health condition.

[0691] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0692] Step 1:

[0693] The device uses a camera and microphone installed in the elderly person's living space to capture their actions and conversations. The input is real-time video and audio data. Specifically, the device continuously captures video frames and records audio signals. This data is used as foundational data for anomaly detection. The output is the captured video and audio data.

[0694] Step 2:

[0695] The terminal acquires biometric data from a portable device worn by the user. The input is biometric information such as heart rate and step count, which are updated in real time. Specifically, the terminal periodically receives data from the portable device using communication functions such as Bluetooth or Wi-Fi. The output is the collected biometric data.

[0696] Step 3:

[0697] The terminal transmits collected video, audio, and biometric data to the server via a secure protocol. The input is all data acquired by the terminal. Specifically, the terminal compresses and encrypts the data before transmitting it over the network. The output is the data transferred to the server.

[0698] Step 4:

[0699] The server receives the transmitted data and begins analysis. The input consists of video, audio, and biometric data sent from the terminal. The server first analyzes the video data to extract behavioral features. Next, it converts the audio data into text and analyzes the content of the conversation using natural language processing techniques. Furthermore, it analyzes the biometric data to detect deviations from normal patterns. The output is the analysis result.

[0700] Step 5:

[0701] The server uses a generative AI model to assess cognitive function risk from voice data. The input is pre-analyzed voice data. Specifically, the server inputs prompt sentences into the generative AI model and retrieves the corresponding risk analysis results. The output is the cognitive function risk assessment result.

[0702] Step 6:

[0703] The server detects anomalies based on the analysis results and notifies the user as necessary. The input consists of the analysis results and risk assessment results generated by the server. Specifically, the server determines the anomaly, generates a warning message using speech synthesis technology, and sends it to the terminal. The output is an audio warning to the user.

[0704] Step 7:

[0705] The server sends detailed information about the anomaly to external parties. The input is all information about the detected anomaly. Specifically, the server notifies parties via email or application. The output is the alert notification sent to the parties.

[0706] (Application Example 1)

[0707] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0708] The goal is to solve the problem of difficulty in quickly detecting abnormal movements or inappropriate postures while ensuring worker safety in work environments such as factories.

[0709] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0710] In this invention, the server includes a data acquisition device, a processing device that analyzes the acquired data to generate behavioral patterns and detect anomalies, a notification device that issues an audible warning to the user when an anomaly is detected, a communication device that transmits details of the anomaly to external parties, a device that learns the worker's movements, and a device that monitors the movements in the work environment. This improves the safety of workers in the work environment and enables rapid response to anomalies.

[0711] A "data acquisition device" is a device that collects data such as images and sounds from the surrounding environment.

[0712] A "processing device" is a device that analyzes acquired data, generates behavioral patterns, and has the function of detecting anomalies.

[0713] A "notification device" is a device that alerts the user via voice when an abnormality is detected.

[0714] A "communication device" is a device used to transmit details of detected anomalies to external parties.

[0715] A "device that learns worker movements" is a device that recognizes the movements of workers in a work environment such as a factory and learns those patterns.

[0716] A "device for monitoring actions" is a device for monitoring the actions of workers in the work environment in real time.

[0717] The system of this invention includes a device for ensuring and improving worker safety in a work environment such as a factory. The server, as a data acquisition device, collects image and audio data from the work environment using a camera and microphone. The processing device analyzes the collected data using image processing libraries such as OpenCV and machine learning frameworks such as TensorFlow to generate worker behavior patterns and detects anomalies by comparing these patterns with actual behavior.

[0718] If an anomaly is detected, the notification device will issue an audio warning to the worker. For example, a message such as, "That posture is dangerous. Take a break," will be issued. Furthermore, details of the anomaly are sent to an external administrator via a communication device, allowing the administrator to respond quickly.

[0719] Furthermore, the device that learns the worker's movements analyzes past data to learn work patterns and promote appropriate work habits. This system, as a movement monitoring device, performs real-time work monitoring and improves safety on site.

[0720] The following is an example of a prompt message for a specific generative AI model.

[0721] "Use image data to analyze worker behavior patterns and design a machine learning model to detect anomalies. Propose how to issue warnings and notify managers when an anomaly is detected."

[0722] This system makes it possible to improve work efficiency and safety within the factory.

[0723] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0724] Step 1:

[0725] The terminal uses a camera and microphone to acquire image and audio data from the factory work environment. This input data includes worker actions and conversations. This data is transmitted to the server in real time.

[0726] Step 2:

[0727] The server processes the received image data using OpenCV to analyze the worker's posture and movements. Specifically, it detects human movement within the image and compares it to normal movement patterns. This process makes it possible to detect abnormal movements based on data from normal operation.

[0728] Step 3:

[0729] The server analyzes audio data using TensorFlow, converting the workers' conversations into text data. Then, it performs feature extraction and pattern analysis. This process enables the detection of abnormal conversation content and unusual sounds.

[0730] Step 4:

[0731] If the server's processing unit detects an anomaly based on the analysis results, it will issue an audible warning to the worker via a notification device. For example, if a worker is in a dangerous position, it will warn, "This is a dangerous position. Please be careful." This output plays a role in enhancing worker safety.

[0732] Step 5:

[0733] The server transmits details of detected anomalies to an external administrator via a communication device. This allows the administrator to take prompt action. This includes a timestamp of the anomaly and environmental data.

[0734] Step 6:

[0735] The server continuously learns work patterns and updates this information using a generative AI model. An example prompt might be, "Use image data to analyze worker behavior patterns and design a machine learning model to detect anomalies." This process ensures optimal anomaly detection at all times.

[0736] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0737] This invention relates to a system for use in the living environments of elderly people, which utilizes various sensors and devices to monitor and analyze lifestyle patterns and emotional states, and to provide appropriate responses. In particular, this system has the function of collecting image and audio data and simultaneously analyzing abnormal lifestyle patterns and emotional changes.

[0738] terminal

[0739] The terminal is installed in the user's living space and has the function of acquiring image and audio data through a camera and microphone. Furthermore, it can communicate with wearable devices and acquire biometric information. The acquired data is sequentially sent to a server for analysis in real time.

[0740] server

[0741] The server first analyzes the received data based on a lifestyle pattern model. This allows it to detect abnormalities if the data deviates from the normal lifestyle pattern. It also uses a generative model to analyze voice data and assess the risk of dementia.

[0742] Furthermore, this invention uses an emotion engine to recognize emotions from the user's voice and facial expression data. For example, if the user is showing signs of stress or anxiety, this engine captures that state and uses it as information for making appropriate decisions. The server analyzes the changes in emotion, generates a customized message based on that analysis, and provides it to the user.

[0743] User

[0744] Users receive notifications via voice messages, and advice based on abnormalities and emotional changes is provided through their devices. For example, if a user is clearly feeling anxious, the system can say, "Are you okay? We'll send help if you need it."

[0745] If emotional changes are deemed persistent or abnormal, the server notifies family members or care managers of the details via email or a dedicated app. This allows for prompt action and enhanced mental and physical health support.

[0746] This system allows for the automatic monitoring of elderly individuals' emotional states and living conditions, and the provision of appropriate support, without the need for constant supervision by specialized care staff. This enables more comprehensive and precise monitoring compared to conventional methods.

[0747] The following describes the processing flow.

[0748] Step 1:

[0749] The device uses cameras and microphones installed within the living space to collect the user's image and voice data in real time. It also acquires biometric information from wearable devices.

[0750] Step 2:

[0751] The device transmits collected images, audio, and biometric information to the server via a predetermined communication protocol. The data is encrypted during transmission, ensuring security.

[0752] Step 3:

[0753] The server converts the received audio data into text using speech recognition technology and analyzes the user's conversation. The analyzed data is then used with a generative model to assess the risk of dementia.

[0754] Step 4:

[0755] The server analyzes image data to identify the user's behavior. Based on the analysis results, it detects an anomaly if the user's lifestyle pattern deviates from the norm.

[0756] Step 5:

[0757] The server uses an emotion engine to analyze voice and facial expression data and evaluate the user's emotional state. It recognizes changes in emotion and signs of anxiety and stress.

[0758] Step 6:

[0759] When the server detects anomalies or changes in the user's emotional state, it generates a customized message. This message includes advice and warnings based on both the user's life and emotions.

[0760] Step 7:

[0761] The device, based on instructions from the server, delivers voice messages to the user through its speaker. For example, it might convey information about unusual circumstances in daily life or emotional anxieties.

[0762] Step 8:

[0763] The server notifies family members or care managers via email or app of details about any anomalies or emotional changes. This allows for a quick response if necessary.

[0764] Step 9:

[0765] The server periodically aggregates the analysis results and generates a report. This report provides important information about users' lifestyle patterns and emotional states and is sent to relevant parties on a regular basis.

[0766] (Example 2)

[0767] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0768] There is a growing need to monitor the living environment of the elderly and appropriately assess their emotions and health status in order to provide prompt and accurate support. However, achieving this requires a system that integrates data from multiple sources and analyzes it efficiently. Conventional methods have made it difficult to quickly capture changes in behavior and emotions, and appropriate responses have sometimes been delayed even when abnormalities occur. This invention aims to solve these problems.

[0769] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0770] In this invention, the server includes information gathering means for collecting images and sounds, analysis means for analyzing the collected image and sound information to create behavioral patterns and detect anomalies, and emotion recognition means for analyzing emotion information to evaluate the user's emotional state. This makes it possible to quickly detect anomalies and changes in emotions in the user's living environment and provide appropriate responses.

[0771] "Information gathering means" refers to devices and technologies for acquiring image and audio information, and has the function of collecting data in the user's living environment.

[0772] "Analysis means" refers to processes and technologies for analyzing collected image and audio information to identify user behavior patterns and anomalies.

[0773] "Emotion recognition means" refers to technologies and algorithms that evaluate a user's emotional state based on voice and facial expression data.

[0774] A "guidance system" is a function that notifies the user of detected anomalies or emotional states and directly conveys messages through voice or other means.

[0775] "Means of communication" refers to communication technologies and infrastructure used to transmit information about abnormal or important emotional states to external parties.

[0776] "Generative technology" refers to methods for performing specific risks and assessments using data generation models, which are used in the analysis of audio data and other similar materials.

[0777] A "wearable device" refers to a device that is attached to a user's body to acquire biometric data.

[0778] This invention is a system for monitoring the living environment of elderly people and evaluating abnormalities and emotional states. The main components necessary for carrying out the invention include a terminal for collecting information, a server for analyzing the data, and means for notifying the user.

[0779] The terminal is installed in the living space and uses a camera and microphone to collect the user's image and voice data in real time. This terminal can connect with wearable devices as needed and transmit information, including biometric data, to a server. For example, if the user wears a smartwatch, additional data such as heart rate and body temperature can be collected.

[0780] The server aggregates and analyzes data sent from the terminal. For image and audio data, advanced analytical algorithms are used to identify behavioral patterns and detect anomalies. Furthermore, a generative AI model is used to analyze audio data and provide prompts to assess the user's risk of cognitive decline. For example, by inputting "Assess the likelihood of cognitive decline based on the user's voice characteristics," the model provides an output.

[0781] Furthermore, the server utilizes an emotion engine to evaluate the user's emotional state. Based on this, if emotional changes such as stress or anxiety are identified, corrective measures are taken.

[0782] Users receive feedback on their health and emotional state through notifications provided by the system. For example, if a detected anomaly is related to everyday stress, the user may receive a voice notification such as, "You need to relax. Take a break," to facilitate appropriate support.

[0783] This system will create an environment where elderly people can live more independently and with greater peace of mind, and will also enable efficient and effective care support through the rapid provision of information to external stakeholders.

[0784] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0785] Step 1:

[0786] The device uses a camera and microphone to collect user image and audio data in real time. Specifically, the camera captures multiple frames of images per second, and the microphone continuously records ambient sounds. The device compresses this data and sends it to the server. The input is images and audio from the living space, and the output is a compressed data file.

[0787] Step 2:

[0788] The terminal acquires the user's biometric data by communicating with a wearable device. This data includes information such as heart rate and body temperature. The terminal receives data from the device via wireless communication such as Bluetooth and transmits it to the server. The input is biometric data from the wearable device, and the output is the data transmitted to the server.

[0789] Step 3:

[0790] The server receives image and audio data transmitted from the terminal and analyzes the behavioral patterns. The server compares this data with historical data and uses a statistical model to identify deviations from normal patterns. The input is image and audio data files, and the output is information about the detected anomalies.

[0791] Step 4:

[0792] The server analyzes voice data using a generative AI model to assess the risk of cognitive decline. It inputs a prompt sentence into the model and calculates a risk score. The input consists of voice data and the prompt sentence "Assess the likelihood of cognitive decline based on the user's voice characteristics," and the output is the generated risk assessment score.

[0793] Step 5:

[0794] The server uses an emotion engine to identify emotional states from voice and facial expression data. Data analysis identifies emotions such as stress or anxiety and determines how to respond to the user. The input is voice and facial expression data, and the output is the evaluation result of the emotional state.

[0795] Step 6:

[0796] The user receives an audio notification from their device. A response message generated by the server is sent to the device and provided to the user as an audio message. Specifically, advice such as "You need to relax. Please take a break." is delivered via audio. The input is the response message from the server, and the output is the audio notification to the user.

[0797] Step 7:

[0798] The server will send information to external parties if emotions or anomalies persist. Details of the situation and recommended actions will be communicated to relevant parties via email or a dedicated app. Input is the detected anomaly or emotional state, and output is the notification of information to external parties.

[0799] (Application Example 2)

[0800] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0801] In environments where elderly people live alone, there is a growing need for effective monitoring systems that can quickly and accurately detect everyday abnormalities and emotional changes, and take appropriate measures. Conventional methods have challenges in that it is difficult to quickly detect abnormalities and there is insufficient notification to family members or caregivers.

[0802] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes an information collection means for collecting images and sounds, a data analysis means for analyzing the collected image and sound information to create a lifestyle pattern model and detect abnormalities, and a notification means for conveying a voice warning to the user when an abnormality is detected. This makes it possible to immediately detect abnormalities in daily life and changes in emotions in the living space of elderly people and to quickly and accurately notify relevant parties.

[0803] "Information gathering means" refers to a device or system for acquiring image and audio information from an object.

[0804] "Data analysis means" refers to a device or method for processing collected image and audio information, creating a lifestyle pattern model, and detecting anomalies within it.

[0805] "Notification means" refers to a device or system for conveying warnings or notifications to users via voice.

[0806] "Information transmission means" refers to a device or system for transmitting collected and analyzed information to external parties.

[0807] "Computation processing means" refers to a device or system that operates on a cloud infrastructure and performs real-time analysis of collected information.

[0808] "Information distribution means" refers to a device or system for notifying external parties based on analysis results.

[0809] An "information processing system" is a device or method for processing information and performing specific analyses or evaluations.

[0810] An "information management system" is a device or system for managing acquired information and monitoring the status of users.

[0811] The system implementing this invention includes a terminal installed in the user's living space and a server operating on a cloud infrastructure. This section provides a detailed explanation of information gathering, information analysis, and notification.

[0812] The device is equipped with a camera and microphone, and collects image and audio information on a daily basis. This allows for the rapid detection of any anomalies occurring in the user's living space. The device transmits the collected data to a server via the internet.

[0813] The servers are located on a cloud infrastructure and analyze the collected data in real time. TensorFlow and other machine learning frameworks are used for data analysis. This analysis method creates a lifestyle pattern model and detects anomalies when users exhibit behavior that deviates from their daily routine. In addition, a generative AI model is used to evaluate emotional states from voice information and detect specific emotional changes.

[0814] Based on the analysis results, the server uses speech synthesis technology to deliver messages to users via their devices, providing voice warnings and advice. Furthermore, information regarding anomalies and emotional changes is sent to external parties via email or SMS using the Twilio API.

[0815] For example, if a user fails to take their usual afternoon walk, the server will generate a message saying, "It appears you haven't completed your morning exercise. If you are feeling unwell, we recommend taking a rest," and notify the user. It will also send a message to family members saying, "The user hasn't done their morning exercise yet. Please check on their health."

[0816] Examples of prompt statements for a generative AI model are as follows:

[0817] "Please create an algorithm that analyzes lifestyle pattern data and emotional voice data of elderly individuals to detect abnormalities."

[0818] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0819] Step 1:

[0820] The device collects image and audio information of the user in their living space using a camera and microphone. The input is raw data obtained from the camera and microphone, and the output is this raw data converted directly into a digital format. Specifically, it records images captured by the camera and audio recorded by the microphone in real time.

[0821] Step 2:

[0822] The device transmits collected image and audio data to a server via the internet. The input is processed digital data, and the output is data upload to a cloud server. Specifically, it reliably transmits data to the cloud using a communication module.

[0823] Step 3:

[0824] The server analyzes received images and audio and detects anomalies based on a lifestyle pattern model. The input is digital data sent from the terminal, and the output is the result of detecting anomalies in the lifestyle pattern. Specifically, it uses TensorFlow to perform pattern analysis with a pre-trained model and identify anomalies.

[0825] Step 4:

[0826] The server utilizes a generative AI model to analyze the user's emotional state from voice data. The input is voice data, and the output is the evaluation result of the emotional state. Specifically, it identifies the user's emotions through the voice data and performs voice analysis to capture emotional changes such as stress and anxiety.

[0827] Step 5:

[0828] The server uses speech synthesis technology to generate messages to be delivered to the user via the terminal, based on the results of anomaly detection and sentiment analysis. The input is the analysis results, and the output is the voice message to the user. Specifically, it generates appropriate advice and warning messages and delivers them to the user via the speaker.

[0829] Step 6:

[0830] The server uses the Twilio API to notify external stakeholders of anomalies and emotional changes obtained through analysis. Input is detailed information about anomalies and emotional changes, and output is the result of sending emails or SMS messages to stakeholders. The specific operation involves generating notification content and delivering the information quickly and reliably to stakeholders.

[0831] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0832] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0833] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0834] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0835] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0836] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0837] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0838] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0839] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0840] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0841] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0842] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0843] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0844] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0845] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0846] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0847] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0848] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0849] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0850] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0851] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0852] The following is further disclosed regarding the embodiments described above.

[0853] (Claim 1)

[0854] A data collection means for collecting images and audio,

[0855] An analysis means for creating a lifestyle pattern model by analyzing collected image and audio data and detecting anomalies,

[0856] A notification system that alerts the user via voice when an anomaly is detected,

[0857] A means of communication to transmit details of the anomaly to external parties,

[0858] A system that includes this.

[0859] (Claim 2)

[0860] The system according to claim 1, comprising a system that uses a generative model to determine the user's risk of dementia from voice data.

[0861] (Claim 3)

[0862] The system according to claim 1, comprising a system for monitoring a user's health status using biometric information acquired by a wearable device.

[0863] "Example 1"

[0864] (Claim 1)

[0865] Information acquisition means for obtaining behavior and conversation,

[0866] An analytical means for analyzing acquired behavioral and conversational data to generate a lifestyle model and detect anomalies,

[0867] A notification system that provides an audio warning to the user when an anomaly is detected,

[0868] A means of transmitting information about the anomaly to other relevant parties,

[0869] A system that includes this.

[0870] (Claim 2)

[0871] The system according to claim 1, comprising a system that uses a generative model to assess the risk of a user's cognitive function from conversational data.

[0872] (Claim 3)

[0873] The system according to claim 1, comprising a system for monitoring the user's health status using biometric information acquired by a portable device.

[0874] "Application Example 1"

[0875] (Claim 1)

[0876] A device for acquiring data,

[0877] A processing device that analyzes acquired data to generate behavioral patterns and detects anomalies,

[0878] A notification device that provides an audible warning to the user if an abnormality is detected,

[0879] A communication device that transmits details of the anomaly to external parties,

[0880] A device that learns the movements of workers,

[0881] A device for monitoring operations in the work environment,

[0882] A system that includes this.

[0883] (Claim 2)

[0884] The system according to claim 1, comprising a device that uses a generative model to evaluate the user's risk of dementia from voice information.

[0885] (Claim 3)

[0886] The system according to claim 1, comprising a device that monitors the user's health status using biometric data acquired by a portable device.

[0887] "Example 2 of combining an emotion engine"

[0888] (Claim 1)

[0889] Information gathering means for collecting images and audio,

[0890] An analysis means for creating behavioral patterns by analyzing collected image and audio information and for detecting anomalies,

[0891] An emotion recognition means that analyzes emotional information to evaluate the user's emotional state,

[0892] A guidance system that notifies the user by voice when an abnormal or specific emotional state is detected,

[0893] A means of communication for transmitting details of abnormal or noteworthy emotional states to external parties,

[0894] A system that includes this.

[0895] (Claim 2)

[0896] The system according to claim 1, which uses generation technology to evaluate the risk of cognitive decline in a user from voice data.

[0897] (Claim 3)

[0898] The system according to claim 1, which monitors the user's health status using biometric data acquired by a wearable device.

[0899] "Application example 2 when combining with an emotional engine"

[0900] (Claim 1)

[0901] Information gathering means for collecting images and audio,

[0902] A data analysis means that analyzes collected image and audio information to create a lifestyle pattern model and detect anomalies,

[0903] A notification system that alerts users with voice when an anomaly is detected,

[0904] A means of transmitting information about the anomaly to external parties,

[0905] A computational processing means that performs real-time analysis using a model that runs on a cloud platform,

[0906] An information distribution means for notifying external stakeholders based on the analysis results,

[0907] A system that includes this.

[0908] (Claim 2)

[0909] The system according to claim 1, comprising an information processing system that uses a generative model to evaluate the risk of dementia of a user from voice information.

[0910] (Claim 3)

[0911] The system according to claim 1, comprising an information management system that monitors the user's health status using biometric information acquired by a wearable device. [Explanation of Symbols]

[0912] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A data collection means for collecting images and audio, An analysis means for creating a lifestyle pattern model by analyzing collected image and audio data and detecting anomalies, A notification system that alerts the user via voice when an anomaly is detected, A means of communication to transmit details of the anomaly to external parties, A system that includes this.

2. The system according to claim 1, comprising a system that uses a generative model to determine the user's risk of dementia from voice data.

3. The system according to claim 1, comprising a system for monitoring a user's health status using biometric information acquired by a wearable device.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A