System

A system integrating home sensors with multimodal AI for real-time health monitoring addresses the challenge of inaccessible health checks by detecting abnormalities through diverse data analysis.

JP2026022397APending Publication Date: 2026-02-12SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024123914
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-30
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

In modern society, there are environments where regular health checks are difficult to conduct, and medical services are hard to access, making it challenging to detect and prevent diseases early, especially in home settings where health changes may go unnoticed.

Method used

A system that integrates data from multiple sensors in the home, including image, voice, text, and biometric sensors, to collect and analyze data using multimodal artificial intelligence, and sends notifications to users when abnormalities are detected.

Benefits of technology

Enables real-time health monitoring at home, allowing for early detection of health risks and abnormalities by comprehensively analyzing various data types, including emotional and cognitive changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026022397000001_ABST
    Figure 2026022397000001_ABST
Patent Text Reader

Abstract

To provide a system capable of easily monitoring a health condition at home and coping with abnormality in an early stage when the abnormality occurs.SOLUTION: The specific processing unit 290 of the data processing device 12 in the system is collected by a sensor used in the home. The transmitted data are received and stored. Then, the stored data is analyzed using multimodal artificial intelligence, the analysis result is integrated to monitor the health condition, and a notification is generated and transmitted to the user when an abnormality is detected based on the monitoring result.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In modern society, there are environments where it is difficult to conduct regular health checks and areas where medical services are difficult to access, which makes it difficult to adequately detect and prevent diseases early. Furthermore, in everyday life at home, people may not notice changes in their health status and miss serious health risks. There is a need for a system that can solve these issues, easily monitor health status at home, and respond quickly when abnormalities occur. [Means for solving the problem]

[0005] The present invention provides a system for integrated analysis of data from multiple sensors used in the home. Specifically, the system includes: (1) means for collecting data from sensors in the home, (2) means for transmitting the collected data to a server, (3) means for receiving and storing the transmitted data, (4) means for analyzing the stored data using multimodal artificial intelligence, (5) means for integrating the analysis results to monitor health status, and (6) means for generating and sending a notification to the user when an abnormality is detected.

[0006] This makes it easier to conduct regular health checks at home, and by analyzing various information in real time, including image data, video data, audio data, text data, and biometric data, it is possible to identify and respond to potential health risks early.In addition, by analyzing audio data, it is possible to detect changes in tone and speaking style and identify signs related to emotional problems and changes in cognitive function, allowing for comprehensive monitoring of the user's health status.

[0007] A "sensor" is a device that detects a physical or chemical change and outputs it as a signal.

[0008] "Data collection means" refers to devices and methods for obtaining data from sensors.

[0009] "Data transmission means" refers to the device and method for transmitting collected data to a unique receiving device.

[0010] A "server" is a computing device that receives, stores, and analyzes data.

[0011] "Data receiving means" refers to a device and method for receiving transmitted data.

[0012] "Data storage means" refers to a device and method for storing received data in a storage device.

[0013] "Multimodal artificial intelligence (AI)" is an artificial intelligence technology that has the ability to comprehensively analyze multiple different data formats.

[0014] "Data analysis means" refers to devices and methods for analyzing stored data using specific algorithms.

[0015] "Health monitoring tool" refers to devices and methods that integrate the results of data analysis to continuously monitor the health status of an individual.

[0016] "Notification generation means" refers to a device and method for generating a warning or information when an abnormality is detected.

[0017] "User" refers to an individual or group of users of this System. [Brief explanation of the drawings]

[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0020] First, the terms used in the following description will be explained.

[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0026] [First embodiment]

[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0039] The present invention is a system for monitoring the health condition of a user by comprehensively analyzing data from multiple sensors used in the home. A specific embodiment of the system will be described.

[0040] System Overview

[0041] The system collects data from multiple sensors (image sensors, voice sensors, text input devices, and biometric sensors) installed in the user's home and sends it to a server. The server stores the received data, analyzes it using multimodal artificial intelligence (AI), and monitors the user's health status. If an abnormality is detected, the server sends a notification to the user.

[0042] Program processing flow

[0043] Data collection

[0044] Devices (users' home sensors and wearable devices) collect data

[0045] Image data: Security cameras in the home periodically take images and videos, recording the user's movements, facial expressions, and daily life.

[0046] Voice data: Smart speakers and smart assistants collect your voice and use it to detect changes in your tone and speaking style.

[0047] Text data: Text data entered by users through smart assistants or dedicated applications is collected and analyzed for text content and signs of stress.

[0048] Biometric data: Wearable devices measure biometric data such as heart rate, blood pressure, oxygen saturation, and sleep patterns, thereby continuously monitoring the user's physical condition.

[0049] Data transmission

[0050] The device sends the collected data to the server.

[0051] The sensor terminals send the collected data to the server at regular intervals, and the data is encrypted using a security protocol.

[0052] Data analysis

[0053] The server analyzes the received data

[0054] The server uses the image data to analyze changes in the user's movements and facial expressions. For example, unusual movements may indicate a health problem.

[0055] Voice data is analyzed using speech recognition technology to detect changes in tone of voice and speech patterns, which can identify emotional issues and changes in cognitive function.

[0056] The text data is analyzed using natural language processing technology to detect changes in the content and expressions of the text, which can then be used to assess stress and mental health.

[0057] Biometric data is analyzed over time to detect abnormalities in heart rate, blood pressure, and sleep patterns, allowing for early detection of cardiac and respiratory problems.

[0058] Health Status Monitoring

[0059] The server consolidates the analysis results and monitors health status.

[0060] By integrating the results of various data analyses, multimodal AI assesses the user's overall health status, making it easier to detect abnormalities.

[0061] Send notifications

[0062] If the server detects an abnormality, it will notify the user.

[0063] If the server detects an anomaly, it generates a notification message and sends it to the user's smartphone or smart assistant, containing details of the anomaly and recommended actions to take.

[0064] Example: Detecting abnormal heart rates at night

[0065] The terminal (wearable device) continues to measure the user's heart rate throughout the night. The server receives and stores the collected heart rate data. The server then analyzes the received data and detects abnormal increases in heart rate beyond the normal range. If an abnormality is detected, the server sends a notification to the user's smartphone stating, "A sudden increase in heart rate has been detected overnight." The user receives this notification and can take prompt action, such as consulting a medical institution.

[0066] In this way, the present invention can protect the user's health by monitoring the user's health status in real time at home and identifying health risks early.

[0067] The processing flow will be explained below.

[0068] Step 1:

[0069] Devices (sensors) collect data. Security cameras installed in homes periodically capture images and videos, and smart speakers and smart assistants record voice data. Text data entered by users through smart assistants and applications is also collected. In addition, wearable devices measure biometric data such as heart rate, blood pressure, oxygen saturation, and sleep patterns in real time.

[0070] Step 2:

[0071] The terminal (sensor) packetizes and encrypts the collected data, and the data packets are temporarily stored in a queue inside the terminal.

[0072] Step 3:

[0073] The device (sensor) sends data packets to the server at regular intervals. The data transmission is performed using a secure communication protocol to ensure the confidentiality and security of the data.

[0074] Step 4:

[0075] The server receives the data packet. The received data packet is decoded and restored to its original format. The data integrity is checked, and if there are any errors or omissions, a retransmission request is sent to the terminal.

[0076] Step 5:

[0077] The server stores the recovered data in a distributed database, where the data is managed for each user and tagged within the database.

[0078] Step 6:

[0079] The server analyzes the received image and video data, using image classification and motion analysis algorithms to detect changes in the user's movement patterns and facial expressions.

[0080] Step 7:

[0081] The server analyzes the audio data, converts it into text using speech recognition technology, and analyzes changes in tone and speaking style to identify signs associated with emotional problems and changes in cognitive function.

[0082] Step 8:

[0083] The server analyzes the text data and uses natural language processing technology to analyze the content of the text data and evaluate stress levels and health conditions.

[0084] Step 9:

[0085] The server analyzes biometric data, using time series analysis techniques to analyze data such as heart rate, blood pressure, oxygen saturation, and sleep patterns to detect abnormalities, especially values ​​that exceed the normal range.

[0086] Step 10:

[0087] The server integrates the results of each analysis, using multimodal AI to comprehensively analyze images, voice, text, and biometric data to assess the user's overall health.

[0088] Step 11:

[0089] The server generates a notification when an anomaly is detected, containing details about the anomaly and a recommended course of action.

[0090] Step 12:

[0091] The server generates a notification and sends it to the user's smartphone or smart assistant. The notification is sent immediately, informing the user that an abnormality has occurred.

[0092] Step 13:

[0093] Users will receive a notification, which they can review and follow the recommended actions, including seeking medical advice if necessary.

[0094] This processing flow allows for real-time health monitoring within the home, and if an abnormality is detected, it is possible to respond immediately.

[0095] Example 1

[0096] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0097] In modern society, people's lifestyles are becoming more diverse, and health management at home is becoming increasingly important. Daily health monitoring is especially important for the elderly and those with chronic illnesses. However, many current health monitoring systems are single-modal, analyzing only specific data, and lack the ability to comprehensively analyze different types of data. This makes it difficult to accurately grasp overall health status and detect abnormalities early.

[0098] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0099] In this invention, the server includes means for collecting data from sensors used in the home, means for transmitting the collected data to the server, means for receiving and storing the transmitted data, means for analyzing the stored data using multimodal artificial intelligence, means for integrating the analyzed data and monitoring the health status, and means for generating and sending a notification to the user when an abnormality is detected. This makes it possible to integrate and analyze information from multiple data sources, accurately monitor the user's overall health status, and detect abnormalities early.

[0100] A "sensor used in the home" is a device that operates in a user's home environment and collects one or more of image data, video data, audio data, text data, and / or biometric data.

[0101] The "server" is a device or system that receives and stores data sent from sensors in the home, analyzes the data using multimodal artificial intelligence, and integrates the results to monitor the user's health condition.

[0102] "Multimodal artificial intelligence" refers to artificial intelligence technology that comprehensively analyzes different types of data (images, audio, text, biometric data, etc.).

[0103] "Analysis results" refers to analytical information generated by multimodal artificial intelligence based on collected data, including data indicating the user's health status.

[0104] A "notification" is a message generated when the server detects an anomaly based on the analysis results, informing the user of the nature of the anomaly and the recommended course of action.

[0105] This invention is a system that monitors a user's health condition by integrating and analyzing data from multiple sensors used in the home. This system collects data from multiple sensors, including image sensors, voice sensors, text input devices, and biometric sensors, and transmits it to a server. The server stores the received data and analyzes it using multimodal artificial intelligence (AI) to monitor the user's health condition. If an abnormality is detected, the server also sends a notification to the user.

[0106] To implement this system, the following hardware and software are required: image sensors (e.g., security cameras), audio sensors (e.g., smart speakers), text input devices (e.g., smart assistants or dedicated applications), and biometric sensors (e.g., wearable devices) installed in the home. The server requires storage to receive and store this data and computing power to run AI models for data analysis.

[0107] This system collects, transmits, analyzes, monitors, and notifies data in the following ways:

[0108] As a concrete example, let's take the case of detecting abnormal heart rates at night. The wearable device, which is the terminal, continuously measures the user's heart rate at night. The measured heart rate data is sent to a server at regular intervals using an encryption protocol. The server stores the received heart rate data and begins analysis. The main process of the analysis is to identify when the heart rate deviates from the normal range.

[0109] If an abnormality is detected, the server generates a notification stating, "A sudden increase in heart rate was detected overnight." The notification includes a graph of heart rate variability and a recommendation to seek medical advice. The notification is sent to the user's smartphone or smart assistant, allowing the user to take appropriate action quickly.

[0110] Additionally, here are some examples of input prompts for the generative AI model:

[0111] Prompt using natural language processing technology: "How were you feeling yesterday? Please let me know if there was anything in particular that bothered you."

[0112] Prompt when using voice recognition technology: "Good morning. How long did you sleep last night?"

[0113] In this way, the present invention makes it possible to monitor the user's health condition in real time at home and identify health risks at an early stage by comprehensively analyzing various data.

[0114] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0115] Step 1:

[0116] Data collection

[0117] The device collects data from various sensors in the home (image sensors, voice sensors, text input devices, and biometric sensors). Specifically, the following operations are performed:

[0118] The image sensor uses a camera to periodically capture image data and video data.

[0119] Input: Real-time video data

[0120] Output: Captured image and video data

[0121] The audio sensor collects the user's voice at any time.

[0122] Input: Real-time audio data

[0123] Output: Collected audio data

[0124] The text input device captures text data entered by the user.

[0125] Input: Text entered by the user

[0126] Output: The retrieved text data

[0127] Biometric sensors measure vital data such as heart rate, blood pressure, oxygen saturation, and sleep patterns.

[0128] Input: Anthropometric data

[0129] Output: Measured biological data

[0130] Step 2:

[0131] Sending data to the server

[0132] The terminal periodically transmits the collected data to the server. Specifically, the following operations are performed.

[0133] The device encrypts the collected data using security protocols (e.g., TLS / SSL).

[0134] Input: Collected data (images, audio, text, biometric data)

[0135] Output: Encrypted data

[0136] The encrypted data is sent to a server over the Internet.

[0137] Input: Encrypted data

[0138] Output: Data sent to server completed

[0139] Step 3:

[0140] Data storage

[0141] The server stores the received data. Specifically, the following operations are performed:

[0142] The received data is saved in the database.

[0143] Input: Received data (decrypted)

[0144] Output: Saved data

[0145] Step 4:

[0146] Data analysis

[0147] The server analyzes the stored data using multimodal artificial intelligence. Specifically, the following operations are performed:

[0148] Image data analysis: Using computer vision technology, image and video data is analyzed to detect changes in the user's movements and facial expressions.

[0149] Input: Image and video data

[0150] Output: Detection results of changes in behavior and facial expressions

[0151] Voice data analysis: Applying voice recognition technology to analyze changes in tone of voice and speaking patterns.

[0152] Input: Audio data

[0153] Output: Detected changes in tone of voice and speech patterns

[0154] Text data analysis: Using natural language processing techniques, we analyze the content and expression patterns of text.

[0155] Input: Text data

[0156] Output: Analysis results of text content and stress symptoms

[0157] Biometric data analysis: Perform time series analysis to detect abnormalities in heart rate, blood pressure, and sleep patterns.

[0158] Input: Biometric data

[0159] Output: Abnormality detection results for biometric data

[0160] Step 5:

[0161] Health monitoring

[0162] The server integrates the analysis results and monitors the user's health status. Specifically, the following operations are performed:

[0163] The results of various data analyses are integrated and a comprehensive health assessment is performed using multimodal AI.

[0164] Input: Detection results of each analysis data

[0165] Output: Overall health assessment results

[0166] Step 6:

[0167] Generate and send notifications when anomalies occur

[0168] If the server detects an abnormality, it generates a notification and sends it to the user. Specifically, the following operations are performed:

[0169] A notification message is generated based on the anomaly detection.

[0170] Input: Anomaly detection result

[0171] Output: The generated notification message

[0172] Send notification messages to users' smartphones or smart assistants.

[0173] Input: Notification message

[0174] Output: Notification sent to user

[0175] This processing flow enables the system to monitor the user's health condition in real time and detect abnormalities early.

[0176] (Application example 1)

[0177] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0178] While self-driving vehicles are becoming increasingly common in modern society, there is currently no system that monitors the driver's health in real time and responds promptly when an abnormality occurs. This may increase the risk of accidents due to driver health abnormalities. The present invention aims to solve this problem and provide an effective means for improving the safety of self-driving vehicles.

[0179] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0180] In this invention, the server includes means for collecting data from sensors used in the home, means for transmitting the collected data to the server, means for receiving and storing the transmitted data, means for analyzing the stored data using multimodal artificial intelligence, means for integrating the analysis results to monitor the health condition, means for generating and sending a notification to the user when an abnormality is detected, and means for collecting data from internal sensors of the autonomous vehicle and monitoring the driver's health condition. This makes it possible to monitor the driver's health condition in the vehicle in real time and send an immediate notification when an abnormality is detected.

[0181] "Domestic sensors" are devices installed inside the home that collect image, audio, text, biometric data, etc.

[0182] "Means for transmitting to the server" refers to a method or device for transmitting collected data to the server via a communication line such as the Internet.

[0183] "Storage means" refers to a method or device for retaining data received by the server for an extended period of time.

[0184] "Multimodal AI" refers to AI that has the ability to analyze multiple types of data (images, audio, text, biometric data, etc.) and integrate them.

[0185] "Health monitoring means" means a method or device that uses multimodal artificial intelligence to monitor analysis results and continuously assess a user's health status.

[0186] "Means for generating and sending a notification to the user when an abnormality is detected" refers to a method or device that automatically generates a warning or alert when an abnormality is discovered from the analysis results and forwards it to the user.

[0187] "Internal sensors in autonomous vehicles" refer to sensors that are placed inside autonomous vehicles and collect biometric and behavioral data from the driver.

[0188] "Means for monitoring the driver's health condition" refers to methods or devices that analyze data obtained from sensors in an autonomous vehicle and monitor the driver's health condition in real time.

[0189] The present invention provides a system for monitoring the health status of a user and a driver by comprehensively analyzing data from multiple sensors used in homes and autonomous vehicles. Specific embodiments are described below.

[0190] System Overview

[0191] The system collects data from multiple sensors (image sensors, voice sensors, text input devices, and biometric sensors) installed in the user's home and autonomous vehicle and sends it to a server. The server stores the received data and analyzes it using multimodal artificial intelligence (AI) to monitor the health status of the user and driver. If an abnormality is detected, the server generates a notification and sends it to the user or driver.

[0192] Hardware and software used

[0193] Hardware:

[0194] Home sensors (cameras, microphones, biometric sensors)

[0195] Internal sensors in autonomous vehicles (heart rate sensors, blood pressure sensors, cameras, etc.)

[0196] software:

[0197] Data Collection Applications

[0198] Communication protocol (HTTP, HTTPS)

[0199] Server (for data storage and analysis)

[0200] Multimodal AI analysis engine

[0201] Details of data processing and calculation

[0202] Data collection:

[0203] The devices (in-home sensors and in-vehicle sensors) collect image data, voice data, text data, biometric data, etc. This data is sent to a server at specific intervals.

[0204] Data transmission:

[0205] The device encrypts the collected data and sends it to the server using HTTPS, ensuring data security.

[0206] Data Analysis:

[0207] The server stores the received data and analyzes it using multimodal AI. The specific analysis content is as follows:

[0208] Image data: Analyzes changes in the user's or driver's behavior and facial expressions to detect abnormalities.

[0209] Voice data: Analyzed using voice recognition technology to detect changes in tone of voice and speaking patterns.

[0210] Text data: Analyze using natural language processing techniques to detect text content and signs of stress.

[0211] Biometric data: Time series analysis of heart rate, blood pressure abnormalities, and disrupted sleep patterns.

[0212] Anomaly detection and notification:

[0213] Multimodal AI integrates the results of various data analyses to assess the health status of the user and driver. If an abnormality is detected, the server generates a notification message and sends it to the user's smartphone or vehicle's infotainment system. The notification includes details of the abnormality and recommended actions to take.

[0214] Specific examples

[0215] For example, if a biometric sensor detects an abnormal heart rate while a driver of an autonomous vehicle is in the vehicle, the sensor will send this information to a server in real time. The server will analyze the heart rate and, if an abnormality is detected, will send a notification to the driver stating, "Your heart rate is outside the normal range. Please consult a medical professional."

[0216] Example prompts for generative AI models

[0217] "Please demonstrate the process flow of an application that analyzes data collected from internal sensors in an autonomous vehicle and monitors the driver's health. The implementation language should be Python, and the following sensors and data types should be assumed:

[0218] image

[0219] audio

[0220] text

[0221] Biometric data (heart rate, blood pressure)

[0222] Also include the ability to send data to a server and provide notifications based on the analysis results.

[0223] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0224] Step 1:

[0225] Devices (sensors in homes and vehicles) collect data. Specifically, cameras capture image and video data, microphones record voice data, text input devices acquire user input text, and biometric sensors measure biometric data such as heart rate and blood pressure. All of this data is collected from each sensor at regular intervals.

[0226] Input: Image data, audio data, text data, and biometric data obtained from each sensor.

[0227] Output: Consolidated sensor data.

[0228] Step 2:

[0229] The terminal sends the collected data to the server. The collected data is encrypted using a communication protocol (e.g., HTTPS) and sent securely to the server. The data is sent at regular intervals.

[0230] Input: Integrated sensor data.

[0231] Output: The data sent to the server.

[0232] Step 3:

[0233] The server stores the received data. The server stores the received data in a database for later analysis. During this storage process, each data type (image, audio, text, biometric data) is recorded in chronological order.

[0234] Input: The data sent to the server.

[0235] Output: Data stored in the database.

[0236] Step 4:

[0237] The server analyzes the stored data using multimodal artificial intelligence. Specifically, image data is analyzed using computer vision technology to detect movements and facial expressions. Audio data is analyzed using voice recognition technology to detect tone of voice and speech patterns. Text data is analyzed using natural language processing technology to detect text content and signs of stress. Biometric data is analyzed over time to detect abnormalities in heart rate and blood pressure.

[0238] Input: Data stored in a database.

[0239] Output: Analysis results.

[0240] Step 5:

[0241] The server integrates the analysis results and monitors the health status of the user and driver. Multimodal AI integrates the analysis results of each data type to assess the overall health status. From this analysis, anomalies are identified.

[0242] Input: Analysis results for each data type.

[0243] Output: Integrated health status assessment.

[0244] Step 6:

[0245] If the server detects an abnormality, it generates a notification and sends it to the user or driver. The server generates a notification message containing the details of the abnormality and recommended actions to take, and sends it to the user's smartphone or the vehicle's infotainment system.

[0246] Input: Integrated Health Status Assessment.

[0247] Output: Informational message.

[0248] Step 7:

[0249] The user or driver receives a notification and takes appropriate action, such as consulting a medical professional.

[0250] Input: Notification message.

[0251] Output: Coping behavior.

[0252] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0253] The present invention is a system that monitors a user's health condition by comprehensively analyzing data from multiple sensors used in the home. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it is possible to monitor the user's emotional health as well. Specific embodiments of the system will now be described.

[0254] System Overview

[0255] The system collects data from multiple sensors (image sensors, voice sensors, text input devices, and biometric sensors) installed in the user's home and sends it to a server. The server stores the received data and analyzes it using a multimodal artificial intelligence (AI) and emotion engine to monitor the user's health and emotional state. If an abnormality is detected, the server sends a notification to the user.

[0256] Program processing flow

[0257] Data collection

[0258] Devices (users' home sensors and wearable devices) collect data

[0259] Image data: Security cameras in the home periodically take images and videos, recording the user's movements, facial expressions, and daily life.

[0260] Voice data: Smart speakers and smart assistants collect your voice and use the recorded audio to analyze your tone of voice and speaking patterns.

[0261] Text data: Collects text data entered by users through smart assistants or dedicated applications, and analyzes changes in the content and expression of the text.

[0262] Biometric data: Wearable devices measure biometric data such as heart rate, blood pressure, oxygen saturation, and sleep patterns to continuously monitor the user's physical condition.

[0263] Data transmission

[0264] The device sends the collected data to the server.

[0265] The sensor terminals send the collected data to the server at regular intervals, and the data is encrypted using a security protocol.

[0266] Data analysis

[0267] The server analyzes the received data

[0268] Image analysis: The server uses image data to analyze changes in the user's movements and facial expressions, thereby detecting physical changes and patterns in daily activities.

[0269] Voice analysis: The server uses speech recognition technology to convert voice data into text and analyzes changes in tone and speaking style to identify emotional issues and changes in cognitive function.

[0270] Text analysis: The server uses natural language processing technology to analyze the text data and detect changes in the content and expressions of the text, which allows for an assessment of stress and mental health.

[0271] Biometric data analysis: The server uses time series analysis technology to analyze data such as heart rate, blood pressure, oxygen saturation, and sleep patterns to detect abnormalities, allowing for early detection of cardiac and respiratory problems.

[0272] Emotion analysis

[0273] The server uses an emotion engine to analyze the user's emotions.

[0274] Emotion analysis using voice data: The server analyzes the voice data and uses an emotion engine to recognize the user's emotional state from changes in their tone of voice and speaking style.

[0275] Emotion analysis using text data: The server analyzes the text data and uses an emotion engine to recognize the emotional state from changes in the content and expression of the text.

[0276] Emotion analysis using image data: The server analyzes the image data and uses an emotion engine to recognize the user's emotional state from their facial expressions.

[0277] Health and emotional state monitoring

[0278] The server integrates the analysis results and monitors the health and emotional state.

[0279] By integrating the results of various data analyses, the multimodal AI and emotion engine assesses the user's overall health and emotional state, making it easier to detect abnormalities.

[0280] Send notifications

[0281] If the server detects an abnormality, it will notify the user.

[0282] If the server detects an anomaly, it generates a notification message and sends it to the user's smartphone or smart assistant, containing details of the anomaly and recommended actions to take.

[0283] Example: Emotion and health monitoring using voice data

[0284] The device (smart speaker) collects the user's voice from everyday conversations. The server analyzes the received voice data and converts it into text data using voice recognition technology. It then uses an emotion engine to analyze subtle changes in voice tone and speaking style to assess whether the user is feeling stressed. The server also uses multimodal AI to analyze heart rate data collected simultaneously with the voice data, detecting physical changes such as an abnormally high heart rate. When an abnormality is detected, the server immediately generates a notification and sends it to the user's smartphone, stating, "Your recent conversations have shown signs of stress, and your heart rate is also high. Please try some relaxation techniques." This allows for a prompt response.

[0285] In this way, the present invention can protect the overall health of users by monitoring both health and emotions in the home in real time and identifying risks early.

[0286] The processing flow will be explained below.

[0287] Step 1:

[0288] Devices (sensors) collect data. Home security cameras periodically capture images and videos, smart speakers and smart assistants record voice data, and text data entered by users through smart assistants and applications is also collected. Wearable devices also measure biometric data such as heart rate, blood pressure, oxygen saturation, and sleep patterns in real time.

[0289] Step 2:

[0290] The terminal (sensor) packetizes and encrypts the collected data, and the data packets are temporarily stored in a queue inside the terminal.

[0291] Step 3:

[0292] The device (sensor) sends data packets to the server at regular intervals. The data transmission is performed using a secure communication protocol to ensure the confidentiality and security of the data.

[0293] Step 4:

[0294] The server receives the data packet. The received data packet is decoded and restored to its original format. The data integrity is checked, and if there are any errors or omissions, a retransmission request is sent to the terminal.

[0295] Step 5:

[0296] The server stores the recovered data in a distributed database, where the data is managed for each user and tagged within the database.

[0297] Step 6:

[0298] The server analyzes the received image and video data. Image classification and motion analysis algorithms are used to detect changes in the user's movement patterns and facial expressions. For example, unusual movements or changes in facial expressions may be signs of an abnormality.

[0299] Step 7:

[0300] The server analyzes the voice data. It uses speech recognition technology to convert the data into text and analyzes changes in tone and speaking style. This allows it to detect emotional issues and changes in cognitive function. For example, a trembling voice or changes in speaking style could be a sign of stress or changes in cognitive function.

[0301] Step 8:

[0302] The server analyzes the text data. It uses natural language processing technology to analyze the content of the text and detect changes in the content and expressions of the sentences. This allows for an assessment of stress and mental health. For example, an increase in negative expressions in the sentences can be determined to be a sign of stress.

[0303] Step 9:

[0304] The server analyzes biometric data, using time series analysis techniques to analyze data such as heart rate, blood pressure, oxygen saturation, and sleep patterns to detect abnormalities. This allows for early detection of cardiac and respiratory problems. For example, a higher-than-normal heart rate could be a sign of cardiac disease.

[0305] Step 10:

[0306] The server uses an emotion engine to analyze the user's emotions. It recognizes the user's emotions from voice, text, and image data. This allows the system to monitor the user's emotional health. For example, if anger, excitement, or sadness is detected from the voice data, it is considered a sign of emotional instability.

[0307] Step 11:

[0308] The server integrates the results of each analysis, and uses a multimodal AI and emotion engine to comprehensively analyze images, voice, text, and biometric data to assess the user's overall health and emotional state, making it easier to detect abnormalities.

[0309] Step 12:

[0310] If the server detects an anomaly, it generates a notification with details about the anomaly and a recommended course of action, such as "Your heart rate has been higher than normal recently, and you appear to be feeling stressed."

[0311] Step 13:

[0312] The server generates a notification and sends it to the user's smartphone or smart assistant. The notification is sent immediately, informing the user that an abnormality has occurred.

[0313] Step 14:

[0314] Users receive notifications, review them, take action based on recommended actions, and, if necessary, seek medical advice. This allows users to quickly respond to changes in their health and emotional state.

[0315] This specific processing flow enables real-time monitoring of both health and emotions within the home, identifying risks early and protecting the overall health of users.

[0316] Example 2

[0317] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0318] There is a need for technology that can comprehensively monitor not only a user's health status but also their emotional state by collecting and analyzing different types of data from multiple sensors used in the home. However, existing systems have faced many challenges in centralized data management, complex data analysis, anomaly detection, and prompt notification to the user. In particular, it has been difficult to integrate information from multiple data sources and perform accurate analysis.

[0319] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for collecting data from sensors used in the home, means for transmitting the collected data to the server, means for receiving and storing the transmitted data, means for analyzing the stored data using multimodal artificial intelligence and an emotion analysis engine, means for integrating the analysis results to monitor the health status and emotional state, and means for generating and sending a notification to the user when an abnormality is detected. This makes it possible to comprehensively monitor the user's health status and emotional state in real time and to provide a prompt and appropriate notification when an abnormality occurs.

[0320] A "sensor" is a device used in the home that collects information such as images, audio, text, and biometric data.

[0321] "Data" refers to all information collected by sensors, such as image data, video data, audio data, text data, and biometric data.

[0322] A "server" is an information processing device that receives and stores data sent from sensors, and analyzes the data to monitor the user's condition.

[0323] "Multimodal AI" is an AI technology that has the ability to integrate and analyze multiple different types of data (e.g., images, audio, text, biometric data) in a unified manner.

[0324] An "emotion analysis engine" is a technology for analyzing a user's emotional state from data such as voice, text, and images.

[0325] "Analyzing" means that the server processes the data it receives and derives meaningful results.

[0326] "Monitoring" means that the server continuously observes and evaluates the user's health and emotional state based on the analysis results.

[0327] A "notification" is a warning or advice message that is generated and sent to the user when the server detects an abnormality.

[0328] An "anomaly" is any unusual change or problem in the user's health or emotional state that is detected as a result of the server's analysis.

[0329] "Integrating" means that the server compiles the results of multiple data analyses and provides a single comprehensive evaluation.

[0330] The present invention provides a system for monitoring a user's health and emotional state by comprehensively analyzing data from multiple sensors used in the home. A specific method for implementing this system is described below.

[0331] System configuration

[0332] The system collects data from multiple sensors installed in the home (image sensors, voice sensors, text input devices, and biometric sensors) and sends it to a server. The server stores the received data and analyzes it using a multimodal artificial intelligence (AI) and emotion analysis engine. The analysis results are integrated to monitor the user's health and emotional state. If an abnormality is detected, the server sends a notification to the user.

[0333] Hardware and software used

[0334] Sensors used in the home include:

[0335] Image sensor: Home security camera

[0336] Voice Sensors: Smart Speakers and Smart Assistants

[0337] Text input devices: smart assistants and dedicated applications

[0338] Biosensors: Wearable devices that measure heart rate, blood pressure, oxygen saturation, and sleep patterns

[0339] The server includes the following analysis modules:

[0340] Multimodal Artificial Intelligence (AI)

[0341] Sentiment Analysis Engine

[0342] Data Storage and Management System

[0343] Notification System

[0344] Data collection and transmission

[0345] 1. Devices (sensors) collect data

[0346] Image sensors capture images and videos of the home and record changes in the user's movements and facial expressions. For example, a camera in the living room can start recording at a set time.

[0347] The voice sensor uses smart speakers to collect everyday conversations and analyze tone and speech patterns, for example, recording the voice you use to ask about the weather.

[0348] The text input device collects text entered by the user through a dedicated application and analyzes changes in the content and expression of the text, for example, collecting diary entries from the user.

[0349] Biometric sensors use wearable devices to continuously measure heart rate, blood pressure, oxygen saturation, and sleep patterns. For example, a smartwatch worn by a user records heart rate.

[0350] 2. The device sends the collected data to the server

[0351] The sensor terminals transmit the collected data to a server at regular intervals. The transmitted data is encrypted using a security protocol. For example, audio data is uploaded to the server every hour.

[0352] Data analysis

[0353] 1. The server analyzes the received data

[0354] Image analysis: The server uses image data to analyze changes in the user's movements and facial expressions, for example, to detect changes in the user's walking pattern.

[0355] Voice analysis: The server uses speech recognition technology to convert the voice data into text and analyzes changes in tone and speaking style, for example detecting if the user's voice is darker than usual.

[0356] Text analysis: The server uses natural language processing technology to analyze the text data and detect changes in the content and expressions of the sentences. For example, it analyzes the frequent use of the expression "tired."

[0357] Biometric data analysis: The server uses time series analysis to analyze heart rate, blood pressure, oxygen saturation, sleep patterns, etc. to detect abnormalities, such as a higher than normal heart rate.

[0358] 2. Emotion analysis

[0359] Emotion analysis from voice data: The server analyzes the voice data and uses an emotion analysis engine to recognize the emotional state from changes in tone of voice and speaking style, for example, detecting an increase in anger.

[0360] Sentiment analysis based on text data: The server analyzes the text data and uses a sentiment analysis engine to recognize the emotional state from changes in the content and expressions of the text. For example, it detects an increase in pessimistic expressions.

[0361] Emotion analysis based on image data: The server analyzes the image data and uses an emotion analysis engine to recognize the user's emotional state from their facial expressions. For example, it detects that they are smiling less.

[0362] Integrated monitoring and abnormality notification

[0363] 1. The server integrates the analysis results and monitors the health and emotional state.

[0364] Using a multimodal AI and emotion analysis engine, the analysis results are integrated to assess the user's overall health and emotional state. For example, if different emotional signals are detected from image analysis, audio analysis, and text analysis, a comprehensive assessment is made.

[0365] 2. Detect anomalies and generate notifications

[0366] If an anomaly is detected, the server generates a real-time notification. For example, if heart rate and voice data indicate strong signs of stress, a notification will be sent to the user stating, "Your recent conversations have shown signs of stress, and your heart rate is high. Please try some relaxation techniques."

[0367] Specific examples

[0368] For example, a smart speaker collects a user's voice from everyday conversations and sends the data to a server. The server analyzes the voice data and converts it into text data using speech recognition technology. It uses an emotion analysis engine to analyze subtle changes in the tone of voice and speaking style to assess whether the user is feeling stressed. At the same time, the server also analyzes heart rate data from a wearable device and detects if the heart rate is higher than normal. If an abnormality is detected, the server sends the user a notification saying, "Your recent conversations have shown signs of stress, and your heart rate is high. Please try some relaxation techniques."

[0369] Example prompt: "Analyze the voice data collected by the smart speaker and assess whether the user is experiencing stress."

[0370] This invention can protect the overall health of users by monitoring both health and emotions in the home in real time and identifying risks early.

[0371] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0372] Step 1: Collect data

[0373] Data is collected by sensors and devices installed in the home by users.

[0374] Input: Various data from the home environment and the user (images, voice, text, biometric data).

[0375] How it works: Security cameras take pictures, smart speakers record everyday conversations, and wearable devices measure heart rate and blood pressure.

[0376] Output: Various data collected.

[0377] Step 2: Sending data

[0378] The terminal transmits the collected data to the server.

[0379] Input: Various types of collected data (images, audio, text, biometric data).

[0380] Specific operation: Image data captured by the security camera is uploaded to the server, and audio data recorded by the smart speaker is sent to the server.

[0381] Output: The data sent to the server.

[0382] Step 3: Save your data

[0383] The server stores the received data.

[0384] Input: The data sent to the server.

[0385] Specific operation: The server stores various data in the database.

[0386] Output: Saved data.

[0387] Step 4: Analyzing the image data

[0388] The server analyzes the image data.

[0389] Input: Saved image data.

[0390] Specific operation: Using multimodal AI to analyze changes in user behavior and facial expressions from image data, for example, using a facial recognition algorithm to analyze user facial expressions.

[0391] Output: Analysis results (changes in user's behavior and facial expressions).

[0392] Step 5: Analyzing the audio data

[0393] The server analyzes the audio data.

[0394] Input: Stored audio data.

[0395] What it does: It uses speech recognition technology to convert voice data into text, and uses an emotion analysis engine to analyze changes in tone and speaking style, for example, to identify changes in tone in a user's everyday conversation.

[0396] Output: Analysis results (changes in tone of voice and speaking style).

[0397] Step 6: Analyze the text data

[0398] The server parses the text data.

[0399] Input: Saved text data.

[0400] Specific operation: Analyzes text data using natural language processing technology and detects changes in the content and expressions of sentences using a sentiment analysis engine. For example, it analyzes the frequent use of the expression "tired."

[0401] Output: Analysis results (changes in the content and expression of the text).

[0402] Step 7: Analyzing the biometric data

[0403] The server analyzes the biometric data.

[0404] Input: Stored biometric data.

[0405] Specific operation: Using time series analysis technology, it analyzes heart rate, blood pressure, oxygen saturation, sleep patterns, etc. to detect abnormalities, such as when the heart rate is higher than normal.

[0406] Output: Analysis results (heart rate, blood pressure, oxygen saturation, changes in sleep patterns).

[0407] Step 8: Integrating the analysis results

[0408] The server consolidates the various analysis results.

[0409] Input: Analysis results of various data (images, audio, text, biometric data).

[0410] How it works: Using a multimodal AI and emotion analysis engine, the system integrates analysis results to assess the user's overall health and emotional state. For example, if different emotional indicators are found from image analysis, audio analysis, and text analysis, it will comprehensively assess them.

[0411] Output: Integrated analysis results (assessment of the user's health and emotional state).

[0412] Step 9: Detect anomalies

[0413] The server detects an abnormality.

[0414] Input: Consolidated analysis results.

[0415] Specific operation: Based on the analysis results, an anomaly detection algorithm is executed to detect abnormalities in the user's health and emotional state, such as stress and heart rate abnormalities.

[0416] Output: Anomaly detection results.

[0417] Step 10: Generate and send notifications

[0418] The server generates the notification and sends it to the user.

[0419] Input: Anomaly detection results.

[0420] What it does: If an abnormality is detected, it will generate a notification message and send it to the user's smartphone or smart assistant, for example, "Your recent conversations have shown signs of stress and your heart rate is elevated. Please try some relaxation techniques."

[0421] Output: The notification message sent to the user.

[0422] (Application example 2)

[0423] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0424] Conventional health monitoring systems in homes and businesses are limited to monitoring physical health conditions, making it difficult to respond quickly and accurately to changes or abnormalities in emotional states. Furthermore, the means of detecting abnormalities are limited, and detection of abnormalities can be delayed under certain conditions. Therefore, there is a need for a system that can monitor both health and emotional states in an integrated manner, quickly detect abnormalities, and provide appropriate notifications.

[0425] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting data from sensors used in homes and business facilities, means for transmitting the collected data to the server, means for receiving and storing the transmitted data, means for analyzing the stored data using multimodal artificial intelligence, means for integrating the analysis results to monitor health and emotional states, means for generating and transmitting a notification to the user when an abnormality is detected, and means for analyzing changes in the movements, facial expressions, and tone of voice of visitors and employees to identify their emotional states. This makes it possible to monitor not only health states but also emotional states in real time, quickly detect abnormalities, and issue appropriate notifications.

[0426] "Inside a home or business premises" refers to the physical environment in which a person lives or conducts business activities, and refers to the location of the system used within that particular space.

[0427] A "sensor" is a device that detects physical phenomena and converts them into electronic data, and includes image sensors, audio sensors, text input devices, and biometric sensors.

[0428] "Means for collecting data" refers to a system element that uses sensors to obtain information from an object and record it in electronic form.

[0429] "Means for transmitting data to a server" refers to the communications devices and protocols used to transfer the collected digital data over a network to a remote server.

[0430] "Means for receiving and storing data" refers to a system element that allows a server to receive data sent via a network and store it in a database or the like.

[0431] "Multimodal AI" refers to an AI technology that simultaneously analyzes multiple types of data (e.g., images, audio, text, biometric data) and makes integrated decisions.

[0432] "Health and emotional monitoring means" refers to system elements that provide continuous monitoring of an individual's physical and emotional health based on collected and analyzed data.

[0433] "Means for generating and sending notifications to users when an anomaly is detected" refers to the ability for the system to automatically create and send alerts or notifications to users' devices when an anomaly is detected.

[0434] "Means of identifying emotional states by analyzing changes in behavior, facial expressions, and tone of voice" refers to artificial intelligence technology that analyzes image and audio data to infer emotional states from behavior, changes in facial expressions, tone of voice, etc.

[0435] An embodiment of the present invention is described in detail below: The system monitors the health and emotional state of visitors and employees by collecting data from multiple sensors installed in homes and businesses, transmitting the data to a server for analysis.

[0436] System Configuration

[0437] Terminal (sensor)

[0438] The devices include image sensors, voice sensors, text input devices, and biometric sensors that are placed within physical stores to collect the following data from visitors and employees:

[0439] Image data: The camera records your movements and expressions.

[0440] Audio data: Microphones collect speech.

[0441] Text data: The digital device captures the input text.

[0442] Biometric data: Wearable devices and thermal cameras collect heart rate and temperature data.

[0443] The data collected by each terminal is encrypted and sent to a server at regular intervals.

[0444] Server Processing

[0445] Data reception and storage

[0446] The server receives the data sent from the terminal via the network in an encrypted form and stores it in a database, using technologies including security protocols and database systems.

[0447] Data analysis

[0448] The server analyzes the various types of stored data (images, audio, text, and biometric data) using multimodal artificial intelligence (AI). Image data is analyzed for changes in movements and facial expressions using libraries such as OpenCV, and audio data is converted into text using pyaudio and librosa, after which changes in tone of voice and speaking style are analyzed. Text data is analyzed for emotional state using natural language processing techniques such as TextBlob. Furthermore, biometric data is analyzed for abnormalities using time series analysis techniques.

[0449] Emotion analysis

[0450] The server uses an emotion engine to integrate data from different modalities and recognize the emotional state of visitors and employees, allowing real-time monitoring of changes in not only physical health but also emotional health.

[0451] Notifications and Feedback

[0452] If the server detects an abnormality, it automatically generates a notification message and sends it to the user's smartphone or other display device. The notification includes details of the detected abnormality and recommended actions to take. For example, a message might be generated that reads, "Employee A has a high body temperature and his tone of voice indicates stress. We recommend that he leave work early."

[0453] Examples and prompts

[0454] As a concrete example, consider a monitoring system for Employee A in a store. The server collects and analyzes Employee A's temperature and conversation data, and if an abnormality is detected, it sends a notification saying, "Employee A's temperature is high and his / her tone of voice also indicates stress. We recommend that you leave early." An example of how such a prompt is input into a generative AI model is as follows:

[0455] "Monitor employee health and emotions, and if you detect a high temperature or stressful tone in the conversation, generate a notification recommending that employees leave early. Could you please give us an example of what this notification should look like?"

[0456] This will enable the realization of a system that monitors health and emotions in an integrated manner and can respond quickly in the event of an abnormality.

[0457] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0458] Step 1:

[0459] The terminal collects image data, audio data, text data, and biometric data. Image data is obtained by recording the movements and facial expressions of visitors and employees with a camera, while audio data is obtained by collecting conversations and environmental sounds with a microphone. Text data is obtained by acquiring input information from digital terminals, and biometric data is obtained by collecting heart rate and body temperature information with a wearable device or a body temperature detection camera. The input is various data obtained from each sensor, and the output is the collected raw data.

[0460] Step 2:

[0461] The terminal sends the collected data to the server at regular intervals. The data is sent encrypted over the network. In this step, the input is the collected data, and the output is the encrypted data sent to the server.

[0462] Step 3:

[0463] The server saves the received data. The received data is stored in the server's database. The input of this step is the encrypted transmitted data, and the output is the saved data.

[0464] Step 4:

[0465] The server analyzes the stored data. The analysis is divided into the following parts:

[0466] Image analysis: Image data is analyzed using OpenCV etc. to detect changes in movements and facial expressions. The input is the saved image data, and the output is the analysis results of movement and facial expression changes.

[0467] Speech analysis: Using pyaudio and librosa, audio data is converted to text and changes in voice tone and speaking style are analyzed. The input is the stored audio data, and the output is the converted text data and changes in voice tone and speaking style.

[0468] Text Analysis: Analyzes text data using natural language processing techniques such as TextBlob to evaluate emotional states. The input is the stored text data, and the output is the analysis result of the emotional state.

[0469] Biometric data analysis: Detects abnormalities in biometric data (heart rate, body temperature, etc.) using time series analysis techniques. The input is stored biometric data, and the output is the anomaly detection results.

[0470] Step 5:

[0471] The server uses the emotion engine to integrate the analysis results of each data and monitor the health and emotional states of visitors and employees. The input of this step is each analysis result, and the output is the integrated evaluation result of the health and emotional states.

[0472] Step 6:

[0473] If the server detects an abnormality, it generates a notification message and sends it to the user's smartphone or other display device. The notification includes details of the abnormality and recommended actions to take. For example, a notification might be sent stating, "Employee A's temperature is high and his / her tone of voice indicates stress. We recommend that he / she leave work early." The input to this step is the integrated evaluation result, and the output is the generated notification message.

[0474] The above is the flow of processing in the embodiment of the present invention.

[0475] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0476] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0477] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0478] [Second embodiment]

[0479] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0480] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0481] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0482] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0483] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0484] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0485] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0486] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0487] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0488] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0489] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0490] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0491] The present invention is a system for monitoring the health condition of a user by comprehensively analyzing data from multiple sensors used in the home. A specific embodiment of the system will be described.

[0492] System Overview

[0493] The system collects data from multiple sensors (image sensors, voice sensors, text input devices, and biometric sensors) installed in the user's home and sends it to a server. The server stores the received data, analyzes it using multimodal artificial intelligence (AI), and monitors the user's health status. If an abnormality is detected, the server sends a notification to the user.

[0494] Program processing flow

[0495] Data collection

[0496] Devices (users' home sensors and wearable devices) collect data

[0497] Image data: Security cameras in the home periodically take images and videos, recording the user's movements, facial expressions, and daily life.

[0498] Voice data: Smart speakers and smart assistants collect your voice and use it to detect changes in your tone and speaking style.

[0499] Text data: Collects text data entered by users through smart assistants or dedicated applications, and analyzes the content of the text and signs of stress.

[0500] Biometric data: Wearable devices measure biometric data such as heart rate, blood pressure, oxygen saturation, and sleep patterns, thereby continuously monitoring the user's physical condition.

[0501] Data transmission

[0502] The device sends the collected data to the server.

[0503] The sensor terminals send the collected data to the server at regular intervals, and the data is encrypted using a security protocol.

[0504] Data analysis

[0505] The server analyzes the received data

[0506] The server uses the image data to analyze changes in the user's movements and facial expressions. For example, unusual movements may indicate a health problem.

[0507] Voice data is analyzed using speech recognition technology to detect changes in tone of voice and speech patterns, which can identify emotional issues and changes in cognitive function.

[0508] The text data is analyzed using natural language processing technology to detect changes in the content and expressions of the text, which can then be used to assess stress and mental health.

[0509] Biometric data is analyzed over time to detect abnormalities in heart rate, blood pressure, and sleep patterns, allowing for early detection of cardiac and respiratory problems.

[0510] Health Status Monitoring

[0511] The server consolidates the analysis results and monitors health status.

[0512] By integrating the results of various data analyses, multimodal AI assesses the user's overall health status, making it easier to detect abnormalities.

[0513] Send notifications

[0514] If the server detects an abnormality, it will notify the user.

[0515] If the server detects an anomaly, it generates a notification message and sends it to the user's smartphone or smart assistant, containing details of the anomaly and recommended actions to take.

[0516] Example: Detecting abnormal heart rates at night

[0517] The terminal (wearable device) continues to measure the user's heart rate throughout the night. The server receives and stores the collected heart rate data. The server then analyzes the received data and detects abnormal increases in heart rate beyond the normal range. If an abnormality is detected, the server sends a notification to the user's smartphone stating, "A sudden increase in heart rate has been detected overnight." The user receives this notification and can take prompt action, such as consulting a medical institution.

[0518] In this way, the present invention can protect the user's health by monitoring the user's health status in real time at home and identifying health risks early.

[0519] The processing flow will be explained below.

[0520] Step 1:

[0521] Devices (sensors) collect data. Security cameras installed in homes periodically capture images and videos, and smart speakers and smart assistants record voice data. Text data entered by users through smart assistants and applications is also collected. In addition, wearable devices measure biometric data such as heart rate, blood pressure, oxygen saturation, and sleep patterns in real time.

[0522] Step 2:

[0523] The terminal (sensor) packetizes and encrypts the collected data, and the data packets are temporarily stored in a queue inside the terminal.

[0524] Step 3:

[0525] The device (sensor) sends data packets to the server at regular intervals. The data transmission is performed using a secure communication protocol to ensure the confidentiality and security of the data.

[0526] Step 4:

[0527] The server receives the data packet. The received data packet is decoded and restored to its original format. The data integrity is checked, and if there are any errors or omissions, a retransmission request is sent to the terminal.

[0528] Step 5:

[0529] The server stores the recovered data in a distributed database, where the data is managed for each user and tagged within the database.

[0530] Step 6:

[0531] The server analyzes the received image and video data, using image classification and motion analysis algorithms to detect changes in the user's movement patterns and facial expressions.

[0532] Step 7:

[0533] The server analyzes the audio data, converts it into text using speech recognition technology, and analyzes changes in tone and speaking style to identify signs associated with emotional problems and changes in cognitive function.

[0534] Step 8:

[0535] The server analyzes the text data and uses natural language processing technology to analyze the content of the text data and evaluate stress levels and health conditions.

[0536] Step 9:

[0537] The server analyzes biometric data, using time series analysis techniques to analyze data such as heart rate, blood pressure, oxygen saturation, and sleep patterns to detect abnormalities, especially values ​​that exceed the normal range.

[0538] Step 10:

[0539] The server integrates the results of each analysis, using multimodal AI to comprehensively analyze images, voice, text, and biometric data to assess the user's overall health.

[0540] Step 11:

[0541] The server generates a notification when an anomaly is detected, containing details about the anomaly and a recommended course of action.

[0542] Step 12:

[0543] The server generates a notification and sends it to the user's smartphone or smart assistant. The notification is sent immediately, informing the user that an abnormality has occurred.

[0544] Step 13:

[0545] Users will receive a notification, which they can review and follow the recommended actions, including seeking medical advice if necessary.

[0546] This processing flow allows for real-time health monitoring within the home, and if an abnormality is detected, it is possible to respond immediately.

[0547] Example 1

[0548] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0549] In modern society, people's lifestyles are becoming more diverse, and health management at home is becoming increasingly important. Daily health monitoring is especially important for the elderly and those with chronic illnesses. However, many current health monitoring systems are single-modal, analyzing only specific data, and lack the ability to comprehensively analyze different types of data. This makes it difficult to accurately grasp overall health status and detect abnormalities early.

[0550] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0551] In this invention, the server includes means for collecting data from sensors used in the home, means for transmitting the collected data to the server, means for receiving and storing the transmitted data, means for analyzing the stored data using multimodal artificial intelligence, means for integrating the analyzed data and monitoring the health status, and means for generating and sending a notification to the user when an abnormality is detected. This makes it possible to integrate and analyze information from multiple data sources, accurately monitor the user's overall health status, and detect abnormalities early.

[0552] A "sensor used in the home" is a device that operates in a user's home environment and collects one or more of image data, video data, audio data, text data, and / or biometric data.

[0553] The "server" is a device or system that receives and stores data sent from sensors in the home, analyzes the data using multimodal artificial intelligence, and integrates the results to monitor the user's health condition.

[0554] "Multimodal artificial intelligence" refers to artificial intelligence technology that comprehensively analyzes different types of data (images, audio, text, biometric data, etc.).

[0555] "Analysis results" refers to analytical information generated by multimodal artificial intelligence based on collected data, including data indicating the user's health status.

[0556] A "notification" is a message generated when the server detects an anomaly based on the analysis results, informing the user of the nature of the anomaly and the recommended course of action.

[0557] This invention is a system that monitors a user's health condition by integrating and analyzing data from multiple sensors used in the home. This system collects data from multiple sensors, including image sensors, voice sensors, text input devices, and biometric sensors, and transmits it to a server. The server stores the received data and analyzes it using multimodal artificial intelligence (AI) to monitor the user's health condition. If an abnormality is detected, the server also sends a notification to the user.

[0558] To implement this system, the following hardware and software are required: image sensors (e.g., security cameras), audio sensors (e.g., smart speakers), text input devices (e.g., smart assistants or dedicated applications), and biometric sensors (e.g., wearable devices) installed in the home. The server requires storage to receive and store this data and computing power to run AI models for data analysis.

[0559] This system collects, transmits, analyzes, monitors, and notifies data in the following ways:

[0560] As a concrete example, let's take the case of detecting abnormal heart rates at night. The wearable device, which is the terminal, continuously measures the user's heart rate at night. The measured heart rate data is sent to a server at regular intervals using an encryption protocol. The server stores the received heart rate data and begins analysis. The main process of the analysis is to identify when the heart rate deviates from the normal range.

[0561] If an abnormality is detected, the server generates a notification stating, "A sudden increase in heart rate was detected overnight." The notification includes a graph of heart rate variability and a recommendation to seek medical advice. The notification is sent to the user's smartphone or smart assistant, allowing the user to take appropriate action quickly.

[0562] Additionally, here are some examples of input prompts for the generative AI model:

[0563] Prompt using natural language processing technology: "How were you feeling yesterday? Please let me know if there was anything in particular that bothered you."

[0564] Prompt when using voice recognition technology: "Good morning. How long did you sleep last night?"

[0565] In this way, the present invention makes it possible to monitor the user's health condition in real time at home and identify health risks at an early stage by comprehensively analyzing various data.

[0566] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0567] Step 1:

[0568] Data collection

[0569] The device collects data from various sensors in the home (image sensors, voice sensors, text input devices, and biometric sensors). Specifically, the following operations are performed:

[0570] The image sensor uses a camera to periodically capture image data and video data.

[0571] Input: Real-time video data

[0572] Output: Captured image and video data

[0573] The audio sensor collects the user's voice at any time.

[0574] Input: Real-time audio data

[0575] Output: Collected audio data

[0576] The text input device captures text data entered by the user.

[0577] Input: Text entered by the user

[0578] Output: The retrieved text data

[0579] Biometric sensors measure vital data such as heart rate, blood pressure, oxygen saturation, and sleep patterns.

[0580] Input: Anthropometric data

[0581] Output: Measured biological data

[0582] Step 2:

[0583] Sending data to the server

[0584] The device periodically sends the collected data to the server. Specifically, the following operations are performed:

[0585] The device encrypts the collected data using security protocols (e.g., TLS / SSL).

[0586] Input: Collected data (images, audio, text, biometric data)

[0587] Output: Encrypted data

[0588] The encrypted data is sent to a server over the Internet.

[0589] Input: Encrypted data

[0590] Output: Data sent to server completed

[0591] Step 3:

[0592] Data storage

[0593] The server stores the received data. Specifically, the following operations are performed:

[0594] The received data is saved in the database.

[0595] Input: Received data (decrypted)

[0596] Output: Saved data

[0597] Step 4:

[0598] Data analysis

[0599] The server analyzes the stored data using multimodal artificial intelligence. Specifically, the following operations are performed:

[0600] Image data analysis: Using computer vision technology, image and video data is analyzed to detect changes in the user's movements and facial expressions.

[0601] Input: Image and video data

[0602] Output: Detection results of changes in behavior and facial expressions

[0603] Voice data analysis: Apply voice recognition technology to analyze changes in voice tone and speaking patterns.

[0604] Input: Audio data

[0605] Output: Detected changes in tone of voice and speech patterns

[0606] Text data analysis: Using natural language processing techniques, we analyze the content and expression patterns of text.

[0607] Input: Text data

[0608] Output: Analysis results of text content and stress symptoms

[0609] Biometric data analysis: Perform time series analysis to detect abnormalities in heart rate, blood pressure, and sleep patterns.

[0610] Input: Biometric data

[0611] Output: Abnormality detection results for biometric data

[0612] Step 5:

[0613] Health monitoring

[0614] The server integrates the analysis results and monitors the user's health status. Specifically, the following operations are performed:

[0615] The results of various data analyses are integrated and a comprehensive health assessment is performed using multimodal AI.

[0616] Input: Detection results of each analysis data

[0617] Output: Overall health assessment results

[0618] Step 6:

[0619] Generate and send notifications when anomalies occur

[0620] If the server detects an abnormality, it generates a notification and sends it to the user. Specifically, the following operations are performed:

[0621] A notification message is generated based on the anomaly detection.

[0622] Input: Anomaly detection result

[0623] Output: The generated notification message

[0624] Send notification messages to users' smartphones or smart assistants.

[0625] Input: Notification message

[0626] Output: Notification sent to user

[0627] This processing flow enables the system to monitor the user's health condition in real time and detect abnormalities early.

[0628] (Application example 1)

[0629] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0630] While self-driving vehicles are becoming increasingly common in modern society, there is currently no system that monitors the driver's health status in real time and responds promptly when an abnormality occurs. This may increase the risk of accidents due to driver health abnormalities. The present invention aims to solve this problem and provide an effective means for improving the safety of self-driving vehicles.

[0631] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0632] In this invention, the server includes means for collecting data from sensors used in the home, means for transmitting the collected data to the server, means for receiving and storing the transmitted data, means for analyzing the stored data using multimodal artificial intelligence, means for integrating the analysis results to monitor the health condition, means for generating and sending a notification to the user when an abnormality is detected, and means for collecting data from internal sensors of the autonomous vehicle and monitoring the driver's health condition. This makes it possible to monitor the driver's health condition in the vehicle in real time and send an immediate notification when an abnormality is detected.

[0633] "Domestic sensors" are devices installed inside the home that collect image, audio, text, biometric data, etc.

[0634] "Means for transmitting to the server" refers to a method or device for transmitting collected data to the server via a communication line such as the Internet.

[0635] "Storage means" refers to a method or device for retaining data received by the server for an extended period of time.

[0636] "Multimodal AI" refers to AI that has the ability to analyze multiple types of data (images, audio, text, biometric data, etc.) and integrate them.

[0637] "Health monitoring means" means a method or device that uses multimodal artificial intelligence to monitor analysis results and continuously assess a user's health status.

[0638] "Means for generating and sending a notification to the user when an abnormality is detected" refers to a method or device that automatically generates a warning or alert when an abnormality is discovered from the analysis results and forwards it to the user.

[0639] "Internal sensors in autonomous vehicles" refer to sensors that are placed inside autonomous vehicles and collect biometric and behavioral data from the driver.

[0640] "Means for monitoring the driver's health condition" refers to methods or devices that analyze data obtained from sensors in an autonomous vehicle and monitor the driver's health condition in real time.

[0641] The present invention provides a system for monitoring the health status of a user and a driver by comprehensively analyzing data from multiple sensors used in homes and autonomous vehicles. Specific embodiments are described below.

[0642] System Overview

[0643] The system collects data from multiple sensors (image sensors, voice sensors, text input devices, and biometric sensors) installed in the user's home and autonomous vehicle and sends it to a server. The server stores the received data and analyzes it using multimodal artificial intelligence (AI) to monitor the health status of the user and driver. If an abnormality is detected, the server generates a notification and sends it to the user or driver.

[0644] Hardware and software used

[0645] Hardware:

[0646] Home sensors (cameras, microphones, biometric sensors)

[0647] Internal sensors in autonomous vehicles (heart rate sensors, blood pressure sensors, cameras, etc.)

[0648] software:

[0649] Data Collection Applications

[0650] Communication protocol (HTTP, HTTPS)

[0651] Server (for data storage and analysis)

[0652] Multimodal AI analysis engine

[0653] Details of data processing and calculation

[0654] Data collection:

[0655] The devices (in-home sensors and in-vehicle sensors) collect image data, voice data, text data, biometric data, etc. This data is sent to a server at specific intervals.

[0656] Data transmission:

[0657] The device encrypts the collected data and sends it to the server using HTTPS, ensuring data security.

[0658] Data Analysis:

[0659] The server stores the received data and analyzes it using multimodal AI. The specific analysis content is as follows:

[0660] Image data: Analyzes changes in the user's or driver's behavior and facial expressions to detect abnormalities.

[0661] Voice data: Analyzed using voice recognition technology to detect changes in tone of voice and speaking patterns.

[0662] Text data: Analyze using natural language processing techniques to detect text content and signs of stress.

[0663] Biometric data: Time series analysis of heart rate, blood pressure abnormalities, and disrupted sleep patterns.

[0664] Anomaly detection and notification:

[0665] Multimodal AI integrates the results of various data analyses to assess the health status of the user and driver. If an abnormality is detected, the server generates a notification message and sends it to the user's smartphone or vehicle's infotainment system. The notification includes details of the abnormality and recommended actions to take.

[0666] Specific examples

[0667] For example, if a biometric sensor detects an abnormal heart rate while a driver of an autonomous vehicle is in the vehicle, the sensor will send this information to a server in real time. The server will analyze the heart rate and, if an abnormality is detected, will send a notification to the driver stating, "Your heart rate is outside the normal range. Please consult a medical professional."

[0668] Example prompts for generative AI models

[0669] "Please demonstrate the process flow of an application that analyzes data collected from internal sensors in an autonomous vehicle and monitors the driver's health. The implementation language should be Python, and the following sensors and data types should be assumed:

[0670] image

[0671] audio

[0672] text

[0673] Biometric data (heart rate, blood pressure)

[0674] Also include the ability to send data to a server and provide notifications based on the analysis results.

[0675] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0676] Step 1:

[0677] Devices (sensors in homes and vehicles) collect data. Specifically, cameras capture image and video data, microphones record voice data, text input devices acquire user input text, and biometric sensors measure biometric data such as heart rate and blood pressure. All of this data is collected from each sensor at regular intervals.

[0678] Input: Image data, audio data, text data, and biometric data obtained from each sensor.

[0679] Output: Consolidated sensor data.

[0680] Step 2:

[0681] The terminal sends the collected data to the server. The collected data is encrypted using a communication protocol (e.g., HTTPS) and sent securely to the server. The data is sent at regular intervals.

[0682] Input: Integrated sensor data.

[0683] Output: The data sent to the server.

[0684] Step 3:

[0685] The server stores the received data. The server stores the received data in a database for later analysis. During this storage process, each data type (image, audio, text, biometric data) is recorded in chronological order.

[0686] Input: The data sent to the server.

[0687] Output: Data stored in the database.

[0688] Step 4:

[0689] The server analyzes the stored data using multimodal artificial intelligence. Specifically, image data is analyzed using computer vision technology to detect movements and facial expressions. Audio data is analyzed using voice recognition technology to detect tone of voice and speech patterns. Text data is analyzed using natural language processing technology to detect text content and signs of stress. Biometric data is analyzed over time to detect abnormalities in heart rate and blood pressure.

[0690] Input: Data stored in a database.

[0691] Output: Analysis results.

[0692] Step 5:

[0693] The server integrates the analysis results and monitors the health of the user and driver. Multimodal AI integrates the analysis results of each data type to assess the overall health status. From this analysis, anomalies are identified.

[0694] Input: Analysis results for each data type.

[0695] Output: Integrated health status assessment.

[0696] Step 6:

[0697] If the server detects an abnormality, it generates a notification and sends it to the user or driver. The server generates a notification message containing the details of the abnormality and recommended actions to take, and sends it to the user's smartphone or the vehicle's infotainment system.

[0698] Input: Integrated Health Status Assessment.

[0699] Output: Informational message.

[0700] Step 7:

[0701] The user or driver receives a notification and takes appropriate action, such as consulting a medical professional.

[0702] Input: Notification message.

[0703] Output: Coping behavior.

[0704] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0705] The present invention is a system that monitors a user's health condition by comprehensively analyzing data from multiple sensors used in the home. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it is possible to monitor the user's emotional health as well. Specific embodiments of the system will now be described.

[0706] System Overview

[0707] The system collects data from multiple sensors (image sensors, voice sensors, text input devices, and biometric sensors) installed in the user's home and sends it to a server. The server stores the received data and analyzes it using a multimodal artificial intelligence (AI) and emotion engine to monitor the user's health and emotional state. If an abnormality is detected, the server sends a notification to the user.

[0708] Program processing flow

[0709] Data collection

[0710] Devices (users' home sensors and wearable devices) collect data

[0711] Image data: Security cameras in the home periodically take images and videos, recording the user's movements, facial expressions, and daily life.

[0712] Voice data: Smart speakers and smart assistants collect your voice and use the recorded audio to analyze your tone of voice and speaking patterns.

[0713] Text data: Collects text data entered by users through smart assistants or dedicated applications, and analyzes changes in the content and expression of the text.

[0714] Biometric data: Wearable devices measure biometric data such as heart rate, blood pressure, oxygen saturation, and sleep patterns to continuously monitor the user's physical condition.

[0715] Data transmission

[0716] The device sends the collected data to the server.

[0717] The sensor terminals send the collected data to the server at regular intervals, and the data is encrypted using a security protocol.

[0718] Data analysis

[0719] The server analyzes the received data

[0720] Image analysis: The server uses image data to analyze changes in the user's movements and facial expressions, thereby detecting physical changes and patterns in daily activities.

[0721] Voice analysis: The server uses speech recognition technology to convert voice data into text and analyzes changes in tone and speaking style to identify emotional issues and changes in cognitive function.

[0722] Text analysis: The server uses natural language processing technology to analyze the text data and detect changes in the content and expressions of the text, which allows for an assessment of stress and mental health.

[0723] Biometric data analysis: The server uses time series analysis techniques to analyze data such as heart rate, blood pressure, oxygen saturation, and sleep patterns to detect abnormalities, allowing for early detection of cardiac and respiratory problems.

[0724] Emotion analysis

[0725] The server uses an emotion engine to analyze the user's emotions.

[0726] Emotion analysis using voice data: The server analyzes the voice data and uses an emotion engine to recognize the user's emotional state from changes in their tone of voice and speaking style.

[0727] Emotion analysis using text data: The server analyzes the text data and uses an emotion engine to recognize the emotional state from changes in the content and expression of the text.

[0728] Emotion analysis using image data: The server analyzes the image data and uses an emotion engine to recognize the user's emotional state from their facial expressions.

[0729] Health and emotional state monitoring

[0730] The server integrates the analysis results and monitors the health and emotional state.

[0731] By integrating the results of various data analyses, the multimodal AI and emotion engine assesses the user's overall health and emotional state, making it easier to detect abnormalities.

[0732] Send notifications

[0733] If the server detects an abnormality, it will notify the user.

[0734] If the server detects an anomaly, it generates a notification message and sends it to the user's smartphone or smart assistant, containing details of the anomaly and recommended actions to take.

[0735] Example: Emotion and health monitoring using voice data

[0736] The device (smart speaker) collects the user's voice from everyday conversations. The server analyzes the received voice data and converts it into text data using voice recognition technology. It then uses an emotion engine to analyze subtle changes in voice tone and speaking style to assess whether the user is feeling stressed. The server also uses multimodal AI to analyze heart rate data collected simultaneously with the voice data, detecting physical changes such as an abnormally high heart rate. When an abnormality is detected, the server immediately generates a notification and sends it to the user's smartphone, stating, "Your recent conversations have shown signs of stress, and your heart rate is also high. Please try some relaxation techniques." This allows for a prompt response.

[0737] In this way, the present invention can protect the overall health of users by monitoring both health and emotions in the home in real time and identifying risks early.

[0738] The processing flow will be explained below.

[0739] Step 1:

[0740] Devices (sensors) collect data. Home security cameras periodically capture images and videos, smart speakers and smart assistants record voice data, and text data entered by users through smart assistants and applications is also collected. Wearable devices also measure biometric data such as heart rate, blood pressure, oxygen saturation, and sleep patterns in real time.

[0741] Step 2:

[0742] The terminal (sensor) packetizes and encrypts the collected data, and the data packets are temporarily stored in a queue inside the terminal.

[0743] Step 3:

[0744] The device (sensor) sends data packets to the server at regular intervals. The data transmission is performed using a secure communication protocol to ensure the confidentiality and security of the data.

[0745] Step 4:

[0746] The server receives the data packet. The received data packet is decoded and restored to its original format. The data integrity is checked, and if there are any errors or omissions, a retransmission request is sent to the terminal.

[0747] Step 5:

[0748] The server stores the recovered data in a distributed database, where the data is managed for each user and tagged within the database.

[0749] Step 6:

[0750] The server analyzes the received image and video data. Image classification and motion analysis algorithms are used to detect changes in the user's movement patterns and facial expressions. For example, unusual movements or changes in facial expressions may be signs of an abnormality.

[0751] Step 7:

[0752] The server analyzes the voice data. It uses speech recognition technology to convert the data into text and analyzes changes in tone and speaking style. This allows it to detect emotional issues and changes in cognitive function. For example, a trembling voice or changes in speaking style could be a sign of stress or changes in cognitive function.

[0753] Step 8:

[0754] The server analyzes the text data. It uses natural language processing technology to analyze the content of the text and detect changes in the content and expressions of the sentences. This allows for an assessment of stress and mental health. For example, an increase in negative expressions in the sentences can be determined to be a sign of stress.

[0755] Step 9:

[0756] The server analyzes biometric data, using time series analysis techniques to analyze data such as heart rate, blood pressure, oxygen saturation, and sleep patterns to detect abnormalities. This allows for early detection of cardiac and respiratory problems. For example, a higher-than-normal heart rate could be a sign of cardiac disease.

[0757] Step 10:

[0758] The server uses an emotion engine to analyze the user's emotions. It recognizes the user's emotions from voice, text, and image data. This allows the system to monitor the user's emotional health. For example, if anger, excitement, or sadness is detected from the voice data, it is considered a sign of emotional instability.

[0759] Step 11:

[0760] The server integrates the results of each analysis, and uses a multimodal AI and emotion engine to comprehensively analyze images, voice, text, and biometric data to assess the user's overall health and emotional state, making it easier to detect abnormalities.

[0761] Step 12:

[0762] If the server detects an anomaly, it generates a notification with details about the anomaly and a recommended course of action, such as "Your heart rate has been higher than normal recently, and you appear to be feeling stressed."

[0763] Step 13:

[0764] The server generates a notification and sends it to the user's smartphone or smart assistant. The notification is sent immediately, informing the user that an abnormality has occurred.

[0765] Step 14:

[0766] Users receive notifications, review them, take action based on recommended actions, and, if necessary, seek medical advice. This allows users to quickly respond to changes in their health and emotional state.

[0767] This specific processing flow enables real-time monitoring of both health and emotions within the home, identifying risks early and protecting the overall health of users.

[0768] Example 2

[0769] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0770] There is a need for technology that can comprehensively monitor not only a user's health status but also their emotional state by collecting and analyzing different types of data from multiple sensors used in the home. However, existing systems have faced many challenges in centralized data management, complex data analysis, anomaly detection, and prompt notification to the user. In particular, it has been difficult to integrate information from multiple data sources and perform accurate analysis.

[0771] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for collecting data from sensors used in the home, means for transmitting the collected data to the server, means for receiving and storing the transmitted data, means for analyzing the stored data using multimodal artificial intelligence and an emotion analysis engine, means for integrating the analysis results to monitor the health status and emotional state, and means for generating and sending a notification to the user when an abnormality is detected. This makes it possible to comprehensively monitor the user's health status and emotional state in real time and to provide a prompt and appropriate notification when an abnormality occurs.

[0772] A "sensor" is a device used in the home that collects information such as images, audio, text, and biometric data.

[0773] "Data" refers to all information collected by sensors, such as image data, video data, audio data, text data, and biometric data.

[0774] A "server" is an information processing device that receives and stores data sent from sensors, and analyzes the data to monitor the user's condition.

[0775] "Multimodal AI" is an AI technology that has the ability to integrate and analyze multiple different types of data (e.g., images, audio, text, biometric data) in a unified manner.

[0776] An "emotion analysis engine" is a technology for analyzing a user's emotional state from data such as voice, text, and images.

[0777] "Analyzing" means that the server processes the data it receives and derives meaningful results.

[0778] "Monitoring" means that the server continuously observes and evaluates the user's health and emotional state based on the analysis results.

[0779] A "notification" is a warning or advice message that is generated and sent to the user when the server detects an abnormality.

[0780] An "anomaly" is any unusual change or problem in the user's health or emotional state that is detected as a result of the server's analysis.

[0781] "Integrating" means that the server compiles the results of multiple data analyses and provides a single comprehensive evaluation.

[0782] The present invention provides a system for monitoring a user's health and emotional state by comprehensively analyzing data from multiple sensors used in the home. A specific method for implementing this system is described below.

[0783] System configuration

[0784] The system collects data from multiple sensors installed in the home (image sensors, voice sensors, text input devices, and biometric sensors) and sends it to a server. The server stores the received data and analyzes it using a multimodal artificial intelligence (AI) and emotion analysis engine. The analysis results are integrated to monitor the user's health and emotional state. If an abnormality is detected, the server sends a notification to the user.

[0785] Hardware and software used

[0786] Sensors used in the home include:

[0787] Image sensor: Home security camera

[0788] Voice Sensors: Smart Speakers and Smart Assistants

[0789] Text input devices: smart assistants and dedicated applications

[0790] Biosensors: Wearable devices that measure heart rate, blood pressure, oxygen saturation, and sleep patterns

[0791] The server includes the following analysis modules:

[0792] Multimodal Artificial Intelligence (AI)

[0793] Sentiment Analysis Engine

[0794] Data Storage and Management System

[0795] Notification System

[0796] Data collection and transmission

[0797] 1. Devices (sensors) collect data

[0798] Image sensors capture images and videos of the home and record changes in the user's movements and facial expressions. For example, a camera in the living room can start recording at a set time.

[0799] The voice sensor uses smart speakers to collect everyday conversations and analyze tone and speech patterns, for example, recording the voice you use to ask about the weather.

[0800] The text input device collects text entered by the user through a dedicated application and analyzes changes in the content and expression of the text, for example, collecting diary entries from the user.

[0801] Biometric sensors use wearable devices to continuously measure heart rate, blood pressure, oxygen saturation, and sleep patterns. For example, a smartwatch worn by a user records heart rate.

[0802] 2. The device sends the collected data to the server

[0803] The sensor terminals transmit the collected data to a server at regular intervals. The transmitted data is encrypted using a security protocol. For example, audio data is uploaded to the server every hour.

[0804] Data analysis

[0805] 1. The server analyzes the received data

[0806] Image analysis: The server uses image data to analyze changes in the user's movements and facial expressions, for example, to detect changes in the user's walking pattern.

[0807] Voice analysis: The server uses speech recognition technology to convert the voice data into text and analyzes changes in tone and speaking style, for example detecting if the user's voice is darker than usual.

[0808] Text analysis: The server uses natural language processing technology to analyze the text data and detect changes in the content and expressions of the sentences. For example, it analyzes the frequent use of the expression "tired."

[0809] Biometric data analysis: The server uses time series analysis to analyze heart rate, blood pressure, oxygen saturation, sleep patterns, etc. to detect abnormalities, such as a higher than normal heart rate.

[0810] 2. Emotion analysis

[0811] Emotion analysis from voice data: The server analyzes the voice data and uses an emotion analysis engine to recognize the emotional state from changes in tone of voice and speaking style, for example, detecting an increase in anger.

[0812] Sentiment analysis based on text data: The server analyzes the text data and uses a sentiment analysis engine to recognize the emotional state from changes in the content and expressions of the text. For example, it detects an increase in pessimistic expressions.

[0813] Emotion analysis based on image data: The server analyzes the image data and uses an emotion analysis engine to recognize the user's emotional state from their facial expressions. For example, it detects that they are smiling less.

[0814] Integrated monitoring and abnormality notification

[0815] 1. The server integrates the analysis results and monitors the health and emotional state.

[0816] Using a multimodal AI and emotion analysis engine, the analysis results are integrated to assess the user's overall health and emotional state. For example, if different emotional signals are detected from image analysis, audio analysis, and text analysis, a comprehensive assessment is made.

[0817] 2. Detect anomalies and generate notifications

[0818] If an anomaly is detected, the server generates a real-time notification. For example, if heart rate and voice data indicate strong signs of stress, a notification will be sent to the user stating, "Your recent conversations have shown signs of stress, and your heart rate is high. Please try some relaxation techniques."

[0819] Specific examples

[0820] For example, a smart speaker collects a user's voice from everyday conversations and sends the data to a server. The server analyzes the voice data and converts it into text data using speech recognition technology. It uses an emotion analysis engine to analyze subtle changes in the tone of voice and speaking style to assess whether the user is feeling stressed. At the same time, the server also analyzes heart rate data from a wearable device and detects if the heart rate is higher than normal. If an abnormality is detected, the server sends the user a notification saying, "Your recent conversations have shown signs of stress, and your heart rate is high. Please try some relaxation techniques."

[0821] Example prompt: "Analyze the voice data collected by the smart speaker and assess whether the user is experiencing stress."

[0822] This invention can protect the overall health of users by monitoring both health and emotions in the home in real time and identifying risks early.

[0823] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0824] Step 1: Collect data

[0825] Data is collected by sensors and devices installed in the home by users.

[0826] Input: Various data from the home environment and the user (images, voice, text, biometric data).

[0827] How it works: Security cameras take pictures, smart speakers record everyday conversations, and wearable devices measure heart rate and blood pressure.

[0828] Output: Various data collected.

[0829] Step 2: Sending data

[0830] The terminal transmits the collected data to the server.

[0831] Input: Various types of collected data (images, audio, text, biometric data).

[0832] Specific operation: Image data captured by the security camera is uploaded to the server, and audio data recorded by the smart speaker is sent to the server.

[0833] Output: The data sent to the server.

[0834] Step 3: Save your data

[0835] The server stores the received data.

[0836] Input: The data sent to the server.

[0837] Specific operation: The server stores various data in the database.

[0838] Output: Saved data.

[0839] Step 4: Analyzing the image data

[0840] The server analyzes the image data.

[0841] Input: Saved image data.

[0842] Specific operation: Using multimodal AI to analyze changes in user behavior and facial expressions from image data, for example, using a facial recognition algorithm to analyze user facial expressions.

[0843] Output: Analysis results (changes in user's behavior and facial expressions).

[0844] Step 5: Analyzing the audio data

[0845] The server analyzes the audio data.

[0846] Input: Stored audio data.

[0847] What it does: It uses speech recognition technology to convert voice data into text, and uses an emotion analysis engine to analyze changes in tone and speaking style, for example, to identify changes in tone in a user's everyday conversation.

[0848] Output: Analysis results (changes in tone of voice and speaking style).

[0849] Step 6: Analyze the text data

[0850] The server parses the text data.

[0851] Input: Saved text data.

[0852] Specific operation: Analyzes text data using natural language processing technology and detects changes in the content and expressions of sentences using a sentiment analysis engine. For example, it analyzes the frequent use of the expression "tired."

[0853] Output: Analysis results (changes in the content and expression of the text).

[0854] Step 7: Analyzing the biometric data

[0855] The server analyzes the biometric data.

[0856] Input: Stored biometric data.

[0857] Specific operation: Using time series analysis technology, it analyzes heart rate, blood pressure, oxygen saturation, sleep patterns, etc. to detect abnormalities, such as when the heart rate is higher than normal.

[0858] Output: Analysis results (heart rate, blood pressure, oxygen saturation, changes in sleep patterns).

[0859] Step 8: Integrating the analysis results

[0860] The server consolidates the various analysis results.

[0861] Input: Analysis results of various data (images, audio, text, biometric data).

[0862] How it works: Using a multimodal AI and emotion analysis engine, the system integrates analysis results to assess the user's overall health and emotional state. For example, if different emotional indicators are found from image analysis, audio analysis, and text analysis, it will comprehensively assess them.

[0863] Output: Integrated analysis results (assessment of the user's health and emotional state).

[0864] Step 9: Detect anomalies

[0865] The server detects an abnormality.

[0866] Input: Consolidated analysis results.

[0867] Specific operation: Based on the analysis results, an anomaly detection algorithm is executed to detect abnormalities in the user's health and emotional state, such as stress and heart rate abnormalities.

[0868] Output: Anomaly detection results.

[0869] Step 10: Generate and send notifications

[0870] The server generates the notification and sends it to the user.

[0871] Input: Anomaly detection results.

[0872] What it does: If an abnormality is detected, it will generate a notification message and send it to the user's smartphone or smart assistant, for example, "Your recent conversations have shown signs of stress and your heart rate is elevated. Please try some relaxation techniques."

[0873] Output: The notification message sent to the user.

[0874] (Application example 2)

[0875] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0876] Conventional health monitoring systems in homes and businesses are limited to monitoring physical health conditions, making it difficult to respond quickly and accurately to changes or abnormalities in emotional states. Furthermore, the means of detecting abnormalities are limited, and detection of abnormalities can be delayed under certain conditions. Therefore, there is a need for a system that can monitor both health and emotional states in an integrated manner, quickly detect abnormalities, and provide appropriate notifications.

[0877] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting data from sensors used in homes and business facilities, means for transmitting the collected data to the server, means for receiving and storing the transmitted data, means for analyzing the stored data using multimodal artificial intelligence, means for integrating the analysis results to monitor health and emotional states, means for generating and transmitting a notification to the user when an abnormality is detected, and means for analyzing changes in the movements, facial expressions, and tone of voice of visitors and employees to identify their emotional states. This makes it possible to monitor not only health states but also emotional states in real time, quickly detect abnormalities, and issue appropriate notifications.

[0878] "Inside a home or business premises" refers to the physical environment in which a person lives or conducts business activities, and refers to the location of the system used within that particular space.

[0879] A "sensor" is a device that detects physical phenomena and converts them into electronic data, and includes image sensors, audio sensors, text input devices, and biometric sensors.

[0880] "Means for collecting data" refers to a system element that uses sensors to obtain information from an object and record it in electronic form.

[0881] "Means for transmitting data to a server" refers to the communications devices and protocols used to transfer the collected digital data over a network to a remote server.

[0882] "Means for receiving and storing data" refers to a system element that allows a server to receive data sent via a network and store it in a database or the like.

[0883] "Multimodal AI" refers to an AI technology that simultaneously analyzes multiple types of data (e.g., images, audio, text, biometric data) and makes integrated decisions.

[0884] "Health and emotional monitoring means" refers to system elements that provide continuous monitoring of an individual's physical and emotional health based on collected and analyzed data.

[0885] "Means for generating and sending notifications to users when an anomaly is detected" refers to the ability for the system to automatically create and send alerts or notifications to users' devices when an anomaly is detected.

[0886] "Means of identifying emotional states by analyzing changes in behavior, facial expressions, and tone of voice" refers to artificial intelligence technology that analyzes image and audio data to infer emotional states from behavior, changes in facial expressions, tone of voice, etc.

[0887] An embodiment of the present invention is described in detail below: The system monitors the health and emotional state of visitors and employees by collecting data from multiple sensors installed in homes and businesses, transmitting the data to a server for analysis.

[0888] System Configuration

[0889] Terminal (sensor)

[0890] The devices include image sensors, voice sensors, text input devices, and biometric sensors that are placed within physical stores to collect the following data from visitors and employees:

[0891] Image data: The camera records your movements and expressions.

[0892] Audio data: Microphones collect speech.

[0893] Text data: The digital device captures the input text.

[0894] Biometric data: Wearable devices and thermal cameras collect heart rate and temperature data.

[0895] The data collected by each terminal is encrypted and sent to a server at regular intervals.

[0896] Server Processing

[0897] Data reception and storage

[0898] The server receives the data sent from the terminal via the network in an encrypted form and stores it in a database, using technologies including security protocols and database systems.

[0899] Data analysis

[0900] The server analyzes the various types of stored data (images, audio, text, and biometric data) using multimodal artificial intelligence (AI). Image data is analyzed for changes in movements and facial expressions using libraries such as OpenCV, and audio data is converted into text using pyaudio and librosa, after which changes in tone of voice and speaking style are analyzed. Text data is analyzed for emotional state using natural language processing techniques such as TextBlob. Furthermore, biometric data is analyzed for abnormalities using time series analysis techniques.

[0901] Emotion analysis

[0902] The server uses an emotion engine to integrate data from different modalities and recognize the emotional state of visitors and employees, allowing real-time monitoring of changes in not only physical health but also emotional health.

[0903] Notifications and Feedback

[0904] If the server detects an abnormality, it automatically generates a notification message and sends it to the user's smartphone or other display device. The notification includes details of the detected abnormality and recommended actions to take. For example, a message might be generated that reads, "Employee A has a high body temperature and his tone of voice indicates stress. We recommend that he leave work early."

[0905] Examples and prompts

[0906] As a concrete example, consider a monitoring system for Employee A in a store. The server collects and analyzes Employee A's temperature and conversation data, and if an abnormality is detected, it sends a notification saying, "Employee A's temperature is high and his / her tone of voice also indicates stress. We recommend that you leave early." An example of how such a prompt is input into a generative AI model is as follows:

[0907] "Monitor employee health and emotions, and if you detect a high temperature or stressful tone in the conversation, generate a notification recommending that employees leave early. Could you please give us an example of what this notification should look like?"

[0908] This will enable the realization of a system that monitors health and emotions in an integrated manner and can respond quickly in the event of an abnormality.

[0909] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0910] Step 1:

[0911] The terminal collects image data, audio data, text data, and biometric data. Image data is obtained by recording the movements and facial expressions of visitors and employees with a camera, while audio data is obtained by collecting conversations and environmental sounds with a microphone. Text data is obtained by acquiring input information from digital terminals, and biometric data is obtained by collecting heart rate and body temperature information with a wearable device or a body temperature detection camera. The input is various data obtained from each sensor, and the output is the collected raw data.

[0912] Step 2:

[0913] The terminal sends the collected data to the server at regular intervals. The data is sent encrypted over the network. In this step, the input is the collected data, and the output is the encrypted data sent to the server.

[0914] Step 3:

[0915] The server saves the received data. The received data is stored in the server's database. The input of this step is the encrypted transmitted data, and the output is the saved data.

[0916] Step 4:

[0917] The server analyzes the stored data. The analysis is divided into the following parts:

[0918] Image analysis: Image data is analyzed using OpenCV etc. to detect changes in movements and facial expressions. The input is the saved image data, and the output is the analysis results of movement and facial expression changes.

[0919] Speech analysis: Using pyaudio and librosa, audio data is converted to text and changes in voice tone and speaking style are analyzed. The input is the stored audio data, and the output is the converted text data and changes in voice tone and speaking style.

[0920] Text Analysis: Analyzes text data using natural language processing techniques such as TextBlob to evaluate emotional states. The input is the stored text data, and the output is the analysis result of the emotional state.

[0921] Biometric data analysis: Detects abnormalities in biometric data (heart rate, body temperature, etc.) using time series analysis techniques. The input is stored biometric data, and the output is the anomaly detection results.

[0922] Step 5:

[0923] The server uses the emotion engine to integrate the analysis results of each data and monitor the health and emotional states of visitors and employees. The input of this step is each analysis result, and the output is the integrated evaluation result of the health and emotional states.

[0924] Step 6:

[0925] If the server detects an abnormality, it generates a notification message and sends it to the user's smartphone or other display device. The notification includes details of the abnormality and recommended actions to take. For example, a notification might be sent stating, "Employee A's temperature is high and his / her tone of voice indicates stress. We recommend that he / she leave work early." The input to this step is the integrated evaluation result, and the output is the generated notification message.

[0926] The above is the flow of processing in the embodiment of the present invention.

[0927] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0928] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0929] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0930] [Third embodiment]

[0931] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0932] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0933] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0934] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0935] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0936] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0937] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0938] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0939] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0940] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0941] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0942] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0943] The present invention is a system for monitoring the health condition of a user by comprehensively analyzing data from multiple sensors used in the home. A specific embodiment of the system will be described.

[0944] System Overview

[0945] The system collects data from multiple sensors (image sensors, voice sensors, text input devices, and biometric sensors) installed in the user's home and sends it to a server. The server stores the received data, analyzes it using multimodal artificial intelligence (AI), and monitors the user's health status. If an abnormality is detected, the server sends a notification to the user.

[0946] Program processing flow

[0947] Data collection

[0948] Devices (users' home sensors and wearable devices) collect data

[0949] Image data: Security cameras in the home periodically take images and videos, recording the user's movements, facial expressions, and daily life.

[0950] Voice data: Smart speakers and smart assistants collect your voice and use it to detect changes in your tone and speaking style.

[0951] Text data: Collects text data entered by users through smart assistants or dedicated applications, and analyzes the content of the text and signs of stress.

[0952] Biometric data: Wearable devices measure biometric data such as heart rate, blood pressure, oxygen saturation, and sleep patterns, thereby continuously monitoring the user's physical condition.

[0953] Data transmission

[0954] The device sends the collected data to the server.

[0955] The sensor terminals send the collected data to the server at regular intervals, and the data is encrypted using a security protocol.

[0956] Data analysis

[0957] The server analyzes the received data

[0958] The server uses the image data to analyze changes in the user's movements and facial expressions. For example, unusual movements may indicate a health problem.

[0959] Voice data is analyzed using speech recognition technology to detect changes in tone of voice and speech patterns, which can identify emotional issues and changes in cognitive function.

[0960] The text data is analyzed using natural language processing technology to detect changes in the content and expressions of the text, which can then be used to assess stress and mental health.

[0961] Biometric data is analyzed over time to detect abnormalities in heart rate, blood pressure, and sleep patterns, allowing for early detection of cardiac and respiratory problems.

[0962] Health Status Monitoring

[0963] The server consolidates the analysis results and monitors health status.

[0964] By integrating the results of various data analyses, multimodal AI assesses the user's overall health status, making it easier to detect abnormalities.

[0965] Send notifications

[0966] If the server detects an abnormality, it will notify the user.

[0967] If the server detects an anomaly, it generates a notification message and sends it to the user's smartphone or smart assistant, containing details of the anomaly and recommended actions to take.

[0968] Example: Detecting abnormal heart rates at night

[0969] The terminal (wearable device) continues to measure the user's heart rate throughout the night. The server receives and stores the collected heart rate data. The server then analyzes the received data and detects abnormal increases in heart rate beyond the normal range. If an abnormality is detected, the server sends a notification to the user's smartphone stating, "A sudden increase in heart rate has been detected overnight." The user receives this notification and can take prompt action, such as consulting a medical institution.

[0970] In this way, the present invention can protect the user's health by monitoring the user's health status in real time at home and identifying health risks early.

[0971] The processing flow will be explained below.

[0972] Step 1:

[0973] Devices (sensors) collect data. Security cameras installed in homes periodically capture images and videos, and smart speakers and smart assistants record voice data. Text data entered by users through smart assistants and applications is also collected. In addition, wearable devices measure biometric data such as heart rate, blood pressure, oxygen saturation, and sleep patterns in real time.

[0974] Step 2:

[0975] The terminal (sensor) packetizes and encrypts the collected data, and the data packets are temporarily stored in a queue inside the terminal.

[0976] Step 3:

[0977] The device (sensor) sends data packets to the server at regular intervals. The data transmission is performed using a secure communication protocol to ensure the confidentiality and security of the data.

[0978] Step 4:

[0979] The server receives the data packet. The received data packet is decoded and restored to its original format. The data integrity is checked, and if there are any errors or omissions, a retransmission request is sent to the terminal.

[0980] Step 5:

[0981] The server stores the recovered data in a distributed database, where the data is managed for each user and tagged within the database.

[0982] Step 6:

[0983] The server analyzes the received image and video data, using image classification and motion analysis algorithms to detect changes in the user's movement patterns and facial expressions.

[0984] Step 7:

[0985] The server analyzes the audio data, converts it into text using speech recognition technology, and analyzes changes in tone and speaking style to identify signs associated with emotional problems and changes in cognitive function.

[0986] Step 8:

[0987] The server analyzes the text data and uses natural language processing technology to analyze the content of the text data and evaluate stress levels and health conditions.

[0988] Step 9:

[0989] The server analyzes biometric data, using time series analysis techniques to analyze data such as heart rate, blood pressure, oxygen saturation, and sleep patterns to detect abnormalities, especially values ​​that exceed the normal range.

[0990] Step 10:

[0991] The server integrates the results of each analysis, using multimodal AI to comprehensively analyze images, voice, text, and biometric data to assess the user's overall health.

[0992] Step 11:

[0993] The server generates a notification when an anomaly is detected, containing details about the anomaly and a recommended course of action.

[0994] Step 12:

[0995] The server generates a notification and sends it to the user's smartphone or smart assistant. The notification is sent immediately, informing the user that an abnormality has occurred.

[0996] Step 13:

[0997] Users will receive a notification, which they can review and follow the recommended actions, including seeking medical advice if necessary.

[0998] This processing flow allows for real-time health monitoring within the home, and if an abnormality is detected, it is possible to respond immediately.

[0999] Example 1

[1000] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1001] In modern society, people's lifestyles are becoming more diverse, and health management at home is becoming increasingly important. Daily health monitoring is especially important for the elderly and those with chronic illnesses. However, many current health monitoring systems are single-modal, analyzing only specific data, and lack the ability to comprehensively analyze different types of data. This makes it difficult to accurately grasp overall health status and detect abnormalities early.

[1002] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1003] In this invention, the server includes means for collecting data from sensors used in the home, means for transmitting the collected data to the server, means for receiving and storing the transmitted data, means for analyzing the stored data using multimodal artificial intelligence, means for integrating the analyzed data and monitoring the health status, and means for generating and sending a notification to the user when an abnormality is detected. This makes it possible to integrate and analyze information from multiple data sources, accurately monitor the user's overall health status, and detect abnormalities early.

[1004] A "sensor used in the home" is a device that operates in a user's home environment and collects one or more of image data, video data, audio data, text data, and / or biometric data.

[1005] The "server" is a device or system that receives and stores data sent from sensors in the home, analyzes the data using multimodal artificial intelligence, and integrates the results to monitor the user's health condition.

[1006] "Multimodal artificial intelligence" refers to artificial intelligence technology that comprehensively analyzes different types of data (images, audio, text, biometric data, etc.).

[1007] "Analysis results" refers to analytical information generated by multimodal artificial intelligence based on collected data, including data indicating the user's health status.

[1008] A "notification" is a message generated when the server detects an anomaly based on the analysis results, informing the user of the nature of the anomaly and the recommended course of action.

[1009] This invention is a system that monitors a user's health condition by integrating and analyzing data from multiple sensors used in the home. This system collects data from multiple sensors, including image sensors, voice sensors, text input devices, and biometric sensors, and transmits it to a server. The server stores the received data and analyzes it using multimodal artificial intelligence (AI) to monitor the user's health condition. If an abnormality is detected, the server also sends a notification to the user.

[1010] To implement this system, the following hardware and software are required: image sensors (e.g., security cameras), audio sensors (e.g., smart speakers), text input devices (e.g., smart assistants or dedicated applications), and biometric sensors (e.g., wearable devices) installed in the home. The server requires storage to receive and store this data and computing power to run AI models for data analysis.

[1011] This system collects, transmits, analyzes, monitors, and notifies data in the following ways:

[1012] As a concrete example, let's take the case of detecting abnormal heart rates at night. The wearable device, which is the terminal, continuously measures the user's heart rate at night. The measured heart rate data is sent to a server at regular intervals using an encryption protocol. The server stores the received heart rate data and begins analysis. The main process of the analysis is to identify when the heart rate deviates from the normal range.

[1013] If an abnormality is detected, the server generates a notification stating, "A sudden increase in heart rate was detected overnight." The notification includes a graph of heart rate variability and a recommendation to seek medical advice. The notification is sent to the user's smartphone or smart assistant, allowing the user to take appropriate action quickly.

[1014] Additionally, here are some examples of input prompts for the generative AI model:

[1015] Prompt using natural language processing technology: "How were you feeling yesterday? Please let me know if there was anything in particular that bothered you."

[1016] Prompt when using voice recognition technology: "Good morning. How long did you sleep last night?"

[1017] In this way, the present invention makes it possible to monitor the user's health condition in real time at home and identify health risks at an early stage by comprehensively analyzing various data.

[1018] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1019] Step 1:

[1020] Data collection

[1021] The device collects data from various sensors in the home (image sensors, voice sensors, text input devices, and biometric sensors). Specifically, the following operations are performed:

[1022] The image sensor uses a camera to periodically capture image data and video data.

[1023] Input: Real-time video data

[1024] Output: Captured image and video data

[1025] The audio sensor collects the user's voice at any time.

[1026] Input: Real-time audio data

[1027] Output: Collected audio data

[1028] The text input device captures text data entered by the user.

[1029] Input: Text entered by the user

[1030] Output: The retrieved text data

[1031] Biometric sensors measure vital data such as heart rate, blood pressure, oxygen saturation, and sleep patterns.

[1032] Input: Anthropometric data

[1033] Output: Measured biological data

[1034] Step 2:

[1035] Sending data to the server

[1036] The device periodically sends the collected data to the server. Specifically, the following operations are performed:

[1037] The device encrypts the collected data using security protocols (e.g., TLS / SSL).

[1038] Input: Collected data (images, audio, text, biometric data)

[1039] Output: Encrypted data

[1040] The encrypted data is sent to a server over the Internet.

[1041] Input: Encrypted data

[1042] Output: Data sent to server completed

[1043] Step 3:

[1044] Data storage

[1045] The server stores the received data. Specifically, the following operations are performed:

[1046] The received data is saved in the database.

[1047] Input: Received data (decrypted)

[1048] Output: Saved data

[1049] Step 4:

[1050] Data analysis

[1051] The server analyzes the stored data using multimodal artificial intelligence. Specifically, the following operations are performed:

[1052] Image data analysis: Using computer vision technology, image and video data is analyzed to detect changes in the user's movements and facial expressions.

[1053] Input: Image and video data

[1054] Output: Detection results of changes in behavior and facial expressions

[1055] Voice data analysis: Apply voice recognition technology to analyze changes in voice tone and speaking patterns.

[1056] Input: Audio data

[1057] Output: Detected changes in tone of voice and speech patterns

[1058] Text data analysis: Using natural language processing techniques, we analyze the content and expression patterns of text.

[1059] Input: Text data

[1060] Output: Analysis results of text content and stress symptoms

[1061] Biometric data analysis: Perform time series analysis to detect abnormalities in heart rate, blood pressure, and sleep patterns.

[1062] Input: Biometric data

[1063] Output: Abnormality detection results for biometric data

[1064] Step 5:

[1065] Health monitoring

[1066] The server integrates the analysis results and monitors the user's health status. Specifically, the following operations are performed:

[1067] The results of various data analyses are integrated and a comprehensive health assessment is performed using multimodal AI.

[1068] Input: Detection results of each analysis data

[1069] Output: Overall health assessment results

[1070] Step 6:

[1071] Generate and send notifications when anomalies occur

[1072] If the server detects an abnormality, it generates a notification and sends it to the user. Specifically, the following operations are performed:

[1073] A notification message is generated based on the anomaly detection.

[1074] Input: Anomaly detection result

[1075] Output: The generated notification message

[1076] Send notification messages to users' smartphones or smart assistants.

[1077] Input: Notification message

[1078] Output: Notification sent to user

[1079] This processing flow enables the system to monitor the user's health condition in real time and detect abnormalities early.

[1080] (Application example 1)

[1081] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1082] While self-driving vehicles are becoming increasingly common in modern society, there is currently no system that monitors the driver's health status in real time and responds promptly when an abnormality occurs. This may increase the risk of accidents due to driver health abnormalities. The present invention aims to solve this problem and provide an effective means for improving the safety of self-driving vehicles.

[1083] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1084] In this invention, the server includes means for collecting data from sensors used in the home, means for transmitting the collected data to the server, means for receiving and storing the transmitted data, means for analyzing the stored data using multimodal artificial intelligence, means for integrating the analysis results to monitor the health condition, means for generating and sending a notification to the user when an abnormality is detected, and means for collecting data from internal sensors of the autonomous vehicle and monitoring the driver's health condition. This makes it possible to monitor the driver's health condition in the vehicle in real time and send an immediate notification when an abnormality is detected.

[1085] "Domestic sensors" are devices installed inside the home that collect image, audio, text, biometric data, etc.

[1086] "Means for transmitting to the server" refers to a method or device for transmitting collected data to the server via a communication line such as the Internet.

[1087] "Storage means" refers to a method or device for retaining data received by the server for an extended period of time.

[1088] "Multimodal AI" refers to AI that has the ability to analyze multiple types of data (images, audio, text, biometric data, etc.) and integrate them.

[1089] "Health monitoring means" means a method or device that uses multimodal artificial intelligence to monitor analysis results and continuously assess a user's health status.

[1090] "Means for generating and sending a notification to the user when an abnormality is detected" refers to a method or device that automatically generates a warning or alert when an abnormality is discovered from the analysis results and forwards it to the user.

[1091] "Internal sensors in autonomous vehicles" refer to sensors that are placed inside autonomous vehicles and collect biometric and behavioral data from the driver.

[1092] "Means for monitoring the driver's health condition" refers to methods or devices that analyze data obtained from sensors in an autonomous vehicle and monitor the driver's health condition in real time.

[1093] The present invention provides a system for monitoring the health status of a user and a driver by comprehensively analyzing data from multiple sensors used in homes and autonomous vehicles. Specific embodiments are described below.

[1094] System Overview

[1095] The system collects data from multiple sensors (image sensors, voice sensors, text input devices, and biometric sensors) installed in the user's home and autonomous vehicle and sends it to a server. The server stores the received data and analyzes it using multimodal artificial intelligence (AI) to monitor the health status of the user and driver. If an abnormality is detected, the server generates a notification and sends it to the user or driver.

[1096] Hardware and software used

[1097] Hardware:

[1098] Home sensors (cameras, microphones, biometric sensors)

[1099] Internal sensors in autonomous vehicles (heart rate sensors, blood pressure sensors, cameras, etc.)

[1100] software:

[1101] Data Collection Applications

[1102] Communication protocol (HTTP, HTTPS)

[1103] Server (for data storage and analysis)

[1104] Multimodal AI analysis engine

[1105] Details of data processing and calculation

[1106] Data collection:

[1107] The devices (in-home sensors and in-vehicle sensors) collect image data, voice data, text data, biometric data, etc. This data is sent to a server at specific intervals.

[1108] Data transmission:

[1109] The device encrypts the collected data and sends it to the server using HTTPS, ensuring data security.

[1110] Data Analysis:

[1111] The server stores the received data and analyzes it using multimodal AI. The specific analysis content is as follows:

[1112] Image data: Analyzes changes in the user's or driver's behavior and facial expressions to detect abnormalities.

[1113] Voice data: Analyzed using voice recognition technology to detect changes in tone of voice and speaking patterns.

[1114] Text data: Analyze using natural language processing techniques to detect text content and signs of stress.

[1115] Biometric data: Time series analysis of heart rate, blood pressure abnormalities, and disrupted sleep patterns.

[1116] Anomaly detection and notification:

[1117] Multimodal AI integrates the results of various data analyses to assess the health status of the user and driver. If an abnormality is detected, the server generates a notification message and sends it to the user's smartphone or vehicle's infotainment system. The notification includes details of the abnormality and recommended actions to take.

[1118] Specific examples

[1119] For example, if a biometric sensor detects an abnormal heart rate while a driver of an autonomous vehicle is in the vehicle, the sensor will send this information to a server in real time. The server will analyze the heart rate and, if an abnormality is detected, will send a notification to the driver stating, "Your heart rate is outside the normal range. Please consult a medical professional."

[1120] Example prompts for generative AI models

[1121] "Please demonstrate the process flow of an application that analyzes data collected from internal sensors in an autonomous vehicle and monitors the driver's health. The implementation language should be Python, and the following sensors and data types should be assumed:

[1122] image

[1123] audio

[1124] text

[1125] Biometric data (heart rate, blood pressure)

[1126] Also include the ability to send data to a server and provide notifications based on the analysis results.

[1127] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1128] Step 1:

[1129] Devices (sensors in homes and vehicles) collect data. Specifically, cameras capture image and video data, microphones record voice data, text input devices acquire user input text, and biometric sensors measure biometric data such as heart rate and blood pressure. All of this data is collected from each sensor at regular intervals.

[1130] Input: Image data, audio data, text data, and biometric data obtained from each sensor.

[1131] Output: Consolidated sensor data.

[1132] Step 2:

[1133] The terminal sends the collected data to the server. The collected data is encrypted using a communication protocol (e.g., HTTPS) and sent securely to the server. The data is sent at regular intervals.

[1134] Input: Integrated sensor data.

[1135] Output: The data sent to the server.

[1136] Step 3:

[1137] The server stores the received data. The server stores the received data in a database for later analysis. During this storage process, each data type (image, audio, text, biometric data) is recorded in chronological order.

[1138] Input: The data sent to the server.

[1139] Output: Data stored in the database.

[1140] Step 4:

[1141] The server analyzes the stored data using multimodal artificial intelligence. Specifically, image data is analyzed using computer vision technology to detect movements and facial expressions. Audio data is analyzed using voice recognition technology to detect tone of voice and speech patterns. Text data is analyzed using natural language processing technology to detect text content and signs of stress. Biometric data is analyzed over time to detect abnormalities in heart rate and blood pressure.

[1142] Input: Data stored in a database.

[1143] Output: Analysis results.

[1144] Step 5:

[1145] The server integrates the analysis results and monitors the health of the user and driver. Multimodal AI integrates the analysis results of each data type to assess the overall health status. From this analysis, anomalies are identified.

[1146] Input: Analysis results for each data type.

[1147] Output: Integrated health status assessment.

[1148] Step 6:

[1149] If the server detects an abnormality, it generates a notification and sends it to the user or driver. The server generates a notification message containing the details of the abnormality and recommended actions to take, and sends it to the user's smartphone or the vehicle's infotainment system.

[1150] Input: Integrated Health Status Assessment.

[1151] Output: Informational message.

[1152] Step 7:

[1153] The user or driver receives a notification and takes appropriate action, such as consulting a medical professional.

[1154] Input: Notification message.

[1155] Output: Coping behavior.

[1156] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1157] The present invention is a system that monitors a user's health condition by comprehensively analyzing data from multiple sensors used in the home. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it is possible to monitor the user's emotional health as well. Specific embodiments of the system will now be described.

[1158] System Overview

[1159] The system collects data from multiple sensors (image sensors, voice sensors, text input devices, and biometric sensors) installed in the user's home and sends it to a server. The server stores the received data and analyzes it using a multimodal artificial intelligence (AI) and emotion engine to monitor the user's health and emotional state. If an abnormality is detected, the server sends a notification to the user.

[1160] Program processing flow

[1161] Data collection

[1162] Devices (users' home sensors and wearable devices) collect data

[1163] Image data: Security cameras in the home periodically take images and videos, recording the user's movements, facial expressions, and daily life.

[1164] Voice data: Smart speakers and smart assistants collect your voice and use the recorded audio to analyze your tone of voice and speaking patterns.

[1165] Text data: Collects text data entered by users through smart assistants or dedicated applications, and analyzes changes in the content and expression of the text.

[1166] Biometric data: Wearable devices measure biometric data such as heart rate, blood pressure, oxygen saturation, and sleep patterns to continuously monitor the user's physical condition.

[1167] Data transmission

[1168] The device sends the collected data to the server.

[1169] The sensor terminals send the collected data to the server at regular intervals, and the data is encrypted using a security protocol.

[1170] Data analysis

[1171] The server analyzes the received data

[1172] Image analysis: The server uses image data to analyze changes in the user's movements and facial expressions, thereby detecting physical changes and patterns in daily activities.

[1173] Voice analysis: The server uses speech recognition technology to convert voice data into text and analyzes changes in tone and speaking style to identify emotional issues and changes in cognitive function.

[1174] Text analysis: The server uses natural language processing technology to analyze the text data and detect changes in the content and expressions of the text, which allows for an assessment of stress and mental health.

[1175] Biometric data analysis: The server uses time series analysis techniques to analyze data such as heart rate, blood pressure, oxygen saturation, and sleep patterns to detect abnormalities, allowing for early detection of cardiac and respiratory problems.

[1176] Emotion analysis

[1177] The server uses an emotion engine to analyze the user's emotions.

[1178] Emotion analysis using voice data: The server analyzes the voice data and uses an emotion engine to recognize the user's emotional state from changes in their tone of voice and speaking style.

[1179] Emotion analysis using text data: The server analyzes the text data and uses an emotion engine to recognize the emotional state from changes in the content and expression of the text.

[1180] Emotion analysis using image data: The server analyzes the image data and uses an emotion engine to recognize the user's emotional state from their facial expressions.

[1181] Health and emotional state monitoring

[1182] The server integrates the analysis results and monitors the health and emotional state.

[1183] By integrating the results of various data analyses, the multimodal AI and emotion engine assesses the user's overall health and emotional state, making it easier to detect abnormalities.

[1184] Send notifications

[1185] If the server detects an abnormality, it will notify the user.

[1186] If the server detects an anomaly, it generates a notification message and sends it to the user's smartphone or smart assistant, containing details of the anomaly and recommended actions to take.

[1187] Example: Emotion and health monitoring using voice data

[1188] The device (smart speaker) collects the user's voice from everyday conversations. The server analyzes the received voice data and converts it into text data using voice recognition technology. It then uses an emotion engine to analyze subtle changes in voice tone and speaking style to assess whether the user is feeling stressed. The server also uses multimodal AI to analyze heart rate data collected simultaneously with the voice data, detecting physical changes such as an abnormally high heart rate. When an abnormality is detected, the server immediately generates a notification and sends it to the user's smartphone, stating, "Your recent conversations have shown signs of stress, and your heart rate is also high. Please try some relaxation techniques." This allows for a prompt response.

[1189] In this way, the present invention can protect the overall health of users by monitoring both health and emotions in the home in real time and identifying risks early.

[1190] The processing flow will be explained below.

[1191] Step 1:

[1192] Devices (sensors) collect data. Home security cameras periodically capture images and videos, smart speakers and smart assistants record voice data, and text data entered by users through smart assistants and applications is also collected. Wearable devices also measure biometric data such as heart rate, blood pressure, oxygen saturation, and sleep patterns in real time.

[1193] Step 2:

[1194] The terminal (sensor) packetizes and encrypts the collected data, and the data packets are temporarily stored in a queue inside the terminal.

[1195] Step 3:

[1196] The device (sensor) sends data packets to the server at regular intervals. The data transmission is performed using a secure communication protocol to ensure the confidentiality and security of the data.

[1197] Step 4:

[1198] The server receives the data packet. The received data packet is decoded and restored to its original format. The data integrity is checked, and if there are any errors or omissions, a retransmission request is sent to the terminal.

[1199] Step 5:

[1200] The server stores the recovered data in a distributed database, where the data is managed for each user and tagged within the database.

[1201] Step 6:

[1202] The server analyzes the received image and video data. Image classification and motion analysis algorithms are used to detect changes in the user's movement patterns and facial expressions. For example, unusual movements or changes in facial expressions may be signs of an abnormality.

[1203] Step 7:

[1204] The server analyzes the voice data. It uses speech recognition technology to convert the data into text and analyzes changes in tone and speaking style. This allows it to detect emotional issues and changes in cognitive function. For example, a trembling voice or changes in speaking style could be a sign of stress or changes in cognitive function.

[1205] Step 8:

[1206] The server analyzes the text data. It uses natural language processing technology to analyze the content of the text and detect changes in the content and expressions of the sentences. This allows for an assessment of stress and mental health. For example, an increase in negative expressions in the sentences can be determined to be a sign of stress.

[1207] Step 9:

[1208] The server analyzes biometric data, using time series analysis techniques to analyze data such as heart rate, blood pressure, oxygen saturation, and sleep patterns to detect abnormalities. This allows for early detection of cardiac and respiratory problems. For example, a higher-than-normal heart rate could be a sign of cardiac disease.

[1209] Step 10:

[1210] The server uses an emotion engine to analyze the user's emotions. It recognizes the user's emotions from voice, text, and image data. This allows the system to monitor the user's emotional health. For example, if anger, excitement, or sadness is detected from the voice data, it is considered a sign of emotional instability.

[1211] Step 11:

[1212] The server integrates the results of each analysis, and uses a multimodal AI and emotion engine to comprehensively analyze images, voice, text, and biometric data to assess the user's overall health and emotional state, making it easier to detect abnormalities.

[1213] Step 12:

[1214] If the server detects an anomaly, it generates a notification with details about the anomaly and a recommended course of action, such as "Your heart rate has been higher than normal recently, and you appear to be feeling stressed."

[1215] Step 13:

[1216] The server generates a notification and sends it to the user's smartphone or smart assistant. The notification is sent immediately, informing the user that an abnormality has occurred.

[1217] Step 14:

[1218] Users receive notifications, review them, take action based on recommended actions, and, if necessary, seek medical advice. This allows users to quickly respond to changes in their health and emotional state.

[1219] This specific processing flow enables real-time monitoring of both health and emotions within the home, identifying risks early and protecting the overall health of users.

[1220] Example 2

[1221] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1222] There is a need for technology that can comprehensively monitor not only a user's health status but also their emotional state by collecting and analyzing different types of data from multiple sensors used in the home. However, existing systems have faced many challenges in centralized data management, complex data analysis, anomaly detection, and prompt notification to the user. In particular, it has been difficult to integrate information from multiple data sources and perform accurate analysis.

[1223] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for collecting data from sensors used in the home, means for transmitting the collected data to the server, means for receiving and storing the transmitted data, means for analyzing the stored data using multimodal artificial intelligence and an emotion analysis engine, means for integrating the analysis results to monitor the health status and emotional state, and means for generating and sending a notification to the user when an abnormality is detected. This makes it possible to comprehensively monitor the user's health status and emotional state in real time and to provide a prompt and appropriate notification when an abnormality occurs.

[1224] A "sensor" is a device used in the home that collects information such as images, audio, text, and biometric data.

[1225] "Data" refers to all information collected by sensors, such as image data, video data, audio data, text data, and biometric data.

[1226] A "server" is an information processing device that receives and stores data sent from sensors, and analyzes the data to monitor the user's condition.

[1227] "Multimodal AI" is an AI technology that has the ability to integrate and analyze multiple different types of data (e.g., images, audio, text, biometric data) in a unified manner.

[1228] An "emotion analysis engine" is a technology for analyzing a user's emotional state from data such as voice, text, and images.

[1229] "Analyzing" means that the server processes the data it receives and derives meaningful results.

[1230] "Monitoring" means that the server continuously observes and evaluates the user's health and emotional state based on the analysis results.

[1231] A "notification" is a warning or advice message that is generated and sent to the user when the server detects an abnormality.

[1232] An "anomaly" is any unusual change or problem in the user's health or emotional state that is detected as a result of the server's analysis.

[1233] "Integrating" means that the server compiles the results of multiple data analyses and provides a single comprehensive evaluation.

[1234] The present invention provides a system for monitoring a user's health and emotional state by comprehensively analyzing data from multiple sensors used in the home. A specific method for implementing this system will be described below.

[1235] System configuration

[1236] The system collects data from multiple sensors installed in the home (image sensors, voice sensors, text input devices, and biometric sensors) and sends it to a server. The server stores the received data and analyzes it using a multimodal artificial intelligence (AI) and emotion analysis engine. The analysis results are integrated to monitor the user's health and emotional state. If an abnormality is detected, the server sends a notification to the user.

[1237] Hardware and software used

[1238] Sensors used in the home include:

[1239] Image sensor: Home security camera

[1240] Voice Sensors: Smart Speakers and Smart Assistants

[1241] Text input devices: smart assistants and dedicated applications

[1242] Biosensors: Wearable devices that measure heart rate, blood pressure, oxygen saturation, and sleep patterns

[1243] The server includes the following analysis modules:

[1244] Multimodal Artificial Intelligence (AI)

[1245] Sentiment Analysis Engine

[1246] Data Storage and Management System

[1247] Notification System

[1248] Data collection and transmission

[1249] 1. Devices (sensors) collect data

[1250] Image sensors capture images and videos of the home and record changes in the user's movements and facial expressions. For example, a camera in the living room can start recording at a set time.

[1251] The voice sensor uses smart speakers to collect everyday conversations and analyze tone and speech patterns, for example, recording the voice you use to ask about the weather.

[1252] The text input device collects text entered by the user through a dedicated application and analyzes changes in the content and expression of the text, for example, collecting diary entries from the user.

[1253] Biometric sensors use wearable devices to continuously measure heart rate, blood pressure, oxygen saturation, and sleep patterns. For example, a smartwatch worn by a user records heart rate.

[1254] 2. The device sends the collected data to the server

[1255] The sensor terminals transmit the collected data to a server at regular intervals. The transmitted data is encrypted using a security protocol. For example, audio data is uploaded to the server every hour.

[1256] Data analysis

[1257] 1. The server analyzes the received data

[1258] Image analysis: The server uses image data to analyze changes in the user's movements and facial expressions, for example, to detect changes in the user's walking pattern.

[1259] Voice analysis: The server uses voice recognition technology to convert the voice data into text and analyzes changes in tone and speaking style, for example, detecting if the user's voice is darker than usual.

[1260] Text analysis: The server uses natural language processing technology to analyze the text data and detect changes in the content and expressions of the sentences. For example, it analyzes the frequent use of the expression "tired."

[1261] Biometric data analysis: The server uses time series analysis to analyze heart rate, blood pressure, oxygen saturation, sleep patterns, etc. to detect abnormalities, such as a higher than normal heart rate.

[1262] 2. Emotion analysis

[1263] Emotion analysis from voice data: The server analyzes the voice data and uses an emotion analysis engine to recognize the emotional state from changes in tone of voice and speaking style, for example, detecting an increase in anger.

[1264] Sentiment analysis based on text data: The server analyzes the text data and uses a sentiment analysis engine to recognize the emotional state from changes in the content and expressions of the text. For example, it detects an increase in pessimistic expressions.

[1265] Emotion analysis based on image data: The server analyzes the image data and uses an emotion analysis engine to recognize the user's emotional state from their facial expressions. For example, it detects that they are smiling less.

[1266] Integrated monitoring and abnormality notification

[1267] 1. The server integrates the analysis results and monitors the health and emotional state.

[1268] Using a multimodal AI and emotion analysis engine, the analysis results are integrated to assess the user's overall health and emotional state. For example, if different emotional indicators are found from image analysis, audio analysis, and text analysis, a comprehensive assessment is made.

[1269] 2. Detect anomalies and generate notifications

[1270] If an anomaly is detected, the server generates a real-time notification. For example, if heart rate and voice data indicate strong signs of stress, a notification will be sent to the user stating, "Your recent conversations have shown signs of stress, and your heart rate is high. Please try some relaxation techniques."

[1271] Specific examples

[1272] For example, a smart speaker collects a user's voice from everyday conversations and sends the data to a server. The server analyzes the voice data and converts it into text data using speech recognition technology. It uses an emotion analysis engine to analyze subtle changes in the tone of voice and speaking style to assess whether the user is feeling stressed. At the same time, the server also analyzes heart rate data from a wearable device and detects if the heart rate is higher than normal. If an abnormality is detected, the server sends the user a notification saying, "Your recent conversations have shown signs of stress, and your heart rate is high. Please try some relaxation techniques."

[1273] Example prompt: "Analyze the voice data collected by the smart speaker and assess whether the user is experiencing stress."

[1274] This invention can protect the overall health of users by monitoring both health and emotions in the home in real time and identifying risks early.

[1275] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1276] Step 1: Collect data

[1277] Data is collected by sensors and devices installed in the home by users.

[1278] Input: Various data from the home environment and the user (images, voice, text, biometric data).

[1279] How it works: Security cameras take pictures, smart speakers record everyday conversations, and wearable devices measure heart rate and blood pressure.

[1280] Output: Various data collected.

[1281] Step 2: Sending data

[1282] The terminal transmits the collected data to the server.

[1283] Input: Various types of collected data (images, audio, text, biometric data).

[1284] Specific operation: Image data captured by the security camera is uploaded to the server, and audio data recorded by the smart speaker is sent to the server.

[1285] Output: The data sent to the server.

[1286] Step 3: Save your data

[1287] The server stores the received data.

[1288] Input: The data sent to the server.

[1289] Specific operation: The server stores various data in the database.

[1290] Output: Saved data.

[1291] Step 4: Analyzing the image data

[1292] The server analyzes the image data.

[1293] Input: Saved image data.

[1294] Specific operation: Multimodal AI is used to analyze changes in user behavior and facial expressions from image data. For example, facial recognition algorithms are used to analyze user facial expressions.

[1295] Output: Analysis results (changes in user's behavior and facial expressions).

[1296] Step 5: Analyzing the audio data

[1297] The server analyzes the audio data.

[1298] Input: Stored audio data.

[1299] What it does: It uses speech recognition technology to convert voice data into text, and uses an emotion analysis engine to analyze changes in tone and speaking style, for example, to identify changes in tone in a user's everyday conversation.

[1300] Output: Analysis results (changes in tone of voice and speaking style).

[1301] Step 6: Analyze the text data

[1302] The server parses the text data.

[1303] Input: Saved text data.

[1304] Specific operation: Analyzes text data using natural language processing technology and detects changes in the content and expressions of sentences using a sentiment analysis engine. For example, it analyzes the frequent use of the expression "tired."

[1305] Output: Analysis results (changes in the content and expression of the text).

[1306] Step 7: Analyzing the biometric data

[1307] The server analyzes the biometric data.

[1308] Input: Stored biometric data.

[1309] Specific operation: Using time series analysis technology, it analyzes heart rate, blood pressure, oxygen saturation, sleep patterns, etc. to detect abnormalities, such as when the heart rate is higher than normal.

[1310] Output: Analysis results (heart rate, blood pressure, oxygen saturation, changes in sleep patterns).

[1311] Step 8: Integrating the analysis results

[1312] The server consolidates the various analysis results.

[1313] Input: Analysis results of various data (images, audio, text, biometric data).

[1314] How it works: Using a multimodal AI and emotion analysis engine, it integrates analysis results to assess the user's overall health and emotional state. For example, if different emotional indicators are found from image analysis, audio analysis, and text analysis, it will comprehensively evaluate them.

[1315] Output: Integrated analysis results (assessment of the user's health and emotional state).

[1316] Step 9: Detect anomalies

[1317] The server detects an abnormality.

[1318] Input: Consolidated analysis results.

[1319] Specific operation: Based on the analysis results, an anomaly detection algorithm is executed to detect abnormalities in the user's health and emotional state, such as stress and heart rate abnormalities.

[1320] Output: Anomaly detection results.

[1321] Step 10: Generate and send notifications

[1322] The server generates the notification and sends it to the user.

[1323] Input: Anomaly detection results.

[1324] What it does: If an abnormality is detected, it will generate a notification message and send it to the user's smartphone or smart assistant, for example, "Your recent conversations have shown signs of stress and your heart rate is elevated. Please try some relaxation techniques."

[1325] Output: The notification message sent to the user.

[1326] (Application example 2)

[1327] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1328] Conventional health monitoring systems in homes and businesses are limited to monitoring physical health conditions, making it difficult to respond quickly and accurately to changes or abnormalities in emotional states. Furthermore, the means of detecting abnormalities are limited, and detection of abnormalities can be delayed under certain conditions. Therefore, there is a need for a system that can monitor both health and emotional states in an integrated manner, quickly detect abnormalities, and provide appropriate notifications.

[1329] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting data from sensors used in homes and business facilities, means for transmitting the collected data to the server, means for receiving and storing the transmitted data, means for analyzing the stored data using multimodal artificial intelligence, means for integrating the analysis results to monitor health and emotional states, means for generating and transmitting a notification to the user when an abnormality is detected, and means for analyzing changes in the movements, facial expressions, and tone of voice of visitors and employees to identify their emotional states. This makes it possible to monitor not only health states but also emotional states in real time, quickly detect abnormalities, and issue appropriate notifications.

[1330] "Inside a home or business premises" refers to the physical environment in which a person lives or conducts business activities, and refers to the location of the system used within that particular space.

[1331] A "sensor" is a device that detects physical phenomena and converts them into electronic data, and includes image sensors, audio sensors, text input devices, and biometric sensors.

[1332] "Means for collecting data" refers to a system element that uses sensors to obtain information from an object and record it in electronic form.

[1333] "Means for transmitting data to a server" refers to the communications devices and protocols used to transfer the collected digital data over a network to a remote server.

[1334] "Means for receiving and storing data" refers to a system element that allows a server to receive data sent via a network and store it in a database or the like.

[1335] "Multimodal AI" refers to an AI technology that simultaneously analyzes multiple types of data (e.g., images, voice, text, biometric data) and makes integrated judgments.

[1336] "Health and emotional monitoring means" refers to system elements that provide continuous monitoring of an individual's physical and emotional health based on collected and analyzed data.

[1337] "Means for generating and sending notifications to users when an anomaly is detected" refers to the ability for the system to automatically create and send alerts or notifications to users' devices when an anomaly is detected.

[1338] "Means of identifying emotional states by analyzing changes in behavior, facial expressions, and tone of voice" refers to artificial intelligence technology that analyzes image and audio data to estimate emotional states from behavior, changes in facial expressions, tone of voice, etc.

[1339] An embodiment of the present invention is described in detail below: The system monitors the health and emotional state of visitors and employees by collecting data from multiple sensors installed in homes and businesses, transmitting the data to a server for analysis.

[1340] System Configuration

[1341] Terminal (sensor)

[1342] The devices include image sensors, voice sensors, text input devices, and biometric sensors that are placed within physical stores to collect the following data from visitors and employees:

[1343] Image data: The camera records your movements and expressions.

[1344] Audio data: Microphones collect speech.

[1345] Text data: The digital device captures the input text.

[1346] Biometric data: Wearable devices and thermal cameras collect heart rate and temperature data.

[1347] The data collected by each terminal is encrypted and sent to a server at regular intervals.

[1348] Server Processing

[1349] Data reception and storage

[1350] The server receives the data sent from the terminal via the network in an encrypted form and stores it in a database, using technologies including security protocols and database systems.

[1351] Data analysis

[1352] The server analyzes the various types of stored data (images, audio, text, and biometric data) using multimodal artificial intelligence (AI). Image data is analyzed for changes in movements and facial expressions using libraries such as OpenCV, and audio data is converted into text using pyaudio and librosa, after which changes in tone of voice and speaking style are analyzed. Text data is analyzed for emotional state using natural language processing techniques such as TextBlob. Furthermore, biometric data is analyzed for abnormalities using time series analysis techniques.

[1353] Emotion analysis

[1354] The server uses an emotion engine to integrate data from different modalities and recognize the emotional state of visitors and employees, allowing real-time monitoring of changes in not only physical health but also emotional health.

[1355] Notifications and Feedback

[1356] If the server detects an abnormality, it automatically generates a notification message and sends it to the user's smartphone or other display device. The notification includes details of the detected abnormality and recommended actions to take. For example, a message might be generated that reads, "Employee A has a high body temperature and his tone of voice indicates stress. We recommend that he leave work early."

[1357] Examples and prompts

[1358] As a concrete example, consider a monitoring system for Employee A in a store. The server collects and analyzes Employee A's temperature and conversation data, and if an abnormality is detected, it sends a notification saying, "Employee A's temperature is high and his / her tone of voice also indicates stress. We recommend that you leave early." An example of how such a prompt is input into a generative AI model is as follows:

[1359] "Monitor employee health and emotions, and if you detect a high temperature or stressful tone in the conversation, generate a notification recommending that they leave early. Could you please give us an example of what this notification should look like?"

[1360] This will enable the realization of a system that monitors health and emotions in an integrated manner and can respond quickly in the event of an abnormality.

[1361] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1362] Step 1:

[1363] The terminal collects image data, audio data, text data, and biometric data. Image data is obtained by recording the movements and facial expressions of visitors and employees with a camera, while audio data is obtained by collecting conversations and environmental sounds with a microphone. Text data is obtained by acquiring input information from digital terminals, and biometric data is obtained by collecting heart rate and body temperature information with a wearable device or a body temperature detection camera. The input is various data obtained from each sensor, and the output is the collected raw data.

[1364] Step 2:

[1365] The terminal sends the collected data to the server at regular intervals. The data is sent encrypted over the network. In this step, the input is the collected data, and the output is the encrypted data sent to the server.

[1366] Step 3:

[1367] The server saves the received data. The received data is stored in the server's database. The input of this step is the encrypted transmitted data, and the output is the saved data.

[1368] Step 4:

[1369] The server analyzes the stored data. The analysis is divided into the following parts:

[1370] Image analysis: Image data is analyzed using OpenCV etc. to detect changes in movements and facial expressions. The input is the saved image data, and the output is the analysis results of movements and facial expression changes.

[1371] Speech analysis: Using pyaudio and librosa, audio data is converted to text and changes in voice tone and speaking style are analyzed. The input is the stored audio data, and the output is the converted text data and changes in voice tone and speaking style.

[1372] Text Analysis: Analyzes text data using natural language processing techniques such as TextBlob to evaluate emotional states. The input is the stored text data, and the output is the analysis result of the emotional state.

[1373] Biometric data analysis: Detects abnormalities in biometric data (heart rate, body temperature, etc.) using time series analysis techniques. The input is stored biometric data, and the output is the anomaly detection results.

[1374] Step 5:

[1375] The server uses the emotion engine to integrate the analysis results of each data and monitor the health and emotional states of visitors and employees. The input of this step is each analysis result, and the output is the integrated evaluation result of the health and emotional states.

[1376] Step 6:

[1377] If the server detects an abnormality, it generates a notification message and sends it to the user's smartphone or other display device. The notification includes details of the abnormality and recommended actions to take. For example, a notification might be sent stating, "Employee A's temperature is high and his / her tone of voice indicates stress. We recommend that he / she leave work early." The input to this step is the integrated evaluation result, and the output is the generated notification message.

[1378] The above is the flow of processing in the embodiment of the present invention.

[1379] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1380] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1381] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1382] [Fourth embodiment]

[1383] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1384] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1385] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1386] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1387] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1388] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1389] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1390] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1391] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1392] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1393] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1394] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1395] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1396] The present invention is a system for monitoring the health condition of a user by comprehensively analyzing data from multiple sensors used in the home. A specific embodiment of the system will be described.

[1397] System Overview

[1398] The system collects data from multiple sensors (image sensors, voice sensors, text input devices, and biometric sensors) installed in the user's home and sends it to a server. The server stores the received data, analyzes it using multimodal artificial intelligence (AI), and monitors the user's health status. If an abnormality is detected, the server sends a notification to the user.

[1399] Program processing flow

[1400] Data collection

[1401] Devices (users' home sensors and wearable devices) collect data

[1402] Image data: Security cameras in the home periodically take images and videos, recording the user's movements, facial expressions, and daily life.

[1403] Voice data: Smart speakers and smart assistants collect your voice and use it to detect changes in your tone and speaking style.

[1404] Text data: Collects text data entered by users through smart assistants or dedicated applications, and analyzes the content of the text and signs of stress.

[1405] Biometric data: Wearable devices measure biometric data such as heart rate, blood pressure, oxygen saturation, and sleep patterns, thereby continuously monitoring the user's physical condition.

[1406] Data transmission

[1407] The device sends the collected data to the server.

[1408] The sensor terminals send the collected data to the server at regular intervals, and the data is encrypted using a security protocol.

[1409] Data analysis

[1410] The server analyzes the received data

[1411] The server uses the image data to analyze changes in the user's movements and facial expressions. For example, unusual movements may indicate a health problem.

[1412] Voice data is analyzed using speech recognition technology to detect changes in tone of voice and speech patterns, which can identify emotional issues and changes in cognitive function.

[1413] The text data is analyzed using natural language processing technology to detect changes in the content and expressions of the text, which can then be used to assess stress and mental health.

[1414] Biometric data is analyzed over time to detect abnormalities in heart rate, blood pressure, and sleep patterns, allowing for early detection of cardiac and respiratory problems.

[1415] Health Status Monitoring

[1416] The server consolidates the analysis results and monitors health status.

[1417] By integrating the results of various data analyses, multimodal AI assesses the user's overall health status, making it easier to detect abnormalities.

[1418] Send notifications

[1419] If the server detects an abnormality, it will notify the user.

[1420] If the server detects an anomaly, it generates a notification message and sends it to the user's smartphone or smart assistant, containing details of the anomaly and recommended actions to take.

[1421] Example: Detecting abnormal heart rates at night

[1422] The terminal (wearable device) continues to measure the user's heart rate throughout the night. The server receives and stores the collected heart rate data. The server then analyzes the received data and detects abnormal increases in heart rate beyond the normal range. If an abnormality is detected, the server sends a notification to the user's smartphone stating, "A sudden increase in heart rate has been detected overnight." The user receives this notification and can take prompt action, such as consulting a medical institution.

[1423] In this way, the present invention can protect the user's health by monitoring the user's health status in real time at home and identifying health risks early.

[1424] The processing flow will be explained below.

[1425] Step 1:

[1426] Devices (sensors) collect data. Security cameras installed in homes periodically capture images and videos, and smart speakers and smart assistants record voice data. Text data entered by users through smart assistants and applications is also collected. In addition, wearable devices measure biometric data such as heart rate, blood pressure, oxygen saturation, and sleep patterns in real time.

[1427] Step 2:

[1428] The terminal (sensor) packetizes and encrypts the collected data, and the data packets are temporarily stored in a queue inside the terminal.

[1429] Step 3:

[1430] The device (sensor) sends data packets to the server at regular intervals. The data transmission is performed using a secure communication protocol to ensure the confidentiality and security of the data.

[1431] Step 4:

[1432] The server receives the data packet. The received data packet is decoded and restored to its original format. The data integrity is checked, and if there are any errors or omissions, a retransmission request is sent to the terminal.

[1433] Step 5:

[1434] The server stores the recovered data in a distributed database, where the data is managed for each user and tagged within the database.

[1435] Step 6:

[1436] The server analyzes the received image and video data, using image classification and motion analysis algorithms to detect changes in the user's movement patterns and facial expressions.

[1437] Step 7:

[1438] The server analyzes the audio data, converts it into text using speech recognition technology, and analyzes changes in tone and speaking style to identify signs associated with emotional problems and changes in cognitive function.

[1439] Step 8:

[1440] The server analyzes the text data and uses natural language processing technology to analyze the content of the text data and evaluate stress levels and health conditions.

[1441] Step 9:

[1442] The server analyzes biometric data, using time series analysis techniques to analyze data such as heart rate, blood pressure, oxygen saturation, and sleep patterns to detect abnormalities, especially values ​​that exceed the normal range.

[1443] Step 10:

[1444] The server integrates the results of each analysis, using multimodal AI to comprehensively analyze images, voice, text, and biometric data to assess the user's overall health.

[1445] Step 11:

[1446] The server generates a notification when an anomaly is detected, containing details about the anomaly and a recommended course of action.

[1447] Step 12:

[1448] The server generates a notification and sends it to the user's smartphone or smart assistant. The notification is sent immediately, informing the user that an abnormality has occurred.

[1449] Step 13:

[1450] Users will receive a notification, which they can review and follow the recommended actions, including seeking medical advice if necessary.

[1451] This processing flow allows for real-time health monitoring within the home, and if an abnormality is detected, it is possible to respond immediately.

[1452] Example 1

[1453] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1454] In modern society, people's lifestyles are becoming more diverse, and health management at home is becoming increasingly important. Daily health monitoring is especially important for the elderly and those with chronic illnesses. However, many current health monitoring systems are single-modal, analyzing only specific data, and lack the ability to comprehensively analyze different types of data. This makes it difficult to accurately grasp overall health status and detect abnormalities early.

[1455] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1456] In this invention, the server includes means for collecting data from sensors used in the home, means for transmitting the collected data to the server, means for receiving and storing the transmitted data, means for analyzing the stored data using multimodal artificial intelligence, means for integrating the analyzed data and monitoring the health status, and means for generating and sending a notification to the user when an abnormality is detected. This makes it possible to integrate and analyze information from multiple data sources, accurately monitor the user's overall health status, and detect abnormalities early.

[1457] A "sensor used in the home" is a device that operates in a user's home environment and collects one or more of image data, video data, audio data, text data, and / or biometric data.

[1458] The "server" is a device or system that receives and stores data sent from sensors in the home, analyzes the data using multimodal artificial intelligence, and integrates the results to monitor the user's health condition.

[1459] "Multimodal artificial intelligence" refers to artificial intelligence technology that comprehensively analyzes different types of data (images, audio, text, biometric data, etc.).

[1460] "Analysis results" refers to analytical information generated by multimodal artificial intelligence based on collected data, including data indicating the user's health status.

[1461] A "notification" is a message generated when the server detects an anomaly based on the analysis results, informing the user of the nature of the anomaly and the recommended course of action.

[1462] This invention is a system that monitors a user's health condition by integrating and analyzing data from multiple sensors used in the home. This system collects data from multiple sensors, including image sensors, voice sensors, text input devices, and biometric sensors, and transmits it to a server. The server stores the received data and analyzes it using multimodal artificial intelligence (AI) to monitor the user's health condition. If an abnormality is detected, the server also sends a notification to the user.

[1463] To implement this system, the following hardware and software are required: image sensors (e.g., security cameras), audio sensors (e.g., smart speakers), text input devices (e.g., smart assistants or dedicated applications), and biometric sensors (e.g., wearable devices) installed in the home. The server requires storage to receive and store this data and computing power to run AI models for data analysis.

[1464] This system collects, transmits, analyzes, monitors, and notifies data in the following ways:

[1465] As a concrete example, let's take the case of detecting abnormal heart rates at night. The wearable device, which is the terminal, continuously measures the user's heart rate at night. The measured heart rate data is sent to a server at regular intervals using an encryption protocol. The server stores the received heart rate data and begins analysis. The main process of the analysis is to identify when the heart rate deviates from the normal range.

[1466] If an abnormality is detected, the server generates a notification stating, "A sudden increase in heart rate was detected overnight." The notification includes a graph of heart rate variability and a recommendation to seek medical advice. The notification is sent to the user's smartphone or smart assistant, allowing the user to take appropriate action quickly.

[1467] Additionally, here are some examples of input prompts for the generative AI model:

[1468] Prompt using natural language processing technology: "How were you feeling yesterday? Please let me know if there was anything in particular that bothered you."

[1469] Prompt when using voice recognition technology: "Good morning. How long did you sleep last night?"

[1470] In this way, the present invention makes it possible to monitor the user's health condition in real time at home and identify health risks at an early stage by comprehensively analyzing various data.

[1471] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1472] Step 1:

[1473] Data collection

[1474] The device collects data from various sensors in the home (image sensors, voice sensors, text input devices, and biometric sensors). Specifically, the following operations are performed:

[1475] The image sensor uses a camera to periodically capture image data and video data.

[1476] Input: Real-time video data

[1477] Output: Captured image and video data

[1478] The audio sensor collects the user's voice at any time.

[1479] Input: Real-time audio data

[1480] Output: Collected audio data

[1481] The text input device captures text data entered by the user.

[1482] Input: Text entered by the user

[1483] Output: The retrieved text data

[1484] Biometric sensors measure vital data such as heart rate, blood pressure, oxygen saturation, and sleep patterns.

[1485] Input: Anthropometric data

[1486] Output: Measured biological data

[1487] Step 2:

[1488] Sending data to the server

[1489] The terminal periodically transmits the collected data to the server. Specifically, the following operations are performed.

[1490] The device encrypts the collected data using security protocols (e.g., TLS / SSL).

[1491] Input: Collected data (images, audio, text, biometric data)

[1492] Output: Encrypted data

[1493] The encrypted data is sent to a server over the Internet.

[1494] Input: Encrypted data

[1495] Output: Data sent to server completed

[1496] Step 3:

[1497] Data storage

[1498] The server stores the received data. Specifically, the following operations are performed:

[1499] The received data is saved in the database.

[1500] Input: Received data (decrypted)

[1501] Output: Saved data

[1502] Step 4:

[1503] Data analysis

[1504] The server analyzes the stored data using multimodal artificial intelligence. Specifically, the following operations are performed:

[1505] Image data analysis: Using computer vision technology, image and video data is analyzed to detect changes in the user's movements and facial expressions.

[1506] Input: Image and video data

[1507] Output: Detection results of changes in behavior and facial expressions

[1508] Voice data analysis: Applying voice recognition technology to analyze changes in tone of voice and speaking patterns.

[1509] Input: Audio data

[1510] Output: Detected changes in tone of voice and speech patterns

[1511] Text data analysis: Using natural language processing techniques, we analyze the content and expression patterns of text.

[1512] Input: Text data

[1513] Output: Analysis results of text content and stress symptoms

[1514] Biometric data analysis: Perform time series analysis to detect abnormalities in heart rate, blood pressure, and sleep patterns.

[1515] Input: Biometric data

[1516] Output: Abnormality detection results for biometric data

[1517] Step 5:

[1518] Health monitoring

[1519] The server integrates the analysis results and monitors the user's health status. Specifically, the following operations are performed:

[1520] The results of various data analyses are integrated and a comprehensive health assessment is performed using multimodal AI.

[1521] Input: Detection results of each analysis data

[1522] Output: Overall health assessment results

[1523] Step 6:

[1524] Generate and send notifications when anomalies occur

[1525] If the server detects an abnormality, it generates a notification and sends it to the user. Specifically, the following operations are performed:

[1526] A notification message is generated based on the anomaly detection.

[1527] Input: Anomaly detection result

[1528] Output: The generated notification message

[1529] Send notification messages to users' smartphones or smart assistants.

[1530] Input: Notification message

[1531] Output: Notification sent to user

[1532] This processing flow enables the system to monitor the user's health condition in real time and detect abnormalities early.

[1533] (Application example 1)

[1534] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1535] While self-driving vehicles are becoming increasingly common in modern society, there is currently no system that monitors the driver's health in real time and responds promptly when an abnormality occurs. This may increase the risk of accidents due to driver health abnormalities. The present invention aims to solve this problem and provide an effective means for improving the safety of self-driving vehicles.

[1536] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1537] In this invention, the server includes means for collecting data from sensors used in the home, means for transmitting the collected data to the server, means for receiving and storing the transmitted data, means for analyzing the stored data using multimodal artificial intelligence, means for integrating the analysis results to monitor the health condition, means for generating and sending a notification to the user when an abnormality is detected, and means for collecting data from internal sensors of the autonomous vehicle and monitoring the driver's health condition. This makes it possible to monitor the driver's health condition in the vehicle in real time and send an immediate notification when an abnormality is detected.

[1538] "Domestic sensors" are devices installed inside the home that collect image, audio, text, biometric data, etc.

[1539] "Means for transmitting to the server" refers to a method or device for transmitting collected data to the server via a communication line such as the Internet.

[1540] "Storage means" refers to a method or device for retaining data received by the server for an extended period of time.

[1541] "Multimodal AI" refers to AI that has the ability to analyze multiple types of data (images, audio, text, biometric data, etc.) and integrate them.

[1542] "Health monitoring means" means a method or device that uses multimodal artificial intelligence to monitor analysis results and continuously assess a user's health status.

[1543] "Means for generating and sending a notification to the user when an abnormality is detected" refers to a method or device that automatically generates a warning or alert when an abnormality is discovered from the analysis results and forwards it to the user.

[1544] "Internal sensors in autonomous vehicles" refer to sensors that are placed inside autonomous vehicles and collect biometric and behavioral data from the driver.

[1545] "Means for monitoring the driver's health condition" refers to methods or devices that analyze data obtained from sensors in an autonomous vehicle and monitor the driver's health condition in real time.

[1546] The present invention provides a system for monitoring the health status of a user and a driver by comprehensively analyzing data from multiple sensors used in homes and autonomous vehicles. Specific embodiments are described below.

[1547] System Overview

[1548] The system collects data from multiple sensors (image sensors, voice sensors, text input devices, and biometric sensors) installed in the user's home and autonomous vehicle and sends it to a server. The server stores the received data and analyzes it using multimodal artificial intelligence (AI) to monitor the health status of the user and driver. If an abnormality is detected, the server generates a notification and sends it to the user or driver.

[1549] Hardware and software used

[1550] Hardware:

[1551] Home sensors (cameras, microphones, biometric sensors)

[1552] Internal sensors in autonomous vehicles (heart rate sensors, blood pressure sensors, cameras, etc.)

[1553] software:

[1554] Data Collection Applications

[1555] Communication protocol (HTTP, HTTPS)

[1556] Server (for data storage and analysis)

[1557] Multimodal AI analysis engine

[1558] Details of data processing and calculation

[1559] Data collection:

[1560] The devices (in-home sensors and in-vehicle sensors) collect image data, voice data, text data, biometric data, etc. This data is sent to a server at specific intervals.

[1561] Data transmission:

[1562] The device encrypts the collected data and sends it to the server using HTTPS, ensuring data security.

[1563] Data Analysis:

[1564] The server stores the received data and analyzes it using multimodal AI. The specific analysis content is as follows:

[1565] Image data: Analyzes changes in the user's or driver's behavior and facial expressions to detect abnormalities.

[1566] Voice data: Analyzed using voice recognition technology to detect changes in tone of voice and speaking patterns.

[1567] Text data: Analyze using natural language processing techniques to detect text content and signs of stress.

[1568] Biometric data: Time series analysis of heart rate, blood pressure abnormalities, and disrupted sleep patterns.

[1569] Anomaly detection and notification:

[1570] Multimodal AI integrates the results of various data analyses to assess the health status of the user and driver. If an abnormality is detected, the server generates a notification message and sends it to the user's smartphone or vehicle's infotainment system. The notification includes details of the abnormality and recommended actions to take.

[1571] Specific examples

[1572] For example, if a biometric sensor detects an abnormal heart rate while a driver of an autonomous vehicle is in the vehicle, the sensor will send this information to a server in real time. The server will analyze the heart rate and, if an abnormality is detected, will send a notification to the driver stating, "Your heart rate is outside the normal range. Please consult a medical professional."

[1573] Example prompts for generative AI models

[1574] "Please demonstrate the process flow of an application that analyzes data collected from internal sensors in an autonomous vehicle and monitors the driver's health. The implementation language should be Python, and the following sensors and data types should be assumed:

[1575] image

[1576] audio

[1577] text

[1578] Biometric data (heart rate, blood pressure)

[1579] Also include the ability to send data to a server and provide notifications based on the analysis results.

[1580] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1581] Step 1:

[1582] Devices (sensors in homes and vehicles) collect data. Specifically, cameras capture image and video data, microphones record voice data, text input devices acquire user input text, and biometric sensors measure biometric data such as heart rate and blood pressure. All of this data is collected from each sensor at regular intervals.

[1583] Input: Image data, audio data, text data, and biometric data obtained from each sensor.

[1584] Output: Consolidated sensor data.

[1585] Step 2:

[1586] The terminal sends the collected data to the server. The collected data is encrypted using a communication protocol (e.g., HTTPS) and sent securely to the server. The data is sent at regular intervals.

[1587] Input: Integrated sensor data.

[1588] Output: The data sent to the server.

[1589] Step 3:

[1590] The server stores the received data. The server stores the received data in a database for later analysis. During this storage process, each data type (image, audio, text, biometric data) is recorded in chronological order.

[1591] Input: The data sent to the server.

[1592] Output: Data stored in the database.

[1593] Step 4:

[1594] The server analyzes the stored data using multimodal artificial intelligence. Specifically, image data is analyzed using computer vision technology to detect movements and facial expressions. Audio data is analyzed using voice recognition technology to detect tone of voice and speech patterns. Text data is analyzed using natural language processing technology to detect text content and signs of stress. Biometric data is analyzed over time to detect abnormalities in heart rate and blood pressure.

[1595] Input: Data stored in a database.

[1596] Output: Analysis results.

[1597] Step 5:

[1598] The server integrates the analysis results and monitors the health status of the user and driver. Multimodal AI integrates the analysis results of each data type to assess the overall health status. From this analysis, anomalies are identified.

[1599] Input: Analysis results for each data type.

[1600] Output: Integrated health status assessment.

[1601] Step 6:

[1602] If the server detects an abnormality, it generates a notification and sends it to the user or driver. The server generates a notification message containing the details of the abnormality and recommended actions to take, and sends it to the user's smartphone or the vehicle's infotainment system.

[1603] Input: Integrated Health Status Assessment.

[1604] Output: Informational message.

[1605] Step 7:

[1606] The user or driver receives a notification and takes appropriate action, such as consulting a medical professional.

[1607] Input: Notification message.

[1608] Output: Coping behavior.

[1609] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1610] The present invention is a system that monitors a user's health condition by comprehensively analyzing data from multiple sensors used in the home. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it is possible to monitor the user's emotional health as well. Specific embodiments of the system will now be described.

[1611] System Overview

[1612] The system collects data from multiple sensors (image sensors, voice sensors, text input devices, and biometric sensors) installed in the user's home and sends it to a server. The server stores the received data and analyzes it using a multimodal artificial intelligence (AI) and emotion engine to monitor the user's health and emotional state. If an abnormality is detected, the server sends a notification to the user.

[1613] Program processing flow

[1614] Data collection

[1615] Devices (users' home sensors and wearable devices) collect data

[1616] Image data: Security cameras in the home periodically take images and videos, recording the user's movements, facial expressions, and daily life.

[1617] Voice data: Smart speakers and smart assistants collect your voice and use the recorded audio to analyze your tone of voice and speaking patterns.

[1618] Text data: Collects text data entered by users through smart assistants or dedicated applications, and analyzes changes in the content and expression of the text.

[1619] Biometric data: Wearable devices measure biometric data such as heart rate, blood pressure, oxygen saturation, and sleep patterns to continuously monitor the user's physical condition.

[1620] Data transmission

[1621] The device sends the collected data to the server.

[1622] The sensor terminals send the collected data to the server at regular intervals, and the data is encrypted using a security protocol.

[1623] Data analysis

[1624] The server analyzes the received data

[1625] Image analysis: The server uses image data to analyze changes in the user's movements and facial expressions, thereby detecting physical changes and patterns in daily activities.

[1626] Voice analysis: The server uses speech recognition technology to convert voice data into text and analyzes changes in tone and speaking style to identify emotional issues and changes in cognitive function.

[1627] Text analysis: The server uses natural language processing technology to analyze the text data and detect changes in the content and expressions of the text, which allows for an assessment of stress and mental health.

[1628] Biometric data analysis: The server uses time series analysis techniques to analyze data such as heart rate, blood pressure, oxygen saturation, and sleep patterns to detect abnormalities, allowing for early detection of cardiac and respiratory problems.

[1629] Emotion analysis

[1630] The server uses an emotion engine to analyze the user's emotions.

[1631] Emotion analysis using voice data: The server analyzes the voice data and uses an emotion engine to recognize the user's emotional state from changes in their tone of voice and speaking style.

[1632] Emotion analysis using text data: The server analyzes the text data and uses an emotion engine to recognize the emotional state from changes in the content and expression of the text.

[1633] Emotion analysis using image data: The server analyzes the image data and uses an emotion engine to recognize the user's emotional state from their facial expressions.

[1634] Health and emotional state monitoring

[1635] The server integrates the analysis results and monitors the health and emotional state.

[1636] By integrating the results of various data analyses, the multimodal AI and emotion engine assesses the user's overall health and emotional state, making it easier to detect abnormalities.

[1637] Send notifications

[1638] If the server detects an abnormality, it will notify the user.

[1639] If the server detects an anomaly, it generates a notification message and sends it to the user's smartphone or smart assistant, containing details of the anomaly and recommended actions to take.

[1640] Example: Emotion and health monitoring using voice data

[1641] The device (smart speaker) collects the user's voice from everyday conversations. The server analyzes the received voice data and converts it into text data using voice recognition technology. It then uses an emotion engine to analyze subtle changes in voice tone and speaking style to assess whether the user is feeling stressed. The server also uses multimodal AI to analyze heart rate data collected simultaneously with the voice data, detecting physical changes such as an abnormally high heart rate. When an abnormality is detected, the server immediately generates a notification and sends it to the user's smartphone, stating, "Your recent conversations have shown signs of stress, and your heart rate is also high. Please try some relaxation techniques." This allows for a prompt response.

[1642] In this way, the present invention can protect the overall health of users by monitoring both health and emotions in the home in real time and identifying risks early.

[1643] The processing flow will be explained below.

[1644] Step 1:

[1645] Devices (sensors) collect data. Home security cameras periodically capture images and videos, smart speakers and smart assistants record voice data, and text data entered by users through smart assistants and applications is also collected. Wearable devices also measure biometric data such as heart rate, blood pressure, oxygen saturation, and sleep patterns in real time.

[1646] Step 2:

[1647] The terminal (sensor) packetizes and encrypts the collected data, and the data packets are temporarily stored in a queue inside the terminal.

[1648] Step 3:

[1649] The device (sensor) sends data packets to the server at regular intervals. The data transmission is performed using a secure communication protocol to ensure the confidentiality and security of the data.

[1650] Step 4:

[1651] The server receives the data packet. The received data packet is decoded and restored to its original format. The data integrity is checked, and if there are any errors or omissions, a retransmission request is sent to the terminal.

[1652] Step 5:

[1653] The server stores the recovered data in a distributed database, where the data is managed for each user and tagged within the database.

[1654] Step 6:

[1655] The server analyzes the received image and video data. Image classification and motion analysis algorithms are used to detect changes in the user's movement patterns and facial expressions. For example, unusual movements or changes in facial expressions may be signs of an abnormality.

[1656] Step 7:

[1657] The server analyzes the voice data. It uses speech recognition technology to convert the data into text and analyzes changes in tone and speaking style. This allows it to detect emotional issues and changes in cognitive function. For example, a trembling voice or changes in speaking style could be a sign of stress or changes in cognitive function.

[1658] Step 8:

[1659] The server analyzes the text data. It uses natural language processing technology to analyze the content of the text and detect changes in the content and expressions of the sentences. This allows for an assessment of stress and mental health. For example, an increase in negative expressions in the sentences can be determined to be a sign of stress.

[1660] Step 9:

[1661] The server analyzes biometric data, using time series analysis techniques to analyze data such as heart rate, blood pressure, oxygen saturation, and sleep patterns to detect abnormalities. This allows for early detection of cardiac and respiratory problems. For example, a higher-than-normal heart rate could be a sign of cardiac disease.

[1662] Step 10:

[1663] The server uses an emotion engine to analyze the user's emotions. It recognizes the user's emotions from voice, text, and image data. This allows the system to monitor the user's emotional health. For example, if anger, excitement, or sadness is detected from the voice data, it is considered a sign of emotional instability.

[1664] Step 11:

[1665] The server integrates the results of each analysis, and uses a multimodal AI and emotion engine to comprehensively analyze images, voice, text, and biometric data to assess the user's overall health and emotional state, making it easier to detect abnormalities.

[1666] Step 12:

[1667] If the server detects an anomaly, it generates a notification with details about the anomaly and a recommended course of action, such as "Your heart rate has been higher than normal recently, and you appear to be feeling stressed."

[1668] Step 13:

[1669] The server generates a notification and sends it to the user's smartphone or smart assistant. The notification is sent immediately, informing the user that an abnormality has occurred.

[1670] Step 14:

[1671] Users receive notifications, review them, take action based on recommended actions, and, if necessary, seek medical advice. This allows users to quickly respond to changes in their health and emotional state.

[1672] This specific processing flow enables real-time monitoring of both health and emotions within the home, identifying risks early and protecting the overall health of users.

[1673] Example 2

[1674] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1675] There is a need for technology that can comprehensively monitor not only a user's health status but also their emotional state by collecting and analyzing different types of data from multiple sensors used in the home. However, existing systems have faced many challenges in centralized data management, complex data analysis, anomaly detection, and prompt notification to the user. In particular, it has been difficult to integrate information from multiple data sources and perform accurate analysis.

[1676] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for collecting data from sensors used in the home, means for transmitting the collected data to the server, means for receiving and storing the transmitted data, means for analyzing the stored data using multimodal artificial intelligence and an emotion analysis engine, means for integrating the analysis results to monitor the health status and emotional state, and means for generating and sending a notification to the user when an abnormality is detected. This makes it possible to comprehensively monitor the user's health status and emotional state in real time and to provide a prompt and appropriate notification when an abnormality occurs.

[1677] A "sensor" is a device used in the home that collects information such as images, audio, text, and biometric data.

[1678] "Data" refers to all information collected by sensors, such as image data, video data, audio data, text data, and biometric data.

[1679] A "server" is an information processing device that receives and stores data sent from sensors, and analyzes the data to monitor the user's condition.

[1680] "Multimodal AI" is an AI technology that has the ability to integrate and analyze multiple different types of data (e.g., images, audio, text, biometric data) in a unified manner.

[1681] An "emotion analysis engine" is a technology for analyzing a user's emotional state from data such as voice, text, and images.

[1682] "Analyzing" means that the server processes the data it receives and derives meaningful results.

[1683] "Monitoring" means that the server continuously observes and evaluates the user's health and emotional state based on the analysis results.

[1684] A "notification" is a warning or advice message that is generated and sent to the user when the server detects an abnormality.

[1685] An "anomaly" is any unusual change or problem in the user's health or emotional state that is detected as a result of the server's analysis.

[1686] "Integrating" means that the server compiles the results of multiple data analyses and provides a single comprehensive evaluation.

[1687] The present invention provides a system for monitoring a user's health and emotional state by comprehensively analyzing data from multiple sensors used in the home. A specific method for implementing this system is described below.

[1688] System configuration

[1689] The system collects data from multiple sensors installed in the home (image sensors, voice sensors, text input devices, and biometric sensors) and sends it to a server. The server stores the received data and analyzes it using a multimodal artificial intelligence (AI) and emotion analysis engine. The analysis results are integrated to monitor the user's health and emotional state. If an abnormality is detected, the server sends a notification to the user.

[1690] Hardware and software used

[1691] Sensors used in the home include:

[1692] Image sensor: Home security camera

[1693] Voice Sensors: Smart Speakers and Smart Assistants

[1694] Text input devices: smart assistants and dedicated applications

[1695] Biosensors: Wearable devices that measure heart rate, blood pressure, oxygen saturation, and sleep patterns

[1696] The server includes the following analysis modules:

[1697] Multimodal Artificial Intelligence (AI)

[1698] Sentiment Analysis Engine

[1699] Data Storage and Management System

[1700] Notification System

[1701] Data collection and transmission

[1702] 1. Devices (sensors) collect data

[1703] Image sensors capture images and videos of the home and record changes in the user's movements and facial expressions. For example, a camera in the living room can start recording at a set time.

[1704] The voice sensor uses smart speakers to collect everyday conversations and analyze tone and speech patterns, for example, recording the voice you use to ask about the weather.

[1705] The text input device collects text entered by the user through a dedicated application and analyzes changes in the content and expression of the text, for example, collecting diary entries from the user.

[1706] Biometric sensors use wearable devices to continuously measure heart rate, blood pressure, oxygen saturation, and sleep patterns. For example, a smartwatch worn by a user records heart rate.

[1707] 2. The device sends the collected data to the server

[1708] The sensor terminals transmit the collected data to a server at regular intervals. The transmitted data is encrypted using a security protocol. For example, audio data is uploaded to the server every hour.

[1709] Data analysis

[1710] 1. The server analyzes the received data

[1711] Image analysis: The server uses image data to analyze changes in the user's movements and facial expressions, for example, to detect changes in the user's walking pattern.

[1712] Voice analysis: The server uses speech recognition technology to convert the voice data into text and analyzes changes in tone and speaking style, for example detecting if the user's voice is darker than usual.

[1713] Text analysis: The server uses natural language processing technology to analyze the text data and detect changes in the content and expressions of the sentences. For example, it analyzes the frequent use of the expression "tired."

[1714] Biometric data analysis: The server uses time series analysis to analyze heart rate, blood pressure, oxygen saturation, sleep patterns, etc. to detect abnormalities, such as a higher than normal heart rate.

[1715] 2. Emotion analysis

[1716] Emotion analysis from voice data: The server analyzes the voice data and uses an emotion analysis engine to recognize the emotional state from changes in tone of voice and speaking style, for example, detecting an increase in anger.

[1717] Sentiment analysis based on text data: The server analyzes the text data and uses a sentiment analysis engine to recognize the emotional state from changes in the content and expressions of the text. For example, it detects an increase in pessimistic expressions.

[1718] Emotion analysis based on image data: The server analyzes the image data and uses an emotion analysis engine to recognize the user's emotional state from their facial expressions. For example, it detects that they are smiling less.

[1719] Integrated monitoring and abnormality notification

[1720] 1. The server integrates the analysis results and monitors the health and emotional state.

[1721] Using a multimodal AI and emotion analysis engine, the analysis results are integrated to assess the user's overall health and emotional state. For example, if different emotional signals are detected from image analysis, audio analysis, and text analysis, a comprehensive assessment is made.

[1722] 2. Detect anomalies and generate notifications

[1723] If an anomaly is detected, the server generates a real-time notification. For example, if heart rate and voice data indicate strong signs of stress, a notification will be sent to the user stating, "Your recent conversations have shown signs of stress, and your heart rate is high. Please try some relaxation techniques."

[1724] Specific examples

[1725] For example, a smart speaker collects a user's voice from everyday conversations and sends the data to a server. The server analyzes the voice data and converts it into text data using speech recognition technology. It uses an emotion analysis engine to analyze subtle changes in the tone of voice and speaking style to assess whether the user is feeling stressed. At the same time, the server also analyzes heart rate data from a wearable device and detects if the heart rate is higher than normal. If an abnormality is detected, the server sends the user a notification saying, "Your recent conversations have shown signs of stress, and your heart rate is high. Please try some relaxation techniques."

[1726] Example prompt: "Analyze the voice data collected by the smart speaker and assess whether the user is experiencing stress."

[1727] This invention can protect the overall health of users by monitoring both health and emotions in the home in real time and identifying risks early.

[1728] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1729] Step 1: Collect data

[1730] Data is collected by sensors and devices installed in the home by users.

[1731] Input: Various data from the home environment and the user (images, voice, text, biometric data).

[1732] How it works: Security cameras take pictures, smart speakers record everyday conversations, and wearable devices measure heart rate and blood pressure.

[1733] Output: Various data collected.

[1734] Step 2: Sending data

[1735] The terminal transmits the collected data to the server.

[1736] Input: Various types of collected data (images, audio, text, biometric data).

[1737] Specific operation: Image data captured by the security camera is uploaded to the server, and audio data recorded by the smart speaker is sent to the server.

[1738] Output: The data sent to the server.

[1739] Step 3: Save your data

[1740] The server stores the received data.

[1741] Input: The data sent to the server.

[1742] Specific operation: The server stores various data in the database.

[1743] Output: Saved data.

[1744] Step 4: Analyzing the image data

[1745] The server analyzes the image data.

[1746] Input: Saved image data.

[1747] Specific operation: Multimodal AI is used to analyze changes in user behavior and facial expressions from image data. For example, facial recognition algorithms are used to analyze user facial expressions.

[1748] Output: Analysis results (changes in user's behavior and facial expressions).

[1749] Step 5: Analyze the audio data

[1750] The server analyzes the audio data.

[1751] Input: Stored audio data.

[1752] What it does: It uses speech recognition technology to convert voice data into text, and uses an emotion analysis engine to analyze changes in tone and speaking style, for example, to identify changes in tone in a user's everyday conversation.

[1753] Output: Analysis results (changes in tone of voice and speaking style).

[1754] Step 6: Analyze the text data

[1755] The server parses the text data.

[1756] Input: Saved text data.

[1757] Specific operation: Analyzes text data using natural language processing technology and detects changes in the content and expressions of sentences using a sentiment analysis engine. For example, it analyzes the frequent use of the expression "tired."

[1758] Output: Analysis results (changes in the content and expression of the text).

[1759] Step 7: Analyzing the biometric data

[1760] The server analyzes the biometric data.

[1761] Input: Stored biometric data.

[1762] Specific operation: Using time series analysis technology, it analyzes heart rate, blood pressure, oxygen saturation, sleep patterns, etc. to detect abnormalities, such as when the heart rate is higher than normal.

[1763] Output: Analysis results (heart rate, blood pressure, oxygen saturation, changes in sleep patterns).

[1764] Step 8: Integrating the analysis results

[1765] The server consolidates the various analysis results.

[1766] Input: Analysis results of various data (images, audio, text, biometric data).

[1767] How it works: Using a multimodal AI and emotion analysis engine, the system integrates analysis results to assess the user's overall health and emotional state. For example, if different emotional indicators are found from image analysis, audio analysis, and text analysis, it will comprehensively assess them.

[1768] Output: Integrated analysis results (assessment of the user's health and emotional state).

[1769] Step 9: Detect anomalies

[1770] The server detects an abnormality.

[1771] Input: Consolidated analysis results.

[1772] Specific operation: Based on the analysis results, an anomaly detection algorithm is executed to detect abnormalities in the user's health and emotional state, such as stress and heart rate abnormalities.

[1773] Output: Anomaly detection results.

[1774] Step 10: Generate and send notifications

[1775] The server generates the notification and sends it to the user.

[1776] Input: Anomaly detection results.

[1777] What it does: If an abnormality is detected, it will generate a notification message and send it to the user's smartphone or smart assistant, for example, "Your recent conversations have shown signs of stress and your heart rate is elevated. Please try some relaxation techniques."

[1778] Output: The notification message sent to the user.

[1779] (Application example 2)

[1780] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1781] Conventional health monitoring systems in homes and businesses are limited to monitoring physical health conditions, making it difficult to respond quickly and accurately to changes or abnormalities in emotional states. Furthermore, the means of detecting abnormalities are limited, and detection of abnormalities can be delayed under certain conditions. Therefore, there is a need for a system that can monitor both health and emotional states in an integrated manner, quickly detect abnormalities, and provide appropriate notifications.

[1782] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting data from sensors used in homes and business facilities, means for transmitting the collected data to the server, means for receiving and storing the transmitted data, means for analyzing the stored data using multimodal artificial intelligence, means for integrating the analysis results to monitor health and emotional states, means for generating and transmitting a notification to the user when an abnormality is detected, and means for analyzing changes in the movements, facial expressions, and tone of voice of visitors and employees to identify their emotional states. This makes it possible to monitor not only health states but also emotional states in real time, quickly detect abnormalities, and issue appropriate notifications.

[1783] "Inside a home or business premises" refers to the physical environment in which a person lives or conducts business activities, and refers to the location of the system used within that particular space.

[1784] A "sensor" is a device that detects physical phenomena and converts them into electronic data, and includes image sensors, audio sensors, text input devices, and biometric sensors.

[1785] "Means for collecting data" refers to a system element that uses sensors to obtain information from an object and record it in electronic form.

[1786] "Means for transmitting data to a server" refers to the communications devices and protocols used to transfer the collected digital data over a network to a remote server.

[1787] "Means for receiving and storing data" refers to a system element that allows a server to receive data sent via a network and store it in a database or the like.

[1788] "Multimodal AI" refers to an AI technology that simultaneously analyzes multiple types of data (e.g., images, audio, text, biometric data) and makes integrated decisions.

[1789] "Health and emotional monitoring means" refers to system elements that provide continuous monitoring of an individual's physical and emotional health based on collected and analyzed data.

[1790] "Means for generating and sending notifications to users when an anomaly is detected" refers to the ability for the system to automatically create and send alerts or notifications to users' devices when an anomaly is detected.

[1791] "Means of identifying emotional states by analyzing changes in behavior, facial expressions, and tone of voice" refers to artificial intelligence technology that analyzes image and audio data to infer emotional states from behavior, changes in facial expressions, tone of voice, etc.

[1792] An embodiment of the present invention is described in detail below: The system monitors the health and emotional state of visitors and employees by collecting data from multiple sensors installed in homes and businesses, transmitting the data to a server for analysis.

[1793] System Configuration

[1794] Terminal (sensor)

[1795] The devices include image sensors, voice sensors, text input devices, and biometric sensors that are placed within physical stores to collect the following data from visitors and employees:

[1796] Image data: The camera records your movements and expressions.

[1797] Audio data: Microphones collect speech.

[1798] Text data: The digital device captures the input text.

[1799] Biometric data: Wearable devices and thermal cameras collect heart rate and temperature data.

[1800] The data collected by each terminal is encrypted and sent to a server at regular intervals.

[1801] Server Processing

[1802] Data reception and storage

[1803] The server receives the data sent from the terminal via the network in an encrypted form and stores it in a database, using technologies including security protocols and database systems.

[1804] Data analysis

[1805] The server analyzes the various types of stored data (images, audio, text, and biometric data) using multimodal artificial intelligence (AI). Image data is analyzed for changes in movements and facial expressions using libraries such as OpenCV, and audio data is converted into text using pyaudio and librosa, after which changes in tone of voice and speaking style are analyzed. Text data is analyzed for emotional state using natural language processing techniques such as TextBlob. Furthermore, biometric data is analyzed for abnormalities using time series analysis techniques.

[1806] Emotion analysis

[1807] The server uses an emotion engine to integrate data from different modalities and recognize the emotional state of visitors and employees, allowing real-time monitoring of changes in not only physical health but also emotional health.

[1808] Notifications and Feedback

[1809] If the server detects an abnormality, it automatically generates a notification message and sends it to the user's smartphone or other display device. The notification includes details of the detected abnormality and recommended actions to take. For example, a message might be generated that reads, "Employee A has a high body temperature and his tone of voice indicates stress. We recommend that he leave work early."

[1810] Examples and prompts

[1811] As a concrete example, consider a monitoring system for Employee A in a store. The server collects and analyzes Employee A's temperature and conversation data, and if an abnormality is detected, it sends a notification saying, "Employee A's temperature is high and his / her tone of voice also indicates stress. We recommend that you leave early." An example of how such a prompt is input into a generative AI model is as follows:

[1812] "Monitor employee health and emotions, and if you detect a high temperature or stressful tone in the conversation, generate a notification recommending that employees leave early. Could you please give us an example of what this notification should look like?"

[1813] This will enable the realization of a system that monitors health and emotions in an integrated manner and can respond quickly in the event of an abnormality.

[1814] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1815] Step 1:

[1816] The terminal collects image data, audio data, text data, and biometric data. Image data is obtained by recording the movements and facial expressions of visitors and employees with a camera, while audio data is obtained by collecting conversations and environmental sounds with a microphone. Text data is obtained by acquiring input information from digital terminals, and biometric data is obtained by collecting heart rate and body temperature information with a wearable device or a body temperature detection camera. The input is various data obtained from each sensor, and the output is the collected raw data.

[1817] Step 2:

[1818] The terminal sends the collected data to the server at regular intervals. The data is sent encrypted over the network. In this step, the input is the collected data, and the output is the encrypted data sent to the server.

[1819] Step 3:

[1820] The server saves the received data. The received data is stored in the server's database. The input of this step is the encrypted transmitted data, and the output is the saved data.

[1821] Step 4:

[1822] The server analyzes the stored data. The analysis is divided into the following parts:

[1823] Image analysis: Image data is analyzed using OpenCV etc. to detect changes in movements and facial expressions. The input is the saved image data, and the output is the analysis results of movement and facial expression changes.

[1824] Speech analysis: Using pyaudio and librosa, audio data is converted to text and changes in voice tone and speaking style are analyzed. The input is the stored audio data, and the output is the converted text data and changes in voice tone and speaking style.

[1825] Text Analysis: Analyzes text data using natural language processing techniques such as TextBlob to evaluate emotional states. The input is the stored text data, and the output is the analysis result of the emotional state.

[1826] Biometric data analysis: Detects abnormalities in biometric data (heart rate, body temperature, etc.) using time series analysis techniques. The input is stored biometric data, and the output is the anomaly detection results.

[1827] Step 5:

[1828] The server uses the emotion engine to integrate the analysis results of each data and monitor the health and emotional states of visitors and employees. The input of this step is each analysis result, and the output is the integrated evaluation result of the health and emotional states.

[1829] Step 6:

[1830] If the server detects an abnormality, it generates a notification message and sends it to the user's smartphone or other display device. The notification includes details of the abnormality and recommended actions to take. For example, a notification might be sent stating, "Employee A's temperature is high and his / her tone of voice indicates stress. We recommend that he / she leave work early." The input to this step is the integrated evaluation result, and the output is the generated notification message.

[1831] The above is the flow of processing in the embodiment of the present invention.

[1832] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1833] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1834] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1835] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1836] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1837] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1838] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1839] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1840] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1841] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1842] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1843] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1844] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1845] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1846] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1847] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1848] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1849] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1850] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1851] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1852] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1853] The following is further disclosed regarding the above embodiment.

[1854] (Claim 1)

[1855] a means of collecting data from sensors used within the home;

[1856] means for transmitting the collected data to a server;

[1857] means for receiving and storing the transmitted data;

[1858] means for analyzing the stored data using multimodal artificial intelligence;

[1859] A means of integrating the analysis results to monitor health status;

[1860] a means for generating and sending notifications to users when an anomaly is detected;

[1861] A system including:

[1862] (Claim 2)

[1863] 10. The system of claim 1, wherein at least one of image data, video data, audio data, text data, and biometric data is collected.

[1864] (Claim 3)

[1865] 10. The system of claim 1, further comprising means for analyzing the audio data to detect changes in tone of voice or speaking style to identify symptoms associated with emotional problems or changes in cognitive function.

[1866] "Example 1"

[1867] (Claim 1)

[1868] a means of collecting data from sensors used within the home;

[1869] means for transmitting the collected data to a server;

[1870] means for receiving and storing the transmitted data;

[1871] means for analyzing the stored data using multimodal artificial intelligence;

[1872] A means of integrating the various analyzed data and monitoring health status;

[1873] a means for generating and sending notifications to users when an anomaly is detected;

[1874] A system including:

[1875] (Claim 2)

[1876] 10. The system of claim 1, wherein at least one of image data, video data, audio data, text data, and biometric data is collected.

[1877] (Claim 3)

[1878] 10. The system of claim 1, further comprising means for analyzing the audio data to detect changes in tone of voice and speech patterns to identify symptoms associated with emotional problems and changes in cognitive function.

[1879] "Application Example 1"

[1880] (Claim 1)

[1881] a means of collecting data from sensors used within the home;

[1882] means for transmitting the collected data to a server;

[1883] means for receiving and storing the transmitted data;

[1884] means for analyzing the stored data using multimodal artificial intelligence;

[1885] A means of integrating the analysis results to monitor health status;

[1886] a means for generating and sending notifications to users when an anomaly is detected;

[1887] A means of collecting data from internal sensors in autonomous vehicles to monitor the health of the driver; and

[1888] A system including:

[1889] (Claim 2)

[1890] 10. The system of claim 1, wherein at least one of image data, video data, audio data, text data, and biometric data is collected.

[1891] (Claim 3)

[1892] 10. The system of claim 1, further comprising means for analyzing the audio data to detect changes in tone of voice or speaking style to identify symptoms associated with emotional problems or changes in cognitive function.

[1893] "Example 2: Combining Emotion Engines"

[1894] (Claim 1)

[1895] a means of collecting data from sensors used within the home;

[1896] means for transmitting the collected data to a server;

[1897] means for receiving and storing the transmitted data;

[1898] means for analyzing the stored data using a multimodal artificial intelligence and sentiment analysis engine;

[1899] a means for integrating the analytical results to monitor health and emotional states;

[1900] a means for generating and sending notifications to users when an anomaly is detected;

[1901] A system including:

[1902] (Claim 2)

[1903] 10. The system of claim 1, wherein at least one of image data, video data, audio data, text data, and biometric data is collected.

[1904] (Claim 3)

[1905] 10. The system of claim 1, further comprising means for analyzing the audio and text data to detect changes in tone of voice, delivery, content and expression of text, and to identify symptoms associated with emotional problems or changes in cognitive function.

[1906] "Application example 2 when combining emotion engines"

[1907] (Claim 1)

[1908] A means of collecting data from sensors used within homes and business premises;

[1909] means for transmitting the collected data to a server;

[1910] means for receiving and storing the transmitted data;

[1911] means for analyzing the stored data using multimodal artificial intelligence;

[1912] A means of integrating the analysis results to monitor health and emotional status;

[1913] a means for generating and sending notifications to users when an anomaly is detected;

[1914] A means for identifying emotional states of visitors and employees by analyzing changes in their movements, facial expressions, and tone of voice;

[1915] A system including:

[1916] (Claim 2)

[1917] 10. The system of claim 1, wherein at least one of image data, video data, audio data, text data, and biometric data is collected.

[1918] (Claim 3)

[1919] 10. The system of claim 1, further comprising means for analyzing the audio data to detect changes in tone of voice or speaking style to identify symptoms associated with emotional problems or changes in cognitive function. [Explanation of symbols]

[1920] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. a means of collecting data from sensors used within the home; means for transmitting the collected data to a server; means for receiving and storing the transmitted data; means for analyzing the stored data using multimodal artificial intelligence; A means of integrating the analysis results to monitor health status; a means for generating and sending notifications to users when an anomaly is detected; A system including:

2. 10. The system of claim 1, wherein at least one of image data, video data, audio data, text data, and biometric data is collected.

3. 10. The system of claim 1, further comprising means for analyzing the audio data to detect changes in tone of voice and speaking style to identify symptoms associated with emotional problems and changes in cognitive function.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A