system

A real-time monitoring system with integrated image and audio analysis, anomaly detection, and notification mechanisms addresses the challenges of childcare and nursing care by enhancing detection accuracy and speed, allowing caregivers to maintain flexible work schedules while ensuring safety.

JP2026074968APending Publication Date: 2026-05-07SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-21
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

The increasing burden of childcare and nursing care due to declining birthrates and aging populations has led to impaired flexible work styles and increased stress, as existing systems struggle to accurately and quickly detect anomalies in the behavior of children and the elderly, necessitating a more efficient monitoring solution.

Method used

A system that integrates real-time image and audio acquisition, analysis, anomaly detection, and notification mechanisms, utilizing machine learning to improve accuracy and speed of anomaly detection, and continuous data learning to enhance monitoring capabilities.

Benefits of technology

The system reduces the burden on caregivers by providing rapid and accurate anomaly detection, enabling flexible working styles and ensuring the safety of children and the elderly through immediate notifications and continuous learning for improved analysis accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026074968000001_ABST
    Figure 2026074968000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] Image acquisition method, An analysis means that analyzes acquired images to identify the subject's actions, A detection means for detecting anomalies based on the analysis results, A notification means that generates and sends a notification based on the detected anomaly, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In modern society, the burden of childcare and nursing care associated with the declining birthrate and aging population is increasing. As a result, there is a problem that the burden on working people has increased, and their flexible work styles and quality of life have been impaired. In addition, there is a continuous situation where one cannot always take one's eyes off to ensure the safety of children and the elderly, which further causes stress. It is required to improve such a situation and reduce the burden.

Means for Solving the Problems

[0005] This invention provides a system that includes an image acquisition means to acquire images of children and the elderly in real time, and an analysis means to identify the behavior of the subject by analyzing the images. Furthermore, it includes a detection means to detect anomalies based on the analysis results, and a notification means to quickly generate and transmit notifications in the event of an anomaly. By acquiring and analyzing audio as well, even more accurate anomaly detection is possible. In addition, by recording the analysis results and managing the history, the accuracy of the analysis can be improved through continuous data learning. In this way, it realizes the effect of reducing the burden in childcare and elderly care settings and supporting flexible working styles for those providing care.

[0006] "Image acquisition means" refers to a device or system for capturing images of a subject's actions and circumstances in real time and acquiring that image data.

[0007] "Analysis means" refers to a device or program that processes acquired image data and identifies and analyzes the subject's behavior and state.

[0008] A "detection means" is a device or mechanism for determining and detecting an anomaly based on the analyzed results and according to pre-set criteria.

[0009] A "notification means" is a device or system that uses detected anomaly information to generate and send alerts or notifications to relevant parties.

[0010] "Voice acquisition means" refers to a device or system for collecting voices emitted by a subject and ambient sounds in real time, and acquiring that voice data.

[0011] A "recording means" is a device or mechanism for saving analyzed information or detected events to a database or similar system so that they can be referenced later.

[0012] A "learning tool" is a device or program that performs continuous machine learning based on recorded data to improve the accuracy of analysis. [Brief explanation of the drawing]

[0013] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]

[0014] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0015] First, the terms used in the following description will be explained.

[0016] In the following embodiments, a labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0017] In the following embodiments, a labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0018] In the following embodiments, a labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0019] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0021] [First Embodiment]

[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0034] This invention provides a monitoring system to reduce the burden of childcare and elder care, and is configured as follows. This system integrates multiple means for acquiring, analyzing, detecting anomalies, and providing notifications for images.

[0035] First, the device continuously collects data using its camera and microphone at the location where the person being monitored is located. This records real-time images and audio of the person being monitored. The device then transmits this data to a server wirelessly or via a wired connection.

[0036] The server performs advanced image analysis based on the received image data. Specifically, it uses object detection and posture estimation algorithms to determine the actions and state of the subject. For example, it can recognize a child getting out of bed or an elderly person falling. For audio data, natural language processing technology is used to detect abnormal sounds and cries for help.

[0037] Next, the server determines whether an abnormal situation has been detected based on the analyzed information. It compares the results against pre-set criteria, and if it detects any actions or sounds that exceed the criteria, it treats them as abnormal. Based on this determination, it determines the designated emergency level and issues instructions for action.

[0038] When an anomaly is detected, the server immediately generates a notification. This notification includes the time, location, and nature of the anomaly, as well as recommended actions. This information is sent to the user via smartphone or a dedicated device. Receiving the notification allows the user to quickly respond to the situation or take appropriate action.

[0039] The system also records all analysis results and detection events on the server. This data is used to understand long-term activity history and as training data to improve analysis accuracy. The server is designed to continuously update its analysis algorithms using the recorded data, enabling more precise anomaly detection.

[0040] For example, if a child rolls over in their sleep at night, the device captures the movement, and if the server determines it to be a safe movement, no notification is sent. However, if the child makes a movement that suggests they might fall out of bed, a notification is immediately generated, an alert is sent to the user, and appropriate action is prompted.

[0041] The following describes the processing flow.

[0042] Step 1:

[0043] The device acquires image data at a rate of multiple frames per second using a camera in the room where the person being monitored is located, and simultaneously collects audio data using a microphone. This allows for the capture of real-time visual and audio information.

[0044] Step 2:

[0045] The terminal compresses the acquired image and audio data and sends it to the server with minimal latency. This is achieved by utilizing a streaming protocol to enable efficient data transfer.

[0046] Step 3:

[0047] The server analyzes the received images using a high-speed processing algorithm to analyze the subject's movements and posture. It utilizes object detection technology to identify actions related to the safety of children and the elderly (e.g., falls, getting out of bed).

[0048] Step 4:

[0049] The server analyzes the audio data and performs natural language processing and acoustic analysis. It detects when the subject makes unusual sounds (e.g., screams, cries for help) or when there are sudden changes in the sound of the environment.

[0050] Step 5:

[0051] Based on the analysis results, the server detects anomalies by referring to pre-configured anomaly criteria. It immediately evaluates whether the criteria are met, and if an anomaly is detected, it generates response instructions according to the urgency.

[0052] Step 6:

[0053] Based on the detection of an anomaly, the server creates a relevant notification message. This message includes the type and location of the anomaly, as well as recommended actions, and is prepared for sending to the user.

[0054] Step 7:

[0055] The server sends the generated notification to the user. The user receives the notification on their smartphone or dedicated device and can view it on the screen. This allows the user to take immediate action.

[0056] Step 8:

[0057] The server meticulously records all detected events and analysis results. This recorded data will be used for future analysis and learning processes.

[0058] Step 9:

[0059] The server uses the recorded data to update its machine learning algorithms. This improves the accuracy of the analysis and enables more appropriate anomaly detection.

[0060] (Example 1)

[0061] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0062] In childcare and eldercare settings, effectively monitoring the safety of those involved is crucial, but conventional methods have struggled to achieve sufficient accuracy and speed in detecting anomalies. Furthermore, efficient data processing and learning are required for long-term data management and improved analysis accuracy. Additionally, establishing criteria for identifying behavioral patterns and performing accurate anomaly detection has been a challenge.

[0063] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0064] In this invention, the server includes an image acquisition means, an audio data collection means for acquiring sound, and an evaluation means for detecting anomalies from the processed data. This enables rapid and accurate anomaly detection using images and sound.

[0065] "Image acquisition means" refers to a device or method for collecting image data for monitoring a subject.

[0066] "Information processing means" refers to a device or method for analyzing raw data acquired from cameras and microphones to identify abnormalities from the behavior and voice of a subject.

[0067] "Sound data collection means" refers to a device or method for acquiring and analyzing sound.

[0068] "Evaluation means" refers to a device or method for identifying behaviors or sounds that exceed a certain standard from analyzed data and for determining abnormalities.

[0069] "Information provision means" refers to a device or system for notifying the user of information regarding detected anomalies.

[0070] "Management means" refers to a device or method for recording analyzed data and maintaining it as a long-term history.

[0071] "Learning methods" refer to algorithms and methods that utilize recorded data to improve the accuracy of analysis.

[0072] "Behavioral identification means" refers to a device or method for analyzing behavioral patterns based on the subject's state.

[0073] "Criteria setting means" refers to a device or method for setting criteria for detecting abnormalities based on information obtained by the behavior identification means.

[0074] This invention is a monitoring system designed to alleviate the burden of childcare and elder care. This system integrates multiple functions to collect, analyze, and notify images and audio in real time for monitoring the target individual. The following describes its specific embodiments.

[0075] The device is installed where the subject is located and continuously collects data using a camera and microphone. A standard digital camera capable of capturing high-resolution images is used as the camera, and a microphone with noise-canceling capabilities is recommended. The collected data is transmitted to a server via wireless or wired communication.

[0076] The server uses a common image analysis library, which is a machine learning model, to analyze the received image data. For example, a framework like TENSORFLOW® is used for object detection and pose estimation. This allows the server to identify specific actions such as a person getting up or sitting down. For audio data, speech recognition technology such as Google® Cloud Speech-to-Text is used to detect abnormal sounds and urgent voices through natural language processing.

[0077] If an anomaly is detected, the server generates a notification and sends it to the user via smartphone or a dedicated device. This notification includes the time and location of the anomaly, detailed information, and recommended actions. This allows the user to quickly understand the situation and take appropriate measures.

[0078] Furthermore, the server records all analysis results and detection events. This recorded data is used to update the server's learning algorithms and improve analysis accuracy. Specifically, by analyzing behavioral patterns using past data, the accuracy of detecting typical anomalies can be improved.

[0079] For example, if a child rolls over in their sleep at night, the device captures this movement, and if the server determines that the action is normal, no notification is sent. However, if the movement is dangerous, an alert is immediately generated and sent to the user.

[0080] Examples of prompt statements include:

[0081] "Please explain a real-time anomaly detection system for monitoring the safety of children and the elderly."

[0082] This is one possible explanation.

[0083] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0084] Step 1:

[0085] The device continuously collects data using a camera and microphone. Input consists of image and audio data containing the subject's activities. The device captures this data in real time and transmits it to a server wirelessly or via a wired connection. Specifically, it photographs and records moving children or elderly people speaking. Output consists of the transmitted high-resolution image and audio files.

[0086] Step 2:

[0087] The server receives image data transmitted from the terminal. Using the received image as input, it performs object detection and pose estimation using information processing tools. Using an image analysis library such as TensorFlow, it performs action identification to determine whether the subject is standing or sitting. Specifically, this action involves detecting a particular pose or movement of the human body within the image. The output is the type of action and action information.

[0088] Step 3:

[0089] The server also receives audio data. Using the received audio file as input, it converts it to text using speech recognition technology and detects anomalies and important phrases using natural language processing. Specifically, it detects abnormal sounds such as "help" or "I fell." The output is the detected key phrases and types of sounds.

[0090] Step 4:

[0091] The server analyzes information obtained from images and audio using evaluation tools to determine anomalies. Inputs are behavioral and audio information. The analyzed data is compared with established criteria to identify behaviors or sounds deemed abnormal. The specific action involves checking whether the behavior or sound anomaly exceeds the specified threshold. The output is the anomaly detection result and its detailed information.

[0092] Step 5:

[0093] The server generates notifications based on detected anomalies. The input is the result of the anomaly detection. The server uses this information to create a notification to send an alert to the user. Specifically, it specifies the details of the anomaly, the time it occurred, and recommended actions. The output is the notification message sent to the user.

[0094] Step 6:

[0095] Users receive notifications on their smartphones or dedicated devices. The input is the notification sent from the server. Users check the notification and respond quickly, such as checking camera footage if necessary. Specifically, actions include checking the notification and rushing to the scene if required. The output is the user's appropriate response.

[0096] Step 7:

[0097] The server records all analysis results and detection events in a database. Inputs are behavioral identification information and anomaly detection results. This data is managed as history and used as training data to improve future analysis accuracy. Specifically, it continuously incorporates the subject's daily behavioral patterns into the analysis. The output is historical activity history data.

[0098] (Application Example 1)

[0099] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0100] In childcare and elder care, ensuring the safety of those being monitored requires the rapid and accurate detection of abnormal behavior or conditions, and immediate notification. Conventional systems have limitations in the accuracy of anomaly detection and the speed of notification, which can lead to delays in on-site responses, thus necessitating further improvements.

[0101] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0102] In this invention, the server includes an image acquisition means, an analysis means for analyzing the acquired images to identify the subject's actions, a detection means for detecting anomalies based on the analysis results, and a display means for outputting notifications visually and audibly. This enables improved accuracy in anomaly detection and rapid notification.

[0103] "Image acquisition means" refers to a device or function for acquiring images of a person being monitored using a visual sensor such as a camera.

[0104] "Analysis means" refers to algorithms and technologies used to analyze acquired image and audio data and identify the subject's behavior or abnormalities.

[0105] "Detection means" refers to the process of identifying anomalies from information obtained by analysis means and determining their nature.

[0106] "Notification means" refers to a function or device for providing visual or audible notifications to the user based on anomalies detected by detection means.

[0107] "Display means" refers to devices such as displays and speakers that provide notifications to the user, and are responsible for outputting information visually and audibly.

[0108] In order to implement this invention, it is first necessary to construct the entire system. The system combines hardware and software to monitor the subject in real time and detect any abnormalities.

[0109] The "server" is a computer equipped with a high-performance processor for image analysis, and has software frameworks such as Python, TensorFlow, and OpenCV installed. The server receives image and audio data sent from terminals and analyzes them. For images, it performs object detection and pose estimation to identify the actions and state of the subject. For audio data, it applies natural language processing techniques to detect abnormal sounds and cries for help.

[0110] A "terminal" is a device equipped with a camera and microphone, which is installed in the environment where the subject is located. The terminal collects the subject's image and audio data in real time and transmits it to a server wirelessly or via a wired connection.

[0111] The "user" carries a display device such as smart glasses or a smartphone and receives notifications from the server. The notification includes the nature of the anomaly, the time and location of the occurrence, and recommended actions. This allows the user to quickly respond to the scene and take appropriate action.

[0112] For example, when used in a nursing home, if a caregiver is wearing smart glasses and unnatural movements are detected in an elderly person, a visual notification and an audible alert will be displayed on the glasses. This allows the caregiver to intervene immediately and prevent injuries.

[0113] An example of a prompt message to be input to the generating AI model is, "Detect any suspicious behavior from elderly individuals. If there is a risk of falling, send an alert." In this way, the present invention enables efficient monitoring and rapid response.

[0114] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0115] Step 1:

[0116] The device acquires images and audio of the subject using its camera and microphone. This data is stored in the device's temporary memory. The input for this step is real-time acquired images and audio, and the output is digital image and audio data.

[0117] Step 2:

[0118] The terminal transmits the acquired image and audio data to the server wirelessly or via a wired connection. Network communication technology is used for transmission. The input for this step is the image and audio data saved in step 1, and the output is the data transmitted to the server.

[0119] Step 3:

[0120] The server performs image analysis using TensorFlow and OpenCV with the received image data. Object detection and pose estimation algorithms are used to identify the subject's actions and posture. The input for this step is image data sent from the terminal, and the output is the analysis results regarding actions and posture.

[0121] Step 4:

[0122] The server analyzes the received audio data using natural language processing techniques to detect abnormal voice patterns and cries for help. The input for this step is audio data sent from the terminal, and the output is the analysis results regarding abnormal sounds and voice commands.

[0123] Step 5:

[0124] The server detects anomalies based on the analysis results of images and audio. It compares these results to pre-set criteria and determines an anomaly if the threshold is exceeded. The input for this step is the analysis results from steps 3 and 4, and the output is the determination of whether an anomaly occurred.

[0125] Step 6:

[0126] The server generates a notification based on the detected anomaly and sends it to the user's device. The notification includes the nature, time, and location of the anomaly, as well as recommended actions. The input for this step is the anomaly detection result from step 5, and the output is the notification information sent to the user.

[0127] Step 7:

[0128] Users can view and confirm notifications using smart glasses or smartphones and take action. The input for this step is notification information from the server, and the output is the user's response.

[0129] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0130] This invention is a system for enhancing monitoring activities in childcare and elderly care, incorporating an emotion engine that analyzes the emotional state of the subject in addition to effective anomaly detection. This system acquires and analyzes image and audio data to detect anomalies from normal behavior and sounds, and further enables complementary anomaly detection by identifying and evaluating the user's emotions using the emotion engine.

[0131] First, the device uses its camera and microphone to acquire images and audio data of the subject in real time. This data is immediately transmitted to the server. High-speed and stable communication methods are used for data transfer.

[0132] Next, the server analyzes the subject's movements from the received image data and detects abnormal patterns or cries for help from the audio data. Based on these analyses, it determines whether the subject is in a dangerous situation or exhibiting abnormal behavior.

[0133] Furthermore, the server uses an emotion engine to analyze the emotional state from the acquired images and audio. For example, through facial recognition and voice tone analysis, it identifies whether the subject is experiencing anxiety, anger, sadness, or other similar emotions. This enables anomaly detection that also addresses the subject's psychological state.

[0134] If an anomaly or a specific emotional state is detected, the server generates a notification based on a series of pieces of information. This notification includes an alert with adjusted urgency, including a risk assessment based on the subject's emotional state. The notification is sent to the user's smartphone or dedicated device, allowing the user to take immediate action.

[0135] For example, if an elderly person requiring special attention falls, the device captures the event, and the server detects the fall based on motion analysis. Simultaneously, an emotion engine senses anxiety from the elderly person's voice and facial expressions. The server then combines this information to generate a high-urgency alert, which is sent to the user to support a rapid response.

[0136] The system of this invention implements a continuous learning process using recorded analysis results to improve analysis accuracy. Furthermore, all history, including emotional data, is managed and utilized for future activities and analyses. This makes it possible to achieve both the safety of the subjects and a reduction in user burden.

[0137] The following describes the processing flow.

[0138] Step 1:

[0139] The device acquires continuous image data of the subject using its camera and collects audio data in real time using its microphone. This allows the subject's situation to be recorded in real time.

[0140] Step 2:

[0141] The terminal compresses the acquired data and sends it to the server using a high-speed communication protocol. This minimizes delays caused by data transfer, enabling real-time analysis.

[0142] Step 3:

[0143] The server uses the received image data to execute a motion analysis algorithm. It classifies the subject's posture and movements and identifies abnormal movements. For example, it can identify movements such as falls and slipping away.

[0144] Step 4:

[0145] The server analyzes the audio data and uses acoustic features and natural language processing to detect unusual sounds and specific calls. For example, loud calls and long silences may be identified as unusual.

[0146] Step 5:

[0147] The server activates an emotion engine, analyzing image and audio data to assess emotional states. Based on facial expression analysis and voice tone, it identifies emotions such as anxiety, anger, and sadness.

[0148] Step 6:

[0149] The server comprehensively evaluates the results of motion analysis, voice analysis, and emotional state to confirm the presence of an anomaly. If an anomaly is detected, the urgency level is set according to the emotional state.

[0150] Step 7:

[0151] The server generates a notification message based on the configured urgency level. The message includes details of the anomaly, the time and location of the occurrence, and recommended actions. The notification is sent to the user immediately.

[0152] Step 8:

[0153] Users can check received notifications on their smartphones or dedicated devices and decide on the necessary actions. For example, they can immediately head to the scene or check the video feed.

[0154] Step 9:

[0155] The server records and manages all analysis and detection results over the long term. This data is used to improve the accuracy of the analysis and to periodically update the machine learning algorithms.

[0156] (Example 2)

[0157] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0158] In monitoring activities for childcare and elderly care, there is a need to not only detect abnormal behavior in the person being monitored, but also to analyze their emotional state, enabling more comprehensive and rapid detection and response to abnormalities. However, conventional systems have the problem of being unable to adequately ensure the safety of the person being monitored because detailed analysis, including emotional state, is difficult.

[0159] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0160] In this invention, the server includes an image acquisition means, an audio acquisition means, and an emotion analysis means for analyzing the emotional state using the acquired image and audio data. This enables analysis of both the subject's behavior and emotional state, allowing for more accurate and rapid anomaly detection and notification.

[0161] "Image acquisition means" refers to a device or method used to acquire images of a subject, such as a camera, which collects data in real time.

[0162] "Voice acquisition means" refers to a device or method used to acquire the voice of a subject, and involves collecting voice data in real time using a microphone or the like.

[0163] "Analysis means" refers to the process and techniques for analyzing acquired data and identifying abnormalities from the subject's behavior and voice.

[0164] "Emotional analysis means" refers to the process and technology for identifying and analyzing a subject's emotional state using acquired image and audio data.

[0165] "Detection means" refers to the processes and technologies used to discover and judge abnormal phenomena or specific emotional states based on analyzed data.

[0166] "Notification means" refers to a mechanism for generating appropriate notifications based on detected anomalies or emotional states and sending relevant information to the user.

[0167] "Recording means" refers to a system for saving analyzed data and managing it as a history.

[0168] "Learning methods" refer to processes and techniques that utilize recorded data to improve the accuracy of system analysis.

[0169] This invention is a system for enhancing monitoring activities in childcare and elderly care. It analyzes image and audio data acquired in real time to determine the behavior and emotional state of the subject, thereby detecting and notifying of abnormalities.

[0170] First, the device uses its camera and microphone to acquire images and audio data of the subject. The camera is high-resolution, capturing the subject's face and movements in detail. The microphone is highly sensitive, capturing the subject's voice clearly. The acquired data is immediately transmitted to the server via a network with high communication speed and stability.

[0171] The server analyzes the subject's movements by executing a deep learning-based image recognition algorithm on the received image data. In addition, it uses speech recognition software to detect abnormal patterns and cries for help in the audio data. Furthermore, an emotion engine is incorporated, which analyzes the subject's emotional state based on the acquired data through facial recognition technology and speech analysis. This analysis allows for the identification of emotions such as anxiety, anger, and sadness, enabling anomaly detection that also takes psychological circumstances into account.

[0172] If an anomaly or a specific emotional state is detected, the server generates a notification based on the detected information and sends an alert to the user's smartphone or dedicated terminal. This notification includes a risk assessment based on the person's condition, allowing the user to understand the urgency of the situation and take immediate action.

[0173] For example, if an elderly person with dementia falls indoors, the device captures changes in their movements, and the server recognizes the fall from the image data. Simultaneously, it analyzes voice data for signs of pain, and the emotion engine detects feelings of anxiety. From this series of pieces of information, the server immediately generates a high-urgency alert and notifies the user.

[0174] This system features a learning function that continuously improves analysis accuracy using recorded analysis data. By using a generative AI model to execute prompts such as, "Analyze the subject's current condition in detail and report immediately if there are any abnormalities," it enables even more precise monitoring and analysis.

[0175] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0176] Step 1:

[0177] The device uses a camera to acquire images of the subject in real time and a microphone to collect audio data. Physical images and audio are provided as input, which are converted into digital data. The output is the digitized image and audio data. This data is then sent to a server for further analysis.

[0178] Step 2:

[0179] The terminal transmits the acquired image and audio data to the server via a high-bandwidth network. The input is the image and audio data acquired by the terminal, which is then packetized and transmitted. The output is the packetized data received on the server side.

[0180] Step 3:

[0181] The server applies a deep learning-based image recognition algorithm based on the received image data. The input is digitized image data, which is analyzed to identify the subject's movements. The output is data containing the results of the movement analysis.

[0182] Step 4:

[0183] The server analyzes audio data using a speech recognition algorithm. The input is digital audio data, which is analyzed to identify unusual patterns and specific speech content. The output is data containing the results of the speech analysis.

[0184] Step 5:

[0185] The server uses an emotion engine to analyze emotional states from image and audio data. The input is the analysis results of the images and audio, which are used to evaluate the emotional state. The output is data containing the results of the emotion analysis.

[0186] Step 6:

[0187] The server integrates motion analysis results, voice analysis results, and emotion analysis results to detect abnormal conditions. The input is a dataset containing the results of each analysis, which is then analyzed to perform a risk assessment. The output is alert information containing the results of the anomaly detection.

[0188] Step 7:

[0189] The server generates notifications based on detected anomalies and emotional states and sends them to the user's smartphone or dedicated terminal. The input is notification generation data based on anomaly detection and emotional evaluation, and the output is the notification information received by the user.

[0190] (Application Example 2)

[0191] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0192] In modern society, there is a need for means to ensure personal safety. In particular, when going out or acting alone, it is important to be able to understand the surrounding environment in real time and quickly detect the occurrence of abnormalities or psychological changes. However, with current technologies, anomaly detection and emotion analysis are performed separately, and there is insufficient means to respond immediately based on integrated information.

[0193] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0194] In this invention, the server includes an image acquisition means, an analysis means for analyzing the acquired images to identify the subject's behavior, a detection means for detecting anomalies based on the analysis results, an evaluation means that works in conjunction with the analysis means to analyze the emotional state, an adjustment means for performing a risk assessment based on the emotional state and adjusting the notification content, and a monitoring means for monitoring the surrounding situation. This ensures the safety of individuals and enables a quick and appropriate response.

[0195] "Image acquisition means" refers to a means of acquiring image data of a target in real time using an imaging device such as a camera.

[0196] "Analysis means" refers to methods for analyzing acquired image and audio data to identify the subject's behavior and emotional state.

[0197] A "detection means" is a means of detecting anomalies based on analyzed data and determining whether or not there is a problem.

[0198] A "notification system" is a means of generating and sending appropriate alerts based on detected anomalies or risks.

[0199] "Evaluation methods" refer to means of analyzing the emotional state of a subject from acquired data and identifying their psychological state.

[0200] "Adjustment measures" refer to methods for evaluating risk based on the results of an analysis of emotional states and adjusting notification content according to urgency.

[0201] "Monitoring means" refers to methods for constantly monitoring the user's surroundings and detecting abnormalities or dangers.

[0202] "Voice acquisition means" refers to a means of acquiring target voice data using a microphone.

[0203] "Display means" refers to means of visually showing users warnings or notifications regarding detected anomalies.

[0204] A "recording means" is a means of saving the analyzed data as a history for later review or learning.

[0205] A "learning method" is a means of continuously optimizing a system to improve the accuracy of analysis based on recorded data.

[0206] The system for implementing this invention includes smart glasses worn by the user and a data analysis device on a server. The smart glasses are equipped with a camera and microphone, which acquire images and audio data of the surroundings in real time. This data is immediately transmitted to the server via a stable communication means.

[0207] The server uses image analysis software and audio analysis software to analyze the received image and audio data. This allows the server to analyze the subject's behavior and surrounding environment and detect anomalies. Furthermore, to analyze emotional states, it uses a dedicated emotion engine to perform facial recognition and voice tone analysis, thereby evaluating the subject's psychological state.

[0208] If an anomaly is detected, or if sentiment analysis determines that there is a high psychological risk, the server will immediately generate a warning using the notification system and display it on the user's smart glasses display. This notification not only encourages the user to take safe actions but also helps them to take prompt action if necessary.

[0209] As a concrete example, imagine a situation where a user is walking alone at night and suddenly hears loud footsteps around them. In such a case, the smart glasses immediately send data to the server, detect the anomaly, and then visually display a warning such as "Watch your back."

[0210] An example of a prompt for the generating AI model might be, "Explain how to analyze anomalies and emotions in real time from audio and image data to ensure safety." This system provides an effective method for improving user safety.

[0211] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0212] Step 1:

[0213] When a user puts on smart glasses, the device activates its camera and microphone to capture images and sounds from the surroundings in real time. Based on the input, it outputs streams of image and audio data. The device periodically sends this data to a server.

[0214] Step 2:

[0215] The server inputs the received image data into image analysis software to identify the subject's actions and surrounding environment. This allows for the analysis of behavioral patterns such as walking and the approach of others. The output is the analyzed behavioral information.

[0216] Step 3:

[0217] Simultaneously, the server processes the audio data using audio analysis software to detect abnormal sounds and volume changes. For example, this could include the sound of footsteps suddenly approaching. The input is audio data, and the output is the detected abnormal sound information.

[0218] Step 4:

[0219] The server integrates the results of image and audio analysis and uses an emotion engine to evaluate the subject's emotional state. It identifies anxiety and surprise through facial recognition and voice tone analysis. The output is information evaluating the emotional state.

[0220] Step 5:

[0221] The server performs a risk assessment based on all analysis results and generates necessary alerts. For example, it might generate a notification such as "Watch your back." The input is the analysis results, and the output is the alert information.

[0222] Step 6:

[0223] The generated notification is displayed instantly on the user's smart glasses using the device's interface. The user can then see it and take action to pay attention to their surroundings. The output is a visual warning.

[0224] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0225] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0226] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0227] [Second Embodiment]

[0228] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0229] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0230] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0231] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0232] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0233] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0234] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0235] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0236] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0237] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0238] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0239] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0240] This invention provides a monitoring system to reduce the burden of childcare and elder care, and is configured as follows. This system integrates multiple means for acquiring, analyzing, detecting anomalies, and providing notifications for images.

[0241] First, the device continuously collects data using its camera and microphone at the location where the person being monitored is located. This records real-time images and audio of the person being monitored. The device then transmits this data to a server wirelessly or via a wired connection.

[0242] The server performs advanced image analysis based on the received image data. Specifically, it uses object detection and posture estimation algorithms to determine the actions and state of the subject. For example, it can recognize a child getting out of bed or an elderly person falling. For audio data, natural language processing technology is used to detect abnormal sounds and cries for help.

[0243] Next, the server determines whether an abnormal situation has been detected based on the analyzed information. It compares the results against pre-set criteria, and if it detects any actions or sounds that exceed the criteria, it treats them as abnormal. Based on this determination, it determines the designated emergency level and issues instructions for action.

[0244] When an anomaly is detected, the server immediately generates a notification. This notification includes the time, location, and nature of the anomaly, as well as recommended actions. This information is sent to the user via smartphone or a dedicated device. Receiving the notification allows the user to quickly respond to the situation or take appropriate action.

[0245] The system also records all analysis results and detection events on the server. This data is used to understand long-term activity history and as training data to improve analysis accuracy. The server is designed to continuously update its analysis algorithms using the recorded data, enabling more precise anomaly detection.

[0246] For example, if a child rolls over in their sleep at night, the device captures the movement, and if the server determines it to be a safe movement, no notification is sent. However, if the child makes a movement that suggests they might fall out of bed, a notification is immediately generated, an alert is sent to the user, and appropriate action is prompted.

[0247] The following describes the processing flow.

[0248] Step 1:

[0249] The device acquires image data at a rate of multiple frames per second using a camera in the room where the person being monitored is located, and simultaneously collects audio data using a microphone. This allows for the capture of real-time visual and audio information.

[0250] Step 2:

[0251] The terminal compresses the acquired image and audio data and sends it to the server with minimal latency. This is achieved by utilizing a streaming protocol to enable efficient data transfer.

[0252] Step 3:

[0253] The server analyzes the received images using a high-speed processing algorithm to analyze the subject's movements and posture. It utilizes object detection technology to identify actions related to the safety of children and the elderly (e.g., falls, getting out of bed).

[0254] Step 4:

[0255] The server analyzes the audio data and performs natural language processing and acoustic analysis. It detects when the subject makes unusual sounds (e.g., screams, cries for help) or when there are sudden changes in the sound of the environment.

[0256] Step 5:

[0257] Based on the analysis results, the server detects anomalies by referring to pre-configured anomaly criteria. It immediately evaluates whether the criteria are met, and if an anomaly is detected, it generates response instructions according to the urgency.

[0258] Step 6:

[0259] Based on the detection of an anomaly, the server creates a relevant notification message. This message includes the type and location of the anomaly, as well as recommended actions, and is prepared for sending to the user.

[0260] Step 7:

[0261] The server sends the generated notification to the user. The user receives the notification on their smartphone or dedicated device and can view it on the screen. This allows the user to take immediate action.

[0262] Step 8:

[0263] The server meticulously records all detected events and analysis results. This recorded data will be used for future analysis and learning processes.

[0264] Step 9:

[0265] The server uses the recorded data to update its machine learning algorithms. This improves the accuracy of the analysis and enables more appropriate anomaly detection.

[0266] (Example 1)

[0267] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0268] In childcare and eldercare settings, effectively monitoring the safety of those involved is crucial, but conventional methods have struggled to achieve sufficient accuracy and speed in detecting anomalies. Furthermore, efficient data processing and learning are required for long-term data management and improved analysis accuracy. Additionally, establishing criteria for identifying behavioral patterns and performing accurate anomaly detection has been a challenge.

[0269] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0270] In this invention, the server includes an image acquisition means, an audio data collection means for acquiring sound, and an evaluation means for detecting anomalies from the processed data. This enables rapid and accurate anomaly detection using images and sound.

[0271] "Image acquisition means" refers to a device or method for collecting image data for monitoring a subject.

[0272] "Information processing means" refers to a device or method for analyzing raw data acquired from cameras and microphones to identify abnormalities from the behavior and voice of a subject.

[0273] "Sound data collection means" refers to a device or method for acquiring and analyzing sound.

[0274] "Evaluation means" refers to a device or method for identifying behaviors or sounds that exceed a certain standard from analyzed data and for determining abnormalities.

[0275] "Information provision means" refers to a device or system for notifying the user of information regarding detected anomalies.

[0276] "Management means" refers to a device or method for recording analyzed data and maintaining it as a long-term history.

[0277] "Learning methods" refer to algorithms and methods that utilize recorded data to improve the accuracy of analysis.

[0278] "Behavioral identification means" refers to a device or method for analyzing behavioral patterns based on the subject's state.

[0279] "Criteria setting means" refers to a device or method for setting criteria for detecting abnormalities based on information obtained by the behavior identification means.

[0280] This invention is a monitoring system designed to alleviate the burden of childcare and elder care. This system integrates multiple functions to collect, analyze, and notify images and audio in real time for monitoring the target individual. The following describes its specific embodiments.

[0281] The terminal is installed at the location where the target person is, and constantly collects data using a camera and a microphone. A general digital camera that can capture high-resolution images is used for the camera, and a type with a noise cancellation function is recommended for the microphone. The collected data is transmitted to the server via wireless or wired communication.

[0282] The server uses a general image analysis library, which is a machine learning model, to analyze the received image data. For example, frameworks such as TensorFlow are used for object detection and pose estimation. This enables the identification of specific actions such as the target person getting up or sitting down. For voice data, speech recognition technologies such as Google Cloud Speech-to-Text are used to detect abnormal sounds or urgent voices through natural language processing.

[0283] When an abnormality is detected, the server generates a notification and sends it to the user via a smartphone or a dedicated device. This notification includes the time and location where the abnormality occurred, the detailed content, as well as recommended countermeasures. This enables the user to quickly grasp the situation and take appropriate measures.

[0284] Furthermore, the server records all the analysis results and detection events. This recorded data is used to update the learning algorithm in the server and is utilized to improve the analysis accuracy. Specifically, by analyzing the operation patterns using past data, the detection accuracy of typical abnormalities can be enhanced.

[0285] As a specific example, when a child turns over at night, if the terminal captures it and the server determines that the action is normal, no notification is sent. However, if a dangerous movement is made, an alert is immediately generated and sent to the user.

[0286] Examples of prompt texts include

[0287] "Please explain a real-time anomaly detection system for monitoring the safety of children and the elderly."

[0288] This is one possible explanation.

[0289] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0290] Step 1:

[0291] The device continuously collects data using a camera and microphone. Input consists of image and audio data containing the subject's activities. The device captures this data in real time and transmits it to a server wirelessly or via a wired connection. Specifically, it photographs and records moving children or elderly people speaking. Output consists of the transmitted high-resolution image and audio files.

[0292] Step 2:

[0293] The server receives image data transmitted from the terminal. Using the received image as input, it performs object detection and pose estimation using information processing tools. Using an image analysis library such as TensorFlow, it performs action identification to determine whether the subject is standing or sitting. Specifically, this action involves detecting a particular pose or movement of the human body within the image. The output is the type of action and action information.

[0294] Step 3:

[0295] The server also receives audio data. Using the received audio file as input, it converts it to text using speech recognition technology and detects anomalies and important phrases using natural language processing. Specifically, it detects abnormal sounds such as "help" or "I fell." The output is the detected key phrases and types of sounds.

[0296] Step 4:

[0297] The server analyzes information obtained from images and audio using evaluation tools to determine anomalies. Inputs are behavioral and audio information. The analyzed data is compared with established criteria to identify behaviors or sounds deemed abnormal. The specific action involves checking whether the behavior or sound anomaly exceeds the specified threshold. The output is the anomaly detection result and its detailed information.

[0298] Step 5:

[0299] The server generates notifications based on detected anomalies. The input is the result of the anomaly detection. The server uses this information to create a notification to send an alert to the user. Specifically, it specifies the details of the anomaly, the time it occurred, and recommended actions. The output is the notification message sent to the user.

[0300] Step 6:

[0301] Users receive notifications on their smartphones or dedicated devices. The input is the notification sent from the server. Users check the notification and respond quickly, such as checking camera footage if necessary. Specifically, actions include checking the notification and rushing to the scene if required. The output is the user's appropriate response.

[0302] Step 7:

[0303] The server records all analysis results and detection events in a database. Inputs are behavioral identification information and anomaly detection results. This data is managed as history and used as training data to improve future analysis accuracy. Specifically, it continuously incorporates the subject's daily behavioral patterns into the analysis. The output is historical activity history data.

[0304] (Application Example 1)

[0305] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0306] In childcare and nursing care, in order to ensure the safety of the person being watched over, it is required to quickly and accurately detect the abnormal behavior and condition of the person and immediately receive a corresponding notification. In conventional systems, there are limitations in the accuracy of abnormal detection and the speed of notification, and especially when the on-site response is delayed, further improvement is necessary.

[0307] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0308] In this invention, the server includes an image acquisition means, an analysis means for analyzing the acquired image to identify the behavior of the person being watched over, a detection means for detecting an abnormality based on the analysis result, and a display means for outputting a notification visually and audibly. Thereby, it becomes possible to improve the accuracy of abnormal detection and to provide a quick notification.

[0309] The "image acquisition means" is a device or function for acquiring an image of the person being watched over using a visual sensor such as a camera.

[0310] The "analysis means" refers to an algorithm or technology for analyzing the acquired image and voice data to identify the behavior and abnormality of the person being watched over.

[0311] The "detection means" is a process for identifying an abnormality from the information obtained by the analysis means and determining the content thereof.

[0312] The "notification means" is a function or device for visually or audibly notifying the user based on the abnormality detected by the detection means.

[0313] The "display means" refers to a device such as a display or a speaker for providing a notification to the user, and plays a role of outputting information visually and audibly.

[0314] In order to implement this invention, it is first necessary to construct the entire system. The system combines hardware and software to monitor the subject in real time and detect any abnormalities.

[0315] The "server" is a computer equipped with a high-performance processor for image analysis, and has software frameworks such as Python, TensorFlow, and OpenCV installed. The server receives image and audio data sent from terminals and analyzes them. For images, it performs object detection and pose estimation to identify the actions and state of the subject. For audio data, it applies natural language processing techniques to detect abnormal sounds and cries for help.

[0316] A "terminal" is a device equipped with a camera and microphone, which is installed in the environment where the subject is located. The terminal collects the subject's image and audio data in real time and transmits it to a server wirelessly or via a wired connection.

[0317] The "user" carries a display device such as smart glasses or a smartphone and receives notifications from the server. The notification includes the nature of the anomaly, the time and location of the occurrence, and recommended actions. This allows the user to quickly respond to the scene and take appropriate action.

[0318] For example, when used in a nursing home, if a caregiver is wearing smart glasses and unnatural movements are detected in an elderly person, a visual notification and an audible alert will be displayed on the glasses. This allows the caregiver to intervene immediately and prevent injuries.

[0319] An example of a prompt message to be input to the generating AI model is, "Detect any suspicious behavior from elderly individuals. If there is a risk of falling, send an alert." In this way, the present invention enables efficient monitoring and rapid response.

[0320] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0321] Step 1:

[0322] The device acquires images and audio of the subject using its camera and microphone. This data is stored in the device's temporary memory. The input for this step is real-time acquired images and audio, and the output is digital image and audio data.

[0323] Step 2:

[0324] The terminal transmits the acquired image and audio data to the server wirelessly or via a wired connection. Network communication technology is used for transmission. The input for this step is the image and audio data saved in step 1, and the output is the data transmitted to the server.

[0325] Step 3:

[0326] The server performs image analysis using TensorFlow and OpenCV with the received image data. Object detection and pose estimation algorithms are used to identify the subject's actions and posture. The input for this step is image data sent from the terminal, and the output is the analysis results regarding actions and posture.

[0327] Step 4:

[0328] The server analyzes the received audio data using natural language processing techniques to detect abnormal voice patterns and cries for help. The input for this step is audio data sent from the terminal, and the output is the analysis results regarding abnormal sounds and voice commands.

[0329] Step 5:

[0330] The server detects anomalies based on the analysis results of images and audio. It compares these results to pre-set criteria and determines an anomaly if the threshold is exceeded. The input for this step is the analysis results from steps 3 and 4, and the output is the determination of whether an anomaly occurred.

[0331] Step 6:

[0332] The server generates a notification based on the detected anomaly and sends it to the user's device. The notification includes the nature, time, and location of the anomaly, as well as recommended actions. The input for this step is the anomaly detection result from step 5, and the output is the notification information sent to the user.

[0333] Step 7:

[0334] Users can view and confirm notifications using smart glasses or smartphones and take action. The input for this step is notification information from the server, and the output is the user's response.

[0335] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0336] This invention is a system for enhancing monitoring activities in childcare and elderly care, incorporating an emotion engine that analyzes the emotional state of the subject in addition to effective anomaly detection. This system acquires and analyzes image and audio data to detect anomalies from normal behavior and sounds, and further enables complementary anomaly detection by identifying and evaluating the user's emotions using the emotion engine.

[0337] First, the device uses its camera and microphone to acquire images and audio data of the subject in real time. This data is immediately transmitted to the server. High-speed and stable communication methods are used for data transfer.

[0338] Next, the server analyzes the subject's movements from the received image data and detects abnormal patterns or cries for help from the audio data. Based on these analyses, it determines whether the subject is in a dangerous situation or exhibiting abnormal behavior.

[0339] Furthermore, the server uses an emotion engine to analyze the emotional state from the acquired images and audio. For example, through facial recognition and voice tone analysis, it identifies whether the subject is experiencing anxiety, anger, sadness, or other similar emotions. This enables anomaly detection that also addresses the subject's psychological state.

[0340] If an anomaly or a specific emotional state is detected, the server generates a notification based on a series of pieces of information. This notification includes an alert with adjusted urgency, including a risk assessment based on the subject's emotional state. The notification is sent to the user's smartphone or dedicated device, allowing the user to take immediate action.

[0341] For example, if an elderly person requiring special attention falls, the device captures the event, and the server detects the fall based on motion analysis. Simultaneously, an emotion engine senses anxiety from the elderly person's voice and facial expressions. The server then combines this information to generate a high-urgency alert, which is sent to the user to support a rapid response.

[0342] The system of this invention implements a continuous learning process using recorded analysis results to improve analysis accuracy. Furthermore, all history, including emotional data, is managed and utilized for future activities and analyses. This makes it possible to achieve both the safety of the subjects and a reduction in user burden.

[0343] The following describes the processing flow.

[0344] Step 1:

[0345] The device acquires continuous image data of the subject using its camera and collects audio data in real time using its microphone. This allows the subject's situation to be recorded in real time.

[0346] Step 2:

[0347] The terminal compresses the acquired data and sends it to the server using a high-speed communication protocol. This minimizes delays caused by data transfer, enabling real-time analysis.

[0348] Step 3:

[0349] The server uses the received image data to execute a motion analysis algorithm. It classifies the subject's posture and movements and identifies abnormal movements. For example, it can identify movements such as falls and slipping away.

[0350] Step 4:

[0351] The server analyzes the audio data and uses acoustic features and natural language processing to detect unusual sounds and specific calls. For example, loud calls and long silences may be identified as unusual.

[0352] Step 5:

[0353] The server activates an emotion engine, analyzing image and audio data to assess emotional states. Based on facial expression analysis and voice tone, it identifies emotions such as anxiety, anger, and sadness.

[0354] Step 6:

[0355] The server comprehensively evaluates the results of motion analysis, voice analysis, and emotional state to confirm the presence of an anomaly. If an anomaly is detected, the urgency level is set according to the emotional state.

[0356] Step 7:

[0357] The server generates a notification message based on the configured urgency level. The message includes details of the anomaly, the time and location of the occurrence, and recommended actions. The notification is sent to the user immediately.

[0358] Step 8:

[0359] Users can check received notifications on their smartphones or dedicated devices and decide on the necessary actions. For example, they can immediately head to the scene or check the video feed.

[0360] Step 9:

[0361] The server records and manages all analysis and detection results over the long term. This data is used to improve the accuracy of the analysis and to periodically update the machine learning algorithms.

[0362] (Example 2)

[0363] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0364] In monitoring activities for childcare and elderly care, there is a need to not only detect abnormal behavior in the person being monitored, but also to analyze their emotional state, enabling more comprehensive and rapid detection and response to abnormalities. However, conventional systems have the problem of being unable to adequately ensure the safety of the person being monitored because detailed analysis, including emotional state, is difficult.

[0365] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0366] In this invention, the server includes an image acquisition means, an audio acquisition means, and an emotion analysis means for analyzing the emotional state using the acquired image and audio data. This enables analysis of both the subject's behavior and emotional state, allowing for more accurate and rapid anomaly detection and notification.

[0367] "Image acquisition means" refers to a device or method used to acquire images of a subject, such as a camera, which collects data in real time.

[0368] "Voice acquisition means" refers to a device or method used to acquire the voice of a subject, and involves collecting voice data in real time using a microphone or the like.

[0369] "Analysis means" refers to the process and techniques for analyzing acquired data and identifying abnormalities from the subject's behavior and voice.

[0370] "Emotional analysis means" refers to the process and technology for identifying and analyzing a subject's emotional state using acquired image and audio data.

[0371] "Detection means" refers to the processes and technologies used to discover and judge abnormal phenomena or specific emotional states based on analyzed data.

[0372] "Notification means" refers to a mechanism for generating appropriate notifications based on detected anomalies or emotional states and sending relevant information to the user.

[0373] "Recording means" refers to a system for saving analyzed data and managing it as a history.

[0374] "Learning methods" refer to processes and techniques that utilize recorded data to improve the accuracy of system analysis.

[0375] This invention is a system for enhancing monitoring activities in childcare and elderly care. It analyzes image and audio data acquired in real time to determine the behavior and emotional state of the subject, thereby detecting and notifying of abnormalities.

[0376] First, the device uses its camera and microphone to acquire images and audio data of the subject. The camera is high-resolution, capturing the subject's face and movements in detail. The microphone is highly sensitive, capturing the subject's voice clearly. The acquired data is immediately transmitted to the server via a network with high communication speed and stability.

[0377] The server analyzes the subject's movements by executing a deep learning-based image recognition algorithm on the received image data. In addition, it uses speech recognition software to detect abnormal patterns and cries for help in the audio data. Furthermore, an emotion engine is incorporated, which analyzes the subject's emotional state based on the acquired data through facial recognition technology and speech analysis. This analysis allows for the identification of emotions such as anxiety, anger, and sadness, enabling anomaly detection that also takes psychological circumstances into account.

[0378] If an anomaly or a specific emotional state is detected, the server generates a notification based on the detected information and sends an alert to the user's smartphone or dedicated terminal. This notification includes a risk assessment based on the person's condition, allowing the user to understand the urgency of the situation and take immediate action.

[0379] For example, if an elderly person with dementia falls indoors, the device captures changes in their movements, and the server recognizes the fall from the image data. Simultaneously, it analyzes voice data for signs of pain, and the emotion engine detects feelings of anxiety. From this series of pieces of information, the server immediately generates a high-urgency alert and notifies the user.

[0380] This system features a learning function that continuously improves analysis accuracy using recorded analysis data. By using a generative AI model to execute prompts such as, "Analyze the subject's current condition in detail and report immediately if there are any abnormalities," it enables even more precise monitoring and analysis.

[0381] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0382] Step 1:

[0383] The device uses a camera to acquire images of the subject in real time and a microphone to collect audio data. Physical images and audio are provided as input, which are converted into digital data. The output is the digitized image and audio data. This data is then sent to a server for further analysis.

[0384] Step 2:

[0385] The terminal transmits the acquired image and audio data to the server via a high-bandwidth network. The input is the image and audio data acquired by the terminal, which is then packetized and transmitted. The output is the packetized data received on the server side.

[0386] Step 3:

[0387] The server applies a deep learning-based image recognition algorithm based on the received image data. The input is digitized image data, which is analyzed to identify the subject's movements. The output is data containing the results of the movement analysis.

[0388] Step 4:

[0389] The server analyzes audio data using a speech recognition algorithm. The input is digital audio data, which is analyzed to identify unusual patterns and specific speech content. The output is data containing the results of the speech analysis.

[0390] Step 5:

[0391] The server uses an emotion engine to analyze emotional states from image and audio data. The input is the analysis results of the images and audio, which are used to evaluate the emotional state. The output is data containing the results of the emotion analysis.

[0392] Step 6:

[0393] The server integrates motion analysis results, voice analysis results, and emotion analysis results to detect abnormal conditions. The input is a dataset containing the results of each analysis, which is then analyzed to perform a risk assessment. The output is alert information containing the results of the anomaly detection.

[0394] Step 7:

[0395] The server generates notifications based on detected anomalies and emotional states and sends them to the user's smartphone or dedicated terminal. The input is notification generation data based on anomaly detection and emotional evaluation, and the output is the notification information received by the user.

[0396] (Application Example 2)

[0397] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0398] In modern society, there is a need for means to ensure personal safety. In particular, when going out or acting alone, it is important to be able to understand the surrounding environment in real time and quickly detect the occurrence of abnormalities or psychological changes. However, with current technologies, anomaly detection and emotion analysis are performed separately, and there is insufficient means to respond immediately based on integrated information.

[0399] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0400] In this invention, the server includes an image acquisition means, an analysis means for analyzing the acquired images to identify the subject's behavior, a detection means for detecting anomalies based on the analysis results, an evaluation means that works in conjunction with the analysis means to analyze the emotional state, an adjustment means for performing a risk assessment based on the emotional state and adjusting the notification content, and a monitoring means for monitoring the surrounding situation. This ensures the safety of individuals and enables a quick and appropriate response.

[0401] "Image acquisition means" refers to a means of acquiring image data of a target in real time using an imaging device such as a camera.

[0402] "Analysis means" refers to methods for analyzing acquired image and audio data to identify the subject's behavior and emotional state.

[0403] A "detection means" is a means of detecting anomalies based on analyzed data and determining whether or not there is a problem.

[0404] A "notification system" is a means of generating and sending appropriate alerts based on detected anomalies or risks.

[0405] "Evaluation methods" refer to means of analyzing the emotional state of a subject from acquired data and identifying their psychological state.

[0406] "Adjustment measures" refer to methods for evaluating risk based on the results of an analysis of emotional states and adjusting notification content according to urgency.

[0407] "Monitoring means" refers to methods for constantly monitoring the user's surroundings and detecting abnormalities or dangers.

[0408] "Voice acquisition means" refers to a means of acquiring target voice data using a microphone.

[0409] "Display means" refers to means of visually showing users warnings or notifications regarding detected anomalies.

[0410] A "recording means" is a means of saving the analyzed data as a history for later review or learning.

[0411] A "learning method" is a means of continuously optimizing a system to improve the accuracy of analysis based on recorded data.

[0412] The system for implementing this invention includes smart glasses worn by the user and a data analysis device on a server. The smart glasses are equipped with a camera and microphone, which acquire images and audio data of the surroundings in real time. This data is immediately transmitted to the server via a stable communication means.

[0413] The server uses image analysis software and audio analysis software to analyze the received image and audio data. This allows the server to analyze the subject's behavior and surrounding environment and detect anomalies. Furthermore, to analyze emotional states, it uses a dedicated emotion engine to perform facial recognition and voice tone analysis, thereby evaluating the subject's psychological state.

[0414] If an anomaly is detected, or if sentiment analysis determines that there is a high psychological risk, the server will immediately generate a warning using the notification system and display it on the user's smart glasses display. This notification not only encourages the user to take safe actions but also helps them to take prompt action if necessary.

[0415] As a concrete example, imagine a situation where a user is walking alone at night and suddenly hears loud footsteps around them. In such a case, the smart glasses immediately send data to the server, detect the anomaly, and then visually display a warning such as "Watch your back."

[0416] An example of a prompt for the generating AI model might be, "Explain how to analyze anomalies and emotions in real time from audio and image data to ensure safety." This system provides an effective method for improving user safety.

[0417] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0418] Step 1:

[0419] When a user puts on smart glasses, the device activates its camera and microphone to capture images and sounds from the surroundings in real time. Based on the input, it outputs streams of image and audio data. The device periodically sends this data to a server.

[0420] Step 2:

[0421] The server inputs the received image data into image analysis software to identify the subject's actions and surrounding environment. This allows for the analysis of behavioral patterns such as walking and the approach of others. The output is the analyzed behavioral information.

[0422] Step 3:

[0423] Simultaneously, the server processes the audio data using audio analysis software to detect abnormal sounds and volume changes. For example, this could include the sound of footsteps suddenly approaching. The input is audio data, and the output is the detected abnormal sound information.

[0424] Step 4:

[0425] The server integrates the results of image and audio analysis and uses an emotion engine to evaluate the subject's emotional state. It identifies anxiety and surprise through facial recognition and voice tone analysis. The output is information evaluating the emotional state.

[0426] Step 5:

[0427] The server performs a risk assessment based on all analysis results and generates necessary alerts. For example, it might generate a notification such as "Watch your back." The input is the analysis results, and the output is the alert information.

[0428] Step 6:

[0429] The generated notification is displayed instantly on the user's smart glasses using the device's interface. The user can then see it and take action to pay attention to their surroundings. The output is a visual warning.

[0430] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0431] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0432] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0433] [Third Embodiment]

[0434] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0435] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0436] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0437] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0438] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0439] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0440] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0441] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0442] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0443] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0444] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0445] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0446] This invention provides a monitoring system to reduce the burden of childcare and elder care, and is configured as follows. This system integrates multiple means for acquiring, analyzing, detecting anomalies, and providing notifications for images.

[0447] First, the device continuously collects data using its camera and microphone at the location where the person being monitored is located. This records real-time images and audio of the person being monitored. The device then transmits this data to a server wirelessly or via a wired connection.

[0448] The server performs advanced image analysis based on the received image data. Specifically, it uses object detection and posture estimation algorithms to determine the actions and state of the subject. For example, it can recognize a child getting out of bed or an elderly person falling. For audio data, natural language processing technology is used to detect abnormal sounds and cries for help.

[0449] Next, the server determines whether an abnormal situation has been detected based on the analyzed information. It compares the results against pre-set criteria, and if it detects any actions or sounds that exceed the criteria, it treats them as abnormal. Based on this determination, it determines the designated emergency level and issues instructions for action.

[0450] When an anomaly is detected, the server immediately generates a notification. This notification includes the time, location, and nature of the anomaly, as well as recommended actions. This information is sent to the user via smartphone or a dedicated device. Receiving the notification allows the user to quickly respond to the situation or take appropriate action.

[0451] The system also records all analysis results and detection events on the server. This data is used to understand long-term activity history and as training data to improve analysis accuracy. The server is designed to continuously update its analysis algorithms using the recorded data, enabling more precise anomaly detection.

[0452] For example, if a child rolls over in their sleep at night, the device captures the movement, and if the server determines it to be a safe movement, no notification is sent. However, if the child makes a movement that suggests they might fall out of bed, a notification is immediately generated, an alert is sent to the user, and appropriate action is prompted.

[0453] The following describes the processing flow.

[0454] Step 1:

[0455] The device acquires image data at a rate of multiple frames per second using a camera in the room where the person being monitored is located, and simultaneously collects audio data using a microphone. This allows for the capture of real-time visual and audio information.

[0456] Step 2:

[0457] The terminal compresses the acquired image and audio data and sends it to the server with minimal latency. This is achieved by utilizing a streaming protocol to enable efficient data transfer.

[0458] Step 3:

[0459] The server analyzes the received images using a high-speed processing algorithm to analyze the subject's movements and posture. It utilizes object detection technology to identify actions related to the safety of children and the elderly (e.g., falls, getting out of bed).

[0460] Step 4:

[0461] The server analyzes the audio data and performs natural language processing and acoustic analysis. It detects when the subject makes unusual sounds (e.g., screams, cries for help) or when there are sudden changes in the sound of the environment.

[0462] Step 5:

[0463] Based on the analysis results, the server detects anomalies by referring to pre-configured anomaly criteria. It immediately evaluates whether the criteria are met, and if an anomaly is detected, it generates response instructions according to the urgency.

[0464] Step 6:

[0465] Based on the detection of an anomaly, the server creates a relevant notification message. This message includes the type and location of the anomaly, as well as recommended actions, and is prepared for sending to the user.

[0466] Step 7:

[0467] The server sends the generated notification to the user. The user receives the notification on their smartphone or dedicated device and can view it on the screen. This allows the user to take immediate action.

[0468] Step 8:

[0469] The server meticulously records all detected events and analysis results. This recorded data will be used for future analysis and learning processes.

[0470] Step 9:

[0471] The server uses the recorded data to update its machine learning algorithms. This improves the accuracy of the analysis and enables more appropriate anomaly detection.

[0472] (Example 1)

[0473] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0474] In childcare and eldercare settings, effectively monitoring the safety of those involved is crucial, but conventional methods have struggled to achieve sufficient accuracy and speed in detecting anomalies. Furthermore, efficient data processing and learning are required for long-term data management and improved analysis accuracy. Additionally, establishing criteria for identifying behavioral patterns and performing accurate anomaly detection has been a challenge.

[0475] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0476] In this invention, the server includes an image acquisition means, an audio data collection means for acquiring sound, and an evaluation means for detecting anomalies from the processed data. This enables rapid and accurate anomaly detection using images and sound.

[0477] "Image acquisition means" refers to a device or method for collecting image data for monitoring a subject.

[0478] "Information processing means" refers to a device or method for analyzing raw data acquired from cameras and microphones to identify abnormalities from the behavior and voice of a subject.

[0479] "Sound data collection means" refers to a device or method for acquiring and analyzing sound.

[0480] "Evaluation means" refers to a device or method for identifying behaviors or sounds that exceed a certain standard from analyzed data and for determining abnormalities.

[0481] "Information provision means" refers to a device or system for notifying the user of information regarding detected anomalies.

[0482] "Management means" refers to a device or method for recording analyzed data and maintaining it as a long-term history.

[0483] "Learning methods" refer to algorithms and methods that utilize recorded data to improve the accuracy of analysis.

[0484] "Behavioral identification means" refers to a device or method for analyzing behavioral patterns based on the subject's state.

[0485] "Criteria setting means" refers to a device or method for setting criteria for detecting abnormalities based on information obtained by the behavior identification means.

[0486] This invention is a monitoring system designed to alleviate the burden of childcare and elder care. This system integrates multiple functions to collect, analyze, and notify images and audio in real time for monitoring the target individual. The following describes its specific embodiments.

[0487] The device is installed where the subject is located and continuously collects data using a camera and microphone. A standard digital camera capable of capturing high-resolution images is used as the camera, and a microphone with noise-canceling capabilities is recommended. The collected data is transmitted to a server via wireless or wired communication.

[0488] The server uses a common image analysis library, which is a machine learning model, to analyze the received image data. For example, a framework like TensorFlow is used for object detection and pose estimation. This allows the server to identify specific actions such as a person standing up or sitting down. For audio data, speech recognition technology such as Google Cloud Speech-to-Text is used to detect abnormal sounds and urgent voices through natural language processing.

[0489] If an anomaly is detected, the server generates a notification and sends it to the user via smartphone or a dedicated device. This notification includes the time and location of the anomaly, detailed information, and recommended actions. This allows the user to quickly understand the situation and take appropriate measures.

[0490] Furthermore, the server records all analysis results and detection events. This recorded data is used to update the server's learning algorithms and improve analysis accuracy. Specifically, by analyzing behavioral patterns using past data, the accuracy of detecting typical anomalies can be improved.

[0491] For example, if a child rolls over in their sleep at night, the device captures this movement, and if the server determines that the action is normal, no notification is sent. However, if the movement is dangerous, an alert is immediately generated and sent to the user.

[0492] Examples of prompt statements include:

[0493] "Please explain a real-time anomaly detection system for monitoring the safety of children and the elderly."

[0494] This is one possible explanation.

[0495] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0496] Step 1:

[0497] The device continuously collects data using a camera and microphone. Input consists of image and audio data containing the subject's activities. The device captures this data in real time and transmits it to a server wirelessly or via a wired connection. Specifically, it photographs and records moving children or elderly people speaking. Output consists of the transmitted high-resolution image and audio files.

[0498] Step 2:

[0499] The server receives image data transmitted from the terminal. Using the received image as input, it performs object detection and pose estimation using information processing tools. Using an image analysis library such as TensorFlow, it performs action identification to determine whether the subject is standing or sitting. Specifically, this action involves detecting a particular pose or movement of the human body within the image. The output is the type of action and action information.

[0500] Step 3:

[0501] The server also receives audio data. Using the received audio file as input, it converts it to text using speech recognition technology and detects anomalies and important phrases using natural language processing. Specifically, it detects abnormal sounds such as "help" or "I fell." The output is the detected key phrases and types of sounds.

[0502] Step 4:

[0503] The server analyzes information obtained from images and audio using evaluation tools to determine anomalies. Inputs are behavioral and audio information. The analyzed data is compared with established criteria to identify behaviors or sounds deemed abnormal. The specific action involves checking whether the behavior or sound anomaly exceeds the specified threshold. The output is the anomaly detection result and its detailed information.

[0504] Step 5:

[0505] The server generates notifications based on detected anomalies. The input is the result of the anomaly detection. The server uses this information to create a notification to send an alert to the user. Specifically, it specifies the details of the anomaly, the time it occurred, and recommended actions. The output is the notification message sent to the user.

[0506] Step 6:

[0507] Users receive notifications on their smartphones or dedicated devices. The input is the notification sent from the server. Users check the notification and respond quickly, such as checking camera footage if necessary. Specifically, actions include checking the notification and rushing to the scene if required. The output is the user's appropriate response.

[0508] Step 7:

[0509] The server records all analysis results and detection events in a database. Inputs are behavioral identification information and anomaly detection results. This data is managed as history and used as training data to improve future analysis accuracy. Specifically, it continuously incorporates the subject's daily behavioral patterns into the analysis. The output is historical activity history data.

[0510] (Application Example 1)

[0511] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0512] In childcare and elder care, ensuring the safety of those being monitored requires the rapid and accurate detection of abnormal behavior or conditions, and immediate notification. Conventional systems have limitations in the accuracy of anomaly detection and the speed of notification, which can lead to delays in on-site responses, thus necessitating further improvements.

[0513] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0514] In this invention, the server includes an image acquisition means, an analysis means for analyzing the acquired images to identify the subject's actions, a detection means for detecting anomalies based on the analysis results, and a display means for outputting notifications visually and audibly. This enables improved accuracy in anomaly detection and rapid notification.

[0515] "Image acquisition means" refers to a device or function for acquiring images of a person being monitored using a visual sensor such as a camera.

[0516] "Analysis means" refers to algorithms and technologies used to analyze acquired image and audio data and identify the subject's behavior or abnormalities.

[0517] "Detection means" refers to the process of identifying anomalies from information obtained by analysis means and determining their nature.

[0518] "Notification means" refers to a function or device for providing visual or audible notifications to the user based on anomalies detected by detection means.

[0519] "Display means" refers to devices such as displays and speakers that provide notifications to the user, and are responsible for outputting information visually and audibly.

[0520] In order to implement this invention, it is first necessary to construct the entire system. The system combines hardware and software to monitor the subject in real time and detect any abnormalities.

[0521] The "server" is a computer equipped with a high-performance processor for image analysis, and has software frameworks such as Python, TensorFlow, and OpenCV installed. The server receives image and audio data sent from terminals and analyzes them. For images, it performs object detection and pose estimation to identify the actions and state of the subject. For audio data, it applies natural language processing techniques to detect abnormal sounds and cries for help.

[0522] A "terminal" is a device equipped with a camera and microphone, which is installed in the environment where the subject is located. The terminal collects the subject's image and audio data in real time and transmits it to a server wirelessly or via a wired connection.

[0523] The "user" carries a display device such as smart glasses or a smartphone and receives notifications from the server. The notification includes the nature of the anomaly, the time and location of the occurrence, and recommended actions. This allows the user to quickly respond to the scene and take appropriate action.

[0524] For example, when used in a nursing home, if a caregiver is wearing smart glasses and unnatural movements are detected in an elderly person, a visual notification and an audible alert will be displayed on the glasses. This allows the caregiver to intervene immediately and prevent injuries.

[0525] An example of a prompt message to be input to the generating AI model is, "Detect any suspicious behavior from elderly individuals. If there is a risk of falling, send an alert." In this way, the present invention enables efficient monitoring and rapid response.

[0526] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0527] Step 1:

[0528] The device acquires images and audio of the subject using its camera and microphone. This data is stored in the device's temporary memory. The input for this step is real-time acquired images and audio, and the output is digital image and audio data.

[0529] Step 2:

[0530] The terminal transmits the acquired image and audio data to the server wirelessly or via a wired connection. Network communication technology is used for transmission. The input for this step is the image and audio data saved in step 1, and the output is the data transmitted to the server.

[0531] Step 3:

[0532] The server performs image analysis using TensorFlow and OpenCV with the received image data. Object detection and pose estimation algorithms are used to identify the subject's actions and posture. The input for this step is image data sent from the terminal, and the output is the analysis results regarding actions and posture.

[0533] Step 4:

[0534] The server analyzes the received audio data using natural language processing techniques to detect abnormal voice patterns and cries for help. The input for this step is audio data sent from the terminal, and the output is the analysis results regarding abnormal sounds and voice commands.

[0535] Step 5:

[0536] The server detects anomalies based on the analysis results of images and audio. It compares these results to pre-set criteria and determines an anomaly if the threshold is exceeded. The input for this step is the analysis results from steps 3 and 4, and the output is the determination of whether an anomaly occurred.

[0537] Step 6:

[0538] The server generates a notification based on the detected anomaly and sends it to the user's device. The notification includes the nature, time, and location of the anomaly, as well as recommended actions. The input for this step is the anomaly detection result from step 5, and the output is the notification information sent to the user.

[0539] Step 7:

[0540] Users can view and confirm notifications using smart glasses or smartphones and take action. The input for this step is notification information from the server, and the output is the user's response.

[0541] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0542] This invention is a system for enhancing monitoring activities in childcare and elderly care, incorporating an emotion engine that analyzes the emotional state of the subject in addition to effective anomaly detection. This system acquires and analyzes image and audio data to detect anomalies from normal behavior and sounds, and further enables complementary anomaly detection by identifying and evaluating the user's emotions using the emotion engine.

[0543] First, the device uses its camera and microphone to acquire images and audio data of the subject in real time. This data is immediately transmitted to the server. High-speed and stable communication methods are used for data transfer.

[0544] Next, the server analyzes the subject's movements from the received image data and detects abnormal patterns or cries for help from the audio data. Based on these analyses, it determines whether the subject is in a dangerous situation or exhibiting abnormal behavior.

[0545] Furthermore, the server uses an emotion engine to analyze the emotional state from the acquired images and audio. For example, through facial recognition and voice tone analysis, it identifies whether the subject is experiencing anxiety, anger, sadness, or other similar emotions. This enables anomaly detection that also addresses the subject's psychological state.

[0546] If an anomaly or a specific emotional state is detected, the server generates a notification based on a series of pieces of information. This notification includes an alert with adjusted urgency, including a risk assessment based on the subject's emotional state. The notification is sent to the user's smartphone or dedicated device, allowing the user to take immediate action.

[0547] For example, if an elderly person requiring special attention falls, the device captures the event, and the server detects the fall based on motion analysis. Simultaneously, an emotion engine senses anxiety from the elderly person's voice and facial expressions. The server then combines this information to generate a high-urgency alert, which is sent to the user to support a rapid response.

[0548] The system of this invention implements a continuous learning process using recorded analysis results to improve analysis accuracy. Furthermore, all history, including emotional data, is managed and utilized for future activities and analyses. This makes it possible to achieve both the safety of the subjects and a reduction in user burden.

[0549] The following describes the processing flow.

[0550] Step 1:

[0551] The device acquires continuous image data of the subject using its camera and collects audio data in real time using its microphone. This allows the subject's situation to be recorded in real time.

[0552] Step 2:

[0553] The terminal compresses the acquired data and sends it to the server using a high-speed communication protocol. This minimizes delays caused by data transfer, enabling real-time analysis.

[0554] Step 3:

[0555] The server uses the received image data to execute a motion analysis algorithm. It classifies the subject's posture and movements and identifies abnormal movements. For example, it can identify movements such as falls and slipping away.

[0556] Step 4:

[0557] The server analyzes the audio data and uses acoustic features and natural language processing to detect unusual sounds and specific calls. For example, loud calls and long silences may be identified as unusual.

[0558] Step 5:

[0559] The server activates an emotion engine, analyzing image and audio data to assess emotional states. Based on facial expression analysis and voice tone, it identifies emotions such as anxiety, anger, and sadness.

[0560] Step 6:

[0561] The server comprehensively evaluates the results of motion analysis, voice analysis, and emotional state to confirm the presence of an anomaly. If an anomaly is detected, the urgency level is set according to the emotional state.

[0562] Step 7:

[0563] The server generates a notification message based on the configured urgency level. The message includes details of the anomaly, the time and location of the occurrence, and recommended actions. The notification is sent to the user immediately.

[0564] Step 8:

[0565] Users can check received notifications on their smartphones or dedicated devices and decide on the necessary actions. For example, they can immediately head to the scene or check the video feed.

[0566] Step 9:

[0567] The server records and manages all analysis and detection results over the long term. This data is used to improve the accuracy of the analysis and to periodically update the machine learning algorithms.

[0568] (Example 2)

[0569] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0570] In monitoring activities for childcare and elderly care, there is a need to not only detect abnormal behavior in the person being monitored, but also to analyze their emotional state, enabling more comprehensive and rapid detection and response to abnormalities. However, conventional systems have the problem of being unable to adequately ensure the safety of the person being monitored because detailed analysis, including emotional state, is difficult.

[0571] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0572] In this invention, the server includes an image acquisition means, an audio acquisition means, and an emotion analysis means for analyzing the emotional state using the acquired image and audio data. This enables analysis of both the subject's behavior and emotional state, allowing for more accurate and rapid anomaly detection and notification.

[0573] "Image acquisition means" refers to a device or method used to acquire images of a subject, such as a camera, which collects data in real time.

[0574] "Voice acquisition means" refers to a device or method used to acquire the voice of a subject, and involves collecting voice data in real time using a microphone or the like.

[0575] "Analysis means" refers to the process and techniques for analyzing acquired data and identifying abnormalities from the subject's behavior and voice.

[0576] "Emotional analysis means" refers to the process and technology for identifying and analyzing a subject's emotional state using acquired image and audio data.

[0577] "Detection means" refers to the processes and technologies used to discover and judge abnormal phenomena or specific emotional states based on analyzed data.

[0578] "Notification means" refers to a mechanism for generating appropriate notifications based on detected anomalies or emotional states and sending relevant information to the user.

[0579] "Recording means" refers to a system for saving analyzed data and managing it as a history.

[0580] "Learning methods" refer to processes and techniques that utilize recorded data to improve the accuracy of system analysis.

[0581] This invention is a system for enhancing monitoring activities in childcare and elderly care. It analyzes image and audio data acquired in real time to determine the behavior and emotional state of the subject, thereby detecting and notifying of abnormalities.

[0582] First, the device uses its camera and microphone to acquire images and audio data of the subject. The camera is high-resolution, capturing the subject's face and movements in detail. The microphone is highly sensitive, capturing the subject's voice clearly. The acquired data is immediately transmitted to the server via a network with high communication speed and stability.

[0583] The server analyzes the subject's movements by executing a deep learning-based image recognition algorithm on the received image data. In addition, it uses speech recognition software to detect abnormal patterns and cries for help in the audio data. Furthermore, an emotion engine is incorporated, which analyzes the subject's emotional state based on the acquired data through facial recognition technology and speech analysis. This analysis allows for the identification of emotions such as anxiety, anger, and sadness, enabling anomaly detection that also takes psychological circumstances into account.

[0584] If an anomaly or a specific emotional state is detected, the server generates a notification based on the detected information and sends an alert to the user's smartphone or dedicated terminal. This notification includes a risk assessment based on the person's condition, allowing the user to understand the urgency of the situation and take immediate action.

[0585] For example, if an elderly person with dementia falls indoors, the device captures changes in their movements, and the server recognizes the fall from the image data. Simultaneously, it analyzes voice data for signs of pain, and the emotion engine detects feelings of anxiety. From this series of pieces of information, the server immediately generates a high-urgency alert and notifies the user.

[0586] This system features a learning function that continuously improves analysis accuracy using recorded analysis data. By using a generative AI model to execute prompts such as, "Analyze the subject's current condition in detail and report immediately if there are any abnormalities," it enables even more precise monitoring and analysis.

[0587] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0588] Step 1:

[0589] The device uses a camera to acquire images of the subject in real time and a microphone to collect audio data. Physical images and audio are provided as input, which are converted into digital data. The output is the digitized image and audio data. This data is then sent to a server for further analysis.

[0590] Step 2:

[0591] The terminal transmits the acquired image and audio data to the server via a high-bandwidth network. The input is the image and audio data acquired by the terminal, which is then packetized and transmitted. The output is the packetized data received on the server side.

[0592] Step 3:

[0593] The server applies a deep learning-based image recognition algorithm based on the received image data. The input is digitized image data, which is analyzed to identify the subject's movements. The output is data containing the results of the movement analysis.

[0594] Step 4:

[0595] The server analyzes audio data using a speech recognition algorithm. The input is digital audio data, which is analyzed to identify unusual patterns and specific speech content. The output is data containing the results of the speech analysis.

[0596] Step 5:

[0597] The server uses an emotion engine to analyze emotional states from image and audio data. The input is the analysis results of the images and audio, which are used to evaluate the emotional state. The output is data containing the results of the emotion analysis.

[0598] Step 6:

[0599] The server integrates motion analysis results, voice analysis results, and emotion analysis results to detect abnormal conditions. The input is a dataset containing the results of each analysis, which is then analyzed to perform a risk assessment. The output is alert information containing the results of the anomaly detection.

[0600] Step 7:

[0601] The server generates notifications based on detected anomalies and emotional states and sends them to the user's smartphone or dedicated terminal. The input is notification generation data based on anomaly detection and emotional evaluation, and the output is the notification information received by the user.

[0602] (Application Example 2)

[0603] Next, we will explain Application Example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0604] In modern society, there is a need for means to ensure personal safety. In particular, when going out or acting alone, it is important to be able to understand the surrounding environment in real time and quickly detect the occurrence of abnormalities or psychological changes. However, with current technologies, anomaly detection and emotion analysis are performed separately, and there is insufficient means to respond immediately based on integrated information.

[0605] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0606] In this invention, the server includes an image acquisition means, an analysis means for analyzing the acquired images to identify the subject's behavior, a detection means for detecting anomalies based on the analysis results, an evaluation means that works in conjunction with the analysis means to analyze the emotional state, an adjustment means for performing a risk assessment based on the emotional state and adjusting the notification content, and a monitoring means for monitoring the surrounding situation. This ensures the safety of individuals and enables a quick and appropriate response.

[0607] "Image acquisition means" refers to a means of acquiring image data of a target in real time using an imaging device such as a camera.

[0608] "Analysis means" refers to methods for analyzing acquired image and audio data to identify the subject's behavior and emotional state.

[0609] A "detection means" is a means of detecting anomalies based on analyzed data and determining whether or not there is a problem.

[0610] A "notification system" is a means of generating and sending appropriate alerts based on detected anomalies or risks.

[0611] "Evaluation methods" refer to means of analyzing the emotional state of a subject from acquired data and identifying their psychological state.

[0612] "Adjustment measures" refer to methods for evaluating risk based on the results of an analysis of emotional states and adjusting notification content according to urgency.

[0613] "Monitoring means" refers to methods for constantly monitoring the user's surroundings and detecting abnormalities or dangers.

[0614] "Voice acquisition means" refers to a means of acquiring target voice data using a microphone.

[0615] "Display means" refers to means of visually showing users warnings or notifications regarding detected anomalies.

[0616] A "recording means" is a means of saving the analyzed data as a history for later review or learning.

[0617] A "learning method" is a means of continuously optimizing a system to improve the accuracy of analysis based on recorded data.

[0618] The system for implementing this invention includes smart glasses worn by the user and a data analysis device on a server. The smart glasses are equipped with a camera and microphone, which acquire images and audio data of the surroundings in real time. This data is immediately transmitted to the server via a stable communication means.

[0619] The server uses image analysis software and audio analysis software to analyze the received image and audio data. This allows the server to analyze the subject's behavior and surrounding environment and detect anomalies. Furthermore, to analyze emotional states, it uses a dedicated emotion engine to perform facial recognition and voice tone analysis, thereby evaluating the subject's psychological state.

[0620] If an anomaly is detected, or if sentiment analysis determines that there is a high psychological risk, the server will immediately generate a warning using the notification system and display it on the user's smart glasses display. This notification not only encourages the user to take safe actions but also helps them to take prompt action if necessary.

[0621] As a concrete example, imagine a situation where a user is walking alone at night and suddenly hears loud footsteps around them. In such a case, the smart glasses immediately send data to the server, detect the anomaly, and then visually display a warning such as "Watch your back."

[0622] An example of a prompt for the generating AI model might be, "Explain how to analyze anomalies and emotions in real time from audio and image data to ensure safety." This system provides an effective method for improving user safety.

[0623] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0624] Step 1:

[0625] When a user puts on smart glasses, the device activates its camera and microphone to capture images and sounds from the surroundings in real time. Based on the input, it outputs streams of image and audio data. The device periodically sends this data to a server.

[0626] Step 2:

[0627] The server inputs the received image data into image analysis software to identify the subject's actions and surrounding environment. This allows for the analysis of behavioral patterns such as walking and the approach of others. The output is the analyzed behavioral information.

[0628] Step 3:

[0629] Simultaneously, the server processes the audio data using audio analysis software to detect abnormal sounds and volume changes. For example, this could include the sound of footsteps suddenly approaching. The input is audio data, and the output is the detected abnormal sound information.

[0630] Step 4:

[0631] The server integrates the results of image and audio analysis and uses an emotion engine to evaluate the subject's emotional state. It identifies anxiety and surprise through facial recognition and voice tone analysis. The output is information evaluating the emotional state.

[0632] Step 5:

[0633] The server performs a risk assessment based on all analysis results and generates necessary alerts. For example, it might generate a notification such as "Watch your back." The input is the analysis results, and the output is the alert information.

[0634] Step 6:

[0635] The generated notification is displayed instantly on the user's smart glasses using the device's interface. The user can then see it and take action to pay attention to their surroundings. The output is a visual warning.

[0636] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0637] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0638] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0639] [Fourth Embodiment]

[0640] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0641] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0642] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0643] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0644] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0645] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0646] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0647] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0648] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0649] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0650] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0651] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0652] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0653] This invention provides a monitoring system to reduce the burden of childcare and elder care, and is configured as follows. This system integrates multiple means for acquiring, analyzing, detecting anomalies, and providing notifications for images.

[0654] First, the device continuously collects data using its camera and microphone at the location where the person being monitored is located. This records real-time images and audio of the person being monitored. The device then transmits this data to a server wirelessly or via a wired connection.

[0655] The server performs advanced image analysis based on the received image data. Specifically, it uses object detection and posture estimation algorithms to determine the actions and state of the subject. For example, it can recognize a child getting out of bed or an elderly person falling. For audio data, natural language processing technology is used to detect abnormal sounds and cries for help.

[0656] Next, the server determines whether an abnormal situation has been detected based on the analyzed information. It compares the results against pre-set criteria, and if it detects any actions or sounds that exceed the criteria, it treats them as abnormal. Based on this determination, it determines the designated emergency level and issues instructions for action.

[0657] When an anomaly is detected, the server immediately generates a notification. This notification includes the time, location, and nature of the anomaly, as well as recommended actions. This information is sent to the user via smartphone or a dedicated device. Receiving the notification allows the user to quickly respond to the situation or take appropriate action.

[0658] The system also records all analysis results and detection events on the server. This data is used to understand long-term activity history and as training data to improve analysis accuracy. The server is designed to continuously update its analysis algorithms using the recorded data, enabling more precise anomaly detection.

[0659] For example, if a child rolls over in their sleep at night, the device captures the movement, and if the server determines it to be a safe movement, no notification is sent. However, if the child makes a movement that suggests they might fall out of bed, a notification is immediately generated, an alert is sent to the user, and appropriate action is prompted.

[0660] The following describes the processing flow.

[0661] Step 1:

[0662] The device acquires image data at a rate of multiple frames per second using a camera in the room where the person being monitored is located, and simultaneously collects audio data using a microphone. This allows for the capture of real-time visual and audio information.

[0663] Step 2:

[0664] The terminal compresses the acquired image and audio data and sends it to the server with minimal latency. This is achieved by utilizing a streaming protocol to enable efficient data transfer.

[0665] Step 3:

[0666] The server analyzes the received images using a high-speed processing algorithm to analyze the subject's movements and posture. It utilizes object detection technology to identify actions related to the safety of children and the elderly (e.g., falls, getting out of bed).

[0667] Step 4:

[0668] The server analyzes the audio data and performs natural language processing and acoustic analysis. It detects when the subject makes unusual sounds (e.g., screams, cries for help) or when there are sudden changes in the sound of the environment.

[0669] Step 5:

[0670] Based on the analysis results, the server detects anomalies by referring to pre-configured anomaly criteria. It immediately evaluates whether the criteria are met, and if an anomaly is detected, it generates response instructions according to the urgency.

[0671] Step 6:

[0672] Based on the detection of an anomaly, the server creates a relevant notification message. This message includes the type and location of the anomaly, as well as recommended actions, and is prepared for sending to the user.

[0673] Step 7:

[0674] The server sends the generated notification to the user. The user receives the notification on their smartphone or dedicated device and can view it on the screen. This allows the user to take immediate action.

[0675] Step 8:

[0676] The server meticulously records all detected events and analysis results. This recorded data will be used for future analysis and learning processes.

[0677] Step 9:

[0678] The server uses the recorded data to update its machine learning algorithms. This improves the accuracy of the analysis and enables more appropriate anomaly detection.

[0679] (Example 1)

[0680] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0681] In childcare and eldercare settings, effectively monitoring the safety of those involved is crucial, but conventional methods have struggled to achieve sufficient accuracy and speed in detecting anomalies. Furthermore, efficient data processing and learning are required for long-term data management and improved analysis accuracy. Additionally, establishing criteria for identifying behavioral patterns and performing accurate anomaly detection has been a challenge.

[0682] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0683] In this invention, the server includes an image acquisition means, an audio data collection means for acquiring sound, and an evaluation means for detecting anomalies from the processed data. This enables rapid and accurate anomaly detection using images and sound.

[0684] "Image acquisition means" refers to a device or method for collecting image data for monitoring a subject.

[0685] "Information processing means" refers to a device or method for analyzing raw data acquired from cameras and microphones to identify abnormalities from the behavior and voice of a subject.

[0686] "Sound data collection means" refers to a device or method for acquiring and analyzing sound.

[0687] "Evaluation means" refers to a device or method for identifying behaviors or sounds that exceed a certain standard from analyzed data and for determining abnormalities.

[0688] "Information provision means" refers to a device or system for notifying the user of information regarding detected anomalies.

[0689] "Management means" refers to a device or method for recording analyzed data and maintaining it as a long-term history.

[0690] "Learning methods" refer to algorithms and methods that utilize recorded data to improve the accuracy of analysis.

[0691] "Behavioral identification means" refers to a device or method for analyzing behavioral patterns based on the subject's state.

[0692] "Criteria setting means" refers to a device or method for setting criteria for detecting abnormalities based on information obtained by the behavior identification means.

[0693] This invention is a monitoring system designed to alleviate the burden of childcare and elder care. This system integrates multiple functions to collect, analyze, and notify images and audio in real time for monitoring the target individual. The following describes its specific embodiments.

[0694] The device is installed where the subject is located and continuously collects data using a camera and microphone. A standard digital camera capable of capturing high-resolution images is used as the camera, and a microphone with noise-canceling capabilities is recommended. The collected data is transmitted to a server via wireless or wired communication.

[0695] The server uses a common image analysis library, which is a machine learning model, to analyze the received image data. For example, a framework like TensorFlow is used for object detection and pose estimation. This allows the server to identify specific actions such as a person standing up or sitting down. For audio data, speech recognition technology such as Google Cloud Speech-to-Text is used to detect abnormal sounds and urgent voices through natural language processing.

[0696] If an anomaly is detected, the server generates a notification and sends it to the user via smartphone or a dedicated device. This notification includes the time and location of the anomaly, detailed information, and recommended actions. This allows the user to quickly understand the situation and take appropriate measures.

[0697] Furthermore, the server records all analysis results and detection events. This recorded data is used to update the server's learning algorithms and improve analysis accuracy. Specifically, by analyzing behavioral patterns using past data, the accuracy of detecting typical anomalies can be improved.

[0698] For example, if a child rolls over in their sleep at night, the device captures this movement, and if the server determines that the action is normal, no notification is sent. However, if the movement is dangerous, an alert is immediately generated and sent to the user.

[0699] Examples of prompt statements include:

[0700] "Please explain a real-time anomaly detection system for monitoring the safety of children and the elderly."

[0701] This is one possible explanation.

[0702] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0703] Step 1:

[0704] The device continuously collects data using a camera and microphone. Input consists of image and audio data containing the subject's activities. The device captures this data in real time and transmits it to a server wirelessly or via a wired connection. Specifically, it photographs and records moving children or elderly people speaking. Output consists of the transmitted high-resolution image and audio files.

[0705] Step 2:

[0706] The server receives image data transmitted from the terminal. Using the received image as input, it performs object detection and pose estimation using information processing tools. Using an image analysis library such as TensorFlow, it performs action identification to determine whether the subject is standing or sitting. Specifically, this action involves detecting a particular pose or movement of the human body within the image. The output is the type of action and action information.

[0707] Step 3:

[0708] The server also receives audio data. Using the received audio file as input, it converts it to text using speech recognition technology and detects anomalies and important phrases using natural language processing. Specifically, it detects abnormal sounds such as "help" or "I fell." The output is the detected key phrases and types of sounds.

[0709] Step 4:

[0710] The server analyzes information obtained from images and audio using evaluation tools to determine anomalies. Inputs are behavioral and audio information. The analyzed data is compared with established criteria to identify behaviors or sounds deemed abnormal. The specific action involves checking whether the behavior or sound anomaly exceeds the specified threshold. The output is the anomaly detection result and its detailed information.

[0711] Step 5:

[0712] The server generates notifications based on detected anomalies. The input is the result of the anomaly detection. The server uses this information to create a notification to send an alert to the user. Specifically, it specifies the details of the anomaly, the time it occurred, and recommended actions. The output is the notification message sent to the user.

[0713] Step 6:

[0714] Users receive notifications on their smartphones or dedicated devices. The input is the notification sent from the server. Users check the notification and respond quickly, such as checking camera footage if necessary. Specifically, actions include checking the notification and rushing to the scene if required. The output is the user's appropriate response.

[0715] Step 7:

[0716] The server records all analysis results and detection events in a database. Inputs are behavioral identification information and anomaly detection results. This data is managed as history and used as training data to improve future analysis accuracy. Specifically, it continuously incorporates the subject's daily behavioral patterns into the analysis. The output is historical activity history data.

[0717] (Application Example 1)

[0718] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0719] In childcare and elder care, ensuring the safety of those being monitored requires the rapid and accurate detection of abnormal behavior or conditions, and immediate notification. Conventional systems have limitations in the accuracy of anomaly detection and the speed of notification, which can lead to delays in on-site responses, thus necessitating further improvements.

[0720] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0721] In this invention, the server includes an image acquisition means, an analysis means for analyzing the acquired images to identify the subject's actions, a detection means for detecting anomalies based on the analysis results, and a display means for outputting notifications visually and audibly. This enables improved accuracy in anomaly detection and rapid notification.

[0722] "Image acquisition means" refers to a device or function for acquiring images of a person being monitored using a visual sensor such as a camera.

[0723] "Analysis means" refers to algorithms and technologies used to analyze acquired image and audio data and identify the subject's behavior or abnormalities.

[0724] "Detection means" refers to the process of identifying anomalies from information obtained by analysis means and determining their nature.

[0725] "Notification means" refers to a function or device for providing visual or audible notifications to the user based on anomalies detected by detection means.

[0726] "Display means" refers to devices such as displays and speakers that provide notifications to the user, and are responsible for outputting information visually and audibly.

[0727] In order to implement this invention, it is first necessary to construct the entire system. The system combines hardware and software to monitor the subject in real time and detect any abnormalities.

[0728] The "server" is a computer equipped with a high-performance processor for image analysis, and has software frameworks such as Python, TensorFlow, and OpenCV installed. The server receives image and audio data sent from terminals and analyzes them. For images, it performs object detection and pose estimation to identify the actions and state of the subject. For audio data, it applies natural language processing techniques to detect abnormal sounds and cries for help.

[0729] A "terminal" is a device equipped with a camera and microphone, which is installed in the environment where the subject is located. The terminal collects the subject's image and audio data in real time and transmits it to a server wirelessly or via a wired connection.

[0730] The "user" carries a display device such as smart glasses or a smartphone and receives notifications from the server. The notification includes the nature of the anomaly, the time and location of the occurrence, and recommended actions. This allows the user to quickly respond to the scene and take appropriate action.

[0731] For example, when used in a nursing home, if a caregiver is wearing smart glasses and unnatural movements are detected in an elderly person, a visual notification and an audible alert will be displayed on the glasses. This allows the caregiver to intervene immediately and prevent injuries.

[0732] An example of a prompt message to be input to the generating AI model is, "Detect any suspicious behavior from elderly individuals. If there is a risk of falling, send an alert." In this way, the present invention enables efficient monitoring and rapid response.

[0733] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0734] Step 1:

[0735] The device acquires images and audio of the subject using its camera and microphone. This data is stored in the device's temporary memory. The input for this step is real-time acquired images and audio, and the output is digital image and audio data.

[0736] Step 2:

[0737] The terminal transmits the acquired image and audio data to the server wirelessly or via a wired connection. Network communication technology is used for transmission. The input for this step is the image and audio data saved in step 1, and the output is the data transmitted to the server.

[0738] Step 3:

[0739] The server performs image analysis using TensorFlow and OpenCV with the received image data. Object detection and pose estimation algorithms are used to identify the subject's actions and posture. The input for this step is image data sent from the terminal, and the output is the analysis results regarding actions and posture.

[0740] Step 4:

[0741] The server analyzes the received audio data using natural language processing techniques to detect abnormal voice patterns and cries for help. The input for this step is audio data sent from the terminal, and the output is the analysis results regarding abnormal sounds and voice commands.

[0742] Step 5:

[0743] The server detects anomalies based on the analysis results of images and audio. It compares these results to pre-set criteria and determines an anomaly if the threshold is exceeded. The input for this step is the analysis results from steps 3 and 4, and the output is the determination of whether an anomaly occurred.

[0744] Step 6:

[0745] The server generates a notification based on the detected anomaly and sends it to the user's device. The notification includes the nature, time, and location of the anomaly, as well as recommended actions. The input for this step is the anomaly detection result from step 5, and the output is the notification information sent to the user.

[0746] Step 7:

[0747] Users can view and confirm notifications using smart glasses or smartphones and take action. The input for this step is notification information from the server, and the output is the user's response.

[0748] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0749] This invention is a system for enhancing monitoring activities in childcare and elderly care, incorporating an emotion engine that analyzes the emotional state of the subject in addition to effective anomaly detection. This system acquires and analyzes image and audio data to detect anomalies from normal behavior and sounds, and further enables complementary anomaly detection by identifying and evaluating the user's emotions using the emotion engine.

[0750] First, the device uses its camera and microphone to acquire images and audio data of the subject in real time. This data is immediately transmitted to the server. High-speed and stable communication methods are used for data transfer.

[0751] Next, the server analyzes the subject's movements from the received image data and detects abnormal patterns or cries for help from the audio data. Based on these analyses, it determines whether the subject is in a dangerous situation or exhibiting abnormal behavior.

[0752] Furthermore, the server uses an emotion engine to analyze the emotional state from the acquired images and audio. For example, through facial recognition and voice tone analysis, it identifies whether the subject is experiencing anxiety, anger, sadness, or other similar emotions. This enables anomaly detection that also addresses the subject's psychological state.

[0753] If an anomaly or a specific emotional state is detected, the server generates a notification based on a series of pieces of information. This notification includes an alert with adjusted urgency, including a risk assessment based on the subject's emotional state. The notification is sent to the user's smartphone or dedicated device, allowing the user to take immediate action.

[0754] For example, if an elderly person requiring special attention falls, the device captures the event, and the server detects the fall based on motion analysis. Simultaneously, an emotion engine senses anxiety from the elderly person's voice and facial expressions. The server then combines this information to generate a high-urgency alert, which is sent to the user to support a rapid response.

[0755] The system of this invention implements a continuous learning process using recorded analysis results to improve analysis accuracy. Furthermore, all history, including emotional data, is managed and utilized for future activities and analyses. This makes it possible to achieve both the safety of the subjects and a reduction in user burden.

[0756] The following describes the processing flow.

[0757] Step 1:

[0758] The device acquires continuous image data of the subject using its camera and collects audio data in real time using its microphone. This allows the subject's situation to be recorded in real time.

[0759] Step 2:

[0760] The terminal compresses the acquired data and sends it to the server using a high-speed communication protocol. This minimizes delays caused by data transfer, enabling real-time analysis.

[0761] Step 3:

[0762] The server uses the received image data to execute a motion analysis algorithm. It classifies the subject's posture and movements and identifies abnormal movements. For example, it can identify movements such as falls and slipping away.

[0763] Step 4:

[0764] The server analyzes the audio data and uses acoustic features and natural language processing to detect unusual sounds and specific calls. For example, loud calls and long silences may be identified as unusual.

[0765] Step 5:

[0766] The server activates an emotion engine, analyzing image and audio data to assess emotional states. Based on facial expression analysis and voice tone, it identifies emotions such as anxiety, anger, and sadness.

[0767] Step 6:

[0768] The server comprehensively evaluates the results of motion analysis, voice analysis, and emotional state to confirm the presence of an anomaly. If an anomaly is detected, the urgency level is set according to the emotional state.

[0769] Step 7:

[0770] The server generates a notification message based on the configured urgency level. The message includes details of the anomaly, the time and location of the occurrence, and recommended actions. The notification is sent to the user immediately.

[0771] Step 8:

[0772] Users can check received notifications on their smartphones or dedicated devices and decide on the necessary actions. For example, they can immediately head to the scene or check the video feed.

[0773] Step 9:

[0774] The server records and manages all analysis and detection results over the long term. This data is used to improve the accuracy of the analysis and to periodically update the machine learning algorithms.

[0775] (Example 2)

[0776] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0777] In monitoring activities for childcare and elderly care, there is a need to not only detect abnormal behavior in the person being monitored, but also to analyze their emotional state, enabling more comprehensive and rapid detection and response to abnormalities. However, conventional systems have the problem of being unable to adequately ensure the safety of the person being monitored because detailed analysis, including emotional state, is difficult.

[0778] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0779] In this invention, the server includes an image acquisition means, an audio acquisition means, and an emotion analysis means for analyzing the emotional state using the acquired image and audio data. This enables analysis of both the subject's behavior and emotional state, allowing for more accurate and rapid anomaly detection and notification.

[0780] "Image acquisition means" refers to a device or method used to acquire images of a subject, such as a camera, which collects data in real time.

[0781] "Voice acquisition means" refers to a device or method used to acquire the voice of a subject, and involves collecting voice data in real time using a microphone or the like.

[0782] "Analysis means" refers to the process and techniques for analyzing acquired data and identifying abnormalities from the subject's behavior and voice.

[0783] "Emotional analysis means" refers to the process and technology for identifying and analyzing a subject's emotional state using acquired image and audio data.

[0784] "Detection means" refers to the processes and technologies used to discover and judge abnormal phenomena or specific emotional states based on analyzed data.

[0785] "Notification means" refers to a mechanism for generating appropriate notifications based on detected anomalies or emotional states and sending relevant information to the user.

[0786] "Recording means" refers to a system for saving analyzed data and managing it as a history.

[0787] "Learning methods" refer to processes and techniques that utilize recorded data to improve the accuracy of system analysis.

[0788] This invention is a system for enhancing monitoring activities in childcare and elderly care. It analyzes image and audio data acquired in real time to determine the behavior and emotional state of the subject, thereby detecting and notifying of abnormalities.

[0789] First, the device uses its camera and microphone to acquire images and audio data of the subject. The camera is high-resolution, capturing the subject's face and movements in detail. The microphone is highly sensitive, capturing the subject's voice clearly. The acquired data is immediately transmitted to the server via a network with high communication speed and stability.

[0790] The server analyzes the subject's movements by executing a deep learning-based image recognition algorithm on the received image data. In addition, it uses speech recognition software to detect abnormal patterns and cries for help in the audio data. Furthermore, an emotion engine is incorporated, which analyzes the subject's emotional state based on the acquired data through facial recognition technology and speech analysis. This analysis allows for the identification of emotions such as anxiety, anger, and sadness, enabling anomaly detection that also takes psychological circumstances into account.

[0791] If an anomaly or a specific emotional state is detected, the server generates a notification based on the detected information and sends an alert to the user's smartphone or dedicated terminal. This notification includes a risk assessment based on the person's condition, allowing the user to understand the urgency of the situation and take immediate action.

[0792] For example, if an elderly person with dementia falls indoors, the device captures changes in their movements, and the server recognizes the fall from the image data. Simultaneously, it analyzes voice data for signs of pain, and the emotion engine detects feelings of anxiety. From this series of pieces of information, the server immediately generates a high-urgency alert and notifies the user.

[0793] This system features a learning function that continuously improves analysis accuracy using recorded analysis data. By using a generative AI model to execute prompts such as, "Analyze the subject's current condition in detail and report immediately if there are any abnormalities," it enables even more precise monitoring and analysis.

[0794] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0795] Step 1:

[0796] The device uses a camera to acquire images of the subject in real time and a microphone to collect audio data. Physical images and audio are provided as input, which are converted into digital data. The output is the digitized image and audio data. This data is then sent to a server for further analysis.

[0797] Step 2:

[0798] The terminal transmits the acquired image and audio data to the server via a high-bandwidth network. The input is the image and audio data acquired by the terminal, which is then packetized and transmitted. The output is the packetized data received on the server side.

[0799] Step 3:

[0800] The server applies a deep learning-based image recognition algorithm based on the received image data. The input is digitized image data, which is analyzed to identify the subject's movements. The output is data containing the results of the movement analysis.

[0801] Step 4:

[0802] The server analyzes audio data using a speech recognition algorithm. The input is digital audio data, which is analyzed to identify unusual patterns and specific speech content. The output is data containing the results of the speech analysis.

[0803] Step 5:

[0804] The server uses an emotion engine to analyze emotional states from image and audio data. The input is the analysis results of the images and audio, which are used to evaluate the emotional state. The output is data containing the results of the emotion analysis.

[0805] Step 6:

[0806] The server integrates motion analysis results, voice analysis results, and emotion analysis results to detect abnormal conditions. The input is a dataset containing the results of each analysis, which is then analyzed to perform a risk assessment. The output is alert information containing the results of the anomaly detection.

[0807] Step 7:

[0808] The server generates notifications based on detected anomalies and emotional states and sends them to the user's smartphone or dedicated terminal. The input is notification generation data based on anomaly detection and emotional evaluation, and the output is the notification information received by the user.

[0809] (Application Example 2)

[0810] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0811] In modern society, there is a need for means to ensure personal safety. In particular, when going out or acting alone, it is important to be able to understand the surrounding environment in real time and quickly detect the occurrence of abnormalities or psychological changes. However, with current technologies, anomaly detection and emotion analysis are performed separately, and there is insufficient means to respond immediately based on integrated information.

[0812] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0813] In this invention, the server includes an image acquisition means, an analysis means for analyzing the acquired images to identify the subject's behavior, a detection means for detecting anomalies based on the analysis results, an evaluation means that works in conjunction with the analysis means to analyze the emotional state, an adjustment means for performing a risk assessment based on the emotional state and adjusting the notification content, and a monitoring means for monitoring the surrounding situation. This ensures the safety of individuals and enables a quick and appropriate response.

[0814] "Image acquisition means" refers to a means of acquiring image data of a target in real time using an imaging device such as a camera.

[0815] "Analysis means" refers to methods for analyzing acquired image and audio data to identify the subject's behavior and emotional state.

[0816] A "detection means" is a means of detecting anomalies based on analyzed data and determining whether or not there is a problem.

[0817] A "notification system" is a means of generating and sending appropriate alerts based on detected anomalies or risks.

[0818] "Evaluation methods" refer to means of analyzing the emotional state of a subject from acquired data and identifying their psychological state.

[0819] "Adjustment measures" refer to methods for evaluating risk based on the results of an analysis of emotional states and adjusting notification content according to urgency.

[0820] "Monitoring means" refers to methods for constantly monitoring the user's surroundings and detecting abnormalities or dangers.

[0821] "Voice acquisition means" refers to a means of acquiring target voice data using a microphone.

[0822] "Display means" refers to means of visually showing users warnings or notifications regarding detected anomalies.

[0823] A "recording means" is a means of saving the analyzed data as a history for later review or learning.

[0824] A "learning method" is a means of continuously optimizing a system to improve the accuracy of analysis based on recorded data.

[0825] The system for implementing this invention includes smart glasses worn by the user and a data analysis device on a server. The smart glasses are equipped with a camera and microphone, which acquire images and audio data of the surroundings in real time. This data is immediately transmitted to the server via a stable communication means.

[0826] The server uses image analysis software and audio analysis software to analyze the received image and audio data. This allows the server to analyze the subject's behavior and surrounding environment and detect anomalies. Furthermore, to analyze emotional states, it uses a dedicated emotion engine to perform facial recognition and voice tone analysis, thereby evaluating the subject's psychological state.

[0827] If an anomaly is detected, or if sentiment analysis determines that there is a high psychological risk, the server will immediately generate a warning using the notification system and display it on the user's smart glasses display. This notification not only encourages the user to take safe actions but also helps them to take prompt action if necessary.

[0828] As a concrete example, imagine a situation where a user is walking alone at night and suddenly hears loud footsteps around them. In such a case, the smart glasses immediately send data to the server, detect the anomaly, and then visually display a warning such as "Watch your back."

[0829] An example of a prompt for the generating AI model might be, "Explain how to analyze anomalies and emotions in real time from audio and image data to ensure safety." This system provides an effective method for improving user safety.

[0830] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0831] Step 1:

[0832] When a user puts on smart glasses, the device activates its camera and microphone to capture images and sounds from the surroundings in real time. Based on the input, it outputs streams of image and audio data. The device periodically sends this data to a server.

[0833] Step 2:

[0834] The server inputs the received image data into image analysis software to identify the subject's actions and surrounding environment. This allows for the analysis of behavioral patterns such as walking and the approach of others. The output is the analyzed behavioral information.

[0835] Step 3:

[0836] Simultaneously, the server processes the audio data using audio analysis software to detect abnormal sounds and volume changes. For example, this could include the sound of footsteps suddenly approaching. The input is audio data, and the output is the detected abnormal sound information.

[0837] Step 4:

[0838] The server integrates the results of image and audio analysis and uses an emotion engine to evaluate the subject's emotional state. It identifies anxiety and surprise through facial recognition and voice tone analysis. The output is information evaluating the emotional state.

[0839] Step 5:

[0840] The server performs a risk assessment based on all analysis results and generates necessary alerts. For example, it might generate a notification such as "Watch your back." The input is the analysis results, and the output is the alert information.

[0841] Step 6:

[0842] The generated notification is displayed instantly on the user's smart glasses using the device's interface. The user can then see it and take action to pay attention to their surroundings. The output is a visual warning.

[0843] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0844] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0845] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0846] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0847] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0848] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0849] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0850] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0851] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0852] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0853] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0854] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0855] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0856] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0857] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0858] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0859] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0860] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0861] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0862] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0863] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0864] The following is further disclosed regarding the embodiments described above.

[0865] (Claim 1)

[0866] Image acquisition method,

[0867] An analysis means that analyzes acquired images to identify the subject's actions,

[0868] A detection means for detecting anomalies based on the analysis results,

[0869] A notification means that generates and sends a notification based on the detected anomaly,

[0870] A system that includes this.

[0871] (Claim 2)

[0872] A means of acquiring sound,

[0873] An analysis means for detecting anomalies by analyzing acquired audio,

[0874] The system according to claim 1, further comprising:

[0875] (Claim 3)

[0876] A recording means for recording the analyzed data and managing it as a history,

[0877] A learning method that utilizes recorded data to improve analysis accuracy,

[0878] The system according to claim 1, further comprising:

[0879] "Example 1"

[0880] (Claim 1)

[0881] Image acquisition method,

[0882] Information processing means for analyzing acquired images to identify the actions of the subject,

[0883] A means for acquiring sound data,

[0884] Information processing means for analyzing acquired audio to detect anomalies,

[0885] An evaluation means for detecting anomalies from processed data,

[0886] An information provision means that generates and sends a notification when an anomaly is detected,

[0887] A system that includes this.

[0888] (Claim 2)

[0889] A management system for recording and managing the analyzed data as a history,

[0890] A learning method that utilizes recorded data to improve analysis accuracy,

[0891] The system according to claim 1, including the following:

[0892] (Claim 3)

[0893] A behavioral identification means for analyzing behavioral patterns in relation to the subject's condition,

[0894] A criteria setting means that sets criteria for detecting abnormalities based on the behavior detected by the behavior identification means,

[0895] The system according to claim 1, including the following:

[0896] "Application Example 1"

[0897] (Claim 1)

[0898] Image acquisition method,

[0899] An analysis means that analyzes acquired images to identify the subject's actions,

[0900] A detection means for detecting anomalies based on the analysis results,

[0901] A notification means that generates and sends a notification based on the detected anomaly,

[0902] A display means for outputting notifications visually and audibly,

[0903] A system that includes this.

[0904] (Claim 2)

[0905] A means of acquiring sound,

[0906] An analysis means for detecting anomalies by analyzing acquired audio,

[0907] The system according to claim 1.

[0908] (Claim 3)

[0909] A recording means for recording the analyzed data and managing it as a history,

[0910] A learning method that utilizes recorded data to improve analysis accuracy,

[0911] The system according to claim 1.

[0912] "Example 2 of combining an emotion engine"

[0913] (Claim 1)

[0914] Image acquisition method,

[0915] A means of acquiring sound,

[0916] An analysis means that analyzes acquired images to identify the subject's actions,

[0917] An analysis means for detecting anomalies by analyzing acquired audio,

[0918] An emotion analysis method that analyzes emotional states using analyzed image and audio data,

[0919] A detection means for detecting abnormalities or specific emotional states based on the analysis results,

[0920] A notification means that generates and sends notifications based on detected anomalies or emotional states,

[0921] A system that includes this.

[0922] (Claim 2)

[0923] A recording means for recording the analyzed data and managing it as a history,

[0924] A learning method that utilizes recorded data to improve analysis accuracy,

[0925] The system according to claim 1, further comprising:

[0926] (Claim 3)

[0927] The system according to claim 1, comprising communication means for a terminal to acquire image and audio data in real time and transmit it to a server.

[0928] "Application example 2 when combining with an emotional engine"

[0929] (Claim 1)

[0930] Image acquisition method,

[0931] An analysis means that analyzes acquired images to identify the subject's actions,

[0932] A detection means for detecting anomalies based on the analysis results,

[0933] A notification means that generates and sends a notification based on the detected anomaly,

[0934] An evaluation means that works in conjunction with an analysis means to analyze emotional states,

[0935] A means of adjusting the content of notifications by conducting a risk assessment based on emotional state,

[0936] A monitoring device for monitoring the surrounding situation,

[0937] A system that includes this.

[0938] (Claim 2)

[0939] A means of acquiring sound,

[0940] An analysis means for detecting anomalies by analyzing acquired audio,

[0941] A display means that displays a warning when an anomaly is detected,

[0942] The system according to claim 1, further comprising:

[0943] (Claim 3)

[0944] A recording means for recording the analyzed data and managing it as a history,

[0945] A learning method that utilizes recorded data to improve analysis accuracy,

[0946] A notification means for notifying the results of detecting abnormalities and emotional states,

[0947] The system according to claim 1, further comprising: [Explanation of Symbols]

[0948] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. Image acquisition method, An analysis means that analyzes acquired images to identify the subject's actions, A detection means for detecting anomalies based on the analysis results, A notification means that generates and sends a notification based on the detected anomaly, A system that includes this.

2. A means of acquiring sound, An analysis means for detecting anomalies by analyzing acquired audio, The system according to claim 1, further comprising:

3. A recording means for recording the analyzed data and managing it as a history, A learning method that utilizes recorded data to improve analysis accuracy, The system according to claim 1, further comprising:

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A