system
A wearable device with AI-powered video and audio analysis addresses the limitations of GPS-based tracking by providing real-time safety alerts to guardians, enhancing child safety.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-21
- Publication Date
- 2026-05-07
AI Technical Summary
Conventional systems that rely solely on GPS for tracking children's location are inadequate in preventing abductions and device loss, and fail to provide real-time safety measures when guardians are distant.
A wearable monitoring device equipped with a camera and microphone that uses artificial intelligence for real-time video and audio analysis, issuing warnings and alerting guardians via communication devices when suspicious situations are detected.
Ensures the safety of children by promptly identifying and alerting guardians to potential dangers, providing real-time monitoring and reducing parental anxiety.
Smart Images

Figure 2026074854000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In recent years, the risks of crimes and accidents against children have been increasing, and ensuring safety is an important issue, especially when moving alone. Conventional systems that only track location information using the GPS function cannot take sufficient measures against abduction by suspicious persons or device abandonment. Also, when the guardian is in a distant location, it is difficult to grasp the situation in real time, and an improvement in safety is required.
Means for Solving the Problems
[0005] This invention provides a means for acquiring surrounding video and audio using a wearable monitoring device. The acquired data is analyzed in real time by artificial intelligence within the device to identify suspicious individuals and determine dangerous situations. As a result, if a suspicious situation is detected, a warning is immediately issued by a warning device, and the alert is notified to the guardian via a communication device. This makes it possible to effectively monitor the safety of children from a distance and avoid risks in advance.
[0006] A "wearable monitoring device" is an electronic device that has a shape that can be worn on the body, is portable, and has monitoring functions.
[0007] A "camera system" is a functional component for acquiring images of the surroundings, and includes an image sensor and lens.
[0008] A "microphone system" is a functional component for acquiring ambient sounds, and includes a microphone that converts sound waves into electrical signals.
[0009] "Artificial intelligence means" refers to functions including algorithms and software for analyzing acquired data to identify suspicious individuals and determine dangerous situations.
[0010] "Warning means" refers to functions that include actuators and audio output devices for issuing warnings when suspicious situations are detected.
[0011] "Communication means" refers to functions including communication modules and interfaces for transmitting acquired information and warnings to a remote location.
[0012] A "face recognition algorithm" is a computer program process that detects human faces in video data and identifies specific individuals.
[0013] An "audio warning" is a warning signal that uses sound to alert those nearby to a suspicious person or a dangerous situation. [Brief explanation of the drawing]
[0014] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
Embodiments for Carrying Out the Invention
[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, a labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0018] In the following embodiments, a labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0019] In the following embodiments, a labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0022] [First Embodiment]
[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0035] The wearable monitoring device of the present invention has a shape that is not uncomfortable for children to wear on a daily basis, and is configured as a scarf-type or necklace-type device. This device has a built-in camera to acquire surrounding video and a microphone to acquire audio. This acquired data is analyzed in real time by an artificial intelligence system built into the device, and dangerous situations are detected using a facial recognition algorithm and natural language processing technology.
[0036] If the device detects a suspicious person or dangerous situation, it will immediately emit a warning sound to alert those nearby. This warning sound is emitted through a speaker installed inside the device. Furthermore, analysis results and warning information are sent to a server via a communication module within the device. The server then uses this information to notify the user (parent / guardian) as an alert and provides relevant data through an application or web interface.
[0037] Specifically, consider a scenario where children are playing in a park after school. The device continuously films the surrounding area of the park, collecting data through its camera and microphone. The artificial intelligence within the device attempts to detect suspicious individuals from this data. If an unfamiliar person approaches the children, it recognizes their face and analyzes what they say. If danger is detected, details are reported to the server, and an immediate alert is sent to the user. In this case, the recorded data is stored as evidence in case of any unforeseen circumstances.
[0038] This system allows parents to monitor their children's safety from home, reducing unnecessary worries. Furthermore, the collected data serves as a valuable source of information for improving safety. Thus, this invention provides a new way to ensure children's safety through the power of technology.
[0039] The following describes the processing flow.
[0040] Step 1:
[0041] The device connects to the server upon startup and performs device authentication. If authentication is successful, the server sends various configuration information to the device and starts monitoring mode.
[0042] Step 2:
[0043] The device uses its built-in camera and microphone to capture surrounding video and audio in real time. This data is temporarily stored in the device's internal memory.
[0044] Step 3:
[0045] The artificial intelligence installed in the device analyzes the acquired video data and applies a facial recognition algorithm to identify human faces. Simultaneously, audio data is analyzed using natural language processing technology to detect specific keywords and unusual conversational content.
[0046] Step 4:
[0047] The device will immediately emit a warning sound based on the information it receives if it detects a suspicious person or dangerous situation. This audio warning is intended to alert those in the surrounding area.
[0048] Step 5:
[0049] The device sends data to the server when it detects important events or suspicious situations. The server receives this data and sends an alert to the user (parent / guardian) via an application or email.
[0050] Step 6:
[0051] The server securely stores audio and video data received from the terminal and makes it accessible to the user. Users can view this data through their logged-in account.
[0052] Step 7:
[0053] Users can use the collected data to understand the safety status of their children and prepare to provide the data to legal authorities or other relevant parties as needed.
[0054] (Example 1)
[0055] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0056] There is a need for a means to monitor dangerous situations that children may encounter in their daily lives in real time and to notify parents quickly and effectively. However, conventional monitoring devices have problems such as insufficient accuracy in detecting anomalies and often delays in information transmission. To address this situation and ensure the safety of children, there is a need to provide more advanced and efficient monitoring systems.
[0057] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0058] In this invention, the server includes sensor means for acquiring video, sensor means for acquiring audio, and information processing means for analyzing the acquired video and audio data in real time. This makes it possible to immediately detect suspicious situations and quickly transmit information to parents.
[0059] A "sensor" is a device that detects video and audio information and acquires it as a digital signal.
[0060] "Information processing means" refers to a computer system that analyzes acquired video and audio data to detect specific patterns or anomalies.
[0061] "Output means" refers to a device that issues a visual or audible warning to the user in order to communicate detected abnormal information.
[0062] "Information and communication means" refers to network devices used to notify parents in remote locations in real time of acquired and analyzed data.
[0063] A "data storage device" is a recording device that securely stores acquired and analyzed information for later review and use.
[0064] A "generative AI model" is a machine learning algorithm trained to analyze diverse data and identify suspicious situations.
[0065] In this embodiment of the invention, the system is composed of three main elements: a terminal, a server, and a user, through their interaction.
[0066] Device functions and configuration
[0067] The device is designed as a wearable device and is equipped with sensors to acquire video and audio. Specifically, it has a built-in high-performance camera and microphone that capture information about the surrounding environment. The acquired data is analyzed in real time by the device's information processing system. This analysis uses generative AI models, including facial recognition and speech recognition.
[0068] Server Functions
[0069] The server has the means of information and communication to receive analysis data sent from the terminal. Based on the received data, the server can perform more complex analysis and data reintegration, thereby providing the user with an accurate situational analysis.
[0070] User operation and interface
[0071] Users communicate with their devices and servers through a dedicated mobile application or web interface. This allows users to receive real-time alerts from the system and take necessary actions quickly.
[0072] Specific example
[0073] For example, imagine a scenario where the user's child is playing in a park. The device monitors the surrounding sounds and instantly detects the appearance of suspicious individuals or suspicious remarks. This triggers an alert to the user via the server. The user can then view live video through the application and, if deemed necessary, report the incident to the police or other authorities.
[0074] Example of a prompt
[0075] An example of a prompt for a generated AI model would be, "Explain step-by-step how the AI system identifies a suspicious person and sends an alert to the user."
[0076] This system can provide a new means of technically ensuring children's safety in their daily lives and giving parents peace of mind.
[0077] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0078] Step 1:
[0079] The device acquires surrounding video and audio using sensory means. The input is the video and audio of the environment, which is converted into digital data. This is done by the camera capturing video frames in real time and the microphone recording audio. The output is the digital video and audio data, which is fed into subsequent processing steps.
[0080] Step 2:
[0081] The terminal transmits the acquired video and audio data to an information processing system for initial processing. The input is the digital data generated in step 1, and the output is clean data that has undergone noise reduction and basic format conversion. This processing involves the application of a filtering algorithm, which plays a role in improving the quality of the data.
[0082] Step 3:
[0083] The terminal's information processing means applies a generated AI model using clean data to analyze the data. The input is the clean data obtained in step 2, and the output is behavior recognition data and speech recognition data as analysis results. Through this analysis, specific actions are taken to identify suspicious facial features and suspicious statements of suspicious individuals using pattern matching.
[0084] Step 4:
[0085] Based on the analysis results, the terminal determines whether to issue a warning through the output means. The input is the analysis results from step 3, and the output is the generation of a warning sound. In action, the audio output device sounds a pre-set warning sound to alert those in the surrounding area of danger.
[0086] Step 5:
[0087] The terminal uses communication methods to send analysis results and warning information to the server. The input is detailed data about the situation in which the warning was issued, and the output is the transmission of data to the remote server. This process includes encryption to ensure data security.
[0088] Step 6:
[0089] The server notifies the user based on the information it receives. Input is analysis and warning data sent from the terminal, and output is an alert notification sent to the user's device. The server sends the notification to the user through a dedicated application, including push notifications on mobile devices and email notifications.
[0090] (Application Example 1)
[0091] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0092] Conventional technologies often make it difficult to monitor children's safety in real time and respond quickly to abnormal situations. Furthermore, methods for monitoring a child's surroundings when a parent is away are often limited, resulting in insufficient safety even under normal circumstances. This makes it difficult to respond immediately when a child actually faces a dangerous situation, causing anxiety for parents.
[0093] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0094] In this invention, the server includes imaging means for acquiring surrounding images, sound acquisition means for acquiring sound, and information processing means for analyzing the acquired data in real time. This makes it possible to recognize suspicious individuals or emergencies approaching children and immediately send warnings to guardians.
[0095] A "wearable monitoring device" is a monitoring device that is designed to be worn on the body and used on a daily basis.
[0096] "Image acquisition means" refers to a device or module that has the function of acquiring images.
[0097] "Sound acquisition means" refers to a device or module that has the function of acquiring sound.
[0098] "Information processing means" refers to a device or program that has the function of analyzing acquired data and extracting necessary information.
[0099] An "alarm system" is a device that has the function of issuing a warning through voice or digital notification when an abnormality is detected.
[0100] "Information and communication means" refers to a device or module that has the function of sending and receiving data with other terminals or systems.
[0101] A "biometric authentication algorithm" is a computational method used to identify individuals using biometric information such as facial features or fingerprints.
[0102] An "acoustic alarm" is a method of issuing a warning using sound.
[0103] "Electronic notification" refers to the use of digital means to notify other devices of warning content.
[0104] This wearable monitoring system is designed to monitor the surrounding environment in real time to ensure the safety of children.
[0105] The server uses information processing tools to analyze data transmitted from terminals. Specifically, these tools utilize software libraries such as TENSORFLOW® and OpenCV to perform facial recognition and natural language processing. The collected data is divided into image data and audio data. Specific patterns and individuals are extracted from the video data, and specific phrases and tones are analyzed from the audio data. This makes it possible to recognize abnormal individuals or situations and issue an immediate alarm.
[0106] Users can receive alerts via mobile devices such as smartphones. Information and communication methods utilize Wi-Fi or mobile networks to quickly transmit notifications from the server. For example, if a child is approached by a stranger in a park, the device captures the child's face, and the AI system immediately analyzes the data, notifying the user's smartphone of the results. The collected data is stored as evidence as needed.
[0107] The generative AI model can use this data to continue learning and achieve more accurate recognition in the future. Furthermore, it's possible to obtain detailed explanations in the format of prompts such as, "Please explain the purpose and functions of the child monitoring app. Please provide a detailed explanation including specific situational examples."
[0108] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0109] Step 1:
[0110] The terminal activates an imaging means for capturing surrounding video and an audio acquisition means for acquiring sound. During this process, the imaging means acquires video data as input, and the audio acquisition means acquires audio data as input. These data are then transmitted to an information processing means as output.
[0111] Step 2:
[0112] The device analyzes the acquired video data using information processing tools and identifies a specific person using a face recognition algorithm. TensorFlow and OpenCV are used to filter the face information within the video. The input is video data, and the output is the recognized facial features.
[0113] Step 3:
[0114] The device analyzes the audio data using a natural language processing algorithm to understand its content. This process uses a speech recognition model to extract the spoken content as text data. The input is audio data, and the output is the analyzed text content.
[0115] Step 4:
[0116] The server integrates data extracted from video and audio to assess the level of risk. Based on the integrated data, it identifies suspicious individuals and detects abnormal situations. The input consists of facial feature data and text data, and the output is the risk assessment result.
[0117] Step 5:
[0118] The server activates alarm mechanisms based on the risk assessment results. If an anomaly is detected, it immediately uses the terminal's alarm mechanism to emit an audible alarm and also sends an electronic notification to the user via information and communication means. The input is the risk assessment result, and the output is the activation of the audible alarm and the transmission of the electronic notification.
[0119] Step 6:
[0120] The user receives a notification on their mobile device and checks its contents. The user's device immediately displays details of the threat using the received notification data. The input is electronic notification data, and the output is the information displayed on the device.
[0121] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0122] The system of the present invention provides even more advanced monitoring capabilities by implementing an emotion engine that recognizes the user's emotions, in addition to conventional wearable monitoring devices. The device is small and portable, in the form of a necklace or scarf, and can be worn naturally by children in their daily lives. The device incorporates a camera means for acquiring surrounding images, a microphone means for acquiring surrounding sounds, and artificial intelligence.
[0123] The emotion engine analyzes acquired voice data to determine the user's emotional state. This engine employs algorithms to identify emotions such as anger and anxiety by analyzing the tone and speaking patterns of the voice. This information is incorporated into the device's decision-making process, and if a specific emotional state is detected, the response through warning measures is adjusted. For example, if the user's voice is judged to be increasingly urgent, measures such as increasing the volume of the warning sound or shortening the interval between repetitions will be taken.
[0124] The device continuously captures surrounding video and audio and analyzes the data in real time. This analysis helps detect suspicious individuals and determine dangerous situations. The data analysis results, which reflect the user's emotions via an emotion engine, are sent to the server to more accurately assess the urgency of the situation. Based on this data, the server sends optimized alerts to the user (parent / guardian).
[0125] For example, if a child is approached by a stranger while out and about, and the emotion engine detects that the child's voice sounds anxious, that information is immediately sent to the server. Based on this, the server sends a high-priority alert to the parent, enabling them to take necessary measures quickly.
[0126] Thus, this system, equipped with an emotion engine, provides a more multifaceted support for children's safety and offers a new way for parents to monitor their children with peace of mind, even from a distance.
[0127] The following describes the processing flow.
[0128] Step 1:
[0129] When the device is powered on, it connects to the server and performs an authentication procedure. If successful, the device receives the necessary configuration information from the server and switches to monitoring mode.
[0130] Step 2:
[0131] The device uses its internal camera and microphone to capture surrounding video and audio in real time, and temporarily stores that data in its internal memory.
[0132] Step 3:
[0133] The artificial intelligence installed in the device uses the acquired video data to execute a facial recognition algorithm to identify and distinguish human faces. Meanwhile, natural language processing techniques are applied to the audio data to analyze the content of the conversation.
[0134] Step 4:
[0135] The emotion engine evaluates the tone and patterns of the user's voice from the analyzed audio data to determine their emotional state (e.g., anger, anxiety, surprise, etc.).
[0136] Step 5:
[0137] The device will emit a warning sound if it detects a suspicious person, a dangerous situation, or if the user's emotions become urgent. In this case, the type and intensity of the warning sound will be adjusted based on the emotion engine's judgment.
[0138] Step 6:
[0139] The device sends information, including important events and emotional data, to the server. Based on the received data, the server assesses the urgency of the events and sends optimized alerts to the user (parent / guardian).
[0140] Step 7:
[0141] Users receive alerts from the server, check the situation via application or email, and take appropriate action. The collected data is stored for detailed situation analysis and necessary follow-up actions.
[0142] (Example 2)
[0143] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0144] In recent years, there has been a growing need for monitoring devices that are both miniaturized and portable. However, conventional technologies have struggled to accurately understand the user's emotional state and assess the urgency of a situation. As a result, there are problems with delays in providing appropriate warnings and alerts. In particular, in monitoring child safety, the lack of emotional recognition hinders a swift and appropriate response.
[0145] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0146] In this invention, the server includes optical means for acquiring surrounding video, acoustic means for acquiring sound, data processing means for analyzing acquired video and audio data in real time, emotion analysis means for analyzing the tone and pattern of sound to determine emotions, warning means for issuing a warning when a suspicious situation or a specific emotional state is detected, and communication means for sending alerts to guardians. This enables more precise monitoring that always takes the user's emotional state into consideration, and immediate alert issuance.
[0147] A "wearable monitoring device" is a monitoring device that a user can wear and carry with them at all times, and is usable in everyday life.
[0148] "Optical means" refers to imaging devices such as cameras used to acquire visual information of the surroundings.
[0149] "Acoustic means" refers to sound acquisition devices such as microphones used to acquire ambient sound information.
[0150] "Data processing means" refers to computer processing functions that analyze acquired video and audio data in real time and extract necessary information.
[0151] An "emotion analysis tool" is a system equipped with an algorithm that determines the user's emotional state by analyzing the tone and patterns of their voice.
[0152] A "warning device" is a device or mechanism that issues a warning to the user or their surroundings when a suspicious situation or a specific emotional state is detected.
[0153] "Communication means" refers to network communication functions used to transmit analysis results and warning information to the user's guardian or other supervisors.
[0154] "Suspicious circumstances" refer to unusual or potentially dangerous behavior or environments, meaning that the monitored behavior or state deviates from normal circumstances.
[0155] A "specific emotional state" is a mental state that requires immediate attention, such as anger or anxiety, which is determined by emotional analysis to be an emotional change that exceeds a predetermined standard.
[0156] The device is designed as a wearable monitoring device that users can wear and use on a daily basis. Specifically, the device is equipped with an optical means, a camera, and an acoustic means, a microphone. These devices are used to acquire surrounding video and audio data. The device also implements data processing means to analyze the acquired data in real time. This processing includes emotion analysis means that analyze the tone and patterns of voice to determine emotions.
[0157] The device is designed as a necklace or scarf that the user wears, making it lightweight, easy to carry, and comfortable for children to wear naturally. The emotion analysis system analyzes voice tone to identify the user's emotional state. If this process detects specific emotions such as anger or anxiety, the warning system will activate accordingly, providing necessary alerts.
[0158] The server receives data transmitted from the device via communication and sends alerts to the parent based on the analysis results. For example, if the device detects a child's anxious voice, that information is sent to the server, which can then send a high-priority alert to the parent's device. This allows the parent to immediately understand their child's situation and take necessary action.
[0159] As a concrete example, when a user is playing in a park, the device can capture ambient sounds and process them using emotion analysis to detect the child's anxiety when a suspicious person suddenly approaches. This information is transmitted to a server in real time, and a warning is sent to the parents.
[0160] Examples of prompts for a generative AI model:
[0161] "Please describe in detail how the device works when a child is making anxious noises."
[0162] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0163] Step 1:
[0164] The device acquires video data from its surroundings through optical means (camera) and audio data through acoustic means (microphone). The input is the surrounding video and audio, and the output is the video and audio data converted into their respective data formats. Specifically, the device constantly collects information through the camera and microphone in accordance with the user's movements.
[0165] Step 2:
[0166] The terminal analyzes the acquired video and audio data in real time using data processing means. The input is the video and audio data acquired in step 1, and the output is the analyzed data. This analysis uses emotion analysis means to identify emotions by analyzing changes in voice tone and speaking patterns. Specifically, the data processing algorithm within the terminal analyzes the voice tone and identifies emotions such as anger and anxiety.
[0167] Step 3:
[0168] The device activates a warning mechanism when a specific emotional state of the user (e.g., anxiety) is detected through emotion analysis. The input is the emotion information analyzed in step 2, and the output is the issuance of a warning. Specifically, the device emits an audio warning through its speaker.
[0169] Step 4:
[0170] The terminal transmits the analyzed data and warning information to the server via a communication method. The input is the analyzed data and warning information resulting from steps 2 and 3, and the output is a data packet containing this information. Specifically, the terminal transmits data to the server via wireless communication.
[0171] Step 5:
[0172] The server receives data sent from the device and processes it to send alerts to the parent. The input is the data packets received from the device, and the output is the alert message sent to the parent. Specifically, the server sends a push notification to the parent's communication device to report the situation.
[0173] (Application Example 2)
[0174] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0175] In modern home environments, threats from intruders and noises outside windows exist, making it especially important to ensure the safety of family members. Furthermore, in emergencies, it is necessary to respond quickly and accurately, taking into account emotional states. However, conventional security systems have struggled to adequately adjust warnings in real time to reflect the user's emotional state, making it difficult to fully guarantee the safety of the home.
[0176] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0177] In this invention, the server includes means equipped with an emotion engine for analyzing voice data to determine emotional states, image acquisition means for acquiring surrounding video, and voice acquisition means for acquiring audio. This enables rapid detection of abnormal emotional changes or emergencies within the home and the sending of optimized alerts to parents.
[0178] An "emotion engine" is a mechanism that includes an algorithm for analyzing voice data to determine the user's emotional state.
[0179] "Image acquisition means" refers to a device or technology for collecting surrounding video data.
[0180] "Voice acquisition means" refers to a device or technology for collecting ambient sound data.
[0181] "Artificial intelligence means" refers to a computer program or system that analyzes acquired video and audio data in real time and makes appropriate decisions.
[0182] A "warning mechanism" is a device that issues a warning by sound or other means when a specific situation occurs.
[0183] An "information transmission means" is a system equipped with communication functions for transmitting analysis results and warning information to other devices or users.
[0184] The system used to realize this application is configured as a smart security system that enhances safety for users within their homes and surrounding environments. The system acquires audio and video data and analyzes them using artificial intelligence. Specifically, for audio data analysis, the Google® Cloud Speech-to-Text API is used to convert speech to text, and then IBM Watson® Tone Analyzer is used to analyze the emotions contained in the speech.
[0185] Based on the acquired sentiment analysis results, the server optimizes alerts regarding the user's emotional state and sends alerts to the user's device using communication methods. In this process, the server uses an emotion engine to detect emotions such as anxiety and fear and adjusts the warning methods accordingly. If an emergency situation that causes the user to feel fear is detected, a real-time alert, including an audio warning, is transmitted to other members of the household.
[0186] For example, if a suspicious sound is heard at home at night, the system analyzes the tone of the sound and immediately issues a warning if fear is detected. This allows the whole family to respond quickly and take safety precautions more easily.
[0187] Examples of prompts to input into a generative AI model:
[0188] "If sounds indicating anxiety or fear are detected within a residence, please generate a statement explaining what kind of alert should be generated based on the emotion recognition results."
[0189] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0190] Step 1:
[0191] The terminal acquires audio and video data from within the home. Audio data is collected via a microphone, and video data is collected using an image acquisition device. This data is transmitted to the server in its original format.
[0192] Step 2:
[0193] The server converts received audio data into text data using the Google Cloud Speech-to-Text API. The input is audio data, and the output is the converted text data. This conversion makes it possible to recognize the content of the audio as text information.
[0194] Step 3:
[0195] The server inputs text data into IBM Watson Tone Analyzer for sentiment analysis. The input is the converted text data, and the output is the analysis result of the emotional state contained in that text (e.g., calm, anxious, fear). This analysis allows for a detailed understanding of the user's emotional state.
[0196] Step 4:
[0197] The server generates and adjusts alerts based on the emotion analysis results when specific emotions are detected. For example, if a strong emotion of fear is detected, the alert's urgency is increased by raising the volume of the warning device. The input is the emotion analysis result, and the output is the adjusted alert information.
[0198] Step 5:
[0199] The server transmits coordinated alert information to the user's terminal using an information transmission method. The input is the coordinated alert information, and the output is a warning message displayed or audibly output on the user's terminal. This allows the user to immediately understand the situation and take necessary actions quickly.
[0200] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0201] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0202] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0203] [Second Embodiment]
[0204] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0205] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0206] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0207] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0208] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0209] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0210] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0211] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0212] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0213] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0214] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0215] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0216] The wearable monitoring device of the present invention has a shape that is not uncomfortable for children to wear on a daily basis, and is configured as a scarf-type or necklace-type device. This device has a built-in camera to acquire surrounding video and a microphone to acquire audio. This acquired data is analyzed in real time by an artificial intelligence system built into the device, and dangerous situations are detected using a facial recognition algorithm and natural language processing technology.
[0217] If the device detects a suspicious person or dangerous situation, it will immediately emit a warning sound to alert those nearby. This warning sound is emitted through a speaker installed inside the device. Furthermore, analysis results and warning information are sent to a server via a communication module within the device. The server then uses this information to notify the user (parent / guardian) as an alert and provides relevant data through an application or web interface.
[0218] Specifically, consider a scenario where children are playing in a park after school. The device continuously films the surrounding area of the park, collecting data through its camera and microphone. The artificial intelligence within the device attempts to detect suspicious individuals from this data. If an unfamiliar person approaches the children, it recognizes their face and analyzes what they say. If danger is detected, details are reported to the server, and an immediate alert is sent to the user. In this case, the recorded data is stored as evidence in case of any unforeseen circumstances.
[0219] This system allows parents to monitor their children's safety from home, reducing unnecessary worries. Furthermore, the collected data serves as a valuable source of information for improving safety. Thus, this invention provides a new way to ensure children's safety through the power of technology.
[0220] The following describes the processing flow.
[0221] Step 1:
[0222] The device connects to the server upon startup and performs device authentication. If authentication is successful, the server sends various configuration information to the device and starts monitoring mode.
[0223] Step 2:
[0224] The device uses its built-in camera and microphone to capture surrounding video and audio in real time. This data is temporarily stored in the device's internal memory.
[0225] Step 3:
[0226] The artificial intelligence installed in the device analyzes the acquired video data and applies a facial recognition algorithm to identify human faces. Simultaneously, audio data is analyzed using natural language processing technology to detect specific keywords and unusual conversational content.
[0227] Step 4:
[0228] The device will immediately emit a warning sound based on the information it receives if it detects a suspicious person or dangerous situation. This audio warning is intended to alert those in the surrounding area.
[0229] Step 5:
[0230] The device sends data to the server when it detects important events or suspicious situations. The server receives this data and sends an alert to the user (parent / guardian) via an application or email.
[0231] Step 6:
[0232] The server securely stores audio and video data received from the terminal and makes it accessible to the user. Users can view this data through their logged-in account.
[0233] Step 7:
[0234] Users can use the collected data to understand the safety status of their children and prepare to provide the data to legal authorities or other relevant parties as needed.
[0235] (Example 1)
[0236] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0237] There is a need for a means to monitor dangerous situations that children may encounter in their daily lives in real time and to notify parents quickly and effectively. However, conventional monitoring devices have problems such as insufficient accuracy in detecting anomalies and often delays in information transmission. To address this situation and ensure the safety of children, there is a need to provide more advanced and efficient monitoring systems.
[0238] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0239] In this invention, the server includes sensor means for acquiring video, sensor means for acquiring audio, and information processing means for analyzing the acquired video and audio data in real time. This makes it possible to immediately detect suspicious situations and quickly transmit information to parents.
[0240] A "sensor" is a device that detects video and audio information and acquires it as a digital signal.
[0241] "Information processing means" refers to a computer system that analyzes acquired video and audio data to detect specific patterns or anomalies.
[0242] "Output means" refers to a device that issues a visual or audible warning to the user in order to communicate detected abnormal information.
[0243] "Information and communication means" refers to network devices used to notify parents in remote locations in real time of acquired and analyzed data.
[0244] A "data storage device" is a recording device that securely stores acquired and analyzed information for later review and use.
[0245] A "generative AI model" is a machine learning algorithm trained to analyze diverse data and identify suspicious situations.
[0246] In this embodiment of the invention, the system is composed of three main elements: a terminal, a server, and a user, through their interaction.
[0247] Device functions and configuration
[0248] The device is designed as a wearable device and is equipped with sensors to acquire video and audio. Specifically, it has a built-in high-performance camera and microphone that capture information about the surrounding environment. The acquired data is analyzed in real time by the device's information processing system. This analysis uses generative AI models, including facial recognition and speech recognition.
[0249] Server Functions
[0250] The server has the means of information and communication to receive analysis data sent from the terminal. Based on the received data, the server can perform more complex analysis and data reintegration, thereby providing the user with an accurate situational analysis.
[0251] User operation and interface
[0252] Users communicate with their devices and servers through a dedicated mobile application or web interface. This allows users to receive real-time alerts from the system and take necessary actions quickly.
[0253] Specific example
[0254] For example, imagine a scenario where the user's child is playing in a park. The device monitors the surrounding sounds and instantly detects the appearance of suspicious individuals or suspicious remarks. This triggers an alert to the user via the server. The user can then view live video through the application and, if deemed necessary, report the incident to the police or other authorities.
[0255] Example of a prompt
[0256] An example of a prompt for a generated AI model would be, "Explain step-by-step how the AI system identifies a suspicious person and sends an alert to the user."
[0257] This system can provide a new means of technically ensuring children's safety in their daily lives and giving parents peace of mind.
[0258] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0259] Step 1:
[0260] The device acquires surrounding video and audio using sensory means. The input is the video and audio of the environment, which is converted into digital data. This is done by the camera capturing video frames in real time and the microphone recording audio. The output is the digital video and audio data, which is fed into subsequent processing steps.
[0261] Step 2:
[0262] The terminal transmits the acquired video and audio data to an information processing system for initial processing. The input is the digital data generated in step 1, and the output is clean data that has undergone noise reduction and basic format conversion. This processing involves the application of a filtering algorithm, which plays a role in improving the quality of the data.
[0263] Step 3:
[0264] The terminal's information processing means applies a generated AI model using clean data to analyze the data. The input is the clean data obtained in step 2, and the output is behavior recognition data and speech recognition data as analysis results. Through this analysis, specific actions are taken to identify suspicious facial features and suspicious statements of suspicious individuals using pattern matching.
[0265] Step 4:
[0266] Based on the analysis results, the terminal determines whether to issue a warning through the output means. The input is the analysis results from step 3, and the output is the generation of a warning sound. In action, the audio output device sounds a pre-set warning sound to alert those in the surrounding area of danger.
[0267] Step 5:
[0268] The terminal uses communication methods to send analysis results and warning information to the server. The input is detailed data about the situation in which the warning was issued, and the output is the transmission of data to the remote server. This process includes encryption to ensure data security.
[0269] Step 6:
[0270] The server notifies the user based on the information it receives. Input is analysis and warning data sent from the terminal, and output is an alert notification sent to the user's device. The server sends the notification to the user through a dedicated application, including push notifications on mobile devices and email notifications.
[0271] (Application Example 1)
[0272] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0273] Conventional technologies often make it difficult to monitor children's safety in real time and respond quickly to abnormal situations. Furthermore, methods for monitoring a child's surroundings when a parent is away are often limited, resulting in insufficient safety even under normal circumstances. This makes it difficult to respond immediately when a child actually faces a dangerous situation, causing anxiety for parents.
[0274] The specific processing by the specific processing unit 290 of the data processing apparatus 12 in Application Example 1 is realized by the following means.
[0275] In this invention, the server includes an imaging means for acquiring surrounding images, a sound collection means for acquiring sounds, and an information processing means for analyzing the acquired data in real time. Thereby, it becomes possible to recognize an abnormal person approaching a child or an emergency situation and immediately send a warning to the guardian.
[0276] The "wearable monitoring device" is a monitoring device that can be worn on the body and used daily.
[0277] The "imaging means" is a device or module having a function for acquiring images.
[0278] The "sound collection means" is a device or module having a function for acquiring sounds.
[0279] The "information processing means" is a device or program having a function for analyzing the acquired data and extracting necessary information.
[0280] The "warning means" is a device having a function of issuing a warning through sound or digital notification when an abnormality is detected.
[0281] The "information communication means" is a device or module having a function for transmitting and receiving data to and from other terminals or systems.
[0282] The "biometric authentication algorithm" is a calculation method for identifying an individual using biometric information such as a face or fingerprint.
[0283] The "acoustic warning" is a means of issuing a warning by sound. [[ID=3z]]
[0284] The "electronic notification" is to notify the warning content to other terminals using digital means.
[0285] This wearable monitoring device system is designed to monitor the surrounding situation in real time in order to keep children safe.
[0286] The server uses information processing means to analyze the data transmitted from the terminal. Specifically, this information processing means uses software libraries such as TensorFlow and OpenCV to perform face recognition and natural language processing. The collected data is divided into image data and audio data. Specific patterns or people are extracted from the video data, and specific phrases or tones are analyzed from the audio data. This makes it possible to recognize abnormal persons or abnormal situations and issue an alarm immediately.
[0287] Users can receive alerts via a mobile information terminal such as a smartphone. The information communication means uses Wi-Fi or a mobile network to quickly transmit notifications from the server. For example, when a child is approached by a stranger in a park, the terminal captures the face and the AI system immediately notifies the user's smartphone of the analysis result. At this time, the collected data is stored as evidence if necessary.
[0288] The generated AI model can continue to learn using this data to achieve higher-precision recognition in the future. Also, as an example of a prompt sentence, it is possible to obtain a detailed explanation in the form of "Please explain the purpose and functions of the child monitoring app. Please describe it in detail with specific situation examples."
[0289] The flow of the specific processing in Application Example 1 will be described using FIG. 12.
[0290] Step 1:
[0291] The terminal activates an imaging means for capturing surrounding video and an audio acquisition means for acquiring sound. During this process, the imaging means acquires video data as input, and the audio acquisition means acquires audio data as input. These data are then transmitted to an information processing means as output.
[0292] Step 2:
[0293] The device analyzes the acquired video data using information processing tools and identifies a specific person using a face recognition algorithm. TensorFlow and OpenCV are used to filter the face information within the video. The input is video data, and the output is the recognized facial features.
[0294] Step 3:
[0295] The device analyzes the audio data using a natural language processing algorithm to understand its content. This process uses a speech recognition model to extract the spoken content as text data. The input is audio data, and the output is the analyzed text content.
[0296] Step 4:
[0297] The server integrates data extracted from video and audio to assess the level of risk. Based on the integrated data, it identifies suspicious individuals and detects abnormal situations. The input consists of facial feature data and text data, and the output is the risk assessment result.
[0298] Step 5:
[0299] The server activates alarm mechanisms based on the risk assessment results. If an anomaly is detected, it immediately uses the terminal's alarm mechanism to emit an audible alarm and also sends an electronic notification to the user via information and communication means. The input is the risk assessment result, and the output is the activation of the audible alarm and the transmission of the electronic notification.
[0300] Step 6:
[0301] The user receives a notification on their mobile device and checks its contents. The user's device immediately displays details of the threat using the received notification data. The input is electronic notification data, and the output is the information displayed on the device.
[0302] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0303] The system of the present invention provides even more advanced monitoring capabilities by implementing an emotion engine that recognizes the user's emotions, in addition to conventional wearable monitoring devices. The device is small and portable, in the form of a necklace or scarf, and can be worn naturally by children in their daily lives. The device incorporates a camera means for acquiring surrounding images, a microphone means for acquiring surrounding sounds, and artificial intelligence.
[0304] The emotion engine analyzes acquired voice data to determine the user's emotional state. This engine employs algorithms to identify emotions such as anger and anxiety by analyzing the tone and speaking patterns of the voice. This information is incorporated into the device's decision-making process, and if a specific emotional state is detected, the response through warning measures is adjusted. For example, if the user's voice is judged to be increasingly urgent, measures such as increasing the volume of the warning sound or shortening the interval between repetitions will be taken.
[0305] The device continuously captures surrounding video and audio and analyzes the data in real time. This analysis helps detect suspicious individuals and determine dangerous situations. The data analysis results, which reflect the user's emotions via an emotion engine, are sent to the server to more accurately assess the urgency of the situation. Based on this data, the server sends optimized alerts to the user (parent / guardian).
[0306] As a specific example, when a child approaches a stranger while out and the emotion engine determines that the tone of voice is uneasy, the information is immediately transmitted to the server. Based on this, the server sends a high-priority alert to the guardian, enabling prompt implementation of necessary measures.
[0307] In this way, the system equipped with the emotion engine provides a new method to more comprehensively support the safety of children and enable guardians to monitor with confidence even from a remote location.
[0308] The following explains the processing flow.
[0309] Step 1:
[0310] When the terminal is powered on, it connects to the server and performs an authentication procedure. If successful, the terminal receives the necessary configuration information from the server and switches to the monitoring mode.
[0311] Step 2:
[0312] The terminal utilizes the internal camera and microphone to obtain the surrounding video and audio in real time and temporarily stores the data in the built-in memory.
[0313] Step 3:
[0314] The artificial intelligence installed on the terminal executes a face recognition algorithm using the acquired video data to identify and distinguish human faces. On the other hand, natural language processing techniques are applied to the audio data to analyze the conversation content.
[0315] Step 4:
[0316] The emotion engine evaluates the tone and pattern of the user's voice from the analyzed audio data and determines the emotional state (e.g., anger, uneasiness, surprise, etc.).
[0317] Step 5:
[0318] The device will emit a warning sound if it detects a suspicious person, a dangerous situation, or if the user's emotions become urgent. In this case, the type and intensity of the warning sound will be adjusted based on the emotion engine's judgment.
[0319] Step 6:
[0320] The device sends information, including important events and emotional data, to the server. Based on the received data, the server assesses the urgency of the events and sends optimized alerts to the user (parent / guardian).
[0321] Step 7:
[0322] Users receive alerts from the server, check the situation via application or email, and take appropriate action. The collected data is stored for detailed situation analysis and necessary follow-up actions.
[0323] (Example 2)
[0324] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0325] In recent years, there has been a growing need for monitoring devices that are both miniaturized and portable. However, conventional technologies have struggled to accurately understand the user's emotional state and assess the urgency of a situation. As a result, there are problems with delays in providing appropriate warnings and alerts. In particular, in monitoring child safety, the lack of emotional recognition hinders a swift and appropriate response.
[0326] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0327] In this invention, the server includes optical means for acquiring surrounding video, acoustic means for acquiring sound, data processing means for analyzing acquired video and audio data in real time, emotion analysis means for analyzing the tone and pattern of sound to determine emotions, warning means for issuing a warning when a suspicious situation or a specific emotional state is detected, and communication means for sending alerts to guardians. This enables more precise monitoring that always takes the user's emotional state into consideration, and immediate alert issuance.
[0328] A "wearable monitoring device" is a monitoring device that a user can wear and carry with them at all times, and is usable in everyday life.
[0329] "Optical means" refers to imaging devices such as cameras used to acquire visual information of the surroundings.
[0330] "Acoustic means" refers to sound acquisition devices such as microphones used to acquire ambient sound information.
[0331] "Data processing means" refers to computer processing functions that analyze acquired video and audio data in real time and extract necessary information.
[0332] An "emotion analysis tool" is a system equipped with an algorithm that determines the user's emotional state by analyzing the tone and patterns of their voice.
[0333] A "warning device" is a device or mechanism that issues a warning to the user or their surroundings when a suspicious situation or a specific emotional state is detected.
[0334] "Communication means" refers to network communication functions used to transmit analysis results and warning information to the user's guardian or other supervisors.
[0335] "Suspicious circumstances" refer to unusual or potentially dangerous behavior or environments, meaning that the monitored behavior or state deviates from normal circumstances.
[0336] A "specific emotional state" is a mental state that requires immediate attention, such as anger or anxiety, which is determined by emotional analysis to be an emotional change that exceeds a predetermined standard.
[0337] The device is designed as a wearable monitoring device that users can wear and use on a daily basis. Specifically, the device is equipped with an optical means, a camera, and an acoustic means, a microphone. These devices are used to acquire surrounding video and audio data. The device also implements data processing means to analyze the acquired data in real time. This processing includes emotion analysis means that analyze the tone and patterns of voice to determine emotions.
[0338] The device is designed as a necklace or scarf that the user wears, making it lightweight, easy to carry, and comfortable for children to wear naturally. The emotion analysis system analyzes voice tone to identify the user's emotional state. If this process detects specific emotions such as anger or anxiety, the warning system will activate accordingly, providing necessary alerts.
[0339] The server receives data transmitted from the device via communication and sends alerts to the parent based on the analysis results. For example, if the device detects a child's anxious voice, that information is sent to the server, which can then send a high-priority alert to the parent's device. This allows the parent to immediately understand their child's situation and take necessary action.
[0340] As a concrete example, when a user is playing in a park, the device can capture ambient sounds and process them using emotion analysis to detect the child's anxiety when a suspicious person suddenly approaches. This information is transmitted to a server in real time, and a warning is sent to the parents.
[0341] Examples of prompts for a generative AI model:
[0342] "Please describe in detail how the device works when a child is making anxious noises."
[0343] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0344] Step 1:
[0345] The device acquires video data from its surroundings through optical means (camera) and audio data through acoustic means (microphone). The input is the surrounding video and audio, and the output is the video and audio data converted into their respective data formats. Specifically, the device constantly collects information through the camera and microphone in accordance with the user's movements.
[0346] Step 2:
[0347] The terminal analyzes the acquired video and audio data in real time using data processing means. The input is the video and audio data acquired in step 1, and the output is the analyzed data. This analysis uses emotion analysis means to identify emotions by analyzing changes in voice tone and speaking patterns. Specifically, the data processing algorithm within the terminal analyzes the voice tone and identifies emotions such as anger and anxiety.
[0348] Step 3:
[0349] The device activates a warning mechanism when a specific emotional state of the user (e.g., anxiety) is detected through emotion analysis. The input is the emotion information analyzed in step 2, and the output is the issuance of a warning. Specifically, the device emits an audio warning through its speaker.
[0350] Step 4:
[0351] The terminal transmits the analyzed data and warning information to the server via a communication method. The input is the analyzed data and warning information resulting from steps 2 and 3, and the output is a data packet containing this information. Specifically, the terminal transmits data to the server via wireless communication.
[0352] Step 5:
[0353] The server receives data sent from the device and processes it to send alerts to the parent. The input is the data packets received from the device, and the output is the alert message sent to the parent. Specifically, the server sends a push notification to the parent's communication device to report the situation.
[0354] (Application Example 2)
[0355] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0356] In modern home environments, threats from intruders and noises outside windows exist, making it especially important to ensure the safety of family members. Furthermore, in emergencies, it is necessary to respond quickly and accurately, taking into account emotional states. However, conventional security systems have struggled to adequately adjust warnings in real time to reflect the user's emotional state, making it difficult to fully guarantee the safety of the home.
[0357] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0358] In this invention, the server includes means equipped with an emotion engine for analyzing voice data to determine emotional states, image acquisition means for acquiring surrounding video, and voice acquisition means for acquiring audio. This enables rapid detection of abnormal emotional changes or emergencies within the home and the sending of optimized alerts to parents.
[0359] An "emotion engine" is a mechanism that includes an algorithm for analyzing voice data to determine the user's emotional state.
[0360] "Image acquisition means" refers to a device or technology for collecting surrounding video data.
[0361] "Voice acquisition means" refers to a device or technology for collecting ambient sound data.
[0362] "Artificial intelligence means" refers to a computer program or system that analyzes acquired video and audio data in real time and makes appropriate decisions.
[0363] A "warning mechanism" is a device that issues a warning by sound or other means when a specific situation occurs.
[0364] An "information transmission means" is a system equipped with communication functions for transmitting analysis results and warning information to other devices or users.
[0365] The system used to realize this application is configured as a smart security system that enhances safety for users within their homes and surrounding environments. The system acquires audio and video data and analyzes them using artificial intelligence. Specifically, for audio data analysis, the Google Cloud Speech-to-Text API is used to convert speech to text, and then IBM Watson Tone Analyzer is used to analyze the emotions contained in the speech.
[0366] Based on the acquired sentiment analysis results, the server optimizes alerts regarding the user's emotional state and sends alerts to the user's device using communication methods. In this process, the server uses an emotion engine to detect emotions such as anxiety and fear and adjusts the warning methods accordingly. If an emergency situation that causes the user to feel fear is detected, a real-time alert, including an audio warning, is transmitted to other members of the household.
[0367] For example, if a suspicious sound is heard at home at night, the system analyzes the tone of the sound and immediately issues a warning if fear is detected. This allows the whole family to respond quickly and take safety precautions more easily.
[0368] Examples of prompts to input into a generative AI model:
[0369] "If sounds indicating anxiety or fear are detected within a residence, please generate a statement explaining what kind of alert should be generated based on the emotion recognition results."
[0370] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0371] Step 1:
[0372] The terminal acquires audio and video data from within the home. Audio data is collected via a microphone, and video data is collected using an image acquisition device. This data is transmitted to the server in its original format.
[0373] Step 2:
[0374] The server converts received audio data into text data using the Google Cloud Speech-to-Text API. The input is audio data, and the output is the converted text data. This conversion makes it possible to recognize the content of the audio as text information.
[0375] Step 3:
[0376] The server inputs text data into IBM Watson Tone Analyzer for sentiment analysis. The input is the converted text data, and the output is the analysis result of the emotional state contained in that text (e.g., calm, anxious, fear). This analysis allows for a detailed understanding of the user's emotional state.
[0377] Step 4:
[0378] The server generates and adjusts alerts based on the emotion analysis results when specific emotions are detected. For example, if a strong emotion of fear is detected, the alert's urgency is increased by raising the volume of the warning device. The input is the emotion analysis result, and the output is the adjusted alert information.
[0379] Step 5:
[0380] The server transmits coordinated alert information to the user's terminal using an information transmission method. The input is the coordinated alert information, and the output is a warning message displayed or audibly output on the user's terminal. This allows the user to immediately understand the situation and take necessary actions quickly.
[0381] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0382] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0383] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0384] [Third Embodiment]
[0385] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0386] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0387] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0388] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0389] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0390] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0391] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0392] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0393] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0394] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0395] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0396] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0397] The wearable monitoring device of the present invention has a shape that is not uncomfortable for children to wear on a daily basis, and is configured as a scarf-type or necklace-type device. This device has a built-in camera to acquire surrounding video and a microphone to acquire audio. This acquired data is analyzed in real time by an artificial intelligence system built into the device, and dangerous situations are detected using a facial recognition algorithm and natural language processing technology.
[0398] If the device detects a suspicious person or dangerous situation, it will immediately emit a warning sound to alert those nearby. This warning sound is emitted through a speaker installed inside the device. Furthermore, analysis results and warning information are sent to a server via a communication module within the device. The server then uses this information to notify the user (parent / guardian) as an alert and provides relevant data through an application or web interface.
[0399] Specifically, consider a scenario where children are playing in a park after school. The device continuously films the surrounding area of the park, collecting data through its camera and microphone. The artificial intelligence within the device attempts to detect suspicious individuals from this data. If an unfamiliar person approaches the children, it recognizes their face and analyzes what they say. If danger is detected, details are reported to the server, and an immediate alert is sent to the user. In this case, the recorded data is stored as evidence in case of any unforeseen circumstances.
[0400] This system allows parents to monitor their children's safety from home, reducing unnecessary worries. Furthermore, the collected data serves as a valuable source of information for improving safety. Thus, this invention provides a new way to ensure children's safety through the power of technology.
[0401] The following describes the processing flow.
[0402] Step 1:
[0403] The device connects to the server upon startup and performs device authentication. If authentication is successful, the server sends various configuration information to the device and starts monitoring mode.
[0404] Step 2:
[0405] The device uses its built-in camera and microphone to capture surrounding video and audio in real time. This data is temporarily stored in the device's internal memory.
[0406] Step 3:
[0407] The artificial intelligence installed in the device analyzes the acquired video data and applies a facial recognition algorithm to identify human faces. Simultaneously, audio data is analyzed using natural language processing technology to detect specific keywords and unusual conversational content.
[0408] Step 4:
[0409] The device will immediately emit a warning sound based on the information it receives if it detects a suspicious person or dangerous situation. This audio warning is intended to alert those in the surrounding area.
[0410] Step 5:
[0411] The device sends data to the server when it detects important events or suspicious situations. The server receives this data and sends an alert to the user (parent / guardian) via an application or email.
[0412] Step 6:
[0413] The server securely stores audio and video data received from the terminal and makes it accessible to the user. Users can view this data through their logged-in account.
[0414] Step 7:
[0415] Users can use the collected data to understand the safety status of their children and prepare to provide the data to legal authorities or other relevant parties as needed.
[0416] (Example 1)
[0417] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0418] There is a need for a means to monitor dangerous situations that children may encounter in their daily lives in real time and to notify parents quickly and effectively. However, conventional monitoring devices have problems such as insufficient accuracy in detecting anomalies and often delays in information transmission. To address this situation and ensure the safety of children, there is a need to provide more advanced and efficient monitoring systems.
[0419] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0420] In this invention, the server includes sensor means for acquiring video, sensor means for acquiring audio, and information processing means for analyzing the acquired video and audio data in real time. This makes it possible to immediately detect suspicious situations and quickly transmit information to parents.
[0421] A "sensor" is a device that detects video and audio information and acquires it as a digital signal.
[0422] "Information processing means" refers to a computer system that analyzes acquired video and audio data to detect specific patterns or anomalies.
[0423] "Output means" refers to a device that issues a visual or audible warning to the user in order to communicate detected abnormal information.
[0424] "Information and communication means" refers to network devices used to notify parents in remote locations in real time of acquired and analyzed data.
[0425] A "data storage device" is a recording device that securely stores acquired and analyzed information for later review and use.
[0426] A "generative AI model" is a machine learning algorithm trained to analyze diverse data and identify suspicious situations.
[0427] In this embodiment of the invention, the system is composed of three main elements: a terminal, a server, and a user, through their interaction.
[0428] Device functions and configuration
[0429] The device is designed as a wearable device and is equipped with sensors to acquire video and audio. Specifically, it has a built-in high-performance camera and microphone that capture information about the surrounding environment. The acquired data is analyzed in real time by the device's information processing system. This analysis uses generative AI models, including facial recognition and speech recognition.
[0430] Server Functions
[0431] The server has the means of information and communication to receive analysis data sent from the terminal. Based on the received data, the server can perform more complex analysis and data reintegration, thereby providing the user with an accurate situational analysis.
[0432] User operation and interface
[0433] Users communicate with their devices and servers through a dedicated mobile application or web interface. This allows users to receive real-time alerts from the system and take necessary actions quickly.
[0434] Specific example
[0435] For example, imagine a scenario where the user's child is playing in a park. The device monitors the surrounding sounds and instantly detects the appearance of suspicious individuals or suspicious remarks. This triggers an alert to the user via the server. The user can then view live video through the application and, if deemed necessary, report the incident to the police or other authorities.
[0436] Example of a prompt
[0437] An example of a prompt for a generated AI model would be, "Explain step-by-step how the AI system identifies a suspicious person and sends an alert to the user."
[0438] This system can provide a new means of technically ensuring children's safety in their daily lives and giving parents peace of mind.
[0439] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0440] Step 1:
[0441] The device acquires surrounding video and audio using sensory means. The input is the video and audio of the environment, which is converted into digital data. This is done by the camera capturing video frames in real time and the microphone recording audio. The output is the digital video and audio data, which is fed into subsequent processing steps.
[0442] Step 2:
[0443] The terminal transmits the acquired video and audio data to an information processing system for initial processing. The input is the digital data generated in step 1, and the output is clean data that has undergone noise reduction and basic format conversion. This processing involves the application of a filtering algorithm, which plays a role in improving the quality of the data.
[0444] Step 3:
[0445] The terminal's information processing means applies a generated AI model using clean data to analyze the data. The input is the clean data obtained in step 2, and the output is behavior recognition data and speech recognition data as analysis results. Through this analysis, specific actions are taken to identify suspicious facial features and suspicious statements of suspicious individuals using pattern matching.
[0446] Step 4:
[0447] Based on the analysis results, the terminal determines whether to issue a warning through the output means. The input is the analysis results from step 3, and the output is the generation of a warning sound. In action, the audio output device sounds a pre-set warning sound to alert those in the surrounding area of danger.
[0448] Step 5:
[0449] The terminal uses communication methods to send analysis results and warning information to the server. The input is detailed data about the situation in which the warning was issued, and the output is the transmission of data to the remote server. This process includes encryption to ensure data security.
[0450] Step 6:
[0451] The server notifies the user based on the information it receives. Input is analysis and warning data sent from the terminal, and output is an alert notification sent to the user's device. The server sends the notification to the user through a dedicated application, including push notifications on mobile devices and email notifications.
[0452] (Application Example 1)
[0453] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0454] Conventional technologies often make it difficult to monitor children's safety in real time and respond quickly to abnormal situations. Furthermore, methods for monitoring a child's surroundings when a parent is away are often limited, resulting in insufficient safety even under normal circumstances. This makes it difficult to respond immediately when a child actually faces a dangerous situation, causing anxiety for parents.
[0455] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0456] In this invention, the server includes imaging means for acquiring surrounding images, sound acquisition means for acquiring sound, and information processing means for analyzing the acquired data in real time. This makes it possible to recognize suspicious individuals or emergencies approaching children and immediately send warnings to guardians.
[0457] A "wearable monitoring device" is a monitoring device that is designed to be worn on the body and used on a daily basis.
[0458] "Image acquisition means" refers to a device or module that has the function of acquiring images.
[0459] "Sound acquisition means" refers to a device or module that has the function of acquiring sound.
[0460] "Information processing means" refers to a device or program that has the function of analyzing acquired data and extracting necessary information.
[0461] An "alarm system" is a device that has the function of issuing a warning through voice or digital notification when an abnormality is detected.
[0462] "Information and communication means" refers to a device or module that has the function of sending and receiving data with other terminals or systems.
[0463] A "biometric authentication algorithm" is a computational method used to identify individuals using biometric information such as facial features or fingerprints.
[0464] An "acoustic alarm" is a method of issuing a warning using sound.
[0465] "Electronic notification" refers to the use of digital means to notify other devices of warning content.
[0466] This wearable monitoring system is designed to monitor the surrounding environment in real time to ensure the safety of children.
[0467] The server uses information processing tools to analyze data transmitted from terminals. Specifically, these tools utilize software libraries such as TensorFlow and OpenCV to perform facial recognition and natural language processing. The collected data is divided into image data and audio data. Specific patterns and individuals are extracted from the video data, and specific phrases and tones are analyzed from the audio data. This makes it possible to recognize abnormal individuals or situations and issue an immediate alarm.
[0468] Users can receive alerts via mobile devices such as smartphones. Information and communication methods utilize Wi-Fi or mobile networks to quickly transmit notifications from the server. For example, if a child is approached by a stranger in a park, the device captures the child's face, and the AI system immediately analyzes the data, notifying the user's smartphone of the results. The collected data is stored as evidence as needed.
[0469] The generative AI model can use this data to continue learning and achieve more accurate recognition in the future. Furthermore, it's possible to obtain detailed explanations in the format of prompts such as, "Please explain the purpose and functions of the child monitoring app. Please provide a detailed explanation including specific situational examples."
[0470] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0471] Step 1:
[0472] The terminal activates an imaging means for capturing surrounding video and an audio acquisition means for acquiring sound. During this process, the imaging means acquires video data as input, and the audio acquisition means acquires audio data as input. These data are then transmitted to an information processing means as output.
[0473] Step 2:
[0474] The device analyzes the acquired video data using information processing tools and identifies a specific person using a face recognition algorithm. TensorFlow and OpenCV are used to filter the face information within the video. The input is video data, and the output is the recognized facial features.
[0475] Step 3:
[0476] The device analyzes the audio data using a natural language processing algorithm to understand its content. This process uses a speech recognition model to extract the spoken content as text data. The input is audio data, and the output is the analyzed text content.
[0477] Step 4:
[0478] The server integrates data extracted from video and audio to assess the level of risk. Based on the integrated data, it identifies suspicious individuals and detects abnormal situations. The input consists of facial feature data and text data, and the output is the risk assessment result.
[0479] Step 5:
[0480] The server activates alarm mechanisms based on the risk assessment results. If an anomaly is detected, it immediately uses the terminal's alarm mechanism to emit an audible alarm and also sends an electronic notification to the user via information and communication means. The input is the risk assessment result, and the output is the activation of the audible alarm and the transmission of the electronic notification.
[0481] Step 6:
[0482] The user receives a notification on their mobile device and checks its contents. The user's device immediately displays details of the threat using the received notification data. The input is electronic notification data, and the output is the information displayed on the device.
[0483] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0484] The system of the present invention provides even more advanced monitoring capabilities by implementing an emotion engine that recognizes the user's emotions, in addition to conventional wearable monitoring devices. The device is small and portable, in the form of a necklace or scarf, and can be worn naturally by children in their daily lives. The device incorporates a camera means for acquiring surrounding images, a microphone means for acquiring surrounding sounds, and artificial intelligence.
[0485] The emotion engine analyzes acquired voice data to determine the user's emotional state. This engine employs algorithms to identify emotions such as anger and anxiety by analyzing the tone and speaking patterns of the voice. This information is incorporated into the device's decision-making process, and if a specific emotional state is detected, the response through warning measures is adjusted. For example, if the user's voice is judged to be increasingly urgent, measures such as increasing the volume of the warning sound or shortening the interval between repetitions will be taken.
[0486] The device continuously captures surrounding video and audio and analyzes the data in real time. This analysis helps detect suspicious individuals and determine dangerous situations. The data analysis results, which reflect the user's emotions via an emotion engine, are sent to the server to more accurately assess the urgency of the situation. Based on this data, the server sends optimized alerts to the user (parent / guardian).
[0487] For example, if a child is approached by a stranger while out and about, and the emotion engine detects that the child's voice sounds anxious, that information is immediately sent to the server. Based on this, the server sends a high-priority alert to the parent, enabling them to take necessary measures quickly.
[0488] Thus, this system, equipped with an emotion engine, provides a more multifaceted support for children's safety and offers a new way for parents to monitor their children with peace of mind, even from a distance.
[0489] The following describes the processing flow.
[0490] Step 1:
[0491] When the device is powered on, it connects to the server and performs an authentication procedure. If successful, the device receives the necessary configuration information from the server and switches to monitoring mode.
[0492] Step 2:
[0493] The device uses its internal camera and microphone to capture surrounding video and audio in real time, and temporarily stores that data in its internal memory.
[0494] Step 3:
[0495] The artificial intelligence installed in the device uses the acquired video data to execute a facial recognition algorithm to identify and distinguish human faces. Meanwhile, natural language processing techniques are applied to the audio data to analyze the content of the conversation.
[0496] Step 4:
[0497] The emotion engine evaluates the tone and patterns of the user's voice from the analyzed audio data to determine their emotional state (e.g., anger, anxiety, surprise, etc.).
[0498] Step 5:
[0499] The device will emit a warning sound if it detects a suspicious person, a dangerous situation, or if the user's emotions become urgent. In this case, the type and intensity of the warning sound will be adjusted based on the emotion engine's judgment.
[0500] Step 6:
[0501] The device sends information, including important events and emotional data, to the server. Based on the received data, the server assesses the urgency of the events and sends optimized alerts to the user (parent / guardian).
[0502] Step 7:
[0503] Users receive alerts from the server, check the situation via application or email, and take appropriate action. The collected data is stored for detailed situation analysis and necessary follow-up actions.
[0504] (Example 2)
[0505] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0506] In recent years, there has been a growing need for monitoring devices that are both miniaturized and portable. However, conventional technologies have struggled to accurately understand the user's emotional state and assess the urgency of a situation. As a result, there are problems with delays in providing appropriate warnings and alerts. In particular, in monitoring child safety, the lack of emotional recognition hinders a swift and appropriate response.
[0507] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0508] In this invention, the server includes optical means for acquiring surrounding video, acoustic means for acquiring sound, data processing means for analyzing acquired video and audio data in real time, emotion analysis means for analyzing the tone and pattern of sound to determine emotions, warning means for issuing a warning when a suspicious situation or a specific emotional state is detected, and communication means for sending alerts to guardians. This enables more precise monitoring that always takes the user's emotional state into consideration, and immediate alert issuance.
[0509] A "wearable monitoring device" is a monitoring device that a user can wear and carry with them at all times, and is usable in everyday life.
[0510] "Optical means" refers to imaging devices such as cameras used to acquire visual information of the surroundings.
[0511] "Acoustic means" refers to sound acquisition devices such as microphones used to acquire ambient sound information.
[0512] "Data processing means" refers to computer processing functions that analyze acquired video and audio data in real time and extract necessary information.
[0513] An "emotion analysis tool" is a system equipped with an algorithm that determines the user's emotional state by analyzing the tone and patterns of their voice.
[0514] A "warning device" is a device or mechanism that issues a warning to the user or their surroundings when a suspicious situation or a specific emotional state is detected.
[0515] "Communication means" refers to network communication functions used to transmit analysis results and warning information to the user's guardian or other supervisors.
[0516] "Suspicious circumstances" refer to unusual or potentially dangerous behavior or environments, meaning that the monitored behavior or state deviates from normal circumstances.
[0517] A "specific emotional state" is a mental state that requires immediate attention, such as anger or anxiety, which is determined by emotional analysis to be an emotional change that exceeds a predetermined standard.
[0518] The device is designed as a wearable monitoring device that users can wear and use on a daily basis. Specifically, the device is equipped with an optical means, a camera, and an acoustic means, a microphone. These devices are used to acquire surrounding video and audio data. The device also implements data processing means to analyze the acquired data in real time. This processing includes emotion analysis means that analyze the tone and patterns of voice to determine emotions.
[0519] The device is designed as a necklace or scarf that the user wears, making it lightweight, easy to carry, and comfortable for children to wear naturally. The emotion analysis system analyzes voice tone to identify the user's emotional state. If this process detects specific emotions such as anger or anxiety, the warning system will activate accordingly, providing necessary alerts.
[0520] The server receives data transmitted from the device via communication and sends alerts to the parent based on the analysis results. For example, if the device detects a child's anxious voice, that information is sent to the server, which can then send a high-priority alert to the parent's device. This allows the parent to immediately understand their child's situation and take necessary action.
[0521] As a concrete example, when a user is playing in a park, the device can capture ambient sounds and process them using emotion analysis to detect the child's anxiety when a suspicious person suddenly approaches. This information is transmitted to a server in real time, and a warning is sent to the parents.
[0522] Examples of prompts for a generative AI model:
[0523] "Please describe in detail how the device works when a child is making anxious noises."
[0524] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0525] Step 1:
[0526] The device acquires video data from its surroundings through optical means (camera) and audio data through acoustic means (microphone). The input is the surrounding video and audio, and the output is the video and audio data converted into their respective data formats. Specifically, the device constantly collects information through the camera and microphone in accordance with the user's movements.
[0527] Step 2:
[0528] The terminal analyzes the acquired video and audio data in real time using data processing means. The input is the video and audio data acquired in step 1, and the output is the analyzed data. This analysis uses emotion analysis means to identify emotions by analyzing changes in voice tone and speaking patterns. Specifically, the data processing algorithm within the terminal analyzes the voice tone and identifies emotions such as anger and anxiety.
[0529] Step 3:
[0530] The device activates a warning mechanism when a specific emotional state of the user (e.g., anxiety) is detected through emotion analysis. The input is the emotion information analyzed in step 2, and the output is the issuance of a warning. Specifically, the device emits an audio warning through its speaker.
[0531] Step 4:
[0532] The terminal transmits the analyzed data and warning information to the server via a communication method. The input is the analyzed data and warning information resulting from steps 2 and 3, and the output is a data packet containing this information. Specifically, the terminal transmits data to the server via wireless communication.
[0533] Step 5:
[0534] The server receives data sent from the device and processes it to send alerts to the parent. The input is the data packets received from the device, and the output is the alert message sent to the parent. Specifically, the server sends a push notification to the parent's communication device to report the situation.
[0535] (Application Example 2)
[0536] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0537] In modern home environments, threats from intruders and noises outside windows exist, making it especially important to ensure the safety of family members. Furthermore, in emergencies, it is necessary to respond quickly and accurately, taking into account emotional states. However, conventional security systems have struggled to adequately adjust warnings in real time to reflect the user's emotional state, making it difficult to fully guarantee the safety of the home.
[0538] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0539] In this invention, the server includes means equipped with an emotion engine for analyzing voice data to determine emotional states, image acquisition means for acquiring surrounding video, and voice acquisition means for acquiring audio. This enables rapid detection of abnormal emotional changes or emergencies within the home and the sending of optimized alerts to parents.
[0540] An "emotion engine" is a mechanism that includes an algorithm for analyzing voice data to determine the user's emotional state.
[0541] "Image acquisition means" refers to a device or technology for collecting surrounding video data.
[0542] "Voice acquisition means" refers to a device or technology for collecting ambient sound data.
[0543] "Artificial intelligence means" refers to a computer program or system that analyzes acquired video and audio data in real time and makes appropriate decisions.
[0544] A "warning mechanism" is a device that issues a warning by sound or other means when a specific situation occurs.
[0545] An "information transmission means" is a system equipped with communication functions for transmitting analysis results and warning information to other devices or users.
[0546] The system used to realize this application is configured as a smart security system that enhances safety for users within their homes and surrounding environments. The system acquires audio and video data and analyzes them using artificial intelligence. Specifically, for audio data analysis, the Google Cloud Speech-to-Text API is used to convert speech to text, and then IBM Watson Tone Analyzer is used to analyze the emotions contained in the speech.
[0547] Based on the acquired sentiment analysis results, the server optimizes alerts regarding the user's emotional state and sends alerts to the user's device using communication methods. In this process, the server uses an emotion engine to detect emotions such as anxiety and fear and adjusts the warning methods accordingly. If an emergency situation that causes the user to feel fear is detected, a real-time alert, including an audio warning, is transmitted to other members of the household.
[0548] For example, if a suspicious sound is heard at home at night, the system analyzes the tone of the sound and immediately issues a warning if fear is detected. This allows the whole family to respond quickly and take safety precautions more easily.
[0549] Examples of prompts to input into a generative AI model:
[0550] "If sounds indicating anxiety or fear are detected within a residence, please generate a statement explaining what kind of alert should be generated based on the emotion recognition results."
[0551] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0552] Step 1:
[0553] The terminal acquires audio and video data from within the home. Audio data is collected via a microphone, and video data is collected using an image acquisition device. This data is transmitted to the server in its original format.
[0554] Step 2:
[0555] The server converts received audio data into text data using the Google Cloud Speech-to-Text API. The input is audio data, and the output is the converted text data. This conversion makes it possible to recognize the content of the audio as text information.
[0556] Step 3:
[0557] The server inputs text data into IBM Watson Tone Analyzer for sentiment analysis. The input is the converted text data, and the output is the analysis result of the emotional state contained in that text (e.g., calm, anxious, fear). This analysis allows for a detailed understanding of the user's emotional state.
[0558] Step 4:
[0559] The server generates and adjusts alerts based on the emotion analysis results when specific emotions are detected. For example, if a strong emotion of fear is detected, the alert's urgency is increased by raising the volume of the warning device. The input is the emotion analysis result, and the output is the adjusted alert information.
[0560] Step 5:
[0561] The server transmits coordinated alert information to the user's terminal using an information transmission method. The input is the coordinated alert information, and the output is a warning message displayed or audibly output on the user's terminal. This allows the user to immediately understand the situation and take necessary actions quickly.
[0562] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0563] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0564] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0565] [Fourth Embodiment]
[0566] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0567] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0568] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0569] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0570] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0571] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0572] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0573] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0574] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0575] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0576] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0577] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0578] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0579] The wearable monitoring device of the present invention has a shape that is not uncomfortable for children to wear on a daily basis, and is configured as a scarf-type or necklace-type device. This device has a built-in camera to acquire surrounding video and a microphone to acquire audio. This acquired data is analyzed in real time by an artificial intelligence system built into the device, and dangerous situations are detected using a facial recognition algorithm and natural language processing technology.
[0580] If the device detects a suspicious person or dangerous situation, it will immediately emit a warning sound to alert those nearby. This warning sound is emitted through a speaker installed inside the device. Furthermore, analysis results and warning information are sent to a server via a communication module within the device. The server then uses this information to notify the user (parent / guardian) as an alert and provides relevant data through an application or web interface.
[0581] Specifically, consider a scenario where children are playing in a park after school. The device continuously films the surrounding area of the park, collecting data through its camera and microphone. The artificial intelligence within the device attempts to detect suspicious individuals from this data. If an unfamiliar person approaches the children, it recognizes their face and analyzes what they say. If danger is detected, details are reported to the server, and an immediate alert is sent to the user. In this case, the recorded data is stored as evidence in case of any unforeseen circumstances.
[0582] This system allows parents to monitor their children's safety from home, reducing unnecessary worries. Furthermore, the collected data serves as a valuable source of information for improving safety. Thus, this invention provides a new way to ensure children's safety through the power of technology.
[0583] The following describes the processing flow.
[0584] Step 1:
[0585] The device connects to the server upon startup and performs device authentication. If authentication is successful, the server sends various configuration information to the device and starts monitoring mode.
[0586] Step 2:
[0587] The device uses its built-in camera and microphone to capture surrounding video and audio in real time. This data is temporarily stored in the device's internal memory.
[0588] Step 3:
[0589] The artificial intelligence installed in the device analyzes the acquired video data and applies a facial recognition algorithm to identify human faces. Simultaneously, audio data is analyzed using natural language processing technology to detect specific keywords and unusual conversational content.
[0590] Step 4:
[0591] The device will immediately emit a warning sound based on the information it receives if it detects a suspicious person or dangerous situation. This audio warning is intended to alert those in the surrounding area.
[0592] Step 5:
[0593] The device sends data to the server when it detects important events or suspicious situations. The server receives this data and sends an alert to the user (parent / guardian) via an application or email.
[0594] Step 6:
[0595] The server securely stores audio and video data received from the terminal and makes it accessible to the user. Users can view this data through their logged-in account.
[0596] Step 7:
[0597] Users can use the collected data to understand the safety status of their children and prepare to provide the data to legal authorities or other relevant parties as needed.
[0598] (Example 1)
[0599] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0600] There is a need for a means to monitor dangerous situations that children may encounter in their daily lives in real time and to notify parents quickly and effectively. However, conventional monitoring devices have problems such as insufficient accuracy in detecting anomalies and often delays in information transmission. To address this situation and ensure the safety of children, there is a need to provide more advanced and efficient monitoring systems.
[0601] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0602] In this invention, the server includes sensor means for acquiring video, sensor means for acquiring audio, and information processing means for analyzing the acquired video and audio data in real time. This makes it possible to immediately detect suspicious situations and quickly transmit information to parents.
[0603] A "sensor" is a device that detects video and audio information and acquires it as a digital signal.
[0604] "Information processing means" refers to a computer system that analyzes acquired video and audio data to detect specific patterns or anomalies.
[0605] "Output means" refers to a device that issues a visual or audible warning to the user in order to communicate detected abnormal information.
[0606] "Information and communication means" refers to network devices used to notify parents in remote locations in real time of acquired and analyzed data.
[0607] A "data storage device" is a recording device that securely stores acquired and analyzed information for later review and use.
[0608] A "generative AI model" is a machine learning algorithm trained to analyze diverse data and identify suspicious situations.
[0609] In this embodiment of the invention, the system is composed of three main elements: a terminal, a server, and a user, through their interaction.
[0610] Device functions and configuration
[0611] The device is designed as a wearable device and is equipped with sensors to acquire video and audio. Specifically, it has a built-in high-performance camera and microphone that capture information about the surrounding environment. The acquired data is analyzed in real time by the device's information processing system. This analysis uses generative AI models, including facial recognition and speech recognition.
[0612] Server Functions
[0613] The server has the means of information and communication to receive analysis data sent from the terminal. Based on the received data, the server can perform more complex analysis and data reintegration, thereby providing the user with an accurate situational analysis.
[0614] User operation and interface
[0615] Users communicate with their devices and servers through a dedicated mobile application or web interface. This allows users to receive real-time alerts from the system and take necessary actions quickly.
[0616] Specific example
[0617] For example, imagine a scenario where the user's child is playing in a park. The device monitors the surrounding sounds and instantly detects the appearance of suspicious individuals or suspicious remarks. This triggers an alert to the user via the server. The user can then view live video through the application and, if deemed necessary, report the incident to the police or other authorities.
[0618] Example of a prompt
[0619] An example of a prompt for a generated AI model would be, "Explain step-by-step how the AI system identifies a suspicious person and sends an alert to the user."
[0620] This system can provide a new means of technically ensuring children's safety in their daily lives and giving parents peace of mind.
[0621] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0622] Step 1:
[0623] The device acquires surrounding video and audio using sensory means. The input is the video and audio of the environment, which is converted into digital data. This is done by the camera capturing video frames in real time and the microphone recording audio. The output is the digital video and audio data, which is fed into subsequent processing steps.
[0624] Step 2:
[0625] The terminal transmits the acquired video and audio data to an information processing system for initial processing. The input is the digital data generated in step 1, and the output is clean data that has undergone noise reduction and basic format conversion. This processing involves the application of a filtering algorithm, which plays a role in improving the quality of the data.
[0626] Step 3:
[0627] The terminal's information processing means applies a generated AI model using clean data to analyze the data. The input is the clean data obtained in step 2, and the output is behavior recognition data and speech recognition data as analysis results. Through this analysis, specific actions are taken to identify suspicious facial features and suspicious statements of suspicious individuals using pattern matching.
[0628] Step 4:
[0629] Based on the analysis results, the terminal determines whether to issue a warning through the output means. The input is the analysis results from step 3, and the output is the generation of a warning sound. In action, the audio output device sounds a pre-set warning sound to alert those in the surrounding area of danger.
[0630] Step 5:
[0631] The terminal uses communication methods to send analysis results and warning information to the server. The input is detailed data about the situation in which the warning was issued, and the output is the transmission of data to the remote server. This process includes encryption to ensure data security.
[0632] Step 6:
[0633] The server notifies the user based on the information it receives. Input is analysis and warning data sent from the terminal, and output is an alert notification sent to the user's device. The server sends the notification to the user through a dedicated application, including push notifications on mobile devices and email notifications.
[0634] (Application Example 1)
[0635] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0636] Conventional technologies often make it difficult to monitor children's safety in real time and respond quickly to abnormal situations. Furthermore, methods for monitoring a child's surroundings when a parent is away are often limited, resulting in insufficient safety even under normal circumstances. This makes it difficult to respond immediately when a child actually faces a dangerous situation, causing anxiety for parents.
[0637] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0638] In this invention, the server includes imaging means for acquiring surrounding images, sound acquisition means for acquiring sound, and information processing means for analyzing the acquired data in real time. This makes it possible to recognize suspicious individuals or emergencies approaching children and immediately send warnings to guardians.
[0639] A "wearable monitoring device" is a monitoring device that is designed to be worn on the body and used on a daily basis.
[0640] "Image acquisition means" refers to a device or module that has the function of acquiring images.
[0641] "Sound acquisition means" refers to a device or module that has the function of acquiring sound.
[0642] "Information processing means" refers to a device or program that has the function of analyzing acquired data and extracting necessary information.
[0643] An "alarm system" is a device that has the function of issuing a warning through voice or digital notification when an abnormality is detected.
[0644] "Information and communication means" refers to a device or module that has the function of sending and receiving data with other terminals or systems.
[0645] A "biometric authentication algorithm" is a computational method used to identify individuals using biometric information such as facial features or fingerprints.
[0646] An "acoustic alarm" is a method of issuing a warning using sound.
[0647] "Electronic notification" refers to the use of digital means to notify other devices of warning content.
[0648] This wearable monitoring system is designed to monitor the surrounding environment in real time to ensure the safety of children.
[0649] The server uses information processing tools to analyze data transmitted from terminals. Specifically, these tools utilize software libraries such as TensorFlow and OpenCV to perform facial recognition and natural language processing. The collected data is divided into image data and audio data. Specific patterns and individuals are extracted from the video data, and specific phrases and tones are analyzed from the audio data. This makes it possible to recognize abnormal individuals or situations and issue an immediate alarm.
[0650] Users can receive alerts via mobile devices such as smartphones. Information and communication methods utilize Wi-Fi or mobile networks to quickly transmit notifications from the server. For example, if a child is approached by a stranger in a park, the device captures the child's face, and the AI system immediately analyzes the data, notifying the user's smartphone of the results. The collected data is stored as evidence as needed.
[0651] The generative AI model can use this data to continue learning and achieve more accurate recognition in the future. Furthermore, it's possible to obtain detailed explanations in the format of prompts such as, "Please explain the purpose and functions of the child monitoring app. Please provide a detailed explanation including specific situational examples."
[0652] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0653] Step 1:
[0654] The terminal activates an imaging means for capturing surrounding video and an audio acquisition means for acquiring sound. During this process, the imaging means acquires video data as input, and the audio acquisition means acquires audio data as input. These data are then transmitted to an information processing means as output.
[0655] Step 2:
[0656] The device analyzes the acquired video data using information processing tools and identifies a specific person using a face recognition algorithm. TensorFlow and OpenCV are used to filter the face information within the video. The input is video data, and the output is the recognized facial features.
[0657] Step 3:
[0658] The device analyzes the audio data using a natural language processing algorithm to understand its content. This process uses a speech recognition model to extract the spoken content as text data. The input is audio data, and the output is the analyzed text content.
[0659] Step 4:
[0660] The server integrates data extracted from video and audio to assess the level of risk. Based on the integrated data, it identifies suspicious individuals and detects abnormal situations. The input consists of facial feature data and text data, and the output is the risk assessment result.
[0661] Step 5:
[0662] The server activates alarm mechanisms based on the risk assessment results. If an anomaly is detected, it immediately uses the terminal's alarm mechanism to emit an audible alarm and also sends an electronic notification to the user via information and communication means. The input is the risk assessment result, and the output is the activation of the audible alarm and the transmission of the electronic notification.
[0663] Step 6:
[0664] The user receives a notification on their mobile device and checks its contents. The user's device immediately displays details of the threat using the received notification data. The input is electronic notification data, and the output is the information displayed on the device.
[0665] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0666] The system of the present invention provides even more advanced monitoring capabilities by implementing an emotion engine that recognizes the user's emotions, in addition to conventional wearable monitoring devices. The device is small and portable, in the form of a necklace or scarf, and can be worn naturally by children in their daily lives. The device incorporates a camera means for acquiring surrounding images, a microphone means for acquiring surrounding sounds, and artificial intelligence.
[0667] The emotion engine analyzes acquired voice data to determine the user's emotional state. This engine employs algorithms to identify emotions such as anger and anxiety by analyzing the tone and speaking patterns of the voice. This information is incorporated into the device's decision-making process, and if a specific emotional state is detected, the response through warning measures is adjusted. For example, if the user's voice is judged to be increasingly urgent, measures such as increasing the volume of the warning sound or shortening the interval between repetitions will be taken.
[0668] The device continuously captures surrounding video and audio and analyzes the data in real time. This analysis helps detect suspicious individuals and determine dangerous situations. The data analysis results, which reflect the user's emotions via an emotion engine, are sent to the server to more accurately assess the urgency of the situation. Based on this data, the server sends optimized alerts to the user (parent / guardian).
[0669] For example, if a child is approached by a stranger while out and about, and the emotion engine detects that the child's voice sounds anxious, that information is immediately sent to the server. Based on this, the server sends a high-priority alert to the parent, enabling them to take necessary measures quickly.
[0670] Thus, this system, equipped with an emotion engine, provides a more multifaceted support for children's safety and offers a new way for parents to monitor their children with peace of mind, even from a distance.
[0671] The following describes the processing flow.
[0672] Step 1:
[0673] When the device is powered on, it connects to the server and performs an authentication procedure. If successful, the device receives the necessary configuration information from the server and switches to monitoring mode.
[0674] Step 2:
[0675] The device uses its internal camera and microphone to capture surrounding video and audio in real time, and temporarily stores that data in its internal memory.
[0676] Step 3:
[0677] The artificial intelligence installed in the device uses the acquired video data to execute a facial recognition algorithm to identify and distinguish human faces. Meanwhile, natural language processing techniques are applied to the audio data to analyze the content of the conversation.
[0678] Step 4:
[0679] The emotion engine evaluates the tone and patterns of the user's voice from the analyzed audio data to determine their emotional state (e.g., anger, anxiety, surprise, etc.).
[0680] Step 5:
[0681] The device will emit a warning sound if it detects a suspicious person, a dangerous situation, or if the user's emotions become urgent. In this case, the type and intensity of the warning sound will be adjusted based on the emotion engine's judgment.
[0682] Step 6:
[0683] The device sends information, including important events and emotional data, to the server. Based on the received data, the server assesses the urgency of the events and sends optimized alerts to the user (parent / guardian).
[0684] Step 7:
[0685] Users receive alerts from the server, check the situation via application or email, and take appropriate action. The collected data is stored for detailed situation analysis and necessary follow-up actions.
[0686] (Example 2)
[0687] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0688] In recent years, there has been a growing need for monitoring devices that are both miniaturized and portable. However, conventional technologies have struggled to accurately understand the user's emotional state and assess the urgency of a situation. As a result, there are problems with delays in providing appropriate warnings and alerts. In particular, in monitoring child safety, the lack of emotional recognition hinders a swift and appropriate response.
[0689] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0690] In this invention, the server includes optical means for acquiring surrounding video, acoustic means for acquiring sound, data processing means for analyzing acquired video and audio data in real time, emotion analysis means for analyzing the tone and pattern of sound to determine emotions, warning means for issuing a warning when a suspicious situation or a specific emotional state is detected, and communication means for sending alerts to guardians. This enables more precise monitoring that always takes the user's emotional state into consideration, and immediate alert issuance.
[0691] A "wearable monitoring device" is a monitoring device that a user can wear and carry with them at all times, and is usable in everyday life.
[0692] "Optical means" refers to imaging devices such as cameras used to acquire visual information of the surroundings.
[0693] "Acoustic means" refers to sound acquisition devices such as microphones used to acquire ambient sound information.
[0694] "Data processing means" refers to computer processing functions that analyze acquired video and audio data in real time and extract necessary information.
[0695] An "emotion analysis tool" is a system equipped with an algorithm that determines the user's emotional state by analyzing the tone and patterns of their voice.
[0696] A "warning device" is a device or mechanism that issues a warning to the user or their surroundings when a suspicious situation or a specific emotional state is detected.
[0697] "Communication means" refers to network communication functions used to transmit analysis results and warning information to the user's guardian or other supervisors.
[0698] "Suspicious circumstances" refer to unusual or potentially dangerous behavior or environments, meaning that the monitored behavior or state deviates from normal circumstances.
[0699] A "specific emotional state" is a mental state that requires immediate attention, such as anger or anxiety, which is determined by emotional analysis to be an emotional change that exceeds a predetermined standard.
[0700] The device is designed as a wearable monitoring device that users can wear and use on a daily basis. Specifically, the device is equipped with an optical means, a camera, and an acoustic means, a microphone. These devices are used to acquire surrounding video and audio data. The device also implements data processing means to analyze the acquired data in real time. This processing includes emotion analysis means that analyze the tone and patterns of voice to determine emotions.
[0701] The device is designed as a necklace or scarf that the user wears, making it lightweight, easy to carry, and comfortable for children to wear naturally. The emotion analysis system analyzes voice tone to identify the user's emotional state. If this process detects specific emotions such as anger or anxiety, the warning system will activate accordingly, providing necessary alerts.
[0702] The server receives data transmitted from the device via communication and sends alerts to the parent based on the analysis results. For example, if the device detects a child's anxious voice, that information is sent to the server, which can then send a high-priority alert to the parent's device. This allows the parent to immediately understand their child's situation and take necessary action.
[0703] As a concrete example, when a user is playing in a park, the device can capture ambient sounds and process them using emotion analysis to detect the child's anxiety when a suspicious person suddenly approaches. This information is transmitted to a server in real time, and a warning is sent to the parents.
[0704] Examples of prompts for a generative AI model:
[0705] "Please describe in detail how the device works when a child is making anxious noises."
[0706] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0707] Step 1:
[0708] The device acquires video data from its surroundings through optical means (camera) and audio data through acoustic means (microphone). The input is the surrounding video and audio, and the output is the video and audio data converted into their respective data formats. Specifically, the device constantly collects information through the camera and microphone in accordance with the user's movements.
[0709] Step 2:
[0710] The terminal analyzes the acquired video and audio data in real time using data processing means. The input is the video and audio data acquired in step 1, and the output is the analyzed data. This analysis uses emotion analysis means to identify emotions by analyzing changes in voice tone and speaking patterns. Specifically, the data processing algorithm within the terminal analyzes the voice tone and identifies emotions such as anger and anxiety.
[0711] Step 3:
[0712] The device activates a warning mechanism when a specific emotional state of the user (e.g., anxiety) is detected through emotion analysis. The input is the emotion information analyzed in step 2, and the output is the issuance of a warning. Specifically, the device emits an audio warning through its speaker.
[0713] Step 4:
[0714] The terminal transmits the analyzed data and warning information to the server via a communication method. The input is the analyzed data and warning information resulting from steps 2 and 3, and the output is a data packet containing this information. Specifically, the terminal transmits data to the server via wireless communication.
[0715] Step 5:
[0716] The server receives data sent from the device and processes it to send alerts to the parent. The input is the data packets received from the device, and the output is the alert message sent to the parent. Specifically, the server sends a push notification to the parent's communication device to report the situation.
[0717] (Application Example 2)
[0718] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0719] In modern home environments, threats from intruders and noises outside windows exist, making it especially important to ensure the safety of family members. Furthermore, in emergencies, it is necessary to respond quickly and accurately, taking into account emotional states. However, conventional security systems have struggled to adequately adjust warnings in real time to reflect the user's emotional state, making it difficult to fully guarantee the safety of the home.
[0720] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0721] In this invention, the server includes means equipped with an emotion engine for analyzing voice data to determine emotional states, image acquisition means for acquiring surrounding video, and voice acquisition means for acquiring audio. This enables rapid detection of abnormal emotional changes or emergencies within the home and the sending of optimized alerts to parents.
[0722] An "emotion engine" is a mechanism that includes an algorithm for analyzing voice data to determine the user's emotional state.
[0723] "Image acquisition means" refers to a device or technology for collecting surrounding video data.
[0724] "Voice acquisition means" refers to a device or technology for collecting ambient sound data.
[0725] "Artificial intelligence means" refers to a computer program or system that analyzes acquired video and audio data in real time and makes appropriate decisions.
[0726] A "warning mechanism" is a device that issues a warning by sound or other means when a specific situation occurs.
[0727] An "information transmission means" is a system equipped with communication functions for transmitting analysis results and warning information to other devices or users.
[0728] The system used to realize this application is configured as a smart security system that enhances safety for users within their homes and surrounding environments. The system acquires audio and video data and analyzes them using artificial intelligence. Specifically, for audio data analysis, the Google Cloud Speech-to-Text API is used to convert speech to text, and then IBM Watson Tone Analyzer is used to analyze the emotions contained in the speech.
[0729] Based on the acquired sentiment analysis results, the server optimizes alerts regarding the user's emotional state and sends alerts to the user's device using communication methods. In this process, the server uses an emotion engine to detect emotions such as anxiety and fear and adjusts the warning methods accordingly. If an emergency situation that causes the user to feel fear is detected, a real-time alert, including an audio warning, is transmitted to other members of the household.
[0730] For example, if a suspicious sound is heard at home at night, the system analyzes the tone of the sound and immediately issues a warning if fear is detected. This allows the whole family to respond quickly and take safety precautions more easily.
[0731] Examples of prompts to input into a generative AI model:
[0732] "If sounds indicating anxiety or fear are detected within a residence, please generate a statement explaining what kind of alert should be generated based on the emotion recognition results."
[0733] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0734] Step 1:
[0735] The terminal acquires audio and video data from within the home. Audio data is collected via a microphone, and video data is collected using an image acquisition device. This data is transmitted to the server in its original format.
[0736] Step 2:
[0737] The server converts received audio data into text data using the Google Cloud Speech-to-Text API. The input is audio data, and the output is the converted text data. This conversion makes it possible to recognize the content of the audio as text information.
[0738] Step 3:
[0739] The server inputs text data into IBM Watson Tone Analyzer for sentiment analysis. The input is the converted text data, and the output is the analysis result of the emotional state contained in that text (e.g., calm, anxious, fear). This analysis allows for a detailed understanding of the user's emotional state.
[0740] Step 4:
[0741] The server generates and adjusts alerts based on the emotion analysis results when specific emotions are detected. For example, if a strong emotion of fear is detected, the alert's urgency is increased by raising the volume of the warning device. The input is the emotion analysis result, and the output is the adjusted alert information.
[0742] Step 5:
[0743] The server transmits coordinated alert information to the user's terminal using an information transmission method. The input is the coordinated alert information, and the output is a warning message displayed or audibly output on the user's terminal. This allows the user to immediately understand the situation and take necessary actions quickly.
[0744] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0745] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0746] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0747] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0748] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0749] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0750] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0751] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0752] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0753] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0754] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0755] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0756] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0757] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0758] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0759] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0760] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0761] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0762] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0763] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0764] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0765] The following is further disclosed regarding the embodiments described above.
[0766] (Claim 1)
[0767] Wearable monitoring devices,
[0768] A camera system for acquiring images of the surroundings,
[0769] A microphone for acquiring sound,
[0770] An artificial intelligence means for analyzing acquired video and audio data in real time,
[0771] A warning system that issues a warning when suspicious circumstances are detected,
[0772] A means of communication to send alerts to parents,
[0773] A system that includes this.
[0774] (Claim 2)
[0775] The system according to claim 1, characterized in that the artificial intelligence means uses a face recognition algorithm to recognize the face of a suspicious person.
[0776] (Claim 3)
[0777] The system according to claim 1, characterized in that the warning means issues an audible warning when it detects a suspicious person or a dangerous situation.
[0778] "Example 1"
[0779] (Claim 1)
[0780] A sensor means for acquiring video,
[0781] A sensor means for acquiring sound,
[0782] Information processing means for analyzing acquired video and audio data in real time,
[0783] An output device that detects suspicious situations and issues warnings,
[0784] Information and communication means for notifying parents of the analysis results,
[0785] A data storage method for securely storing acquired data,
[0786] An information processing system that performs data analysis using a generative AI model,
[0787] A system that includes this.
[0788] (Claim 2)
[0789] The system according to claim 1, characterized in that the information processing means uses authentication technology for identifying a specific individual.
[0790] (Claim 3)
[0791] The system according to claim 1, characterized in that the output means issues an audible and visual warning when it detects a risk.
[0792] "Application Example 1"
[0793] (Claim 1)
[0794] Wearable monitoring devices,
[0795] An imaging means for acquiring images of the surroundings,
[0796] A means for acquiring sound,
[0797] Information processing means for analyzing acquired data in real time,
[0798] An alarm system that provides voice and digital notifications when an abnormal situation is detected,
[0799] A means of communication for sending alerts to the parent's mobile device,
[0800] A system that includes this.
[0801] (Claim 2)
[0802] The system according to claim 1, characterized in that the information processing means uses a biometric authentication algorithm for identifying abnormal persons.
[0803] (Claim 3)
[0804] The system according to claim 1, characterized in that the alarm means provides an audible alarm and an electronic notification when it detects an abnormal person or an emergency.
[0805] "Example 2 of combining an emotion engine"
[0806] (Claim 1)
[0807] Wearable monitoring devices,
[0808] Optical means for acquiring images of the surroundings,
[0809] Acoustic means for acquiring sound,
[0810] A data processing means for analyzing acquired video and audio data in real time,
[0811] An emotion analysis method that analyzes the tone and patterns of voice to determine emotions,
[0812] A warning system that issues a warning when suspicious circumstances or specific emotional states are detected,
[0813] A means of communication to send alerts to parents,
[0814] A system that includes this.
[0815] (Claim 2)
[0816] The system according to claim 1, characterized in that the data processing means uses an identification algorithm for recognizing the face of a suspicious person.
[0817] (Claim 3)
[0818] The system according to claim 1, characterized in that the warning means issues an audio warning when a suspicious person or a specific emotional state is detected.
[0819] "Application example 2 when combining with an emotional engine"
[0820] (Claim 1)
[0821] A means equipped with an emotion engine that analyzes voice data to determine emotional state,
[0822] An image acquisition means for obtaining surrounding video footage,
[0823] A means for acquiring sound,
[0824] An artificial intelligence means for analyzing acquired video and audio data in real time,
[0825] A warning mechanism that adjusts the warning when a specific emotional state is detected by the emotion engine,
[0826] A means of communication for sending alerts optimized for parents,
[0827] A system that includes this.
[0828] (Claim 2)
[0829] The system according to claim 1, characterized in that the artificial intelligence means uses a face recognition algorithm to recognize the face of a suspicious person.
[0830] (Claim 3)
[0831] The system according to claim 1, characterized in that the warning means issues an audible warning in response to a highly urgent state detected by the emotion engine. [Explanation of symbols]
[0832] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. Wearable monitoring devices, A camera system for acquiring images of the surroundings, A microphone for acquiring sound, An artificial intelligence means for analyzing acquired video and audio data in real time, A warning system that issues a warning when suspicious circumstances are detected, A means of communication to send alerts to parents, A system that includes this.
2. The system according to claim 1, characterized in that the artificial intelligence means uses a face recognition algorithm to recognize the face of a suspicious person.
3. The system according to claim 1, characterized in that the warning means issues an audible warning when it detects a suspicious person or a dangerous situation.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A