system

The system uses integrated data from cameras, acoustic devices, and sensors to enhance playground safety by detecting anomalies and sending alerts, addressing limitations in existing monitoring systems.

JP2026070966APending Publication Date: 2026-04-28SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-16
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing systems for monitoring children in playgrounds and parks lack effective and efficient methods to prevent accidents and ensure safety, as they are limited by human oversight and struggle with anomaly detection accuracy and response speed.

Method used

A system that integrates data from cameras, acoustic devices, and sensors to detect abnormal behavior and environmental conditions in real-time, sending alerts to notification devices for quick countermeasures, and suggests safe play areas.

Benefits of technology

Enhances safety by accurately detecting hazards and providing timely alerts and suggestions, ensuring children can play safely while easing restrictions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026070966000001_ABST
    Figure 2026070966000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means for processing video data acquired using a camera to detect abnormal behavior, A means for processing acoustic data acquired using an acoustic collection device and identifying a specific acoustic pattern, A means for detecting physical changes using a group of sensors and determining abnormal environmental conditions, A means for evaluating the degree of risk and determining countermeasures based on the information of the abnormal behavior and abnormal condition, A means of sending an alert to a notification device using the aforementioned countermeasure, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] It is important to create an environment where children can safely spend time in indoor and outdoor playgrounds and parks, but there are limitations to human monitoring by administrators and guardians. An object of the present invention is to realize a more effective and efficient monitoring and management system in order to prevent accidents and dangers and support the healthy growth of children.

Means for Solving the Problems

[0005] This invention provides a system that detects abnormal behavior and dangerous environmental conditions in real time by integrating and analyzing data from various sensors, cameras, and acoustic devices, and sends alerts to notification devices. This allows administrators and guardians to take quick countermeasures, thereby improving safety. Furthermore, by including a function to select suitable play areas and visually suggest them to the user, it provides a safe environment for children to play while easing restrictions on their play.

[0006] A "camera" is a device used to acquire moving images and plays a role in generating video data of the subject being monitored.

[0007] An "acoustic data acquisition device" is a device that captures sound waves and acquires acoustic data, and is used to identify specified sound patterns.

[0008] A "sensor group" is a group of devices composed of multiple sensors used to detect changes in the environment or physical conditions.

[0009] "Abnormal behavior" refers to actions or states that deviate from normal behavioral patterns, and includes actions that are judged to be risky.

[0010] An "abnormal condition" refers to a situation that deviates from the normal state of the environment, and includes conditions that may be dangerous.

[0011] "Risk level" represents the degree of risk inferred from abnormal behavior or abnormal conditions, and is an indicator of whether a rapid response is necessary.

[0012] "Countermeasures" refer to specific action guidelines and procedures that should be taken according to the detected level of risk.

[0013] A "notification device" is a device used to transmit alerts and information to users, and includes smartphones and dedicated information terminals.

[0014] "Alert" refers to information that is transmitted as a warning to the user when abnormal behavior or an abnormal state is detected.

[0015] "Play area" indicates a space where safe play is possible and refers to an area that is appropriate for children to engage in activities.

Brief Explanation of Drawings

[0016] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13]It is a sequence diagram showing the processing flow of the data processing system in Example 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.

Embodiments for Carrying Out the Invention

[0017] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.

[0020] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0021] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disk (e.g., hard disk), or magnetic tape, etc.

[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0024] [First Embodiment]

[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] As shown in Figure 2, in the data processing device 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0037] The system according to the present invention has the function of detecting dangers and abnormalities and issuing alerts by acquiring environmental data using a camera, an acoustic data collection device, and a group of sensors, and analyzing it in real time, in order to enhance safety in environments where children play.

[0038] Specifically, the server collects video data from multiple cameras installed in the child's location, and also aggregates environmental data from acoustic devices that capture ambient sound and a group of sensors that detect movement and temperature. This data is analyzed in real time and integrated on the server. The analysis uses pre-trained AI algorithms to quickly recognize abnormal behavior and abnormal environmental conditions.

[0039] A key feature of this system is that the server instantly assesses the level of risk based on detected abnormal behavior or conditions and automatically determines countermeasures. As a result of this process, the server sends alert information, including proposed countermeasures, to the terminal. This allows nearby parents and management staff to take swift action to protect children from danger.

[0040] For example, consider a scenario where too many children are gathered on one piece of playground equipment, creating a potential danger. In this case, the server analyzes video data acquired from the camera to detect the abnormal crowding and assesses the risk based on the increased risk of falls and other incidents. It then sends an alert to the terminal, including suggestions for appropriate limits on the number of children or suggestions to guide the children to a different, safer play area.

[0041] Furthermore, if weather changes affect the use of playground equipment, the system can consider sensor information that detects humidity and wind speed to issue appropriate warnings. For example, if playground equipment becomes slippery due to rain, it can send an alert in a timely manner stating that "caution is required when using the playground equipment."

[0042] As a result, parents and administrators, who are the users, can receive these alerts, follow the suggested actions, and take appropriate measures for the environment, thus ensuring that children can play safely.

[0043] The following describes the processing flow.

[0044] Step 1:

[0045] The server collects data in real time from the imaging device, acoustic data acquisition device, and sensor array. This includes video data, acoustic data, and environmental data such as temperature and humidity.

[0046] Step 2:

[0047] The server preprocesses the collected raw data. Specifically, it performs noise reduction, resolution adjustment, and audio data filtering to prepare the data for analysis.

[0048] Step 3:

[0049] The server analyzes pre-processed data using AI algorithms to detect abnormal behavior and conditions. For example, it uses video analysis to understand the density of children around playground equipment.

[0050] Step 4:

[0051] The server assesses the level of risk based on detected anomalies. Factors such as the frequency of the activity and the speed of change are considered in the assessment.

[0052] Step 5:

[0053] The server determines appropriate countermeasures based on the assessed risk level. Specifically, these include limiting the number of people for safety reasons and suggesting alternative play areas.

[0054] Step 6:

[0055] The server sends an alert to the device along with the decided course of action. This notifies parents and administrative staff.

[0056] Step 7:

[0057] The device displays received alerts to the user, using visual and auditory means to encourage prompt action.

[0058] Step 8:

[0059] Based on alert information from their devices, users take action to ensure the safety of children. Specifically, they follow instructions and take measures such as guiding children to safe play areas.

[0060] (Example 1)

[0061] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0062] In today's play environments for children, there is a need for safety monitoring systems to prevent accidents and injuries. However, existing systems have limitations in terms of anomaly detection accuracy and response speed, making it difficult to provide prompt and appropriate countermeasures. This invention aims to provide a safe play environment by detecting hazards in real time with high accuracy and proposing appropriate countermeasures.

[0063] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0064] In this invention, the server includes means for processing video data acquired using a camera and detecting abnormal behavior, means for processing acoustic data acquired using an acoustic collection device and identifying specific acoustic patterns, and means for detecting physical changes using a group of sensors and determining abnormal environmental conditions. This makes it possible to analyze environmental data in real time, quickly and accurately detect anomalies, and immediately derive countermeasures.

[0065] A "recording device" is a device used to acquire video data of children's activities and their surroundings.

[0066] An "acoustic acquisition device" is a device that captures ambient sounds and acquires data to identify specific acoustic patterns.

[0067] A "sensor group" is a group of devices composed of multiple sensors used to detect physical changes such as temperature, humidity, and movement.

[0068] "Abnormal behavior" refers to actions that deviate from normal behavioral patterns and may indicate danger.

[0069] An "abnormal environmental condition" refers to a state in which physical environmental factors such as temperature, humidity, and acoustics deviate significantly from normal values.

[0070] "Risk assessment" is the process of analyzing the risks of detected abnormal behaviors or abnormal environmental conditions and setting priorities according to their urgency.

[0071] "Countermeasures" refer to specific action plans or proposals implemented to mitigate or eliminate a particular risk.

[0072] A "notification device" is a device that receives alert information sent from a server and transmits it to the user.

[0073] A "generative AI model" is a model that implements artificial intelligence algorithms used to analyze acquired data.

[0074] The system of this invention aims to enhance safety in children's play environments by collecting and analyzing data using multiple devices to detect anomalies and promptly issue alerts.

[0075] The server collects video data using cameras installed at the child's location. These cameras have a wide field of view and high resolution, allowing them to capture the child's movements in real time. In addition, an acoustic sound capture device captures ambient sound. This acoustic data is used to quickly identify abnormal or warning sounds. Environmental data from sensors detecting temperature, humidity, and motion is also collected.

[0076] The server integrates this data and performs real-time analysis using a generative AI model. This analysis process utilizes AI algorithms to quickly recognize abnormal behavior and abnormal environmental conditions and assess the level of risk. This allows the server to immediately determine appropriate countermeasures. For example, if multiple children are crowded together on a single piece of playground equipment, the server assesses the risk of overcrowding and generates suggestions for guiding the children.

[0077] The server generates an alert based on the decided response and sends it to the device. The device notifies parents and administrative staff of the alert, quickly sharing important information. This allows nearby parents and administrative staff to take appropriate action quickly and ensure the safety of the child.

[0078] A concrete example would be a prompt message such as, "Monitor the concentration level of children and generate an alert if the density is high." Based on this prompt, the AI ​​model performs data analysis and generates an appropriate alert.

[0079] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0080] Step 1:

[0081] The server acquires video data from the camera in real time. This data includes children's movements and the usage of playground equipment. The input video data is extracted frame by frame and converted into a format that is easy to analyze. This standardizes the image resolution and format, making it suitable for the next processing step.

[0082] Step 2:

[0083] The server acquires acoustic data from the acoustic data acquisition device and identifies specific acoustic patterns. In this process, the input acoustic data is analyzed in terms of time and frequency to extract abnormal sounds and patterns. Specifically, noise filtering and Fourier transforms are used to highlight important features and prepare for the next analysis stage.

[0084] Step 3:

[0085] The server collects environmental data such as temperature, humidity, and motion from a group of sensors. This data is used as input to capture changes in the physical environment. The server normalizes the sensor data and processes it as multidimensional data points to prepare for the detection of abnormal conditions.

[0086] Step 4:

[0087] The server analyzes this integrated data using a generating AI model. It combines input visual, auditory, and environmental data to rapidly recognize abnormal behavior and environmental anomalies. The AI ​​algorithm applies a pre-trained model to classify anomalies and assess risk. The output of this process is a list of anomalous events and their severity ratings.

[0088] Step 5:

[0089] The server assesses the level of risk based on the analysis results and determines countermeasures. Based on the input list of anomalies, the server prioritizes the urgency and necessity of action. The output generates specific recommended actions and response plans.

[0090] Step 6:

[0091] The server generates an alert based on the decided response and sends it to the terminal. The terminal notifies the user of the received alert and prompts them to take the necessary action. Specifically, this could include messages such as "Please temporarily stop using the playground equipment" or "Please move the children to a safe place." This notification enables swift action to ensure the safety of the children.

[0092] (Application Example 1)

[0093] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0094] For autonomous vehicles to operate safely, they must monitor their surroundings in real time, quickly detect potential hazards, and take appropriate action. However, existing autonomous driving technologies are not sufficient to respond to specific environmental conditions or unexpected situations, and more sophisticated methods for ensuring safety are needed.

[0095] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0096] In this invention, the server includes means for processing video data acquired using a camera and detecting abnormal behavior; means for processing acoustic data acquired using an acoustic collection device and identifying specific acoustic patterns; means for detecting physical changes using a group of sensors and determining abnormal environmental conditions; and automatic control means for monitoring the surrounding environment and performing emergency stops or route changes to enhance the safety of vehicle operation. This enables autonomous vehicles to improve safety in real time.

[0097] A "filming device" is a device used to acquire video data and can record video information of objects or landscapes.

[0098] An "acoustic data acquisition device" is a device for acquiring acoustic data, which can sense ambient sound waves and record that information as data.

[0099] A "sensor cluster" is a collection of multiple sensors used to acquire various environmental information and to detect physical changes and environmental conditions.

[0100] "Abnormal behavior" refers to unexpected actions that deviate from normal behavioral patterns, indicating a potentially dangerous situation within a system.

[0101] "Risk level" is a measure that assesses the degree of risk in a particular situation, indicating how much attention and response a system requires.

[0102] "Countermeasures" refer to the actions to be taken in response to detected abnormalities or dangers, and are specific action plans to ensure safety.

[0103] A "notification device" is a device that receives alert information from a system and informs the user of its contents.

[0104] "Automatic control means" refers to a function that autonomously adjusts the operation of a vehicle in response to changes in the surrounding conditions, and can perform emergency stops and route changes.

[0105] The system implementing this invention aims to improve the safety of autonomous vehicles by analyzing environmental data acquired from various devices in real time. This system constantly monitors the surrounding environment using a camera, an acoustic data collection device, and a group of sensors.

[0106] The server processes video data acquired from the camera and detects abnormal behavior using AI algorithms. Specifically, it applies pre-trained models using machine learning libraries such as TENSORFLOW® to identify anomalies from the data. Similarly, it analyzes acoustic data from the acoustic collection device using AI models to identify specific acoustic patterns.

[0107] In addition, the server integrates data acquired from the sensor array, detects physical changes such as temperature, humidity, and vibration, and determines abnormal environmental conditions. This process rapidly analyzes input from the sensors and assesses potential hazards in real time.

[0108] If an anomaly is detected, the server adjusts the vehicle's operation via automatic control mechanisms. Specific safety measures include emergency stops and route changes. These control signals are transmitted directly to the vehicle's control system and executed immediately.

[0109] Users receive alerts via notification devices, along with suggestions for specific actions to take. This allows for appropriate responses even from outside the system.

[0110] As a concrete example, consider a scenario where a pedestrian suddenly appears in front of the vehicle while it is operating in the rain. In this case, cameras and sensors work together to understand the situation, AI assesses the danger in real time, and issues a deceleration command to the vehicle control system.

[0111] Examples of prompts for a generative AI model:

[0112] "When a pedestrian appears ahead, what kind of sensor data should be used, and how should it be analyzed, to safely slow down?"

[0113] In this way, the system enhances safety in the operation of autonomous vehicles and can respond quickly and appropriately to changes in the environment.

[0114] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0115] Step 1:

[0116] The server receives video data from the camera as input. The video data is analyzed in real time using an AI algorithm to detect abnormal behavior. Specifically, an image recognition model is used to compare the data with pre-trained abnormal patterns to determine whether an abnormality has occurred. As a result, information regarding the abnormal behavior is output.

[0117] Step 2:

[0118] The server receives acoustic data from an acoustic data collection device as input. This data is analyzed by an AI model, and specific acoustic patterns are identified. For example, speech recognition technology is used to detect sudden braking sounds or warning sounds. The detected acoustic patterns are then output and used to evaluate the safety of the system.

[0119] Step 3:

[0120] The server receives data from a group of sensors as input. Physical changes such as temperature, humidity, and vibration are evaluated in real time, and calculations are performed to determine abnormal environmental conditions. By integrating the sensor information and analyzing deviations from normal patterns, anomalies can be identified. As a result, anomaly information regarding the environmental condition is output.

[0121] Step 4:

[0122] The server assesses the vehicle's risk level based on the abnormal behavior, sound patterns, and environmental conditions obtained in steps 1-3. It calculates a risk score based on this information and determines the necessary countermeasures. This includes specific measures such as whether an emergency stop is necessary or whether the route should be changed. The risk assessment results and countermeasures are then output.

[0123] Step 5:

[0124] Based on the determined response, the server outputs control signals to the autonomous vehicle's control system to adjust its operation. These signals may include emergency stop orders or new route instructions. These control signals encourage a rapid response from the vehicle and ensure safety.

[0125] Step 6:

[0126] The device receives alert information sent from the server and notifies the user. The notification includes security-related information and action suggestions to help the user take prompt action. This allows the user to understand the situation and take appropriate action as needed.

[0127] An example of a prompt for the generated AI model is, "When a pedestrian appears ahead, what sensor data should be used and how should it be analyzed to safely slow down?"

[0128] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0129] The system according to the present invention is an advanced monitoring and management system that not only provides a safe and effective play environment but also recognizes the user's emotional state and combines corresponding functions. By combining an emotion engine with a basic system that acquires environmental data using a camera, an acoustic data collection device, and a group of sensors, the system analyzes the user's emotional state in real time.

[0130] Specifically, the server captures the user's facial expressions and voice tone through camera footage and audio data, and the emotion engine uses this information to recognize the user's emotions (joy, anger, fear, etc.). The emotion engine is designed to identify a wide range of emotions by utilizing AI-powered machine learning algorithms.

[0131] The server uses the collected emotional information to determine appropriate countermeasures necessary to ensure the safety of children in the play environment. In this process, emotional information is taken into account in addition to the usual risk assessment, and suggestions for rest to alleviate excessive stress or recommendations for appropriate activities to increase the variety of play are made.

[0132] For example, if a child shows signs of stress during play, the server suggests a relaxing activity and sends an alert to the device saying, "Let's take a break and read a picture book." If the child continues to play and enjoys themselves, the server may suggest, "Let's try a new athletic activity," to further enhance positive emotions.

[0133] By integrating emotional information with conventional environmental data, users such as parents and administrators can implement more appropriate and emotionally sensitive safety measures for children. This provides a safe and comfortable play environment, promoting the healthy development and emotional growth of children.

[0134] The following describes the processing flow.

[0135] Step 1:

[0136] The server acquires user facial expression data and voice data from the camera and sound acquisition devices. This collects raw data that forms the basis of emotions.

[0137] Step 2:

[0138] The server preprocesses the acquired data, performing noise reduction and conversion to the required format. This prepares the data for analysis.

[0139] Step 3:

[0140] The server inputs pre-processed data into the emotion engine to identify the user's emotional state. In this process, a trained AI model identifies emotions such as joy, sadness, and anger.

[0141] Step 4:

[0142] The server aggregates a series of emotional data based on the analysis results of the emotion engine and evaluates the user's stress level and emotional state.

[0143] Step 5:

[0144] The server integrates emotional data and environmental data (abnormal behavior, weather changes, etc.) to reassess the level of risk in the play environment. Based on this assessment, it determines the priority of countermeasures.

[0145] Step 6:

[0146] The server will formulate appropriate countermeasures based on the evaluation results. For example, if emotional data indicates a high-stress state, it will suggest "moving to a rest area."

[0147] Step 7:

[0148] The server sends alert information, including the formulated countermeasures, to the device and notifies parents or administrators.

[0149] Step 8:

[0150] The device displays received alerts to the user visually and audibly, prompting them to take specific action.

[0151] Step 9:

[0152] Based on alert information from the device, the user takes the most appropriate action according to the child's situation. Specifically, they follow the suggestions to guide the child into a safe and positive play environment.

[0153] (Example 2)

[0154] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0155] Existing safety management systems in play environments have struggled to adequately understand and respond to the psychological state of users. Therefore, it has been difficult to implement safety measures that directly address the stress and anxiety experienced by children in particular, and a flexible approach that responds to users' emotional needs is required.

[0156] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0157] In this invention, the server includes means for processing video information acquired using a camera and analyzing the user's emotional state, means for processing acoustic information acquired using an acoustic collection device and analyzing emotions from voice tone, and means for acquiring and analyzing physical environment data using a group of sensors. This enables immediate and appropriate safety measures and action suggestions in accordance with the user's psychological state.

[0158] A "filming device" is a device used to acquire video information and is responsible for capturing the user's movements and facial expressions.

[0159] An "acoustic data acquisition device" is a device used to acquire acoustic information, and its role is to record the user's voice and ambient sounds.

[0160] A "sensor group" is a collection of multiple sensor devices that acquire physical environmental data, such as temperature, motion, and location information.

[0161] "Emotional state" refers to information that indicates the user's psychological and emotional state, and includes psychological elements such as joy, anger, fear, and sadness.

[0162] "Psychological state" refers to the mental and emotional state a user is experiencing, and is a concept that includes various emotions such as stress, relief, and tension.

[0163] A "suggestion alert" is information that the server uses to notify the user of recommended actions, such as relaxing activities or challenging new activities.

[0164] "Environmental data" refers to information that describes the physical state of the environment, and includes data such as temperature, humidity, sound level, and brightness.

[0165] This invention provides a safe and effective play environment, as well as an advanced monitoring and management system for recognizing and responding to the emotional state of users. This system is realized by collecting and analyzing data in real time using a camera, an acoustic data collection device, and a group of sensors.

[0166] The server has the function of processing video information acquired from the camera. It analyzes the user's facial expressions from the video information and recognizes their emotional state in real time. This makes it possible to accurately determine the user's emotions, such as joy, anger, sadness, and happiness. The server also has the function of processing acoustic data acquired from the acoustic collection device, and analyzes the tone of voice to estimate the user's emotions. This uses an AI-based generative model and machine learning algorithms. Furthermore, the server analyzes physical environment data acquired from the sensor group and takes the surrounding environmental conditions into consideration, thereby more comprehensively evaluating the user's psychological state.

[0167] Based on the results of the emotion analysis performed by the server, appropriate countermeasures to meet the user's current emotional needs are quickly determined. This includes emotionally sensitive suggestions such as selecting a play environment and suggesting breaks to reduce stress. For example, if the user is experiencing excessive stress, the server may suggest relaxing activities such as "Let's take a break and read a picture book." If the user is actively engaged, it may suggest actions such as "Let's try a new athletic activity."

[0168] These suggestions are communicated to the user via the device. The device displays the suggestions intuitively through visual and auditory means, making them easy for the user to understand. This allows the user to choose an appropriate action based on their emotional state.

[0169] For example, by inputting a prompt such as, "Based on the user's emotional data, suggest relaxing activities," into the AI ​​model, the system can accurately provide suggestions tailored to the user's individual needs. This enables the provision of a safer and more comfortable play environment.

[0170] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0171] Step 1:

[0172] The server acquires data from the camera and sound acquisition equipment. It receives video data, including the user's facial expressions, from the camera and audio data from the microphone in real time. This provides input to understand the user's current state. Specifically, the camera periodically captures the user's face and sends this as digital video data to the server. The audio data captures the tone and volume of the user's voice and is sent for acoustic analysis.

[0173] Step 2:

[0174] The server analyzes the acquired video and audio data using a generating AI model. From the video data, a machine learning algorithm is used to recognize facial expressions and extract the user's emotional state. From the audio data, additional emotional insights are obtained through voice tone analysis. This data processing outputs results that show the user's overall emotional state. Specifically, the AI ​​model analyzes facial features and generates emotion labels such as anger and joy. From the audio data, it analyzes the pitch and tempo of the voice to derive indicators of stress and calmness.

[0175] Step 3:

[0176] The server takes the analyzed emotional data into account to generate suggestions for a play environment suitable for the user. The main inputs here are the results of the emotional analysis and known safety standards. The server integrates this information to determine appropriate countermeasures and activities. Specifically, if relaxation is determined to be needed, it will recommend activities in a quiet environment. Conversely, if high energy levels are determined, it will suggest vigorous athletic activities.

[0177] Step 4:

[0178] The device notifies the user of action suggestions received from the server. This uses both visual and auditory information to ensure intuitive understanding. The device presents prompts to the user in natural language to encourage action. This notification is output as feedback to the user and influences which activity they choose. For example, the device displays "Let's relax by reading a picture book" on its screen and also conveys the same message aloud.

[0179] Step 5:

[0180] The user receives suggestions from the device and selects an action that aligns with their emotional state. In this final step, the user's choices are returned to the system as feedback, which is used for further analysis and system improvement. Specifically, the user checks an alert and begins taking action according to the suggested activity. This allows the user to continue their activities in a safer and more comfortable environment.

[0181] (Application Example 2)

[0182] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0183] Traditional in-store shopping experiences often lack individuality because they don't adequately consider customers' emotional states. This can lead to decreased customer satisfaction and purchase intent. Therefore, it's crucial to improve the quality of the shopping experience by providing more personalized product and service suggestions based on customer emotions.

[0184] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0185] In this invention, the server includes means for processing video data acquired using a camera and detecting abnormal behavior, means for processing acoustic data acquired using an acoustic collection device and identifying specific acoustic patterns, and means for detecting physical changes using a group of sensors and determining abnormal environmental conditions. This enables real-time analysis of the user's emotional state and the suggestion of suitable products or services based on those emotions.

[0186] A "recording device" is a device used to acquire visual information, such as the conditions in the environment or the user's facial expressions.

[0187] An "acoustic acquisition device" is a device that acquires voice and ambient sounds, and is used to analyze the tone of the user's voice and the surrounding acoustic patterns.

[0188] A "sensor cluster" is a collection of various sensors installed to detect diverse physical changes and is used to determine abnormal environmental conditions.

[0189] "Real-time analysis of emotional state" means instantly processing data such as the user's facial expressions and voice to identify their emotions at that moment.

[0190] "Suggesting suitable products or services" refers to the act of selecting and presenting products or services that match the emotions of the analyzed user.

[0191] The system for realizing this invention is a mechanism that improves the in-store purchasing experience by analyzing the user's emotional state in real time and suggesting products and services based on those emotions.

[0192] The server uses a camera, sound acquisition device, and sensor array to collect data such as the user's facial expressions, voice tone, and surrounding environmental conditions. Specifically, the camera records the user's facial movements and expressions, while the microphone records the user's voice and background sounds. This data is integrated with environmental data collected through the sensor array and transmitted to the server.

[0193] On the server, the collected data is processed using a cloud-based sentiment analysis engine (e.g., Amazon Rekognition or Google® Cloud Vision API). This analyzes and identifies the emotions the user is experiencing. Based on this analysis, the server selects and visually presents products and services that are suitable for the user.

[0194] Furthermore, it is possible to utilize generative AI models to create prompt messages that respond to the user's emotions. These prompt messages would serve as information that encourages specific actions from the user, such as, "Identify the user's emotions from the facial expression analysis data and create a prompt to customize the shopping experience based on those emotions."

[0195] As a concrete example, suppose a user is browsing in a store and the server recognizes that the user has made a surprised expression upon seeing a particular product. In this case, the server will encourage a purchase by displaying special offers or new information related to that product on the user's smartphone or display.

[0196] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0197] Step 1:

[0198] The server collects user facial and audio data using a camera and audio acquisition device. Inputs are camera video and audio from a microphone. These inputs are acquired in real time, and initial processing is performed using facial recognition technology and audio analysis techniques. Outputs are the user's basic facial features and audio patterns.

[0199] Step 2:

[0200] The server collects environmental data from a group of sensors. The input consists of physical data such as temperature, humidity, and light intensity within the store. This sensor data is aggregated to evaluate the environmental conditions. The output is an evaluation value indicating abnormal environmental conditions.

[0201] Step 3:

[0202] The server integrates the collected facial expression data, acoustic data, and environmental data to perform emotion analysis. The input consists of the data obtained in Step 1 and Step 2. This data is integrated, and the emotion analysis engine is used to identify the user's emotions. The output is the user's emotional state (e.g., joy, surprise, anger).

[0203] Step 4:

[0204] The server selects suitable products and services based on the results of sentiment analysis. The input is information about the user's emotional state. It references a database to extract products and services that match the emotion. The output is information about the suggested products and services.

[0205] Step 5:

[0206] The terminal receives suggestion information from the server and presents it visually to the user. The input is suggestion information about products and services sent from the server. This information is displayed on a display or smartphone screen to attract the user's attention. The output is the visual information provided to the user.

[0207] Step 6:

[0208] The server utilizes a generative AI model to create emotion-based prompts. The input consists of the user's emotional state and suggested product information. Based on this data, the generative AI model generates prompts. The output is a prompt that encourages the user to take a specific action.

[0209] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0210] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0211] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0212] [Second Embodiment]

[0213] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0214] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0215] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0216] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0217] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0218] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0219] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0220] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0221] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0222] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0223] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0224] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0225] The system according to the present invention has the function of detecting dangers and abnormalities and issuing alerts by acquiring environmental data using a camera, an acoustic data collection device, and a group of sensors, and analyzing it in real time, in order to enhance safety in environments where children play.

[0226] Specifically, the server collects video data from multiple cameras installed in the child's location, and also aggregates environmental data from acoustic devices that capture ambient sound and a group of sensors that detect movement and temperature. This data is analyzed in real time and integrated on the server. The analysis uses pre-trained AI algorithms to quickly recognize abnormal behavior and abnormal environmental conditions.

[0227] A key feature of this system is that the server instantly assesses the level of risk based on detected abnormal behavior or conditions and automatically determines countermeasures. As a result of this process, the server sends alert information, including proposed countermeasures, to the terminal. This allows nearby parents and management staff to take swift action to protect children from danger.

[0228] For example, consider a scenario where too many children are gathered on one piece of playground equipment, creating a potential danger. In this case, the server analyzes video data acquired from the camera to detect the abnormal crowding and assesses the risk based on the increased risk of falls and other incidents. It then sends an alert to the terminal, including suggestions for appropriate limits on the number of children or suggestions to guide the children to a different, safer play area.

[0229] Furthermore, if weather changes affect the use of playground equipment, the system can consider sensor information that detects humidity and wind speed to issue appropriate warnings. For example, if playground equipment becomes slippery due to rain, it can send an alert in a timely manner stating that "caution is required when using the playground equipment."

[0230] As a result, parents and administrators, who are the users, can receive these alerts, follow the suggested actions, and take appropriate measures for the environment, thus ensuring that children can play safely.

[0231] The following describes the processing flow.

[0232] Step 1:

[0233] The server collects data in real time from the imaging device, acoustic data acquisition device, and sensor array. This includes video data, acoustic data, and environmental data such as temperature and humidity.

[0234] Step 2:

[0235] The server preprocesses the collected raw data. Specifically, it performs noise reduction, resolution adjustment, and audio data filtering to prepare the data for analysis.

[0236] Step 3:

[0237] The server analyzes pre-processed data using AI algorithms to detect abnormal behavior and conditions. For example, it uses video analysis to understand the density of children around playground equipment.

[0238] Step 4:

[0239] The server assesses the level of risk based on detected anomalies. Factors such as the frequency of the activity and the speed of change are considered in the assessment.

[0240] Step 5:

[0241] The server determines appropriate countermeasures based on the assessed risk level. Specifically, these include limiting the number of people for safety reasons and suggesting alternative play areas.

[0242] Step 6:

[0243] The server sends an alert to the device along with the decided course of action. This notifies parents and administrative staff.

[0244] Step 7:

[0245] The device displays received alerts to the user, using visual and auditory means to encourage prompt action.

[0246] Step 8:

[0247] Based on alert information from their devices, users take action to ensure the safety of children. Specifically, they follow instructions and take measures such as guiding children to safe play areas.

[0248] (Example 1)

[0249] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0250] In today's play environments for children, there is a need for safety monitoring systems to prevent accidents and injuries. However, existing systems have limitations in terms of anomaly detection accuracy and response speed, making it difficult to provide prompt and appropriate countermeasures. This invention aims to provide a safe play environment by detecting hazards in real time with high accuracy and proposing appropriate countermeasures.

[0251] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0252] In this invention, the server includes means for processing video data acquired using a camera and detecting abnormal behavior, means for processing acoustic data acquired using an acoustic collection device and identifying specific acoustic patterns, and means for detecting physical changes using a group of sensors and determining abnormal environmental conditions. This makes it possible to analyze environmental data in real time, quickly and accurately detect anomalies, and immediately derive countermeasures.

[0253] A "recording device" is a device used to acquire video data of children's activities and their surroundings.

[0254] An "acoustic acquisition device" is a device that captures ambient sounds and acquires data to identify specific acoustic patterns.

[0255] A "sensor group" is a group of devices composed of multiple sensors used to detect physical changes such as temperature, humidity, and movement.

[0256] "Abnormal behavior" refers to actions that deviate from normal behavioral patterns and may indicate danger.

[0257] An "abnormal environmental condition" refers to a state in which physical environmental factors such as temperature, humidity, and acoustics deviate significantly from normal values.

[0258] "Risk assessment" is the process of analyzing the risks of detected abnormal behaviors or abnormal environmental conditions and setting priorities according to their urgency.

[0259] "Countermeasures" refer to specific action plans or proposals implemented to mitigate or eliminate a particular risk.

[0260] A "notification device" is a device that receives alert information sent from a server and transmits it to the user.

[0261] A "generative AI model" is a model that implements artificial intelligence algorithms used to analyze acquired data.

[0262] The system of this invention aims to enhance safety in children's play environments by collecting and analyzing data using multiple devices to detect anomalies and promptly issue alerts.

[0263] The server collects video data using cameras installed at the child's location. These cameras have a wide field of view and high resolution, allowing them to capture the child's movements in real time. In addition, an acoustic sound capture device captures ambient sound. This acoustic data is used to quickly identify abnormal or warning sounds. Environmental data from sensors detecting temperature, humidity, and motion is also collected.

[0264] The server integrates this data and performs real-time analysis using a generative AI model. This analysis process utilizes AI algorithms to quickly recognize abnormal behavior and abnormal environmental conditions and assess the level of risk. This allows the server to immediately determine appropriate countermeasures. For example, if multiple children are crowded together on a single piece of playground equipment, the server assesses the risk of overcrowding and generates suggestions for guiding the children.

[0265] The server generates an alert based on the decided response and sends it to the device. The device notifies parents and administrative staff of the alert, quickly sharing important information. This allows nearby parents and administrative staff to take appropriate action quickly and ensure the safety of the child.

[0266] A concrete example would be a prompt message such as, "Monitor the concentration level of children and generate an alert if the density is high." Based on this prompt, the AI ​​model performs data analysis and generates an appropriate alert.

[0267] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0268] Step 1:

[0269] The server acquires video data from the camera in real time. This data includes children's movements and the usage of playground equipment. The input video data is extracted frame by frame and converted into a format that is easy to analyze. This standardizes the image resolution and format, making it suitable for the next processing step.

[0270] Step 2:

[0271] The server acquires acoustic data from the acoustic data acquisition device and identifies specific acoustic patterns. In this process, the input acoustic data is analyzed in terms of time and frequency to extract abnormal sounds and patterns. Specifically, noise filtering and Fourier transforms are used to highlight important features and prepare for the next analysis stage.

[0272] Step 3:

[0273] The server collects environmental data such as temperature, humidity, and motion from a group of sensors. This data is used as input to capture changes in the physical environment. The server normalizes the sensor data and processes it as multidimensional data points to prepare for the detection of abnormal conditions.

[0274] Step 4:

[0275] The server analyzes this integrated data using a generating AI model. It combines input visual, auditory, and environmental data to rapidly recognize abnormal behavior and environmental anomalies. The AI ​​algorithm applies a pre-trained model to classify anomalies and assess risk. The output of this process is a list of anomalous events and their severity ratings.

[0276] Step 5:

[0277] The server assesses the level of risk based on the analysis results and determines countermeasures. Based on the input list of anomalies, the server prioritizes the urgency and necessity of action. The output generates specific recommended actions and response plans.

[0278] Step 6:

[0279] The server generates an alert based on the decided response and sends it to the terminal. The terminal notifies the user of the received alert and prompts them to take the necessary action. Specifically, this could include messages such as "Please temporarily stop using the playground equipment" or "Please move the children to a safe place." This notification enables swift action to ensure the safety of the children.

[0280] (Application Example 1)

[0281] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as a "server", and the smart glasses 214 are referred to as a "terminal".

[0282] For an autonomous vehicle to operate safely, it is required to monitor the surrounding environment in real time, quickly detect potential dangers, and take appropriate actions. However, existing autonomous driving technologies do not respond sufficiently to specific environmental conditions or unexpected situations, and a more refined method for ensuring safety is required.

[0283] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0284] In this invention, the server includes means for processing video data acquired using a photographing device to detect abnormal behavior, means for processing acoustic data acquired using an acoustic collecting device to identify a specific acoustic pattern, means for detecting physical changes using a sensor group to determine an abnormal state of the environment, and automatic control means for monitoring the surrounding environment and performing an emergency stop or route change to enhance the safety in the operation of the vehicle. Thereby, the autonomous vehicle can improve safety in real time.

[0285] The "photographing device" is a device for acquiring video data and can record video information of an object or a landscape.

[0286] The "acoustic collecting device" is a device for acquiring acoustic data and can sense surrounding sound waves and record the information as data.

[0287] The "sensor group" is a collection of a plurality of sensors for acquiring various environmental information and is used to detect physical changes and the state of the environment.

[0288] "Abnormal behavior" refers to unexpected behavior deviating from the normal behavior pattern and indicates a situation that may be dangerous in the system.

[0289] "Risk level" is a measure that assesses the degree of risk in a particular situation, indicating how much attention and response a system requires.

[0290] "Countermeasures" refer to the actions to be taken in response to detected abnormalities or dangers, and are specific action plans to ensure safety.

[0291] A "notification device" is a device that receives alert information from a system and informs the user of its contents.

[0292] "Automatic control means" refers to a function that autonomously adjusts the operation of a vehicle in response to changes in the surrounding conditions, and can perform emergency stops and route changes.

[0293] The system implementing this invention aims to improve the safety of autonomous vehicles by analyzing environmental data acquired from various devices in real time. This system constantly monitors the surrounding environment using a camera, an acoustic data collection device, and a group of sensors.

[0294] The server processes video data acquired from the camera and detects abnormal behavior using AI algorithms. Specifically, it applies pre-trained models using machine learning libraries such as TensorFlow to identify anomalies from the data. Similarly, it analyzes acoustic data from the acoustic collection device using AI models to identify specific acoustic patterns.

[0295] In addition, the server integrates data acquired from the sensor array, detects physical changes such as temperature, humidity, and vibration, and determines abnormal environmental conditions. This process rapidly analyzes input from the sensors and assesses potential hazards in real time.

[0296] If an anomaly is detected, the server adjusts the vehicle's operation via automatic control mechanisms. Specific safety measures include emergency stops and route changes. These control signals are transmitted directly to the vehicle's control system and executed immediately.

[0297] Users receive alerts via notification devices, along with suggestions for specific actions to take. This allows for appropriate responses even from outside the system.

[0298] As a concrete example, consider a scenario where a pedestrian suddenly appears in front of the vehicle while it is operating in the rain. In this case, cameras and sensors work together to understand the situation, AI assesses the danger in real time, and issues a deceleration command to the vehicle control system.

[0299] Examples of prompts for a generative AI model:

[0300] "When a pedestrian appears ahead, what kind of sensor data should be used, and how should it be analyzed, to safely slow down?"

[0301] In this way, the system enhances safety in the operation of autonomous vehicles and can respond quickly and appropriately to changes in the environment.

[0302] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0303] Step 1:

[0304] The server receives video data from the camera as input. The video data is analyzed in real time using an AI algorithm to detect abnormal behavior. Specifically, an image recognition model is used to compare the data with pre-trained abnormal patterns to determine whether an abnormality has occurred. As a result, information regarding the abnormal behavior is output.

[0305] Step 2:

[0306] The server receives acoustic data from the acoustic collection device as input. This data is analyzed by an AI model to identify specific acoustic patterns. For example, speech recognition technology is utilized to detect sudden braking sounds or warning sounds. As a result, the detected acoustic patterns are output and used for evaluating the safety level of the system.

[0307] Step 3:

[0308] The server obtains data from the sensor group as input. Physical changes such as temperature, humidity, and vibration are evaluated in real-time, and calculations are performed to determine abnormal states of the environment. By integrating sensor information and analyzing deviations from normal patterns, abnormalities can be identified. As a result, abnormal information regarding the environmental state is output.

[0309] Step 4:

[0310] Based on the information of abnormal behavior, acoustic patterns, and environmental conditions obtained in Steps 1 to 3, the server evaluates the risk level of the vehicle. A risk score is calculated based on each piece of information, and necessary countermeasures are determined. This includes specific countermeasure plans such as whether an emergency stop is required or the operating route should be changed. The risk assessment result and countermeasures are output.

[0311] Step 5:

[0312] Based on the determined countermeasures, the server outputs a control signal to adjust the operation to the control system of the autonomous vehicle. For example, it includes an emergency stop command or a new route instruction. This control signal prompts the vehicle to react quickly and ensures safety.

[0313] Step 6:

[0314] The terminal receives the alert information sent from the server and notifies the user. The notification content includes information regarding safety and action proposals, assisting the user to take appropriate actions promptly. As a result, the user can recognize the current situation and take proper actions as needed.

[0315] An example of a prompt for the generated AI model is, "When a pedestrian appears ahead, what sensor data should be used and how should it be analyzed to safely slow down?"

[0316] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0317] The system according to the present invention is an advanced monitoring and management system that not only provides a safe and effective play environment but also recognizes the user's emotional state and combines corresponding functions. By combining an emotion engine with a basic system that acquires environmental data using a camera, an acoustic data collection device, and a group of sensors, the system analyzes the user's emotional state in real time.

[0318] Specifically, the server captures the user's facial expressions and voice tone through camera footage and audio data, and the emotion engine uses this information to recognize the user's emotions (joy, anger, fear, etc.). The emotion engine is designed to identify a wide range of emotions by utilizing AI-powered machine learning algorithms.

[0319] The server uses the collected emotional information to determine appropriate countermeasures necessary to ensure the safety of children in the play environment. In this process, emotional information is taken into account in addition to the usual risk assessment, and suggestions for rest to alleviate excessive stress or recommendations for appropriate activities to increase the variety of play are made.

[0320] For example, if a child shows signs of stress during play, the server suggests a relaxing activity and sends an alert to the device saying, "Let's take a break and read a picture book." If the child continues to play and enjoys themselves, the server may suggest, "Let's try a new athletic activity," to further enhance positive emotions.

[0321] By integrating emotional information with conventional environmental data, users such as parents and administrators can implement more appropriate and emotionally sensitive safety measures for children. This provides a safe and comfortable play environment, promoting the healthy development and emotional growth of children.

[0322] The following describes the processing flow.

[0323] Step 1:

[0324] The server acquires user facial expression data and voice data from the camera and sound acquisition devices. This collects raw data that forms the basis of emotions.

[0325] Step 2:

[0326] The server preprocesses the acquired data, performing noise reduction and conversion to the required format. This prepares the data for analysis.

[0327] Step 3:

[0328] The server inputs pre-processed data into the emotion engine to identify the user's emotional state. In this process, a trained AI model identifies emotions such as joy, sadness, and anger.

[0329] Step 4:

[0330] The server aggregates a series of emotional data based on the analysis results of the emotion engine and evaluates the user's stress level and emotional state.

[0331] Step 5:

[0332] The server integrates emotional data and environmental data (abnormal behavior, weather changes, etc.) to reassess the level of risk in the play environment. Based on this assessment, it determines the priority of countermeasures.

[0333] Step 6:

[0334] The server will formulate appropriate countermeasures based on the evaluation results. For example, if emotional data indicates a high-stress state, it will suggest "moving to a rest area."

[0335] Step 7:

[0336] The server sends alert information, including the formulated countermeasures, to the device and notifies parents or administrators.

[0337] Step 8:

[0338] The device displays received alerts to the user visually and audibly, prompting them to take specific action.

[0339] Step 9:

[0340] Based on alert information from the device, the user takes the most appropriate action according to the child's situation. Specifically, they follow the suggestions to guide the child into a safe and positive play environment.

[0341] (Example 2)

[0342] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0343] Existing safety management systems in play environments have struggled to adequately understand and respond to the psychological state of users. Therefore, it has been difficult to implement safety measures that directly address the stress and anxiety experienced by children in particular, and a flexible approach that responds to users' emotional needs is required.

[0344] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0345] In this invention, the server includes means for processing video information acquired using a camera and analyzing the user's emotional state, means for processing acoustic information acquired using an acoustic collection device and analyzing emotions from voice tone, and means for acquiring and analyzing physical environment data using a group of sensors. This enables immediate and appropriate safety measures and action suggestions in accordance with the user's psychological state.

[0346] A "filming device" is a device used to acquire video information and is responsible for capturing the user's movements and facial expressions.

[0347] An "acoustic data acquisition device" is a device used to acquire acoustic information, and its role is to record the user's voice and ambient sounds.

[0348] A "sensor group" is a collection of multiple sensor devices that acquire physical environmental data, such as temperature, motion, and location information.

[0349] "Emotional state" refers to information that indicates the user's psychological and emotional state, and includes psychological elements such as joy, anger, fear, and sadness.

[0350] "Psychological state" refers to the mental and emotional state a user is experiencing, and is a concept that includes various emotions such as stress, relief, and tension.

[0351] A "suggestion alert" is information that the server uses to notify the user of recommended actions, such as relaxing activities or challenging new activities.

[0352] "Environmental data" refers to information that describes the physical state of the environment, and includes data such as temperature, humidity, sound level, and brightness.

[0353] This invention provides a safe and effective play environment, as well as an advanced monitoring and management system for recognizing and responding to the emotional state of users. This system is realized by collecting and analyzing data in real time using a camera, an acoustic data collection device, and a group of sensors.

[0354] The server has the function of processing video information acquired from the camera. It analyzes the user's facial expressions from the video information and recognizes their emotional state in real time. This makes it possible to accurately determine the user's emotions, such as joy, anger, sadness, and happiness. The server also has the function of processing acoustic data acquired from the acoustic collection device, and analyzes the tone of voice to estimate the user's emotions. This uses an AI-based generative model and machine learning algorithms. Furthermore, the server analyzes physical environment data acquired from the sensor group and takes the surrounding environmental conditions into consideration, thereby more comprehensively evaluating the user's psychological state.

[0355] Based on the results of the emotion analysis performed by the server, appropriate countermeasures to meet the user's current emotional needs are quickly determined. This includes emotionally sensitive suggestions such as selecting a play environment and suggesting breaks to reduce stress. For example, if the user is experiencing excessive stress, the server may suggest relaxing activities such as "Let's take a break and read a picture book." If the user is actively engaged, it may suggest actions such as "Let's try a new athletic activity."

[0356] These suggestions are communicated to the user via the device. The device displays the suggestions intuitively through visual and auditory means, making them easy for the user to understand. This allows the user to choose an appropriate action based on their emotional state.

[0357] For example, by inputting a prompt such as, "Based on the user's emotional data, suggest relaxing activities," into the AI ​​model, the system can accurately provide suggestions tailored to the user's individual needs. This enables the provision of a safer and more comfortable play environment.

[0358] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0359] Step 1:

[0360] The server acquires data from the camera and sound acquisition equipment. It receives video data, including the user's facial expressions, from the camera and audio data from the microphone in real time. This provides input to understand the user's current state. Specifically, the camera periodically captures the user's face and sends this as digital video data to the server. The audio data captures the tone and volume of the user's voice and is sent for acoustic analysis.

[0361] Step 2:

[0362] The server analyzes the acquired video and audio data using a generating AI model. From the video data, a machine learning algorithm is used to recognize facial expressions and extract the user's emotional state. From the audio data, additional emotional insights are obtained through voice tone analysis. This data processing outputs results that show the user's overall emotional state. Specifically, the AI ​​model analyzes facial features and generates emotion labels such as anger and joy. From the audio data, it analyzes the pitch and tempo of the voice to derive indicators of stress and calmness.

[0363] Step 3:

[0364] The server takes the analyzed emotional data into account to generate suggestions for a play environment suitable for the user. The main inputs here are the results of the emotional analysis and known safety standards. The server integrates this information to determine appropriate countermeasures and activities. Specifically, if relaxation is determined to be needed, it will recommend activities in a quiet environment. Conversely, if high energy levels are determined, it will suggest vigorous athletic activities.

[0365] Step 4:

[0366] The device notifies the user of action suggestions received from the server. This uses both visual and auditory information to ensure intuitive understanding. The device presents prompts to the user in natural language to encourage action. This notification is output as feedback to the user and influences which activity they choose. For example, the device displays "Let's relax by reading a picture book" on its screen and also conveys the same message aloud.

[0367] Step 5:

[0368] The user receives suggestions from the device and selects an action that aligns with their emotional state. In this final step, the user's choices are returned to the system as feedback, which is used for further analysis and system improvement. Specifically, the user checks an alert and begins taking action according to the suggested activity. This allows the user to continue their activities in a safer and more comfortable environment.

[0369] (Application Example 2)

[0370] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".

[0371] Traditional in-store shopping experiences often lack individuality because they don't adequately consider customers' emotional states. This can lead to decreased customer satisfaction and purchase intent. Therefore, it's crucial to improve the quality of the shopping experience by providing more personalized product and service suggestions based on customer emotions.

[0372] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0373] In this invention, the server includes means for processing video data acquired using a camera and detecting abnormal behavior, means for processing acoustic data acquired using an acoustic collection device and identifying specific acoustic patterns, and means for detecting physical changes using a group of sensors and determining abnormal environmental conditions. This enables real-time analysis of the user's emotional state and the suggestion of suitable products or services based on those emotions.

[0374] A "recording device" is a device used to acquire visual information, such as the conditions in the environment or the user's facial expressions.

[0375] An "acoustic acquisition device" is a device that acquires voice and ambient sounds, and is used to analyze the tone of the user's voice and the surrounding acoustic patterns.

[0376] A "sensor cluster" is a collection of various sensors installed to detect diverse physical changes and is used to determine abnormal environmental conditions.

[0377] "Real-time analysis of emotional state" means instantly processing data such as the user's facial expressions and voice to identify their emotions at that moment.

[0378] "Suggesting suitable products or services" refers to the act of selecting and presenting products or services that match the emotions of the analyzed user.

[0379] The system for realizing this invention is a mechanism that improves the in-store purchasing experience by analyzing the user's emotional state in real time and suggesting products and services based on those emotions.

[0380] The server uses a camera, sound acquisition device, and sensor array to collect data such as the user's facial expressions, voice tone, and surrounding environmental conditions. Specifically, the camera records the user's facial movements and expressions, while the microphone records the user's voice and background sounds. This data is integrated with environmental data collected through the sensor array and transmitted to the server.

[0381] On the server, the collected data is processed using a cloud-based sentiment analysis engine (e.g., Amazon Rekognition or Google Cloud Vision API). This analyzes and identifies the emotions the user is experiencing. Based on this analysis, the server selects and visually presents products and services that are suitable for the user.

[0382] Furthermore, it is possible to utilize generative AI models to create prompt messages that respond to the user's emotions. These prompt messages would serve as information that encourages specific actions from the user, such as, "Identify the user's emotions from the facial expression analysis data and create a prompt to customize the shopping experience based on those emotions."

[0383] As a concrete example, suppose a user is browsing in a store and the server recognizes that the user has made a surprised expression upon seeing a particular product. In this case, the server will encourage a purchase by displaying special offers or new information related to that product on the user's smartphone or display.

[0384] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0385] Step 1:

[0386] The server collects user facial and audio data using a camera and audio acquisition device. Inputs are camera video and audio from a microphone. These inputs are acquired in real time, and initial processing is performed using facial recognition technology and audio analysis techniques. Outputs are the user's basic facial features and audio patterns.

[0387] Step 2:

[0388] The server collects environmental data from a group of sensors. The input consists of physical data such as temperature, humidity, and light intensity within the store. This sensor data is aggregated to evaluate the environmental conditions. The output is an evaluation value indicating abnormal environmental conditions.

[0389] Step 3:

[0390] The server integrates the collected facial expression data, acoustic data, and environmental data to perform emotion analysis. The input consists of the data obtained in Step 1 and Step 2. This data is integrated, and the emotion analysis engine is used to identify the user's emotions. The output is the user's emotional state (e.g., joy, surprise, anger).

[0391] Step 4:

[0392] The server selects suitable products and services based on the results of sentiment analysis. The input is information about the user's emotional state. It references a database to extract products and services that match the emotion. The output is information about the suggested products and services.

[0393] Step 5:

[0394] The terminal receives suggestion information from the server and presents it visually to the user. The input is suggestion information about products and services sent from the server. This information is displayed on a display or smartphone screen to attract the user's attention. The output is the visual information provided to the user.

[0395] Step 6:

[0396] The server utilizes a generative AI model to create emotion-based prompts. The input consists of the user's emotional state and suggested product information. Based on this data, the generative AI model generates prompts. The output is a prompt that encourages the user to take a specific action.

[0397] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0398] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0399] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0400] [Third Embodiment]

[0401] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0402] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0403] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0404] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0405] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0406] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0407] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0408] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0409] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0410] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0411] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0412] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0413] The system according to the present invention has the function of detecting dangers and abnormalities and issuing alerts by acquiring environmental data using a camera, an acoustic data collection device, and a group of sensors, and analyzing it in real time, in order to enhance safety in environments where children play.

[0414] Specifically, the server collects video data from multiple cameras installed in the child's location, and also aggregates environmental data from acoustic devices that capture ambient sound and a group of sensors that detect movement and temperature. This data is analyzed in real time and integrated on the server. The analysis uses pre-trained AI algorithms to quickly recognize abnormal behavior and abnormal environmental conditions.

[0415] A key feature of this system is that the server instantly assesses the level of risk based on detected abnormal behavior or conditions and automatically determines countermeasures. As a result of this process, the server sends alert information, including proposed countermeasures, to the terminal. This allows nearby parents and management staff to take swift action to protect children from danger.

[0416] For example, consider a scenario where too many children are gathered on one piece of playground equipment, creating a potential danger. In this case, the server analyzes video data acquired from the camera to detect the abnormal crowding and assesses the risk based on the increased risk of falls and other incidents. It then sends an alert to the terminal, including suggestions for appropriate limits on the number of children or suggestions to guide the children to a different, safer play area.

[0417] Furthermore, if weather changes affect the use of playground equipment, the system can consider sensor information that detects humidity and wind speed to issue appropriate warnings. For example, if playground equipment becomes slippery due to rain, it can send an alert in a timely manner stating that "caution is required when using the playground equipment."

[0418] As a result, parents and administrators, who are the users, can receive these alerts, follow the suggested actions, and take appropriate measures for the environment, thus ensuring that children can play safely.

[0419] The following describes the processing flow.

[0420] Step 1:

[0421] The server collects data in real time from the imaging device, acoustic data acquisition device, and sensor array. This includes video data, acoustic data, and environmental data such as temperature and humidity.

[0422] Step 2:

[0423] The server preprocesses the collected raw data. Specifically, it performs noise reduction, resolution adjustment, and audio data filtering to prepare the data for analysis.

[0424] Step 3:

[0425] The server analyzes pre-processed data using AI algorithms to detect abnormal behavior and conditions. For example, it uses video analysis to understand the density of children around playground equipment.

[0426] Step 4:

[0427] The server assesses the level of risk based on detected anomalies. Factors such as the frequency of the activity and the speed of change are considered in the assessment.

[0428] Step 5:

[0429] The server determines appropriate countermeasures based on the assessed risk level. Specifically, these include limiting the number of people for safety reasons and suggesting alternative play areas.

[0430] Step 6:

[0431] The server sends an alert to the device along with the decided course of action. This notifies parents and administrative staff.

[0432] Step 7:

[0433] The device displays received alerts to the user, using visual and auditory means to encourage prompt action.

[0434] Step 8:

[0435] Based on alert information from their devices, users take action to ensure the safety of children. Specifically, they follow instructions and take measures such as guiding children to safe play areas.

[0436] (Example 1)

[0437] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0438] In today's play environments for children, there is a need for safety monitoring systems to prevent accidents and injuries. However, existing systems have limitations in terms of anomaly detection accuracy and response speed, making it difficult to provide prompt and appropriate countermeasures. This invention aims to provide a safe play environment by detecting hazards in real time with high accuracy and proposing appropriate countermeasures.

[0439] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0440] In this invention, the server includes means for processing video data acquired using a camera and detecting abnormal behavior, means for processing acoustic data acquired using an acoustic collection device and identifying specific acoustic patterns, and means for detecting physical changes using a group of sensors and determining abnormal environmental conditions. This makes it possible to analyze environmental data in real time, quickly and accurately detect anomalies, and immediately derive countermeasures.

[0441] A "recording device" is a device used to acquire video data of children's activities and their surroundings.

[0442] An "acoustic acquisition device" is a device that captures ambient sounds and acquires data to identify specific acoustic patterns.

[0443] A "sensor group" is a group of devices composed of multiple sensors used to detect physical changes such as temperature, humidity, and movement.

[0444] "Abnormal behavior" refers to actions that deviate from normal behavioral patterns and may indicate danger.

[0445] An "abnormal environmental condition" refers to a state in which physical environmental factors such as temperature, humidity, and acoustics deviate significantly from normal values.

[0446] "Risk assessment" is the process of analyzing the risks of detected abnormal behaviors or abnormal environmental conditions and setting priorities according to their urgency.

[0447] "Countermeasures" refer to specific action plans or proposals implemented to mitigate or eliminate a particular risk.

[0448] A "notification device" is a device that receives alert information sent from a server and transmits it to the user.

[0449] A "generative AI model" is a model that implements artificial intelligence algorithms used to analyze acquired data.

[0450] The system of this invention aims to enhance safety in children's play environments by collecting and analyzing data using multiple devices to detect anomalies and promptly issue alerts.

[0451] The server collects video data using cameras installed at the child's location. These cameras have a wide field of view and high resolution, allowing them to capture the child's movements in real time. In addition, an acoustic sound capture device captures ambient sound. This acoustic data is used to quickly identify abnormal or warning sounds. Environmental data from sensors detecting temperature, humidity, and motion is also collected.

[0452] The server integrates this data and performs real-time analysis using a generative AI model. This analysis process utilizes AI algorithms to quickly recognize abnormal behavior and abnormal environmental conditions and assess the level of risk. This allows the server to immediately determine appropriate countermeasures. For example, if multiple children are crowded together on a single piece of playground equipment, the server assesses the risk of overcrowding and generates suggestions for guiding the children.

[0453] The server generates an alert based on the decided response and sends it to the device. The device notifies parents and administrative staff of the alert, quickly sharing important information. This allows nearby parents and administrative staff to take appropriate action quickly and ensure the safety of the child.

[0454] A concrete example would be a prompt message such as, "Monitor the concentration level of children and generate an alert if the density is high." Based on this prompt, the AI ​​model performs data analysis and generates an appropriate alert.

[0455] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0456] Step 1:

[0457] The server acquires video data from the camera in real time. This data includes children's movements and the usage of playground equipment. The input video data is extracted frame by frame and converted into a format that is easy to analyze. This standardizes the image resolution and format, making it suitable for the next processing step.

[0458] Step 2:

[0459] The server acquires acoustic data from the acoustic data acquisition device and identifies specific acoustic patterns. In this process, the input acoustic data is analyzed in terms of time and frequency to extract abnormal sounds and patterns. Specifically, noise filtering and Fourier transforms are used to highlight important features and prepare for the next analysis stage.

[0460] Step 3:

[0461] The server collects environmental data such as temperature, humidity, and motion from a group of sensors. This data is used as input to capture changes in the physical environment. The server normalizes the sensor data and processes it as multidimensional data points to prepare for the detection of abnormal conditions.

[0462] Step 4:

[0463] The server analyzes this integrated data using a generating AI model. It combines input visual, auditory, and environmental data to rapidly recognize abnormal behavior and environmental anomalies. The AI ​​algorithm applies a pre-trained model to classify anomalies and assess risk. The output of this process is a list of anomalous events and their severity ratings.

[0464] Step 5:

[0465] The server assesses the level of risk based on the analysis results and determines countermeasures. Based on the input list of anomalies, the server prioritizes the urgency and necessity of action. The output generates specific recommended actions and response plans.

[0466] Step 6:

[0467] The server generates an alert based on the decided response and sends it to the terminal. The terminal notifies the user of the received alert and prompts them to take the necessary action. Specifically, this could include messages such as "Please temporarily stop using the playground equipment" or "Please move the children to a safe place." This notification enables swift action to ensure the safety of the children.

[0468] (Application Example 1)

[0469] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0470] For autonomous vehicles to operate safely, they must monitor their surroundings in real time, quickly detect potential hazards, and take appropriate action. However, existing autonomous driving technologies are not sufficient to respond to specific environmental conditions or unexpected situations, and more sophisticated methods for ensuring safety are needed.

[0471] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0472] In this invention, the server includes means for processing video data acquired using a camera and detecting abnormal behavior; means for processing acoustic data acquired using an acoustic collection device and identifying specific acoustic patterns; means for detecting physical changes using a group of sensors and determining abnormal environmental conditions; and automatic control means for monitoring the surrounding environment and performing emergency stops or route changes to enhance the safety of vehicle operation. This enables autonomous vehicles to improve safety in real time.

[0473] A "filming device" is a device used to acquire video data and can record video information of objects or landscapes.

[0474] An "acoustic data acquisition device" is a device for acquiring acoustic data, which can sense ambient sound waves and record that information as data.

[0475] A "sensor cluster" is a collection of multiple sensors used to acquire various environmental information and to detect physical changes and environmental conditions.

[0476] "Abnormal behavior" refers to unexpected actions that deviate from normal behavioral patterns, indicating a potentially dangerous situation within a system.

[0477] "Risk level" is a measure that assesses the degree of risk in a particular situation, indicating how much attention and response a system requires.

[0478] "Countermeasures" refer to the actions to be taken in response to detected abnormalities or dangers, and are specific action plans to ensure safety.

[0479] A "notification device" is a device that receives alert information from a system and informs the user of its contents.

[0480] "Automatic control means" refers to a function that autonomously adjusts the operation of a vehicle in response to changes in the surrounding conditions, and can perform emergency stops and route changes.

[0481] The system implementing this invention aims to improve the safety of autonomous vehicles by analyzing environmental data acquired from various devices in real time. This system constantly monitors the surrounding environment using a camera, an acoustic data collection device, and a group of sensors.

[0482] The server processes video data acquired from the camera and detects abnormal behavior using AI algorithms. Specifically, it applies pre-trained models using machine learning libraries such as TensorFlow to identify anomalies from the data. Similarly, it analyzes acoustic data from the acoustic collection device using AI models to identify specific acoustic patterns.

[0483] In addition, the server integrates data acquired from the sensor array, detects physical changes such as temperature, humidity, and vibration, and determines abnormal environmental conditions. This process rapidly analyzes input from the sensors and assesses potential hazards in real time.

[0484] If an anomaly is detected, the server adjusts the vehicle's operation via automatic control mechanisms. Specific safety measures include emergency stops and route changes. These control signals are transmitted directly to the vehicle's control system and executed immediately.

[0485] Users receive alerts via notification devices, along with suggestions for specific actions to take. This allows for appropriate responses even from outside the system.

[0486] As a concrete example, consider a scenario where a pedestrian suddenly appears in front of the vehicle while it is operating in the rain. In this case, cameras and sensors work together to understand the situation, AI assesses the danger in real time, and issues a deceleration command to the vehicle control system.

[0487] Examples of prompts for a generative AI model:

[0488] "When a pedestrian appears ahead, what kind of sensor data should be used, and how should it be analyzed, to safely slow down?"

[0489] In this way, the system enhances safety in the operation of autonomous vehicles and can respond quickly and appropriately to changes in the environment.

[0490] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0491] Step 1:

[0492] The server receives video data from the camera as input. The video data is analyzed in real time using an AI algorithm to detect abnormal behavior. Specifically, an image recognition model is used to compare the data with pre-trained abnormal patterns to determine whether an abnormality has occurred. As a result, information regarding the abnormal behavior is output.

[0493] Step 2:

[0494] The server receives acoustic data from an acoustic data collection device as input. This data is analyzed by an AI model, and specific acoustic patterns are identified. For example, speech recognition technology is used to detect sudden braking sounds or warning sounds. The detected acoustic patterns are then output and used to evaluate the safety of the system.

[0495] Step 3:

[0496] The server receives data from a group of sensors as input. Physical changes such as temperature, humidity, and vibration are evaluated in real time, and calculations are performed to determine abnormal environmental conditions. By integrating the sensor information and analyzing deviations from normal patterns, anomalies can be identified. As a result, anomaly information regarding the environmental condition is output.

[0497] Step 4:

[0498] The server assesses the vehicle's risk level based on the abnormal behavior, sound patterns, and environmental conditions obtained in steps 1-3. It calculates a risk score based on this information and determines the necessary countermeasures. This includes specific measures such as whether an emergency stop is necessary or whether the route should be changed. The risk assessment results and countermeasures are then output.

[0499] Step 5:

[0500] Based on the determined response, the server outputs control signals to the autonomous vehicle's control system to adjust its operation. These signals may include emergency stop orders or new route instructions. These control signals encourage a rapid response from the vehicle and ensure safety.

[0501] Step 6:

[0502] The device receives alert information sent from the server and notifies the user. The notification includes security-related information and action suggestions to help the user take prompt action. This allows the user to understand the situation and take appropriate action as needed.

[0503] An example of a prompt for the generated AI model is, "When a pedestrian appears ahead, what sensor data should be used and how should it be analyzed to safely slow down?"

[0504] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0505] The system according to the present invention is an advanced monitoring and management system that not only provides a safe and effective play environment but also recognizes the user's emotional state and combines corresponding functions. By combining an emotion engine with a basic system that acquires environmental data using a camera, an acoustic data collection device, and a group of sensors, the system analyzes the user's emotional state in real time.

[0506] Specifically, the server captures the user's facial expressions and voice tone through camera footage and audio data, and the emotion engine uses this information to recognize the user's emotions (joy, anger, fear, etc.). The emotion engine is designed to identify a wide range of emotions by utilizing AI-powered machine learning algorithms.

[0507] The server uses the collected emotional information to determine appropriate countermeasures necessary to ensure the safety of children in the play environment. In this process, emotional information is taken into account in addition to the usual risk assessment, and suggestions for rest to alleviate excessive stress or recommendations for appropriate activities to increase the variety of play are made.

[0508] For example, if a child shows signs of stress during play, the server suggests a relaxing activity and sends an alert to the device saying, "Let's take a break and read a picture book." If the child continues to play and enjoys themselves, the server may suggest, "Let's try a new athletic activity," to further enhance positive emotions.

[0509] By integrating emotional information with conventional environmental data, users such as parents and administrators can implement more appropriate and emotionally sensitive safety measures for children. This provides a safe and comfortable play environment, promoting the healthy development and emotional growth of children.

[0510] The following describes the processing flow.

[0511] Step 1:

[0512] The server acquires user facial expression data and voice data from the camera and sound acquisition devices. This collects raw data that forms the basis of emotions.

[0513] Step 2:

[0514] The server preprocesses the acquired data, performing noise reduction and conversion to the required format. This prepares the data for analysis.

[0515] Step 3:

[0516] The server inputs pre-processed data into the emotion engine to identify the user's emotional state. In this process, a trained AI model identifies emotions such as joy, sadness, and anger.

[0517] Step 4:

[0518] The server aggregates a series of emotional data based on the analysis results of the emotion engine and evaluates the user's stress level and emotional state.

[0519] Step 5:

[0520] The server integrates emotional data and environmental data (abnormal behavior, weather changes, etc.) to reassess the level of risk in the play environment. Based on this assessment, it determines the priority of countermeasures.

[0521] Step 6:

[0522] The server will formulate appropriate countermeasures based on the evaluation results. For example, if emotional data indicates a high-stress state, it will suggest "moving to a rest area."

[0523] Step 7:

[0524] The server sends alert information, including the formulated countermeasures, to the device and notifies parents or administrators.

[0525] Step 8:

[0526] The device displays received alerts to the user visually and audibly, prompting them to take specific action.

[0527] Step 9:

[0528] Based on alert information from the device, the user takes the most appropriate action according to the child's situation. Specifically, they follow the suggestions to guide the child into a safe and positive play environment.

[0529] (Example 2)

[0530] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0531] Existing safety management systems in play environments have struggled to adequately understand and respond to the psychological state of users. Therefore, it has been difficult to implement safety measures that directly address the stress and anxiety experienced by children in particular, and a flexible approach that responds to users' emotional needs is required.

[0532] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0533] In this invention, the server includes means for processing video information acquired using a camera and analyzing the user's emotional state, means for processing acoustic information acquired using an acoustic collection device and analyzing emotions from voice tone, and means for acquiring and analyzing physical environment data using a group of sensors. This enables immediate and appropriate safety measures and action suggestions in accordance with the user's psychological state.

[0534] A "filming device" is a device used to acquire video information and is responsible for capturing the user's movements and facial expressions.

[0535] An "acoustic data acquisition device" is a device used to acquire acoustic information, and its role is to record the user's voice and ambient sounds.

[0536] A "sensor group" is a collection of multiple sensor devices that acquire physical environmental data, such as temperature, motion, and location information.

[0537] "Emotional state" refers to information that indicates the user's psychological and emotional state, and includes psychological elements such as joy, anger, fear, and sadness.

[0538] "Psychological state" refers to the mental and emotional state a user is experiencing, and is a concept that includes various emotions such as stress, relief, and tension.

[0539] A "suggestion alert" is information that the server uses to notify the user of recommended actions, such as relaxing activities or challenging new activities.

[0540] "Environmental data" refers to information that describes the physical state of the environment, and includes data such as temperature, humidity, sound level, and brightness.

[0541] This invention provides a safe and effective play environment, as well as an advanced monitoring and management system for recognizing and responding to the emotional state of users. This system is realized by collecting and analyzing data in real time using a camera, an acoustic data collection device, and a group of sensors.

[0542] The server has the function of processing video information acquired from the camera. It analyzes the user's facial expressions from the video information and recognizes their emotional state in real time. This makes it possible to accurately determine the user's emotions, such as joy, anger, sadness, and happiness. The server also has the function of processing acoustic data acquired from the acoustic collection device, and analyzes the tone of voice to estimate the user's emotions. This uses an AI-based generative model and machine learning algorithms. Furthermore, the server analyzes physical environment data acquired from the sensor group and takes the surrounding environmental conditions into consideration, thereby more comprehensively evaluating the user's psychological state.

[0543] Based on the results of the emotion analysis performed by the server, appropriate countermeasures to meet the user's current emotional needs are quickly determined. This includes emotionally sensitive suggestions such as selecting a play environment and suggesting breaks to reduce stress. For example, if the user is experiencing excessive stress, the server may suggest relaxing activities such as "Let's take a break and read a picture book." If the user is actively engaged, it may suggest actions such as "Let's try a new athletic activity."

[0544] These suggestions are communicated to the user via the device. The device displays the suggestions intuitively through visual and auditory means, making them easy for the user to understand. This allows the user to choose an appropriate action based on their emotional state.

[0545] For example, by inputting a prompt such as, "Based on the user's emotional data, suggest relaxing activities," into the AI ​​model, the system can accurately provide suggestions tailored to the user's individual needs. This enables the provision of a safer and more comfortable play environment.

[0546] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0547] Step 1:

[0548] The server acquires data from the camera and sound acquisition equipment. It receives video data, including the user's facial expressions, from the camera and audio data from the microphone in real time. This provides input to understand the user's current state. Specifically, the camera periodically captures the user's face and sends this as digital video data to the server. The audio data captures the tone and volume of the user's voice and is sent for acoustic analysis.

[0549] Step 2:

[0550] The server analyzes the acquired video and audio data using a generating AI model. From the video data, a machine learning algorithm is used to recognize facial expressions and extract the user's emotional state. From the audio data, additional emotional insights are obtained through voice tone analysis. This data processing outputs results that show the user's overall emotional state. Specifically, the AI ​​model analyzes facial features and generates emotion labels such as anger and joy. From the audio data, it analyzes the pitch and tempo of the voice to derive indicators of stress and calmness.

[0551] Step 3:

[0552] The server takes the analyzed emotional data into account to generate suggestions for a play environment suitable for the user. The main inputs here are the results of the emotional analysis and known safety standards. The server integrates this information to determine appropriate countermeasures and activities. Specifically, if relaxation is determined to be needed, it will recommend activities in a quiet environment. Conversely, if high energy levels are determined, it will suggest vigorous athletic activities.

[0553] Step 4:

[0554] The device notifies the user of action suggestions received from the server. This uses both visual and auditory information to ensure intuitive understanding. The device presents prompts to the user in natural language to encourage action. This notification is output as feedback to the user and influences which activity they choose. For example, the device displays "Let's relax by reading a picture book" on its screen and also conveys the same message aloud.

[0555] Step 5:

[0556] The user receives suggestions from the device and selects an action that aligns with their emotional state. In this final step, the user's choices are returned to the system as feedback, which is used for further analysis and system improvement. Specifically, the user checks an alert and begins taking action according to the suggested activity. This allows the user to continue their activities in a safer and more comfortable environment.

[0557] (Application Example 2)

[0558] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0559] Traditional in-store shopping experiences often lack individuality because they don't adequately consider customers' emotional states. This can lead to decreased customer satisfaction and purchase intent. Therefore, it's crucial to improve the quality of the shopping experience by providing more personalized product and service suggestions based on customer emotions.

[0560] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0561] In this invention, the server includes means for processing video data acquired using a camera and detecting abnormal behavior, means for processing acoustic data acquired using an acoustic collection device and identifying specific acoustic patterns, and means for detecting physical changes using a group of sensors and determining abnormal environmental conditions. This enables real-time analysis of the user's emotional state and the suggestion of suitable products or services based on those emotions.

[0562] A "recording device" is a device used to acquire visual information, such as the conditions in the environment or the user's facial expressions.

[0563] An "acoustic acquisition device" is a device that acquires voice and ambient sounds, and is used to analyze the tone of the user's voice and the surrounding acoustic patterns.

[0564] A "sensor cluster" is a collection of various sensors installed to detect diverse physical changes and is used to determine abnormal environmental conditions.

[0565] "Real-time analysis of emotional state" means instantly processing data such as the user's facial expressions and voice to identify their emotions at that moment.

[0566] "Suggesting suitable products or services" refers to the act of selecting and presenting products or services that match the emotions of the analyzed user.

[0567] The system for realizing this invention is a mechanism that improves the in-store purchasing experience by analyzing the user's emotional state in real time and suggesting products and services based on those emotions.

[0568] The server uses a camera, sound acquisition device, and sensor array to collect data such as the user's facial expressions, voice tone, and surrounding environmental conditions. Specifically, the camera records the user's facial movements and expressions, while the microphone records the user's voice and background sounds. This data is integrated with environmental data collected through the sensor array and transmitted to the server.

[0569] On the server, the collected data is processed using a cloud-based sentiment analysis engine (e.g., Amazon Rekognition or Google Cloud Vision API). This analyzes and identifies the emotions the user is experiencing. Based on this analysis, the server selects and visually presents products and services that are suitable for the user.

[0570] Furthermore, it is possible to utilize generative AI models to create prompt messages that respond to the user's emotions. These prompt messages would serve as information that encourages specific actions from the user, such as, "Identify the user's emotions from the facial expression analysis data and create a prompt to customize the shopping experience based on those emotions."

[0571] As a concrete example, suppose a user is browsing in a store and the server recognizes that the user has made a surprised expression upon seeing a particular product. In this case, the server will encourage a purchase by displaying special offers or new information related to that product on the user's smartphone or display.

[0572] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0573] Step 1:

[0574] The server collects user facial and audio data using a camera and audio acquisition device. Inputs are camera video and audio from a microphone. These inputs are acquired in real time, and initial processing is performed using facial recognition technology and audio analysis techniques. Outputs are the user's basic facial features and audio patterns.

[0575] Step 2:

[0576] The server collects environmental data from a group of sensors. The input consists of physical data such as temperature, humidity, and light intensity within the store. This sensor data is aggregated to evaluate the environmental conditions. The output is an evaluation value indicating abnormal environmental conditions.

[0577] Step 3:

[0578] The server integrates the collected facial expression data, acoustic data, and environmental data to perform emotion analysis. The input consists of the data obtained in Step 1 and Step 2. This data is integrated, and the emotion analysis engine is used to identify the user's emotions. The output is the user's emotional state (e.g., joy, surprise, anger).

[0579] Step 4:

[0580] The server selects suitable products and services based on the results of sentiment analysis. The input is information about the user's emotional state. It references a database to extract products and services that match the emotion. The output is information about the suggested products and services.

[0581] Step 5:

[0582] The terminal receives suggestion information from the server and presents it visually to the user. The input is suggestion information about products and services sent from the server. This information is displayed on a display or smartphone screen to attract the user's attention. The output is the visual information provided to the user.

[0583] Step 6:

[0584] The server utilizes a generative AI model to create emotion-based prompts. The input consists of the user's emotional state and suggested product information. Based on this data, the generative AI model generates prompts. The output is a prompt that encourages the user to take a specific action.

[0585] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0586] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0587] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0588] [Fourth Embodiment]

[0589] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0590] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0591] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0592] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0593] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0594] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0595] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0596] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0597] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0598] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0599] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0600] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0601] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0602] The system according to the present invention has the function of detecting dangers and abnormalities and issuing alerts by acquiring environmental data using a camera, an acoustic data collection device, and a group of sensors, and analyzing it in real time, in order to enhance safety in environments where children play.

[0603] Specifically, the server collects video data from multiple cameras installed in the child's location, and also aggregates environmental data from acoustic devices that capture ambient sound and a group of sensors that detect movement and temperature. This data is analyzed in real time and integrated on the server. The analysis uses pre-trained AI algorithms to quickly recognize abnormal behavior and abnormal environmental conditions.

[0604] A key feature of this system is that the server instantly assesses the level of risk based on detected abnormal behavior or conditions and automatically determines countermeasures. As a result of this process, the server sends alert information, including proposed countermeasures, to the terminal. This allows nearby parents and management staff to take swift action to protect children from danger.

[0605] For example, consider a scenario where too many children are gathered on one piece of playground equipment, creating a potential danger. In this case, the server analyzes video data acquired from the camera to detect the abnormal crowding and assesses the risk based on the increased risk of falls and other incidents. It then sends an alert to the terminal, including suggestions for appropriate limits on the number of children or suggestions to guide the children to a different, safer play area.

[0606] Furthermore, if weather changes affect the use of playground equipment, the system can consider sensor information that detects humidity and wind speed to issue appropriate warnings. For example, if playground equipment becomes slippery due to rain, it can send an alert in a timely manner stating that "caution is required when using the playground equipment."

[0607] As a result, parents and administrators, who are the users, can receive these alerts, follow the suggested actions, and take appropriate measures for the environment, thus ensuring that children can play safely.

[0608] The following describes the processing flow.

[0609] Step 1:

[0610] The server collects data in real time from the imaging device, acoustic data acquisition device, and sensor array. This includes video data, acoustic data, and environmental data such as temperature and humidity.

[0611] Step 2:

[0612] The server preprocesses the collected raw data. Specifically, it performs noise reduction, resolution adjustment, and audio data filtering to prepare the data for analysis.

[0613] Step 3:

[0614] The server analyzes pre-processed data using AI algorithms to detect abnormal behavior and conditions. For example, it uses video analysis to understand the density of children around playground equipment.

[0615] Step 4:

[0616] The server assesses the level of risk based on detected anomalies. Factors such as the frequency of the activity and the speed of change are considered in the assessment.

[0617] Step 5:

[0618] The server determines appropriate countermeasures based on the assessed risk level. Specifically, these include limiting the number of people for safety reasons and suggesting alternative play areas.

[0619] Step 6:

[0620] The server sends an alert to the device along with the decided course of action. This notifies parents and administrative staff.

[0621] Step 7:

[0622] The device displays received alerts to the user, using visual and auditory means to encourage prompt action.

[0623] Step 8:

[0624] Based on alert information from their devices, users take action to ensure the safety of children. Specifically, they follow instructions and take measures such as guiding children to safe play areas.

[0625] (Example 1)

[0626] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0627] In today's play environments for children, there is a need for safety monitoring systems to prevent accidents and injuries. However, existing systems have limitations in terms of anomaly detection accuracy and response speed, making it difficult to provide prompt and appropriate countermeasures. This invention aims to provide a safe play environment by detecting hazards in real time with high accuracy and proposing appropriate countermeasures.

[0628] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0629] In this invention, the server includes means for processing video data acquired using a camera and detecting abnormal behavior, means for processing acoustic data acquired using an acoustic collection device and identifying specific acoustic patterns, and means for detecting physical changes using a group of sensors and determining abnormal environmental conditions. This makes it possible to analyze environmental data in real time, quickly and accurately detect anomalies, and immediately derive countermeasures.

[0630] A "recording device" is a device used to acquire video data of children's activities and their surroundings.

[0631] An "acoustic acquisition device" is a device that captures ambient sounds and acquires data to identify specific acoustic patterns.

[0632] A "sensor group" is a group of devices composed of multiple sensors used to detect physical changes such as temperature, humidity, and movement.

[0633] "Abnormal behavior" refers to actions that deviate from normal behavioral patterns and may indicate danger.

[0634] An "abnormal environmental condition" refers to a state in which physical environmental factors such as temperature, humidity, and acoustics deviate significantly from normal values.

[0635] "Risk assessment" is the process of analyzing the risks of detected abnormal behaviors or abnormal environmental conditions and setting priorities according to their urgency.

[0636] "Countermeasures" refer to specific action plans or proposals implemented to mitigate or eliminate a particular risk.

[0637] A "notification device" is a device that receives alert information sent from a server and transmits it to the user.

[0638] A "generative AI model" is a model that implements artificial intelligence algorithms used to analyze acquired data.

[0639] The system of this invention aims to enhance safety in children's play environments by collecting and analyzing data using multiple devices to detect anomalies and promptly issue alerts.

[0640] The server collects video data using cameras installed at the child's location. These cameras have a wide field of view and high resolution, allowing them to capture the child's movements in real time. In addition, an acoustic sound capture device captures ambient sound. This acoustic data is used to quickly identify abnormal or warning sounds. Environmental data from sensors detecting temperature, humidity, and motion is also collected.

[0641] The server integrates this data and performs real-time analysis using a generative AI model. This analysis process utilizes AI algorithms to quickly recognize abnormal behavior and abnormal environmental conditions and assess the level of risk. This allows the server to immediately determine appropriate countermeasures. For example, if multiple children are crowded together on a single piece of playground equipment, the server assesses the risk of overcrowding and generates suggestions for guiding the children.

[0642] The server generates an alert based on the decided response and sends it to the device. The device notifies parents and administrative staff of the alert, quickly sharing important information. This allows nearby parents and administrative staff to take appropriate action quickly and ensure the safety of the child.

[0643] A concrete example would be a prompt message such as, "Monitor the concentration level of children and generate an alert if the density is high." Based on this prompt, the AI ​​model performs data analysis and generates an appropriate alert.

[0644] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0645] Step 1:

[0646] The server acquires video data from the camera in real time. This data includes children's movements and the usage of playground equipment. The input video data is extracted frame by frame and converted into a format that is easy to analyze. This standardizes the image resolution and format, making it suitable for the next processing step.

[0647] Step 2:

[0648] The server acquires acoustic data from the acoustic data acquisition device and identifies specific acoustic patterns. In this process, the input acoustic data is analyzed in terms of time and frequency to extract abnormal sounds and patterns. Specifically, noise filtering and Fourier transforms are used to highlight important features and prepare for the next analysis stage.

[0649] Step 3:

[0650] The server collects environmental data such as temperature, humidity, and motion from a group of sensors. This data is used as input to capture changes in the physical environment. The server normalizes the sensor data and processes it as multidimensional data points to prepare for the detection of abnormal conditions.

[0651] Step 4:

[0652] The server analyzes this integrated data using a generating AI model. It combines input visual, auditory, and environmental data to rapidly recognize abnormal behavior and environmental anomalies. The AI ​​algorithm applies a pre-trained model to classify anomalies and assess risk. The output of this process is a list of anomalous events and their severity ratings.

[0653] Step 5:

[0654] The server assesses the level of risk based on the analysis results and determines countermeasures. Based on the input list of anomalies, the server prioritizes the urgency and necessity of action. The output generates specific recommended actions and response plans.

[0655] Step 6:

[0656] The server generates an alert based on the decided response and sends it to the terminal. The terminal notifies the user of the received alert and prompts them to take the necessary action. Specifically, this could include messages such as "Please temporarily stop using the playground equipment" or "Please move the children to a safe place." This notification enables swift action to ensure the safety of the children.

[0657] (Application Example 1)

[0658] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0659] For autonomous vehicles to operate safely, they must monitor their surroundings in real time, quickly detect potential hazards, and take appropriate action. However, existing autonomous driving technologies are not sufficient to respond to specific environmental conditions or unexpected situations, and more sophisticated methods for ensuring safety are needed.

[0660] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0661] In this invention, the server includes means for processing video data acquired using a camera and detecting abnormal behavior; means for processing acoustic data acquired using an acoustic collection device and identifying specific acoustic patterns; means for detecting physical changes using a group of sensors and determining abnormal environmental conditions; and automatic control means for monitoring the surrounding environment and performing emergency stops or route changes to enhance the safety of vehicle operation. This enables autonomous vehicles to improve safety in real time.

[0662] A "filming device" is a device used to acquire video data and can record video information of objects or landscapes.

[0663] An "acoustic data acquisition device" is a device for acquiring acoustic data, which can sense ambient sound waves and record that information as data.

[0664] A "sensor cluster" is a collection of multiple sensors used to acquire various environmental information and to detect physical changes and environmental conditions.

[0665] "Abnormal behavior" refers to unexpected actions that deviate from normal behavioral patterns, indicating a potentially dangerous situation within a system.

[0666] "Risk level" is a measure that assesses the degree of risk in a particular situation, indicating how much attention and response a system requires.

[0667] "Countermeasures" refer to the actions to be taken in response to detected abnormalities or dangers, and are specific action plans to ensure safety.

[0668] A "notification device" is a device that receives alert information from a system and informs the user of its contents.

[0669] "Automatic control means" refers to a function that autonomously adjusts the operation of a vehicle in response to changes in the surrounding conditions, and can perform emergency stops and route changes.

[0670] The system implementing this invention aims to improve the safety of autonomous vehicles by analyzing environmental data acquired from various devices in real time. This system constantly monitors the surrounding environment using a camera, an acoustic data collection device, and a group of sensors.

[0671] The server processes video data acquired from the camera and detects abnormal behavior using AI algorithms. Specifically, it applies pre-trained models using machine learning libraries such as TensorFlow to identify anomalies from the data. Similarly, it analyzes acoustic data from the acoustic collection device using AI models to identify specific acoustic patterns.

[0672] In addition, the server integrates data acquired from the sensor array, detects physical changes such as temperature, humidity, and vibration, and determines abnormal environmental conditions. This process rapidly analyzes input from the sensors and assesses potential hazards in real time.

[0673] If an anomaly is detected, the server adjusts the vehicle's operation via automatic control mechanisms. Specific safety measures include emergency stops and route changes. These control signals are transmitted directly to the vehicle's control system and executed immediately.

[0674] Users receive alerts via notification devices, along with suggestions for specific actions to take. This allows for appropriate responses even from outside the system.

[0675] As a concrete example, consider a scenario where a pedestrian suddenly appears in front of the vehicle while it is operating in the rain. In this case, cameras and sensors work together to understand the situation, AI assesses the danger in real time, and issues a deceleration command to the vehicle control system.

[0676] Examples of prompts for a generative AI model:

[0677] "When a pedestrian appears ahead, what kind of sensor data should be used, and how should it be analyzed, to safely slow down?"

[0678] In this way, the system enhances safety in the operation of autonomous vehicles and can respond quickly and appropriately to changes in the environment.

[0679] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0680] Step 1:

[0681] The server receives video data from the camera as input. The video data is analyzed in real time using an AI algorithm to detect abnormal behavior. Specifically, an image recognition model is used to compare the data with pre-trained abnormal patterns to determine whether an abnormality has occurred. As a result, information regarding the abnormal behavior is output.

[0682] Step 2:

[0683] The server receives acoustic data from an acoustic data collection device as input. This data is analyzed by an AI model, and specific acoustic patterns are identified. For example, speech recognition technology is used to detect sudden braking sounds or warning sounds. The detected acoustic patterns are then output and used to evaluate the safety of the system.

[0684] Step 3:

[0685] The server receives data from a group of sensors as input. Physical changes such as temperature, humidity, and vibration are evaluated in real time, and calculations are performed to determine abnormal environmental conditions. By integrating the sensor information and analyzing deviations from normal patterns, anomalies can be identified. As a result, anomaly information regarding the environmental condition is output.

[0686] Step 4:

[0687] The server assesses the vehicle's risk level based on the abnormal behavior, sound patterns, and environmental conditions obtained in steps 1-3. It calculates a risk score based on this information and determines the necessary countermeasures. This includes specific measures such as whether an emergency stop is necessary or whether the route should be changed. The risk assessment results and countermeasures are then output.

[0688] Step 5:

[0689] Based on the determined response, the server outputs control signals to the autonomous vehicle's control system to adjust its operation. These signals may include emergency stop orders or new route instructions. These control signals encourage a rapid response from the vehicle and ensure safety.

[0690] Step 6:

[0691] The device receives alert information sent from the server and notifies the user. The notification includes security-related information and action suggestions to help the user take prompt action. This allows the user to understand the situation and take appropriate action as needed.

[0692] An example of a prompt for the generated AI model is, "When a pedestrian appears ahead, what sensor data should be used and how should it be analyzed to safely slow down?"

[0693] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0694] The system according to the present invention is an advanced monitoring and management system that not only provides a safe and effective play environment but also recognizes the user's emotional state and combines corresponding functions. By combining an emotion engine with a basic system that acquires environmental data using a camera, an acoustic data collection device, and a group of sensors, the system analyzes the user's emotional state in real time.

[0695] Specifically, the server captures the user's facial expressions and voice tone through camera footage and audio data, and the emotion engine uses this information to recognize the user's emotions (joy, anger, fear, etc.). The emotion engine is designed to identify a wide range of emotions by utilizing AI-powered machine learning algorithms.

[0696] The server uses the collected emotional information to determine appropriate countermeasures necessary to ensure the safety of children in the play environment. In this process, emotional information is taken into account in addition to the usual risk assessment, and suggestions for rest to alleviate excessive stress or recommendations for appropriate activities to increase the variety of play are made.

[0697] For example, if a child shows signs of stress during play, the server suggests a relaxing activity and sends an alert to the device saying, "Let's take a break and read a picture book." If the child continues to play and enjoys themselves, the server may suggest, "Let's try a new athletic activity," to further enhance positive emotions.

[0698] By integrating emotional information with conventional environmental data, users such as parents and administrators can implement more appropriate and emotionally sensitive safety measures for children. This provides a safe and comfortable play environment, promoting the healthy development and emotional growth of children.

[0699] The following describes the processing flow.

[0700] Step 1:

[0701] The server acquires user facial expression data and voice data from the camera and sound acquisition devices. This collects raw data that forms the basis of emotions.

[0702] Step 2:

[0703] The server preprocesses the acquired data, performing noise reduction and conversion to the required format. This prepares the data for analysis.

[0704] Step 3:

[0705] The server inputs pre-processed data into the emotion engine to identify the user's emotional state. In this process, a trained AI model identifies emotions such as joy, sadness, and anger.

[0706] Step 4:

[0707] The server aggregates a series of emotional data based on the analysis results of the emotion engine and evaluates the user's stress level and emotional state.

[0708] Step 5:

[0709] The server integrates emotional data and environmental data (abnormal behavior, weather changes, etc.) to reassess the level of risk in the play environment. Based on this assessment, it determines the priority of countermeasures.

[0710] Step 6:

[0711] The server will formulate appropriate countermeasures based on the evaluation results. For example, if emotional data indicates a high-stress state, it will suggest "moving to a rest area."

[0712] Step 7:

[0713] The server sends alert information, including the formulated countermeasures, to the device and notifies parents or administrators.

[0714] Step 8:

[0715] The device displays received alerts to the user visually and audibly, prompting them to take specific action.

[0716] Step 9:

[0717] Based on alert information from the device, the user takes the most appropriate action according to the child's situation. Specifically, they follow the suggestions to guide the child into a safe and positive play environment.

[0718] (Example 2)

[0719] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0720] Existing safety management systems in play environments have struggled to adequately understand and respond to the psychological state of users. Therefore, it has been difficult to implement safety measures that directly address the stress and anxiety experienced by children in particular, and a flexible approach that responds to users' emotional needs is required.

[0721] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0722] In this invention, the server includes means for processing video information acquired using a camera and analyzing the user's emotional state, means for processing acoustic information acquired using an acoustic collection device and analyzing emotions from voice tone, and means for acquiring and analyzing physical environment data using a group of sensors. This enables immediate and appropriate safety measures and action suggestions in accordance with the user's psychological state.

[0723] A "filming device" is a device used to acquire video information and is responsible for capturing the user's movements and facial expressions.

[0724] An "acoustic data acquisition device" is a device used to acquire acoustic information, and its role is to record the user's voice and ambient sounds.

[0725] A "sensor group" is a collection of multiple sensor devices that acquire physical environmental data, such as temperature, motion, and location information.

[0726] "Emotional state" refers to information that indicates the user's psychological and emotional state, and includes psychological elements such as joy, anger, fear, and sadness.

[0727] "Psychological state" refers to the mental and emotional state a user is experiencing, and is a concept that includes various emotions such as stress, relief, and tension.

[0728] A "suggestion alert" is information that the server uses to notify the user of recommended actions, such as relaxing activities or challenging new activities.

[0729] "Environmental data" refers to information that describes the physical state of the environment, and includes data such as temperature, humidity, sound level, and brightness.

[0730] This invention provides a safe and effective play environment, as well as an advanced monitoring and management system for recognizing and responding to the emotional state of users. This system is realized by collecting and analyzing data in real time using a camera, an acoustic data collection device, and a group of sensors.

[0731] The server has the function of processing video information acquired from the camera. It analyzes the user's facial expressions from the video information and recognizes their emotional state in real time. This makes it possible to accurately determine the user's emotions, such as joy, anger, sadness, and happiness. The server also has the function of processing acoustic data acquired from the acoustic collection device, and analyzes the tone of voice to estimate the user's emotions. This uses an AI-based generative model and machine learning algorithms. Furthermore, the server analyzes physical environment data acquired from the sensor group and takes the surrounding environmental conditions into consideration, thereby more comprehensively evaluating the user's psychological state.

[0732] Based on the results of the emotion analysis performed by the server, appropriate countermeasures to meet the user's current emotional needs are quickly determined. This includes emotionally sensitive suggestions such as selecting a play environment and suggesting breaks to reduce stress. For example, if the user is experiencing excessive stress, the server may suggest relaxing activities such as "Let's take a break and read a picture book." If the user is actively engaged, it may suggest actions such as "Let's try a new athletic activity."

[0733] These suggestions are communicated to the user via the device. The device displays the suggestions intuitively through visual and auditory means, making them easy for the user to understand. This allows the user to choose an appropriate action based on their emotional state.

[0734] For example, by inputting a prompt such as, "Based on the user's emotional data, suggest relaxing activities," into the AI ​​model, the system can accurately provide suggestions tailored to the user's individual needs. This enables the provision of a safer and more comfortable play environment.

[0735] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0736] Step 1:

[0737] The server acquires data from the camera and sound acquisition equipment. It receives video data, including the user's facial expressions, from the camera and audio data from the microphone in real time. This provides input to understand the user's current state. Specifically, the camera periodically captures the user's face and sends this as digital video data to the server. The audio data captures the tone and volume of the user's voice and is sent for acoustic analysis.

[0738] Step 2:

[0739] The server analyzes the acquired video and audio data using a generating AI model. From the video data, a machine learning algorithm is used to recognize facial expressions and extract the user's emotional state. From the audio data, additional emotional insights are obtained through voice tone analysis. This data processing outputs results that show the user's overall emotional state. Specifically, the AI ​​model analyzes facial features and generates emotion labels such as anger and joy. From the audio data, it analyzes the pitch and tempo of the voice to derive indicators of stress and calmness.

[0740] Step 3:

[0741] The server takes the analyzed emotional data into account to generate suggestions for a play environment suitable for the user. The main inputs here are the results of the emotional analysis and known safety standards. The server integrates this information to determine appropriate countermeasures and activities. Specifically, if relaxation is determined to be needed, it will recommend activities in a quiet environment. Conversely, if high energy levels are determined, it will suggest vigorous athletic activities.

[0742] Step 4:

[0743] The device notifies the user of action suggestions received from the server. This uses both visual and auditory information to ensure intuitive understanding. The device presents prompts to the user in natural language to encourage action. This notification is output as feedback to the user and influences which activity they choose. For example, the device displays "Let's relax by reading a picture book" on its screen and also conveys the same message aloud.

[0744] Step 5:

[0745] The user receives suggestions from the device and selects an action that aligns with their emotional state. In this final step, the user's choices are returned to the system as feedback, which is used for further analysis and system improvement. Specifically, the user checks an alert and begins taking action according to the suggested activity. This allows the user to continue their activities in a safer and more comfortable environment.

[0746] (Application Example 2)

[0747] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0748] Traditional in-store shopping experiences often lack individuality because they don't adequately consider customers' emotional states. This can lead to decreased customer satisfaction and purchase intent. Therefore, it's crucial to improve the quality of the shopping experience by providing more personalized product and service suggestions based on customer emotions.

[0749] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0750] In this invention, the server includes means for processing video data acquired using a camera and detecting abnormal behavior, means for processing acoustic data acquired using an acoustic collection device and identifying specific acoustic patterns, and means for detecting physical changes using a group of sensors and determining abnormal environmental conditions. This enables real-time analysis of the user's emotional state and the suggestion of suitable products or services based on those emotions.

[0751] A "recording device" is a device used to acquire visual information, such as the conditions in the environment or the user's facial expressions.

[0752] An "acoustic acquisition device" is a device that acquires voice and ambient sounds, and is used to analyze the tone of the user's voice and the surrounding acoustic patterns.

[0753] A "sensor cluster" is a collection of various sensors installed to detect diverse physical changes and is used to determine abnormal environmental conditions.

[0754] "Real-time analysis of emotional state" means instantly processing data such as the user's facial expressions and voice to identify their emotions at that moment.

[0755] "Suggesting suitable products or services" refers to the act of selecting and presenting products or services that match the emotions of the analyzed user.

[0756] The system for realizing this invention is a mechanism that improves the in-store purchasing experience by analyzing the user's emotional state in real time and suggesting products and services based on those emotions.

[0757] The server uses a camera, sound acquisition device, and sensor array to collect data such as the user's facial expressions, voice tone, and surrounding environmental conditions. Specifically, the camera records the user's facial movements and expressions, while the microphone records the user's voice and background sounds. This data is integrated with environmental data collected through the sensor array and transmitted to the server.

[0758] On the server, the collected data is processed using a cloud-based sentiment analysis engine (e.g., Amazon Rekognition or Google Cloud Vision API). This analyzes and identifies the emotions the user is experiencing. Based on this analysis, the server selects and visually presents products and services that are suitable for the user.

[0759] Furthermore, it is possible to utilize generative AI models to create prompt messages that respond to the user's emotions. These prompt messages would serve as information that encourages specific actions from the user, such as, "Identify the user's emotions from the facial expression analysis data and create a prompt to customize the shopping experience based on those emotions."

[0760] As a concrete example, suppose a user is browsing in a store and the server recognizes that the user has made a surprised expression upon seeing a particular product. In this case, the server will encourage a purchase by displaying special offers or new information related to that product on the user's smartphone or display.

[0761] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0762] Step 1:

[0763] The server collects user facial and audio data using a camera and audio acquisition device. Inputs are camera video and audio from a microphone. These inputs are acquired in real time, and initial processing is performed using facial recognition technology and audio analysis techniques. Outputs are the user's basic facial features and audio patterns.

[0764] Step 2:

[0765] The server collects environmental data from a group of sensors. The input consists of physical data such as temperature, humidity, and light intensity within the store. This sensor data is aggregated to evaluate the environmental conditions. The output is an evaluation value indicating abnormal environmental conditions.

[0766] Step 3:

[0767] The server integrates the collected facial expression data, acoustic data, and environmental data to perform emotion analysis. The input consists of the data obtained in Step 1 and Step 2. This data is integrated, and the emotion analysis engine is used to identify the user's emotions. The output is the user's emotional state (e.g., joy, surprise, anger).

[0768] Step 4:

[0769] The server selects suitable products and services based on the results of sentiment analysis. The input is information about the user's emotional state. It references a database to extract products and services that match the emotion. The output is information about the suggested products and services.

[0770] Step 5:

[0771] The terminal receives suggestion information from the server and presents it visually to the user. The input is suggestion information about products and services sent from the server. This information is displayed on a display or smartphone screen to attract the user's attention. The output is the visual information provided to the user.

[0772] Step 6:

[0773] The server utilizes a generative AI model to create emotion-based prompts. The input consists of the user's emotional state and suggested product information. Based on this data, the generative AI model generates prompts. The output is a prompt that encourages the user to take a specific action.

[0774] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0775] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0776] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0777] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0778] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. In the upper and lower directions of the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. Also, the upper side of the concentric circles is where "pleasant" emotions are located, and the lower side is where "unpleasant" emotions are located. In this way, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0779] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0780] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0781] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0782] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0783] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0784] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0785] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0786] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0787] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0788] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0789] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0790] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0791] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0792] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0793] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0794] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0795] The following is further disclosed regarding the embodiments described above.

[0796] (Claim 1)

[0797] A means for processing video data acquired using a camera to detect abnormal behavior,

[0798] A means for processing acoustic data acquired using an acoustic collection device and identifying a specific acoustic pattern,

[0799] A means for detecting physical changes using a group of sensors and determining abnormal environmental conditions,

[0800] A means for evaluating the degree of risk and determining countermeasures based on the information of the abnormal behavior and abnormal condition,

[0801] A means of sending an alert to a notification device using the aforementioned countermeasure,

[0802] A system that includes this.

[0803] (Claim 2)

[0804] The system according to claim 1, which displays a suggestion for action based on an alert in response to sending an alert to the notification device.

[0805] (Claim 3)

[0806] The system according to claim 1, comprising a function to select a suitable play area and visually present the area to the user by integrally analyzing data obtained from the aforementioned imaging device, sound collection device, and sensor group.

[0807] "Example 1"

[0808] (Claim 1)

[0809] A means for processing video data acquired using a camera to detect abnormal behavior,

[0810] A means for processing acoustic data acquired using an acoustic collection device and identifying a specific acoustic pattern,

[0811] A means for detecting physical changes using a group of sensors and determining abnormal environmental conditions,

[0812] A means for evaluating the degree of risk and determining countermeasures based on the information of the abnormal behavior and abnormal condition,

[0813] A means for analyzing the aforementioned data using a generating AI model,

[0814] A means of sending an alert to a notification device using the aforementioned countermeasure,

[0815] A system that includes this.

[0816] (Claim 2)

[0817] The system according to claim 1, which displays a suggestion for action based on an alert in response to sending an alert to the notification device.

[0818] (Claim 3)

[0819] The system according to claim 1, comprising a function to select a suitable play area and visually present the area to the user by integrally analyzing data obtained from the aforementioned imaging device, sound collection device, and sensor group.

[0820] "Application Example 1"

[0821] (Claim 1)

[0822] A means for processing video data acquired using a camera to detect abnormal behavior,

[0823] A means for processing acoustic data acquired using an acoustic collection device and identifying a specific acoustic pattern,

[0824] A means for detecting physical changes using a group of sensors and determining abnormal environmental conditions,

[0825] A means for evaluating the degree of risk and determining countermeasures based on the information of the abnormal behavior and abnormal condition,

[0826] A means of sending an alert to a notification device using the aforementioned countermeasure,

[0827] Automatic control means that monitors the surrounding environment and performs emergency stops or route changes to enhance the safety of vehicle operation,

[0828] A system that includes this.

[0829] (Claim 2)

[0830] The system according to claim 1, which, in response to sending an alert to the notification device, displays a suggested action based on the alert and presents safety measures related to vehicle operation.

[0831] (Claim 3)

[0832] The system according to claim 1, comprising a function to select a suitable vehicle route by comprehensively analyzing data obtained from the aforementioned imaging device, acoustic collection device, and sensor group, and to present the route to an automatic control device.

[0833] "Example 2 of combining an emotion engine"

[0834] (Claim 1)

[0835] A means of processing video information acquired using a camera and analyzing the user's emotional state,

[0836] A method for processing acoustic information acquired using an acoustic collection device and analyzing emotions from voice tone,

[0837] A means for acquiring and analyzing physical environmental data using a group of sensors,

[0838] A means for evaluating the user's psychological state and determining appropriate countermeasures based on the aforementioned emotion analysis information and environmental data,

[0839] A means for sending an action suggestion alert to a notification device based on the aforementioned countermeasures,

[0840] A system that includes this.

[0841] (Claim 2)

[0842] The system according to claim 1, which provides the user with suggestions for relaxation or activity in response to an action suggestion alert transmitted to the notification device.

[0843] (Claim 3)

[0844] The system according to claim 1, which performs modifications and adjustments to the gaming environment in consideration of the user's emotions by comprehensively analyzing the data obtained from the aforementioned shooting device, sound collection device, and sensor group.

[0845] "Application example 2 when combining with an emotional engine"

[0846] (Claim 1)

[0847] A means for processing video data acquired using a camera to detect abnormal behavior,

[0848] A means for processing acoustic data acquired using an acoustic collection device and identifying a specific acoustic pattern,

[0849] A means for detecting physical changes using a group of sensors and determining abnormal environmental conditions,

[0850] A means for evaluating the degree of risk and determining countermeasures based on the information of the abnormal behavior and abnormal condition,

[0851] A means of sending an alert to a notification device using the aforementioned countermeasure,

[0852] A means of analyzing a user's emotional state in real time and suggesting suitable products or services based on those emotions,

[0853] A system that includes this.

[0854] (Claim 2)

[0855] The system according to claim 1, which displays a suggestion for action based on an alert in response to sending an alert to the notification device.

[0856] (Claim 3)

[0857] The system according to claim 1, comprising a function to select a suitable purchasing experience and visually present that experience to the user by comprehensively analyzing data obtained from the aforementioned imaging device, sound collection device, and sensor group. [Explanation of symbols]

[0858] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for processing video data acquired using a camera to detect abnormal behavior, A means for processing acoustic data acquired using an acoustic collection device and identifying a specific acoustic pattern, A means for detecting physical changes using a group of sensors and determining abnormal environmental conditions, A means for evaluating the degree of risk and determining countermeasures based on the information of the abnormal behavior and abnormal condition, A means of sending an alert to a notification device using the aforementioned countermeasure, A system that includes this.

2. The system according to claim 1, which displays a suggested action based on an alert in response to sending an alert to the notification device.

3. The system according to claim 1, comprising a function to select a suitable play area and visually present the area to the user by integrally analyzing data obtained from the aforementioned imaging device, sound collection device, and sensor group.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A