system

JP7912632B1Active Publication Date: 2026-08-28SOFTBANK GROUP CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2025044951
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2026-08-28
Estimated Expiration
2045-03-19

AI Technical Summary

Benefits of technology

【0005】 本発明は、画像認識AIを用いて子供の表情や行動を分析し、いじめの可能性を検出する。さらに、検出したいじめの可能性に基づいて警告を発し、保護者や教育関係者に対して対策を提案する。これにより、いじめの早期発見と対策が可能となる。また、子供の表情や行動の分析結果を保護者や教育関係者に提供し、いじめの可能性の検出精度を向上させるための学習機能を含む。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007912632000001_ABST
    Figure 0007912632000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A system that includes means for analyzing a child's facial expressions and behavior using image recognition AI to detect the possibility of bullying, means for issuing a warning based on the detected possibility of bullying, and means for proposing countermeasures to parents and educators who have received the warning.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background Art]

[0002] Patent Document 1 discloses a persona chatbot control method executed by at least one processor, the method comprising: receiving a user utterance; adding the user utterance to a prompt including an instruction associated with a description of a character of the chatbot; encoding the prompt; and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior Art Document] [Patent Document]

[0003] [Patent Document 1] Japanese Unexamined Patent Application Publication No. 2022-180282 [Summary of the Invention] [Problem to be Solved by the Invention]

[0004] Currently, the problem of bullying in schools and homes is becoming more serious, and early detection and countermeasures therefor are required. However, it is difficult to directly witness the scene of bullying, and it is also difficult for the victim themselves to report bullying. For this reason, early detection and countermeasures for bullying are difficult. [Means for Solving the Problem]

[0005] The present invention uses image recognition AI to analyze children's facial expressions and behaviors, and detects the possibility of bullying. Furthermore, the present invention issues a warning based on the detected possibility of bullying, and proposes countermeasures to guardians and education personnel. This enables early detection and countermeasures for bullying. The present invention also includes a learning function for providing analysis results of children's facial expressions and behaviors to guardians and education personnel, and improving the detection accuracy of the possibility of bullying. [Brief explanation of the drawing]

[0006] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Embodiment 1 of Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1 of Form Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2 of Embodiment 2. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2 of Form Example 2. [Figure 15]It is a sequence diagram showing the processing flow of the data processing system in Example 3 of Embodiment 3. [Figure 16] It is a sequence diagram showing the processing flow of the data processing system in Application Example 3 of Embodiment 3. [Figure 17] It is a sequence diagram showing the processing flow of the data processing system in Example 1 of Embodiment 1 when an emotion engine is combined. [Figure 18] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1 of Embodiment 1 when an emotion engine is combined. [Figure 19] It is a sequence diagram showing the processing flow of the data processing system in Example 2 of Embodiment 2 when an emotion engine is combined. [Figure 20] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 of Embodiment 2 when an emotion engine is combined. [Figure 21] It is a sequence diagram showing the processing flow of the data processing system in Example 3 of Embodiment 3 when an emotion engine is combined. [Figure 22] It is a sequence diagram showing the processing flow of the data processing system in Application Example 3 of Embodiment 3 when an emotion engine is combined. [Figure 23] It is a sequence diagram showing the processing flow of the data processing system in another embodiment. MODE FOR CARRYING OUT THE INVENTION

[0007] Hereinafter, an example embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0008] First, terms used in the following description will be explained.

[0009] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of a plurality of arithmetic units. Further, the processor may be one type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of the arithmetic unit include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), and TPU (TENSOR PROCESSING UNIT (registered trademark)).

[0010] In the following embodiments, the labeled RAM (Random Access Memory) is a memory that temporarily stores information, and is used as a working memory by the processor.

[0011] In the following embodiments, the labeled storage is one or more non-volatile storage devices that store various programs, various parameters and the like. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0012] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F manages communication between a plurality of computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), and Bluetooth (registered trademark).

[0013] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0014] [First Embodiment]

[0015] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0016] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0017] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0018] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0019] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0020] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0021] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0022] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0023] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0024] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0025] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0026] Next, the identification process performed by the identification processing unit 290 of the data processing device 12 will be described.

[0027] "Example of form 1"

[0028] One embodiment of the present invention involves equipping a camera or smartphone, or other imaging device, with image recognition AI for use in schools or homes. This image recognition AI analyzes a child's facial expressions and behavior in real time to detect the possibility of bullying. For example, it detects the possibility of bullying if it determines that a child's facial expression indicates fear or sadness, or if it determines that the child is being subjected to aggressive behavior from other children.

[0029] "Example of form 2"

[0030] Based on the detected potential bullying, warnings are issued to parents and educators. These warnings are sent via email, app notifications, etc. The warnings include details such as the date, time, and location where the potential bullying was detected, as well as details of the child's facial expressions and behavior. Specific countermeasures are also suggested. For example, parents may be advised to take measures to encourage communication with their children, and educators may be advised to strengthen school monitoring systems and offer counseling.

[0031] "Example of form 3"

[0032] Furthermore, the system provides parents and educators with analysis results of children's facial expressions and behavior. This allows parents and educators to understand the child's situation more concretely and take appropriate action. It also includes a learning function to improve the accuracy of detecting potential bullying. This learning function allows the image recognition AI to self-learn based on new data and improve detection accuracy.

[0033] The following describes the processing flow for each example of the form.

[0034] "Example of form 1"

[0035] Step 1: Equipate image recognition AI into cameras, smartphones, and other imaging devices for use in schools and homes.

[0036] Step 2: Capture the child's facial expressions and actions in real time through the camera and send them to the image recognition AI.

[0037] Step 3: The image recognition AI analyzes the transmitted image data and detects the possibility of bullying from the child's facial expressions and behavior.

[0038] "Example of form 2"

[0039] Step 1: If the image recognition AI detects a potential bullying incident, it issues a warning to parents and educators.

[0040] Step 2: The warning will be sent via email, app notification, etc.

[0041] Step 3: The warning will include details such as the date and location where the bullying was detected, as well as the child's facial expressions and behavior.

[0042] Step 4: Specific measures are proposed. For example, parents may be advised to take steps to encourage communication with their children, and educators may be advised to strengthen school supervision systems and provide counseling.

[0043] "Example of form 3"

[0044] Step 1: Provide parents and educators with the results of the analysis of the child's facial expressions and behavior.

[0045] Step 2: Parents and educators will be able to understand the child's situation more concretely and take appropriate action.

[0046] Step 3: Includes a learning function to improve the accuracy of detecting potential bullying.

[0047] Step 4: This learning function allows the image recognition AI to self-learn based on new data and improve detection accuracy.

[0048] (Example 1)

[0049] Next, we will describe Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0050] In modern schools and homes, children are increasingly likely to encounter bullying, but there is a challenge in detecting signs of bullying early and taking appropriate measures. In particular, there is a need for technology that can monitor changes in children's facial expressions and behavior in real time and quickly detect the possibility of bullying.

[0051] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0052] In this invention, the server includes means for collecting video data using a camera, means for transmitting the collected video data via communication means, and means for analyzing a child's facial expressions and behavior using image recognition technology to detect the possibility of bullying. This makes it possible to detect signs of bullying early and take appropriate measures.

[0053] A "recording device" is a device used to collect video data, and includes devices such as cameras and smartphones.

[0054] "Video data" refers to digital data containing visual information collected by a camera or camera.

[0055] "Communication methods" refer to the technologies and protocols used to send and receive data, and include methods such as the internet and wireless communication.

[0056] "Image recognition technology" is a technology that allows computers to analyze image data and identify specific patterns or features.

[0057] "Analysis of facial expressions and behavior" is the process of using image recognition technology to analyze a child's facial expressions and body movements to identify specific emotions and behaviors.

[0058] "Detecting the possibility of bullying" is a process of determining the likelihood of bullying occurring based on analyzed facial expression and behavioral data.

[0059] "Means of issuing warnings" refer to methods or devices used to alert those involved when the possibility of bullying is detected.

[0060] "Means of proposing countermeasures" refer to methods and devices for presenting specific countermeasures to parents and educators based on the detected possibility of bullying.

[0061] "Learning function" refers to a function that allows the system to improve the accuracy of its analysis based on past data, and includes machine learning algorithms.

[0062] This invention is a system for use in schools and homes that analyzes children's facial expressions and behavior in real time to detect potential bullying. The system consists of a camera, communication means, and image recognition technology.

[0063] The terminal uses cameras, smartphones, and other recording devices to collect video data of children in real time. The collected video data is transmitted to a server via a secure communication method. The communication method uses the internet or wireless communication technology.

[0064] The server uses image recognition technology to analyze the received video data. This technology leverages generative AI models built using machine learning frameworks such as TENSORFLOW® and PyTorch. Specifically, it uses OpenCV to extract facial features and analyze changes in facial expressions. It also applies motion recognition algorithms to detect aggressive behavior.

[0065] Users receive analysis results from the server and are warned if bullying is suspected. Furthermore, the system suggests specific measures to parents and educators based on the detected results. This makes it possible to detect signs of bullying early and take appropriate action.

[0066] For example, a user might monitor children's behavior through a camera installed in the classroom. The AI ​​would warn of potential bullying if it detects that a child frequently looks frightened or if it detects behavior such as another child unilaterally pushing or shoving. Examples of prompts to input into the generating AI model could be, "What should I do if a child looks frightened?" or "Please tell me how to deal with aggressive behavior that is detected."

[0067] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0068] Step 1:

[0069] The device uses a camera or smartphone to collect video data of children in real time. The input is video data obtained from the recording device. Specifically, the device ensures high-resolution video so that facial expressions and actions can be clearly identified. The output is the collected video data.

[0070] Step 2:

[0071] The terminal sends the collected video data to the server. The input is the video data obtained in step 1. Specifically, the terminal uses a secure communication protocol (e.g., HTTPS) to send the data and protect privacy. The output is the video data sent to the server.

[0072] Step 3:

[0073] The server analyzes the received video data using image recognition technology. The input is the video data sent in step 2. Specifically, the server utilizes generative AI models using TensorFlow and PyTorch, extracts facial features using OpenCV, and analyzes changes in facial expressions. It also applies a motion recognition algorithm to detect aggressive behavior. The output is the analyzed facial expression and behavior data.

[0074] Step 4:

[0075] The server determines the possibility of bullying based on the analysis results and issues a warning if necessary. The input is the analysis data obtained in step 3. Specifically, if the server determines that bullying is possible, it generates a warning message. The output is the warning message.

[0076] Step 5:

[0077] The user receives a warning from the server and considers appropriate action. The input is the warning message generated in step 4. Specifically, the user checks the warning on their smartphone or PC and takes action, such as meeting with the child if necessary. The output is the user's response.

[0078] (Application Example 1)

[0079] Next, we will describe Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server," and the smart device 14 will be referred to as a "terminal."

[0080] In modern society, bullying of children at school and at home is a serious problem. However, it is not easy to detect signs of bullying early and take appropriate measures. In particular, there is a lack of means to monitor children's emotions and behavior in real time and quickly detect the possibility of bullying, which often leads to delays in early detection and response to bullying. To solve this problem, a system is needed that analyzes children's emotions and behavior in real time, immediately detects signs of bullying, and notifies those involved.

[0081] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0082] In this invention, the server includes means for analyzing a child's emotions and actions using image recognition technology to detect the possibility of bullying, means for monitoring a child's emotions and actions in real time using a mobile information terminal, and means for sending a notification when signs of bullying are detected. This makes it possible to detect signs of bullying in a child early and quickly notify the relevant parties.

[0083] "Image recognition technology" is a technique that analyzes image data acquired using cameras and sensors to identify specific patterns and features.

[0084] "Children's emotions and actions" refer to the psychological state and behaviors that can be interpreted from a child's facial expressions and body movements.

[0085] "Possible bullying" refers to signs that indicate a child may be experiencing aggressive behavior or psychological pressure from others.

[0086] A "portable information terminal" is an electronic device that is portable and capable of processing information, such as a smartphone or tablet.

[0087] "Real-time monitoring" means observing and analyzing the ongoing situation immediately.

[0088] "Sending a notification" means immediately informing relevant parties of specific information.

[0089] The system for implementing this invention mainly consists of a server and a mobile information terminal. The server utilizes image recognition technology to analyze children's emotions and actions and detect the possibility of bullying. Specifically, the server uses image recognition libraries such as TensorFlow and OpenCV to process image data acquired from cameras and sensors in real time. This makes it possible to identify signs of bullying from children's facial expressions and actions.

[0090] Mobile information devices refer to smartphones, tablets, and other similar devices. These devices work in conjunction with a server to monitor children's emotions and behavior in real time. The devices receive analysis results sent from the server, and if signs of bullying are detected, they immediately send notifications to parents and educators. Notifications are sent in the form of push notifications or email.

[0091] As a concrete example, suppose a smartphone is used to film children playing in a school classroom. This video data is sent to a server, which uses image recognition technology to analyze the children's facial expressions and actions. For example, if the server detects that a child is frightened or being subjected to aggressive behavior from other children, it immediately sends a notification to parents or educators.

[0092] An example of a prompt for a generative AI model is, "Based on this video, analyze whether the child may be being bullied." Using this prompt, the AI ​​model can analyze the video data and assess the likelihood of bullying.

[0093] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0094] Step 1:

[0095] The device uses its camera to capture images of children in real time. The captured video data is temporarily stored on the device. This video data becomes the input for the next processing step.

[0096] Step 2:

[0097] The device sends the captured video data to the server. The server takes the received video data as input and begins analysis using image recognition technology. Specifically, it uses libraries such as TensorFlow and OpenCV to analyze the facial expressions and movements of children in the video and extract features. The results of this analysis become the input for the next processing step.

[0098] Step 3:

[0099] The server uses a generative AI model to evaluate the possibility of bullying based on the extracted features. The prompt message, "Based on this video, analyze whether the child may be being bullied," is input to the AI ​​model, and signs of bullying are detected. This evaluation result becomes the input for the next processing step.

[0100] Step 4:

[0101] If signs of bullying are detected, the server sends the assessment results to the device. Based on the received assessment results, the device sends a notification to parents or educators. The notification is sent via push notification or email and informs them of the possibility of bullying. This notification is the final output.

[0102] (Example 2)

[0103] Next, we will describe Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0104] While early detection and rapid response to bullying are crucial, traditional methods often miss early signs of bullying, making it difficult to implement appropriate measures. In particular, there is a lack of mechanisms to monitor changes in children's facial expressions and behavior in real time and promptly notify relevant parties.

[0105] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0106] In this invention, the server includes means for collecting video and audio data, means for analyzing the collected data and analyzing the child's facial expressions and behavior, means for evaluating the possibility of bullying based on the analysis results, means for generating warning messages using a generative AI model, means for sending the generated warning messages to parents and educators, and means for proposing specific countermeasures to parents and educators. This makes it possible to detect signs of bullying early and notify relevant parties quickly.

[0107] "Video data" refers to visual information acquired using cameras and other recording devices, and is fundamental information for analyzing children's facial expressions and behavior.

[0108] "Audio data" refers to auditory information acquired using sound devices such as microphones, and is fundamental information for analyzing conversation content and tone.

[0109] "Analysis" refers to the process of processing collected video and audio data to analyze children's facial expressions and behavior.

[0110] A "generative AI model" refers to a program that uses artificial intelligence technology to perform natural language processing and generate warning messages based on the analysis results.

[0111] A "warning message" is a notification generated when potential bullying is detected, and it contains information to inform parents and educators about the situation.

[0112] "Parents and educators" refers to adults involved in the upbringing and education of children, and who have the responsibility to take appropriate action against signs of bullying.

[0113] "Specific measures" refer to action guidelines proposed when the possibility of bullying is detected, outlining the countermeasures that parents and educators should take.

[0114] A description of embodiments for carrying out this invention will be given.

[0115] The user generates a program for an bullying detection system. This program uses cameras and sensors installed within the school to monitor children's facial expressions and behavior in real time. The server collects video and audio data from these devices. Specifically, it processes video data acquired using cameras with image recognition software (e.g., OpenCV) to analyze children's facial expressions. It also processes audio data acquired using microphones with speech analysis software (e.g., Google Cloud Speech-to-Text) to analyze conversation content and tone.

[0116] The server uses a generative AI model (e.g., OpenAI®'s GPT-4®) based on the analysis results to assess the likelihood of bullying. The generative AI model converts the analysis data into natural language and calculates a score indicating the likelihood of bullying. If signs of bullying are detected, the server generates a warning message. This warning message includes the date, time, and location where the bullying was detected, as well as details of the child's facial expressions and behavior.

[0117] The device sends generated warning messages to parents and educators via email or app notifications. This allows stakeholders to quickly understand the situation and take appropriate action. For example, if a specific child exhibits aggressive behavior towards another child in the schoolyard at 3 PM on a given day, a warning detailing the incident will be sent. An example of a prompt message might be: "Potential bullying detected in the schoolyard. Date and time: 3 PM, October 5, 2023. Location: Schoolyard. The child appeared angry and acted aggressively. What actions would you suggest to the parents?"

[0118] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0119] Step 1:

[0120] The server collects video and audio data from cameras and microphones installed within the school. It receives real-time visual and auditory information as input. This data is used as foundational information for monitoring children's facial expressions and behavior. Specifically, the server receives data streams from each device and stores them in a database.

[0121] Step 2:

[0122] The server processes the collected video data using image recognition software (e.g., OpenCV) to analyze the children's facial expressions. The video data collected in step 1 is used as input. For data processing, a face detection algorithm is applied to extract facial features. The output is the analysis results for each child's facial expression. Specifically, the server detects faces in each video frame and identifies expressions such as smiles and anger.

[0123] Step 3:

[0124] The server processes the collected audio data using speech analysis software (e.g., Google Cloud Speech-to-Text) to analyze the conversation content and tone. The audio data collected in step 1 is used as input. Data processing involves converting the audio to text and performing sentiment analysis. The output is the analysis results regarding the conversation content and tone. Specifically, the server converts the audio data to text and detects aggressive words and tones.

[0125] Step 4:

[0126] Based on the analysis results from steps 2 and 3, the server uses a generative AI model (e.g., OpenAI's GPT-4) to assess the likelihood of bullying. The inputs used are the analysis results of facial expressions and voice. The data calculation involves the generative AI model calculating a score indicating the likelihood of bullying. The output is an evaluation result regarding the likelihood of bullying. Specifically, the server prompts the generative AI model with the analysis results and calculates the score.

[0127] Step 5:

[0128] The server generates warning messages using a generative AI model. It uses the bullying possibility assessment results obtained in step 4 as input. For data processing, it creates warning messages in natural language. The output is a warning message to be sent to parents and educators. Specifically, the server generates a message that includes details such as the date, time, location, facial expressions, and actions.

[0129] Step 6:

[0130] The device sends the warning message received from the server to parents and educators via email or app notification. It uses the warning message generated in step 5 as input. Its output is a notification to the relevant parties. Specifically, the device sends the message to the designated contacts to prompt a quick response.

[0131] (Application Example 2)

[0132] Next, we will describe application example 2 of form example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0133] In recent years, bullying among children in schools and homes has become a serious social problem. However, it is difficult to detect signs of bullying early and take appropriate measures. In particular, detecting the possibility of bullying from a child's facial expressions and behavior is a great burden for educators and parents. To solve this problem, a system is needed that can monitor children's behavior in real time and quickly detect signs of bullying.

[0134] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0135] In this invention, the server includes means for analyzing a child's facial expressions and behavior using image recognition technology to detect the possibility of bullying, means for monitoring a child's facial expressions and behavior in real time using a visual device, and means for analyzing signs of bullying using a generative AI model. This makes it possible to detect signs of bullying early and take appropriate measures quickly.

[0136] "Image recognition technology" is a technique that analyzes image data acquired using cameras and sensors to identify specific patterns and features.

[0137] "Children" refers to children of school age or those receiving education at home.

[0138] "Facial expression" refers to the emotions and reactions expressed through the movement of the facial muscles.

[0139] "Behavior" refers to actions and reactions that children take in their daily lives or in specific situations.

[0140] "Possible bullying" refers to signs of aggressive or inappropriate behavior towards others that can be inferred from a child's facial expressions and actions.

[0141] A "warning" is a notification or message issued to alert those involved when a potential bullying situation is detected.

[0142] "Guardian" refers to a child's parent or legal supervisor, who is responsible for the child's welfare and education.

[0143] "Education-related personnel" refers to individuals who hold positions related to the education and guidance of children in schools or educational institutions.

[0144] "Countermeasures" refers to specific actions or measures that should be taken when there is a possibility of bullying.

[0145] "Visual devices" refer to devices used to acquire and display images and videos.

[0146] A "generative AI model" refers to an algorithm or system that uses artificial intelligence technology to analyze data and generate specific patterns or results.

[0147] "Analysis" is the process of examining data in detail and deriving specific information or conclusions.

[0148] A "notification" is a message or alert used to convey specific information to relevant parties.

[0149] The system for carrying out this invention mainly consists of a server, a visual device, and a generative AI model. The server uses image recognition technology to analyze data on children's facial expressions and behavior acquired from the visual device and detect the possibility of bullying. The visual device is a device such as smart glasses and is responsible for monitoring children's behavior in real time. The generative AI model is built using frameworks such as TensorFlow or PyTorch and is used to analyze patterns in children's facial expressions and behavior.

[0150] The server receives video data transmitted from the visual device and analyzes the data using a generative AI model. If the analysis detects signs of bullying, the server generates a warning and sends a notification to parents and educators. This notification includes the date, time, and location of the detection, details of the child's facial expressions and behavior, and further specific countermeasures are suggested.

[0151] As a concrete example, imagine a school recess where a visual device monitors a group of students. If the generative AI model detects that a particular student frequently displays a sad expression, the server sends a notification to educators stating, "This may be bullying. Please investigate further." This notification allows educators to take prompt action.

[0152] An example of a prompt message for a generative AI model is: "Analyze the child's facial expression data and detect signs of bullying. If detected, output the date, time, location, and facial expression details, and suggest appropriate countermeasures."

[0153] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0154] Step 1:

[0155] The server receives video data transmitted in real time from the visual device. This video data includes the children's facial expressions and actions. The server preprocesses the video data and converts it into a format to which image recognition technology can be applied.

[0156] Step 2:

[0157] The server inputs pre-processed video data into a generative AI model. The generative AI model, built using TensorFlow, analyzes the children's facial expressions and behavior. The model extracts facial features and detects patterns that indicate the possibility of bullying. The analysis results in an output indicating whether or not bullying is likely.

[0158] Step 3:

[0159] Based on the output from the generating AI model, the server generates an alert if potential bullying is detected. This alert includes the date and time of detection, location, and details of the child's facial expressions and behavior. The server organizes this information and creates a notification message.

[0160] Step 4:

[0161] The server sends the created notification message to parents and educators. The notification is sent via email or app push notification. Recipients can check the notification and take appropriate action based on the child's situation.

[0162] Step 5:

[0163] After sending a notification, the server updates the training database for the generated AI model. Newly detected cases are added to the training data to improve the model's accuracy. This process continuously improves the accuracy of bullying detection.

[0164] (Example 3)

[0165] Next, we will describe Embodiment 3 of Embodiment Example 3. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0166] There is a need to accurately analyze children's emotional states and behavioral patterns to detect potential abnormal behavior at an early stage. However, conventional systems suffer from insufficient analytical accuracy, resulting in a lack of information for parents and educators to take appropriate action. Furthermore, the absence of self-learning capabilities based on analysis results makes it difficult to improve detection accuracy.

[0167] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 3 is realized by the following means.

[0168] In this invention, the server includes means for analyzing a child's emotional state and behavioral patterns using an image processing device, means for detecting the possibility of abnormal behavior based on the analyzed data, and means for providing the analysis results through an information display device. This enables more accurate analysis of a child's behavior, early detection of abnormal behavior, and rapid provision of information to parents and educators.

[0169] An "image processing device" is a device used to analyze image data and extract specific features.

[0170] "Children" refers to children attending educational institutions or minors.

[0171] "Emotional state" refers to the psychological state that can be inferred from a child's facial expressions and behavior.

[0172] "Behavioral patterns" refer to the tendencies of a series of actions and behaviors that a child exhibits.

[0173] "Abnormal behavior" refers to actions that deviate from normal behavioral patterns and includes signs of bullying and problematic behavior.

[0174] An "information display device" is a device used to visually display analysis results, and includes monitors, smartphones, and other similar devices.

[0175] The "self-learning function" is a feature that allows the system to automatically learn from newly collected data and improve the accuracy of its analysis.

[0176] A description of embodiments for carrying out this invention will be given.

[0177] The user builds a system to analyze children's emotional states and behavioral patterns. This system has the capability to analyze children's facial expressions and behavior in real time using an image processing device. Specifically, a camera installed on the terminal captures video of the children and sends the data to a server. The server preprocesses the received image data and extracts features using an image recognition AI model. This AI model is built using machine learning frameworks such as TensorFlow and PyTorch.

[0178] The server detects potential abnormal behavior based on extracted characteristic data. The detected results are provided to parents and educators via an information display device. This allows users to gain a concrete understanding of the child's situation and take appropriate action as needed.

[0179] For example, when a user monitors children's activities through cameras installed in school classrooms, the server receives image data in real time, and an AI model analyzes that data. The analysis results are provided to parents and educators through a dashboard.

[0180] An example of a prompt to be input to the generating AI model is, "Analyze the facial expressions and behavior of children in the classroom and create a report to detect signs of abnormal behavior." In response to this prompt, the AI ​​performs the specified task and provides the results. The specific processing flow in Example 3 is explained using Figure 15.

[0181] Step 1:

[0182] The device uses a camera to capture images of the children. The input is real-time video data. The device sends this video data to a server. Specifically, the camera continuously captures the facial expressions and movements of the children in the classroom.

[0183] Step 2:

[0184] The server preprocesses the video data received from the terminal. The input is the video data transmitted from the terminal. The server uses image processing libraries to adjust the image resolution and remove noise. The output is image data that has been prepared for analysis. Specifically, the server uses OpenCV to clear the image.

[0185] Step 3:

[0186] The server extracts features from pre-processed image data. The input is pre-processed image data. The server uses an image recognition AI model to extract the facial expressions and behavioral features of children as numerical data. The output is feature data. Specifically, the AI ​​model performs facial landmark detection and motion analysis.

[0187] Step 4:

[0188] The server detects potential abnormal behavior based on extracted feature data. The input is feature data. The server uses an AI model to execute an algorithm to detect signs of bullying and abnormal behavior. The output is an analysis result indicating potential abnormal behavior. Specifically, the AI ​​model identifies patterns of abnormal behavior.

[0189] Step 5:

[0190] The server provides analysis results to parents and educators via an information display device. The input is the analysis results indicating potential abnormal behavior. The server displays the results on a dashboard and mobile app. The output is the analysis results in a format that users can review. Specifically, the server displays the results via a web interface.

[0191] Step 6:

[0192] The server uses newly collected data to train an AI model. The input consists of analysis results and newly collected data. The server executes the AI ​​model's learning algorithm to improve its accuracy. The output is the improved AI model. Specifically, the server updates the model's parameters.

[0193] (Application Example 3)

[0194] Next, we will describe application example 3 of form example 3. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0195] Bullying and isolation among minors are serious challenges in educational settings. Traditional methods make it difficult to detect these problems early and respond appropriately. In particular, there is a lack of means to monitor changes in minors' facial expressions and behavior in real time and respond quickly. Therefore, a new system is needed to ensure the safety and healthy development of minors.

[0196] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 3 is realized by the following means.

[0197] This invention includes a server that uses an image recognition algorithm to analyze the facial expressions and behavior of minors and detect the possibility of bullying; a server that issues a warning based on the detected possibility of bullying; a server that proposes countermeasures to parents and educators who have received a warning; a server that analyzes the facial expressions and behavior of minors in real time and sends a notification when an abnormality is detected; and a server that can immediately check the situation of minors via a smart device. This makes it possible to detect problems involving minors early and respond quickly and appropriately.

[0198] An "image recognition algorithm" is a method that uses computer vision technology to identify specific patterns or features from image data.

[0199] "Minor" refers to a person who has not reached the legal age of majority, and usually includes children and adolescents under the age of 18.

[0200] "Facial expression" is a visual representation of emotions and reactions shown through the movement and arrangement of facial muscles.

[0201] "Behavior" is a general term for actions and reactions that an individual exhibits in specific circumstances.

[0202] "Potential bullying" refers to a situation in which a particular behavior or situation has the potential to lead to aggressive or harmful behavior towards others.

[0203] A "warning" is a notification or alert that informs you of a potential risk or problem.

[0204] "Guardian" refers to a parent or legal guardian who is responsible for the life and education of a minor.

[0205] "Educational personnel" refers to teachers and staff involved in the education and guidance of minors in educational institutions.

[0206] "Measures" refer to specific actions or plans taken to address a particular problem or issue.

[0207] "Real-time" refers to a state where data processing and information provision occur immediately, with virtually no delay.

[0208] A "notification" is a means of conveying specific information or a message to a recipient.

[0209] A "smart device" refers to a portable or stationary electronic device that is capable of connecting to the internet and running applications.

[0210] The system for implementing this invention primarily uses a server, a smart device, and an image recognition algorithm. The server executes the image recognition algorithm to analyze the facial expressions and behavior of minors in real time. Specifically, it processes video data acquired from a camera using machine learning frameworks such as TensorFlow and PyTorch. This allows for the detection of changes in the facial expressions and behavior of minors and the assessment of the possibility of bullying.

[0211] The server issues a warning based on the detected potential bullying. The warning is notified to parents and educators via smart devices. Smart devices such as smartphones and smart glasses receive notifications in real time, allowing them to immediately check on the minor's situation.

[0212] As a concrete example, during school breaks, cameras monitor the behavior of minors. A server analyzes the video data and detects potential bullying if a particular minor is isolated or having trouble with other minors. If bullying is detected, the server immediately sends a notification to parents or educators and suggests appropriate measures.

[0213] By utilizing generative AI models, the server can self-learn based on new data, improving the accuracy of bullying detection. An example of a prompt would be, "I want to develop an application that analyzes children's facial expressions and behavior and sends notifications if abnormalities are detected. What algorithms and technologies should I use?"

[0214] The flow of the specific processing in Application Example 3 will be explained using Figure 16.

[0215] Step 1:

[0216] The server acquires video data from the camera. The input is real-time video footage of minors. The server preprocesses this video data and converts it into a format that is easy for image recognition algorithms to process. Specifically, it adjusts the video resolution and removes noise.

[0217] Step 2:

[0218] The server inputs pre-processed video data into an image recognition algorithm. The algorithm analyzes the facial expressions and behaviors of minors and extracts features. The output is feature data related to facial expressions and behaviors. Specifically, it performs facial landmark detection and behavioral pattern recognition.

[0219] Step 3:

[0220] The server evaluates the likelihood of bullying based on extracted feature data. The input is feature data, and the output is a score indicating the likelihood of bullying. The server uses a generative AI model to compare with past data and detect anomalies. Specifically, it applies an anomaly detection algorithm to evaluate deviations from normal behavioral patterns.

[0221] Step 4:

[0222] The server generates a warning if it determines there is a high probability of bullying. The input is a score indicating the likelihood of bullying, and the output is a warning message. The server sends this warning to smart devices. Specifically, it generates push notifications and sends alerts to parents and educators.

[0223] Step 5:

[0224] The device (smart device) receives a warning from the server. The input is the warning message, and the output is a notification displayed to the user. The device displays the notification on the screen so that the user can immediately check the situation. Specifically, it plays a notification sound and displays a pop-up on the screen.

[0225] Step 6:

[0226] The user checks the notification and takes action as necessary. The input is the warning message displayed on the device, and the output is the user's response. Based on the notification, the user checks the minor's situation and takes appropriate action. Specific actions might include educators going to the scene or parents contacting the school.

[0227] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0228] "Example of form 1"

[0229] This invention is a system that combines an emotion engine. Specifically, in addition to a device that captures a child's facial expressions and behavior, it is equipped with an emotion engine that recognizes the user's emotions. The emotion engine analyzes emotions from, for example, the user's tone of voice, word choice, and facial expressions. This emotion engine can determine that if a child is showing negative emotions such as anger or sadness, it is more likely to be a contributing factor to bullying, thereby improving the accuracy of detecting the possibility of bullying.

[0230] "Example of form 2"

[0231] This invention is a system that improves the accuracy of detecting the possibility of bullying based on the user's emotions detected by an emotion engine. Specifically, the emotion engine feeds back the detected emotion data to an image recognition AI, which then learns from that data. For example, if the emotion engine detects that a child is showing emotions such as anger or sadness, the AI ​​analyzes the child's facial expressions and behavior based on that information, enabling it to detect the possibility of bullying with greater accuracy.

[0232] "Example of form 3"

[0233] This invention is a system that adjusts the content of warnings based on the user's emotions detected by an emotion engine. Specifically, it adjusts the content of warnings and suggested countermeasures based on the emotion data detected by the emotion engine. For example, if the emotion engine detects that a child is showing strong anger, it can use that information to strengthen the content of the warning and suggest more specific countermeasures to parents and educators.

[0234] The following describes the processing flow for each example of the form.

[0235] "Example of form 1"

[0236] Step 1: A device that captures the child's facial expressions and actions records the child's actions.

[0237] Step 2: The captured data is sent to the emotion engine, which analyzes the user's emotions based on their voice tone, word choice, facial expressions, etc.

[0238] Step 3: The emotional engine determines that if a child is exhibiting negative emotions such as anger or sadness, it is more likely to be a contributing factor to bullying.

[0239] "Example of form 2"

[0240] Step 1: The emotion engine analyzes the child's emotions and feeds the results back to the image recognition AI.

[0241] Step 2: The image recognition AI learns based on feedback from the emotion engine.

[0242] Step 3: Based on the learning results, the AI ​​analyzes the child's facial expressions and behavior to detect the possibility of bullying with greater accuracy.

[0243] "Example of form 3"

[0244] Step 1: The emotion engine analyzes the child's emotions and adjusts the content of the warnings and suggested countermeasures based on the results.

[0245] Step 2: For example, if the emotion engine detects that the child is showing strong anger, it uses that information to strengthen the content of the warning.

[0246] Step 3: Based on that information, propose more specific measures to parents and educators.

[0247] (Example 1)

[0248] Next, we will describe Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0249] In modern society, bullying is a serious social problem, particularly affecting the physical and mental health of children. However, it is difficult to detect the signs of bullying early and to respond appropriately. Traditional methods can lead to delays in detecting bullying and escalating the damage, so there is a need to quickly and accurately detect the possibility of bullying and take appropriate measures.

[0250] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0251] In this invention, the server includes means for acquiring a child's facial expressions and behavior using a camera; means for analyzing the acquired data using image recognition technology to detect the possibility of bullying; and means for analyzing audio data to identify emotions. This makes it possible to quickly and accurately detect the possibility of bullying and take appropriate measures.

[0252] A "recording device" is a device used to capture a child's facial expressions and behavior, and includes devices such as cameras and smartphones.

[0253] "Image recognition technology" is a technique that analyzes acquired image data to identify specific patterns and features, and is used to analyze children's facial expressions and behavior.

[0254] "Audio data" refers to data that records children's voices and surrounding sounds, and is used for analysis to identify emotions.

[0255] "Emotion analysis technology" is a technology that analyzes voice data to identify the speaker's emotions, and it is a technology that determines emotions by analyzing the tone of voice and word choice.

[0256] "Potential bullying" is an indicator that shows the possibility that a child is being bullied or is involved in bullying, and it is evaluated by integrating the results of image recognition technology and emotion analysis technology.

[0257] A "warning" is a notification issued to parents and educators when the possibility of bullying is detected, and it contains information to encourage appropriate measures.

[0258] The "means of proposing countermeasures" refer to a function that provides specific countermeasures to parents and educators based on the detected possibility of bullying.

[0259] This invention is a system for quickly and accurately detecting the possibility of bullying and taking appropriate countermeasures. The system consists of a camera, image recognition technology, and voice data analysis technology.

[0260] Users utilize cameras and smartphones as recording devices for use at school or home. These devices capture children's facial expressions and behavior in real time. The devices then transmit the acquired image data to image recognition technology. This technology analyzes children's facial expressions and behavior using libraries such as TensorFlow and OpenCV. Specifically, it extracts facial feature points and determines whether the expression indicates fear or sadness.

[0261] The server receives audio data transmitted from the terminal and analyzes it using emotion analysis technology. This technology analyzes the tone and word choice of the voice to identify the emotions the user is expressing. For example, it analyzes the pitch and speed of the voice to detect negative emotions such as anger or sadness.

[0262] The server integrates data obtained from image recognition and sentiment analysis technologies to assess the likelihood of bullying. If the assessment determines that there is a high probability of bullying, the server issues a warning to the user. After receiving the warning, the user can take appropriate measures at school or home as needed.

[0263] As a concrete example, suppose a user is using a smartphone to film their child. If image recognition technology analyzes the child's facial expression and determines that the child is depressed, the server will use that information to notify the user that bullying may be occurring. Similarly, if emotion analysis technology analyzes the tone of the child's voice and determines that the child is angry, the system will also determine that bullying may be occurring.

[0264] An example of a prompt to input into a generative AI model is, "Please tell me how to analyze a child's facial expressions and tone of voice to detect the possibility of bullying." This prompt allows the AI ​​model to provide information about specific analysis methods and the techniques to be used.

[0265] The flow of the specific processing in Example 1 will be explained using Figure 17.

[0266] Step 1:

[0267] The device uses a camera or smartphone to capture the child's facial expressions and actions in real time. The input is the captured image data. The device prepares this image data to be transmitted to image recognition technology. Specifically, the device converts the image data to an appropriate format and performs preprocessing for analysis.

[0268] Step 2:

[0269] The device analyzes captured image data using image recognition technology. The input is the image data acquired in step 1. The image recognition technology uses libraries such as TensorFlow and OpenCV to extract facial feature points and determine whether the expression is frightened or depressed. The output is the result of the facial expression analysis. Specifically, the device applies a face detection algorithm and quantifies the facial features.

[0270] Step 3:

[0271] The server receives audio data transmitted from a terminal and analyzes the data using emotion analysis technology. The input is audio data. Emotion analysis technology analyzes the tone and wording of speech to identify the emotion expressed by the user. The output is an emotion analysis result. As a specific operation, the server performs spectrum analysis on the audio data and detects changes in pitch and speed.

[0272] Step 4:

[0273] The server integrates data obtained from image recognition technology and emotion analysis technology to evaluate the possibility of bullying. The inputs are a facial expression analysis result and an emotion analysis result. The server integrates these pieces of data to calculate an index indicating the possibility of bullying. The output is an evaluation result of the possibility of bullying. As a specific operation, the server uses a statistical method to weight each analysis result and performs a comprehensive evaluation.

[0274] Step 5:

[0275] If it is determined that the possibility of bullying is high, the server issues a warning to the user. The input is the evaluation result of the possibility of bullying. The server generates a warning message based on the evaluation result and notifies the user thereof. The output is a warning message. As a specific operation, the server transmits the message to the user's terminal via a notification system.

[0276] (Application Example 1)

[0277] Next, Application Example 1 of Embodiment 1 will be described. In the following description, the data processing device 12 is referred to as a "server", and the smart device 14 is referred to as a "terminal".

[0278] When there is a possibility that a child is bullied at school or at home, it is required to detect such signs at an early stage and take appropriate countermeasures. However, conventional methods often miss signs of bullying, which poses a problem that it is difficult to sufficiently ensure the safety of children.

[0279] The identification processing by the identification processing unit 290 of the data processing apparatus 12 in Application Example 1 is implemented by the following means.

[0280] In the present invention, the server includes: means for analyzing children's facial expressions and behaviors using image recognition technology to detect the possibility of bullying; means for analyzing the tone and wording of children's voices using emotion analysis technology to evaluate the possibility of bullying; and means for monitoring children's safety in real time via smart devices. This makes it possible to detect signs of bullying at an early stage and promptly notify guardians and education-related personnel.

[0281] "Image recognition technology" refers to technology that analyzes image data acquired using cameras and sensors to identify specific patterns and features.

[0282] "Children" refers to children of an age to receive education at school or at home.

[0283] "Facial expression" refers to the emotions and reactions expressed by movements of the facial muscles.

[0284] "Behavior" refers to the movements and conduct of children in their daily lives.

[0285] "Possibility of bullying" indicates the possibility that a child suffers psychological or physical harm from aggressive behaviors or words from others.

[0286] "Emotion analysis technology" refers to technology that analyzes voice and text data to estimate the speaker's emotional state.

[0287] "Voice tone" refers to the characteristics of a speaker's voice expressed by the pitch, intensity, inflection and other aspects of the voice.

[0288] "Wording" refers to the word selection and expression methods used by a speaker.

[0289] A "smart device" refers to a portable electronic device that has internet connectivity and can run various applications.

[0290] "Real-time" refers to a state where data acquisition and processing occur instantly, and results are obtained without delay.

[0291] "Guardian" refers to a child's parent or the person legally responsible for caring for that child.

[0292] "Education professionals" refers to those who are responsible for duties related to the education and guidance of children in schools and educational institutions.

[0293] The system for implementing this invention mainly consists of three elements: a server, a terminal, and a user. The server uses image recognition technology and emotion analysis technology to analyze the child's facial expressions, behavior, tone of voice, and speech. Specifically, the server uses TensorFlow to analyze image data acquired from cameras and sensors to identify the child's facial expressions and behavior. It also uses IBM Watson® to analyze audio data and estimate the child's emotional state.

[0294] The device functions as a smart device, monitoring children's safety in real time. It receives analysis results from the server and immediately sends notifications to parents and educators if potential bullying is detected. This enables a rapid response.

[0295] As a parent or educator, the user can receive notifications from the device and take appropriate measures to ensure the safety of children. For example, if the device detects a situation during school recess where a child is being verbally abused by another child, the server will determine from the child's facial expression that they are "scared" and analyze the tone of their voice as "angry." Based on these results, the device will determine that there is a "possible case of bullying" and send a notification to the parent or guardian.

[0296] As an example of a prompt sentence to be input to a generative AI model, "Please analyze the facial expressions and voice tones of children shown in this video and evaluate the possibility of bullying." can be cited. With this prompt sentence, the server can perform appropriate analysis and evaluate the possibility of bullying.

[0297] The flow of the specifying process in Application Example 1 will be described with reference to FIG. 18.

[0298] Step 1:

[0299] The server receives the image data and audio data transmitted from the terminal. The input is data of children's facial expressions and voices acquired via a camera or a microphone. In order to analyze these data, the server first preprocesses the data to remove noise.

[0300] Step 2:

[0301] The server analyzes the preprocessed image data using TensorFlow. The input is the noise-removed image data. The server uses image recognition technology to identify the children's facial expressions and behaviors, and extracts features indicating the possibility of bullying. The output is an analysis result regarding the facial expressions and behaviors.

[0302] Step 3:

[0303] The server analyzes the preprocessed audio data using IBM Watson. The input is the noise-removed audio data. The server uses sentiment analysis technology to analyze the children's voice tones and wording, and estimates their emotional states. The output is an analysis result regarding emotions.

[0304] Step 4:

[0305] The server integrates the analysis results of the image data and the audio data, and evaluates the possibility of bullying. The input is the analysis results regarding facial expressions, behaviors and emotions. The server combines these data to comprehensively determine the possibility of bullying. The output is an evaluation result regarding the possibility of bullying.

[0306] Step 5:

[0307] The device receives the assessment results regarding the possibility of bullying, which are sent from the server. The input is the assessment results from the server. Based on these results, if the device determines that there is a high possibility of bullying, it sends a notification to parents or educators. The output is the notification message.

[0308] Step 6:

[0309] The user receives notifications from the device and takes appropriate measures to ensure the child's safety. The input is the notification message from the device. Based on this information, the user checks the child's situation and intervenes as needed. The output is the specific action taken to ensure the child's safety.

[0310] (Example 2)

[0311] Next, we will describe Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0312] There are challenges in the early detection of bullying and the proposal of appropriate countermeasures. In particular, there is a need to accurately capture changes in children's emotions and behavior and detect the possibility of bullying with high precision. It is also important to quickly provide concrete countermeasures based on the detection results.

[0313] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0314] In this invention, the server includes means for collecting emotional data, means for feeding the collected emotional data back to image recognition technology, and means for analyzing a child's facial expressions and behavior using image recognition technology to detect the possibility of bullying. This makes it possible to detect the possibility of bullying with high accuracy and to provide appropriate countermeasures quickly.

[0315] "Emotional data" refers to information about a child's emotions obtained from their facial expressions and behavior.

[0316] "Image recognition technology" is a technique that analyzes image data to identify specific patterns or features.

[0317] "Possibility of bullying" refers to the possibility of bullying occurring, inferred from the analysis of a child's facial expressions and behavior.

[0318] A "warning" is a cautionary message sent to parents and educators when the possibility of bullying is detected.

[0319] "Generative AI technology" is a technology that uses artificial intelligence to generate new information and proposals.

[0320] A "learning function" is a feature that allows a system to improve itself based on past data.

[0321] This invention is a system that detects the possibility of bullying with high accuracy and provides appropriate countermeasures. Specific embodiments are shown below.

[0322] The user uses their device to record the child's facial expressions and behavior using an emotion engine. The emotion engine uses sensors such as cameras and microphones to detect the child's emotions, such as anger or sadness, in real time. For example, if a child looks sad at school, video and audio data of that moment are collected.

[0323] The device sends the collected emotional data to a server. The server receives this data and feeds it back into image recognition technology. The image recognition technology learns from the received emotional data and analyzes the child's facial expressions and behavior in detail. This analysis is used to assess the possibility of bullying.

[0324] If potential bullying is detected, the server will issue a warning to parents or educators. This warning will be sent as an email or app notification and will include details such as the date, time, and location of the detection, as well as the child's facial expressions and behavior.

[0325] Furthermore, the server uses generative AI technology to suggest specific countermeasures. For example, if you input a prompt message such as "What should you do if your child looks sad at school?", the AI ​​will suggest ways for parents to encourage communication with their children and suggest ways for educators to strengthen school monitoring systems.

[0326] In this way, the system can quickly and effectively detect potential bullying and provide appropriate countermeasures.

[0327] The flow of the specific processing in Example 2 will be explained using Figure 19.

[0328] Step 1:

[0329] The user uses a device to record the child's facial expressions and behavior using an emotion engine. Video and audio data from the camera and microphone are used as input. The emotion engine analyzes this data to detect the child's emotional state (e.g., anger, sadness). The detected emotion data is generated as output.

[0330] Step 2:

[0331] The device sends the collected emotion data to the server. The emotion data generated in step 1 is used as input. The server receives this data and feeds it back to the image recognition technology. The feedback emotion data is passed back to the image recognition technology as output.

[0332] Step 3:

[0333] The server uses image recognition technology to learn from the feedbacked emotional data. The feedbacked emotional data is used as input. The image recognition technology analyzes the data, examining the child's facial expressions and behavior in detail. The output generates an assessment of the likelihood of bullying.

[0334] Step 4:

[0335] If the server detects a potential bullying incident, it will issue a warning to parents or educators. The evaluation results generated in step 3 are used as input. The server will send a warning via email or app notification, including the date, time, and location of the detection, as well as details of the child's facial expressions and behavior. An alert message is generated and sent as output.

[0336] Step 5:

[0337] The server uses AI generation technology to propose specific countermeasures. The warning message generated in step 4 is used as input. Based on the prompt "What should be done if a child looks sad at school?", the AI ​​generation technology proposes specific countermeasures for parents and educators. The proposed countermeasures are generated and provided as output.

[0338] (Application Example 2)

[0339] Next, we will describe application example 2 of form example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0340] While early detection and appropriate response are crucial in addressing bullying among children, conventional methods struggle with detection and risk overlooking early signs of bullying. Furthermore, there is a need for a system that can accurately detect potential bullying and promptly notify relevant parties.

[0341] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0342] In this invention, the server includes means for analyzing a child's facial expressions and behavior using an image recognition algorithm to detect the possibility of bullying, means for analyzing video in real time and extracting emotional data, and means for learning based on the extracted emotional data and evaluating the possibility of bullying. This makes it possible to detect the possibility of bullying with high accuracy and to quickly notify parents and educators.

[0343] An "image recognition algorithm" is a computational method used to analyze video data acquired from cameras and sensors and identify specific patterns or features.

[0344] "Children's facial expressions and behavior" refers to the visual characteristics of children, such as their facial expressions, body movements, and posture.

[0345] "Potential bullying" refers to the probability or risk of a child exhibiting signs of aggressive or inappropriate behavior towards other children.

[0346] "Emotional data" refers to information that indicates emotional states such as anger, sadness, and joy, extracted from a child's facial expressions and behavior.

[0347] "Learning" is the process by which an algorithm improves its performance based on past data and newly acquired data.

[0348] "Evaluating" means judging the possibility of bullying or the emotional state based on specific criteria and deriving a result.

[0349] "Generating and sending notifications" is the process of creating and sending messages to relevant parties to provide warnings or information based on detected data.

[0350] The system for implementing this invention consists of a server, a terminal, and a user. The server acquires video data from surveillance cameras installed within the school and analyzes the children's facial expressions and behavior in real time using an image recognition algorithm. Specifically, the server uses software such as OpenCV and TensorFlow to extract facial features from the video data and generates emotion data using an emotion engine.

[0351] The generated emotion data is used in a learning process within the server to update the model for evaluating the possibility of bullying. Based on this evaluation, the server generates a warning and sends a notification to the devices of parents and educators. The notification includes the date, time, and location of the detection, details of the child's facial expressions and behavior, and specific countermeasures are suggested.

[0352] The terminal is a device such as a smartphone or tablet that receives notifications and displays warnings to the user. The user can check the notification and take the suggested action.

[0353] For example, if a camera captures a scene in a school hallway where a student is showing anger towards another student, the server will detect that anger using an emotion engine and analyze the scene with an image recognition algorithm. If it is determined that there is a high probability of bullying, the server will generate a notification stating, "Possible bullying has been detected in the hallway. Please check the details," and send it to the relevant parties.

[0354] An example of a prompt message is: "Analyze the facial expressions of children in the hallway. If feelings of anger or sadness are detected, use that information to assess the possibility of bullying and generate a notification."

[0355] The flow of a specific process in Application Example 2 will be explained using Figure 20.

[0356] Step 1:

[0357] The server acquires video data in real time from surveillance cameras installed within the school. The input is a video stream from the cameras, and the output is analyzable image data. To process this data, the server divides the video into frames.

[0358] Step 2:

[0359] The server uses an image recognition algorithm to detect children's faces from the acquired frames. The input is the image data obtained in step 1, and the output is the face's position information and features. The server uses OpenCV to extract facial features and determine the face's position.

[0360] Step 3:

[0361] The server uses an emotion engine to generate emotion data from detected facial features. The input is the facial features obtained in step 2, and the output is data indicating the type and intensity of emotion. The server feeds this data back into the generating AI model to perform emotion analysis.

[0362] Step 4:

[0363] The server inputs the generated emotion data into an image recognition algorithm to evaluate the possibility of bullying. The input is emotion data, and the output is a score indicating the possibility of bullying. The server uses TensorFlow to analyze the emotion data and quantify the possibility of bullying.

[0364] Step 5:

[0365] The server generates and sends notifications to parents and educators based on the bullying possibility score. The input is the score obtained in step 3, and the output is a warning message. The server uses a generative AI model to create notifications that suggest specific actions based on the prompt text.

[0366] Example prompt: "Analyze students' facial expressions in the hallway. If anger or sadness is detected, use this information to assess the possibility of bullying and generate a notification."

[0367] (Example 3)

[0368] Next, we will describe Embodiment 3 of Embodiment Example 3. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0369] Conventional technologies have struggled to accurately analyze an individual's emotional state and provide appropriate warnings and countermeasures. Furthermore, they lacked the learning capabilities necessary to detect potential bullying with high accuracy, hindering the ability of stakeholders to respond quickly and appropriately.

[0370] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 3 is realized by the following means.

[0371] In this invention, the server includes means for analyzing an individual's emotional state using image processing technology, means for adjusting the content of a warning based on the analyzed emotional data, and means for proposing specific countermeasures to the person who received the warning. This makes it possible to accurately grasp an individual's emotional state and quickly provide appropriate warnings and countermeasures.

[0372] "Image processing technology" refers to techniques for analyzing digital images and extracting or recognizing specific information.

[0373] "Emotional state" refers to the psychological or emotional state an individual exhibits at a particular moment in time.

[0374] "Warning content" refers to information that serves as a warning or instruction to relevant parties, generated based on the analyzed data.

[0375] "Specific countermeasures" refer to proposals that instruct relevant parties on the actions and measures they should take depending on the detected situation.

[0376] "Self-learning" is the process by which a system automatically learns from new data and improves its performance and accuracy.

[0377] "Stakeholders" refers to individuals such as parents and educators who are the recipients of the analysis results and warnings.

[0378] A description of embodiments for carrying out this invention will be given.

[0379] The server receives data on an individual's facial expressions and actions transmitted from the device. The device collects data in real time using cameras and sensors and sends it to the server. This data is acquired when the individual is active in a specific environment.

[0380] The server analyzes the received data using image processing technology. This analysis utilizes image recognition AI based on frameworks such as TensorFlow and PyTorch. The image recognition AI infers emotional states from an individual's facial expressions and identifies emotions such as anger, sadness, and joy.

[0381] The analyzed emotion data is passed to the emotion engine on the server. The emotion engine quantifies the individual's emotional state based on the analysis results and adjusts the warning content. For example, if the emotion engine detects that an individual is showing strong anger, the server uses that information to strengthen the warning and propose specific countermeasures to those involved.

[0382] Furthermore, the server performs self-learning based on new data, improving the detection accuracy of the image recognition AI. This allows the server to continuously improve its accuracy and provide more precise analysis results.

[0383] As a concrete example, let's look at an example of a prompt sentence to be input to the generating AI model: "Analyze the facial expressions an individual displays during a specific activity and detect emotions such as anger or sadness. Based on the results, propose appropriate measures for those involved." In response to this prompt sentence, the server utilizes image recognition AI and emotion engine to analyze the individual's emotional state in detail and provide appropriate information. The flow of specific processing in Example 3 will be explained using Figure 21.

[0384] Step 1:

[0385] The device uses cameras and sensors to collect data on an individual's facial expressions and behavior in real time. The collected data is transmitted to a server in image or video format. The input is raw data of the individual's facial expressions and behavior, and the output is the transmission of data to the server.

[0386] Step 2:

[0387] The server analyzes the data received from the terminal using image processing technology. Specifically, an image recognition AI using frameworks such as TensorFlow or PyTorch analyzes the received image data and infers the emotional state from the individual's facial expressions. The input is image data sent from the terminal, and the output is the analyzed emotional data.

[0388] Step 3:

[0389] The server passes the analyzed emotional data to the emotion engine, which quantifies the individual's emotional state. Based on the analysis results, the emotion engine identifies emotions such as anger, sadness, and joy, and adjusts the warning content accordingly. The input is the analysis results from the image recognition AI, and the output is the quantified emotional data.

[0390] Step 4:

[0391] The server generates warnings based on quantified emotion data and suggests specific actions for those involved. For example, if strong anger is detected, the server generates a warning such as, "An individual is exhibiting strong anger. Assess the situation and take steps to calm them down." The input is quantified emotion data, and the output is the generated warning and suggested actions.

[0392] Step 5:

[0393] The server performs self-learning based on new data to improve the detection accuracy of the image recognition AI. Specifically, it compares past data with new data and adjusts the model to reduce misrecognition. The input is past data and new data, and the output is the adjusted AI model.

[0394] Step 6:

[0395] The server sends the generated warnings and suggestions to the terminal, notifying the relevant user. The terminal displays the received information on its screen, allowing the user to understand the situation. The input is the generated warnings and suggestions, and the output is the information displayed on the terminal.

[0396] (Application Example 3)

[0397] Next, we will describe application example 3 of form example 3. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0398] There is a need to detect bullying and emotional changes among minors early and to take appropriate action. However, conventional methods may miss signs of bullying, and there are challenges in responding appropriately to emotional changes.

[0399] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 3 is realized by the following means.

[0400] This invention includes a server that uses image recognition technology to analyze the facial expressions and behavior of minors and detect the possibility of bullying; a server that issues a warning based on the detected possibility of bullying; a server that proposes countermeasures to parents and educators who receive the warning; a server that adjusts the content of the warning based on the minor's emotional data; and a server that captures the minor's facial expressions in real time and notifies them of the analysis results. This enables a swift and appropriate response to protect the safety and welfare of minors.

[0401] "Image recognition technology" is a technique that uses computer vision to extract and analyze specific information from image data.

[0402] A "minor" refers to a person who has not reached the legal age of majority.

[0403] "Facial expression" refers to the emotions and reactions shown through the movement of the facial muscles.

[0404] "Action" refers to a series of actions or reactions performed by an individual.

[0405] "Possible bullying" refers to the possibility that a minor is experiencing inappropriate behavior or remarks from others.

[0406] A "warning" is a notification or message intended to draw attention to a specific situation.

[0407] A "guardian" refers to a person who has the responsibility to supervise the life and education of a minor.

[0408] "Education-related personnel" refers to individuals who hold positions in educational institutions that involve the education of minors.

[0409] "Measures" refer to specific actions or means taken to address a particular problem or situation.

[0410] "Emotional data" refers to information that expresses an individual's emotional state using numerical values ​​or categories.

[0411] "Real-time" refers to processing or responding instantly without delay.

[0412] A "notification" is a message or alert used to convey specific information to a recipient.

[0413] The system for implementing this invention mainly consists of a server and a terminal. The server uses image recognition technology to analyze the facial expressions and behavior of minors and detect the possibility of bullying. Specifically, the terminal's camera is used to capture the minor's facial expressions in real time, and the data is sent to the server. The server uses OpenCV as the image recognition technology and Microsoft® Azure® Emotion API for sentiment analysis. This allows the system to extract the minor's sentiment data and evaluate the possibility of bullying.

[0414] The server issues warnings based on detected bullying and suggests countermeasures to parents and educators. The content of the warnings is adjusted based on the minor's emotional data. For example, if the minor is showing strong anger, the warning is strengthened and specific countermeasures are suggested.

[0415] The device functions as a smartphone or smart glasses, capturing the facial expressions of minors in real time. The analysis results are sent as notifications to parents and educators. This enables quick and appropriate responses to protect the safety and well-being of minors.

[0416] As a concrete example, if a minor displays strong anger in a classroom, the device captures their facial expression and sends it to a server. The server performs an emotion analysis, and if it determines that the expression is "anger," it sends a notification to the parent saying, "Your child is showing strong anger. Please listen to them."

[0417] An example of a prompt message is: "Analyze the child's facial expression data and determine their emotion. If anger is detected, generate a message to send a notification to the parent / guardian."

[0418] The flow of the specific processing in Application Example 3 will be explained using Figure 22.

[0419] Step 1:

[0420] The device uses a camera to capture the facial expressions of minors in real time. The input is video data from the camera, and the output is image data. The device then prepares to send this image data to a server.

[0421] Step 2:

[0422] The server receives image data transmitted from the terminal. The input is image data from the terminal, and the output is analyzable image information. The server processes this image information using image recognition technology (OpenCV) to analyze the facial expressions of minors.

[0423] Step 3:

[0424] The server inputs facial expression data extracted using image recognition technology into an emotion analysis engine (Microsoft Azure Emotion API). The input is facial expression data, and the output is emotion data. Based on this emotion data, the server determines the emotional state of the minor.

[0425] Step 4:

[0426] The server analyzes sentiment data and assesses the likelihood of bullying. The input is sentiment data, and the output is an assessment result: "bullying is likely" or "bullying is not likely." Based on this assessment, the server generates warnings as needed.

[0427] Step 5:

[0428] The server notifies parents and educators of the generated warnings. The input is the evaluation result, and the output is the warning message. The server adjusts the content of the warnings and suggests specific countermeasures based on the minor's emotional data.

[0429] Step 6:

[0430] The user receives notifications from the server and takes appropriate action based on the minor's situation. The input is the notification from the server, and the output is the user's response. Based on the notification, the user interacts with the minor and, if necessary, collaborates with educators to resolve the problem.

[0431] (Other examples)

[0432] Next, other embodiments will be described. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0433] In recent years, bullying among children in schools and homes has become a serious problem, and there is a need to detect signs of bullying early and take appropriate measures. However, conventional methods have a high probability of missing signs of bullying, and countermeasures are often delayed. Therefore, there is a need for a system that analyzes children's facial expressions and behavior in real time to quickly detect the possibility of bullying and provide appropriate warnings and countermeasures.

[0434] The identification process performed by the identification processing unit 290 of the data processing device 12 in other embodiments is realized by the following means.

[0435] In this invention, the server includes means for analyzing a child's facial expressions and behavior using image recognition technology and emotion analysis technology to detect the possibility of bullying, means for generating warning messages using a generative AI model and adjusting the content of the warnings, and means for providing specific countermeasures to parents and educators based on the analysis results. This enables early detection of signs of bullying and prompt and appropriate responses.

[0436] "Image recognition technology" is a technology that analyzes video data acquired from a camera device to identify specific patterns or features.

[0437] "Emotional analysis technology" is a technique that analyzes audio and text data to identify the emotional state of the speaker.

[0438] A "generative AI model" is an artificial intelligence model that generates natural language text based on an input prompt sentence, and specific examples include OpenAI's GPT-3 (registered trademark).

[0439] A "prompt statement" is an input statement used to instruct a generative AI model to produce a specific output.

[0440] A "warning message" is a message generated to alert parents and educators when potential bullying is detected.

[0441] A "countermeasure plan" is a proposal outlining specific actions and measures that parents and educators should take when the possibility of bullying is detected.

[0442] "Self-learning" is the process by which a machine learning model automatically learns and improves its accuracy based on newly collected data.

[0443] TensorFlow is an open-source software library for building and training machine learning models.

[0444] This invention is a system for early detection of potential bullying among children and for taking appropriate countermeasures. Specific embodiments of this system are described below.

[0445] The server receives video and audio data in real time from cameras and microphones installed in classrooms and homes. The cameras capture children's facial expressions, and the microphones record children's conversations and voice tones. This data is transmitted to the server using a secure communication protocol.

[0446] The server analyzes the received video data using the OpenCV library to identify emotions from the children's facial expressions. For example, it can identify smiles and angry expressions. For audio data, it converts it to text using the Google Cloud Speech-to-Text API and performs emotion analysis. This allows the server to evaluate the emotional state of the children based on their voice tone and word choice.

[0447] The analyzed facial expression and emotion data are integrated using a machine learning model based on TensorFlow to assess the likelihood of bullying. The model is trained on historical data and identifies signs of bullying by detecting specific patterns.

[0448] If potential bullying is detected, the server inputs a prompt message to the generative AI model to generate a warning message. The generative AI model used is OpenAI's GPT-3, which generates an appropriate warning message. Example prompt message: "Signs of bullying have been detected from child A's facial expressions and emotion data. Please generate an appropriate warning message."

[0449] Furthermore, the server inputs prompts to the AI ​​model to generate proposed countermeasures. These generated countermeasures are then provided to parents and educators. Example of a prompt: "Possible bullying of child A has been detected. Please generate specific countermeasures to propose to parents and educators."

[0450] The server provides parents and educators with analysis results, generated warning messages, and suggested countermeasures. This is done by sending emails via the mail server or by providing notifications through a dedicated application.

[0451] The device provides an interface accessible to parents and educators, visually displaying analysis results and suggested countermeasures in a dashboard format.

[0452] The server uses newly collected data to self-learn and improve the accuracy of detecting potential bullying. Specifically, it compares past analysis results with new data and updates the model parameters using TensorFlow. This allows the system to continuously improve its accuracy.

[0453] The flow of specific processing in other embodiments will be explained using Figure 23.

[0454] Step 1:

[0455] The device uses cameras and microphones installed in classrooms and homes to acquire video and audio data of children in real time. The cameras capture children's facial expressions, and the microphones record children's conversations and voice tones. This data is transmitted to a server using a secure communication protocol. The input is raw data from the cameras and microphones, and the output is the video and audio data transmitted to the server.

[0456] Step 2:

[0457] The server analyzes the received video data using the OpenCV library to identify emotions from the children's facial expressions. Specifically, it divides the video data into frames and applies a face detection algorithm to identify facial expressions. The input is video data, and the output is the identified facial expression data.

[0458] Step 3:

[0459] The server uses the Google Cloud Speech-to-Text API to convert the audio data into text and then performs sentiment analysis. After the audio data is converted to text, natural language processing techniques are used to evaluate the emotional state based on the tone of voice and word choice. The input is audio data, and the output is data indicating the emotional state.

[0460] Step 4:

[0461] The server integrates the analyzed facial expression and emotion data and uses a machine learning model based on TensorFlow to assess the likelihood of bullying. The model is trained on historical data and determines signs of bullying by detecting specific patterns. The input is facial expression and emotion data, and the output is an assessment result indicating the likelihood of bullying.

[0462] Step 5:

[0463] If the possibility of bullying is detected, the server inputs a prompt message to the AI ​​model that generates a warning message. Specifically, it uses the prompt message, "Signs of bullying have been detected from Child A's facial expressions and emotion data. Please generate an appropriate warning message." The input is the prompt message, and the output is the generated warning message.

[0464] Step 6:

[0465] The server receives a prompt message from the generating AI model to generate proposed countermeasures. For example, it might use the prompt message, "Possible bullying of child A has been detected. Please generate specific countermeasures to propose to parents and educators." The input is the prompt message, and the output is the generated countermeasures.

[0466] Step 7:

[0467] The server provides parents and educators with analysis results, generated warning messages, and suggested countermeasures. This is done by sending emails via a mail server or by providing notifications through a dedicated application. The input is the analysis results, generated messages, and suggested countermeasures, and the output is the sent notifications.

[0468] Step 8:

[0469] The server uses newly collected data to self-learn and improve the accuracy of detecting potential bullying. Specifically, it compares past analysis results with the new data and updates the model parameters using TensorFlow. The input is the newly collected data, and the output is the updated model.

[0470] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0471] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0472] Other examples of generative AI include Gemini® (registered trademark) (Internet search). <url: https: gemini.google.com ?hl="ja">) are examples.

[0473] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0474] [Second Embodiment]

[0475] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0476] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0477] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0478] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0479] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0480] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0481] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0482] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0483] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0484] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0485] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0486] Next, the identification process performed by the identification processing unit 290 of the data processing device 12 will be described.

[0487] "Example of form 1"

[0488] One embodiment of the present invention involves equipping a camera or smartphone, or other imaging device, with image recognition AI for use in schools or homes. This image recognition AI analyzes a child's facial expressions and behavior in real time to detect the possibility of bullying. For example, it detects the possibility of bullying if it determines that a child's facial expression indicates fear or sadness, or if it determines that the child is being subjected to aggressive behavior from other children.

[0489] "Example of form 2"

[0490] Based on the detected potential bullying, warnings are issued to parents and educators. These warnings are sent via email, app notifications, etc. The warnings include details such as the date, time, and location where the potential bullying was detected, as well as details of the child's facial expressions and behavior. Specific countermeasures are also suggested. For example, parents may be advised to take measures to encourage communication with their children, and educators may be advised to strengthen school monitoring systems and offer counseling.

[0491] "Example of form 3"

[0492] Furthermore, the system provides parents and educators with analysis results of children's facial expressions and behavior. This allows parents and educators to understand the child's situation more concretely and take appropriate action. It also includes a learning function to improve the accuracy of detecting potential bullying. This learning function allows the image recognition AI to self-learn based on new data and improve detection accuracy.

[0493] The following describes the processing flow for each example of the form.

[0494] "Example of form 1"

[0495] Step 1: Equipate image recognition AI into cameras, smartphones, and other imaging devices for use in schools and homes.

[0496] Step 2: Capture the child's facial expressions and actions in real time through the camera and send them to the image recognition AI.

[0497] Step 3: The image recognition AI analyzes the transmitted image data and detects the possibility of bullying from the child's facial expressions and behavior.

[0498] "Example of form 2"

[0499] Step 1: If the image recognition AI detects a potential bullying incident, it issues a warning to parents and educators.

[0500] Step 2: The warning will be sent via email, app notification, etc.

[0501] Step 3: The warning will include details such as the date and location where the bullying was detected, as well as the child's facial expressions and behavior.

[0502] Step 4: Specific measures are proposed. For example, parents may be advised to take steps to encourage communication with their children, and educators may be advised to strengthen school supervision systems and provide counseling.

[0503] "Example of form 3"

[0504] Step 1: Provide parents and educators with the results of the analysis of the child's facial expressions and behavior.

[0505] Step 2: Parents and educators will be able to understand the child's situation more concretely and take appropriate action.

[0506] Step 3: Includes a learning function to improve the accuracy of detecting potential bullying.

[0507] Step 4: This learning function allows the image recognition AI to self-learn based on new data and improve detection accuracy.

[0508] (Example 1)

[0509] Next, we will describe Embodiment 1 of Embodiment Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0510] In modern schools and homes, children are increasingly likely to encounter bullying, but there is a challenge in detecting signs of bullying early and taking appropriate measures. In particular, there is a need for technology that can monitor changes in children's facial expressions and behavior in real time and quickly detect the possibility of bullying.

[0511] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0512] In this invention, the server includes means for collecting video data using a camera, means for transmitting the collected video data via communication means, and means for analyzing a child's facial expressions and behavior using image recognition technology to detect the possibility of bullying. This makes it possible to detect signs of bullying early and take appropriate measures.

[0513] A "recording device" is a device used to collect video data, and includes devices such as cameras and smartphones.

[0514] "Video data" refers to digital data containing visual information collected by a camera or camera.

[0515] "Communication methods" refer to the technologies and protocols used to send and receive data, and include methods such as the internet and wireless communication.

[0516] "Image recognition technology" is a technology that allows computers to analyze image data and identify specific patterns or features.

[0517] "Analysis of facial expressions and behavior" is the process of using image recognition technology to analyze a child's facial expressions and body movements to identify specific emotions and behaviors.

[0518] "Detecting the possibility of bullying" is a process of determining the likelihood of bullying occurring based on analyzed facial expression and behavioral data.

[0519] "Means of issuing warnings" refer to methods or devices used to alert those involved when the possibility of bullying is detected.

[0520] "Means of proposing countermeasures" refer to methods and devices for presenting specific countermeasures to parents and educators based on the detected possibility of bullying.

[0521] "Learning function" refers to a function that allows the system to improve the accuracy of its analysis based on past data, and includes machine learning algorithms.

[0522] This invention is a system for use in schools and homes that analyzes children's facial expressions and behavior in real time to detect potential bullying. The system consists of a camera, communication means, and image recognition technology.

[0523] The terminal uses cameras, smartphones, and other recording devices to collect video data of children in real time. The collected video data is transmitted to a server via a secure communication method. The communication method uses the internet or wireless communication technology.

[0524] The server uses image recognition technology to analyze the received video data. This technology leverages generative AI models built using machine learning frameworks such as TensorFlow and PyTorch. Specifically, it uses OpenCV to extract facial features and analyze changes in facial expressions. It also applies motion recognition algorithms to detect aggressive behavior.

[0525] Users receive analysis results from the server and are warned if bullying is suspected. Furthermore, the system suggests specific measures to parents and educators based on the detected results. This makes it possible to detect signs of bullying early and take appropriate action.

[0526] For example, a user might monitor children's behavior through a camera installed in the classroom. The AI ​​would warn of potential bullying if it detects that a child frequently looks frightened or if it detects behavior such as another child unilaterally pushing or shoving. Examples of prompts to input into the generating AI model could be, "What should I do if a child looks frightened?" or "Please tell me how to deal with aggressive behavior that is detected."

[0527] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0528] Step 1:

[0529] The device uses a camera or smartphone to collect video data of children in real time. The input is video data obtained from the recording device. Specifically, the device ensures high-resolution video so that facial expressions and actions can be clearly identified. The output is the collected video data.

[0530] Step 2:

[0531] The terminal sends the collected video data to the server. The input is the video data obtained in step 1. Specifically, the terminal uses a secure communication protocol (e.g., HTTPS) to send the data and protect privacy. The output is the video data sent to the server.

[0532] Step 3:

[0533] The server analyzes the received video data using image recognition technology. The input is the video data sent in step 2. Specifically, the server utilizes generative AI models using TensorFlow and PyTorch, extracts facial features using OpenCV, and analyzes changes in facial expressions. It also applies a motion recognition algorithm to detect aggressive behavior. The output is the analyzed facial expression and behavior data.

[0534] Step 4:

[0535] The server determines the possibility of bullying based on the analysis results and issues a warning if necessary. The input is the analysis data obtained in step 3. Specifically, if the server determines that bullying is possible, it generates a warning message. The output is the warning message.

[0536] Step 5:

[0537] The user receives a warning from the server and considers appropriate action. The input is the warning message generated in step 4. Specifically, the user checks the warning on their smartphone or PC and takes action, such as meeting with the child if necessary. The output is the user's response.

[0538] (Application Example 1)

[0539] Next, we will describe Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0540] In modern society, bullying of children at school and at home is a serious problem. However, it is not easy to detect signs of bullying early and take appropriate measures. In particular, there is a lack of means to monitor children's emotions and behavior in real time and quickly detect the possibility of bullying, which often leads to delays in early detection and response to bullying. To solve this problem, a system is needed that analyzes children's emotions and behavior in real time, immediately detects signs of bullying, and notifies those involved.

[0541] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0542] In this invention, the server includes means for analyzing a child's emotions and actions using image recognition technology to detect the possibility of bullying, means for monitoring a child's emotions and actions in real time using a mobile information terminal, and means for sending a notification when signs of bullying are detected. This makes it possible to detect signs of bullying in a child early and quickly notify the relevant parties.

[0543] "Image recognition technology" is a technique that analyzes image data acquired using cameras and sensors to identify specific patterns and features.

[0544] "Children's emotions and actions" refer to the psychological state and behaviors that can be interpreted from a child's facial expressions and body movements.

[0545] "Possible bullying" refers to signs that indicate a child may be experiencing aggressive behavior or psychological pressure from others.

[0546] A "portable information terminal" is an electronic device that is portable and capable of processing information, such as a smartphone or tablet.

[0547] "Real-time monitoring" means observing and analyzing the ongoing situation immediately.

[0548] "Sending a notification" means immediately informing relevant parties of specific information.

[0549] The system for implementing this invention mainly consists of a server and a mobile information terminal. The server utilizes image recognition technology to analyze children's emotions and actions and detect the possibility of bullying. Specifically, the server uses image recognition libraries such as TensorFlow and OpenCV to process image data acquired from cameras and sensors in real time. This makes it possible to identify signs of bullying from children's facial expressions and actions.

[0550] Mobile information devices refer to smartphones, tablets, and other similar devices. These devices work in conjunction with a server to monitor children's emotions and behavior in real time. The devices receive analysis results sent from the server, and if signs of bullying are detected, they immediately send notifications to parents and educators. Notifications are sent in the form of push notifications or email.

[0551] As a concrete example, suppose a smartphone is used to film children playing in a school classroom. This video data is sent to a server, which uses image recognition technology to analyze the children's facial expressions and actions. For example, if the server detects that a child is frightened or being subjected to aggressive behavior from other children, it immediately sends a notification to parents or educators.

[0552] An example of a prompt for a generative AI model is, "Based on this video, analyze whether the child may be being bullied." Using this prompt, the AI ​​model can analyze the video data and assess the likelihood of bullying.

[0553] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0554] Step 1:

[0555] The device uses its camera to capture images of children in real time. The captured video data is temporarily stored on the device. This video data becomes the input for the next processing step.

[0556] Step 2:

[0557] The device sends the captured video data to the server. The server takes the received video data as input and begins analysis using image recognition technology. Specifically, it uses libraries such as TensorFlow and OpenCV to analyze the facial expressions and movements of children in the video and extract features. The results of this analysis become the input for the next processing step.

[0558] Step 3:

[0559] The server uses a generative AI model to evaluate the possibility of bullying based on the extracted features. The prompt message, "Based on this video, analyze whether the child may be being bullied," is input to the AI ​​model, and signs of bullying are detected. This evaluation result becomes the input for the next processing step.

[0560] Step 4:

[0561] If signs of bullying are detected, the server sends the assessment results to the device. Based on the received assessment results, the device sends a notification to parents or educators. The notification is sent via push notification or email and informs them of the possibility of bullying. This notification is the final output.

[0562] (Example 2)

[0563] Next, we will describe Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0564] While early detection and rapid response to bullying are crucial, traditional methods often miss early signs of bullying, making it difficult to implement appropriate measures. In particular, there is a lack of mechanisms to monitor changes in children's facial expressions and behavior in real time and promptly notify relevant parties.

[0565] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0566] In this invention, the server includes means for collecting video and audio data, means for analyzing the collected data and analyzing the child's facial expressions and behavior, means for evaluating the possibility of bullying based on the analysis results, means for generating warning messages using a generative AI model, means for sending the generated warning messages to parents and educators, and means for proposing specific countermeasures to parents and educators. This makes it possible to detect signs of bullying early and notify relevant parties quickly.

[0567] "Video data" refers to visual information acquired using cameras and other recording devices, and is fundamental information for analyzing children's facial expressions and behavior.

[0568] "Audio data" refers to auditory information acquired using sound devices such as microphones, and is fundamental information for analyzing conversation content and tone.

[0569] "Analysis" refers to the process of processing collected video and audio data to analyze children's facial expressions and behavior.

[0570] A "generative AI model" refers to a program that uses artificial intelligence technology to perform natural language processing and generate warning messages based on the analysis results.

[0571] A "warning message" is a notification generated when potential bullying is detected, and it contains information to inform parents and educators about the situation.

[0572] "Parents and educators" refers to adults involved in the upbringing and education of children, and who have the responsibility to take appropriate action against signs of bullying.

[0573] "Specific measures" refer to action guidelines proposed when the possibility of bullying is detected, outlining the countermeasures that parents and educators should take.

[0574] A description of embodiments for carrying out this invention will be given.

[0575] The user generates a program for a bullying detection system. This program uses cameras and sensors installed within the school to monitor children's facial expressions and behavior in real time. The server collects video and audio data from these devices. Specifically, it processes video data acquired using cameras with image recognition software (e.g., OpenCV) to analyze children's facial expressions. It also processes audio data acquired using microphones with speech analysis software (e.g., Google Cloud Speech-to-Text) to analyze conversation content and tone.

[0576] The server uses a generative AI model (e.g., OpenAI's GPT-4) based on the analysis results to assess the likelihood of bullying. The generative AI model converts the analysis data into natural language and calculates a score indicating the likelihood of bullying. If signs of bullying are detected, the server generates a warning message. This warning message includes the date, time, and location where the bullying was detected, as well as details of the child's facial expressions and behavior.

[0577] The device sends generated warning messages to parents and educators via email or app notifications. This allows stakeholders to quickly understand the situation and take appropriate action. For example, if a specific child exhibits aggressive behavior towards another child in the schoolyard at 3 PM on a given day, a warning detailing the incident will be sent. An example of a prompt message might be: "Potential bullying detected in the schoolyard. Date and time: 3 PM, October 5, 2023. Location: Schoolyard. The child appeared angry and acted aggressively. What actions would you suggest to the parents?"

[0578] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0579] Step 1:

[0580] The server collects video and audio data from cameras and microphones installed within the school. It receives real-time visual and auditory information as input. This data is used as foundational information for monitoring children's facial expressions and behavior. Specifically, the server receives data streams from each device and stores them in a database.

[0581] Step 2:

[0582] The server processes the collected video data using image recognition software (e.g., OpenCV) to analyze the children's facial expressions. The video data collected in step 1 is used as input. For data processing, a face detection algorithm is applied to extract facial features. The output is the analysis results for each child's facial expression. Specifically, the server detects faces in each video frame and identifies expressions such as smiles and anger.

[0583] Step 3:

[0584] The server processes the collected audio data using speech analysis software (e.g., Google Cloud Speech-to-Text) to analyze the conversation content and tone. The audio data collected in step 1 is used as input. Data processing involves converting the audio to text and performing sentiment analysis. The output is the analysis results regarding the conversation content and tone. Specifically, the server converts the audio data to text and detects aggressive words and tones.

[0585] Step 4:

[0586] Based on the analysis results from steps 2 and 3, the server uses a generative AI model (e.g., OpenAI's GPT-4) to assess the likelihood of bullying. The inputs used are the analysis results of facial expressions and voice. The data calculation involves the generative AI model calculating a score indicating the likelihood of bullying. The output is an evaluation result regarding the likelihood of bullying. Specifically, the server prompts the generative AI model with the analysis results and calculates the score.

[0587] Step 5:

[0588] The server generates warning messages using a generative AI model. It uses the bullying possibility assessment results obtained in step 4 as input. For data processing, it creates warning messages in natural language. The output is a warning message to be sent to parents and educators. Specifically, the server generates a message that includes details such as the date, time, location, facial expressions, and actions.

[0589] Step 6:

[0590] The device sends the warning message received from the server to parents and educators via email or app notification. It uses the warning message generated in step 5 as input. Its output is a notification to the relevant parties. Specifically, the device sends the message to the designated contacts to prompt a quick response.

[0591] (Application Example 2)

[0592] Next, we will describe application example 2 of form example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".

[0593] In recent years, bullying among children in schools and homes has become a serious social problem. However, it is difficult to detect signs of bullying early and take appropriate measures. In particular, detecting the possibility of bullying from a child's facial expressions and behavior is a great burden for educators and parents. To solve this problem, a system is needed that can monitor children's behavior in real time and quickly detect signs of bullying.

[0594] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0595] In this invention, the server includes means for analyzing a child's facial expressions and behavior using image recognition technology to detect the possibility of bullying, means for monitoring a child's facial expressions and behavior in real time using a visual device, and means for analyzing signs of bullying using a generative AI model. This makes it possible to detect signs of bullying early and take appropriate measures quickly.

[0596] "Image recognition technology" is a technique that analyzes image data acquired using cameras and sensors to identify specific patterns and features.

[0597] "Children" refers to children of school age or those receiving education at home.

[0598] "Facial expression" refers to the emotions and reactions expressed through the movement of the facial muscles.

[0599] "Behavior" refers to actions and reactions that children take in their daily lives or in specific situations.

[0600] "Possible bullying" refers to signs of aggressive or inappropriate behavior towards others that can be inferred from a child's facial expressions and actions.

[0601] A "warning" is a notification or message issued to alert those involved when a potential bullying situation is detected.

[0602] "Guardian" refers to a child's parent or legal supervisor, who is responsible for the child's welfare and education.

[0603] "Education-related personnel" refers to individuals who hold positions related to the education and guidance of children in schools or educational institutions.

[0604] "Countermeasures" refers to specific actions or measures that should be taken when there is a possibility of bullying.

[0605] "Visual devices" refer to devices used to acquire and display images and videos.

[0606] A "generative AI model" refers to an algorithm or system that uses artificial intelligence technology to analyze data and generate specific patterns or results.

[0607] "Analysis" is the process of examining data in detail and deriving specific information or conclusions.

[0608] A "notification" is a message or alert used to convey specific information to relevant parties.

[0609] The system for carrying out this invention mainly consists of a server, a visual device, and a generative AI model. The server uses image recognition technology to analyze data on children's facial expressions and behavior acquired from the visual device and detect the possibility of bullying. The visual device is a device such as smart glasses and is responsible for monitoring children's behavior in real time. The generative AI model is built using frameworks such as TensorFlow or PyTorch and is used to analyze patterns in children's facial expressions and behavior.

[0610] The server receives video data transmitted from the visual device and analyzes the data using a generative AI model. If the analysis detects signs of bullying, the server generates a warning and sends a notification to parents and educators. This notification includes the date, time, and location of the detection, details of the child's facial expressions and behavior, and further specific countermeasures are suggested.

[0611] As a concrete example, imagine a school recess where a visual device monitors a group of students. If the generative AI model detects that a particular student frequently displays a sad expression, the server sends a notification to educators stating, "This may be bullying. Please investigate further." This notification allows educators to take prompt action.

[0612] An example of a prompt message for a generative AI model is: "Analyze the child's facial expression data and detect signs of bullying. If detected, output the date, time, location, and facial expression details, and suggest appropriate countermeasures."

[0613] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0614] Step 1:

[0615] The server receives video data transmitted in real time from the visual device. This video data includes the children's facial expressions and actions. The server preprocesses the video data and converts it into a format to which image recognition technology can be applied.

[0616] Step 2:

[0617] The server inputs pre-processed video data into a generative AI model. The generative AI model, built using TensorFlow, analyzes the children's facial expressions and behavior. The model extracts facial features and detects patterns that indicate the possibility of bullying. The analysis results in an output indicating whether or not bullying is likely.

[0618] Step 3:

[0619] Based on the output from the generating AI model, the server generates an alert if potential bullying is detected. This alert includes the date and time of detection, location, and details of the child's facial expressions and behavior. The server organizes this information and creates a notification message.

[0620] Step 4:

[0621] The server sends the created notification message to parents and educators. The notification is sent via email or app push notification. Recipients can check the notification and take appropriate action based on the child's situation.

[0622] Step 5:

[0623] After sending a notification, the server updates the training database for the generated AI model. Newly detected cases are added to the training data to improve the model's accuracy. This process continuously improves the accuracy of bullying detection.

[0624] (Example 3)

[0625] Next, we will describe Embodiment 3 of Embodiment Example 3. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0626] There is a need to accurately analyze children's emotional states and behavioral patterns to detect potential abnormal behavior at an early stage. However, conventional systems suffer from insufficient analytical accuracy, resulting in a lack of information for parents and educators to take appropriate action. Furthermore, the absence of self-learning capabilities based on analysis results makes it difficult to improve detection accuracy.

[0627] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 3 is realized by the following means.

[0628] In this invention, the server includes means for analyzing a child's emotional state and behavioral patterns using an image processing device, means for detecting the possibility of abnormal behavior based on the analyzed data, and means for providing the analysis results through an information display device. This enables more accurate analysis of a child's behavior, early detection of abnormal behavior, and rapid provision of information to parents and educators.

[0629] An "image processing device" is a device used to analyze image data and extract specific features.

[0630] "Children" refers to children attending educational institutions or minors.

[0631] "Emotional state" refers to the psychological state that can be inferred from a child's facial expressions and behavior.

[0632] "Behavioral patterns" refer to the tendencies of a series of actions and behaviors that a child exhibits.

[0633] "Abnormal behavior" refers to actions that deviate from normal behavioral patterns and includes signs of bullying and problematic behavior.

[0634] An "information display device" is a device used to visually display analysis results, and includes monitors, smartphones, and other similar devices.

[0635] The "self-learning function" is a feature that allows the system to automatically learn from newly collected data and improve the accuracy of its analysis.

[0636] A description of embodiments for carrying out this invention will be given.

[0637] The user builds a system to analyze children's emotional states and behavioral patterns. This system has the capability to analyze children's facial expressions and behavior in real time using an image processing device. Specifically, a camera installed on the terminal captures video of the children and sends the data to a server. The server preprocesses the received image data and extracts features using an image recognition AI model. This AI model is built using machine learning frameworks such as TensorFlow and PyTorch.

[0638] The server detects potential abnormal behavior based on extracted characteristic data. The detected results are provided to parents and educators via an information display device. This allows users to gain a concrete understanding of the child's situation and take appropriate action as needed.

[0639] For example, when a user monitors children's activities through cameras installed in school classrooms, the server receives image data in real time, and an AI model analyzes that data. The analysis results are provided to parents and educators through a dashboard.

[0640] An example of a prompt to be input to the generating AI model is, "Analyze the facial expressions and behavior of children in the classroom and create a report to detect signs of abnormal behavior." In response to this prompt, the AI ​​performs the specified task and provides the results. The specific processing flow in Example 3 is explained using Figure 15.

[0641] Step 1:

[0642] The device uses a camera to capture images of the children. The input is real-time video data. The device sends this video data to a server. Specifically, the camera continuously captures the facial expressions and movements of the children in the classroom.

[0643] Step 2:

[0644] The server preprocesses the video data received from the terminal. The input is the video data transmitted from the terminal. The server uses image processing libraries to adjust the image resolution and remove noise. The output is image data that has been prepared for analysis. Specifically, the server uses OpenCV to clear the image.

[0645] Step 3:

[0646] The server extracts features from pre-processed image data. The input is pre-processed image data. The server uses an image recognition AI model to extract the facial expressions and behavioral features of children as numerical data. The output is feature data. Specifically, the AI ​​model performs facial landmark detection and motion analysis.

[0647] Step 4:

[0648] The server detects potential abnormal behavior based on extracted feature data. The input is feature data. The server uses an AI model to execute an algorithm to detect signs of bullying and abnormal behavior. The output is an analysis result indicating potential abnormal behavior. Specifically, the AI ​​model identifies patterns of abnormal behavior.

[0649] Step 5:

[0650] The server provides analysis results to parents and educators via an information display device. The input is the analysis results indicating potential abnormal behavior. The server displays the results on a dashboard and mobile app. The output is the analysis results in a format that users can review. Specifically, the server displays the results via a web interface.

[0651] Step 6:

[0652] The server uses newly collected data to train an AI model. The input consists of analysis results and newly collected data. The server executes the AI ​​model's learning algorithm to improve its accuracy. The output is the improved AI model. Specifically, the server updates the model's parameters.

[0653] (Application Example 3)

[0654] Next, we will describe application example 3 of form example 3. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".

[0655] Bullying and isolation among minors are serious challenges in educational settings. Traditional methods make it difficult to detect these problems early and respond appropriately. In particular, there is a lack of means to monitor changes in minors' facial expressions and behavior in real time and respond quickly. Therefore, a new system is needed to ensure the safety and healthy development of minors.

[0656] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 3 is realized by the following means.

[0657] This invention includes a server that uses an image recognition algorithm to analyze the facial expressions and behavior of minors and detect the possibility of bullying; a server that issues a warning based on the detected possibility of bullying; a server that proposes countermeasures to parents and educators who have received a warning; a server that analyzes the facial expressions and behavior of minors in real time and sends a notification when an abnormality is detected; and a server that can immediately check the situation of minors via a smart device. This makes it possible to detect problems involving minors early and respond quickly and appropriately.

[0658] An "image recognition algorithm" is a method that uses computer vision technology to identify specific patterns or features from image data.

[0659] "Minor" refers to a person who has not reached the legal age of majority, and usually includes children and adolescents under the age of 18.

[0660] "Facial expression" is a visual representation of emotions and reactions shown through the movement and arrangement of facial muscles.

[0661] "Behavior" is a general term for actions and reactions that an individual exhibits in specific circumstances.

[0662] "Potential bullying" refers to a situation in which a particular behavior or situation has the potential to lead to aggressive or harmful behavior towards others.

[0663] A "warning" is a notification or alert that informs you of a potential risk or problem.

[0664] "Guardian" refers to a parent or legal guardian who is responsible for the life and education of a minor.

[0665] "Educational personnel" refers to teachers and staff involved in the education and guidance of minors in educational institutions.

[0666] "Measures" refer to specific actions or plans taken to address a particular problem or issue.

[0667] "Real-time" refers to a state where data processing and information provision occur immediately, with virtually no delay.

[0668] A "notification" is a means of conveying specific information or a message to a recipient.

[0669] A "smart device" refers to a portable or stationary electronic device that is capable of connecting to the internet and running applications.

[0670] The system for implementing this invention primarily uses a server, a smart device, and an image recognition algorithm. The server executes the image recognition algorithm to analyze the facial expressions and behavior of minors in real time. Specifically, it processes video data acquired from a camera using machine learning frameworks such as TensorFlow and PyTorch. This allows for the detection of changes in the facial expressions and behavior of minors and the assessment of the possibility of bullying.

[0671] The server issues a warning based on the detected potential bullying. The warning is notified to parents and educators via smart devices. Smart devices such as smartphones and smart glasses receive notifications in real time, allowing them to immediately check on the minor's situation.

[0672] As a concrete example, during school breaks, cameras monitor the behavior of minors. A server analyzes the video data and detects potential bullying if a particular minor is isolated or having trouble with other minors. If bullying is detected, the server immediately sends a notification to parents or educators and suggests appropriate measures.

[0673] By utilizing generative AI models, the server can self-learn based on new data, improving the accuracy of bullying detection. An example of a prompt would be, "I want to develop an application that analyzes children's facial expressions and behavior and sends notifications if abnormalities are detected. What algorithms and technologies should I use?"

[0674] The flow of the specific processing in Application Example 3 will be explained using Figure 16.

[0675] Step 1:

[0676] The server acquires video data from the camera. The input is real-time video footage of minors. The server preprocesses this video data and converts it into a format that is easy for image recognition algorithms to process. Specifically, it adjusts the video resolution and removes noise.

[0677] Step 2:

[0678] The server inputs pre-processed video data into an image recognition algorithm. The algorithm analyzes the facial expressions and behaviors of minors and extracts features. The output is feature data related to facial expressions and behaviors. Specifically, it performs facial landmark detection and behavioral pattern recognition.

[0679] Step 3:

[0680] The server evaluates the likelihood of bullying based on extracted feature data. The input is feature data, and the output is a score indicating the likelihood of bullying. The server uses a generative AI model to compare with past data and detect anomalies. Specifically, it applies an anomaly detection algorithm to evaluate deviations from normal behavioral patterns.

[0681] Step 4:

[0682] The server generates a warning if it determines there is a high probability of bullying. The input is a score indicating the likelihood of bullying, and the output is a warning message. The server sends this warning to smart devices. Specifically, it generates push notifications and sends alerts to parents and educators.

[0683] Step 5:

[0684] The device (smart device) receives a warning from the server. The input is the warning message, and the output is a notification displayed to the user. The device displays the notification on the screen so that the user can immediately check the situation. Specifically, it plays a notification sound and displays a pop-up on the screen.

[0685] Step 6:

[0686] The user checks the notification and takes action as necessary. The input is the warning message displayed on the device, and the output is the user's response. Based on the notification, the user checks the minor's situation and takes appropriate action. Specific actions might include educators going to the scene or parents contacting the school.

[0687] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0688] "Example of form 1"

[0689] This invention is a system that combines an emotion engine. Specifically, in addition to a device that captures a child's facial expressions and behavior, it is equipped with an emotion engine that recognizes the user's emotions. The emotion engine analyzes emotions from, for example, the user's tone of voice, word choice, and facial expressions. This emotion engine can determine that if a child is showing negative emotions such as anger or sadness, it is more likely to be a contributing factor to bullying, thereby improving the accuracy of detecting the possibility of bullying.

[0690] "Example of form 2"

[0691] This invention is a system that improves the accuracy of detecting the possibility of bullying based on the user's emotions detected by an emotion engine. Specifically, the emotion engine feeds back the detected emotion data to an image recognition AI, which then learns from that data. For example, if the emotion engine detects that a child is showing emotions such as anger or sadness, the AI ​​analyzes the child's facial expressions and behavior based on that information, enabling it to detect the possibility of bullying with greater accuracy.

[0692] "Example of form 3"

[0693] This invention is a system that adjusts the content of warnings based on the user's emotions detected by an emotion engine. Specifically, it adjusts the content of warnings and suggested countermeasures based on the emotion data detected by the emotion engine. For example, if the emotion engine detects that a child is showing strong anger, it can use that information to strengthen the content of the warning and suggest more specific countermeasures to parents and educators.

[0694] The following describes the processing flow for each example of the form.

[0695] "Example of form 1"

[0696] Step 1: A device that captures the child's facial expressions and actions records the child's actions.

[0697] Step 2: The captured data is sent to the emotion engine, which analyzes the user's emotions based on their voice tone, word choice, facial expressions, etc.

[0698] Step 3: The emotional engine determines that if a child is exhibiting negative emotions such as anger or sadness, it is more likely to be a contributing factor to bullying.

[0699] "Example of form 2"

[0700] Step 1: The emotion engine analyzes the child's emotions and feeds the results back to the image recognition AI.

[0701] Step 2: The image recognition AI learns based on feedback from the emotion engine.

[0702] Step 3: Based on the learning results, the AI ​​analyzes the child's facial expressions and behavior to detect the possibility of bullying with greater accuracy.

[0703] "Example of form 3"

[0704] Step 1: The emotion engine analyzes the child's emotions and adjusts the content of the warnings and suggested countermeasures based on the results.

[0705] Step 2: For example, if the emotion engine detects that the child is showing strong anger, it uses that information to strengthen the content of the warning.

[0706] Step 3: Based on that information, propose more specific measures to parents and educators.

[0707] (Example 1)

[0708] Next, we will describe Embodiment 1 of Embodiment Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0709] In modern society, bullying is a serious social problem, particularly affecting the physical and mental health of children. However, it is difficult to detect the signs of bullying early and to respond appropriately. Traditional methods can lead to delays in detecting bullying and escalating the damage, so there is a need to quickly and accurately detect the possibility of bullying and take appropriate measures.

[0710] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0711] In this invention, the server includes means for acquiring a child's facial expressions and behavior using a camera; means for analyzing the acquired data using image recognition technology to detect the possibility of bullying; and means for analyzing audio data to identify emotions. This makes it possible to quickly and accurately detect the possibility of bullying and take appropriate measures.

[0712] A "recording device" is a device used to capture a child's facial expressions and behavior, and includes devices such as cameras and smartphones.

[0713] "Image recognition technology" is a technique that analyzes acquired image data to identify specific patterns and features, and is used to analyze children's facial expressions and behavior.

[0714] "Audio data" refers to data that records children's voices and surrounding sounds, and is used for analysis to identify emotions.

[0715] "Emotion analysis technology" is a technology that analyzes voice data to identify the speaker's emotions, and it is a technology that determines emotions by analyzing the tone of voice and word choice.

[0716] "Potential bullying" is an indicator that shows the possibility that a child is being bullied or is involved in bullying, and it is evaluated by integrating the results of image recognition technology and emotion analysis technology.

[0717] A "warning" is a notification issued to parents and educators when the possibility of bullying is detected, and it contains information to encourage appropriate measures.

[0718] The "means of proposing countermeasures" refer to a function that provides specific countermeasures to parents and educators based on the detected possibility of bullying.

[0719] This invention is a system for quickly and accurately detecting the possibility of bullying and taking appropriate countermeasures. The system consists of a camera, image recognition technology, and voice data analysis technology.

[0720] Users utilize cameras and smartphones as recording devices for use at school or home. These devices capture children's facial expressions and behavior in real time. The devices then transmit the acquired image data to image recognition technology. This technology analyzes children's facial expressions and behavior using libraries such as TensorFlow and OpenCV. Specifically, it extracts facial feature points and determines whether the expression indicates fear or sadness.

[0721] The server receives audio data transmitted from the terminal and analyzes it using emotion analysis technology. This technology analyzes the tone and word choice of the voice to identify the emotions the user is expressing. For example, it analyzes the pitch and speed of the voice to detect negative emotions such as anger or sadness.

[0722] The server integrates data obtained from image recognition and sentiment analysis technologies to assess the likelihood of bullying. If the assessment determines that there is a high probability of bullying, the server issues a warning to the user. After receiving the warning, the user can take appropriate measures at school or home as needed.

[0723] As a concrete example, suppose a user is using a smartphone to film their child. If image recognition technology analyzes the child's facial expression and determines that the child is depressed, the server will use that information to notify the user that bullying may be occurring. Similarly, if emotion analysis technology analyzes the tone of the child's voice and determines that the child is angry, the system will also determine that bullying may be occurring.

[0724] An example of a prompt to input into a generative AI model is, "Please tell me how to analyze a child's facial expressions and tone of voice to detect the possibility of bullying." This prompt allows the AI ​​model to provide information about specific analysis methods and the techniques to be used.

[0725] The flow of the specific processing in Example 1 will be explained using Figure 17.

[0726] Step 1:

[0727] The device uses a camera or smartphone to capture the child's facial expressions and actions in real time. The input is the captured image data. The device prepares this image data to be transmitted to image recognition technology. Specifically, the device converts the image data to an appropriate format and performs preprocessing for analysis.

[0728] Step 2:

[0729] The device analyzes captured image data using image recognition technology. The input is the image data acquired in step 1. The image recognition technology uses libraries such as TensorFlow and OpenCV to extract facial feature points and determine whether the expression is frightened or depressed. The output is the result of the facial expression analysis. Specifically, the device applies a face detection algorithm and quantifies the facial features.

[0730] Step 3:

[0731] The server receives audio data transmitted from the terminal and analyzes it using sentiment analysis technology. The input is audio data. Sentiment analysis technology analyzes the tone and word choice of the voice to identify the emotions expressed by the user. The output is the result of the sentiment analysis. Specifically, the server performs spectral analysis on the audio data to detect changes in pitch and speed.

[0732] Step 4:

[0733] The server integrates data obtained from image recognition and emotion analysis technologies to assess the likelihood of bullying. The input consists of facial expression analysis results and emotion analysis results. The server integrates this data to calculate an index indicating the likelihood of bullying. The output is the assessment result of the likelihood of bullying. Specifically, the server uses statistical methods to weight each analysis result and perform an overall evaluation.

[0734] Step 5:

[0735] The server issues a warning to the user if it determines that bullying is highly likely. The input is the evaluation result of the likelihood of bullying. Based on the evaluation result, the server generates a warning message and notifies the user. The output is the warning message. Specifically, the server sends the message to the user's terminal via the notification system.

[0736] (Application Example 1)

[0737] Next, we will describe Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0738] When children are at risk of being bullied at school or at home, it is crucial to detect early signs of bullying and take appropriate measures. However, traditional methods often miss signs of bullying, making it difficult to adequately ensure the safety of children.

[0739] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0740] This invention includes a server that uses image recognition technology to analyze a child's facial expressions and behavior to detect the possibility of bullying, means of analyzing a child's tone of voice and language use using emotion analysis technology to evaluate the possibility of bullying, and means of monitoring the child's safety in real time via a smart device. This makes it possible to detect signs of bullying early and quickly notify parents and educators.

[0741] "Image recognition technology" is a technique that analyzes image data acquired using cameras and sensors to identify specific patterns and features.

[0742] "Children" refers to children of school age or those receiving education at home.

[0743] "Facial expression" refers to the emotions and reactions expressed through the movement of the facial muscles.

[0744] "Behavior" refers to actions and behaviors that children perform in their daily lives.

[0745] "Potential bullying" refers to the possibility that a child may suffer mental or physical harm due to aggressive behavior or words from others.

[0746] "Emotion analysis technology" is a technique that analyzes voice and text data to estimate the emotional state of the speaker.

[0747] "Voice tone" refers to the characteristics of a speaker's voice, expressed through factors such as pitch, volume, and intonation.

[0748] "Word choice" refers to the way a speaker selects and expresses themselves using words.

[0749] A "smart device" refers to a portable electronic device that has internet connectivity and can run various applications.

[0750] "Real-time" refers to a state where data acquisition and processing occur instantly, and results are obtained without delay.

[0751] "Guardian" refers to a child's parent or the person legally responsible for caring for that child.

[0752] "Education professionals" refers to those who are responsible for duties related to the education and guidance of children in schools and educational institutions.

[0753] The system for implementing this invention mainly consists of three elements: a server, a terminal, and a user. The server uses image recognition technology and emotion analysis technology to analyze the child's facial expressions, behavior, tone of voice, and speech. Specifically, the server uses TensorFlow to analyze image data acquired from cameras and sensors to identify the child's facial expressions and behavior. It also uses IBM Watson to analyze audio data and estimate the child's emotional state.

[0754] The device functions as a smart device, monitoring children's safety in real time. It receives analysis results from the server and immediately sends notifications to parents and educators if potential bullying is detected. This enables a rapid response.

[0755] As a parent or educator, the user can receive notifications from the device and take appropriate measures to ensure the safety of children. For example, if the device detects a situation during school recess where a child is being verbally abused by another child, the server will determine from the child's facial expression that they are "scared" and analyze the tone of their voice as "angry." Based on these results, the device will determine that there is a "possible case of bullying" and send a notification to the parent or guardian.

[0756] An example of a prompt to input into a generative AI model is, "Analyze the facial expressions and tone of voice of the child in this video and assess the possibility of bullying." This prompt allows the server to perform an appropriate analysis and assess the possibility of bullying.

[0757] The flow of a specific process in Application Example 1 will be explained using Figure 18.

[0758] Step 1:

[0759] The server receives image and audio data transmitted from the terminal. The input consists of facial expressions and voice data of the children acquired through the camera and microphone. To analyze this data, the server first preprocesses the data and removes noise.

[0760] Step 2:

[0761] The server analyzes preprocessed image data using TensorFlow. The input is denoised image data. The server uses image recognition technology to identify the children's facial expressions and behaviors and extract features that indicate the possibility of bullying. The output is the analysis results regarding facial expressions and behaviors.

[0762] Step 3:

[0763] The server analyzes pre-processed audio data using IBM Watson. The input is audio data with noise removed. The server uses emotion analysis technology to analyze the tone of the child's voice and word choice, and estimates their emotional state. The output is the analysis result regarding emotions.

[0764] Step 4:

[0765] The server integrates the analysis results of image and audio data to assess the possibility of bullying. The input consists of analysis results regarding facial expressions, behavior, and emotions. The server combines this data to make a comprehensive judgment on the possibility of bullying. The output is the assessment result regarding the possibility of bullying.

[0766] Step 5:

[0767] The device receives the assessment results regarding the possibility of bullying, which are sent from the server. The input is the assessment results from the server. Based on these results, if the device determines that there is a high possibility of bullying, it sends a notification to parents or educators. The output is the notification message.

[0768] Step 6:

[0769] The user receives notifications from the device and takes appropriate measures to ensure the child's safety. The input is the notification message from the device. Based on this information, the user checks the child's situation and intervenes as needed. The output is the specific action taken to ensure the child's safety.

[0770] (Example 2)

[0771] Next, we will describe Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0772] There are challenges in the early detection of bullying and the proposal of appropriate countermeasures. In particular, there is a need to accurately capture changes in children's emotions and behavior and detect the possibility of bullying with high precision. It is also important to quickly provide concrete countermeasures based on the detection results.

[0773] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0774] In this invention, the server includes means for collecting emotional data, means for feeding the collected emotional data back to image recognition technology, and means for analyzing a child's facial expressions and behavior using image recognition technology to detect the possibility of bullying. This makes it possible to detect the possibility of bullying with high accuracy and to provide appropriate countermeasures quickly.

[0775] "Emotional data" refers to information about a child's emotions obtained from their facial expressions and behavior.

[0776] "Image recognition technology" is a technique that analyzes image data to identify specific patterns or features.

[0777] "Possibility of bullying" refers to the possibility of bullying occurring, inferred from the analysis of a child's facial expressions and behavior.

[0778] A "warning" is a cautionary message sent to parents and educators when the possibility of bullying is detected.

[0779] "Generative AI technology" is a technology that uses artificial intelligence to generate new information and proposals.

[0780] A "learning function" is a feature that allows a system to improve itself based on past data.

[0781] This invention is a system that detects the possibility of bullying with high accuracy and provides appropriate countermeasures. Specific embodiments are shown below.

[0782] The user uses their device to record the child's facial expressions and behavior using an emotion engine. The emotion engine uses sensors such as cameras and microphones to detect the child's emotions, such as anger or sadness, in real time. For example, if a child looks sad at school, video and audio data of that moment are collected.

[0783] The device sends the collected emotional data to a server. The server receives this data and feeds it back into image recognition technology. The image recognition technology learns from the received emotional data and analyzes the child's facial expressions and behavior in detail. This analysis is used to assess the possibility of bullying.

[0784] If potential bullying is detected, the server will issue a warning to parents or educators. This warning will be sent as an email or app notification and will include details such as the date, time, and location of the detection, as well as the child's facial expressions and behavior.

[0785] Furthermore, the server uses generative AI technology to suggest specific countermeasures. For example, if you input a prompt message such as "What should you do if your child looks sad at school?", the AI ​​will suggest ways for parents to encourage communication with their children and suggest ways for educators to strengthen school monitoring systems.

[0786] In this way, the system can quickly and effectively detect potential bullying and provide appropriate countermeasures.

[0787] The flow of the specific processing in Example 2 will be explained using Figure 19.

[0788] Step 1:

[0789] The user uses a device to record the child's facial expressions and behavior using an emotion engine. Video and audio data from the camera and microphone are used as input. The emotion engine analyzes this data to detect the child's emotional state (e.g., anger, sadness). The detected emotion data is generated as output.

[0790] Step 2:

[0791] The device sends the collected emotion data to the server. The emotion data generated in step 1 is used as input. The server receives this data and feeds it back to the image recognition technology. The feedback emotion data is passed back to the image recognition technology as output.

[0792] Step 3:

[0793] The server uses image recognition technology to learn from the feedbacked emotional data. The feedbacked emotional data is used as input. The image recognition technology analyzes the data, examining the child's facial expressions and behavior in detail. The output generates an assessment of the likelihood of bullying.

[0794] Step 4:

[0795] If the server detects a potential bullying incident, it will issue a warning to parents or educators. The evaluation results generated in step 3 are used as input. The server will send a warning via email or app notification, including the date, time, and location of the detection, as well as details of the child's facial expressions and behavior. An alert message is generated and sent as output.

[0796] Step 5:

[0797] The server uses AI generation technology to propose specific countermeasures. The warning message generated in step 4 is used as input. Based on the prompt "What should be done if a child looks sad at school?", the AI ​​generation technology proposes specific countermeasures for parents and educators. The proposed countermeasures are generated and provided as output.

[0798] (Application Example 2)

[0799] Next, we will describe application example 2 of form example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".

[0800] While early detection and appropriate response are crucial in addressing bullying among children, conventional methods struggle with detection and risk overlooking early signs of bullying. Furthermore, there is a need for a system that can accurately detect potential bullying and promptly notify relevant parties.

[0801] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0802] In this invention, the server includes means for analyzing a child's facial expressions and behavior using an image recognition algorithm to detect the possibility of bullying, means for analyzing video in real time and extracting emotional data, and means for learning based on the extracted emotional data and evaluating the possibility of bullying. This makes it possible to detect the possibility of bullying with high accuracy and to quickly notify parents and educators.

[0803] An "image recognition algorithm" is a computational method used to analyze video data acquired from cameras and sensors and identify specific patterns or features.

[0804] "Children's facial expressions and behavior" refers to the visual characteristics of children, such as their facial expressions, body movements, and posture.

[0805] "Potential bullying" refers to the probability or risk of a child exhibiting signs of aggressive or inappropriate behavior towards other children.

[0806] "Emotional data" refers to information that indicates emotional states such as anger, sadness, and joy, extracted from a child's facial expressions and behavior.

[0807] "Learning" is the process by which an algorithm improves its performance based on past data and newly acquired data.

[0808] "Evaluating" means judging the possibility of bullying or the emotional state based on specific criteria and deriving a result.

[0809] "Generating and sending notifications" is the process of creating and sending messages to relevant parties to provide warnings or information based on detected data.

[0810] The system for implementing this invention consists of a server, a terminal, and a user. The server acquires video data from surveillance cameras installed within the school and analyzes the children's facial expressions and behavior in real time using an image recognition algorithm. Specifically, the server uses software such as OpenCV and TensorFlow to extract facial features from the video data and generates emotion data using an emotion engine.

[0811] The generated emotion data is used in a learning process within the server to update the model for evaluating the possibility of bullying. Based on this evaluation, the server generates a warning and sends a notification to the devices of parents and educators. The notification includes the date, time, and location of the detection, details of the child's facial expressions and behavior, and specific countermeasures are suggested.

[0812] The terminal is a device such as a smartphone or tablet that receives notifications and displays warnings to the user. The user can check the notification and take the suggested action.

[0813] For example, if a camera captures a scene in a school hallway where a student is showing anger towards another student, the server will detect that anger using an emotion engine and analyze the scene with an image recognition algorithm. If it is determined that there is a high probability of bullying, the server will generate a notification stating, "Possible bullying has been detected in the hallway. Please check the details," and send it to the relevant parties.

[0814] An example of a prompt message is: "Analyze the facial expressions of children in the hallway. If feelings of anger or sadness are detected, use that information to assess the possibility of bullying and generate a notification."

[0815] The flow of a specific process in Application Example 2 will be explained using Figure 20.

[0816] Step 1:

[0817] The server acquires video data in real time from surveillance cameras installed within the school. The input is a video stream from the cameras, and the output is analyzable image data. To process this data, the server divides the video into frames.

[0818] Step 2:

[0819] The server uses an image recognition algorithm to detect children's faces from the acquired frames. The input is the image data obtained in step 1, and the output is the face's position information and features. The server uses OpenCV to extract facial features and determine the face's position.

[0820] Step 3:

[0821] The server uses an emotion engine to generate emotion data from detected facial features. The input is the facial features obtained in step 2, and the output is data indicating the type and intensity of emotion. The server feeds this data back into the generating AI model to perform emotion analysis.

[0822] Step 4:

[0823] The server inputs the generated emotion data into an image recognition algorithm to evaluate the possibility of bullying. The input is emotion data, and the output is a score indicating the possibility of bullying. The server uses TensorFlow to analyze the emotion data and quantify the possibility of bullying.

[0824] Step 5:

[0825] The server generates and sends notifications to parents and educators based on the bullying possibility score. The input is the score obtained in step 3, and the output is a warning message. The server uses a generative AI model to create notifications that suggest specific actions based on the prompt text.

[0826] Example prompt: "Analyze students' facial expressions in the hallway. If anger or sadness is detected, use this information to assess the possibility of bullying and generate a notification."

[0827] (Example 3)

[0828] Next, we will describe Embodiment 3 of Embodiment Example 3. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0829] Conventional technologies have struggled to accurately analyze an individual's emotional state and provide appropriate warnings and countermeasures. Furthermore, they lacked the learning capabilities necessary to detect potential bullying with high accuracy, hindering the ability of stakeholders to respond quickly and appropriately.

[0830] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 3 is realized by the following means.

[0831] In this invention, the server includes means for analyzing an individual's emotional state using image processing technology, means for adjusting the content of a warning based on the analyzed emotional data, and means for proposing specific countermeasures to the person who received the warning. This makes it possible to accurately grasp an individual's emotional state and quickly provide appropriate warnings and countermeasures.

[0832] "Image processing technology" refers to techniques for analyzing digital images and extracting or recognizing specific information.

[0833] "Emotional state" refers to the psychological or emotional state an individual exhibits at a particular moment in time.

[0834] "Warning content" refers to information that serves as a warning or instruction to relevant parties, generated based on the analyzed data.

[0835] "Specific countermeasures" refer to proposals that instruct relevant parties on the actions and measures they should take depending on the detected situation.

[0836] "Self-learning" is the process by which a system automatically learns from new data and improves its performance and accuracy.

[0837] "Stakeholders" refers to individuals such as parents and educators who are the recipients of the analysis results and warnings.

[0838] A description of embodiments for carrying out this invention will be given.

[0839] The server receives data on an individual's facial expressions and actions transmitted from the device. The device collects data in real time using cameras and sensors and sends it to the server. This data is acquired when the individual is active in a specific environment.

[0840] The server analyzes the received data using image processing technology. This analysis utilizes image recognition AI based on frameworks such as TensorFlow and PyTorch. The image recognition AI infers emotional states from an individual's facial expressions and identifies emotions such as anger, sadness, and joy.

[0841] The analyzed emotion data is passed to the emotion engine on the server. The emotion engine quantifies the individual's emotional state based on the analysis results and adjusts the warning content. For example, if the emotion engine detects that an individual is showing strong anger, the server uses that information to strengthen the warning and propose specific countermeasures to those involved.

[0842] Furthermore, the server performs self-learning based on new data, improving the detection accuracy of the image recognition AI. This allows the server to continuously improve its accuracy and provide more precise analysis results.

[0843] As a concrete example, let's look at an example of a prompt sentence to be input to the generating AI model: "Analyze the facial expressions an individual displays during a specific activity and detect emotions such as anger or sadness. Based on the results, propose appropriate measures for those involved." In response to this prompt sentence, the server utilizes image recognition AI and emotion engine to analyze the individual's emotional state in detail and provide appropriate information. The flow of specific processing in Example 3 will be explained using Figure 21.

[0844] Step 1:

[0845] The device uses cameras and sensors to collect data on an individual's facial expressions and behavior in real time. The collected data is transmitted to a server in image or video format. The input is raw data of the individual's facial expressions and behavior, and the output is the transmission of data to the server.

[0846] Step 2:

[0847] The server analyzes the data received from the terminal using image processing technology. Specifically, an image recognition AI using frameworks such as TensorFlow or PyTorch analyzes the received image data and infers the emotional state from the individual's facial expressions. The input is image data sent from the terminal, and the output is the analyzed emotional data.

[0848] Step 3:

[0849] The server passes the analyzed emotional data to the emotion engine, which quantifies the individual's emotional state. Based on the analysis results, the emotion engine identifies emotions such as anger, sadness, and joy, and adjusts the warning content accordingly. The input is the analysis results from the image recognition AI, and the output is the quantified emotional data.

[0850] Step 4:

[0851] The server generates warnings based on quantified emotion data and suggests specific actions for those involved. For example, if strong anger is detected, the server generates a warning such as, "An individual is exhibiting strong anger. Assess the situation and take steps to calm them down." The input is quantified emotion data, and the output is the generated warning and suggested actions.

[0852] Step 5:

[0853] The server performs self-learning based on new data to improve the detection accuracy of the image recognition AI. Specifically, it compares past data with new data and adjusts the model to reduce misrecognition. The input is past data and new data, and the output is the adjusted AI model.

[0854] Step 6:

[0855] The server sends the generated warnings and suggestions to the terminal, notifying the relevant user. The terminal displays the received information on its screen, allowing the user to understand the situation. The input is the generated warnings and suggestions, and the output is the information displayed on the terminal.

[0856] (Application Example 3)

[0857] Next, we will describe application example 3 of form example 3. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".

[0858] There is a need to detect bullying and emotional changes among minors early and to take appropriate action. However, conventional methods may miss signs of bullying, and there are challenges in responding appropriately to emotional changes.

[0859] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 3 is realized by the following means.

[0860] This invention includes a server that uses image recognition technology to analyze the facial expressions and behavior of minors and detect the possibility of bullying; a server that issues a warning based on the detected possibility of bullying; a server that proposes countermeasures to parents and educators who receive the warning; a server that adjusts the content of the warning based on the minor's emotional data; and a server that captures the minor's facial expressions in real time and notifies them of the analysis results. This enables a swift and appropriate response to protect the safety and welfare of minors.

[0861] "Image recognition technology" is a technique that uses computer vision to extract and analyze specific information from image data.

[0862] A "minor" refers to a person who has not reached the legal age of majority.

[0863] "Facial expression" refers to the emotions and reactions shown through the movement of the facial muscles.

[0864] "Action" refers to a series of actions or reactions performed by an individual.

[0865] "Possible bullying" refers to the possibility that a minor is experiencing inappropriate behavior or remarks from others.

[0866] A "warning" is a notification or message intended to draw attention to a specific situation.

[0867] A "guardian" refers to a person who has the responsibility to supervise the life and education of a minor.

[0868] "Education-related personnel" refers to individuals who hold positions in educational institutions that involve the education of minors.

[0869] "Measures" refer to specific actions or means taken to address a particular problem or situation.

[0870] "Emotional data" refers to information that expresses an individual's emotional state using numerical values ​​or categories.

[0871] "Real-time" refers to processing or responding instantly without delay.

[0872] A "notification" is a message or alert used to convey specific information to a recipient.

[0873] The system for implementing this invention mainly consists of a server and a terminal. The server uses image recognition technology to analyze the facial expressions and behavior of minors and detect the possibility of bullying. Specifically, the terminal's camera is used to capture the minor's facial expressions in real time, and the data is sent to the server. The server uses OpenCV as the image recognition technology and the Microsoft Azure Emotion API for emotion analysis. This allows the system to extract emotional data from minors and evaluate the possibility of bullying.

[0874] The server issues warnings based on detected bullying and suggests countermeasures to parents and educators. The content of the warnings is adjusted based on the minor's emotional data. For example, if the minor is showing strong anger, the warning is strengthened and specific countermeasures are suggested.

[0875] The device functions as a smartphone or smart glasses, capturing the facial expressions of minors in real time. The analysis results are sent as notifications to parents and educators. This enables quick and appropriate responses to protect the safety and well-being of minors.

[0876] As a concrete example, if a minor displays strong anger in a classroom, the device captures their facial expression and sends it to a server. The server performs an emotion analysis, and if it determines that the expression is "anger," it sends a notification to the parent saying, "Your child is showing strong anger. Please listen to them."

[0877] An example of a prompt message is: "Analyze the child's facial expression data and determine their emotion. If anger is detected, generate a message to send a notification to the parent / guardian."

[0878] The flow of the specific processing in Application Example 3 will be explained using Figure 22.

[0879] Step 1:

[0880] The device uses a camera to capture the facial expressions of minors in real time. The input is video data from the camera, and the output is image data. The device then prepares to send this image data to a server.

[0881] Step 2:

[0882] The server receives image data transmitted from the terminal. The input is image data from the terminal, and the output is analyzable image information. The server processes this image information using image recognition technology (OpenCV) to analyze the facial expressions of minors.

[0883] Step 3:

[0884] The server inputs facial expression data extracted using image recognition technology into an emotion analysis engine (Microsoft Azure Emotion API). The input is facial expression data, and the output is emotion data. Based on this emotion data, the server determines the emotional state of the minor.

[0885] Step 4:

[0886] The server analyzes sentiment data and assesses the likelihood of bullying. The input is sentiment data, and the output is an assessment result: "bullying is likely" or "bullying is not likely." Based on this assessment, the server generates warnings as needed.

[0887] Step 5:

[0888] The server notifies parents and educators of the generated warnings. The input is the evaluation result, and the output is the warning message. The server adjusts the content of the warnings and suggests specific countermeasures based on the minor's emotional data.

[0889] Step 6:

[0890] The user receives notifications from the server and takes appropriate action based on the minor's situation. The input is the notification from the server, and the output is the user's response. Based on the notification, the user interacts with the minor and, if necessary, collaborates with educators to resolve the problem.

[0891] (Other examples)

[0892] Since this is the same as the specific processing described in the other embodiments of the first embodiment above, the explanation will be omitted.

[0893] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0894] The data generation model 58 is a form of so-called generative AI (Artificial Intelligence). One example of the data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0895] Other examples of generative AI include Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) are examples.

[0896] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0897] [Third Embodiment]

[0898] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0899] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0900] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0901] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0902] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0903] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0904] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0905] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0906] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0907] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0908] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0909] Next, the identification process performed by the identification processing unit 290 of the data processing device 12 will be described.

[0910] "Example of form 1"

[0911] One embodiment of the present invention involves equipping a camera or smartphone, or other imaging device, with image recognition AI for use in schools or homes. This image recognition AI analyzes a child's facial expressions and behavior in real time to detect the possibility of bullying. For example, it detects the possibility of bullying if it determines that a child's facial expression indicates fear or sadness, or if it determines that the child is being subjected to aggressive behavior from other children.

[0912] "Example of form 2"

[0913] Based on the detected potential bullying, warnings are issued to parents and educators. These warnings are sent via email, app notifications, etc. The warnings include details such as the date, time, and location where the potential bullying was detected, as well as details of the child's facial expressions and behavior. Specific countermeasures are also suggested. For example, parents may be advised to take measures to encourage communication with their children, and educators may be advised to strengthen school monitoring systems and offer counseling.

[0914] "Example of form 3"

[0915] Furthermore, the system provides parents and educators with analysis results of children's facial expressions and behavior. This allows parents and educators to understand the child's situation more concretely and take appropriate action. It also includes a learning function to improve the accuracy of detecting potential bullying. This learning function allows the image recognition AI to self-learn based on new data and improve detection accuracy.

[0916] The following describes the processing flow for each example of the form.

[0917] "Example of form 1"

[0918] Step 1: Equipate image recognition AI into cameras, smartphones, and other imaging devices for use in schools and homes.

[0919] Step 2: Capture the child's facial expressions and actions in real time through the camera and send them to the image recognition AI.

[0920] Step 3: The image recognition AI analyzes the transmitted image data and detects the possibility of bullying from the child's facial expressions and behavior.

[0921] "Example of form 2"

[0922] Step 1: If the image recognition AI detects a potential bullying incident, it issues a warning to parents and educators.

[0923] Step 2: The warning will be sent via email, app notification, etc.

[0924] Step 3: The warning will include details such as the date and location where the bullying was detected, as well as the child's facial expressions and behavior.

[0925] Step 4: Specific measures are proposed. For example, parents may be advised to take steps to encourage communication with their children, and educators may be advised to strengthen school supervision systems and provide counseling.

[0926] "Example of form 3"

[0927] Step 1: Provide parents and educators with the results of the analysis of the child's facial expressions and behavior.

[0928] Step 2: Parents and educators will be able to understand the child's situation more concretely and take appropriate action.

[0929] Step 3: Includes a learning function to improve the accuracy of detecting potential bullying.

[0930] Step 4: This learning function allows the image recognition AI to self-learn based on new data and improve detection accuracy.

[0931] (Example 1)

[0932] Next, we will describe Embodiment 1 of Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0933] In modern schools and homes, children are increasingly likely to encounter bullying, but there is a challenge in detecting signs of bullying early and taking appropriate measures. In particular, there is a need for technology that can monitor changes in children's facial expressions and behavior in real time and quickly detect the possibility of bullying.

[0934] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0935] In this invention, the server includes means for collecting video data using a camera, means for transmitting the collected video data via communication means, and means for analyzing a child's facial expressions and behavior using image recognition technology to detect the possibility of bullying. This makes it possible to detect signs of bullying early and take appropriate measures.

[0936] A "recording device" is a device used to collect video data, and includes devices such as cameras and smartphones.

[0937] "Video data" refers to digital data containing visual information collected by a camera or camera.

[0938] "Communication methods" refer to the technologies and protocols used to send and receive data, and include methods such as the internet and wireless communication.

[0939] "Image recognition technology" is a technology that allows computers to analyze image data and identify specific patterns or features.

[0940] "Analysis of facial expressions and behavior" is the process of using image recognition technology to analyze a child's facial expressions and body movements to identify specific emotions and behaviors.

[0941] "Detecting the possibility of bullying" is a process of determining the likelihood of bullying occurring based on analyzed facial expression and behavioral data.

[0942] "Means of issuing warnings" refer to methods or devices used to alert those involved when the possibility of bullying is detected.

[0943] "Means of proposing countermeasures" refer to methods and devices for presenting specific countermeasures to parents and educators based on the detected possibility of bullying.

[0944] "Learning function" refers to a function that allows the system to improve the accuracy of its analysis based on past data, and includes machine learning algorithms.

[0945] This invention is a system for use in schools and homes that analyzes children's facial expressions and behavior in real time to detect potential bullying. The system consists of a camera, communication means, and image recognition technology.

[0946] The terminal uses cameras, smartphones, and other recording devices to collect video data of children in real time. The collected video data is transmitted to a server via a secure communication method. The communication method uses the internet or wireless communication technology.

[0947] The server uses image recognition technology to analyze the received video data. This technology leverages generative AI models built using machine learning frameworks such as TensorFlow and PyTorch. Specifically, it uses OpenCV to extract facial features and analyze changes in facial expressions. It also applies motion recognition algorithms to detect aggressive behavior.

[0948] Users receive analysis results from the server and are warned if bullying is suspected. Furthermore, the system suggests specific measures to parents and educators based on the detected results. This makes it possible to detect signs of bullying early and take appropriate action.

[0949] For example, a user might monitor children's behavior through a camera installed in the classroom. The AI ​​would warn of potential bullying if it detects that a child frequently looks frightened or if it detects behavior such as another child unilaterally pushing or shoving. Examples of prompts to input into the generating AI model could be, "What should I do if a child looks frightened?" or "Please tell me how to deal with aggressive behavior that is detected."

[0950] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0951] Step 1:

[0952] The device uses a camera or smartphone to collect video data of children in real time. The input is video data obtained from the recording device. Specifically, the device ensures high-resolution video so that facial expressions and actions can be clearly identified. The output is the collected video data.

[0953] Step 2:

[0954] The terminal sends the collected video data to the server. The input is the video data obtained in step 1. Specifically, the terminal uses a secure communication protocol (e.g., HTTPS) to send the data and protect privacy. The output is the video data sent to the server.

[0955] Step 3:

[0956] The server analyzes the received video data using image recognition technology. The input is the video data sent in step 2. Specifically, the server utilizes generative AI models using TensorFlow and PyTorch, extracts facial features using OpenCV, and analyzes changes in facial expressions. It also applies a motion recognition algorithm to detect aggressive behavior. The output is the analyzed facial expression and behavior data.

[0957] Step 4:

[0958] The server determines the possibility of bullying based on the analysis results and issues a warning if necessary. The input is the analysis data obtained in step 3. Specifically, if the server determines that bullying is possible, it generates a warning message. The output is the warning message.

[0959] Step 5:

[0960] The user receives a warning from the server and considers appropriate action. The input is the warning message generated in step 4. Specifically, the user checks the warning on their smartphone or PC and takes action, such as meeting with the child if necessary. The output is the user's response.

[0961] (Application Example 1)

[0962] Next, we will describe Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0963] In modern society, bullying of children at school and at home is a serious problem. However, it is not easy to detect signs of bullying early and take appropriate measures. In particular, there is a lack of means to monitor children's emotions and behavior in real time and quickly detect the possibility of bullying, which often leads to delays in early detection and response to bullying. To solve this problem, a system is needed that analyzes children's emotions and behavior in real time, immediately detects signs of bullying, and notifies those involved.

[0964] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0965] In this invention, the server includes means for analyzing a child's emotions and actions using image recognition technology to detect the possibility of bullying, means for monitoring a child's emotions and actions in real time using a mobile information terminal, and means for sending a notification when signs of bullying are detected. This makes it possible to detect signs of bullying in a child early and quickly notify the relevant parties.

[0966] "Image recognition technology" is a technique that analyzes image data acquired using cameras and sensors to identify specific patterns and features.

[0967] "Children's emotions and actions" refer to the psychological state and behaviors that can be interpreted from a child's facial expressions and body movements.

[0968] "Possible bullying" refers to signs that indicate a child may be experiencing aggressive behavior or psychological pressure from others.

[0969] A "portable information terminal" is an electronic device that is portable and capable of processing information, such as a smartphone or tablet.

[0970] "Real-time monitoring" means observing and analyzing the ongoing situation immediately.

[0971] "Sending a notification" means immediately informing relevant parties of specific information.

[0972] The system for implementing this invention mainly consists of a server and a mobile information terminal. The server utilizes image recognition technology to analyze children's emotions and actions and detect the possibility of bullying. Specifically, the server uses image recognition libraries such as TensorFlow and OpenCV to process image data acquired from cameras and sensors in real time. This makes it possible to identify signs of bullying from children's facial expressions and actions.

[0973] Mobile information devices refer to smartphones, tablets, and other similar devices. These devices work in conjunction with a server to monitor children's emotions and behavior in real time. The devices receive analysis results sent from the server, and if signs of bullying are detected, they immediately send notifications to parents and educators. Notifications are sent in the form of push notifications or email.

[0974] As a concrete example, suppose a smartphone is used to film children playing in a school classroom. This video data is sent to a server, which uses image recognition technology to analyze the children's facial expressions and actions. For example, if the server detects that a child is frightened or being subjected to aggressive behavior from other children, it immediately sends a notification to parents or educators.

[0975] An example of a prompt for a generative AI model is, "Based on this video, analyze whether the child may be being bullied." Using this prompt, the AI ​​model can analyze the video data and assess the likelihood of bullying.

[0976] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0977] Step 1:

[0978] The device uses its camera to capture images of children in real time. The captured video data is temporarily stored on the device. This video data becomes the input for the next processing step.

[0979] Step 2:

[0980] The device sends the captured video data to the server. The server takes the received video data as input and begins analysis using image recognition technology. Specifically, it uses libraries such as TensorFlow and OpenCV to analyze the facial expressions and movements of children in the video and extract features. The results of this analysis become the input for the next processing step.

[0981] Step 3:

[0982] The server uses a generative AI model to evaluate the possibility of bullying based on the extracted features. The prompt message, "Based on this video, analyze whether the child may be being bullied," is input to the AI ​​model, and signs of bullying are detected. This evaluation result becomes the input for the next processing step.

[0983] Step 4:

[0984] If signs of bullying are detected, the server sends the assessment results to the device. Based on the received assessment results, the device sends a notification to parents or educators. The notification is sent via push notification or email and informs them of the possibility of bullying. This notification is the final output.

[0985] (Example 2)

[0986] Next, we will describe Example 2 of the morphological example. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0987] While early detection and rapid response to bullying are crucial, traditional methods often miss early signs of bullying, making it difficult to implement appropriate measures. In particular, there is a lack of mechanisms to monitor changes in children's facial expressions and behavior in real time and promptly notify relevant parties.

[0988] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0989] In this invention, the server includes means for collecting video and audio data, means for analyzing the collected data and analyzing the child's facial expressions and behavior, means for evaluating the possibility of bullying based on the analysis results, means for generating warning messages using a generative AI model, means for sending the generated warning messages to parents and educators, and means for proposing specific countermeasures to parents and educators. This makes it possible to detect signs of bullying early and notify relevant parties quickly.

[0990] "Video data" refers to visual information acquired using cameras and other recording devices, and is fundamental information for analyzing children's facial expressions and behavior.

[0991] "Audio data" refers to auditory information acquired using sound devices such as microphones, and is fundamental information for analyzing conversation content and tone.

[0992] "Analysis" refers to the process of processing collected video and audio data to analyze children's facial expressions and behavior.

[0993] A "generative AI model" refers to a program that uses artificial intelligence technology to perform natural language processing and generate warning messages based on the analysis results.

[0994] A "warning message" is a notification generated when potential bullying is detected, and it contains information to inform parents and educators about the situation.

[0995] "Parents and educators" refers to adults involved in the upbringing and education of children, and who have the responsibility to take appropriate action against signs of bullying.

[0996] "Specific measures" refer to action guidelines proposed when the possibility of bullying is detected, outlining the countermeasures that parents and educators should take.

[0997] A description of embodiments for carrying out this invention will be given.

[0998] The user generates a program for a bullying detection system. This program uses cameras and sensors installed within the school to monitor children's facial expressions and behavior in real time. The server collects video and audio data from these devices. Specifically, it processes video data acquired using cameras with image recognition software (e.g., OpenCV) to analyze children's facial expressions. It also processes audio data acquired using microphones with speech analysis software (e.g., Google Cloud Speech-to-Text) to analyze conversation content and tone.

[0999] The server uses a generative AI model (e.g., OpenAI's GPT-4) based on the analysis results to assess the likelihood of bullying. The generative AI model converts the analysis data into natural language and calculates a score indicating the likelihood of bullying. If signs of bullying are detected, the server generates a warning message. This warning message includes the date, time, and location where the bullying was detected, as well as details of the child's facial expressions and behavior.

[1000] The device sends generated warning messages to parents and educators via email or app notifications. This allows stakeholders to quickly understand the situation and take appropriate action. For example, if a specific child exhibits aggressive behavior towards another child in the schoolyard at 3 PM on a given day, a warning detailing the incident will be sent. An example of a prompt message might be: "Potential bullying detected in the schoolyard. Date and time: 3 PM, October 5, 2023. Location: Schoolyard. The child appeared angry and acted aggressively. What actions would you suggest to the parents?"

[1001] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1002] Step 1:

[1003] The server collects video and audio data from cameras and microphones installed within the school. It receives real-time visual and auditory information as input. This data is used as foundational information for monitoring children's facial expressions and behavior. Specifically, the server receives data streams from each device and stores them in a database.

[1004] Step 2:

[1005] The server processes the collected video data using image recognition software (e.g., OpenCV) to analyze the children's facial expressions. The video data collected in step 1 is used as input. For data processing, a face detection algorithm is applied to extract facial features. The output is the analysis results for each child's facial expression. Specifically, the server detects faces in each video frame and identifies expressions such as smiles and anger.

[1006] Step 3:

[1007] The server processes the collected audio data using speech analysis software (e.g., Google Cloud Speech-to-Text) to analyze the conversation content and tone. The audio data collected in step 1 is used as input. Data processing involves converting the audio to text and performing sentiment analysis. The output is the analysis results regarding the conversation content and tone. Specifically, the server converts the audio data to text and detects aggressive words and tones.

[1008] Step 4:

[1009] Based on the analysis results from steps 2 and 3, the server uses a generative AI model (e.g., OpenAI's GPT-4) to assess the likelihood of bullying. The inputs used are the analysis results of facial expressions and voice. The data calculation involves the generative AI model calculating a score indicating the likelihood of bullying. The output is an evaluation result regarding the likelihood of bullying. Specifically, the server prompts the generative AI model with the analysis results and calculates the score.

[1010] Step 5:

[1011] The server generates warning messages using a generative AI model. It uses the bullying possibility assessment results obtained in step 4 as input. For data processing, it creates warning messages in natural language. The output is a warning message to be sent to parents and educators. Specifically, the server generates a message that includes details such as the date, time, location, facial expressions, and actions.

[1012] Step 6:

[1013] The device sends the warning message received from the server to parents and educators via email or app notification. It uses the warning message generated in step 5 as input. Its output is a notification to the relevant parties. Specifically, the device sends the message to the designated contacts to prompt a quick response.

[1014] (Application Example 2)

[1015] Next, we will describe application example 2 of form example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1016] In recent years, bullying among children in schools and homes has become a serious social problem. However, it is difficult to detect signs of bullying early and take appropriate measures. In particular, detecting the possibility of bullying from a child's facial expressions and behavior is a great burden for educators and parents. To solve this problem, a system is needed that can monitor children's behavior in real time and quickly detect signs of bullying.

[1017] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[1018] In this invention, the server includes means for analyzing a child's facial expressions and behavior using image recognition technology to detect the possibility of bullying, means for monitoring a child's facial expressions and behavior in real time using a visual device, and means for analyzing signs of bullying using a generative AI model. This makes it possible to detect signs of bullying early and take appropriate measures quickly.

[1019] "Image recognition technology" is a technique that analyzes image data acquired using cameras and sensors to identify specific patterns and features.

[1020] "Children" refers to children of school age or those receiving education at home.

[1021] "Facial expression" refers to the emotions and reactions expressed through the movement of the facial muscles.

[1022] "Behavior" refers to actions and reactions that children take in their daily lives or in specific situations.

[1023] "Possible bullying" refers to signs of aggressive or inappropriate behavior towards others that can be inferred from a child's facial expressions and actions.

[1024] A "warning" is a notification or message issued to alert those involved when a potential bullying situation is detected.

[1025] "Guardian" refers to a child's parent or legal supervisor, who is responsible for the child's welfare and education.

[1026] "Education-related personnel" refers to individuals who hold positions related to the education and guidance of children in schools or educational institutions.

[1027] "Countermeasures" refers to specific actions or measures that should be taken when there is a possibility of bullying.

[1028] "Visual devices" refer to devices used to acquire and display images and videos.

[1029] A "generative AI model" refers to an algorithm or system that uses artificial intelligence technology to analyze data and generate specific patterns or results.

[1030] "Analysis" is the process of examining data in detail and deriving specific information or conclusions.

[1031] A "notification" is a message or alert used to convey specific information to relevant parties.

[1032] The system for carrying out this invention mainly consists of a server, a visual device, and a generative AI model. The server uses image recognition technology to analyze data on children's facial expressions and behavior acquired from the visual device and detect the possibility of bullying. The visual device is a device such as smart glasses and is responsible for monitoring children's behavior in real time. The generative AI model is built using frameworks such as TensorFlow or PyTorch and is used to analyze patterns in children's facial expressions and behavior.

[1033] The server receives video data transmitted from the visual device and analyzes the data using a generative AI model. If the analysis detects signs of bullying, the server generates a warning and sends a notification to parents and educators. This notification includes the date, time, and location of the detection, details of the child's facial expressions and behavior, and further specific countermeasures are suggested.

[1034] As a concrete example, imagine a school recess where a visual device monitors a group of students. If the generative AI model detects that a particular student frequently displays a sad expression, the server sends a notification to educators stating, "This may be bullying. Please investigate further." This notification allows educators to take prompt action.

[1035] An example of a prompt message for a generative AI model is: "Analyze the child's facial expression data and detect signs of bullying. If detected, output the date, time, location, and facial expression details, and suggest appropriate countermeasures."

[1036] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1037] Step 1:

[1038] The server receives video data transmitted in real time from the visual device. This video data includes the children's facial expressions and actions. The server preprocesses the video data and converts it into a format to which image recognition technology can be applied.

[1039] Step 2:

[1040] The server inputs pre-processed video data into a generative AI model. The generative AI model, built using TensorFlow, analyzes the children's facial expressions and behavior. The model extracts facial features and detects patterns that indicate the possibility of bullying. The analysis results in an output indicating whether or not bullying is likely.

[1041] Step 3:

[1042] Based on the output from the generating AI model, the server generates an alert if potential bullying is detected. This alert includes the date and time of detection, location, and details of the child's facial expressions and behavior. The server organizes this information and creates a notification message.

[1043] Step 4:

[1044] The server sends the created notification message to parents and educators. The notification is sent via email or app push notification. Recipients can check the notification and take appropriate action based on the child's situation.

[1045] Step 5:

[1046] After sending a notification, the server updates the training database for the generated AI model. Newly detected cases are added to the training data to improve the model's accuracy. This process continuously improves the accuracy of bullying detection.

[1047] (Example 3)

[1048] Next, we will describe Embodiment 3 of Embodiment Example 3. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1049] There is a need to accurately analyze children's emotional states and behavioral patterns to detect potential abnormal behavior at an early stage. However, conventional systems suffer from insufficient analytical accuracy, resulting in a lack of information for parents and educators to take appropriate action. Furthermore, the absence of self-learning capabilities based on analysis results makes it difficult to improve detection accuracy.

[1050] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 3 is realized by the following means.

[1051] In this invention, the server includes means for analyzing a child's emotional state and behavioral patterns using an image processing device, means for detecting the possibility of abnormal behavior based on the analyzed data, and means for providing the analysis results through an information display device. This enables more accurate analysis of a child's behavior, early detection of abnormal behavior, and rapid provision of information to parents and educators.

[1052] An "image processing device" is a device used to analyze image data and extract specific features.

[1053] "Children" refers to children attending educational institutions or minors.

[1054] "Emotional state" refers to the psychological state that can be inferred from a child's facial expressions and behavior.

[1055] "Behavioral patterns" refer to the tendencies of a series of actions and behaviors that a child exhibits.

[1056] "Abnormal behavior" refers to actions that deviate from normal behavioral patterns and includes signs of bullying and problematic behavior.

[1057] An "information display device" is a device used to visually display analysis results, and includes monitors, smartphones, and other similar devices.

[1058] The "self-learning function" is a feature that allows the system to automatically learn from newly collected data and improve the accuracy of its analysis.

[1059] A description of embodiments for carrying out this invention will be given.

[1060] The user builds a system to analyze children's emotional states and behavioral patterns. This system has the capability to analyze children's facial expressions and behavior in real time using an image processing device. Specifically, a camera installed on the terminal captures video of the children and sends the data to a server. The server preprocesses the received image data and extracts features using an image recognition AI model. This AI model is built using machine learning frameworks such as TensorFlow and PyTorch.

[1061] The server detects potential abnormal behavior based on extracted characteristic data. The detected results are provided to parents and educators via an information display device. This allows users to gain a concrete understanding of the child's situation and take appropriate action as needed.

[1062] For example, when a user monitors children's activities through cameras installed in school classrooms, the server receives image data in real time, and an AI model analyzes that data. The analysis results are provided to parents and educators through a dashboard.

[1063] An example of a prompt to be input to the generating AI model is, "Analyze the facial expressions and behavior of children in the classroom and create a report to detect signs of abnormal behavior." In response to this prompt, the AI ​​performs the specified task and provides the results. The specific processing flow in Example 3 is explained using Figure 15.

[1064] Step 1:

[1065] The device uses a camera to capture images of the children. The input is real-time video data. The device sends this video data to a server. Specifically, the camera continuously captures the facial expressions and movements of the children in the classroom.

[1066] Step 2:

[1067] The server preprocesses the video data received from the terminal. The input is the video data transmitted from the terminal. The server uses image processing libraries to adjust the image resolution and remove noise. The output is image data that has been prepared for analysis. Specifically, the server uses OpenCV to clear the image.

[1068] Step 3:

[1069] The server extracts features from pre-processed image data. The input is pre-processed image data. The server uses an image recognition AI model to extract the facial expressions and behavioral features of children as numerical data. The output is feature data. Specifically, the AI ​​model performs facial landmark detection and motion analysis.

[1070] Step 4:

[1071] The server detects potential abnormal behavior based on extracted feature data. The input is feature data. The server uses an AI model to execute an algorithm to detect signs of bullying and abnormal behavior. The output is an analysis result indicating potential abnormal behavior. Specifically, the AI ​​model identifies patterns of abnormal behavior.

[1072] Step 5:

[1073] The server provides analysis results to parents and educators via an information display device. The input is the analysis results indicating potential abnormal behavior. The server displays the results on a dashboard and mobile app. The output is the analysis results in a format that users can review. Specifically, the server displays the results via a web interface.

[1074] Step 6:

[1075] The server uses newly collected data to train an AI model. The input consists of analysis results and newly collected data. The server executes the AI ​​model's learning algorithm to improve its accuracy. The output is the improved AI model. Specifically, the server updates the model's parameters.

[1076] (Application Example 3)

[1077] Next, we will describe application example 3 of form example 3. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1078] Bullying and isolation among minors are serious challenges in educational settings. Traditional methods make it difficult to detect these problems early and respond appropriately. In particular, there is a lack of means to monitor changes in minors' facial expressions and behavior in real time and respond quickly. Therefore, a new system is needed to ensure the safety and healthy development of minors.

[1079] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 3 is realized by the following means.

[1080] This invention includes a server that uses an image recognition algorithm to analyze the facial expressions and behavior of minors and detect the possibility of bullying; a server that issues a warning based on the detected possibility of bullying; a server that proposes countermeasures to parents and educators who have received a warning; a server that analyzes the facial expressions and behavior of minors in real time and sends a notification when an abnormality is detected; and a server that can immediately check the situation of minors via a smart device. This makes it possible to detect problems involving minors early and respond quickly and appropriately.

[1081] An "image recognition algorithm" is a method that uses computer vision technology to identify specific patterns or features from image data.

[1082] "Minor" refers to a person who has not reached the legal age of majority, and usually includes children and adolescents under the age of 18.

[1083] "Facial expression" is a visual representation of emotions and reactions shown through the movement and arrangement of facial muscles.

[1084] "Behavior" is a general term for actions and reactions that an individual exhibits in specific circumstances.

[1085] "Potential bullying" refers to a situation in which a particular behavior or situation has the potential to lead to aggressive or harmful behavior towards others.

[1086] A "warning" is a notification or alert that informs you of a potential risk or problem.

[1087] "Guardian" refers to a parent or legal guardian who is responsible for the life and education of a minor.

[1088] "Educational personnel" refers to teachers and staff involved in the education and guidance of minors in educational institutions.

[1089] "Measures" refer to specific actions or plans taken to address a particular problem or issue.

[1090] "Real-time" refers to a state where data processing and information provision occur immediately, with virtually no delay.

[1091] A "notification" is a means of conveying specific information or a message to a recipient.

[1092] A "smart device" refers to a portable or stationary electronic device that is capable of connecting to the internet and running applications.

[1093] The system for implementing this invention primarily uses a server, a smart device, and an image recognition algorithm. The server executes the image recognition algorithm to analyze the facial expressions and behavior of minors in real time. Specifically, it processes video data acquired from a camera using machine learning frameworks such as TensorFlow and PyTorch. This allows for the detection of changes in the facial expressions and behavior of minors and the assessment of the possibility of bullying.

[1094] The server issues a warning based on the detected potential bullying. The warning is notified to parents and educators via smart devices. Smart devices such as smartphones and smart glasses receive notifications in real time, allowing them to immediately check on the minor's situation.

[1095] As a concrete example, during school breaks, cameras monitor the behavior of minors. A server analyzes the video data and detects potential bullying if a particular minor is isolated or having trouble with other minors. If bullying is detected, the server immediately sends a notification to parents or educators and suggests appropriate measures.

[1096] By utilizing generative AI models, the server can self-learn based on new data, improving the accuracy of bullying detection. An example of a prompt would be, "I want to develop an application that analyzes children's facial expressions and behavior and sends notifications if abnormalities are detected. What algorithms and technologies should I use?"

[1097] The flow of the specific processing in Application Example 3 will be explained using Figure 16.

[1098] Step 1:

[1099] The server acquires video data from the camera. The input is real-time video footage of minors. The server preprocesses this video data and converts it into a format that is easy for image recognition algorithms to process. Specifically, it adjusts the video resolution and removes noise.

[1100] Step 2:

[1101] The server inputs pre-processed video data into an image recognition algorithm. The algorithm analyzes the facial expressions and behaviors of minors and extracts features. The output is feature data related to facial expressions and behaviors. Specifically, it performs facial landmark detection and behavioral pattern recognition.

[1102] Step 3:

[1103] The server evaluates the likelihood of bullying based on extracted feature data. The input is feature data, and the output is a score indicating the likelihood of bullying. The server uses a generative AI model to compare with past data and detect anomalies. Specifically, it applies an anomaly detection algorithm to evaluate deviations from normal behavioral patterns.

[1104] Step 4:

[1105] The server generates a warning if it determines there is a high probability of bullying. The input is a score indicating the likelihood of bullying, and the output is a warning message. The server sends this warning to smart devices. Specifically, it generates push notifications and sends alerts to parents and educators.

[1106] Step 5:

[1107] The device (smart device) receives a warning from the server. The input is the warning message, and the output is a notification displayed to the user. The device displays the notification on the screen so that the user can immediately check the situation. Specifically, it plays a notification sound and displays a pop-up on the screen.

[1108] Step 6:

[1109] The user checks the notification and takes action as necessary. The input is the warning message displayed on the device, and the output is the user's response. Based on the notification, the user checks the minor's situation and takes appropriate action. Specific actions might include educators going to the scene or parents contacting the school.

[1110] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1111] "Example of form 1"

[1112] This invention is a system that combines an emotion engine. Specifically, in addition to a device that captures a child's facial expressions and behavior, it is equipped with an emotion engine that recognizes the user's emotions. The emotion engine analyzes emotions from, for example, the user's tone of voice, word choice, and facial expressions. This emotion engine can determine that if a child is showing negative emotions such as anger or sadness, it is more likely to be a contributing factor to bullying, thereby improving the accuracy of detecting the possibility of bullying.

[1113] "Example of form 2"

[1114] This invention is a system that improves the accuracy of detecting the possibility of bullying based on the user's emotions detected by an emotion engine. Specifically, the emotion engine feeds back the detected emotion data to an image recognition AI, which then learns from that data. For example, if the emotion engine detects that a child is showing emotions such as anger or sadness, the AI ​​analyzes the child's facial expressions and behavior based on that information, enabling it to detect the possibility of bullying with greater accuracy.

[1115] "Example of form 3"

[1116] This invention is a system that adjusts the content of warnings based on the user's emotions detected by an emotion engine. Specifically, it adjusts the content of warnings and suggested countermeasures based on the emotion data detected by the emotion engine. For example, if the emotion engine detects that a child is showing strong anger, it can use that information to strengthen the content of the warning and suggest more specific countermeasures to parents and educators.

[1117] The following describes the processing flow for each example of the form.

[1118] "Example of form 1"

[1119] Step 1: A device that captures the child's facial expressions and actions records the child's actions.

[1120] Step 2: The captured data is sent to the emotion engine, which analyzes the user's emotions based on their voice tone, word choice, facial expressions, etc.

[1121] Step 3: The emotional engine determines that if a child is exhibiting negative emotions such as anger or sadness, it is more likely to be a contributing factor to bullying.

[1122] "Example of form 2"

[1123] Step 1: The emotion engine analyzes the child's emotions and feeds the results back to the image recognition AI.

[1124] Step 2: The image recognition AI learns based on feedback from the emotion engine.

[1125] Step 3: Based on the learning results, the AI ​​analyzes the child's facial expressions and behavior to detect the possibility of bullying with greater accuracy.

[1126] "Example of form 3"

[1127] Step 1: The emotion engine analyzes the child's emotions and adjusts the content of the warnings and suggested countermeasures based on the results.

[1128] Step 2: For example, if the emotion engine detects that the child is showing strong anger, it uses that information to strengthen the content of the warning.

[1129] Step 3: Based on that information, propose more specific measures to parents and educators.

[1130] (Example 1)

[1131] Next, we will describe Embodiment 1 of Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1132] In modern society, bullying is a serious social problem, particularly affecting the physical and mental health of children. However, it is difficult to detect the signs of bullying early and to respond appropriately. Traditional methods can lead to delays in detecting bullying and escalating the damage, so there is a need to quickly and accurately detect the possibility of bullying and take appropriate measures.

[1133] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1134] In this invention, the server includes means for acquiring a child's facial expressions and behavior using a camera; means for analyzing the acquired data using image recognition technology to detect the possibility of bullying; and means for analyzing audio data to identify emotions. This makes it possible to quickly and accurately detect the possibility of bullying and take appropriate measures.

[1135] A "recording device" is a device used to capture a child's facial expressions and behavior, and includes devices such as cameras and smartphones.

[1136] "Image recognition technology" is a technique that analyzes acquired image data to identify specific patterns and features, and is used to analyze children's facial expressions and behavior.

[1137] "Audio data" refers to data that records children's voices and surrounding sounds, and is used for analysis to identify emotions.

[1138] "Emotion analysis technology" is a technology that analyzes voice data to identify the speaker's emotions, and it is a technology that determines emotions by analyzing the tone of voice and word choice.

[1139] "Potential bullying" is an indicator that shows the possibility that a child is being bullied or is involved in bullying, and it is evaluated by integrating the results of image recognition technology and emotion analysis technology.

[1140] A "warning" is a notification issued to parents and educators when the possibility of bullying is detected, and it contains information to encourage appropriate measures.

[1141] The "means of proposing countermeasures" refer to a function that provides specific countermeasures to parents and educators based on the detected possibility of bullying.

[1142] This invention is a system for quickly and accurately detecting the possibility of bullying and taking appropriate countermeasures. The system consists of a camera, image recognition technology, and voice data analysis technology.

[1143] Users utilize cameras and smartphones as recording devices for use at school or home. These devices capture children's facial expressions and behavior in real time. The devices then transmit the acquired image data to image recognition technology. This technology analyzes children's facial expressions and behavior using libraries such as TensorFlow and OpenCV. Specifically, it extracts facial feature points and determines whether the expression indicates fear or sadness.

[1144] The server receives audio data transmitted from the terminal and analyzes it using emotion analysis technology. This technology analyzes the tone and word choice of the voice to identify the emotions the user is expressing. For example, it analyzes the pitch and speed of the voice to detect negative emotions such as anger or sadness.

[1145] The server integrates data obtained from image recognition and sentiment analysis technologies to assess the likelihood of bullying. If the assessment determines that there is a high probability of bullying, the server issues a warning to the user. After receiving the warning, the user can take appropriate measures at school or home as needed.

[1146] As a concrete example, suppose a user is using a smartphone to film their child. If image recognition technology analyzes the child's facial expression and determines that the child is depressed, the server will use that information to notify the user that bullying may be occurring. Similarly, if emotion analysis technology analyzes the tone of the child's voice and determines that the child is angry, the system will also determine that bullying may be occurring.

[1147] An example of a prompt to input into a generative AI model is, "Please tell me how to analyze a child's facial expressions and tone of voice to detect the possibility of bullying." This prompt allows the AI ​​model to provide information about specific analysis methods and the techniques to be used.

[1148] The flow of the specific processing in Example 1 will be explained using Figure 17.

[1149] Step 1:

[1150] The device uses a camera or smartphone to capture the child's facial expressions and actions in real time. The input is the captured image data. The device prepares this image data to be transmitted to image recognition technology. Specifically, the device converts the image data to an appropriate format and performs preprocessing for analysis.

[1151] Step 2:

[1152] The device analyzes captured image data using image recognition technology. The input is the image data acquired in step 1. The image recognition technology uses libraries such as TensorFlow and OpenCV to extract facial feature points and determine whether the expression is frightened or depressed. The output is the result of the facial expression analysis. Specifically, the device applies a face detection algorithm and quantifies the facial features.

[1153] Step 3:

[1154] The server receives audio data transmitted from the terminal and analyzes it using sentiment analysis technology. The input is audio data. Sentiment analysis technology analyzes the tone and word choice of the voice to identify the emotions expressed by the user. The output is the result of the sentiment analysis. Specifically, the server performs spectral analysis on the audio data to detect changes in pitch and speed.

[1155] Step 4:

[1156] The server integrates data obtained from image recognition and emotion analysis technologies to assess the likelihood of bullying. The input consists of facial expression analysis results and emotion analysis results. The server integrates this data to calculate an index indicating the likelihood of bullying. The output is the assessment result of the likelihood of bullying. Specifically, the server uses statistical methods to weight each analysis result and perform an overall evaluation.

[1157] Step 5:

[1158] The server issues a warning to the user if it determines that bullying is highly likely. The input is the evaluation result of the likelihood of bullying. Based on the evaluation result, the server generates a warning message and notifies the user. The output is the warning message. Specifically, the server sends the message to the user's terminal via the notification system.

[1159] (Application Example 1)

[1160] Next, we will describe Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1161] When children are at risk of being bullied at school or at home, it is crucial to detect early signs of bullying and take appropriate measures. However, traditional methods often miss signs of bullying, making it difficult to adequately ensure the safety of children.

[1162] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1163] This invention includes a server that uses image recognition technology to analyze a child's facial expressions and behavior to detect the possibility of bullying, means of analyzing a child's tone of voice and language use using emotion analysis technology to evaluate the possibility of bullying, and means of monitoring the child's safety in real time via a smart device. This makes it possible to detect signs of bullying early and quickly notify parents and educators.

[1164] "Image recognition technology" is a technique that analyzes image data acquired using cameras and sensors to identify specific patterns and features.

[1165] "Children" refers to children of school age or those receiving education at home.

[1166] "Facial expression" refers to the emotions and reactions expressed through the movement of the facial muscles.

[1167] "Behavior" refers to actions and behaviors that children perform in their daily lives.

[1168] "Potential bullying" refers to the possibility that a child may suffer mental or physical harm due to aggressive behavior or words from others.

[1169] "Emotion analysis technology" is a technique that analyzes voice and text data to estimate the emotional state of the speaker.

[1170] "Voice tone" refers to the characteristics of a speaker's voice, expressed through factors such as pitch, volume, and intonation.

[1171] "Word choice" refers to the way a speaker selects and expresses themselves using words.

[1172] A "smart device" refers to a portable electronic device that has internet connectivity and can run various applications.

[1173] "Real-time" refers to a state where data acquisition and processing occur instantly, and results are obtained without delay.

[1174] "Guardian" refers to a child's parent or the person legally responsible for caring for that child.

[1175] "Education professionals" refers to those who are responsible for duties related to the education and guidance of children in schools and educational institutions.

[1176] The system for implementing this invention mainly consists of three elements: a server, a terminal, and a user. The server uses image recognition technology and emotion analysis technology to analyze the child's facial expressions, behavior, tone of voice, and speech. Specifically, the server uses TensorFlow to analyze image data acquired from cameras and sensors to identify the child's facial expressions and behavior. It also uses IBM Watson to analyze audio data and estimate the child's emotional state.

[1177] The device functions as a smart device, monitoring children's safety in real time. It receives analysis results from the server and immediately sends notifications to parents and educators if potential bullying is detected. This enables a rapid response.

[1178] As a parent or educator, the user can receive notifications from the device and take appropriate measures to ensure the safety of children. For example, if the device detects a situation during school recess where a child is being verbally abused by another child, the server will determine from the child's facial expression that they are "scared" and analyze the tone of their voice as "angry." Based on these results, the device will determine that there is a "possible case of bullying" and send a notification to the parent or guardian.

[1179] An example of a prompt to input into a generative AI model is, "Analyze the facial expressions and tone of voice of the child in this video and assess the possibility of bullying." This prompt allows the server to perform an appropriate analysis and assess the possibility of bullying.

[1180] The flow of a specific process in Application Example 1 will be explained using Figure 18.

[1181] Step 1:

[1182] The server receives image and audio data transmitted from the terminal. The input consists of facial expressions and voice data of the children acquired through the camera and microphone. To analyze this data, the server first preprocesses the data and removes noise.

[1183] Step 2:

[1184] The server analyzes preprocessed image data using TensorFlow. The input is denoised image data. The server uses image recognition technology to identify the children's facial expressions and behaviors and extract features that indicate the possibility of bullying. The output is the analysis results regarding facial expressions and behaviors.

[1185] Step 3:

[1186] The server analyzes pre-processed audio data using IBM Watson. The input is audio data with noise removed. The server uses emotion analysis technology to analyze the tone of the child's voice and word choice, and estimates their emotional state. The output is the analysis result regarding emotions.

[1187] Step 4:

[1188] The server integrates the analysis results of image and audio data to assess the possibility of bullying. The input consists of analysis results regarding facial expressions, behavior, and emotions. The server combines this data to make a comprehensive judgment on the possibility of bullying. The output is the assessment result regarding the possibility of bullying.

[1189] Step 5:

[1190] The device receives the assessment results regarding the possibility of bullying, which are sent from the server. The input is the assessment results from the server. Based on these results, if the device determines that there is a high possibility of bullying, it sends a notification to parents or educators. The output is the notification message.

[1191] Step 6:

[1192] The user receives notifications from the device and takes appropriate measures to ensure the child's safety. The input is the notification message from the device. Based on this information, the user checks the child's situation and intervenes as needed. The output is the specific action taken to ensure the child's safety.

[1193] (Example 2)

[1194] Next, we will describe Example 2 of the morphological example. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1195] There are challenges in the early detection of bullying and the proposal of appropriate countermeasures. In particular, there is a need to accurately capture changes in children's emotions and behavior and detect the possibility of bullying with high precision. It is also important to quickly provide concrete countermeasures based on the detection results.

[1196] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1197] In this invention, the server includes means for collecting emotional data, means for feeding the collected emotional data back to image recognition technology, and means for analyzing a child's facial expressions and behavior using image recognition technology to detect the possibility of bullying. This makes it possible to detect the possibility of bullying with high accuracy and to provide appropriate countermeasures quickly.

[1198] "Emotional data" refers to information about a child's emotions obtained from their facial expressions and behavior.

[1199] "Image recognition technology" is a technique that analyzes image data to identify specific patterns or features.

[1200] "Possibility of bullying" refers to the possibility of bullying occurring, inferred from the analysis of a child's facial expressions and behavior.

[1201] A "warning" is a cautionary message sent to parents and educators when the possibility of bullying is detected.

[1202] "Generative AI technology" is a technology that uses artificial intelligence to generate new information and proposals.

[1203] A "learning function" is a feature that allows a system to improve itself based on past data.

[1204] This invention is a system that detects the possibility of bullying with high accuracy and provides appropriate countermeasures. Specific embodiments are shown below.

[1205] The user uses their device to record the child's facial expressions and behavior using an emotion engine. The emotion engine uses sensors such as cameras and microphones to detect the child's emotions, such as anger or sadness, in real time. For example, if a child looks sad at school, video and audio data of that moment are collected.

[1206] The device sends the collected emotional data to a server. The server receives this data and feeds it back into image recognition technology. The image recognition technology learns from the received emotional data and analyzes the child's facial expressions and behavior in detail. This analysis is used to assess the possibility of bullying.

[1207] If potential bullying is detected, the server will issue a warning to parents or educators. This warning will be sent as an email or app notification and will include details such as the date, time, and location of the detection, as well as the child's facial expressions and behavior.

[1208] Furthermore, the server uses generative AI technology to suggest specific countermeasures. For example, if you input a prompt message such as "What should you do if your child looks sad at school?", the AI ​​will suggest ways for parents to encourage communication with their children and suggest ways for educators to strengthen school monitoring systems.

[1209] In this way, the system can quickly and effectively detect potential bullying and provide appropriate countermeasures.

[1210] The flow of the specific processing in Example 2 will be explained using Figure 19.

[1211] Step 1:

[1212] The user uses a device to record the child's facial expressions and behavior using an emotion engine. Video and audio data from the camera and microphone are used as input. The emotion engine analyzes this data to detect the child's emotional state (e.g., anger, sadness). The detected emotion data is generated as output.

[1213] Step 2:

[1214] The device sends the collected emotion data to the server. The emotion data generated in step 1 is used as input. The server receives this data and feeds it back to the image recognition technology. The feedback emotion data is passed back to the image recognition technology as output.

[1215] Step 3:

[1216] The server uses image recognition technology to learn from the feedbacked emotional data. The feedbacked emotional data is used as input. The image recognition technology analyzes the data, examining the child's facial expressions and behavior in detail. The output generates an assessment of the likelihood of bullying.

[1217] Step 4:

[1218] If the server detects a potential bullying incident, it will issue a warning to parents or educators. The evaluation results generated in step 3 are used as input. The server will send a warning via email or app notification, including the date, time, and location of the detection, as well as details of the child's facial expressions and behavior. An alert message is generated and sent as output.

[1219] Step 5:

[1220] The server uses AI generation technology to propose specific countermeasures. The warning message generated in step 4 is used as input. Based on the prompt "What should be done if a child looks sad at school?", the AI ​​generation technology proposes specific countermeasures for parents and educators. The proposed countermeasures are generated and provided as output.

[1221] (Application Example 2)

[1222] Next, we will describe application example 2 of form example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1223] While early detection and appropriate response are crucial in addressing bullying among children, conventional methods struggle with detection and risk overlooking early signs of bullying. Furthermore, there is a need for a system that can accurately detect potential bullying and promptly notify relevant parties.

[1224] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[1225] In this invention, the server includes means for analyzing a child's facial expressions and behavior using an image recognition algorithm to detect the possibility of bullying, means for analyzing video in real time and extracting emotional data, and means for learning based on the extracted emotional data and evaluating the possibility of bullying. This makes it possible to detect the possibility of bullying with high accuracy and to quickly notify parents and educators.

[1226] An "image recognition algorithm" is a computational method used to analyze video data acquired from cameras and sensors and identify specific patterns or features.

[1227] "Children's facial expressions and behavior" refers to the visual characteristics of children, such as their facial expressions, body movements, and posture.

[1228] "Potential bullying" refers to the probability or risk of a child exhibiting signs of aggressive or inappropriate behavior towards other children.

[1229] "Emotional data" refers to information that indicates emotional states such as anger, sadness, and joy, extracted from a child's facial expressions and behavior.

[1230] "Learning" is the process by which an algorithm improves its performance based on past data and newly acquired data.

[1231] "Evaluating" means judging the possibility of bullying or the emotional state based on specific criteria and deriving a result.

[1232] "Generating and sending notifications" is the process of creating and sending messages to relevant parties to provide warnings or information based on detected data.

[1233] The system for implementing this invention consists of a server, a terminal, and a user. The server acquires video data from surveillance cameras installed within the school and analyzes the children's facial expressions and behavior in real time using an image recognition algorithm. Specifically, the server uses software such as OpenCV and TensorFlow to extract facial features from the video data and generates emotion data using an emotion engine.

[1234] The generated emotion data is used in a learning process within the server to update the model for evaluating the possibility of bullying. Based on this evaluation, the server generates a warning and sends a notification to the devices of parents and educators. The notification includes the date, time, and location of the detection, details of the child's facial expressions and behavior, and specific countermeasures are suggested.

[1235] The terminal is a device such as a smartphone or tablet that receives notifications and displays warnings to the user. The user can check the notification and take the suggested action.

[1236] For example, if a camera captures a scene in a school hallway where a student is showing anger towards another student, the server will detect that anger using an emotion engine and analyze the scene with an image recognition algorithm. If it is determined that there is a high probability of bullying, the server will generate a notification stating, "Possible bullying has been detected in the hallway. Please check the details," and send it to the relevant parties.

[1237] An example of a prompt message is: "Analyze the facial expressions of children in the hallway. If feelings of anger or sadness are detected, use that information to assess the possibility of bullying and generate a notification."

[1238] The flow of a specific process in Application Example 2 will be explained using Figure 20.

[1239] Step 1:

[1240] The server acquires video data in real time from surveillance cameras installed within the school. The input is a video stream from the cameras, and the output is analyzable image data. To process this data, the server divides the video into frames.

[1241] Step 2:

[1242] The server uses an image recognition algorithm to detect children's faces from the acquired frames. The input is the image data obtained in step 1, and the output is the face's position information and features. The server uses OpenCV to extract facial features and determine the face's position.

[1243] Step 3:

[1244] The server uses an emotion engine to generate emotion data from detected facial features. The input is the facial features obtained in step 2, and the output is data indicating the type and intensity of emotion. The server feeds this data back into the generating AI model to perform emotion analysis.

[1245] Step 4:

[1246] The server inputs the generated emotion data into an image recognition algorithm to evaluate the possibility of bullying. The input is emotion data, and the output is a score indicating the possibility of bullying. The server uses TensorFlow to analyze the emotion data and quantify the possibility of bullying.

[1247] Step 5:

[1248] The server generates and sends notifications to parents and educators based on the bullying possibility score. The input is the score obtained in step 3, and the output is a warning message. The server uses a generative AI model to create notifications that suggest specific actions based on the prompt text.

[1249] Example prompt: "Analyze students' facial expressions in the hallway. If anger or sadness is detected, use this information to assess the possibility of bullying and generate a notification."

[1250] (Example 3)

[1251] Next, we will describe Embodiment 3 of Embodiment Example 3. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1252] Conventional technologies have struggled to accurately analyze an individual's emotional state and provide appropriate warnings and countermeasures. Furthermore, they lacked the learning capabilities necessary to detect potential bullying with high accuracy, hindering the ability of stakeholders to respond quickly and appropriately.

[1253] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 3 is realized by the following means.

[1254] In this invention, the server includes means for analyzing an individual's emotional state using image processing technology, means for adjusting the content of a warning based on the analyzed emotional data, and means for proposing specific countermeasures to the person who received the warning. This makes it possible to accurately grasp an individual's emotional state and quickly provide appropriate warnings and countermeasures.

[1255] "Image processing technology" refers to techniques for analyzing digital images and extracting or recognizing specific information.

[1256] "Emotional state" refers to the psychological or emotional state an individual exhibits at a particular moment in time.

[1257] "Warning content" refers to information that serves as a warning or instruction to relevant parties, generated based on the analyzed data.

[1258] "Specific countermeasures" refer to proposals that instruct relevant parties on the actions and measures they should take depending on the detected situation.

[1259] "Self-learning" is the process by which a system automatically learns from new data and improves its performance and accuracy.

[1260] "Stakeholders" refers to individuals such as parents and educators who are the recipients of the analysis results and warnings.

[1261] A description of embodiments for carrying out this invention will be given.

[1262] The server receives data on an individual's facial expressions and actions transmitted from the device. The device collects data in real time using cameras and sensors and sends it to the server. This data is acquired when the individual is active in a specific environment.

[1263] The server analyzes the received data using image processing technology. This analysis utilizes image recognition AI based on frameworks such as TensorFlow and PyTorch. The image recognition AI infers emotional states from an individual's facial expressions and identifies emotions such as anger, sadness, and joy.

[1264] The analyzed emotion data is passed to the emotion engine on the server. The emotion engine quantifies the individual's emotional state based on the analysis results and adjusts the warning content. For example, if the emotion engine detects that an individual is showing strong anger, the server uses that information to strengthen the warning and propose specific countermeasures to those involved.

[1265] Furthermore, the server performs self-learning based on new data, improving the detection accuracy of the image recognition AI. This allows the server to continuously improve its accuracy and provide more precise analysis results.

[1266] As a concrete example, let's look at an example of a prompt sentence to be input to the generating AI model: "Analyze the facial expressions an individual displays during a specific activity and detect emotions such as anger or sadness. Based on the results, propose appropriate measures for those involved." In response to this prompt sentence, the server utilizes image recognition AI and emotion engine to analyze the individual's emotional state in detail and provide appropriate information. The flow of specific processing in Example 3 will be explained using Figure 21.

[1267] Step 1:

[1268] The device uses cameras and sensors to collect data on an individual's facial expressions and behavior in real time. The collected data is transmitted to a server in image or video format. The input is raw data of the individual's facial expressions and behavior, and the output is the transmission of data to the server.

[1269] Step 2:

[1270] The server analyzes the data received from the terminal using image processing technology. Specifically, an image recognition AI using frameworks such as TensorFlow or PyTorch analyzes the received image data and infers the emotional state from the individual's facial expressions. The input is image data sent from the terminal, and the output is the analyzed emotional data.

[1271] Step 3:

[1272] The server passes the analyzed emotional data to the emotion engine, which quantifies the individual's emotional state. Based on the analysis results, the emotion engine identifies emotions such as anger, sadness, and joy, and adjusts the warning content accordingly. The input is the analysis results from the image recognition AI, and the output is the quantified emotional data.

[1273] Step 4:

[1274] The server generates warnings based on quantified emotion data and suggests specific actions for those involved. For example, if strong anger is detected, the server generates a warning such as, "An individual is exhibiting strong anger. Assess the situation and take steps to calm them down." The input is quantified emotion data, and the output is the generated warning and suggested actions.

[1275] Step 5:

[1276] The server performs self-learning based on new data to improve the detection accuracy of the image recognition AI. Specifically, it compares past data with new data and adjusts the model to reduce misrecognition. The input is past data and new data, and the output is the adjusted AI model.

[1277] Step 6:

[1278] The server sends the generated warnings and suggestions to the terminal, notifying the relevant user. The terminal displays the received information on its screen, allowing the user to understand the situation. The input is the generated warnings and suggestions, and the output is the information displayed on the terminal.

[1279] (Application Example 3)

[1280] Next, we will describe application example 3 of form example 3. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1281] There is a need to detect bullying and emotional changes among minors early and to take appropriate action. However, conventional methods may miss signs of bullying, and there are challenges in responding appropriately to emotional changes.

[1282] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 3 is realized by the following means.

[1283] This invention includes a server that uses image recognition technology to analyze the facial expressions and behavior of minors and detect the possibility of bullying; a server that issues a warning based on the detected possibility of bullying; a server that proposes countermeasures to parents and educators who receive the warning; a server that adjusts the content of the warning based on the minor's emotional data; and a server that captures the minor's facial expressions in real time and notifies them of the analysis results. This enables a swift and appropriate response to protect the safety and welfare of minors.

[1284] "Image recognition technology" is a technique that uses computer vision to extract and analyze specific information from image data.

[1285] A "minor" refers to a person who has not reached the legal age of majority.

[1286] "Facial expression" refers to the emotions and reactions shown through the movement of the facial muscles.

[1287] "Action" refers to a series of actions or reactions performed by an individual.

[1288] "Possible bullying" refers to the possibility that a minor is experiencing inappropriate behavior or remarks from others.

[1289] A "warning" is a notification or message intended to draw attention to a specific situation.

[1290] A "guardian" refers to a person who has the responsibility to supervise the life and education of a minor.

[1291] "Education-related personnel" refers to individuals who hold positions in educational institutions that involve the education of minors.

[1292] "Measures" refer to specific actions or means taken to address a particular problem or situation.

[1293] "Emotional data" refers to information that expresses an individual's emotional state using numerical values ​​or categories.

[1294] "Real-time" refers to processing or responding instantly without delay.

[1295] A "notification" is a message or alert used to convey specific information to a recipient.

[1296] The system for implementing this invention mainly consists of a server and a terminal. The server uses image recognition technology to analyze the facial expressions and behavior of minors and detect the possibility of bullying. Specifically, the terminal's camera is used to capture the minor's facial expressions in real time, and the data is sent to the server. The server uses OpenCV as the image recognition technology and the Microsoft Azure Emotion API for emotion analysis. This allows the system to extract emotional data from minors and evaluate the possibility of bullying.

[1297] The server issues warnings based on detected bullying and suggests countermeasures to parents and educators. The content of the warnings is adjusted based on the minor's emotional data. For example, if the minor is showing strong anger, the warning is strengthened and specific countermeasures are suggested.

[1298] The device functions as a smartphone or smart glasses, capturing the facial expressions of minors in real time. The analysis results are sent as notifications to parents and educators. This enables quick and appropriate responses to protect the safety and well-being of minors.

[1299] As a concrete example, if a minor displays strong anger in a classroom, the device captures their facial expression and sends it to a server. The server performs an emotion analysis, and if it determines that the expression is "anger," it sends a notification to the parent saying, "Your child is showing strong anger. Please listen to them."

[1300] An example of a prompt message is: "Analyze the child's facial expression data and determine their emotion. If anger is detected, generate a message to send a notification to the parent / guardian."

[1301] The flow of the specific processing in Application Example 3 will be explained using Figure 22.

[1302] Step 1:

[1303] The device uses a camera to capture the facial expressions of minors in real time. The input is video data from the camera, and the output is image data. The device then prepares to send this image data to a server.

[1304] Step 2:

[1305] The server receives image data transmitted from the terminal. The input is image data from the terminal, and the output is analyzable image information. The server processes this image information using image recognition technology (OpenCV) to analyze the facial expressions of minors.

[1306] Step 3:

[1307] The server inputs facial expression data extracted using image recognition technology into an emotion analysis engine (Microsoft Azure Emotion API). The input is facial expression data, and the output is emotion data. Based on this emotion data, the server determines the emotional state of the minor.

[1308] Step 4:

[1309] The server analyzes sentiment data and assesses the likelihood of bullying. The input is sentiment data, and the output is an assessment result: "bullying is likely" or "bullying is not likely." Based on this assessment, the server generates warnings as needed.

[1310] Step 5:

[1311] The server notifies parents and educators of the generated warnings. The input is the evaluation result, and the output is the warning message. The server adjusts the content of the warnings and suggests specific countermeasures based on the minor's emotional data.

[1312] Step 6:

[1313] The user receives notifications from the server and takes appropriate action based on the minor's situation. The input is the notification from the server, and the output is the user's response. Based on the notification, the user interacts with the minor and, if necessary, collaborates with educators to resolve the problem.

[1314] (Other examples)

[1315] Since this is the same as the specific processing described in the other embodiments of the first embodiment above, the explanation will be omitted.

[1316] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1317] The data generation model 58 is a form of so-called generative AI (Artificial Intelligence). One example of the data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1318] Other examples of generative AI include Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) are examples.

[1319] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[1320] [Fourth Embodiment]

[1321] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[1322] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1323] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1324] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[1325] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1326] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1327] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1328] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[1329] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1330] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1331] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1332] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1333] Next, the identification process performed by the identification processing unit 290 of the data processing device 12 will be described.

[1334] "Example of form 1"

[1335] One embodiment of the present invention involves equipping a camera or smartphone, or other imaging device, with image recognition AI for use in schools or homes. This image recognition AI analyzes a child's facial expressions and behavior in real time to detect the possibility of bullying. For example, it detects the possibility of bullying if it determines that a child's facial expression indicates fear or sadness, or if it determines that the child is being subjected to aggressive behavior from other children.

[1336] "Example of form 2"

[1337] Based on the detected potential bullying, warnings are issued to parents and educators. These warnings are sent via email, app notifications, etc. The warnings include details such as the date, time, and location where the potential bullying was detected, as well as details of the child's facial expressions and behavior. Specific countermeasures are also suggested. For example, parents may be advised to take measures to encourage communication with their children, and educators may be advised to strengthen school monitoring systems and offer counseling.

[1338] "Example of form 3"

[1339] Furthermore, the system provides parents and educators with analysis results of children's facial expressions and behavior. This allows parents and educators to understand the child's situation more concretely and take appropriate action. It also includes a learning function to improve the accuracy of detecting potential bullying. This learning function allows the image recognition AI to self-learn based on new data and improve detection accuracy.

[1340] The following describes the processing flow for each example of the form.

[1341] "Example of form 1"

[1342] Step 1: Equipate image recognition AI into cameras, smartphones, and other imaging devices for use in schools and homes.

[1343] Step 2: Capture the child's facial expressions and actions in real time through the camera and send them to the image recognition AI.

[1344] Step 3: The image recognition AI analyzes the transmitted image data and detects the possibility of bullying from the child's facial expressions and behavior.

[1345] "Example of form 2"

[1346] Step 1: If the image recognition AI detects a potential bullying incident, it issues a warning to parents and educators.

[1347] Step 2: The warning will be sent via email, app notification, etc.

[1348] Step 3: The warning will include details such as the date and location where the bullying was detected, as well as the child's facial expressions and behavior.

[1349] Step 4: Specific measures are proposed. For example, parents may be advised to take steps to encourage communication with their children, and educators may be advised to strengthen school supervision systems and provide counseling.

[1350] "Example of form 3"

[1351] Step 1: Provide parents and educators with the results of the analysis of the child's facial expressions and behavior.

[1352] Step 2: Parents and educators will be able to understand the child's situation more concretely and take appropriate action.

[1353] Step 3: Includes a learning function to improve the accuracy of detecting potential bullying.

[1354] Step 4: This learning function allows the image recognition AI to self-learn based on new data and improve detection accuracy.

[1355] (Example 1)

[1356] Next, we will describe Embodiment 1 of Example Form 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1357] In modern schools and homes, children are increasingly likely to encounter bullying, but there is a challenge in detecting signs of bullying early and taking appropriate measures. In particular, there is a need for technology that can monitor changes in children's facial expressions and behavior in real time and quickly detect the possibility of bullying.

[1358] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1359] In this invention, the server includes means for collecting video data using a camera, means for transmitting the collected video data via communication means, and means for analyzing a child's facial expressions and behavior using image recognition technology to detect the possibility of bullying. This makes it possible to detect signs of bullying early and take appropriate measures.

[1360] A "recording device" is a device used to collect video data, and includes devices such as cameras and smartphones.

[1361] "Video data" refers to digital data containing visual information collected by a camera or camera.

[1362] "Communication methods" refer to the technologies and protocols used to send and receive data, and include methods such as the internet and wireless communication.

[1363] "Image recognition technology" is a technology that allows computers to analyze image data and identify specific patterns or features.

[1364] "Analysis of facial expressions and behavior" is the process of using image recognition technology to analyze a child's facial expressions and body movements to identify specific emotions and behaviors.

[1365] "Detecting the possibility of bullying" is a process of determining the likelihood of bullying occurring based on analyzed facial expression and behavioral data.

[1366] "Means of issuing warnings" refer to methods or devices used to alert those involved when the possibility of bullying is detected.

[1367] "Means of proposing countermeasures" refer to methods and devices for presenting specific countermeasures to parents and educators based on the detected possibility of bullying.

[1368] "Learning function" refers to a function that allows the system to improve the accuracy of its analysis based on past data, and includes machine learning algorithms.

[1369] This invention is a system for use in schools and homes that analyzes children's facial expressions and behavior in real time to detect potential bullying. The system consists of a camera, communication means, and image recognition technology.

[1370] The terminal uses cameras, smartphones, and other recording devices to collect video data of children in real time. The collected video data is transmitted to a server via a secure communication method. The communication method uses the internet or wireless communication technology.

[1371] The server uses image recognition technology to analyze the received video data. This technology leverages generative AI models built using machine learning frameworks such as TensorFlow and PyTorch. Specifically, it uses OpenCV to ...

Claims

1. The processor comprises, Using image recognition technology to analyze the child's facial expressions and behavior, and using emotion analysis technology to analyze the child's tone of voice and word choice, the possibility of bullying is detected. If the possibility of bullying is detected, a prompt message is input to the generating AI model to generate a warning message indicating that signs of bullying have been detected, and a warning message regarding the bullying is generated. Based on the analyzed emotion data, the warning content of the warning message is adjusted, and if the emotion data indicates strong anger, the warning content is strengthened. The warning message includes the date and time and location where the potential bullying was detected, as well as details of the child's facial expression and behavior. Using the aforementioned generation AI model, specific countermeasures against bullying are generated based on the analyzed data and emotional data. The aforementioned warning message and the aforementioned specific countermeasures shall be provided to at least one of the parents and educators. system.

2. The processor provides a dashboard for providing the results of the analysis of the child's facial expressions and behavior to at least one of the parents and educators, and enables the information to be viewed through an interface accessible to at least one of the parents and educators. The system according to claim 1.

3. The processor performs self-learning using newly collected data, periodically evaluates the model to improve the accuracy of detecting potential bullying, and forms a feedback loop for accuracy improvement. The system according to claim 1.

Citation Information

Patent Citations

  • Information processing apparatus, method of finding bullying, information processing system, and computer program

    JP2018112831A

  • Extraction program, extraction method, and extraction device

    JP2020106948A

  • Bullying indication determination program and system, and bullying determination program and system

    JP2021096836A

  • Persona chatbot control method and system

    JP2022180282A

  • Abuse response support device, abuse response support method, program, and recording medium

    JP2024084910A