System

The system uses AI for real-time detection and reporting of child abuse and bullying through emotion recognition, surveillance, and data analysis, facilitating early intervention.

JP2026014911APending Publication Date: 2026-01-29SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024116385
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-19
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Conventional methods struggle to detect child abuse and bullying in real-time, making it difficult to implement effective early countermeasures.

Method used

A system utilizing AI technology for emotion recognition, surveillance cameras, audio devices, 24-hour chat systems, and data collection from social media to analyze facial expressions, voice, and online content to identify abnormalities and alert appropriate authorities.

Benefits of technology

Enables real-time detection and reporting of child abuse and bullying, allowing for prompt intervention and support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026014911000001_ABST
    Figure 2026014911000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for acquiring a user's face or voice data by an emotion recognition device; means for analyzing a user's emotion from the acquired data by using a AI model; and means for reporting to an appropriate institution when an abnormal emotion is detected based on the analyzed emotion information.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In recent years, cases of child abuse and bullying have been increasing, making early detection and countermeasures urgently needed. Conventional methods make it difficult to detect problems early, so a rapid response is required. This invention aims to achieve early detection and countermeasures for abuse and bullying by utilizing AI technology to detect abnormalities in children's emotions and behavior in real time and notify the appropriate authorities. [Means for solving the problem]

[0005] The present invention solves the above problems by providing a system including the following means.

[0006] The emotion recognition device acquires the user's face and voice data,

[0007] The acquired data is used to analyze the user's emotions using an AI model,

[0008] Based on the analyzed emotional information, if abnormal emotions are detected, a report is sent to the appropriate authorities.

[0009] Additionally, we have installed surveillance cameras and audio devices to detect signs of bullying and abuse.

[0010] The captured camera and audio data is analyzed in real time by an AI model,

[0011] The analysis results are used to detect abnormal behavior or sounds and issue an alert.

[0012] In addition, we receive messages from users via chat or phone 24 hours a day,

[0013] Analyze received messages using natural language processing technology,

[0014] Based on the analysis results, appropriate support will be provided.

[0015] Finally, we collect text, image, and video data from social media and online platforms.

[0016] The collected data is analyzed using natural language processing and image recognition technology.

[0017] Detect anomalous posting content and report it to the appropriate authorities.

[0018] Real-time data and analysis results are centrally managed,

[0019] Record details of detected anomalies and notify administrators to take necessary measures.

[0020] An "emotion recognition device" is a device that identifies a person's emotions based on facial expressions and tone of voice.

[0021] An "AI model" is an algorithm that uses artificial intelligence to analyze data and predict specific outcomes.

[0022] A "surveillance camera" is a device for monitoring a specific area and capturing video data in real time.

[0023] An "audio device" is a device for capturing audio and providing that data for analysis.

[0024] "Real-time" refers to data acquisition and processing occurring almost instantaneously.

[0025] "Natural language processing technology" is a technology for analyzing human language and understanding its meaning and emotions.

[0026] A "warning" is a notification issued to alert administrators or related organizations when an abnormal situation is detected.

[0027] "Chat" is a means of communication in which text messages are exchanged in real time.

[0028] "Support content" refers to specific advice or services provided to users who are facing problems or difficulties.

[0029] "SNS" stands for Social Networking Service, a platform for people to interact online.

[0030] "Online platform" refers to all services and applications provided via the Internet.

[0031] "Text data" is data expressed as character information.

[0032] "Image data" is data that contains visual information.

[0033] "Video data" is data that includes moving images.

[0034] "Unusual posts" are posts that contain content that is out of the ordinary or suggests a problem.

[0035] "Appropriate institutions" are organizations and organisations involved in law and welfare that can respond promptly to problems or irregularities.

[0036] "Centralized management" refers to the centralized and unified management of multiple pieces of information and data.

[0037] "Administrator" means a person or organization that operates and manages a system or service. [Brief explanation of the drawings]

[0038] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0039] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0040] First, the terms used in the following description will be explained.

[0041] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0042] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0043] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0044] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0045] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0046] [First embodiment]

[0047] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0048] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0049] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0050] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0051] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0052] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0053] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0054] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0055] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0056] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0057] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0058] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0059] The present invention is a system that uses AI technology to detect child abuse and bullying in real time and report it to the appropriate authorities. The system includes emotion recognition devices, surveillance cameras, audio devices, 24-hour chat and phone systems, and data collection devices from social media and online platforms.

[0060] AI-powered emotion recognition

[0061] The server initializes the camera and microphone devices, and the device captures the user's face and voice data in real time. This data is analyzed using an AI model to detect whether the child is experiencing negative emotions such as fear or stress. If the detected emotional information shows abnormal values, the server sends a report to the appropriate child protection agency.

[0062] AI surveillance system

[0063] The server initializes surveillance cameras and audio devices installed in schools and homes, and the devices capture video and audio data in real time. The captured data is sent to the server and analyzed by an AI model. If abnormal behavior or audio is detected, the server will alert the administrator. This system makes it possible to detect signs of everyday bullying and abuse at an early stage.

[0064] AI-powered helpline

[0065] The server provides a chat system and telephone line that users can use 24 hours a day. When users input their concerns or questions, the device analyzes the message using natural language processing technology and determines the appropriate response. Based on the analysis results, the server provides encouraging messages and contact information for appropriate support centers.

[0066] AI-powered early warning system

[0067] The server collects data from social media and other online platforms, including text, images, and video data. The device then analyzes this data using AI models to detect unusual posts and signs of bullying or abuse. If an anomaly is detected, the server sends a report to the appropriate authorities.

[0068] Specific examples

[0069] For example, suppose a user captures a video of a child's face and voice. The device analyzes the data and detects that the child is experiencing high levels of stress. The server immediately reports this information to a child welfare center. Also, if a school's surveillance camera captures a student being bullied, the device will detect this and the server will alert the school administrator.

[0070] For example, if a user sends a message of advice such as "I don't want to go to school" through the helpline chat system, the device analyzes the content and the server replies with an encouraging message and contact information for a child consultation center. Furthermore, if signs of bullying are detected in a post on social media, the server will use that information to report the matter to the appropriate authorities.

[0071] As a result, the present invention makes it possible to detect child abuse and bullying at an early stage and to take prompt measures against it.

[0072] The processing flow will be explained below.

[0073] AI emotion recognition processing flow

[0074] Step 1:

[0075] The server initializes the camera and microphone devices, which allows face and voice data to be captured in real time.

[0076] Step 2:

[0077] The device captures face and voice data at regular intervals: face data from the camera and voice data from the microphone.

[0078] Step 3:

[0079] The device converts captured facial data into a grayscale image and preprocesses audio data based on the sample rate, making it easier for AI models to analyze.

[0080] Step 4:

[0081] The device then inputs the preprocessed data into an AI model to analyze emotions, which can identify multiple emotions such as fear, stress, and joy.

[0082] Step 5:

[0083] The server receives the analyzed emotional information and reports to the appropriate authorities if abnormal emotions (e.g., high levels of fear or stress) are detected.

[0084] AI monitoring system processing flow

[0085] Step 1:

[0086] The server initializes the surveillance cameras and audio devices installed in schools and homes, enabling continuous monitoring.

[0087] Step 2:

[0088] The device captures video and audio data in real time and transmits it to the server.

[0089] Step 3:

[0090] The server analyzes the video and audio data received in real time using an AI model designed to detect abnormal behavior and audio.

[0091] Step 4:

[0092] If the server detects any unusual activity or sound, it will immediately alert administrators, including details of when and where the anomaly occurred.

[0093] AI-powered helpline process

[0094] Step 1:

[0095] The server provides a 24-hour chat and phone system, which users can use to input their concerns and questions.

[0096] Step 2:

[0097] A user sends a message through the chat or phone system, and the message arrives at the server.

[0098] Step 3:

[0099] The server analyzes the messages it receives using natural language processing technology and classifies their emotions and content.

[0100] Step 4:

[0101] Based on the analysis results, the server provides the user with necessary support and encouraging messages, and in some cases, contact information for specialized helplines.

[0102] Processing flow of an AI-based early warning system

[0103] Step 1:

[0104] The server uses APIs to collect data from social media and other online platforms, retrieving relevant text, image, and video data.

[0105] Step 2:

[0106] The device analyzes the collected data using natural language processing and image recognition techniques, conducting sentiment analysis on text data and detecting specific abnormal behaviors on images and videos.

[0107] Step 3:

[0108] If the server detects any anomalous posting content, it will send a report to the appropriate authorities, which will include details of the posting where the anomaly occurred.

[0109] Step 4:

[0110] The server centrally manages the data and analysis results acquired in real time, records the details of detected anomalies, and notifies the administrator so that necessary measures can be taken.

[0111] The above are the specific processing steps for each system.

[0112] Example 1

[0113] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0114] Child abuse and bullying are serious social issues, and early detection and reporting to appropriate authorities is essential. However, current systems have difficulty analyzing emotions and behaviors in real time and detecting abnormalities, preventing effective countermeasures. Furthermore, even 24-hour helplines have difficulty providing prompt support. An efficient and highly accurate system is needed to solve these problems.

[0115] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0116] In this invention, the server includes: [means for acquiring data on the user's face and voice using an emotion recognition device]; [means for analyzing the user's emotions from the acquired data using a generative AI model]; and [means for reporting to an appropriate institution if an abnormal emotion is detected based on the analyzed emotion information]. This enables real-time analysis of children's emotions and early detection and reporting of abnormalities.

[0117] The server also includes a means for initializing the device, a means for the user to capture face and voice data, and a means for evaluating the analysis results and sending a report. This allows for efficient use of the device and improved accuracy of the analysis.

[0118] Furthermore, the server includes: [means for installing surveillance cameras and audio devices to detect signs of bullying and abuse; [means for analyzing acquired camera and audio data in real time using a generative AI model; and [means for detecting abnormal behavior or audio from the analysis results and issuing an alert.] This enables early detection and rapid response to bullying and abuse using surveillance cameras and audio devices.

[0119] Furthermore, the server includes a means for receiving messages from users via 24-hour chat or telephone, a means for analyzing the received messages using natural language processing technology, and a means for providing appropriate support based on the analysis results. This makes it possible to provide prompt and accurate support in response to consultation messages from users.

[0120] A "server" is a central control device that processes and manages data, communicates with other devices via a network, analyzes various data, and sends reports.

[0121] A "terminal" is a device that is directly operated by a user, and is a device that captures data through a camera or microphone and transmits it to a server.

[0122] A "user" is a person who operates or inputs data using a system, and in particular, who captures data or sends messages.

[0123] An "emotion recognition device" is a device that uses sensors such as cameras and microphones to acquire facial and voice data and analyzes emotions from that data.

[0124] A "generative AI model" is an artificial intelligence model developed using machine learning and deep learning techniques, and includes algorithms for emotion analysis and anomaly detection.

[0125] "Abnormal emotions" are emotions that are outside the normal range, and primarily refer to negative emotions such as fear, stress, and anger.

[0126] "Report" means a report or notification sent to an appropriate institution when an abnormality is detected, and includes the results of the analysis and its specific contents.

[0127] A "surveillance camera" is a photographic device that captures video in real time and transmits the data to a server.

[0128] An "audio device" is a recording device that captures audio, converts it into data, and sends it to a server.

[0129] A "chat system" is a software system for exchanging text-based messages in real time, allowing for 24-hour support.

[0130] "Natural language processing technology" refers to technology for understanding and analyzing human language, and includes algorithms that perform semantic and emotional analysis of text data.

[0131] "Appropriate support content" refers to solutions and support information provided to users in response to their inquiries and problems, including encouraging messages and contact information for support desks.

[0132] The present invention is a system that uses AI technology to detect child abuse and bullying in real time and report it to the appropriate authorities. The system includes emotion recognition devices, surveillance cameras, audio devices, 24-hour chat and phone systems, and data collection devices from social media and online platforms.

[0133] An embodiment of an AI-based emotion recognition system

[0134] First, the server initializes the camera and microphone connected to the emotion recognition device. This initialization includes installing device drivers and verifying the connection. Next, the device uses these devices to allow the user (parent or teacher) to capture the child's face and voice. This captured data is analyzed in real time using a generative AI model.

[0135] As a specific example, if a user captures video of their child's face and voice at school and the analysis results indicate that the child is experiencing severe stress, the server will immediately send a report to a child consultation center.

[0136] AI surveillance system implementation

[0137] The server then initializes the surveillance cameras and audio devices installed in schools and homes. The devices capture video and audio data from these devices in real time. This data is sent to the server and analyzed by the generative AI model. If any abnormal behavior or audio is detected, the server alerts administrators.

[0138] As a concrete example, if a school's surveillance camera captures a scene in which a particular student is being bullied, and if abnormal behavior is detected as a result of analyzing this data, the server will immediately send an alert to the school administrator.

[0139] AI-powered helpline implementation

[0140] In addition, the server provides a chat system and telephone line that users can use 24 hours a day. When users enter their concerns or questions into the chat system, the device analyzes the message using natural language processing technology. Based on the analysis, the server provides encouraging messages and contact information for appropriate support centers.

[0141] As a specific example, if a user types "I don't want to go to school" in a chat, the device analyzes the content and the server replies with an encouraging message and contact information for a child consultation center.

[0142] An example of an early warning system using AI

[0143] The server also collects text, image, and video data from social media and online platforms. The device analyzes this data using generative AI models to detect signs of bullying or abuse. If an anomaly is detected, the server sends a report to the appropriate authorities.

[0144] As a specific example, if a post such as "I'm being bullied by a classmate" is detected on a social networking site, the device will analyze the information and the server will send a report to the appropriate authority.

[0145] Prompt Sentence Examples

[0146] Detect posts on social media that say things like "I'm being bullied at school."

[0147] If your analysis results indicate that "this child is experiencing severe stress," please send a report to the child consultation center.

[0148] Please send an encouraging reply to the consultation message "I don't want to go to school."

[0149] The purpose of this invention is to provide a safe environment by using this system to enable early detection of child abuse and bullying and prompt countermeasures.

[0150] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0151] Processing steps of AI emotion recognition system

[0152] Step 1: Initialize your device

[0153] The server initializes the camera and microphone connected to the emotion recognition device, which includes installing the device drivers and checking the connection.

[0154] Input: Device connection information

[0155] Data processing: Installing device drivers and checking connections

[0156] Output: Device initialization completion notification

[0157] What happens: The server logs the message "Initializing camera and microphone."

[0158] Step 2: Capture the data

[0159] The device uses the initialized camera and microphone to allow the user (parent or teacher) to capture the child's face and voice, and the captured data is temporarily stored on the device.

[0160] Input: Data captured by camera and microphone

[0161] Data processing: collection and temporary storage of face and voice data

[0162] Output: Notification that capture data has been saved

[0163] Specific behavior: The device will display a pop-up saying "Data capturing."

[0164] Step 3: Analyze the data

[0165] The device then inputs the captured facial and voice data into generative AI models for real-time analysis, including convolutional neural networks (CNNs) and recurrent neural networks (RNNs).

[0166] Input: Capture data

[0167] Data processing: Analyzing face and voice data with AI models

[0168] Output: Emotion analysis results

[0169] Specific behavior: The device will display the status "Analyzing data". Once the analysis is complete, the result will be displayed as "Negative emotion detected".

[0170] Step 4: Evaluate the analysis results and submit a report

[0171] The server receives the analysis results sent from the device and, if any abnormal values ​​are detected, sends a report to the appropriate child protection agency. This report includes the analysis data and the results.

[0172] Input: Sentiment analysis results

[0173] Data processing: Evaluation of analysis results and report generation

[0174] Output: Report sent to child protective services

[0175] Specific behavior: The server logs "Report sending" and displays "Report sent successfully" after the report has been sent.

[0176] AI surveillance system processing steps

[0177] Step 1: Initialize your device

[0178] The server initializes the surveillance cameras and audio devices installed in schools and homes.

[0179] Input: Monitoring device connection information

[0180] Data processing: Installing device drivers and checking connections

[0181] Output: Device initialization completion notification

[0182] Specific behavior: The server logs "Monitoring device initializing."

[0183] Step 2: Capture the data

[0184] The device uses surveillance cameras and audio devices to capture video and audio data in real time.

[0185] Input: Data captured by security cameras and audio devices

[0186] Data processing: collection and temporary storage of video and audio data

[0187] Output: Notification that capture data has been saved

[0188] Specific behavior: The camera light will turn on and the message "Capturing video and audio" will be displayed.

[0189] Step 3: Send the data

[0190] The device transmits the captured video and audio data to the server.

[0191] Input: Capture data

[0192] Data processing: Packetizing and sending data

[0193] Output: Notification of completion of data transmission to the server

[0194] Specific behavior: The status will be displayed as "Data sending."

[0195] Step 4: Analyze the data

[0196] The server analyzes the received data using a generative AI model to detect abnormal behavior or sounds.

[0197] Input: Submitted capture data

[0198] Data processing: Analyze video and audio data using AI models

[0199] Output: Analysis of abnormal behavior and audio

[0200] Specific behavior: The server will log "Analyzing data." After the analysis is complete, it will display "Abnormal behavior detected."

[0201] Step 5: Evaluate the analysis results and issue a warning

[0202] The server will alert administrators if any unusual activity or sound is detected. These alerts will be sent via email and SMS.

[0203] Input: Anomalous behavior and audio analysis results

[0204] Data processing: generating and sending warning messages

[0205] Output: Warning notice to administrator

[0206] Specific operation: The server records "Notifying administrator of warning" and displays "Warning sent successfully."

[0207] AI-powered helpline process

[0208] Step 1: Provide a chat system and phone line

[0209] The server provides a 24-hour chat system and telephone lines.

[0210] Input: System Operation Information

[0211] Data processing: Initialization of chat system and phone line

[0212] Output: System operation completion notification

[0213] Specific behavior: Display "Chat system and phone line are operational."

[0214] Step 2: User message input

[0215] The user inputs the content of the consultation into the chat system.

[0216] Input: Message from the user

[0217] Data processing: receiving and storing messages

[0218] Output: Message reception notification

[0219] What happens: The chat window will say "Please enter a message."

[0220] Step 3: Parse the message

[0221] The terminal analyzes messages sent by users using natural language processing technology.

[0222] Input: User's message

[0223] Data processing: Message content analysis

[0224] Output: Analysis results

[0225] Specific behavior: The status will be displayed as "Message parsing."

[0226] Step 4: Determine the appropriate response

[0227] Based on the results of message analysis, the server determines an encouraging message and contact information for a support center.

[0228] Input: Analysis results

[0229] Data processing: choosing the appropriate response

[0230] Output: Notification of support decision

[0231] What happens: The server logs "Determining appropriate action."

[0232] Step 5: Send a reply message

[0233] The server sends the determined reply message to the user.

[0234] Input: Support details

[0235] Data processing: Creating and sending a response message

[0236] Output: Reply notification to user

[0237] Specific operation: Displays "Sending reply message" and displays "Reply completed" after sending.

[0238] Processing steps for an AI-powered early warning system

[0239] Step 1: Collect data

[0240] The server collects text, image, and video data from social media and online platforms.

[0241] Input: Data from online platform

[0242] Data processing: data collection and storage

[0243] Output: Data collection completion notification

[0244] What happens: The server logs "Collecting data."

[0245] Step 2: Analyze the data

[0246] The device analyzes the collected data using a generative AI model to detect abnormal posts and signs of bullying or abuse.

[0247] Input: Collected data

[0248] Data processing: Data content analysis

[0249] Output: Analysis results

[0250] Specific operation: The device will display "Analyzing data". After the analysis is complete, it will display "Anomaly detected".

[0251] Step 3: Evaluate the analysis results

[0252] The server evaluates the analysis results sent from the terminal.

[0253] Input: Analysis results

[0254] Data processing: Evaluation of analysis results

[0255] Output: Anomaly detection evaluation results

[0256] Specific behavior: The server logs "Evaluating analysis results."

[0257] Step 4: Send the report

[0258] The server will send a report to the appropriate authorities if an abnormal value is detected.

[0259] Input: Anomaly detection evaluation results

[0260] Data Processing: Report creation and sending

[0261] Output: Report to child protective agencies

[0262] Specific operation: The server displays "Report sending" and then displays "Report sent successfully" after sending.

[0263] (Application example 1)

[0264] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0265] In modern society, problems of abuse and bullying are becoming more serious in the environment surrounding children. Early detection and rapid response to these problems are required, but conventional methods often lack real-time monitoring and emotion analysis, making effective responses impossible. The present invention aims to provide a system that solves these problems and ensures the safety of children.

[0266] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0267] In this invention, the server includes means for acquiring data on the face and voice of a target using an emotion recognition device, means for analyzing the target's emotions from the acquired data using an AI model, means for reporting to an appropriate organization if an abnormal emotion is detected based on the analyzed emotional information, means for detecting the target's behavior and voice in real time and issuing an alert if an abnormality is detected, and means for acquiring video and audio using a wearable device and notifying the user of the analysis results. This makes it possible to monitor children's emotions and behavior in real time, and if an abnormality is detected, to quickly notify the appropriate organization and take early action.

[0268] An "emotion recognition device" is a device that acquires facial and voice data of a subject and analyzes emotions from that data.

[0269] An "AI model" is an algorithm that uses artificial intelligence to analyze data and recognize specific patterns and emotions.

[0270] "Report to appropriate authorities" means reporting any detected abnormal emotional information or abnormal behavior to designated protective authorities or related agencies.

[0271] A "surveillance camera" is a device that captures, observes, and records images within a designated area in real time.

[0272] An "audio device" is an acoustic collection device that captures environmental sounds and conversation sounds and analyzes them as data.

[0273] A "wearable device" is a device that is worn by a user and has a built-in camera and microphone for capturing video and audio.

[0274] "Abnormal emotions" refer to negative emotions such as stress, fear, or anger that exceed the normal range.

[0275] "Abnormal behavior" refers to unnatural behavior that is clearly different from previous patterns or behavior that indicates a serious problem.

[0276] "24-hour chat or phone" refers to a means of communication designed to allow users to make inquiries or receive advice at any time.

[0277] "Natural language processing technology" refers to computer technology for understanding and analyzing human language.

[0278] This invention is a system that uses AI technology to detect child abuse and bullying in real time and promptly report it to the appropriate authorities. The system includes emotion recognition devices, surveillance cameras, audio devices, wearable devices (smart glasses), a 24-hour chat system, and data collection devices from social media and online platforms.

[0279] First, the server acquires the subject's facial and voice data using an emotion recognition device. This is done by capturing data using the camera and microphone installed in the smart glasses and sending it to the server in real time. The server then analyzes the data using an AI model to analyze the subject's emotional information. If abnormal emotional information (such as strong stress or fear) is detected, the server will promptly send a report to the appropriate authorities.

[0280] In addition, the server initializes surveillance cameras and audio devices installed in schools and homes, capturing video and audio data in real time. The captured data is sent to the server and analyzed by an AI model. If abnormal behavior or audio is detected, the server will alert administrators. This system makes it possible to detect signs of bullying and abuse early on.

[0281] The wearable device (smart glasses) is worn by the user (e.g., teacher or parent) on a daily basis and captures the child's face and voice in real time. The device has a built-in camera and microphone, and sends the data to a server that notifies the user of the analysis results. This allows the user to understand the child's emotions and condition on the spot.

[0282] The server also provides a chat system and telephone line that users can use 24 hours a day. When users input their concerns or questions, the device analyzes the message using natural language processing technology and determines the appropriate response. Based on the analysis results, the server provides encouraging messages and contact information for appropriate support centers.

[0283] As a concrete example, there is a series of steps: "Capture camera footage in the program's main loop," "Convert the captured facial area to grayscale and analyze it with an emotion recognition model," and "If a negative emotion is detected, send a warning message to a specified email address." This example would be input as a prompt to the generative AI model as follows:

[0284] Example prompt sentence:

[0285] "Create a program that monitors a child's emotions in real time. Capture camera footage, convert the facial area to grayscale, and analyze it with an emotion recognition model. If a negative emotion is detected, send a warning message to a specified email address."

[0286] In this way, the present invention provides a system that enables early detection of child abuse and bullying and prompt countermeasures.

[0287] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0288] Step 1:

[0289] The server initializes the camera and microphone of the smart glasses and acquires video and audio data from the smart glasses worn by the user. Specifically, the server captures real-time video (face images) and audio (conversation and environmental sounds) transmitted from the smart glasses. The input is video and audio data from the smart glasses, and the output is receiving these data in digital form.

[0290] Step 2:

[0291] The server processes the captured video data, crops out the facial area, and converts it to grayscale. Specifically, the server uses OpenCV to perform facial recognition, extract the facial area, and apply grayscale conversion. The input is the video data from the smart glasses, and the output is grayscale-converted facial image data.

[0292] Step 3:

[0293] The server inputs the grayscale-converted facial image data into an AI model to analyze emotions. Specifically, the server uses TensorFlow to pass the facial image data to an emotion recognition model, and obtains an emotion score (e.g., stress, fear, joy, etc.) as the result. The input is the grayscale-converted facial image data, and the output is the emotion score.

[0294] Step 4:

[0295] The server determines whether an abnormal emotion has been detected based on the acquired emotion score. Specifically, the server checks whether the emotion score exceeds a set threshold, and sets a flag if it is determined to be abnormal. The input is the emotion score, and the output is an anomaly detection flag (true / false).

[0296] Step 5:

[0297] If an abnormal emotion is detected, the server will send a warning message to the specified email address. Specifically, the server uses smtplib to construct an email and send a message stating that an abnormality has occurred. The input is the abnormality detection flag and the child's emotional information, and the output is a warning email.

[0298] Step 6:

[0299] The server stores the voice data acquired in real time and performs voice analysis. Specifically, the voice data is preprocessed using Librosa to extract specific voice features (e.g., voice tone and patterns). The input is the voice data from the smart glasses, and the output is voice feature data.

[0300] Step 7:

[0301] The server inputs the extracted voice features into an AI model and analyzes voice anomalies. Specifically, the voice feature data is passed to a voice recognition model, which determines whether the voice is abnormal (e.g., screaming or crying). The input is the voice feature data, and the output is the voice anomaly detection result.

[0302] Step 8:

[0303] If an abnormal sound is detected, the server issues a warning to the administrator. Specifically, a real-time notification is sent to the administrator's device, prompting them to check the situation on-site. The input is the result of the sound abnormality judgment, and the output is a warning notification to the administrator.

[0304] Step 9:

[0305] The server receives messages from users of a 24-hour chat system. Specifically, the server receives and stores text messages sent through the chat interface. The input is the message from the user, and the output is the stored message data.

[0306] Step 10:

[0307] The server analyzes the received message using natural language processing technology and determines its content. Specifically, it uses an NLP algorithm to classify the message content and determine the appropriate response. The input is the stored message data, and the output is the response content as the analysis result.

[0308] Step 11:

[0309] Based on the analysis results, the server provides appropriate support, such as encouraging messages and information on relevant support centers. The input is the response content based on the analysis results, and the output is a support message to the user.

[0310] Through these steps, the system can monitor children's emotions and behaviors in real time, respond quickly if an abnormality is detected, and provide support as needed to resolve the problem early.

[0311] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0312] This invention is a system that combines an emotion recognition device and an emotion engine to detect signs of child abuse and bullying at an early stage and report them to the appropriate authorities. This enables real-time emotion analysis and understanding of long-term emotional changes, allowing for more effective countermeasures.

[0313] Combining AI emotion recognition and emotion engine

[0314] The server initializes the camera and microphone devices, and the device captures the user's face and voice data in real time. This data is then analyzed by the emotion recognition device and emotion engine to detect whether the child is experiencing negative emotions such as fear or stress.

[0315] The emotion engine has the ability to cumulatively record a user's emotional data and analyze long-term changes in emotions. This makes it possible to grasp not only short-term emotions but also emotional patterns. For example, if a child is continuously feeling stressed, the emotion engine will recognize this as a pattern and issue an alert if necessary.

[0316] Collaboration between AI surveillance system and emotion engine

[0317] The server initializes surveillance cameras and audio devices installed in schools and homes, and the devices capture video and audio data in real time. The captured data is sent to the server and analyzed in real time by an AI model via an emotion engine. If abnormal behavior or audio is detected, the emotion engine records it and analyzes the pattern. If the abnormal pattern continues, the server issues an alert to the administrator.

[0318] AI-powered helpline and emotion engine integration

[0319] The server provides a chat system and telephone line that users can use 24 hours a day. When users input their worries or concerns, the device analyzes the message using natural language processing technology through an emotion engine. Based on the analysis results, the emotion engine can provide customized support content to the user from the server as needed. For example, if a user complains of stress over a long period of time, the emotion engine will understand this and can respond by providing them with a consultation service with a professional counselor.

[0320] Social Media and Online Platform Integration

[0321] The server collects data from social media and other online platforms, including text, images, and video data. The device then uses an emotion engine to analyze this data using natural language processing and image recognition technology to detect abnormal posts and signs of bullying or abuse. If an anomaly is detected, the emotion engine records it and sends a report to the appropriate authorities.

[0322] Specific examples

[0323] For example, suppose a user captures a child's face and voice at home. The device analyzes the data with an emotion engine and detects that the child is experiencing high levels of stress. The emotion engine records this and, if this pattern continues, the server sends a report to a child welfare center.

[0324] The same applies if a school's surveillance camera captures a student being bullied. The device analyzes this with its emotion engine and records it as an abnormal pattern. If the abnormality persists, the server immediately issues an alert to the school administrator.

[0325] In addition, a user can send a message of concern, such as "I don't want to go to school," through the helpline's chat system. The device analyzes the message using an emotion engine, which then accumulates and records it. When the accumulated stress exceeds a certain level, the server provides the user with an encouraging message and the contact information for a child consultation hotline.

[0326] As a result, the present invention makes it possible to detect child abuse and bullying at an early stage and take prompt measures. By utilizing the emotion engine, it is possible to focus on not only short-term emotion recognition but also long-term emotional changes, allowing for more appropriate responses.

[0327] The processing flow will be explained below.

[0328] Processing flow of AI-based emotion recognition and emotion engine combination

[0329] Step 1:

[0330] The server initializes the camera and microphone devices so that face and voice data can be captured.

[0331] Step 2:

[0332] The device captures the user's face and voice data in real time at regular intervals. The face data is obtained from the camera and the voice data is obtained from the microphone.

[0333] Step 3:

[0334] The device converts captured facial data into a grayscale image and preprocesses audio data based on the sample rate, making it easier for AI models to analyze.

[0335] Step 4:

[0336] The device inputs the preprocessed data into an emotion engine to analyze the user's emotions, which can identify emotions such as fear, stress, and joy.

[0337] Step 5:

[0338] The emotion engine cumulatively records the analyzed emotion information and analyzes short-term and long-term emotion patterns.

[0339] Step 6:

[0340] The server sends a report to the appropriate authorities if the emotion engine detects abnormal emotions (e.g., high levels of fear or stress).

[0341] Processing flow of collaboration between AI monitoring system and emotion engine

[0342] Step 1:

[0343] The server initializes the surveillance cameras and audio devices installed in schools and homes, enabling real-time monitoring.

[0344] Step 2:

[0345] The device captures video and audio data in real time and transmits it to the server.

[0346] Step 3:

[0347] The server receives the captured video and audio data and analyzes it in real time using an AI model via an emotion engine, which detects signs of bullying or abuse.

[0348] Step 4:

[0349] When the emotion engine detects abnormal behavior or voice, it records the analysis results cumulatively and analyzes consistent abnormal patterns.

[0350] Step 5:

[0351] If the server continues to exhibit abnormal patterns, it will alert the administrator.

[0352] AI-powered helpline and emotion engine integration process flow

[0353] Step 1:

[0354] The server provides a chat system and telephone line that users can use 24 hours a day, so that users can consult with the service at any time.

[0355] Step 2:

[0356] A user sends a message through the chat or phone system, and the message arrives at the server.

[0357] Step 3:

[0358] The server analyzes the messages it receives using natural language processing technology through an emotion engine and classifies the user's emotions and content.

[0359] Step 4:

[0360] The emotion engine cumulatively records the analysis results and determines customized support content.

[0361] Step 5:

[0362] Based on the analysis results, the server provides users with appropriate support, encouraging messages, contact information for specialist helplines, and more.

[0363] Social Media and Online Platform Integration Process Flow

[0364] Step 1:

[0365] The server uses APIs to collect data from social media and other online platforms, retrieving relevant text, image, and video data.

[0366] Step 2:

[0367] The device analyzes the collected data using an emotion engine, natural language processing, and image recognition technology to detect abnormal posts and signs of bullying or abuse.

[0368] Step 3:

[0369] The emotion engine cumulatively records the analysis results and analyzes the content of abnormal posts.

[0370] Step 4:

[0371] If the server detects an anomaly, it will send a report to the appropriate authorities, which will include details of the post where the anomaly occurred.

[0372] Step 5:

[0373] The server centrally manages the data and analysis results acquired in real time, records the details of detected anomalies, and notifies the administrator so that necessary measures can be taken.

[0374] The above are the specific processing steps for each system.

[0375] Example 2

[0376] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0377] Child abuse and bullying are serious social issues that require early detection and appropriate countermeasures. However, current systems have difficulty detecting changes in children's emotions in real time and identifying long-term trends. Furthermore, there is a lack of effective means to accurately detect signs of bullying and abuse and send reports to the appropriate authorities.

[0378] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0379] In this invention, the server includes: [means for initializing a device that acquires user face and voice data and capturing data in real time; [means for analyzing the user's emotions from the acquired data using a generative AI model and sending the data to an emotion engine; [means for accumulating and recording the analysis results and analyzing long-term changes in emotions; and [means for detecting abnormal emotional patterns from the accumulated emotional data and sending a report to an appropriate institution if the abnormality persists.] This makes it possible [to detect signs of child abuse and bullying early, analyze emotions in real time and understand long-term changes in emotions, and quickly take appropriate measures].

[0380] "User" refers to a person such as a child or student who uses the emotion recognition system, or their guardian or educator.

[0381] The "server" refers to a machine that serves as the central hub of the entire emotion recognition system and performs data collection, analysis, accumulating records, and anomaly detection processing.

[0382] "Terminal" refers to a device used by a user, equipped with a camera and microphone, that captures the user's face and voice data in real time.

[0383] An "emotion recognition device" refers to a device that analyzes a user's emotions from acquired facial and voice data.

[0384] An "emotion engine" refers to software or hardware that accumulates and records emotional data analyzed by an emotion recognition device and analyzes long-term changes in emotions.

[0385] "Generative AI model" refers to an artificial intelligence model used in emotion recognition devices to analyze a user's facial and voice data and identify emotions.

[0386] An "abnormal emotional pattern" refers to a state in which negative emotions such as anxiety, stress, or fear that exceed the normal range are continuously observed in the user's emotional data.

[0387] "Appropriate agencies" refers to specialized agencies needed to provide early intervention and support, such as child guidance centers, school administrators, and counselors.

[0388] "Chat System" means an online platform that provides users with messaging services available 24 hours a day.

[0389] "Surveillance cameras and audio devices" refers to devices installed in schools or homes that capture video and audio data.

[0390] "Natural language processing technology" refers to technology that enables computer programs to understand and analyze human language.

[0391] "Report" refers to a report that detects abnormal emotional patterns, signs of bullying or abuse, and notifies the appropriate authorities.

[0392] "Real-time" refers to data being processed and analyzed as soon as it is acquired.

[0393] This invention relates to a system that combines an emotion recognition device and an emotion engine to detect signs of child abuse and bullying at an early stage and report them to the appropriate authorities. This enables real-time emotion analysis and understanding of long-term emotional changes, allowing for more effective countermeasures to be taken.

[0394] System configuration

[0395] 1. Initialize the camera and microphone

[0396] The server initializes the camera and microphone devices, including loading the necessary device drivers. The device then activates the camera and microphone to capture the user's face and voice data in real time. This functionality can be used at home or in schools.

[0397] 2. Data capture and transmission

[0398] The device collects video data from the camera and audio data from the microphone and sends them to a server, where they are passed to an emotion recognition device in real time.

[0399] 3. Emotion analysis

[0400] The server analyzes the received data using an emotion recognition device. This analysis process uses a generative AI model, which identifies emotions from the user's facial expressions and tone of voice. At this point, negative emotions such as stress, fear, and sadness may be detected.

[0401] 4. Accumulative Recording of Emotional Data

[0402] The emotion engine records the analysis results cumulatively. The recorded data is used to analyze long-term changes in emotions. For example, by graphing the fluctuations in emotions over the past week, it can determine whether the user is experiencing persistent stress.

[0403] 5. Detecting abnormal patterns and generating alerts

[0404] The emotion engine has the ability to detect abnormal emotion patterns based on accumulated data. If the abnormal pattern persists, the emotion engine notifies the server, which then sends an alert to the appropriate authorities. This alert may be sent via email or SMS.

[0405] Specific examples

[0406] Examples of use within the home

[0407] Suppose a user captures their child's face and voice at home. The device sends this data to a server in real time and analyzes it with an emotion engine. The emotion engine detects when the child is experiencing high levels of stress and records this cumulatively. If this pattern continues, the server automatically sends a report to a child welfare center.

[0408] Examples of use within schools

[0409] At school, surveillance cameras capture footage of a specific student being bullied. Devices transmit this video data to a server in real time, where an emotion recognition device identifies the bullying. If the abnormal behavior continues, the server immediately issues an alert to school administrators.

[0410] Helpline usage example

[0411] A user sends a message such as "I don't want to go to school" through the helpline chat system. The device passes the message to the emotion engine, which analyzes it using natural language processing technology. The emotion engine detects continued stress and notifies the server. The server then provides the user with an encouraging message and contact information for a child consultation center.

[0412] Prompt Sentence Examples

[0413] Describe the implementation of a system that analyzes conversations in the home in real time to detect signs of stress.

[0414] This makes it possible to detect early signs of child abuse and bullying and quickly implement appropriate countermeasures. The use of an emotion engine allows for not only short-term emotion recognition but also long-term emotional changes, enabling more accurate responses.

[0415] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0416] Step 1:

[0417] The server initializes the camera and microphone devices. The initialized devices are ready to capture the user's face and voice data in real time. The input includes the server recognizing the camera and microphone devices and loading their drivers. The output is that the devices are successfully enabled and capture begins.

[0418] Specific operation:

[0419] The server detects the connected cameras and microphones, loads their device drivers, and captures test video and audio once to ensure the devices are initialized correctly.

[0420] Step 2:

[0421] The device uses the initialized camera and microphone to capture the user's face and voice data in real time. The input includes the video data captured by the camera and the audio data captured by the microphone. The output is sent to the server.

[0422] Specific operation:

[0423] The device captures video images via a camera installed in front of the user and ambient conversations and sounds via a microphone, and transmits the captured data as packets to a server at a fixed frame rate and sampling rate.

[0424] Step 3:

[0425] The server passes the acquired video and audio data to the emotion recognition device. The input includes the raw data sent from the device. The output is converted into a data format that the emotion recognition device can analyze.

[0426] Specific operation:

[0427] The server buffers the video and audio data received from the device and converts it into a format suitable for input to the emotion recognition device, including data compression and noise filtering.

[0428] Step 4:

[0429] The emotion recognizer uses a generative AI model to analyze video and audio data to identify the user's emotions. The input includes formatted data passed from the server. The output is a tag for the identified emotion, which is sent to the emotion engine.

[0430] Specific operation:

[0431] The emotion recognition device analyzes facial expressions from video data and vocal tone and tempo from audio data, and the generative AI model uses this information to tag emotions such as "stress," "joy," and "fear."

[0432] Step 5:

[0433] The emotion engine accumulates and records the analyzed emotion data. The input includes emotion tags sent from the emotion recognizer. The output is to add new emotion data to the accumulated emotion database.

[0434] Specific operation:

[0435] The emotion engine stores date- and time-stamped emotion data for each user in a cumulative database, allowing for tracking of emotional changes over time.

[0436] Step 6:

[0437] The emotion engine analyzes the accumulated data and detects anomalous emotion patterns. It includes the accumulated database as input and generates anomalous emotion pattern detection results as output.

[0438] Specific operation:

[0439] The emotion engine analyzes a user's emotional data over time to detect abnormal fluctuations and continuous negative emotions, such as when the "stress" tag is applied for more than a week.

[0440] Step 7:

[0441] When an anomalous emotional pattern is detected, the emotion engine notifies the server, which includes the anomaly detection results as input and sends an alert to the appropriate authorities as output.

[0442] Specific operation:

[0443] When the emotion engine detects an abnormal pattern, it sends an alert message to the server, which then sends an email or SMS to the child consultation center or school administrator based on the message.

[0444] The above is the flow of processing of the program of this system.

[0445] (Application example 2)

[0446] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0447] Conventional emotion recognition systems are limited to simply recognizing temporary emotions, making them unable to detect serious problems such as child abuse and bullying early on. They also lack the functionality to respond immediately when negative emotions persist, making it difficult to issue immediate warnings or notify appropriate authorities. Furthermore, they lack sufficient mechanisms for analyzing long-term emotional changes and identifying abnormal patterns, resulting in no fundamental solution. Therefore, a more comprehensive and effective emotion recognition system is needed.

[0448] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring data of the user's face and voice using an emotion recognition device, means for analyzing the user's emotions from the acquired data using an AI model, means for reporting to an appropriate agency if an abnormal emotion is detected based on the analyzed emotion information, means for recording cumulative emotion data and analyzing long-term emotion patterns, and means for sending notifications and alerts to the management device in real time if many negative emotions are detected. This makes it possible to monitor changes in children's emotions in real time and understand short-term and long-term emotion patterns, enabling early detection of bullying and abuse and rapid response.

[0449] An "emotion recognition device" is a device that acquires data on a user's face and voice and determines the user's emotions based on that data.

[0450] "Data acquisition" refers to the act of collecting information about a user's face and voice in real time using input devices such as a camera or microphone.

[0451] "AI Model" means a computational model that uses artificial intelligence technology and includes algorithms that analyze user data for emotion recognition.

[0452] "Emotional information" is data that indicates the user's emotional state, obtained as a result of analysis using an AI model.

[0453] "Reporting to appropriate authorities" refers to the act of reporting information to relevant authorities, such as child welfare centers or school administrators, when abnormal emotions are detected based on emotional information.

[0454] "Cumulative emotion data" is historical emotion data that is accumulated by recording the user's emotional changes over a long period of time.

[0455] "Long-term emotional pattern analysis" refers to the act of using accumulated emotional data to analyze trends in user emotional fluctuations and abnormal patterns.

[0456] "Sending notifications and alerts to management devices in real time" refers to the act of immediately sending warnings or notification messages to management devices such as smartphones and PCs when the emotion recognition device detects an abnormal emotion.

[0457] A "surveillance camera" is a camera device installed to acquire video data.

[0458] An "audio device" is a microphone or related device installed to capture audio data.

[0459] "Natural language processing technology" is an artificial intelligence technology that analyzes natural language data such as text and speech and understands its meaning and emotions.

[0460] This invention is a comprehensive emotion recognition system for early detection of signs of bullying and abuse and taking appropriate countermeasures. The system combines emotion recognition devices, AI models, long-term emotion pattern analysis, and real-time notification and alert functions.

[0461] Overall system configuration

[0462] 1. Hardware

[0463] Surveillance cameras: High-resolution cameras are installed in classrooms and around the school, such as generic high-resolution cameras (e.g., Hikvision DS-2CD2143G0-I).

[0464] Audio Device: A high-sensitivity microphone is installed in the designated location. For example, a generic high-sensitivity microphone (e.g., Blue Yeti USB Microphone).

[0465] Server: A high-performance server will be installed to perform data analysis and recording. For example, a generic high-performance server (e.g., Dell PowerEdge T340) will be installed.

[0466] Managed devices: Smartphones and PCs used by administrators. Smartphones run on iOS or Android, and PCs run on Windows or MacOS.

[0467] 2. Software

[0468] Sentiment Analysis: We use Python and TensorFlow to build a custom AI model that analyzes the user's facial and voice data in real time to recognize emotions.

[0469] Database: A MySQL database is used to manage the emotion data.

[0470] Notifications: Firebase Cloud Messaging and email notification systems are used to send real-time notifications to managed devices when an anomaly is detected.

[0471] Natural Language Processing: Incorporates natural language processing technology to analyze chats and messages.

[0472] Specific examples to realize

[0473] 1. Data Acquisition

[0474] The server collects video and audio data from surveillance cameras and audio devices within the school, recording the children's daily activities and comments.

[0475] 2. Data Analysis

[0476] The collected data is analyzed in real time by an emotion recognition system and AI models running on a server. For example, a custom AI model built with Python and TensorFlow analyzes facial expressions and tone of voice to detect negative emotions such as stress or fear.

[0477] 3. Notification of abnormalities

[0478] The collected and analyzed emotion data is cumulatively recorded in a MySQL database. If a large number of negative emotions are detected within a certain period of time, an alert is sent to the administrator's smartphone or PC using Firebase Cloud Messaging. For example, a message stating "Stress detected" is sent in real time.

[0479] 4. Long-term emotional pattern analysis

[0480] The server analyzes long-term emotional patterns using the accumulated emotional data, and if an abnormal pattern is detected, it sends a report to the appropriate organization, such as a child welfare center or parent.

[0481] 5. Natural Language Processing Support

[0482] Chat messages and phone consultations from users are analyzed using natural language processing technology. For example, if a message such as "I don't want to go to school" is sent, the content is analyzed by an emotion engine. If the accumulated stress exceeds a threshold, the system will connect the user to a professional counselor and provide customized support.

[0483] Examples and Prompts

[0484] An example of a prompt is as follows:

[0485] "From the video and audio of this child, you can determine if he is stressed or scared."

[0486] "Please conduct a long-term analysis to see if the students in this video are experiencing continuous stress."

[0487] This allows the emotion recognition system of the present invention to monitor changes in children's emotions in real time and respond quickly if an abnormality is detected.

[0488] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0489] Step 1:

[0490] The server initializes the surveillance cameras and audio devices and captures video and audio data in real time. The input is real-time data from the surveillance cameras and microphones, and the output is raw data stored on the server. This allows us to collect users' daily actions and comments.

[0491] Step 2:

[0492] The server sends the captured video and audio data to the emotion recognition device and AI model for data analysis. The input is the raw data acquired in step 1, and the output is the analyzed emotion information. Specifically, it uses Python and TensorFlow to analyze facial expressions and tone of voice to determine the user's emotion.

[0493] Step 3:

[0494] The server cumulatively records the analyzed emotional information in a MySQL database. The input is the emotional information obtained in step 2, and the output is the cumulative data stored in the database. This data is used to understand long-term emotional patterns.

[0495] Step 4:

[0496] The server analyzes the accumulated emotional data and sends real-time notifications and alerts to managed devices if abnormal emotional patterns are detected. The input is the accumulated data retrieved from the MySQL database, and the output is a notification sent to the administrator's smartphone or PC. For example, a message saying "Stress detected" is sent using Firebase Cloud Messaging.

[0497] Step 5:

[0498] When a user sends a message via 24-hour chat or phone, the server analyzes the message using natural language processing technology. The input is the user's message, and the output is the analyzed result. Specifically, natural language processing technology is used to analyze the content of the message and understand the user's emotional state.

[0499] Step 6:

[0500] The server records cumulative emotional data based on the analysis results and provides appropriate support. The input is the emotion analysis results obtained in step 5, and the output is the provision of support content. For example, if the accumulated stress exceeds a threshold, it will connect the user to a professional counselor or provide customized support.

[0501] Step 7:

[0502] The server uses the accumulated emotional data to further analyze long-term emotional patterns and, if necessary, sends reports to appropriate agencies. The input is long-term emotional data, and the output is a report to child consultation centers and parents. Specifically, the AI ​​model recognizes abnormal patterns, and creates and sends reports to appropriate agencies.

[0503] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0504] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0505] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0506] [Second embodiment]

[0507] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0508] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0509] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0510] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0511] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0512] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0513] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0514] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0515] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0516] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0517] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0518] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0519] The present invention is a system that uses AI technology to detect child abuse and bullying in real time and report it to the appropriate authorities. The system includes emotion recognition devices, surveillance cameras, audio devices, 24-hour chat and phone systems, and data collection devices from social media and online platforms.

[0520] AI-powered emotion recognition

[0521] The server initializes the camera and microphone devices, and the device captures the user's face and voice data in real time. This data is analyzed using an AI model to detect whether the child is experiencing negative emotions such as fear or stress. If the detected emotional information shows abnormal values, the server sends a report to the appropriate child protection agency.

[0522] AI surveillance system

[0523] The server initializes surveillance cameras and audio devices installed in schools and homes, and the devices capture video and audio data in real time. The captured data is sent to the server and analyzed by an AI model. If abnormal behavior or audio is detected, the server will alert the administrator. This system makes it possible to detect signs of everyday bullying and abuse at an early stage.

[0524] AI-powered helpline

[0525] The server provides a chat system and telephone line that users can use 24 hours a day. When users input their concerns or questions, the device analyzes the message using natural language processing technology and determines the appropriate response. Based on the analysis results, the server provides encouraging messages and contact information for appropriate support centers.

[0526] AI-powered early warning system

[0527] The server collects data from social media and other online platforms, including text, images, and video data. The device then analyzes this data using AI models to detect unusual posts and signs of bullying or abuse. If an anomaly is detected, the server sends a report to the appropriate authorities.

[0528] Specific examples

[0529] For example, suppose a user captures a video of a child's face and voice. The device analyzes the data and detects that the child is experiencing high levels of stress. The server immediately reports this information to a child welfare center. Also, if a school's surveillance camera captures a student being bullied, the device will detect this and the server will alert the school administrator.

[0530] For example, if a user sends a message of advice such as "I don't want to go to school" through the helpline chat system, the device analyzes the content and the server replies with an encouraging message and contact information for a child consultation center. Furthermore, if signs of bullying are detected in a post on social media, the server will use that information to report the matter to the appropriate authorities.

[0531] As a result, the present invention makes it possible to detect child abuse and bullying at an early stage and to take prompt measures against it.

[0532] The processing flow will be explained below.

[0533] AI emotion recognition processing flow

[0534] Step 1:

[0535] The server initializes the camera and microphone devices, which allows face and voice data to be captured in real time.

[0536] Step 2:

[0537] The device captures face and voice data at regular intervals: face data from the camera and voice data from the microphone.

[0538] Step 3:

[0539] The device converts captured facial data into a grayscale image and preprocesses audio data based on the sample rate, making it easier for AI models to analyze.

[0540] Step 4:

[0541] The device then inputs the preprocessed data into an AI model to analyze emotions, which can identify multiple emotions such as fear, stress, and joy.

[0542] Step 5:

[0543] The server receives the analyzed emotional information and reports to the appropriate authorities if abnormal emotions (e.g., high levels of fear or stress) are detected.

[0544] AI monitoring system processing flow

[0545] Step 1:

[0546] The server initializes the surveillance cameras and audio devices installed in schools and homes, enabling continuous monitoring.

[0547] Step 2:

[0548] The device captures video and audio data in real time and transmits it to the server.

[0549] Step 3:

[0550] The server analyzes the video and audio data received in real time using an AI model designed to detect abnormal behavior and audio.

[0551] Step 4:

[0552] If the server detects any unusual activity or sound, it will immediately alert administrators, including details of when and where the anomaly occurred.

[0553] AI-powered helpline process

[0554] Step 1:

[0555] The server provides a 24-hour chat and phone system, which users can use to input their concerns and questions.

[0556] Step 2:

[0557] A user sends a message through the chat or phone system, and the message arrives at the server.

[0558] Step 3:

[0559] The server analyzes the messages it receives using natural language processing technology and classifies their emotions and content.

[0560] Step 4:

[0561] Based on the analysis results, the server provides the user with necessary support and encouraging messages, and in some cases, contact information for specialized helplines.

[0562] Processing flow of an AI-based early warning system

[0563] Step 1:

[0564] The server uses APIs to collect data from social media and other online platforms, retrieving relevant text, image, and video data.

[0565] Step 2:

[0566] The device analyzes the collected data using natural language processing and image recognition techniques, conducting sentiment analysis on text data and detecting specific abnormal behaviors on images and videos.

[0567] Step 3:

[0568] If the server detects any anomalous posting content, it will send a report to the appropriate authorities, which will include details of the posting where the anomaly occurred.

[0569] Step 4:

[0570] The server centrally manages the data and analysis results acquired in real time, records the details of detected anomalies, and notifies the administrator so that necessary measures can be taken.

[0571] The above are the specific processing steps for each system.

[0572] Example 1

[0573] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0574] Child abuse and bullying are serious social issues, and early detection and reporting to appropriate authorities is essential. However, current systems have difficulty analyzing emotions and behaviors in real time and detecting abnormalities, preventing effective countermeasures. Furthermore, even 24-hour helplines have difficulty providing prompt support. An efficient and highly accurate system is needed to solve these problems.

[0575] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0576] In this invention, the server includes: [means for acquiring data on the user's face and voice using an emotion recognition device]; [means for analyzing the user's emotions from the acquired data using a generative AI model]; and [means for reporting to an appropriate institution if an abnormal emotion is detected based on the analyzed emotion information]. This enables real-time analysis of children's emotions and early detection and reporting of abnormalities.

[0577] The server also includes a means for initializing the device, a means for the user to capture face and voice data, and a means for evaluating the analysis results and sending a report. This allows for efficient use of the device and improved accuracy of the analysis.

[0578] Furthermore, the server includes: [means for installing surveillance cameras and audio devices to detect signs of bullying and abuse; [means for analyzing acquired camera and audio data in real time using a generative AI model; and [means for detecting abnormal behavior or audio from the analysis results and issuing an alert.] This enables early detection and rapid response to bullying and abuse using surveillance cameras and audio devices.

[0579] Furthermore, the server includes a means for receiving messages from users via 24-hour chat or telephone, a means for analyzing the received messages using natural language processing technology, and a means for providing appropriate support based on the analysis results. This makes it possible to provide prompt and accurate support in response to consultation messages from users.

[0580] A "server" is a central control device that processes and manages data, communicates with other devices via a network, analyzes various data, and sends reports.

[0581] A "terminal" is a device that is directly operated by a user, and is a device that captures data through a camera or microphone and transmits it to a server.

[0582] A "user" is a person who operates or inputs data using a system, and in particular, who captures data or sends messages.

[0583] An "emotion recognition device" is a device that uses sensors such as cameras and microphones to acquire facial and voice data and analyzes emotions from that data.

[0584] A "generative AI model" is an artificial intelligence model developed using machine learning and deep learning techniques, and includes algorithms for emotion analysis and anomaly detection.

[0585] "Abnormal emotions" are emotions that are outside the normal range, and primarily refer to negative emotions such as fear, stress, and anger.

[0586] "Report" means a report or notification sent to an appropriate institution when an abnormality is detected, and includes the results of the analysis and its specific contents.

[0587] A "surveillance camera" is a photographic device that captures video in real time and transmits the data to a server.

[0588] An "audio device" is a recording device that captures audio, converts it into data, and sends it to a server.

[0589] A "chat system" is a software system for exchanging text-based messages in real time, allowing for 24-hour support.

[0590] "Natural language processing technology" refers to technology for understanding and analyzing human language, and includes algorithms that perform semantic and emotional analysis of text data.

[0591] "Appropriate support content" refers to solutions and support information provided to users in response to their inquiries and problems, including encouraging messages and contact information for support desks.

[0592] The present invention is a system that uses AI technology to detect child abuse and bullying in real time and report it to the appropriate authorities. The system includes emotion recognition devices, surveillance cameras, audio devices, 24-hour chat and phone systems, and data collection devices from social media and online platforms.

[0593] An embodiment of an AI-based emotion recognition system

[0594] First, the server initializes the camera and microphone connected to the emotion recognition device. This initialization includes installing device drivers and verifying the connection. Next, the device uses these devices to allow the user (parent or teacher) to capture the child's face and voice. This captured data is analyzed in real time using a generative AI model.

[0595] As a specific example, if a user captures video of their child's face and voice at school and the analysis results indicate that the child is experiencing severe stress, the server will immediately send a report to a child consultation center.

[0596] AI surveillance system implementation

[0597] The server then initializes the surveillance cameras and audio devices installed in schools and homes. The devices capture video and audio data from these devices in real time. This data is sent to the server and analyzed by the generative AI model. If any abnormal behavior or audio is detected, the server alerts administrators.

[0598] As a concrete example, if a school's surveillance camera captures a scene in which a particular student is being bullied, and if abnormal behavior is detected as a result of analyzing this data, the server will immediately send an alert to the school administrator.

[0599] AI-powered helpline implementation

[0600] In addition, the server provides a chat system and telephone line that users can use 24 hours a day. When users enter their concerns or questions into the chat system, the device analyzes the message using natural language processing technology. Based on the analysis, the server provides encouraging messages and contact information for appropriate support centers.

[0601] As a specific example, if a user types "I don't want to go to school" in a chat, the device analyzes the content and the server replies with an encouraging message and contact information for a child consultation center.

[0602] An example of an early warning system using AI

[0603] The server also collects text, image, and video data from social media and online platforms. The device analyzes this data using generative AI models to detect signs of bullying or abuse. If an anomaly is detected, the server sends a report to the appropriate authorities.

[0604] As a specific example, if a post such as "I'm being bullied by a classmate" is detected on a social networking site, the device will analyze the information and the server will send a report to the appropriate authority.

[0605] Prompt Sentence Examples

[0606] Detect posts on social media that say things like "I'm being bullied at school."

[0607] If your analysis results indicate that "this child is experiencing severe stress," please send a report to the child consultation center.

[0608] Please send an encouraging reply to the consultation message "I don't want to go to school."

[0609] The purpose of this invention is to provide a safe environment by using this system to enable early detection of child abuse and bullying and prompt countermeasures.

[0610] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0611] Processing steps of AI emotion recognition system

[0612] Step 1: Initialize your device

[0613] The server initializes the camera and microphone connected to the emotion recognition device, which includes installing the device drivers and checking the connection.

[0614] Input: Device connection information

[0615] Data processing: Installing device drivers and checking connections

[0616] Output: Device initialization completion notification

[0617] What happens: The server logs the message "Initializing camera and microphone."

[0618] Step 2: Capture the data

[0619] The device uses the initialized camera and microphone to allow the user (parent or teacher) to capture the child's face and voice, and the captured data is temporarily stored on the device.

[0620] Input: Data captured by camera and microphone

[0621] Data processing: collection and temporary storage of face and voice data

[0622] Output: Notification that capture data has been saved

[0623] Specific behavior: The device will display a pop-up saying "Data capturing."

[0624] Step 3: Analyze the data

[0625] The device then inputs the captured facial and voice data into generative AI models for real-time analysis, including convolutional neural networks (CNNs) and recurrent neural networks (RNNs).

[0626] Input: Capture data

[0627] Data processing: Analyzing face and voice data with AI models

[0628] Output: Emotion analysis results

[0629] Specific behavior: The device will display the status "Analyzing data". Once the analysis is complete, the result will be displayed as "Negative emotion detected".

[0630] Step 4: Evaluate the analysis results and submit a report

[0631] The server receives the analysis results sent from the device and, if any abnormal values ​​are detected, sends a report to the appropriate child protection agency. This report includes the analysis data and the results.

[0632] Input: Sentiment analysis results

[0633] Data processing: Evaluation of analysis results and report generation

[0634] Output: Report sent to child protective services

[0635] Specific behavior: The server logs "Report sending" and displays "Report sent successfully" after the report has been sent.

[0636] AI surveillance system processing steps

[0637] Step 1: Initialize your device

[0638] The server initializes the surveillance cameras and audio devices installed in schools and homes.

[0639] Input: Monitoring device connection information

[0640] Data processing: Installing device drivers and checking connections

[0641] Output: Device initialization completion notification

[0642] Specific behavior: The server logs "Monitoring device initializing."

[0643] Step 2: Capture the data

[0644] The device uses surveillance cameras and audio devices to capture video and audio data in real time.

[0645] Input: Data captured by security cameras and audio devices

[0646] Data processing: collection and temporary storage of video and audio data

[0647] Output: Notification that capture data has been saved

[0648] Specific behavior: The camera light will turn on and the message "Capturing video and audio" will be displayed.

[0649] Step 3: Send the data

[0650] The device transmits the captured video and audio data to the server.

[0651] Input: Capture data

[0652] Data processing: Packetizing and sending data

[0653] Output: Notification of completion of data transmission to the server

[0654] Specific behavior: The status will be displayed as "Data sending."

[0655] Step 4: Analyze the data

[0656] The server analyzes the received data using a generative AI model to detect abnormal behavior or sounds.

[0657] Input: Submitted capture data

[0658] Data processing: Analyze video and audio data using AI models

[0659] Output: Analysis of abnormal behavior and audio

[0660] Specific behavior: The server will log "Analyzing data." After the analysis is complete, it will display "Abnormal behavior detected."

[0661] Step 5: Evaluate the analysis results and issue a warning

[0662] The server will alert administrators if any unusual activity or sound is detected. These alerts will be sent via email and SMS.

[0663] Input: Anomalous behavior and audio analysis results

[0664] Data processing: generating and sending warning messages

[0665] Output: Warning notice to administrator

[0666] Specific operation: The server records "Notifying administrator of warning" and displays "Warning sent successfully."

[0667] AI-powered helpline process

[0668] Step 1: Provide a chat system and phone line

[0669] The server provides a 24-hour chat system and telephone lines.

[0670] Input: System Operation Information

[0671] Data processing: Initialization of chat system and phone line

[0672] Output: System operation completion notification

[0673] Specific behavior: Display "Chat system and phone line are operational."

[0674] Step 2: User message input

[0675] The user inputs the content of the consultation into the chat system.

[0676] Input: Message from the user

[0677] Data processing: receiving and storing messages

[0678] Output: Message reception notification

[0679] What happens: The chat window will say "Please enter a message."

[0680] Step 3: Parse the message

[0681] The terminal analyzes messages sent by users using natural language processing technology.

[0682] Input: User's message

[0683] Data processing: Message content analysis

[0684] Output: Analysis results

[0685] Specific behavior: The status will be displayed as "Message parsing."

[0686] Step 4: Determine the appropriate response

[0687] Based on the results of message analysis, the server determines an encouraging message and contact information for a support center.

[0688] Input: Analysis results

[0689] Data processing: choosing the appropriate response

[0690] Output: Notification of support decision

[0691] What happens: The server logs "Determining appropriate action."

[0692] Step 5: Send a reply message

[0693] The server sends the determined reply message to the user.

[0694] Input: Support details

[0695] Data processing: Creating and sending a response message

[0696] Output: Reply notification to user

[0697] Specific operation: Displays "Sending reply message" and displays "Reply completed" after sending.

[0698] Processing steps for an AI-powered early warning system

[0699] Step 1: Collect data

[0700] The server collects text, image, and video data from social media and online platforms.

[0701] Input: Data from online platform

[0702] Data processing: data collection and storage

[0703] Output: Data collection completion notification

[0704] What happens: The server logs "Collecting data."

[0705] Step 2: Analyze the data

[0706] The device analyzes the collected data using a generative AI model to detect abnormal posts and signs of bullying or abuse.

[0707] Input: Collected data

[0708] Data processing: Data content analysis

[0709] Output: Analysis results

[0710] Specific operation: The device will display "Analyzing data". After the analysis is complete, it will display "Anomaly detected".

[0711] Step 3: Evaluate the analysis results

[0712] The server evaluates the analysis results sent from the terminal.

[0713] Input: Analysis results

[0714] Data processing: Evaluation of analysis results

[0715] Output: Anomaly detection evaluation results

[0716] Specific behavior: The server logs "Evaluating analysis results."

[0717] Step 4: Send the report

[0718] The server will send a report to the appropriate authorities if an abnormal value is detected.

[0719] Input: Anomaly detection evaluation results

[0720] Data Processing: Report creation and sending

[0721] Output: Report to child protective agencies

[0722] Specific operation: The server displays "Report sending" and then displays "Report sent successfully" after sending.

[0723] (Application example 1)

[0724] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0725] In modern society, problems of abuse and bullying are becoming more serious in the environment surrounding children. Early detection and rapid response to these problems are required, but conventional methods often lack real-time monitoring and emotion analysis, making effective responses impossible. The present invention aims to provide a system that solves these problems and ensures the safety of children.

[0726] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0727] In this invention, the server includes means for acquiring data on the face and voice of a target using an emotion recognition device, means for analyzing the target's emotions from the acquired data using an AI model, means for reporting to an appropriate organization if an abnormal emotion is detected based on the analyzed emotional information, means for detecting the target's behavior and voice in real time and issuing an alert if an abnormality is detected, and means for acquiring video and audio using a wearable device and notifying the user of the analysis results. This makes it possible to monitor children's emotions and behavior in real time, and if an abnormality is detected, to quickly notify the appropriate organization and take early action.

[0728] An "emotion recognition device" is a device that acquires facial and voice data of a subject and analyzes emotions from that data.

[0729] An "AI model" is an algorithm that uses artificial intelligence to analyze data and recognize specific patterns and emotions.

[0730] "Report to appropriate authorities" means reporting any detected abnormal emotional information or abnormal behavior to designated protective authorities or related agencies.

[0731] A "surveillance camera" is a device that captures, observes, and records images within a designated area in real time.

[0732] An "audio device" is an acoustic collection device that captures environmental sounds and conversation sounds and analyzes them as data.

[0733] A "wearable device" is a device that is worn by a user and has a built-in camera and microphone for capturing video and audio.

[0734] "Abnormal emotions" refer to negative emotions such as stress, fear, or anger that exceed the normal range.

[0735] "Abnormal behavior" refers to unnatural behavior that is clearly different from previous patterns or behavior that indicates a serious problem.

[0736] "24-hour chat or phone" refers to a means of communication designed to allow users to make inquiries or receive advice at any time.

[0737] "Natural language processing technology" refers to computer technology for understanding and analyzing human language.

[0738] This invention is a system that uses AI technology to detect child abuse and bullying in real time and promptly report it to the appropriate authorities. The system includes emotion recognition devices, surveillance cameras, audio devices, wearable devices (smart glasses), a 24-hour chat system, and data collection devices from social media and online platforms.

[0739] First, the server acquires the subject's facial and voice data using an emotion recognition device. This is done by capturing data using the camera and microphone installed in the smart glasses and sending it to the server in real time. The server then analyzes the data using an AI model to analyze the subject's emotional information. If abnormal emotional information (such as strong stress or fear) is detected, the server will promptly send a report to the appropriate authorities.

[0740] In addition, the server initializes surveillance cameras and audio devices installed in schools and homes, capturing video and audio data in real time. The captured data is sent to the server and analyzed by an AI model. If abnormal behavior or audio is detected, the server will alert administrators. This system makes it possible to detect signs of bullying and abuse early on.

[0741] The wearable device (smart glasses) is worn by the user (e.g., teacher or parent) on a daily basis and captures the child's face and voice in real time. The device has a built-in camera and microphone, and sends the data to a server that notifies the user of the analysis results. This allows the user to understand the child's emotions and condition on the spot.

[0742] The server also provides a chat system and telephone line that users can use 24 hours a day. When users input their concerns or questions, the device analyzes the message using natural language processing technology and determines the appropriate response. Based on the analysis results, the server provides encouraging messages and contact information for appropriate support centers.

[0743] As a concrete example, there is a series of steps: "Capture camera footage in the program's main loop," "Convert the captured facial area to grayscale and analyze it with an emotion recognition model," and "If a negative emotion is detected, send a warning message to a specified email address." This example would be input as a prompt to the generative AI model as follows:

[0744] Example prompt sentence:

[0745] "Create a program that monitors a child's emotions in real time. Capture camera footage, convert the facial area to grayscale, and analyze it with an emotion recognition model. If a negative emotion is detected, send a warning message to a specified email address."

[0746] In this way, the present invention provides a system that enables early detection of child abuse and bullying and prompt countermeasures.

[0747] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0748] Step 1:

[0749] The server initializes the camera and microphone of the smart glasses and acquires video and audio data from the smart glasses worn by the user. Specifically, the server captures real-time video (face images) and audio (conversation and environmental sounds) transmitted from the smart glasses. The input is video and audio data from the smart glasses, and the output is receiving these data in digital form.

[0750] Step 2:

[0751] The server processes the captured video data, crops out the facial area, and converts it to grayscale. Specifically, the server uses OpenCV to perform facial recognition, extract the facial area, and apply grayscale conversion. The input is the video data from the smart glasses, and the output is grayscale-converted facial image data.

[0752] Step 3:

[0753] The server inputs the grayscale-converted facial image data into an AI model to analyze emotions. Specifically, the server uses TensorFlow to pass the facial image data to an emotion recognition model, and obtains an emotion score (e.g., stress, fear, joy, etc.) as the result. The input is the grayscale-converted facial image data, and the output is the emotion score.

[0754] Step 4:

[0755] The server determines whether an abnormal emotion has been detected based on the acquired emotion score. Specifically, the server checks whether the emotion score exceeds a set threshold, and sets a flag if it is determined to be abnormal. The input is the emotion score, and the output is an anomaly detection flag (true / false).

[0756] Step 5:

[0757] If an abnormal emotion is detected, the server will send a warning message to the specified email address. Specifically, the server uses smtplib to construct an email and send a message stating that an abnormality has occurred. The input is the abnormality detection flag and the child's emotional information, and the output is a warning email.

[0758] Step 6:

[0759] The server stores the voice data acquired in real time and performs voice analysis. Specifically, the voice data is preprocessed using Librosa to extract specific voice features (e.g., voice tone and patterns). The input is the voice data from the smart glasses, and the output is voice feature data.

[0760] Step 7:

[0761] The server inputs the extracted voice features into an AI model and analyzes voice anomalies. Specifically, the voice feature data is passed to a voice recognition model, which determines whether the voice is abnormal (e.g., screaming or crying). The input is the voice feature data, and the output is the voice anomaly detection result.

[0762] Step 8:

[0763] If an abnormal sound is detected, the server issues a warning to the administrator. Specifically, a real-time notification is sent to the administrator's device, prompting them to check the situation on-site. The input is the result of the sound abnormality judgment, and the output is a warning notification to the administrator.

[0764] Step 9:

[0765] The server receives messages from users of a 24-hour chat system. Specifically, the server receives and stores text messages sent through the chat interface. The input is the message from the user, and the output is the stored message data.

[0766] Step 10:

[0767] The server analyzes the received message using natural language processing technology and determines its content. Specifically, it uses an NLP algorithm to classify the message content and determine the appropriate response. The input is the stored message data, and the output is the response content as the analysis result.

[0768] Step 11:

[0769] Based on the analysis results, the server provides appropriate support, such as encouraging messages and information on relevant support centers. The input is the response content based on the analysis results, and the output is a support message to the user.

[0770] Through these steps, the system can monitor children's emotions and behaviors in real time, respond quickly if an abnormality is detected, and provide support as needed to resolve the problem early.

[0771] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0772] This invention is a system that combines an emotion recognition device and an emotion engine to detect signs of child abuse and bullying at an early stage and report them to the appropriate authorities. This enables real-time emotion analysis and understanding of long-term emotional changes, allowing for more effective countermeasures.

[0773] Combining AI emotion recognition and emotion engine

[0774] The server initializes the camera and microphone devices, and the device captures the user's face and voice data in real time. This data is then analyzed by the emotion recognition device and emotion engine to detect whether the child is experiencing negative emotions such as fear or stress.

[0775] The emotion engine has the ability to cumulatively record a user's emotional data and analyze long-term changes in emotions. This makes it possible to grasp not only short-term emotions but also emotional patterns. For example, if a child is continuously feeling stressed, the emotion engine will recognize this as a pattern and issue an alert if necessary.

[0776] Collaboration between AI surveillance system and emotion engine

[0777] The server initializes surveillance cameras and audio devices installed in schools and homes, and the devices capture video and audio data in real time. The captured data is sent to the server and analyzed in real time by an AI model via an emotion engine. If abnormal behavior or audio is detected, the emotion engine records it and analyzes the pattern. If the abnormal pattern continues, the server issues an alert to the administrator.

[0778] AI-powered helpline and emotion engine integration

[0779] The server provides a chat system and telephone line that users can use 24 hours a day. When users input their worries or concerns, the device analyzes the message using natural language processing technology through an emotion engine. Based on the analysis results, the emotion engine can provide customized support content to the user from the server as needed. For example, if a user complains of stress over a long period of time, the emotion engine will understand this and can respond by providing them with a consultation service with a professional counselor.

[0780] Social Media and Online Platform Integration

[0781] The server collects data from social media and other online platforms, including text, images, and video data. The device then uses an emotion engine to analyze this data using natural language processing and image recognition technology to detect abnormal posts and signs of bullying or abuse. If an anomaly is detected, the emotion engine records it and sends a report to the appropriate authorities.

[0782] Specific examples

[0783] For example, suppose a user captures a child's face and voice at home. The device analyzes the data with an emotion engine and detects that the child is experiencing high levels of stress. The emotion engine records this and, if this pattern continues, the server sends a report to a child welfare center.

[0784] The same applies if a school's surveillance camera captures a student being bullied. The device analyzes this with its emotion engine and records it as an abnormal pattern. If the abnormality persists, the server immediately issues an alert to the school administrator.

[0785] In addition, a user can send a message of concern, such as "I don't want to go to school," through the helpline's chat system. The device analyzes the message using an emotion engine, which then accumulates and records it. When the accumulated stress exceeds a certain level, the server provides the user with an encouraging message and the contact information for a child consultation hotline.

[0786] As a result, the present invention makes it possible to detect child abuse and bullying at an early stage and take prompt measures. By utilizing the emotion engine, it is possible to focus on not only short-term emotion recognition but also long-term emotional changes, allowing for more appropriate responses.

[0787] The processing flow will be explained below.

[0788] Processing flow of AI-based emotion recognition and emotion engine combination

[0789] Step 1:

[0790] The server initializes the camera and microphone devices so that face and voice data can be captured.

[0791] Step 2:

[0792] The device captures the user's face and voice data in real time at regular intervals. The face data is obtained from the camera and the voice data is obtained from the microphone.

[0793] Step 3:

[0794] The device converts captured facial data into a grayscale image and preprocesses audio data based on the sample rate, making it easier for AI models to analyze.

[0795] Step 4:

[0796] The device inputs the preprocessed data into an emotion engine to analyze the user's emotions, which can identify emotions such as fear, stress, and joy.

[0797] Step 5:

[0798] The emotion engine cumulatively records the analyzed emotion information and analyzes short-term and long-term emotion patterns.

[0799] Step 6:

[0800] The server sends a report to the appropriate authorities if the emotion engine detects abnormal emotions (e.g., high levels of fear or stress).

[0801] Processing flow of collaboration between AI monitoring system and emotion engine

[0802] Step 1:

[0803] The server initializes the surveillance cameras and audio devices installed in schools and homes, enabling real-time monitoring.

[0804] Step 2:

[0805] The device captures video and audio data in real time and transmits it to the server.

[0806] Step 3:

[0807] The server receives the captured video and audio data and analyzes it in real time using an AI model via an emotion engine, which detects signs of bullying or abuse.

[0808] Step 4:

[0809] When the emotion engine detects abnormal behavior or voice, it records the analysis results cumulatively and analyzes consistent abnormal patterns.

[0810] Step 5:

[0811] If the server continues to exhibit abnormal patterns, it will alert the administrator.

[0812] AI-powered helpline and emotion engine integration process flow

[0813] Step 1:

[0814] The server provides a chat system and telephone line that users can use 24 hours a day, so that users can consult with the service at any time.

[0815] Step 2:

[0816] A user sends a message through the chat or phone system, and the message arrives at the server.

[0817] Step 3:

[0818] The server analyzes the messages it receives using natural language processing technology through an emotion engine and classifies the user's emotions and content.

[0819] Step 4:

[0820] The emotion engine cumulatively records the analysis results and determines customized support content.

[0821] Step 5:

[0822] Based on the analysis results, the server provides users with appropriate support, encouraging messages, contact information for specialist helplines, and more.

[0823] Social Media and Online Platform Integration Process Flow

[0824] Step 1:

[0825] The server uses APIs to collect data from social media and other online platforms, retrieving relevant text, image, and video data.

[0826] Step 2:

[0827] The device analyzes the collected data using an emotion engine, natural language processing, and image recognition technology to detect abnormal posts and signs of bullying or abuse.

[0828] Step 3:

[0829] The emotion engine cumulatively records the analysis results and analyzes the content of abnormal posts.

[0830] Step 4:

[0831] If the server detects an anomaly, it will send a report to the appropriate authorities, which will include details of the post where the anomaly occurred.

[0832] Step 5:

[0833] The server centrally manages the data and analysis results acquired in real time, records the details of detected anomalies, and notifies the administrator so that necessary measures can be taken.

[0834] The above are the specific processing steps for each system.

[0835] Example 2

[0836] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0837] Child abuse and bullying are serious social issues that require early detection and appropriate countermeasures. However, current systems have difficulty detecting changes in children's emotions in real time and identifying long-term trends. Furthermore, there is a lack of effective means to accurately detect signs of bullying and abuse and send reports to the appropriate authorities.

[0838] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0839] In this invention, the server includes: [means for initializing a device that acquires user face and voice data and capturing data in real time; [means for analyzing the user's emotions from the acquired data using a generative AI model and sending the data to an emotion engine; [means for accumulating and recording the analysis results and analyzing long-term changes in emotions; and [means for detecting abnormal emotional patterns from the accumulated emotional data and sending a report to an appropriate institution if the abnormality persists.] This makes it possible [to detect signs of child abuse and bullying early, analyze emotions in real time and understand long-term changes in emotions, and quickly take appropriate measures].

[0840] "User" refers to a person such as a child or student who uses the emotion recognition system, or their guardian or educator.

[0841] The "server" refers to a machine that serves as the central hub of the entire emotion recognition system and performs data collection, analysis, accumulating records, and anomaly detection processing.

[0842] "Terminal" refers to a device used by a user, equipped with a camera and microphone, that captures the user's face and voice data in real time.

[0843] An "emotion recognition device" refers to a device that analyzes a user's emotions from acquired facial and voice data.

[0844] An "emotion engine" refers to software or hardware that accumulates and records emotional data analyzed by an emotion recognition device and analyzes long-term changes in emotions.

[0845] "Generative AI model" refers to an artificial intelligence model used in emotion recognition devices to analyze a user's facial and voice data and identify emotions.

[0846] An "abnormal emotional pattern" refers to a state in which negative emotions such as anxiety, stress, or fear that exceed the normal range are continuously observed in the user's emotional data.

[0847] "Appropriate agencies" refers to specialized agencies needed to provide early intervention and support, such as child guidance centers, school administrators, and counselors.

[0848] "Chat System" means an online platform that provides users with messaging services available 24 hours a day.

[0849] "Surveillance cameras and audio devices" refers to devices installed in schools or homes that capture video and audio data.

[0850] "Natural language processing technology" refers to technology that enables computer programs to understand and analyze human language.

[0851] "Report" refers to a report that detects abnormal emotional patterns, signs of bullying or abuse, and notifies the appropriate authorities.

[0852] "Real-time" refers to data being processed and analyzed as soon as it is acquired.

[0853] This invention relates to a system that combines an emotion recognition device and an emotion engine to detect signs of child abuse and bullying at an early stage and report them to the appropriate authorities. This enables real-time emotion analysis and understanding of long-term emotional changes, allowing for more effective countermeasures to be taken.

[0854] System configuration

[0855] 1. Initialize the camera and microphone

[0856] The server initializes the camera and microphone devices, including loading the necessary device drivers. The device then activates the camera and microphone to capture the user's face and voice data in real time. This functionality can be used at home or in schools.

[0857] 2. Data capture and transmission

[0858] The device collects video data from the camera and audio data from the microphone and sends them to a server, where they are passed to an emotion recognition device in real time.

[0859] 3. Emotion analysis

[0860] The server analyzes the received data using an emotion recognition device. This analysis process uses a generative AI model, which identifies emotions from the user's facial expressions and tone of voice. At this point, negative emotions such as stress, fear, and sadness may be detected.

[0861] 4. Accumulative Recording of Emotional Data

[0862] The emotion engine records the analysis results cumulatively. The recorded data is used to analyze long-term changes in emotions. For example, by graphing the fluctuations in emotions over the past week, it can determine whether the user is experiencing persistent stress.

[0863] 5. Detecting abnormal patterns and generating alerts

[0864] The emotion engine has the ability to detect abnormal emotion patterns based on accumulated data. If the abnormal pattern persists, the emotion engine notifies the server, which then sends an alert to the appropriate authorities. This alert may be sent via email or SMS.

[0865] Specific examples

[0866] Examples of use within the home

[0867] Suppose a user captures their child's face and voice at home. The device sends this data to a server in real time and analyzes it with an emotion engine. The emotion engine detects when the child is experiencing high levels of stress and records this cumulatively. If this pattern continues, the server automatically sends a report to a child welfare center.

[0868] Examples of use within schools

[0869] At school, surveillance cameras capture footage of a specific student being bullied. Devices transmit this video data to a server in real time, where an emotion recognition device identifies the bullying. If the abnormal behavior continues, the server immediately issues an alert to school administrators.

[0870] Helpline usage example

[0871] A user sends a message such as "I don't want to go to school" through the helpline chat system. The device passes the message to the emotion engine, which analyzes it using natural language processing technology. The emotion engine detects continued stress and notifies the server. The server then provides the user with an encouraging message and contact information for a child consultation center.

[0872] Prompt Sentence Examples

[0873] Describe the implementation of a system that analyzes conversations in the home in real time to detect signs of stress.

[0874] This makes it possible to detect early signs of child abuse and bullying and quickly implement appropriate countermeasures. The use of an emotion engine allows for not only short-term emotion recognition but also long-term emotional changes, enabling more accurate responses.

[0875] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0876] Step 1:

[0877] The server initializes the camera and microphone devices. The initialized devices are ready to capture the user's face and voice data in real time. The input includes the server recognizing the camera and microphone devices and loading their drivers. The output is that the devices are successfully enabled and capture begins.

[0878] Specific operation:

[0879] The server detects the connected cameras and microphones, loads their device drivers, and captures test video and audio once to ensure the devices are initialized correctly.

[0880] Step 2:

[0881] The device uses the initialized camera and microphone to capture the user's face and voice data in real time. The input includes the video data captured by the camera and the audio data captured by the microphone. The output is sent to the server.

[0882] Specific operation:

[0883] The device captures video images via a camera installed in front of the user and ambient conversations and sounds via a microphone, and transmits the captured data as packets to a server at a fixed frame rate and sampling rate.

[0884] Step 3:

[0885] The server passes the acquired video and audio data to the emotion recognition device. The input includes the raw data sent from the device. The output is converted into a data format that the emotion recognition device can analyze.

[0886] Specific operation:

[0887] The server buffers the video and audio data received from the device and converts it into a format suitable for input to the emotion recognition device, including data compression and noise filtering.

[0888] Step 4:

[0889] The emotion recognizer uses a generative AI model to analyze video and audio data to identify the user's emotions. The input includes formatted data passed from the server. The output is a tag for the identified emotion, which is sent to the emotion engine.

[0890] Specific operation:

[0891] The emotion recognition device analyzes facial expressions from video data and vocal tone and tempo from audio data, and the generative AI model uses this information to tag emotions such as "stress," "joy," and "fear."

[0892] Step 5:

[0893] The emotion engine accumulates and records the analyzed emotion data. The input includes emotion tags sent from the emotion recognizer. The output is to add new emotion data to the accumulated emotion database.

[0894] Specific operation:

[0895] The emotion engine stores date- and time-stamped emotion data for each user in a cumulative database, allowing for tracking of emotional changes over time.

[0896] Step 6:

[0897] The emotion engine analyzes the accumulated data and detects anomalous emotion patterns. It includes the accumulated database as input and generates anomalous emotion pattern detection results as output.

[0898] Specific operation:

[0899] The emotion engine analyzes a user's emotional data over time to detect abnormal fluctuations and continuous negative emotions, such as when the "stress" tag is applied for more than a week.

[0900] Step 7:

[0901] When an anomalous emotional pattern is detected, the emotion engine notifies the server, which includes the anomaly detection results as input and sends an alert to the appropriate authorities as output.

[0902] Specific operation:

[0903] When the emotion engine detects an abnormal pattern, it sends an alert message to the server, which then sends an email or SMS to the child consultation center or school administrator based on the message.

[0904] The above is the flow of processing of the program of this system.

[0905] (Application example 2)

[0906] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0907] Conventional emotion recognition systems are limited to simply recognizing temporary emotions, making them unable to detect serious problems such as child abuse and bullying early on. They also lack the functionality to respond immediately when negative emotions persist, making it difficult to issue immediate warnings or notify appropriate authorities. Furthermore, they lack sufficient mechanisms for analyzing long-term emotional changes and identifying abnormal patterns, resulting in no fundamental solution. Therefore, a more comprehensive and effective emotion recognition system is needed.

[0908] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring data of the user's face and voice using an emotion recognition device, means for analyzing the user's emotions from the acquired data using an AI model, means for reporting to an appropriate agency if an abnormal emotion is detected based on the analyzed emotion information, means for recording cumulative emotion data and analyzing long-term emotion patterns, and means for sending notifications and alerts to the management device in real time if many negative emotions are detected. This makes it possible to monitor changes in children's emotions in real time and understand short-term and long-term emotion patterns, enabling early detection of bullying and abuse and rapid response.

[0909] An "emotion recognition device" is a device that acquires data on a user's face and voice and determines the user's emotions based on that data.

[0910] "Data acquisition" refers to the act of collecting information about a user's face and voice in real time using input devices such as a camera or microphone.

[0911] "AI Model" means a computational model that uses artificial intelligence technology and includes algorithms that analyze user data for emotion recognition.

[0912] "Emotional information" is data that indicates the user's emotional state, obtained as a result of analysis using an AI model.

[0913] "Reporting to appropriate authorities" refers to the act of reporting information to relevant authorities, such as child welfare centers or school administrators, when abnormal emotions are detected based on emotional information.

[0914] "Cumulative emotion data" is historical emotion data that is accumulated by recording the user's emotional changes over a long period of time.

[0915] "Long-term emotional pattern analysis" refers to the act of using accumulated emotional data to analyze trends in user emotional fluctuations and abnormal patterns.

[0916] "Sending notifications and alerts to management devices in real time" refers to the act of immediately sending warnings or notification messages to management devices such as smartphones and PCs when the emotion recognition device detects an abnormal emotion.

[0917] A "surveillance camera" is a camera device installed to acquire video data.

[0918] An "audio device" is a microphone or related device installed to capture audio data.

[0919] "Natural language processing technology" is an artificial intelligence technology that analyzes natural language data such as text and speech and understands its meaning and emotions.

[0920] This invention is a comprehensive emotion recognition system for early detection of signs of bullying and abuse and taking appropriate countermeasures. The system combines emotion recognition devices, AI models, long-term emotion pattern analysis, and real-time notification and alert functions.

[0921] Overall system configuration

[0922] 1. Hardware

[0923] Surveillance cameras: High-resolution cameras are installed in classrooms and around the school, such as generic high-resolution cameras (e.g., Hikvision DS-2CD2143G0-I).

[0924] Audio Device: A high-sensitivity microphone is installed in the designated location. For example, a generic high-sensitivity microphone (e.g., Blue Yeti USB Microphone).

[0925] Server: A high-performance server will be installed to perform data analysis and recording. For example, a generic high-performance server (e.g., Dell PowerEdge T340) will be installed.

[0926] Managed devices: Smartphones and PCs used by administrators. Smartphones run on iOS or Android, and PCs run on Windows or MacOS.

[0927] 2. Software

[0928] Sentiment Analysis: We use Python and TensorFlow to build a custom AI model that analyzes the user's facial and voice data in real time to recognize emotions.

[0929] Database: A MySQL database is used to manage the emotion data.

[0930] Notifications: Firebase Cloud Messaging and email notification systems are used to send real-time notifications to managed devices when an anomaly is detected.

[0931] Natural Language Processing: Incorporates natural language processing technology to analyze chats and messages.

[0932] Specific examples to realize

[0933] 1. Data Acquisition

[0934] The server collects video and audio data from surveillance cameras and audio devices within the school, recording the children's daily activities and comments.

[0935] 2. Data Analysis

[0936] The collected data is analyzed in real time by an emotion recognition system and AI models running on a server. For example, a custom AI model built with Python and TensorFlow analyzes facial expressions and tone of voice to detect negative emotions such as stress or fear.

[0937] 3. Notification of abnormalities

[0938] The collected and analyzed emotion data is cumulatively recorded in a MySQL database. If a large number of negative emotions are detected within a certain period of time, an alert is sent to the administrator's smartphone or PC using Firebase Cloud Messaging. For example, a message stating "Stress detected" is sent in real time.

[0939] 4. Long-term emotional pattern analysis

[0940] The server analyzes long-term emotional patterns using the accumulated emotional data, and if an abnormal pattern is detected, it sends a report to the appropriate organization, such as a child welfare center or parent.

[0941] 5. Natural Language Processing Support

[0942] Chat messages and phone consultations from users are analyzed using natural language processing technology. For example, if a message such as "I don't want to go to school" is sent, the content is analyzed by an emotion engine. If the accumulated stress exceeds a threshold, the system will connect the user to a professional counselor and provide customized support.

[0943] Examples and Prompts

[0944] An example of a prompt is as follows:

[0945] "From the video and audio of this child, you can determine if he is stressed or scared."

[0946] "Please conduct a long-term analysis to see if the students in this video are experiencing continuous stress."

[0947] This allows the emotion recognition system of the present invention to monitor changes in children's emotions in real time and respond quickly if an abnormality is detected.

[0948] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0949] Step 1:

[0950] The server initializes the surveillance cameras and audio devices and captures video and audio data in real time. The input is real-time data from the surveillance cameras and microphones, and the output is raw data stored on the server. This allows us to collect users' daily actions and comments.

[0951] Step 2:

[0952] The server sends the captured video and audio data to the emotion recognition device and AI model for data analysis. The input is the raw data acquired in step 1, and the output is the analyzed emotion information. Specifically, it uses Python and TensorFlow to analyze facial expressions and tone of voice to determine the user's emotion.

[0953] Step 3:

[0954] The server cumulatively records the analyzed emotional information in a MySQL database. The input is the emotional information obtained in step 2, and the output is the cumulative data stored in the database. This data is used to understand long-term emotional patterns.

[0955] Step 4:

[0956] The server analyzes the accumulated emotional data and sends real-time notifications and alerts to managed devices if abnormal emotional patterns are detected. The input is the accumulated data retrieved from the MySQL database, and the output is a notification sent to the administrator's smartphone or PC. For example, a message saying "Stress detected" is sent using Firebase Cloud Messaging.

[0957] Step 5:

[0958] When a user sends a message via 24-hour chat or phone, the server analyzes the message using natural language processing technology. The input is the user's message, and the output is the analyzed result. Specifically, natural language processing technology is used to analyze the content of the message and understand the user's emotional state.

[0959] Step 6:

[0960] The server records cumulative emotional data based on the analysis results and provides appropriate support. The input is the emotion analysis results obtained in step 5, and the output is the provision of support content. For example, if the accumulated stress exceeds a threshold, it will connect the user to a professional counselor or provide customized support.

[0961] Step 7:

[0962] The server uses the accumulated emotional data to further analyze long-term emotional patterns and, if necessary, sends reports to appropriate agencies. The input is long-term emotional data, and the output is a report to child consultation centers and parents. Specifically, the AI ​​model recognizes abnormal patterns, and creates and sends reports to appropriate agencies.

[0963] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0964] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0965] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0966] [Third embodiment]

[0967] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0968] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0969] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0970] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0971] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0972] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0973] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0974] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0975] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0976] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0977] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0978] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0979] The present invention is a system that uses AI technology to detect child abuse and bullying in real time and report it to the appropriate authorities. The system includes emotion recognition devices, surveillance cameras, audio devices, 24-hour chat and phone systems, and data collection devices from social media and online platforms.

[0980] AI-powered emotion recognition

[0981] The server initializes the camera and microphone devices, and the device captures the user's face and voice data in real time. This data is analyzed using an AI model to detect whether the child is experiencing negative emotions such as fear or stress. If the detected emotional information shows abnormal values, the server sends a report to the appropriate child protection agency.

[0982] AI surveillance system

[0983] The server initializes surveillance cameras and audio devices installed in schools and homes, and the devices capture video and audio data in real time. The captured data is sent to the server and analyzed by an AI model. If abnormal behavior or audio is detected, the server will alert the administrator. This system makes it possible to detect signs of everyday bullying and abuse at an early stage.

[0984] AI-powered helpline

[0985] The server provides a chat system and telephone line that users can use 24 hours a day. When users input their concerns or questions, the device analyzes the message using natural language processing technology and determines the appropriate response. Based on the analysis results, the server provides encouraging messages and contact information for appropriate support centers.

[0986] AI-powered early warning system

[0987] The server collects data from social media and other online platforms, including text, images, and video data. The device then analyzes this data using AI models to detect unusual posts and signs of bullying or abuse. If an anomaly is detected, the server sends a report to the appropriate authorities.

[0988] Specific examples

[0989] For example, suppose a user captures a video of a child's face and voice. The device analyzes the data and detects that the child is experiencing high levels of stress. The server immediately reports this information to a child welfare center. Also, if a school's surveillance camera captures a student being bullied, the device will detect this and the server will alert the school administrator.

[0990] For example, if a user sends a message of advice such as "I don't want to go to school" through the helpline chat system, the device analyzes the content and the server replies with an encouraging message and contact information for a child consultation center. Furthermore, if signs of bullying are detected in a post on social media, the server will use that information to report the matter to the appropriate authorities.

[0991] As a result, the present invention makes it possible to detect child abuse and bullying at an early stage and to take prompt measures against it.

[0992] The processing flow will be explained below.

[0993] AI emotion recognition processing flow

[0994] Step 1:

[0995] The server initializes the camera and microphone devices, which allows face and voice data to be captured in real time.

[0996] Step 2:

[0997] The device captures face and voice data at regular intervals: face data from the camera and voice data from the microphone.

[0998] Step 3:

[0999] The device converts captured facial data into a grayscale image and preprocesses audio data based on the sample rate, making it easier for AI models to analyze.

[1000] Step 4:

[1001] The device then inputs the preprocessed data into an AI model to analyze emotions, which can identify multiple emotions such as fear, stress, and joy.

[1002] Step 5:

[1003] The server receives the analyzed emotional information and reports to the appropriate authorities if abnormal emotions (e.g., high levels of fear or stress) are detected.

[1004] AI monitoring system processing flow

[1005] Step 1:

[1006] The server initializes the surveillance cameras and audio devices installed in schools and homes, enabling continuous monitoring.

[1007] Step 2:

[1008] The device captures video and audio data in real time and transmits it to the server.

[1009] Step 3:

[1010] The server analyzes the video and audio data received in real time using an AI model designed to detect abnormal behavior and audio.

[1011] Step 4:

[1012] If the server detects any unusual activity or sound, it will immediately alert administrators, including details of when and where the anomaly occurred.

[1013] AI-powered helpline process

[1014] Step 1:

[1015] The server provides a 24-hour chat and phone system, which users can use to input their concerns and questions.

[1016] Step 2:

[1017] A user sends a message through the chat or phone system, and the message arrives at the server.

[1018] Step 3:

[1019] The server analyzes the messages it receives using natural language processing technology and classifies their emotions and content.

[1020] Step 4:

[1021] Based on the analysis results, the server provides the user with necessary support and encouraging messages, and in some cases, contact information for specialized helplines.

[1022] Processing flow of an AI-based early warning system

[1023] Step 1:

[1024] The server uses APIs to collect data from social media and other online platforms, retrieving relevant text, image, and video data.

[1025] Step 2:

[1026] The device analyzes the collected data using natural language processing and image recognition techniques, conducting sentiment analysis on text data and detecting specific abnormal behaviors on images and videos.

[1027] Step 3:

[1028] If the server detects any anomalous posting content, it will send a report to the appropriate authorities, which will include details of the posting where the anomaly occurred.

[1029] Step 4:

[1030] The server centrally manages the data and analysis results acquired in real time, records the details of detected anomalies, and notifies the administrator so that necessary measures can be taken.

[1031] The above are the specific processing steps for each system.

[1032] Example 1

[1033] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1034] Child abuse and bullying are serious social issues, and early detection and reporting to appropriate authorities is essential. However, current systems have difficulty analyzing emotions and behaviors in real time and detecting abnormalities, preventing effective countermeasures. Furthermore, even 24-hour helplines have difficulty providing prompt support. An efficient and highly accurate system is needed to solve these problems.

[1035] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1036] In this invention, the server includes: [means for acquiring data on the user's face and voice using an emotion recognition device]; [means for analyzing the user's emotions from the acquired data using a generative AI model]; and [means for reporting to an appropriate institution if an abnormal emotion is detected based on the analyzed emotion information]. This enables real-time analysis of children's emotions and early detection and reporting of abnormalities.

[1037] The server also includes a means for initializing the device, a means for the user to capture face and voice data, and a means for evaluating the analysis results and sending a report. This allows for efficient use of the device and improved accuracy of the analysis.

[1038] Furthermore, the server includes: [means for installing surveillance cameras and audio devices to detect signs of bullying and abuse; [means for analyzing acquired camera and audio data in real time using a generative AI model; and [means for detecting abnormal behavior or audio from the analysis results and issuing an alert.] This enables early detection and rapid response to bullying and abuse using surveillance cameras and audio devices.

[1039] Furthermore, the server includes a means for receiving messages from users via 24-hour chat or telephone, a means for analyzing the received messages using natural language processing technology, and a means for providing appropriate support based on the analysis results. This makes it possible to provide prompt and accurate support in response to consultation messages from users.

[1040] A "server" is a central control device that processes and manages data, communicates with other devices via a network, analyzes various data, and sends reports.

[1041] A "terminal" is a device that is directly operated by a user, and is a device that captures data through a camera or microphone and transmits it to a server.

[1042] A "user" is a person who operates or inputs data using a system, and in particular, who captures data or sends messages.

[1043] An "emotion recognition device" is a device that uses sensors such as cameras and microphones to acquire facial and voice data and analyzes emotions from that data.

[1044] A "generative AI model" is an artificial intelligence model developed using machine learning and deep learning techniques, and includes algorithms for emotion analysis and anomaly detection.

[1045] "Abnormal emotions" are emotions that are outside the normal range, and primarily refer to negative emotions such as fear, stress, and anger.

[1046] "Report" means a report or notification sent to an appropriate institution when an abnormality is detected, and includes the results of the analysis and its specific contents.

[1047] A "surveillance camera" is a photographic device that captures video in real time and transmits the data to a server.

[1048] An "audio device" is a recording device that captures audio, converts it into data, and sends it to a server.

[1049] A "chat system" is a software system for exchanging text-based messages in real time, allowing for 24-hour support.

[1050] "Natural language processing technology" refers to technology for understanding and analyzing human language, and includes algorithms that perform semantic and emotional analysis of text data.

[1051] "Appropriate support content" refers to solutions and support information provided to users in response to their inquiries and problems, including encouraging messages and contact information for support desks.

[1052] The present invention is a system that uses AI technology to detect child abuse and bullying in real time and report it to the appropriate authorities. The system includes emotion recognition devices, surveillance cameras, audio devices, 24-hour chat and phone systems, and data collection devices from social media and online platforms.

[1053] An embodiment of an AI-based emotion recognition system

[1054] First, the server initializes the camera and microphone connected to the emotion recognition device. This initialization includes installing device drivers and verifying the connection. Next, the device uses these devices to allow the user (parent or teacher) to capture the child's face and voice. This captured data is analyzed in real time using a generative AI model.

[1055] As a specific example, if a user captures video of their child's face and voice at school and the analysis results indicate that the child is experiencing severe stress, the server will immediately send a report to a child consultation center.

[1056] AI surveillance system implementation

[1057] The server then initializes the surveillance cameras and audio devices installed in schools and homes. The devices capture video and audio data from these devices in real time. This data is sent to the server and analyzed by the generative AI model. If any abnormal behavior or audio is detected, the server alerts administrators.

[1058] As a concrete example, if a school's surveillance camera captures a scene in which a particular student is being bullied, and if abnormal behavior is detected as a result of analyzing this data, the server will immediately send an alert to the school administrator.

[1059] AI-powered helpline implementation

[1060] In addition, the server provides a chat system and telephone line that users can use 24 hours a day. When users enter their concerns or questions into the chat system, the device analyzes the message using natural language processing technology. Based on the analysis, the server provides encouraging messages and contact information for appropriate support centers.

[1061] As a specific example, if a user types "I don't want to go to school" in a chat, the device analyzes the content and the server replies with an encouraging message and contact information for a child consultation center.

[1062] An example of an early warning system using AI

[1063] The server also collects text, image, and video data from social media and online platforms. The device analyzes this data using generative AI models to detect signs of bullying or abuse. If an anomaly is detected, the server sends a report to the appropriate authorities.

[1064] As a specific example, if a post such as "I'm being bullied by a classmate" is detected on a social networking site, the device will analyze the information and the server will send a report to the appropriate authority.

[1065] Prompt Sentence Examples

[1066] Detect posts on social media that say things like "I'm being bullied at school."

[1067] If your analysis results indicate that "this child is experiencing severe stress," please send a report to the child consultation center.

[1068] Please send an encouraging reply to the consultation message "I don't want to go to school."

[1069] The purpose of this invention is to provide a safe environment by using this system to enable early detection of child abuse and bullying and prompt countermeasures.

[1070] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1071] Processing steps of AI emotion recognition system

[1072] Step 1: Initialize your device

[1073] The server initializes the camera and microphone connected to the emotion recognition device, which includes installing the device drivers and checking the connection.

[1074] Input: Device connection information

[1075] Data processing: Installing device drivers and checking connections

[1076] Output: Device initialization completion notification

[1077] What happens: The server logs the message "Initializing camera and microphone."

[1078] Step 2: Capture the data

[1079] The device uses the initialized camera and microphone to allow the user (parent or teacher) to capture the child's face and voice, and the captured data is temporarily stored on the device.

[1080] Input: Data captured by camera and microphone

[1081] Data processing: collection and temporary storage of face and voice data

[1082] Output: Notification that capture data has been saved

[1083] Specific behavior: The device will display a pop-up saying "Data capturing."

[1084] Step 3: Analyze the data

[1085] The device then inputs the captured facial and voice data into generative AI models for real-time analysis, including convolutional neural networks (CNNs) and recurrent neural networks (RNNs).

[1086] Input: Capture data

[1087] Data processing: Analyzing face and voice data with AI models

[1088] Output: Emotion analysis results

[1089] Specific behavior: The device will display the status "Analyzing data". Once the analysis is complete, the result will be displayed as "Negative emotion detected".

[1090] Step 4: Evaluate the analysis results and submit a report

[1091] The server receives the analysis results sent from the device and, if any abnormal values ​​are detected, sends a report to the appropriate child protection agency. This report includes the analysis data and the results.

[1092] Input: Sentiment analysis results

[1093] Data processing: Evaluation of analysis results and report generation

[1094] Output: Report sent to child protective services

[1095] Specific behavior: The server logs "Report sending" and displays "Report sent successfully" after the report has been sent.

[1096] AI surveillance system processing steps

[1097] Step 1: Initialize your device

[1098] The server initializes the surveillance cameras and audio devices installed in schools and homes.

[1099] Input: Monitoring device connection information

[1100] Data processing: Installing device drivers and checking connections

[1101] Output: Device initialization completion notification

[1102] Specific behavior: The server logs "Monitoring device initializing."

[1103] Step 2: Capture the data

[1104] The device uses surveillance cameras and audio devices to capture video and audio data in real time.

[1105] Input: Data captured by security cameras and audio devices

[1106] Data processing: collection and temporary storage of video and audio data

[1107] Output: Notification that capture data has been saved

[1108] Specific behavior: The camera light will turn on and the message "Capturing video and audio" will be displayed.

[1109] Step 3: Send the data

[1110] The device transmits the captured video and audio data to the server.

[1111] Input: Capture data

[1112] Data processing: Packetizing and sending data

[1113] Output: Notification of completion of data transmission to the server

[1114] Specific behavior: The status will be displayed as "Data sending."

[1115] Step 4: Analyze the data

[1116] The server analyzes the received data using a generative AI model to detect abnormal behavior or sounds.

[1117] Input: Submitted capture data

[1118] Data processing: Analyze video and audio data using AI models

[1119] Output: Analysis of abnormal behavior and audio

[1120] Specific behavior: The server will log "Analyzing data." After the analysis is complete, it will display "Abnormal behavior detected."

[1121] Step 5: Evaluate the analysis results and issue a warning

[1122] The server will alert administrators if any unusual activity or sound is detected. These alerts will be sent via email and SMS.

[1123] Input: Anomalous behavior and audio analysis results

[1124] Data processing: generating and sending warning messages

[1125] Output: Warning notice to administrator

[1126] Specific operation: The server records "Notifying administrator of warning" and displays "Warning sent successfully."

[1127] AI-powered helpline process

[1128] Step 1: Provide a chat system and phone line

[1129] The server provides a 24-hour chat system and telephone lines.

[1130] Input: System Operation Information

[1131] Data processing: Initialization of chat system and phone line

[1132] Output: System operation completion notification

[1133] Specific behavior: Display "Chat system and phone line are operational."

[1134] Step 2: User message input

[1135] The user inputs the content of the consultation into the chat system.

[1136] Input: Message from the user

[1137] Data processing: receiving and storing messages

[1138] Output: Message reception notification

[1139] What happens: The chat window will say "Please enter a message."

[1140] Step 3: Parse the message

[1141] The terminal analyzes messages sent by users using natural language processing technology.

[1142] Input: User's message

[1143] Data processing: Message content analysis

[1144] Output: Analysis results

[1145] Specific behavior: The status will be displayed as "Message parsing."

[1146] Step 4: Determine the appropriate response

[1147] Based on the results of message analysis, the server determines an encouraging message and contact information for a support center.

[1148] Input: Analysis results

[1149] Data processing: choosing the appropriate response

[1150] Output: Notification of support decision

[1151] What happens: The server logs "Determining appropriate action."

[1152] Step 5: Send a reply message

[1153] The server sends the determined reply message to the user.

[1154] Input: Support details

[1155] Data processing: Creating and sending a response message

[1156] Output: Reply notification to user

[1157] Specific operation: Displays "Sending reply message" and displays "Reply completed" after sending.

[1158] Processing steps for an AI-powered early warning system

[1159] Step 1: Collect data

[1160] The server collects text, image, and video data from social media and online platforms.

[1161] Input: Data from online platform

[1162] Data processing: data collection and storage

[1163] Output: Data collection completion notification

[1164] What happens: The server logs "Collecting data."

[1165] Step 2: Analyze the data

[1166] The device analyzes the collected data using a generative AI model to detect abnormal posts and signs of bullying or abuse.

[1167] Input: Collected data

[1168] Data processing: Data content analysis

[1169] Output: Analysis results

[1170] Specific operation: The device will display "Analyzing data". After the analysis is complete, it will display "Anomaly detected".

[1171] Step 3: Evaluate the analysis results

[1172] The server evaluates the analysis results sent from the terminal.

[1173] Input: Analysis results

[1174] Data processing: Evaluation of analysis results

[1175] Output: Anomaly detection evaluation results

[1176] Specific behavior: The server logs "Evaluating analysis results."

[1177] Step 4: Send the report

[1178] The server will send a report to the appropriate authorities if an abnormal value is detected.

[1179] Input: Anomaly detection evaluation results

[1180] Data Processing: Report creation and sending

[1181] Output: Report to child protective agencies

[1182] Specific operation: The server displays "Report sending" and then displays "Report sent successfully" after sending.

[1183] (Application example 1)

[1184] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1185] In modern society, problems of abuse and bullying are becoming more serious in the environment surrounding children. Early detection and rapid response to these problems are required, but conventional methods often lack real-time monitoring and emotion analysis, making effective responses impossible. The present invention aims to provide a system that solves these problems and ensures the safety of children.

[1186] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1187] In this invention, the server includes means for acquiring data on the face and voice of a target using an emotion recognition device, means for analyzing the target's emotions from the acquired data using an AI model, means for reporting to an appropriate organization if an abnormal emotion is detected based on the analyzed emotional information, means for detecting the target's behavior and voice in real time and issuing an alert if an abnormality is detected, and means for acquiring video and audio using a wearable device and notifying the user of the analysis results. This makes it possible to monitor children's emotions and behavior in real time, and if an abnormality is detected, to quickly notify the appropriate organization and take early action.

[1188] An "emotion recognition device" is a device that acquires facial and voice data of a subject and analyzes emotions from that data.

[1189] An "AI model" is an algorithm that uses artificial intelligence to analyze data and recognize specific patterns and emotions.

[1190] "Report to appropriate authorities" means reporting any detected abnormal emotional information or abnormal behavior to designated protective authorities or related agencies.

[1191] A "surveillance camera" is a device that captures, observes, and records images within a designated area in real time.

[1192] An "audio device" is an acoustic collection device that captures environmental sounds and conversation sounds and analyzes them as data.

[1193] A "wearable device" is a device that is worn by a user and has a built-in camera and microphone for capturing video and audio.

[1194] "Abnormal emotions" refer to negative emotions such as stress, fear, or anger that exceed the normal range.

[1195] "Abnormal behavior" refers to unnatural behavior that is clearly different from previous patterns or behavior that indicates a serious problem.

[1196] "24-hour chat or phone" refers to a means of communication designed to allow users to make inquiries or receive advice at any time.

[1197] "Natural language processing technology" refers to computer technology for understanding and analyzing human language.

[1198] This invention is a system that uses AI technology to detect child abuse and bullying in real time and promptly report it to the appropriate authorities. The system includes emotion recognition devices, surveillance cameras, audio devices, wearable devices (smart glasses), a 24-hour chat system, and data collection devices from social media and online platforms.

[1199] First, the server acquires the subject's facial and voice data using an emotion recognition device. This is done by capturing data using the camera and microphone installed in the smart glasses and sending it to the server in real time. The server then analyzes the data using an AI model to analyze the subject's emotional information. If abnormal emotional information (such as strong stress or fear) is detected, the server will promptly send a report to the appropriate authorities.

[1200] In addition, the server initializes surveillance cameras and audio devices installed in schools and homes, capturing video and audio data in real time. The captured data is sent to the server and analyzed by an AI model. If abnormal behavior or audio is detected, the server will alert administrators. This system makes it possible to detect signs of bullying and abuse early on.

[1201] The wearable device (smart glasses) is worn by the user (e.g., teacher or parent) on a daily basis and captures the child's face and voice in real time. The device has a built-in camera and microphone, and sends the data to a server that notifies the user of the analysis results. This allows the user to understand the child's emotions and condition on the spot.

[1202] The server also provides a chat system and telephone line that users can use 24 hours a day. When users input their concerns or questions, the device analyzes the message using natural language processing technology and determines the appropriate response. Based on the analysis results, the server provides encouraging messages and contact information for appropriate support centers.

[1203] As a concrete example, there is a series of steps: "Capture camera footage in the program's main loop," "Convert the captured facial area to grayscale and analyze it with an emotion recognition model," and "If a negative emotion is detected, send a warning message to a specified email address." This example would be input as a prompt to the generative AI model as follows:

[1204] Example prompt sentence:

[1205] "Create a program that monitors a child's emotions in real time. Capture camera footage, convert the facial area to grayscale, and analyze it with an emotion recognition model. If a negative emotion is detected, send a warning message to a specified email address."

[1206] In this way, the present invention provides a system that enables early detection of child abuse and bullying and prompt countermeasures.

[1207] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1208] Step 1:

[1209] The server initializes the camera and microphone of the smart glasses and acquires video and audio data from the smart glasses worn by the user. Specifically, the server captures real-time video (face images) and audio (conversation and environmental sounds) transmitted from the smart glasses. The input is video and audio data from the smart glasses, and the output is receiving these data in digital form.

[1210] Step 2:

[1211] The server processes the captured video data, crops out the facial area, and converts it to grayscale. Specifically, the server uses OpenCV to perform facial recognition, extract the facial area, and apply grayscale conversion. The input is the video data from the smart glasses, and the output is grayscale-converted facial image data.

[1212] Step 3:

[1213] The server inputs the grayscale-converted facial image data into an AI model to analyze emotions. Specifically, the server uses TensorFlow to pass the facial image data to an emotion recognition model, and obtains an emotion score (e.g., stress, fear, joy, etc.) as the result. The input is the grayscale-converted facial image data, and the output is the emotion score.

[1214] Step 4:

[1215] The server determines whether an abnormal emotion has been detected based on the acquired emotion score. Specifically, the server checks whether the emotion score exceeds a set threshold, and sets a flag if it is determined to be abnormal. The input is the emotion score, and the output is an anomaly detection flag (true / false).

[1216] Step 5:

[1217] If an abnormal emotion is detected, the server will send a warning message to the specified email address. Specifically, the server uses smtplib to construct an email and send a message stating that an abnormality has occurred. The input is the abnormality detection flag and the child's emotional information, and the output is a warning email.

[1218] Step 6:

[1219] The server stores the voice data acquired in real time and performs voice analysis. Specifically, the voice data is preprocessed using Librosa to extract specific voice features (e.g., voice tone and patterns). The input is the voice data from the smart glasses, and the output is voice feature data.

[1220] Step 7:

[1221] The server inputs the extracted voice features into an AI model and analyzes voice anomalies. Specifically, the voice feature data is passed to a voice recognition model, which determines whether the voice is abnormal (e.g., screaming or crying). The input is the voice feature data, and the output is the voice anomaly detection result.

[1222] Step 8:

[1223] If an abnormal sound is detected, the server issues a warning to the administrator. Specifically, a real-time notification is sent to the administrator's device, prompting them to check the situation on-site. The input is the result of the sound abnormality judgment, and the output is a warning notification to the administrator.

[1224] Step 9:

[1225] The server receives messages from users of a 24-hour chat system. Specifically, the server receives and stores text messages sent through the chat interface. The input is the message from the user, and the output is the stored message data.

[1226] Step 10:

[1227] The server analyzes the received message using natural language processing technology and determines its content. Specifically, it uses an NLP algorithm to classify the message content and determine the appropriate response. The input is the stored message data, and the output is the response content as the analysis result.

[1228] Step 11:

[1229] Based on the analysis results, the server provides appropriate support, such as encouraging messages and information on relevant support centers. The input is the response content based on the analysis results, and the output is a support message to the user.

[1230] Through these steps, the system can monitor children's emotions and behaviors in real time, respond quickly if an abnormality is detected, and provide support as needed to resolve the problem early.

[1231] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1232] This invention is a system that combines an emotion recognition device and an emotion engine to detect signs of child abuse and bullying at an early stage and report them to the appropriate authorities. This enables real-time emotion analysis and understanding of long-term emotional changes, allowing for more effective countermeasures.

[1233] Combining AI emotion recognition and emotion engine

[1234] The server initializes the camera and microphone devices, and the device captures the user's face and voice data in real time. This data is then analyzed by the emotion recognition device and emotion engine to detect whether the child is experiencing negative emotions such as fear or stress.

[1235] The emotion engine has the ability to cumulatively record a user's emotional data and analyze long-term changes in emotions. This makes it possible to grasp not only short-term emotions but also emotional patterns. For example, if a child is continuously feeling stressed, the emotion engine will recognize this as a pattern and issue an alert if necessary.

[1236] Collaboration between AI surveillance system and emotion engine

[1237] The server initializes surveillance cameras and audio devices installed in schools and homes, and the devices capture video and audio data in real time. The captured data is sent to the server and analyzed in real time by an AI model via an emotion engine. If abnormal behavior or audio is detected, the emotion engine records it and analyzes the pattern. If the abnormal pattern continues, the server issues an alert to the administrator.

[1238] AI-powered helpline and emotion engine integration

[1239] The server provides a chat system and telephone line that users can use 24 hours a day. When users input their worries or concerns, the device analyzes the message using natural language processing technology through an emotion engine. Based on the analysis results, the emotion engine can provide customized support content to the user from the server as needed. For example, if a user complains of stress over a long period of time, the emotion engine will understand this and can respond by providing them with a consultation service with a professional counselor.

[1240] Social Media and Online Platform Integration

[1241] The server collects data from social media and other online platforms, including text, images, and video data. The device then uses an emotion engine to analyze this data using natural language processing and image recognition technology to detect abnormal posts and signs of bullying or abuse. If an anomaly is detected, the emotion engine records it and sends a report to the appropriate authorities.

[1242] Specific examples

[1243] For example, suppose a user captures a child's face and voice at home. The device analyzes the data with an emotion engine and detects that the child is experiencing high levels of stress. The emotion engine records this and, if this pattern continues, the server sends a report to a child welfare center.

[1244] The same applies if a school's surveillance camera captures a student being bullied. The device analyzes this with its emotion engine and records it as an abnormal pattern. If the abnormality persists, the server immediately issues an alert to the school administrator.

[1245] In addition, a user can send a message of concern, such as "I don't want to go to school," through the helpline's chat system. The device analyzes the message using an emotion engine, which then accumulates and records it. When the accumulated stress exceeds a certain level, the server provides the user with an encouraging message and the contact information for a child consultation hotline.

[1246] As a result, the present invention makes it possible to detect child abuse and bullying at an early stage and take prompt measures. By utilizing the emotion engine, it is possible to focus on not only short-term emotion recognition but also long-term emotional changes, allowing for more appropriate responses.

[1247] The processing flow will be explained below.

[1248] Processing flow of AI-based emotion recognition and emotion engine combination

[1249] Step 1:

[1250] The server initializes the camera and microphone devices so that face and voice data can be captured.

[1251] Step 2:

[1252] The device captures the user's face and voice data in real time at regular intervals. The face data is obtained from the camera and the voice data is obtained from the microphone.

[1253] Step 3:

[1254] The device converts captured facial data into a grayscale image and preprocesses audio data based on the sample rate, making it easier for AI models to analyze.

[1255] Step 4:

[1256] The device inputs the preprocessed data into an emotion engine to analyze the user's emotions, which can identify emotions such as fear, stress, and joy.

[1257] Step 5:

[1258] The emotion engine cumulatively records the analyzed emotion information and analyzes short-term and long-term emotion patterns.

[1259] Step 6:

[1260] The server sends a report to the appropriate authorities if the emotion engine detects abnormal emotions (e.g., high levels of fear or stress).

[1261] Processing flow of collaboration between AI monitoring system and emotion engine

[1262] Step 1:

[1263] The server initializes the surveillance cameras and audio devices installed in schools and homes, enabling real-time monitoring.

[1264] Step 2:

[1265] The device captures video and audio data in real time and transmits it to the server.

[1266] Step 3:

[1267] The server receives the captured video and audio data and analyzes it in real time using an AI model via an emotion engine, which detects signs of bullying or abuse.

[1268] Step 4:

[1269] When the emotion engine detects abnormal behavior or voice, it records the analysis results cumulatively and analyzes consistent abnormal patterns.

[1270] Step 5:

[1271] If the server continues to exhibit abnormal patterns, it will alert the administrator.

[1272] AI-powered helpline and emotion engine integration process flow

[1273] Step 1:

[1274] The server provides a chat system and telephone line that users can use 24 hours a day, so that users can consult with the service at any time.

[1275] Step 2:

[1276] A user sends a message through the chat or phone system, and the message arrives at the server.

[1277] Step 3:

[1278] The server analyzes the messages it receives using natural language processing technology through an emotion engine and classifies the user's emotions and content.

[1279] Step 4:

[1280] The emotion engine cumulatively records the analysis results and determines customized support content.

[1281] Step 5:

[1282] Based on the analysis results, the server provides users with appropriate support, encouraging messages, contact information for specialist helplines, and more.

[1283] Social Media and Online Platform Integration Process Flow

[1284] Step 1:

[1285] The server uses APIs to collect data from social media and other online platforms, retrieving relevant text, image, and video data.

[1286] Step 2:

[1287] The device analyzes the collected data using an emotion engine, natural language processing, and image recognition technology to detect abnormal posts and signs of bullying or abuse.

[1288] Step 3:

[1289] The emotion engine cumulatively records the analysis results and analyzes the content of abnormal posts.

[1290] Step 4:

[1291] If the server detects an anomaly, it will send a report to the appropriate authorities, which will include details of the post where the anomaly occurred.

[1292] Step 5:

[1293] The server centrally manages the data and analysis results acquired in real time, records the details of detected anomalies, and notifies the administrator so that necessary measures can be taken.

[1294] The above are the specific processing steps for each system.

[1295] Example 2

[1296] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1297] Child abuse and bullying are serious social issues that require early detection and appropriate countermeasures. However, current systems have difficulty detecting changes in children's emotions in real time and identifying long-term trends. Furthermore, there is a lack of effective means to accurately detect signs of bullying and abuse and send reports to the appropriate authorities.

[1298] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1299] In this invention, the server includes: [means for initializing a device that acquires user face and voice data and capturing data in real time; [means for analyzing the user's emotions from the acquired data using a generative AI model and sending the data to an emotion engine; [means for accumulating and recording the analysis results and analyzing long-term changes in emotions; and [means for detecting abnormal emotional patterns from the accumulated emotional data and sending a report to an appropriate institution if the abnormality persists.] This makes it possible [to detect signs of child abuse and bullying early, analyze emotions in real time and understand long-term changes in emotions, and quickly take appropriate measures].

[1300] "User" refers to a person such as a child or student who uses the emotion recognition system, or their guardian or educator.

[1301] The "server" refers to a machine that serves as the central hub of the entire emotion recognition system and performs data collection, analysis, accumulating records, and anomaly detection processing.

[1302] "Terminal" refers to a device used by a user, equipped with a camera and microphone, that captures the user's face and voice data in real time.

[1303] An "emotion recognition device" refers to a device that analyzes a user's emotions from acquired facial and voice data.

[1304] An "emotion engine" refers to software or hardware that accumulates and records emotional data analyzed by an emotion recognition device and analyzes long-term changes in emotions.

[1305] "Generative AI model" refers to an artificial intelligence model used in emotion recognition devices to analyze a user's facial and voice data and identify emotions.

[1306] An "abnormal emotional pattern" refers to a state in which negative emotions such as anxiety, stress, or fear that exceed the normal range are continuously observed in the user's emotional data.

[1307] "Appropriate agencies" refers to specialized agencies needed to provide early intervention and support, such as child guidance centers, school administrators, and counselors.

[1308] "Chat System" means an online platform that provides users with messaging services available 24 hours a day.

[1309] "Surveillance cameras and audio devices" refers to devices installed in schools or homes that capture video and audio data.

[1310] "Natural language processing technology" refers to technology that enables computer programs to understand and analyze human language.

[1311] "Report" refers to a report that detects abnormal emotional patterns, signs of bullying or abuse, and notifies the appropriate authorities.

[1312] "Real-time" refers to data being processed and analyzed as soon as it is acquired.

[1313] This invention relates to a system that combines an emotion recognition device and an emotion engine to detect signs of child abuse and bullying at an early stage and report them to the appropriate authorities. This enables real-time emotion analysis and understanding of long-term emotional changes, allowing for more effective countermeasures to be taken.

[1314] System configuration

[1315] 1. Initialize the camera and microphone

[1316] The server initializes the camera and microphone devices, including loading the necessary device drivers. The device then activates the camera and microphone to capture the user's face and voice data in real time. This functionality can be used at home or in schools.

[1317] 2. Data capture and transmission

[1318] The device collects video data from the camera and audio data from the microphone and sends them to a server, where they are passed to an emotion recognition device in real time.

[1319] 3. Emotion analysis

[1320] The server analyzes the received data using an emotion recognition device. This analysis process uses a generative AI model, which identifies emotions from the user's facial expressions and tone of voice. At this point, negative emotions such as stress, fear, and sadness may be detected.

[1321] 4. Accumulative Recording of Emotional Data

[1322] The emotion engine records the analysis results cumulatively. The recorded data is used to analyze long-term changes in emotions. For example, by graphing the fluctuations in emotions over the past week, it can determine whether the user is experiencing persistent stress.

[1323] 5. Detecting abnormal patterns and generating alerts

[1324] The emotion engine has the ability to detect abnormal emotion patterns based on accumulated data. If the abnormal pattern persists, the emotion engine notifies the server, which then sends an alert to the appropriate authorities. This alert may be sent via email or SMS.

[1325] Specific examples

[1326] Examples of use within the home

[1327] Suppose a user captures their child's face and voice at home. The device sends this data to a server in real time and analyzes it with an emotion engine. The emotion engine detects when the child is experiencing high levels of stress and records this cumulatively. If this pattern continues, the server automatically sends a report to a child welfare center.

[1328] Examples of use within schools

[1329] At school, surveillance cameras capture footage of a specific student being bullied. Devices transmit this video data to a server in real time, where an emotion recognition device identifies the bullying. If the abnormal behavior continues, the server immediately issues an alert to school administrators.

[1330] Helpline usage example

[1331] A user sends a message such as "I don't want to go to school" through the helpline chat system. The device passes the message to the emotion engine, which analyzes it using natural language processing technology. The emotion engine detects continued stress and notifies the server. The server then provides the user with an encouraging message and contact information for a child consultation center.

[1332] Prompt Sentence Examples

[1333] Describe the implementation of a system that analyzes conversations in the home in real time to detect signs of stress.

[1334] This makes it possible to detect early signs of child abuse and bullying and quickly implement appropriate countermeasures. The use of an emotion engine allows for not only short-term emotion recognition but also long-term emotional changes, enabling more accurate responses.

[1335] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1336] Step 1:

[1337] The server initializes the camera and microphone devices. The initialized devices are ready to capture the user's face and voice data in real time. The input includes the server recognizing the camera and microphone devices and loading their drivers. The output is that the devices are successfully enabled and capture begins.

[1338] Specific operation:

[1339] The server detects the connected cameras and microphones, loads their device drivers, and captures test video and audio once to ensure the devices are initialized correctly.

[1340] Step 2:

[1341] The device uses the initialized camera and microphone to capture the user's face and voice data in real time. The input includes the video data captured by the camera and the audio data captured by the microphone. The output is sent to the server.

[1342] Specific operation:

[1343] The device captures video images via a camera installed in front of the user and ambient conversations and sounds via a microphone, and transmits the captured data as packets to a server at a fixed frame rate and sampling rate.

[1344] Step 3:

[1345] The server passes the acquired video and audio data to the emotion recognition device. The input includes the raw data sent from the device. The output is converted into a data format that the emotion recognition device can analyze.

[1346] Specific operation:

[1347] The server buffers the video and audio data received from the device and converts it into a format suitable for input to the emotion recognition device, including data compression and noise filtering.

[1348] Step 4:

[1349] The emotion recognizer uses a generative AI model to analyze video and audio data to identify the user's emotions. The input includes formatted data passed from the server. The output is a tag for the identified emotion, which is sent to the emotion engine.

[1350] Specific operation:

[1351] The emotion recognition device analyzes facial expressions from video data and vocal tone and tempo from audio data, and the generative AI model uses this information to tag emotions such as "stress," "joy," and "fear."

[1352] Step 5:

[1353] The emotion engine accumulates and records the analyzed emotion data. The input includes emotion tags sent from the emotion recognizer. The output is to add new emotion data to the accumulated emotion database.

[1354] Specific operation:

[1355] The emotion engine stores date- and time-stamped emotion data for each user in a cumulative database, allowing for tracking of emotional changes over time.

[1356] Step 6:

[1357] The emotion engine analyzes the accumulated data and detects anomalous emotion patterns. It includes the accumulated database as input and generates anomalous emotion pattern detection results as output.

[1358] Specific operation:

[1359] The emotion engine analyzes a user's emotional data over time to detect abnormal fluctuations and continuous negative emotions, such as when the "stress" tag is applied for more than a week.

[1360] Step 7:

[1361] When an anomalous emotional pattern is detected, the emotion engine notifies the server, which includes the anomaly detection results as input and sends an alert to the appropriate authorities as output.

[1362] Specific operation:

[1363] When the emotion engine detects an abnormal pattern, it sends an alert message to the server, which then sends an email or SMS to the child consultation center or school administrator based on the message.

[1364] The above is the flow of processing of the program of this system.

[1365] (Application example 2)

[1366] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1367] Conventional emotion recognition systems are limited to simply recognizing temporary emotions, making them unable to detect serious problems such as child abuse and bullying early on. They also lack the functionality to respond immediately when negative emotions persist, making it difficult to issue immediate warnings or notify appropriate authorities. Furthermore, they lack sufficient mechanisms for analyzing long-term emotional changes and identifying abnormal patterns, resulting in no fundamental solution. Therefore, a more comprehensive and effective emotion recognition system is needed.

[1368] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring data of the user's face and voice using an emotion recognition device, means for analyzing the user's emotions from the acquired data using an AI model, means for reporting to an appropriate agency if an abnormal emotion is detected based on the analyzed emotion information, means for recording cumulative emotion data and analyzing long-term emotion patterns, and means for sending notifications and alerts to the management device in real time if many negative emotions are detected. This makes it possible to monitor changes in children's emotions in real time and understand short-term and long-term emotion patterns, enabling early detection of bullying and abuse and rapid response.

[1369] An "emotion recognition device" is a device that acquires data on a user's face and voice and determines the user's emotions based on that data.

[1370] "Data acquisition" refers to the act of collecting information about a user's face and voice in real time using input devices such as a camera or microphone.

[1371] "AI Model" means a computational model that uses artificial intelligence technology and includes algorithms that analyze user data for emotion recognition.

[1372] "Emotional information" is data that indicates the user's emotional state, obtained as a result of analysis using an AI model.

[1373] "Reporting to appropriate authorities" refers to the act of reporting information to relevant authorities, such as child welfare centers or school administrators, when abnormal emotions are detected based on emotional information.

[1374] "Cumulative emotion data" is historical emotion data that is accumulated by recording the user's emotional changes over a long period of time.

[1375] "Long-term emotional pattern analysis" refers to the act of using accumulated emotional data to analyze trends in user emotional fluctuations and abnormal patterns.

[1376] "Sending notifications and alerts to management devices in real time" refers to the act of immediately sending warnings or notification messages to management devices such as smartphones and PCs when the emotion recognition device detects an abnormal emotion.

[1377] A "surveillance camera" is a camera device installed to acquire video data.

[1378] An "audio device" is a microphone or related device installed to capture audio data.

[1379] "Natural language processing technology" is an artificial intelligence technology that analyzes natural language data such as text and speech and understands its meaning and emotions.

[1380] This invention is a comprehensive emotion recognition system for early detection of signs of bullying and abuse and taking appropriate countermeasures. The system combines emotion recognition devices, AI models, long-term emotion pattern analysis, and real-time notification and alert functions.

[1381] Overall system configuration

[1382] 1. Hardware

[1383] Surveillance cameras: High-resolution cameras are installed in classrooms and around the school, such as generic high-resolution cameras (e.g., Hikvision DS-2CD2143G0-I).

[1384] Audio Device: A high-sensitivity microphone is installed in the designated location. For example, a generic high-sensitivity microphone (e.g., Blue Yeti USB Microphone).

[1385] Server: A high-performance server will be installed to perform data analysis and recording. For example, a generic high-performance server (e.g., Dell PowerEdge T340) will be installed.

[1386] Managed devices: Smartphones and PCs used by administrators. Smartphones run on iOS or Android, and PCs run on Windows or MacOS.

[1387] 2. Software

[1388] Sentiment Analysis: We use Python and TensorFlow to build a custom AI model that analyzes the user's facial and voice data in real time to recognize emotions.

[1389] Database: A MySQL database is used to manage the emotion data.

[1390] Notifications: Firebase Cloud Messaging and email notification systems are used to send real-time notifications to managed devices when an anomaly is detected.

[1391] Natural Language Processing: Incorporates natural language processing technology to analyze chats and messages.

[1392] Specific examples to realize

[1393] 1. Data Acquisition

[1394] The server collects video and audio data from surveillance cameras and audio devices within the school, recording the children's daily activities and comments.

[1395] 2. Data Analysis

[1396] The collected data is analyzed in real time by an emotion recognition system and AI models running on a server. For example, a custom AI model built with Python and TensorFlow analyzes facial expressions and tone of voice to detect negative emotions such as stress or fear.

[1397] 3. Notification of abnormalities

[1398] The collected and analyzed emotion data is cumulatively recorded in a MySQL database. If a large number of negative emotions are detected within a certain period of time, an alert is sent to the administrator's smartphone or PC using Firebase Cloud Messaging. For example, a message stating "Stress detected" is sent in real time.

[1399] 4. Long-term emotional pattern analysis

[1400] The server analyzes long-term emotional patterns using the accumulated emotional data, and if an abnormal pattern is detected, it sends a report to the appropriate organization, such as a child welfare center or parent.

[1401] 5. Natural Language Processing Support

[1402] Chat messages and phone consultations from users are analyzed using natural language processing technology. For example, if a message such as "I don't want to go to school" is sent, the content is analyzed by an emotion engine. If the accumulated stress exceeds a threshold, the system will connect the user to a professional counselor and provide customized support.

[1403] Examples and Prompts

[1404] An example of a prompt is as follows:

[1405] "From the video and audio of this child, you can determine if he is stressed or scared."

[1406] "Please conduct a long-term analysis to see if the students in this video are experiencing continuous stress."

[1407] This allows the emotion recognition system of the present invention to monitor changes in children's emotions in real time and respond quickly if an abnormality is detected.

[1408] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1409] Step 1:

[1410] The server initializes the surveillance cameras and audio devices and captures video and audio data in real time. The input is real-time data from the surveillance cameras and microphones, and the output is raw data stored on the server. This allows us to collect users' daily actions and comments.

[1411] Step 2:

[1412] The server sends the captured video and audio data to the emotion recognition device and AI model for data analysis. The input is the raw data acquired in step 1, and the output is the analyzed emotion information. Specifically, it uses Python and TensorFlow to analyze facial expressions and tone of voice to determine the user's emotion.

[1413] Step 3:

[1414] The server cumulatively records the analyzed emotional information in a MySQL database. The input is the emotional information obtained in step 2, and the output is the cumulative data stored in the database. This data is used to understand long-term emotional patterns.

[1415] Step 4:

[1416] The server analyzes the accumulated emotional data and sends real-time notifications and alerts to managed devices if abnormal emotional patterns are detected. The input is the accumulated data retrieved from the MySQL database, and the output is a notification sent to the administrator's smartphone or PC. For example, a message saying "Stress detected" is sent using Firebase Cloud Messaging.

[1417] Step 5:

[1418] When a user sends a message via 24-hour chat or phone, the server analyzes the message using natural language processing technology. The input is the user's message, and the output is the analyzed result. Specifically, natural language processing technology is used to analyze the content of the message and understand the user's emotional state.

[1419] Step 6:

[1420] The server records cumulative emotional data based on the analysis results and provides appropriate support. The input is the emotion analysis results obtained in step 5, and the output is the provision of support content. For example, if the accumulated stress exceeds a threshold, it will connect the user to a professional counselor or provide customized support.

[1421] Step 7:

[1422] The server uses the accumulated emotional data to further analyze long-term emotional patterns and, if necessary, sends reports to appropriate agencies. The input is long-term emotional data, and the output is a report to child consultation centers and parents. Specifically, the AI ​​model recognizes abnormal patterns, and creates and sends reports to appropriate agencies.

[1423] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1424] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1425] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1426] [Fourth embodiment]

[1427] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1428] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1429] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1430] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1431] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1432] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1433] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1434] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1435] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1436] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1437] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1438] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1439] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1440] The present invention is a system that uses AI technology to detect child abuse and bullying in real time and report it to the appropriate authorities. The system includes emotion recognition devices, surveillance cameras, audio devices, 24-hour chat and phone systems, and data collection devices from social media and online platforms.

[1441] AI-powered emotion recognition

[1442] The server initializes the camera and microphone devices, and the device captures the user's face and voice data in real time. This data is analyzed using an AI model to detect whether the child is experiencing negative emotions such as fear or stress. If the detected emotional information shows abnormal values, the server sends a report to the appropriate child protection agency.

[1443] AI surveillance system

[1444] The server initializes surveillance cameras and audio devices installed in schools and homes, and the devices capture video and audio data in real time. The captured data is sent to the server and analyzed by an AI model. If abnormal behavior or audio is detected, the server will alert the administrator. This system makes it possible to detect signs of everyday bullying and abuse at an early stage.

[1445] AI-powered helpline

[1446] The server provides a chat system and telephone line that users can use 24 hours a day. When users input their concerns or questions, the device analyzes the message using natural language processing technology and determines the appropriate response. Based on the analysis results, the server provides encouraging messages and contact information for appropriate support centers.

[1447] AI-powered early warning system

[1448] The server collects data from social media and other online platforms, including text, images, and video data. The device then analyzes this data using AI models to detect unusual posts and signs of bullying or abuse. If an anomaly is detected, the server sends a report to the appropriate authorities.

[1449] Specific examples

[1450] For example, suppose a user captures a video of a child's face and voice. The device analyzes the data and detects that the child is experiencing high levels of stress. The server immediately reports this information to a child welfare center. Also, if a school's surveillance camera captures a student being bullied, the device will detect this and the server will alert the school administrator.

[1451] For example, if a user sends a message of advice such as "I don't want to go to school" through the helpline chat system, the device analyzes the content and the server replies with an encouraging message and contact information for a child consultation center. Furthermore, if signs of bullying are detected in a post on social media, the server will use that information to report the matter to the appropriate authorities.

[1452] As a result, the present invention makes it possible to detect child abuse and bullying at an early stage and to take prompt measures against it.

[1453] The processing flow will be explained below.

[1454] AI emotion recognition processing flow

[1455] Step 1:

[1456] The server initializes the camera and microphone devices, which allows face and voice data to be captured in real time.

[1457] Step 2:

[1458] The device captures face and voice data at regular intervals: face data from the camera and voice data from the microphone.

[1459] Step 3:

[1460] The device converts captured facial data into a grayscale image and preprocesses audio data based on the sample rate, making it easier for AI models to analyze.

[1461] Step 4:

[1462] The device then inputs the preprocessed data into an AI model to analyze emotions, which can identify multiple emotions such as fear, stress, and joy.

[1463] Step 5:

[1464] The server receives the analyzed emotional information and reports to the appropriate authorities if abnormal emotions (e.g., high levels of fear or stress) are detected.

[1465] AI monitoring system processing flow

[1466] Step 1:

[1467] The server initializes the surveillance cameras and audio devices installed in schools and homes, enabling continuous monitoring.

[1468] Step 2:

[1469] The device captures video and audio data in real time and transmits it to the server.

[1470] Step 3:

[1471] The server analyzes the video and audio data received in real time using an AI model designed to detect abnormal behavior and audio.

[1472] Step 4:

[1473] If the server detects any unusual activity or sound, it will immediately alert administrators, including details of when and where the anomaly occurred.

[1474] AI-powered helpline process

[1475] Step 1:

[1476] The server provides a 24-hour chat and phone system, which users can use to input their concerns and questions.

[1477] Step 2:

[1478] A user sends a message through the chat or phone system, and the message arrives at the server.

[1479] Step 3:

[1480] The server analyzes the messages it receives using natural language processing technology and classifies their emotions and content.

[1481] Step 4:

[1482] Based on the analysis results, the server provides the user with necessary support and encouraging messages, and in some cases, contact information for specialized helplines.

[1483] Processing flow of an AI-based early warning system

[1484] Step 1:

[1485] The server uses APIs to collect data from social media and other online platforms, retrieving relevant text, image, and video data.

[1486] Step 2:

[1487] The device analyzes the collected data using natural language processing and image recognition techniques, conducting sentiment analysis on text data and detecting specific abnormal behaviors on images and videos.

[1488] Step 3:

[1489] If the server detects any anomalous posting content, it will send a report to the appropriate authorities, which will include details of the posting where the anomaly occurred.

[1490] Step 4:

[1491] The server centrally manages the data and analysis results acquired in real time, records the details of detected anomalies, and notifies the administrator so that necessary measures can be taken.

[1492] The above are the specific processing steps for each system.

[1493] Example 1

[1494] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1495] Child abuse and bullying are serious social issues, and early detection and reporting to appropriate authorities is essential. However, current systems have difficulty analyzing emotions and behaviors in real time and detecting abnormalities, preventing effective countermeasures. Furthermore, even 24-hour helplines have difficulty providing prompt support. An efficient and highly accurate system is needed to solve these problems.

[1496] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1497] In this invention, the server includes: [means for acquiring data on the user's face and voice using an emotion recognition device]; [means for analyzing the user's emotions from the acquired data using a generative AI model]; and [means for reporting to an appropriate institution if an abnormal emotion is detected based on the analyzed emotion information]. This enables real-time analysis of children's emotions and early detection and reporting of abnormalities.

[1498] The server also includes a means for initializing the device, a means for the user to capture face and voice data, and a means for evaluating the analysis results and sending a report. This allows for efficient use of the device and improved accuracy of the analysis.

[1499] Furthermore, the server includes: [means for installing surveillance cameras and audio devices to detect signs of bullying and abuse; [means for analyzing acquired camera and audio data in real time using a generative AI model; and [means for detecting abnormal behavior or audio from the analysis results and issuing an alert.] This enables early detection and rapid response to bullying and abuse using surveillance cameras and audio devices.

[1500] Furthermore, the server includes a means for receiving messages from users via 24-hour chat or telephone, a means for analyzing the received messages using natural language processing technology, and a means for providing appropriate support based on the analysis results. This makes it possible to provide prompt and accurate support in response to consultation messages from users.

[1501] A "server" is a central control device that processes and manages data, communicates with other devices via a network, analyzes various data, and sends reports.

[1502] A "terminal" is a device that is directly operated by a user, and is a device that captures data through a camera or microphone and transmits it to a server.

[1503] A "user" is a person who operates or inputs data using a system, and in particular, who captures data or sends messages.

[1504] An "emotion recognition device" is a device that uses sensors such as cameras and microphones to acquire facial and voice data and analyzes emotions from that data.

[1505] A "generative AI model" is an artificial intelligence model developed using machine learning and deep learning techniques, and includes algorithms for emotion analysis and anomaly detection.

[1506] "Abnormal emotions" are emotions that are outside the normal range, and primarily refer to negative emotions such as fear, stress, and anger.

[1507] "Report" means a report or notification sent to an appropriate institution when an abnormality is detected, and includes the results of the analysis and its specific contents.

[1508] A "surveillance camera" is a photographic device that captures video in real time and transmits the data to a server.

[1509] An "audio device" is a recording device that captures audio, converts it into data, and sends it to a server.

[1510] A "chat system" is a software system for exchanging text-based messages in real time, allowing for 24-hour support.

[1511] "Natural language processing technology" refers to technology for understanding and analyzing human language, and includes algorithms that perform semantic and emotional analysis of text data.

[1512] "Appropriate support content" refers to solutions and support information provided to users in response to their inquiries and problems, including encouraging messages and contact information for support desks.

[1513] The present invention is a system that uses AI technology to detect child abuse and bullying in real time and report it to the appropriate authorities. The system includes emotion recognition devices, surveillance cameras, audio devices, 24-hour chat and phone systems, and data collection devices from social media and online platforms.

[1514] An embodiment of an AI-based emotion recognition system

[1515] First, the server initializes the camera and microphone connected to the emotion recognition device. This initialization includes installing device drivers and verifying the connection. Next, the device uses these devices to allow the user (parent or teacher) to capture the child's face and voice. This captured data is analyzed in real time using a generative AI model.

[1516] As a specific example, if a user captures video of their child's face and voice at school and the analysis results indicate that the child is experiencing severe stress, the server will immediately send a report to a child consultation center.

[1517] AI surveillance system implementation

[1518] The server then initializes the surveillance cameras and audio devices installed in schools and homes. The devices capture video and audio data from these devices in real time. This data is sent to the server and analyzed by the generative AI model. If any abnormal behavior or audio is detected, the server alerts administrators.

[1519] As a concrete example, if a school's surveillance camera captures a scene in which a particular student is being bullied, and if abnormal behavior is detected as a result of analyzing this data, the server will immediately send an alert to the school administrator.

[1520] AI-powered helpline implementation

[1521] In addition, the server provides a chat system and telephone line that users can use 24 hours a day. When users enter their concerns or questions into the chat system, the device analyzes the message using natural language processing technology. Based on the analysis, the server provides encouraging messages and contact information for appropriate support centers.

[1522] As a specific example, if a user types "I don't want to go to school" in a chat, the device analyzes the content and the server replies with an encouraging message and contact information for a child consultation center.

[1523] An example of an early warning system using AI

[1524] The server also collects text, image, and video data from social media and online platforms. The device analyzes this data using generative AI models to detect signs of bullying or abuse. If an anomaly is detected, the server sends a report to the appropriate authorities.

[1525] As a specific example, if a post such as "I'm being bullied by a classmate" is detected on a social networking site, the device will analyze the information and the server will send a report to the appropriate authority.

[1526] Prompt Sentence Examples

[1527] Detect posts on social media that say things like "I'm being bullied at school."

[1528] If your analysis results indicate that "this child is experiencing severe stress," please send a report to the child consultation center.

[1529] Please send an encouraging reply to the consultation message "I don't want to go to school."

[1530] The purpose of this invention is to provide a safe environment by using this system to enable early detection of child abuse and bullying and prompt countermeasures.

[1531] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1532] Processing steps of AI emotion recognition system

[1533] Step 1: Initialize your device

[1534] The server initializes the camera and microphone connected to the emotion recognition device, which includes installing the device drivers and checking the connection.

[1535] Input: Device connection information

[1536] Data processing: Installing device drivers and checking connections

[1537] Output: Device initialization completion notification

[1538] What happens: The server logs the message "Initializing camera and microphone."

[1539] Step 2: Capture the data

[1540] The device uses the initialized camera and microphone to allow the user (parent or teacher) to capture the child's face and voice, and the captured data is temporarily stored on the device.

[1541] Input: Data captured by camera and microphone

[1542] Data processing: collection and temporary storage of face and voice data

[1543] Output: Notification that capture data has been saved

[1544] Specific behavior: The device will display a pop-up saying "Data capturing."

[1545] Step 3: Analyze the data

[1546] The device then inputs the captured facial and voice data into generative AI models for real-time analysis, including convolutional neural networks (CNNs) and recurrent neural networks (RNNs).

[1547] Input: Capture data

[1548] Data processing: Analyzing face and voice data with AI models

[1549] Output: Emotion analysis results

[1550] Specific behavior: The device will display the status "Analyzing data". Once the analysis is complete, the result will be displayed as "Negative emotion detected".

[1551] Step 4: Evaluate the analysis results and submit a report

[1552] The server receives the analysis results sent from the device and, if any abnormal values ​​are detected, sends a report to the appropriate child protection agency. This report includes the analysis data and the results.

[1553] Input: Sentiment analysis results

[1554] Data processing: Evaluation of analysis results and report generation

[1555] Output: Report sent to child protective services

[1556] Specific behavior: The server logs "Report sending" and displays "Report sent successfully" after the report has been sent.

[1557] AI surveillance system processing steps

[1558] Step 1: Initialize your device

[1559] The server initializes the surveillance cameras and audio devices installed in schools and homes.

[1560] Input: Monitoring device connection information

[1561] Data processing: Installing device drivers and checking connections

[1562] Output: Device initialization completion notification

[1563] Specific behavior: The server logs "Monitoring device initializing."

[1564] Step 2: Capture the data

[1565] The device uses surveillance cameras and audio devices to capture video and audio data in real time.

[1566] Input: Data captured by security cameras and audio devices

[1567] Data processing: collection and temporary storage of video and audio data

[1568] Output: Notification that capture data has been saved

[1569] Specific behavior: The camera light will turn on and the message "Capturing video and audio" will be displayed.

[1570] Step 3: Send the data

[1571] The device transmits the captured video and audio data to the server.

[1572] Input: Capture data

[1573] Data processing: Packetizing and sending data

[1574] Output: Notification of completion of data transmission to the server

[1575] Specific behavior: The status will be displayed as "Data sending."

[1576] Step 4: Analyze the data

[1577] The server analyzes the received data using a generative AI model to detect abnormal behavior or sounds.

[1578] Input: Submitted capture data

[1579] Data processing: Analyze video and audio data using AI models

[1580] Output: Analysis of abnormal behavior and audio

[1581] Specific behavior: The server will log "Analyzing data." After the analysis is complete, it will display "Abnormal behavior detected."

[1582] Step 5: Evaluate the analysis results and issue a warning

[1583] The server will alert administrators if any unusual activity or sound is detected. These alerts will be sent via email and SMS.

[1584] Input: Anomalous behavior and audio analysis results

[1585] Data processing: generating and sending warning messages

[1586] Output: Warning notice to administrator

[1587] Specific operation: The server records "Notifying administrator of warning" and displays "Warning sent successfully."

[1588] AI-powered helpline process

[1589] Step 1: Provide a chat system and phone line

[1590] The server provides a 24-hour chat system and telephone lines.

[1591] Input: System Operation Information

[1592] Data processing: Initialization of chat system and phone line

[1593] Output: System operation completion notification

[1594] Specific behavior: Display "Chat system and phone line are operational."

[1595] Step 2: User message input

[1596] The user inputs the content of the consultation into the chat system.

[1597] Input: Message from the user

[1598] Data processing: receiving and storing messages

[1599] Output: Message reception notification

[1600] What happens: The chat window will say "Please enter a message."

[1601] Step 3: Parse the message

[1602] The terminal analyzes messages sent by users using natural language processing technology.

[1603] Input: User's message

[1604] Data processing: Message content analysis

[1605] Output: Analysis results

[1606] Specific behavior: The status will be displayed as "Message parsing."

[1607] Step 4: Determine the appropriate response

[1608] Based on the results of message analysis, the server determines an encouraging message and contact information for a support center.

[1609] Input: Analysis results

[1610] Data processing: choosing the appropriate response

[1611] Output: Notification of support decision

[1612] What happens: The server logs "Determining appropriate action."

[1613] Step 5: Send a reply message

[1614] The server sends the determined reply message to the user.

[1615] Input: Support details

[1616] Data processing: Creating and sending a response message

[1617] Output: Reply notification to user

[1618] Specific operation: Displays "Sending reply message" and displays "Reply completed" after sending.

[1619] Processing steps for an AI-powered early warning system

[1620] Step 1: Collect data

[1621] The server collects text, image, and video data from social media and online platforms.

[1622] Input: Data from online platform

[1623] Data processing: data collection and storage

[1624] Output: Data collection completion notification

[1625] What happens: The server logs "Collecting data."

[1626] Step 2: Analyze the data

[1627] The device analyzes the collected data using a generative AI model to detect abnormal posts and signs of bullying or abuse.

[1628] Input: Collected data

[1629] Data processing: Data content analysis

[1630] Output: Analysis results

[1631] Specific operation: The device will display "Analyzing data". After the analysis is complete, it will display "Anomaly detected".

[1632] Step 3: Evaluate the analysis results

[1633] The server evaluates the analysis results sent from the terminal.

[1634] Input: Analysis results

[1635] Data processing: Evaluation of analysis results

[1636] Output: Anomaly detection evaluation results

[1637] Specific behavior: The server logs "Evaluating analysis results."

[1638] Step 4: Send the report

[1639] The server will send a report to the appropriate authorities if an abnormal value is detected.

[1640] Input: Anomaly detection evaluation results

[1641] Data Processing: Report creation and sending

[1642] Output: Report to child protective agencies

[1643] Specific operation: The server displays "Report sending" and then displays "Report sent successfully" after sending.

[1644] (Application example 1)

[1645] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1646] In modern society, problems of abuse and bullying are becoming more serious in the environment surrounding children. Early detection and rapid response to these problems are required, but conventional methods often lack real-time monitoring and emotion analysis, making effective responses impossible. The present invention aims to provide a system that solves these problems and ensures the safety of children.

[1647] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1648] In this invention, the server includes means for acquiring data on the face and voice of a target using an emotion recognition device, means for analyzing the target's emotions from the acquired data using an AI model, means for reporting to an appropriate organization if an abnormal emotion is detected based on the analyzed emotional information, means for detecting the target's behavior and voice in real time and issuing an alert if an abnormality is detected, and means for acquiring video and audio using a wearable device and notifying the user of the analysis results. This makes it possible to monitor children's emotions and behavior in real time, and if an abnormality is detected, to quickly notify the appropriate organization and take early action.

[1649] An "emotion recognition device" is a device that acquires facial and voice data of a subject and analyzes emotions from that data.

[1650] An "AI model" is an algorithm that uses artificial intelligence to analyze data and recognize specific patterns and emotions.

[1651] "Report to appropriate authorities" means reporting any detected abnormal emotional information or abnormal behavior to designated protective authorities or related agencies.

[1652] A "surveillance camera" is a device that captures, observes, and records images within a designated area in real time.

[1653] An "audio device" is an acoustic collection device that captures environmental sounds and conversation sounds and analyzes them as data.

[1654] A "wearable device" is a device that is worn by a user and has a built-in camera and microphone for capturing video and audio.

[1655] "Abnormal emotions" refer to negative emotions such as stress, fear, or anger that exceed the normal range.

[1656] "Abnormal behavior" refers to unnatural behavior that is clearly different from previous patterns or behavior that indicates a serious problem.

[1657] "24-hour chat or phone" refers to a means of communication designed to allow users to make inquiries or receive advice at any time.

[1658] "Natural language processing technology" refers to computer technology for understanding and analyzing human language.

[1659] This invention is a system that uses AI technology to detect child abuse and bullying in real time and promptly report it to the appropriate authorities. The system includes emotion recognition devices, surveillance cameras, audio devices, wearable devices (smart glasses), a 24-hour chat system, and data collection devices from social media and online platforms.

[1660] First, the server acquires the subject's facial and voice data using an emotion recognition device. This is done by capturing data using the camera and microphone installed in the smart glasses and sending it to the server in real time. The server then analyzes the data using an AI model to analyze the subject's emotional information. If abnormal emotional information (such as strong stress or fear) is detected, the server will promptly send a report to the appropriate authorities.

[1661] In addition, the server initializes surveillance cameras and audio devices installed in schools and homes, capturing video and audio data in real time. The captured data is sent to the server and analyzed by an AI model. If abnormal behavior or audio is detected, the server will alert administrators. This system makes it possible to detect signs of bullying and abuse early on.

[1662] The wearable device (smart glasses) is worn by the user (e.g., teacher or parent) on a daily basis and captures the child's face and voice in real time. The device has a built-in camera and microphone, and sends the data to a server that notifies the user of the analysis results. This allows the user to understand the child's emotions and condition on the spot.

[1663] The server also provides a chat system and telephone line that users can use 24 hours a day. When users input their concerns or questions, the device analyzes the message using natural language processing technology and determines the appropriate response. Based on the analysis results, the server provides encouraging messages and contact information for appropriate support centers.

[1664] As a concrete example, there is a series of steps: "Capture camera footage in the program's main loop," "Convert the captured facial area to grayscale and analyze it with an emotion recognition model," and "If a negative emotion is detected, send a warning message to a specified email address." This example would be input as a prompt to the generative AI model as follows:

[1665] Example prompt sentence:

[1666] "Create a program that monitors a child's emotions in real time. Capture camera footage, convert the facial area to grayscale, and analyze it with an emotion recognition model. If a negative emotion is detected, send a warning message to a specified email address."

[1667] In this way, the present invention provides a system that enables early detection of child abuse and bullying and prompt countermeasures.

[1668] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1669] Step 1:

[1670] The server initializes the camera and microphone of the smart glasses and acquires video and audio data from the smart glasses worn by the user. Specifically, the server captures real-time video (face images) and audio (conversation and environmental sounds) transmitted from the smart glasses. The input is video and audio data from the smart glasses, and the output is receiving these data in digital form.

[1671] Step 2:

[1672] The server processes the captured video data, crops out the facial area, and converts it to grayscale. Specifically, the server uses OpenCV to perform facial recognition, extract the facial area, and apply grayscale conversion. The input is the video data from the smart glasses, and the output is grayscale-converted facial image data.

[1673] Step 3:

[1674] The server inputs the grayscale-converted facial image data into an AI model to analyze emotions. Specifically, the server uses TensorFlow to pass the facial image data to an emotion recognition model, and obtains an emotion score (e.g., stress, fear, joy, etc.) as the result. The input is the grayscale-converted facial image data, and the output is the emotion score.

[1675] Step 4:

[1676] The server determines whether an abnormal emotion has been detected based on the acquired emotion score. Specifically, the server checks whether the emotion score exceeds a set threshold, and sets a flag if it is determined to be abnormal. The input is the emotion score, and the output is an anomaly detection flag (true / false).

[1677] Step 5:

[1678] If an abnormal emotion is detected, the server will send a warning message to the specified email address. Specifically, the server uses smtplib to construct an email and send a message stating that an abnormality has occurred. The input is the abnormality detection flag and the child's emotional information, and the output is a warning email.

[1679] Step 6:

[1680] The server stores the voice data acquired in real time and performs voice analysis. Specifically, the voice data is preprocessed using Librosa to extract specific voice features (e.g., voice tone and patterns). The input is the voice data from the smart glasses, and the output is voice feature data.

[1681] Step 7:

[1682] The server inputs the extracted voice features into an AI model and analyzes voice anomalies. Specifically, the voice feature data is passed to a voice recognition model, which determines whether the voice is abnormal (e.g., screaming or crying). The input is the voice feature data, and the output is the voice anomaly detection result.

[1683] Step 8:

[1684] If an abnormal sound is detected, the server issues a warning to the administrator. Specifically, a real-time notification is sent to the administrator's device, prompting them to check the situation on-site. The input is the result of the sound abnormality judgment, and the output is a warning notification to the administrator.

[1685] Step 9:

[1686] The server receives messages from users of a 24-hour chat system. Specifically, the server receives and stores text messages sent through the chat interface. The input is the message from the user, and the output is the stored message data.

[1687] Step 10:

[1688] The server analyzes the received message using natural language processing technology and determines its content. Specifically, it uses an NLP algorithm to classify the message content and determine the appropriate response. The input is the stored message data, and the output is the response content as the analysis result.

[1689] Step 11:

[1690] Based on the analysis results, the server provides appropriate support, such as encouraging messages and information on relevant support centers. The input is the response content based on the analysis results, and the output is a support message to the user.

[1691] Through these steps, the system can monitor children's emotions and behaviors in real time, respond quickly if an abnormality is detected, and provide support as needed to resolve the problem early.

[1692] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1693] This invention is a system that combines an emotion recognition device and an emotion engine to detect signs of child abuse and bullying at an early stage and report them to the appropriate authorities. This enables real-time emotion analysis and understanding of long-term emotional changes, allowing for more effective countermeasures.

[1694] Combining AI emotion recognition and emotion engine

[1695] The server initializes the camera and microphone devices, and the device captures the user's face and voice data in real time. This data is then analyzed by the emotion recognition device and emotion engine to detect whether the child is experiencing negative emotions such as fear or stress.

[1696] The emotion engine has the ability to cumulatively record a user's emotional data and analyze long-term changes in emotions. This makes it possible to grasp not only short-term emotions but also emotional patterns. For example, if a child is continuously feeling stressed, the emotion engine will recognize this as a pattern and issue an alert if necessary.

[1697] Collaboration between AI surveillance system and emotion engine

[1698] The server initializes surveillance cameras and audio devices installed in schools and homes, and the devices capture video and audio data in real time. The captured data is sent to the server and analyzed in real time by an AI model via an emotion engine. If abnormal behavior or audio is detected, the emotion engine records it and analyzes the pattern. If the abnormal pattern continues, the server issues an alert to the administrator.

[1699] AI-powered helpline and emotion engine integration

[1700] The server provides a chat system and telephone line that users can use 24 hours a day. When users input their worries or concerns, the device analyzes the message using natural language processing technology through an emotion engine. Based on the analysis results, the emotion engine can provide customized support content to the user from the server as needed. For example, if a user complains of stress over a long period of time, the emotion engine will understand this and can respond by providing them with a consultation service with a professional counselor.

[1701] Social Media and Online Platform Integration

[1702] The server collects data from social media and other online platforms, including text, images, and video data. The device then uses an emotion engine to analyze this data using natural language processing and image recognition technology to detect abnormal posts and signs of bullying or abuse. If an anomaly is detected, the emotion engine records it and sends a report to the appropriate authorities.

[1703] Specific examples

[1704] For example, suppose a user captures a child's face and voice at home. The device analyzes the data with an emotion engine and detects that the child is experiencing high levels of stress. The emotion engine records this and, if this pattern continues, the server sends a report to a child welfare center.

[1705] The same applies if a school's surveillance camera captures a student being bullied. The device analyzes this with its emotion engine and records it as an abnormal pattern. If the abnormality persists, the server immediately issues an alert to the school administrator.

[1706] In addition, a user can send a message of concern, such as "I don't want to go to school," through the helpline's chat system. The device analyzes the message using an emotion engine, which then accumulates and records it. When the accumulated stress exceeds a certain level, the server provides the user with an encouraging message and the contact information for a child consultation hotline.

[1707] As a result, the present invention makes it possible to detect child abuse and bullying at an early stage and take prompt measures. By utilizing the emotion engine, it is possible to focus on not only short-term emotion recognition but also long-term emotional changes, allowing for more appropriate responses.

[1708] The processing flow will be explained below.

[1709] Processing flow of AI-based emotion recognition and emotion engine combination

[1710] Step 1:

[1711] The server initializes the camera and microphone devices so that face and voice data can be captured.

[1712] Step 2:

[1713] The device captures the user's face and voice data in real time at regular intervals. The face data is obtained from the camera and the voice data is obtained from the microphone.

[1714] Step 3:

[1715] The device converts captured facial data into a grayscale image and preprocesses audio data based on the sample rate, making it easier for AI models to analyze.

[1716] Step 4:

[1717] The device inputs the preprocessed data into an emotion engine to analyze the user's emotions, which can identify emotions such as fear, stress, and joy.

[1718] Step 5:

[1719] The emotion engine cumulatively records the analyzed emotion information and analyzes short-term and long-term emotion patterns.

[1720] Step 6:

[1721] The server sends a report to the appropriate authorities if the emotion engine detects abnormal emotions (e.g., high levels of fear or stress).

[1722] Processing flow of collaboration between AI monitoring system and emotion engine

[1723] Step 1:

[1724] The server initializes the surveillance cameras and audio devices installed in schools and homes, enabling real-time monitoring.

[1725] Step 2:

[1726] The device captures video and audio data in real time and transmits it to the server.

[1727] Step 3:

[1728] The server receives the captured video and audio data and analyzes it in real time using an AI model via an emotion engine, which detects signs of bullying or abuse.

[1729] Step 4:

[1730] When the emotion engine detects abnormal behavior or voice, it records the analysis results cumulatively and analyzes consistent abnormal patterns.

[1731] Step 5:

[1732] If the server continues to exhibit abnormal patterns, it will alert the administrator.

[1733] AI-powered helpline and emotion engine integration process flow

[1734] Step 1:

[1735] The server provides a chat system and telephone line that users can use 24 hours a day, so that users can consult with the service at any time.

[1736] Step 2:

[1737] A user sends a message through the chat or phone system, and the message arrives at the server.

[1738] Step 3:

[1739] The server analyzes the messages it receives using natural language processing technology through an emotion engine and classifies the user's emotions and content.

[1740] Step 4:

[1741] The emotion engine cumulatively records the analysis results and determines customized support content.

[1742] Step 5:

[1743] Based on the analysis results, the server provides users with appropriate support, encouraging messages, contact information for specialist helplines, and more.

[1744] Social Media and Online Platform Integration Process Flow

[1745] Step 1:

[1746] The server uses APIs to collect data from social media and other online platforms, retrieving relevant text, image, and video data.

[1747] Step 2:

[1748] The device analyzes the collected data using an emotion engine, natural language processing, and image recognition technology to detect abnormal posts and signs of bullying or abuse.

[1749] Step 3:

[1750] The emotion engine cumulatively records the analysis results and analyzes the content of abnormal posts.

[1751] Step 4:

[1752] If the server detects an anomaly, it will send a report to the appropriate authorities, which will include details of the post where the anomaly occurred.

[1753] Step 5:

[1754] The server centrally manages the data and analysis results acquired in real time, records the details of detected anomalies, and notifies the administrator so that necessary measures can be taken.

[1755] The above are the specific processing steps for each system.

[1756] Example 2

[1757] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1758] Child abuse and bullying are serious social issues that require early detection and appropriate countermeasures. However, current systems have difficulty detecting changes in children's emotions in real time and identifying long-term trends. Furthermore, there is a lack of effective means to accurately detect signs of bullying and abuse and send reports to the appropriate authorities.

[1759] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1760] In this invention, the server includes: [means for initializing a device that acquires user face and voice data and capturing data in real time; [means for analyzing the user's emotions from the acquired data using a generative AI model and sending the data to an emotion engine; [means for accumulating and recording the analysis results and analyzing long-term changes in emotions; and [means for detecting abnormal emotional patterns from the accumulated emotional data and sending a report to an appropriate institution if the abnormality persists.] This makes it possible [to detect signs of child abuse and bullying early, analyze emotions in real time and understand long-term changes in emotions, and quickly take appropriate measures].

[1761] "User" refers to a person such as a child or student who uses the emotion recognition system, or their guardian or educator.

[1762] The "server" refers to a machine that serves as the central hub of the entire emotion recognition system and performs data collection, analysis, accumulating records, and anomaly detection processing.

[1763] "Terminal" refers to a device used by a user, equipped with a camera and microphone, that captures the user's face and voice data in real time.

[1764] An "emotion recognition device" refers to a device that analyzes a user's emotions from acquired facial and voice data.

[1765] An "emotion engine" refers to software or hardware that accumulates and records emotional data analyzed by an emotion recognition device and analyzes long-term changes in emotions.

[1766] "Generative AI model" refers to an artificial intelligence model used in emotion recognition devices to analyze a user's facial and voice data and identify emotions.

[1767] An "abnormal emotional pattern" refers to a state in which negative emotions such as anxiety, stress, or fear that exceed the normal range are continuously observed in the user's emotional data.

[1768] "Appropriate agencies" refers to specialized agencies needed to provide early intervention and support, such as child guidance centers, school administrators, and counselors.

[1769] "Chat System" means an online platform that provides users with messaging services available 24 hours a day.

[1770] "Surveillance cameras and audio devices" refers to devices installed in schools or homes that capture video and audio data.

[1771] "Natural language processing technology" refers to technology that enables computer programs to understand and analyze human language.

[1772] "Report" refers to a report that detects abnormal emotional patterns, signs of bullying or abuse, and notifies the appropriate authorities.

[1773] "Real-time" refers to data being processed and analyzed as soon as it is acquired.

[1774] This invention relates to a system that combines an emotion recognition device and an emotion engine to detect signs of child abuse and bullying at an early stage and report them to the appropriate authorities. This enables real-time emotion analysis and understanding of long-term emotional changes, allowing for more effective countermeasures to be taken.

[1775] System configuration

[1776] 1. Initialize the camera and microphone

[1777] The server initializes the camera and microphone devices, including loading the necessary device drivers. The device then activates the camera and microphone to capture the user's face and voice data in real time. This functionality can be used at home or in schools.

[1778] 2. Data capture and transmission

[1779] The device collects video data from the camera and audio data from the microphone and sends them to a server, where they are passed to an emotion recognition device in real time.

[1780] 3. Emotion analysis

[1781] The server analyzes the received data using an emotion recognition device. This analysis process uses a generative AI model, which identifies emotions from the user's facial expressions and tone of voice. At this point, negative emotions such as stress, fear, and sadness may be detected.

[1782] 4. Accumulative Recording of Emotional Data

[1783] The emotion engine records the analysis results cumulatively. The recorded data is used to analyze long-term changes in emotions. For example, by graphing the fluctuations in emotions over the past week, it can determine whether the user is experiencing persistent stress.

[1784] 5. Detecting abnormal patterns and generating alerts

[1785] The emotion engine has the ability to detect abnormal emotion patterns based on accumulated data. If the abnormal pattern persists, the emotion engine notifies the server, which then sends an alert to the appropriate authorities. This alert may be sent via email or SMS.

[1786] Specific examples

[1787] Examples of use within the home

[1788] Suppose a user captures their child's face and voice at home. The device sends this data to a server in real time and analyzes it with an emotion engine. The emotion engine detects when the child is experiencing high levels of stress and records this cumulatively. If this pattern continues, the server automatically sends a report to a child welfare center.

[1789] Examples of use within schools

[1790] At school, surveillance cameras capture footage of a specific student being bullied. Devices transmit this video data to a server in real time, where an emotion recognition device identifies the bullying. If the abnormal behavior continues, the server immediately issues an alert to school administrators.

[1791] Helpline usage example

[1792] A user sends a message such as "I don't want to go to school" through the helpline chat system. The device passes the message to the emotion engine, which analyzes it using natural language processing technology. The emotion engine detects continued stress and notifies the server. The server then provides the user with an encouraging message and contact information for a child consultation center.

[1793] Prompt Sentence Examples

[1794] Describe the implementation of a system that analyzes conversations in the home in real time to detect signs of stress.

[1795] This makes it possible to detect early signs of child abuse and bullying and quickly implement appropriate countermeasures. The use of an emotion engine allows for not only short-term emotion recognition but also long-term emotional changes, enabling more accurate responses.

[1796] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1797] Step 1:

[1798] The server initializes the camera and microphone devices. The initialized devices are ready to capture the user's face and voice data in real time. The input includes the server recognizing the camera and microphone devices and loading their drivers. The output is that the devices are successfully enabled and capture begins.

[1799] Specific operation:

[1800] The server detects the connected cameras and microphones, loads their device drivers, and captures test video and audio once to ensure the devices are initialized correctly.

[1801] Step 2:

[1802] The device uses the initialized camera and microphone to capture the user's face and voice data in real time. The input includes the video data captured by the camera and the audio data captured by the microphone. The output is sent to the server.

[1803] Specific operation:

[1804] The device captures video images via a camera installed in front of the user and ambient conversations and sounds via a microphone, and transmits the captured data as packets to a server at a fixed frame rate and sampling rate.

[1805] Step 3:

[1806] The server passes the acquired video and audio data to the emotion recognition device. The input includes the raw data sent from the device. The output is converted into a data format that the emotion recognition device can analyze.

[1807] Specific operation:

[1808] The server buffers the video and audio data received from the device and converts it into a format suitable for input to the emotion recognition device, including data compression and noise filtering.

[1809] Step 4:

[1810] The emotion recognizer uses a generative AI model to analyze video and audio data to identify the user's emotions. The input includes formatted data passed from the server. The output is a tag for the identified emotion, which is sent to the emotion engine.

[1811] Specific operation:

[1812] The emotion recognition device analyzes facial expressions from video data and vocal tone and tempo from audio data, and the generative AI model uses this information to tag emotions such as "stress," "joy," and "fear."

[1813] Step 5:

[1814] The emotion engine accumulates and records the analyzed emotion data. The input includes emotion tags sent from the emotion recognizer. The output is to add new emotion data to the accumulated emotion database.

[1815] Specific operation:

[1816] The emotion engine stores date- and time-stamped emotion data for each user in a cumulative database, allowing for tracking of emotional changes over time.

[1817] Step 6:

[1818] The emotion engine analyzes the accumulated data and detects anomalous emotion patterns. It includes the accumulated database as input and generates anomalous emotion pattern detection results as output.

[1819] Specific operation:

[1820] The emotion engine analyzes a user's emotional data over time to detect abnormal fluctuations and continuous negative emotions, such as when the "stress" tag is applied for more than a week.

[1821] Step 7:

[1822] When an anomalous emotional pattern is detected, the emotion engine notifies the server, which includes the anomaly detection results as input and sends an alert to the appropriate authorities as output.

[1823] Specific operation:

[1824] When the emotion engine detects an abnormal pattern, it sends an alert message to the server, which then sends an email or SMS to the child consultation center or school administrator based on the message.

[1825] The above is the flow of processing of the program of this system.

[1826] (Application example 2)

[1827] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1828] Conventional emotion recognition systems are limited to simply recognizing temporary emotions, making them unable to detect serious problems such as child abuse and bullying early on. They also lack the functionality to respond immediately when negative emotions persist, making it difficult to issue immediate warnings or notify appropriate authorities. Furthermore, they lack sufficient mechanisms for analyzing long-term emotional changes and identifying abnormal patterns, resulting in no fundamental solution. Therefore, a more comprehensive and effective emotion recognition system is needed.

[1829] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring data of the user's face and voice using an emotion recognition device, means for analyzing the user's emotions from the acquired data using an AI model, means for reporting to an appropriate agency if an abnormal emotion is detected based on the analyzed emotion information, means for recording cumulative emotion data and analyzing long-term emotion patterns, and means for sending notifications and alerts to the management device in real time if many negative emotions are detected. This makes it possible to monitor changes in children's emotions in real time and understand short-term and long-term emotion patterns, enabling early detection of bullying and abuse and rapid response.

[1830] An "emotion recognition device" is a device that acquires data on a user's face and voice and determines the user's emotions based on that data.

[1831] "Data acquisition" refers to the act of collecting information about a user's face and voice in real time using input devices such as a camera or microphone.

[1832] "AI Model" means a computational model that uses artificial intelligence technology and includes algorithms that analyze user data for emotion recognition.

[1833] "Emotional information" is data that indicates the user's emotional state, obtained as a result of analysis using an AI model.

[1834] "Reporting to appropriate authorities" refers to the act of reporting information to relevant authorities, such as child welfare centers or school administrators, when abnormal emotions are detected based on emotional information.

[1835] "Cumulative emotion data" is historical emotion data that is accumulated by recording the user's emotional changes over a long period of time.

[1836] "Long-term emotional pattern analysis" refers to the act of using accumulated emotional data to analyze trends in user emotional fluctuations and abnormal patterns.

[1837] "Sending notifications and alerts to management devices in real time" refers to the act of immediately sending warnings or notification messages to management devices such as smartphones and PCs when the emotion recognition device detects an abnormal emotion.

[1838] A "surveillance camera" is a camera device installed to acquire video data.

[1839] An "audio device" is a microphone or related device installed to capture audio data.

[1840] "Natural language processing technology" is an artificial intelligence technology that analyzes natural language data such as text and speech and understands its meaning and emotions.

[1841] This invention is a comprehensive emotion recognition system for early detection of signs of bullying and abuse and taking appropriate countermeasures. The system combines emotion recognition devices, AI models, long-term emotion pattern analysis, and real-time notification and alert functions.

[1842] Overall system configuration

[1843] 1. Hardware

[1844] Surveillance cameras: High-resolution cameras are installed in classrooms and around the school, such as generic high-resolution cameras (e.g., Hikvision DS-2CD2143G0-I).

[1845] Audio Device: A high-sensitivity microphone is installed in the designated location. For example, a generic high-sensitivity microphone (e.g., Blue Yeti USB Microphone).

[1846] Server: A high-performance server will be installed to perform data analysis and recording. For example, a generic high-performance server (e.g., Dell PowerEdge T340) will be installed.

[1847] Managed devices: Smartphones and PCs used by administrators. Smartphones run on iOS or Android, and PCs run on Windows or MacOS.

[1848] 2. Software

[1849] Sentiment Analysis: We use Python and TensorFlow to build a custom AI model that analyzes the user's facial and voice data in real time to recognize emotions.

[1850] Database: A MySQL database is used to manage the emotion data.

[1851] Notifications: Firebase Cloud Messaging and email notification systems are used to send real-time notifications to managed devices when an anomaly is detected.

[1852] Natural Language Processing: Incorporates natural language processing technology to analyze chats and messages.

[1853] Specific examples to realize

[1854] 1. Data Acquisition

[1855] The server collects video and audio data from surveillance cameras and audio devices within the school, recording the children's daily activities and comments.

[1856] 2. Data Analysis

[1857] The collected data is analyzed in real time by an emotion recognition system and AI models running on a server. For example, a custom AI model built with Python and TensorFlow analyzes facial expressions and tone of voice to detect negative emotions such as stress or fear.

[1858] 3. Notification of abnormalities

[1859] The collected and analyzed emotion data is cumulatively recorded in a MySQL database. If a large number of negative emotions are detected within a certain period of time, an alert is sent to the administrator's smartphone or PC using Firebase Cloud Messaging. For example, a message stating "Stress detected" is sent in real time.

[1860] 4. Long-term emotional pattern analysis

[1861] The server analyzes long-term emotional patterns using the accumulated emotional data, and if an abnormal pattern is detected, it sends a report to the appropriate organization, such as a child welfare center or parent.

[1862] 5. Natural Language Processing Support

[1863] Chat messages and phone consultations from users are analyzed using natural language processing technology. For example, if a message such as "I don't want to go to school" is sent, the content is analyzed by an emotion engine. If the accumulated stress exceeds a threshold, the system will connect the user to a professional counselor and provide customized support.

[1864] Examples and Prompts

[1865] An example of a prompt is as follows:

[1866] "From the video and audio of this child, you can determine if he is stressed or scared."

[1867] "Please conduct a long-term analysis to see if the students in this video are experiencing continuous stress."

[1868] This allows the emotion recognition system of the present invention to monitor changes in children's emotions in real time and respond quickly if an abnormality is detected.

[1869] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1870] Step 1:

[1871] The server initializes the surveillance cameras and audio devices and captures video and audio data in real time. The input is real-time data from the surveillance cameras and microphones, and the output is raw data stored on the server. This allows us to collect users' daily actions and comments.

[1872] Step 2:

[1873] The server sends the captured video and audio data to the emotion recognition device and AI model for data analysis. The input is the raw data acquired in step 1, and the output is the analyzed emotion information. Specifically, it uses Python and TensorFlow to analyze facial expressions and tone of voice to determine the user's emotion.

[1874] Step 3:

[1875] The server cumulatively records the analyzed emotional information in a MySQL database. The input is the emotional information obtained in step 2, and the output is the cumulative data stored in the database. This data is used to understand long-term emotional patterns.

[1876] Step 4:

[1877] The server analyzes the accumulated emotional data and sends real-time notifications and alerts to managed devices if abnormal emotional patterns are detected. The input is the accumulated data retrieved from the MySQL database, and the output is a notification sent to the administrator's smartphone or PC. For example, a message saying "Stress detected" is sent using Firebase Cloud Messaging.

[1878] Step 5:

[1879] When a user sends a message via 24-hour chat or phone, the server analyzes the message using natural language processing technology. The input is the user's message, and the output is the analyzed result. Specifically, natural language processing technology is used to analyze the content of the message and understand the user's emotional state.

[1880] Step 6:

[1881] The server records cumulative emotional data based on the analysis results and provides appropriate support. The input is the emotion analysis results obtained in step 5, and the output is the provision of support content. For example, if the accumulated stress exceeds a threshold, it will connect the user to a professional counselor or provide customized support.

[1882] Step 7:

[1883] The server uses the accumulated emotional data to further analyze long-term emotional patterns and, if necessary, sends reports to appropriate agencies. The input is long-term emotional data, and the output is a report to child consultation centers and parents. Specifically, the AI ​​model recognizes abnormal patterns, and creates and sends reports to appropriate agencies.

[1884] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1885] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1886] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1887] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1888] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1889] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1890] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1891] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1892] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1893] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1894] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1895] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1896] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1897] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1898] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1899] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1900] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1901] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1902] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1903] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1904] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1905] The following is further disclosed regarding the above embodiment.

[1906] (Claim 1)

[1907] means for acquiring face and voice data of a user by an emotion recognition device;

[1908] A means for analyzing user emotions from the acquired data using an AI model;

[1909] A means for reporting to appropriate authorities if abnormal emotions are detected based on the analyzed emotional information; and

[1910] A system including:

[1911] (Claim 2)

[1912] measures to install surveillance cameras and audio devices to detect signs of bullying or abuse;

[1913] A means to analyze captured camera and audio data in ...

Claims

1. means for acquiring face and voice data of a user by an emotion recognition device; A means for analyzing user emotions from the acquired data using an AI model; A means for reporting to appropriate authorities if abnormal emotions are detected based on the analyzed emotional information; and A system including:

2. Installing and implementing surveillance cameras and audio devices to detect signs of bullying and abuse; A means to analyze captured camera and audio data in real time using AI models; A means of detecting abnormal behavior or sounds from the analysis results and issuing an alert; The system of claim 1 , comprising:

3. A means to receive messages from users via chat or phone 24 hours a day; means for analyzing the received message using natural language processing technology; Based on the analysis results, a means of providing appropriate support; The system of claim 1 , comprising:

4. means of collecting text, image, and video data from social media and online platforms; A means for analyzing the collected data using natural language processing and image recognition technology; A means of detecting and reporting anomalous postings to the appropriate authorities; and The system of claim 1 , comprising:

5. A means of centrally managing data acquired in real time and analysis results; A means of recording details of detected anomalies and notifying administrators to take necessary measures; The system of claim 1 , comprising:

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A