System

A portable device with front and rear cameras and AI analysis provides real-time warnings to prevent accidents and robberies for vulnerable individuals.

JP2026034103APending Publication Date: 2026-02-27SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024137224
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-16
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

Existing systems fail to provide real-time risk detection and warning for vulnerable individuals such as elementary school children and the elderly, particularly against traffic accidents and robberies.

Method used

A portable device equipped with front and rear cameras captures images in real-time, transmitting them to a server for AI analysis to detect pedestrians and suspicious individuals, generating warnings through audio or vibration.

Benefits of technology

The system enables real-time detection and prevention of traffic accidents and robberies by providing timely warnings to users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026034103000001_ABST
    Figure 2026034103000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for acquiring videos of a front side and a rear side by using a mobile device including a front camera and a rear camera; means for transmitting the videos acquired from the mobile device to a server; means for analyzing the videos of the front side and the rear side by AI in the server and detecting running-out of pedestrians and suspicious persons around; means for generating a warning message based on the detection result; and means for transmitting the warning message to the mobile device and notifying a user of a warning.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] The purpose of this invention is to prevent traffic accidents and robberies that elementary school children and the elderly may encounter in their daily lives. Specifically, the purpose is to realize a system that uses a portable device equipped with front and rear cameras to detect accidents where a person suddenly jumps out into the street while walking or a snatch-and-grab attempt by a suspicious person from behind, and to issue an appropriate warning.

[0005] The rise in traffic accidents caused by elementary school children running out into the street and crimes such as snatching of elderly people calls for new measures to protect particularly vulnerable individuals. However, existing methods do not adequately provide real-time risk detection and warning, so it is necessary to provide an effective system to address these issues. [Means for solving the problem]

[0006] The present invention provides a means for capturing front and rear images using a portable device equipped with a front camera and a rear camera. The device captures images in real time and transmits them to a server at regular intervals. The server receives the transmitted image data and analyzes the images using an AI analysis module that utilizes deep learning.

[0007] The analysis process detects pedestrians and suspicious people in the vicinity, and generates a warning message based on the detection results. The generated warning message is then sent back to the mobile device to notify the user. Specifically, it issues an audio warning using a speech synthesis module or a physical warning using vibration.

[0008] This system allows users to detect risks in real time and take appropriate action, thereby preventing traffic accidents and robberies and improving the safety of elementary school children and the elderly.

[0009] A "front camera" is a camera installed in front of the mobile device to capture video in the direction in which a pedestrian is moving.

[0010] A "rear camera" is a camera installed at the rear of a mobile device to capture video behind a pedestrian.

[0011] A "portable device" is a small piece of equipment that can be carried by a user and that is equipped with a front camera and a rear camera.

[0012] "Means for acquiring video" refers to the functions and processes for capturing and recording video in real time using the front and rear cameras.

[0013] A "server" is a computer connected to a network that receives and analyzes video data sent from a portable device.

[0014] "Transmitting means" refers to the function and process of transferring video data from a portable device to a server using a certain communication protocol.

[0015] The "AI analysis module" is a combination of software and hardware that uses artificial intelligence to analyze video data and detect pedestrians and suspicious individuals.

[0016] "Pedestrians rushing out" refers to the act of pedestrians suddenly entering the road, ignoring designated crosswalks and safety zones.

[0017] A "suspicious person" is someone who exhibits unusual behavior toward people and objects around them and shows signs of criminal activity.

[0018] "Warning message" refers to the notification content that is generated based on detected risk information and is used to notify the user of danger.

[0019] "Means for notifying a warning" refers to the functions and processes for conveying the generated warning message to the user by voice synthesis or vibration.

[0020] A "deep learning model" refers to an algorithm that uses a multi-layer neural network to learn from large amounts of data and detect objects in video with high accuracy.

[0021] "Tracking" refers to the process of continuously tracking the movement of a recognized object and determining its position and speed in real time. [Brief explanation of the drawings]

[0022] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0023] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0024] First, the terms used in the following description will be explained.

[0025] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0026] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0027] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0028] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0029] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0030] [First embodiment]

[0031] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0032] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0033] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0034] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0035] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0036] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0037] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0038] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0039] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0040] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0041] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0042] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0043] MODE FOR CARRYING OUT THE INVENTION

[0044] This invention is a system to prevent traffic accidents and robberies that elementary school children and the elderly may encounter in their daily lives. Specifically, it uses a portable device equipped with a front and rear camera and utilizes AI technology to detect surrounding risks and issue a warning.

[0045] System configuration

[0046] This system consists of the following main components:

[0047] 1. Handheld devices

[0048] Front and rear cameras: These are necessary elements for capturing images in real time.

[0049] Communication module: Has the function to send captured images to a server.

[0050] Warning notification function: This module notifies the user of warning messages received from the server. Specifically, it has audio output and vibration functions.

[0051] 2. Server

[0052] Video receiving module: Has the function to receive video data transmitted from a portable device.

[0053] AI analysis module: Analyzes the received video data and executes algorithms to detect sudden jumps in and suspicious individuals.

[0054] Alert generation module: Based on the detection results, it generates an alert message and sends it to the mobile device.

[0055] System Operation

[0056] The operation of this system will be explained below.

[0057] 1. Video acquisition and transmission

[0058] Terminal: The mobile device captures video in real time using front and rear cameras and transmits the data to the server at regular intervals.

[0059] 2. Video Analysis

[0060] Server: The server receives the video data sent from the mobile device and inputs it into the AI ​​analysis module, which uses deep learning to detect objects in the video (e.g., vehicles, pedestrians, suspicious individuals).

[0061] 3. Risk detection and prediction

[0062] Server: Using AI analysis, it assesses the risk of pedestrians running out into the road and the approach of suspicious individuals. For example, it detects when a pedestrian is about to run out into the road or when a suspicious individual is approaching an elderly person.

[0063] 4. Generating and sending warning messages

[0064] Server: Automatically generates warning messages based on each risk and sends them to the mobile device.

[0065] 5. Warning Notification

[0066] Terminal: The terminal notifies the user of warning messages received from the server by voice or vibration. Specifically, it uses voice synthesis to issue warnings such as "Danger! Watch out for people jumping out into the street!" or "There's a suspicious person behind you, so be careful."

[0067] Specific examples

[0068] Detecting elementary school students jumping out

[0069] Device: As elementary school students walk to school, the front-facing camera captures footage of the road.

[0070] Terminal: Video data is sent to the server.

[0071] Server: The video received by the server is analyzed by AI, and any movement of the elementary school student trying to jump out is detected.

[0072] Server: Generates a "Watch out for people jumping out!" warning message and sends it to the handheld device.

[0073] Device: Elementary school students will hear a voice warning saying, "Danger! Watch out for people jumping out!"

[0074] User: Elementary school student listens to the warning and stops to prevent a traffic accident.

[0075] Detecting suspicious elderly people

[0076] Device: The rear camera captures rearward footage as the elderly person walks home.

[0077] Terminal: The acquired video data is sent to the server.

[0078] Server: The server analyzes the video and detects suspicious individuals approaching the elderly.

[0079] Server: Generates a "Beware of suspicious person!" warning message and sends it to the mobile device.

[0080] Terminal: A voice warning is issued to the elderly person saying, "There is a suspicious person behind you, so please be careful."

[0081] Users: Elderly people receive a warning and pay attention to their surroundings to prevent harm from suspicious people.

[0082] As described above, the present invention improves the safety of elementary school children and the elderly in their daily lives through a system that combines a portable device and a server. The real-time risk detection and warning functions provided by this system prevent traffic accidents and robberies.

[0083] The processing flow will be explained below.

[0084] Program processing flow

[0085] Video acquisition and transmission

[0086] Device:

[0087] 1. Step 1: Activate the front and rear cameras on your mobile device.

[0088] 2. Step 2: Capture video in real time using the front and rear cameras. Video frames are acquired at a rate of, for example, 30 frames per second.

[0089] 3. Step 3: Temporarily save the captured video data in memory.

[0090] 4. Step 4: At regular intervals (for example, every second), the video data stored in memory is divided into packets and organized.

[0091] 5. Step 5: The organized data packets are sent to the server via the communication module using HTTP or WebSocket.

[0092] Video analysis

[0093] server:

[0094] 1. Step 1: The server receives the front and rear camera image data.

[0095] 2. Step 2: The received video data is decoded and input into the AI ​​analysis module.

[0096] 3. Step 3: The AI ​​analysis module uses a deep learning model (e.g., YOLO, SSD) to detect objects (e.g., vehicles, pedestrians, suspicious people) in the video frame.

[0097] 4. Step 4: Track the movement of the detected object and analyze its position and velocity.

[0098] 5. Step 5: Evaluate the risk of pedestrians running out into the street or the approach of suspicious people. Specifically, predict whether certain movements pose a risk of traffic accidents or robberies.

[0099] Risk detection and warning

[0100] server:

[0101] 1. Step 1: Determine whether there is a risk based on the results of AI analysis. For example, evaluate whether a pedestrian is about to run out into the road or whether a suspicious person is approaching.

[0102] 2. Step 2: If a risk is detected, generate an appropriate warning message, which includes the warning type (e.g., watch out for a sudden jump, watch out for a suspicious person) and a recommended action (e.g., stop, turn around).

[0103] 3. Step 3: Send the generated warning message to the terminal.

[0104] Warning Notification

[0105] Device:

[0106] 1. Step 1: Decode the warning message received from the server.

[0107] 2. Step 2: The decoded warning message is notified to the user. Specifically, the voice synthesis module generates a warning voice and outputs it from the speaker or activates the vibration motor.

[0108] User Feedback

[0109] User:

[0110] 1. Step 1: Receive an alert notification from your device, for example, hear a sound notification or feel a vibration.

[0111] 2. Step 2: Follow the warning and take appropriate action. Specifically, if you receive a warning about people jumping out into the street, stop, and if you receive a warning about suspicious people, turn around and check your surroundings.

[0112] Specific examples

[0113] Detecting elementary school students jumping out

[0114] Device:

[0115] 1. Step 1: When an elementary school student is walking to school, the front camera captures images of the road.

[0116] 2. Step 2: The video data is sent to the server.

[0117] server:

[0118] 1. Step 1: The video received by the server is analyzed by AI, and any movement of the elementary school student trying to jump out is detected.

[0119] 2. Step 2: Generate a warning message "Watch out for people jumping out!" and send it to the mobile device.

[0120] Device:

[0121] 1. Step 1: Elementary school students will hear a voice warning saying, "Danger! Watch out for people jumping out!"

[0122] User:

[0123] 1. Step 1: Elementary school students listen to the warning and stop to prevent traffic accidents.

[0124] Detecting suspicious elderly people

[0125] Device:

[0126] 1. Step 1: When an elderly person is returning home, the rear camera captures the rear view.

[0127] 2. Step 2: The acquired video data is sent to the server.

[0128] server:

[0129] 1. Step 1: The server analyzes the video and detects suspicious individuals approaching the elderly person.

[0130] 2. Step 2: Generate a warning message saying "Beware of suspicious person!" and send it to the mobile device.

[0131] Device:

[0132] 1. Step 1: The elderly person receives a voice warning saying, "There is a suspicious person behind you, so please be careful."

[0133] User:

[0134] 1. Step 1: The elderly person receives a warning and pays attention to their surroundings to prevent harm from suspicious individuals.

[0135] Example 1

[0136] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0137] This invention relates to a system that prevents traffic accidents and robberies that elementary school students and the elderly may encounter in their daily lives. In particular, it aims to provide a system that can detect surrounding risks in real time and issue accurate warnings. Conventional systems sometimes experience delays in detecting risks and generating warning messages, making them insufficient to avoid actual danger. Therefore, the present invention aims to prevent traffic accidents and robberies by quickly and accurately detecting risks and issuing warnings to users.

[0138] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0139] In this invention, the server includes means for receiving the acquired video data with a video receiving module, analyzing the front and rear video using an AI analysis module, and detecting pedestrians running out into the road and suspicious people in the vicinity, means for generating a warning message with a warning generation module based on the detection results, and means for transmitting the warning message to the mobile terminal device via a communication module and notifying the user of the warning using a warning notification function. This makes it possible to detect risks quickly and accurately and notify the user of appropriate warnings in real time.

[0140] A "portable terminal device" is a device that has a front camera and a rear camera and can be carried by a user.

[0141] A "communication module" is a combination of hardware and software for transmitting data from a mobile terminal device to a server.

[0142] A "server" is a computing system that receives, analyzes, and processes data sent from a mobile terminal device.

[0143] The "video receiving module" is a module having a function for receiving video data transmitted from a portable terminal device at a server.

[0144] The "AI analysis module" is a module that uses artificial intelligence technology to analyze received video data and detect specific patterns and risks.

[0145] The "warning generation module" is a module for generating appropriate warning messages based on the detection results of the AI ​​analysis module.

[0146] The "warning notification function" is a function in which the portable terminal device notifies the user of a warning message by voice or vibration.

[0147] "Pedestrian running out" refers to a pedestrian unintentionally running out into a dangerous place such as the roadway.

[0148] A "suspicious individual" is an individual who deviates from normal patterns of behavior and is perceived as a threat to users.

[0149] A "deep learning model" is an artificial intelligence technology that learns complex patterns from large amounts of data and uses them for analysis.

[0150] This invention is a system for preventing traffic accidents and robberies in the daily lives of elementary school children and the elderly. This system is constructed by combining a portable terminal device equipped with a front camera and a rear camera with a server. The configuration and operation of this system are described in detail below.

[0151] System configuration

[0152] 1. Portable terminal device

[0153] Front and rear cameras:

[0154] The mobile terminal device is equipped with a camera that captures images of the front and rear in real time, allowing the user to monitor the situation around them in detail.

[0155] Communication Module:

[0156] The mobile terminal device includes a communication module for transmitting the acquired video data to a server. Specifically, a 4G LTE module is used.

[0157] Warning notification function:

[0158] The terminal is equipped with a voice output and vibration function as a warning notification function, which allows the terminal to convey the warning message received from the server to the user.

[0159] 2. Server

[0160] Video Receiving Module:

[0161] The server uses a video receiving module to receive video data sent from the mobile terminal device, specifically, Apache (registered trademark) Kafka.

[0162] AI analysis module:

[0163] The server is equipped with an AI analysis module that analyzes the received video data and uses a deep learning model (e.g., TENSORFLOW®) to detect objects in the video (e.g., vehicles, pedestrians, suspicious individuals).

[0164] Warning generation modules:

[0165] The server includes a warning generation module for generating a warning message based on the detection result by the AI ​​analysis module, and the generated warning message is transmitted to the mobile terminal device again.

[0166] System Operation

[0167] Video acquisition and transmission

[0168] Device: Front and rear cameras capture images of the user's surroundings in real time. The image data is sent to a server at regular intervals.

[0169] Receiving and analyzing video data

[0170] Server: The video receiving module receives the video data and inputs it into the AI ​​analysis module. Analysis detects pedestrians running out into the street and suspicious people approaching.

[0171] Evaluating risks and generating warning messages

[0172] Server: Based on the results from the AI ​​analysis module, a risk assessment is performed and an appropriate warning message is generated in the warning generation module. For example, messages such as "Watch out for people jumping out!" or "Watch out for suspicious people!" are generated.

[0173] Warning message notification

[0174] Terminal: A mobile terminal device receives the warning message and notifies the user by voice or vibration. A warning message is generated using voice synthesis technology (e.g., Google® Cloud Text-to-Speech) and output from the speaker.

[0175] Specific examples

[0176] Detecting elementary school students jumping out

[0177] Device: When an elementary school student walks to school, the front camera captures footage of the road.

[0178] Terminal: Video data is sent to the server.

[0179] Server: The video received by the server is analyzed by AI, and any movement of the elementary school student trying to jump out is detected.

[0180] Server: Generates a warning message saying "Watch out for people jumping out!" and sends it to the mobile terminal device.

[0181] Device: Elementary school students will hear a voice warning saying, "Danger! Watch out for people jumping out!"

[0182] User: Elementary school student listens to the warning and stops to prevent a traffic accident.

[0183] Detecting suspicious elderly people

[0184] Device: The rear camera captures rearward footage as the elderly person walks home.

[0185] Terminal: The acquired video data is sent to the server.

[0186] Server: The server analyzes the video and detects suspicious individuals approaching the elderly.

[0187] Server: Generates a warning message saying "Beware of suspicious person!" and sends it to the mobile terminal device.

[0188] Terminal: A voice warning is issued to the elderly person saying, "There is a suspicious person behind you, so please be careful."

[0189] Users: Elderly people receive a warning and pay attention to their surroundings to prevent harm from suspicious people.

[0190] Prompt Sentence Examples

[0191] "While elementary school students are walking to school, please capture road footage with a front camera, use AI to detect the risk of them running out into the street, and use a notification system to generate a warning message. Please also tell us details about the hardware and software required."

[0192] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0193] Step 1: Acquire footage

[0194] Terminal: The mobile terminal device captures video in real time using the front and rear cameras, capturing video at 30 frames per second and storing it in an internal buffer.

[0195] Input: Real-time video of the user's surroundings.

[0196] Output: Front and rear video data stored in internal buffers.

[0197] Step 2: Sending the video

[0198] Terminal: At regular intervals (e.g., every 5 seconds), the video data stored in the buffer is sent to the server via the communications module. Specifically, the data is encoded via the 4G LTE module and transferred to the server via TCP / IP protocol.

[0199] Input: Video data stored in the internal buffer.

[0200] Output: Video data sent to the server.

[0201] Step 3: Receiving video data

[0202] Server: The server-side video receiving module receives the video data sent from the mobile terminal device and stores it in a database. Specifically, it uses Apache Kafka to manage the data stream and store it appropriately.

[0203] Input: Video data transmitted from a mobile terminal device.

[0204] Output: Video data stored in a database.

[0205] Step 4: Analyzing the footage

[0206] Server: The received video data is input into the AI ​​analysis module and analyzed using a deep learning model (e.g., TensorFlow). Specifically, objects in the video (vehicles, pedestrians, suspicious individuals, etc.) are detected and their movements are tracked.

[0207] Input: Video data stored in a database.

[0208] Output: Analysis results (e.g., JSON format) containing information about detected objects.

[0209] Step 5: Risk detection and prediction

[0210] Server: Based on the results of the AI ​​analysis module, it evaluates the risk of pedestrians running out into the street and the approach of suspicious individuals. Specifically, it calculates the speed and distance of objects and executes an algorithm to assess the risk level.

[0211] Input: Analysis results of the AI ​​analysis module.

[0212] Output: Evaluated risk information (e.g., risk of people jumping out, suspicious person warning).

[0213] Step 6: Generate a warning message

[0214] Server: Based on the detected risks, the warning generation module generates warning messages. Specifically, a Python script is used to generate warning messages (e.g., "Watch out for people jumping out!", "Watch out for suspicious people!") as text and audio data.

[0215] Input: Assessed risk information.

[0216] Output: Warning message in text and audio format.

[0217] Step 7: Sending a warning message

[0218] Server: The server transmits the generated warning message to the mobile terminal device. Specifically, the server encodes the warning message via the communication module and transmits it.

[0219] Input: Warning message in text and audio format.

[0220] Output: The warning message sent to the handheld device.

[0221] Step 8: Notification of warnings

[0222] Terminal: The mobile terminal device notifies the user of the received warning message. Specifically, it generates the warning message using voice synthesis technology and outputs it from the speaker. It also issues a physical warning using the vibration function.

[0223] Input: The warning message received.

[0224] Output: Audio and vibration alert notifications.

[0225] (Application example 1)

[0226] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0227] Conventional safety monitoring systems in factories mainly use fixed cameras, making it difficult to effectively monitor the movements of dynamically moving robots and people in real time. Furthermore, advanced analysis technology is required to detect abnormal or suspicious behavior, and it is not possible to issue appropriate warnings immediately. This results in low accuracy in improving safety in factories, and there is a high possibility of delayed response in emergencies.

[0228] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0229] In this invention, the server includes a means for acquiring video from the front and rear, a means for analyzing the acquired video using AI to detect abnormal or suspicious behavior, and a means for generating a warning message based on the detection results and transmitting it to the portable device or robot. This makes it possible to detect abnormal or suspicious behavior with high accuracy in real time and to immediately issue warnings to robots and workers in the factory by voice and vibration.

[0230] "Front camera and rear camera" are imaging devices for capturing images in front of and behind the device.

[0231] A "portable device" is a device that can be carried around, has front and rear cameras, and has the ability to capture and transmit images.

[0232] A "server" is a computer system that receives and analyzes video sent from a portable device.

[0233] "AI analysis" is the process of using artificial intelligence technology to analyze video data and detect specific patterns or anomalies.

[0234] A "warning message" is a message that is generated to warn the user when an abnormality or risk is detected.

[0235] A "factory robot" is a mechanical device that operates autonomously or semi-autonomously in a factory to perform designated tasks.

[0236] An "audio output device" is a device for producing electronically generated speech.

[0237] The "vibration function" is a function that causes a device or robot to vibrate to provide a physical notification to the user.

[0238] "Real-time" refers to the fact that information acquisition, processing, and output of results are carried out almost simultaneously.

[0239] "Abnormal and suspicious behavior" refers to behavior that deviates from normal patterns of behavior or that appears suspicious.

[0240] A "deep learning model" is an artificial intelligence model that uses a multi-layer neural network to learn the characteristics of data and perform advanced predictions and classifications.

[0241] "Tracking" is a technology that continuously tracks a specific object or movement and evaluates its position and status.

[0242] The present invention provides a system for monitoring the safety of robots and workers in a factory, detecting abnormal or suspicious behavior in real time, and issuing a warning. A specific embodiment of this system will be described below.

[0243] Main components of the system

[0244] 1. Handheld devices

[0245] It is equipped with a front-facing camera and a rear-facing camera, which capture images in real time. It also has the function of transmitting the images to a server via a communication module and receiving warning messages.

[0246] 2. Server

[0247] It receives video data sent from a portable device, performs AI analysis, and generates a warning message based on the analysis results, which is then sent back to the portable device.

[0248] 3. Robots in factories

[0249] Equipped with front and rear cameras, it captures video footage to detect abnormal or suspicious activity, and receives warning messages and issues audio and vibration alerts.

[0250] Hardware and Software Used

[0251] Cameras: Front and rear facing cameras

[0252] Communication module: Wi-Fi or Bluetooth

[0253] Server: High-performance server equipped with GPU

[0254] AI analysis software: TensorFlow, PyTorch

[0255] Video capture software: OpenCV

[0256] Warning notification software: paho-mqtt (communication), pyttsx3 (voice synthesis)

[0257] System Operation

[0258] The server receives video data sent from the mobile devices and factory robots. It analyzes this data using a deep learning model to detect abnormal or suspicious behavior in real time. Based on the results, it generates a warning message and sends it back to the mobile devices and robots. The robots then immediately notify them of the received warning message by voice and vibration.

[0259] Specific examples

[0260] Abnormal behavior detection

[0261] 1. Terminal (handheld device): Robots in the factory capture front and rear images in real time and send them to the server.

[0262] 2. Server: Analyzes the received video using a deep learning model to detect abnormal behavior, such as when a robot deviates from its normal behavior pattern.

[0263] 3. Server: Generates a warning message saying "Danger! Caution! Watch the machine!" and sends it to the robots in the factory.

[0264] 4. Terminal (robot in factory): Notifies warnings by voice and vibration.

[0265] Suspicious Activity Detection

[0266] 1. Terminal (portable device): A robot moving around the factory captures rearward-facing images and sends them to a server.

[0267] 2. Server: AI analyzes the received video and detects suspicious movements, such as when a suspicious person approaches the robot.

[0268] 3. Server: Generates a warning message saying "Suspicious behavior detected!" and sends it to the robots in the factory.

[0269] 4. Terminal (robot in factory): Notifies warnings by voice and vibration.

[0270] Prompt Sentence Examples

[0271] Factory robots capture images from cameras in front and behind them in real time and use AI to monitor the safety of the work area. For example, if a person or machine behaves abnormally, it will automatically issue a voice warning saying, "Danger! Watch out for machine operation!" What will happen if an abnormality is detected?

[0272] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0273] Step 1:

[0274] The terminal (a robot in a factory) acquires real-time images from the front and rear. It uses a front camera and a rear camera to capture the images. The input obtained is raw image data.

[0275] Step 2:

[0276] The video data captured by the terminal (a robot in a factory) is sent to the server via a communication module. Using a communication protocol (e.g., MQTT or HTTP), the video data is divided into packets and sent to the server. The server then receives the video data as input.

[0277] Step 3:

[0278] The video data received by the server is input into an AI analysis module, which uses a deep learning model (e.g., TensorFlow or PyTorch) to recognize objects in the video and analyze their behavior. This outputs intermediate data (e.g., the position of each object and its movement vector) that can be used to identify abnormal or suspicious behavior.

[0279] Step 4:

[0280] The server performs risk assessment based on the results of AI analysis. It evaluates identified abnormal or suspicious behavior and runs an algorithm to calculate the risk level. This generates a warning message for any detected risks.

[0281] Step 5:

[0282] The server generates a warning message and sends it back to the terminal (the robot in the factory) via the communication module, which then receives the warning message.

[0283] Step 6:

[0284] Based on the warning message received by the terminal (a robot in the factory), a warning is issued by voice and vibration. The voice output device is used to issue a voice warning, and the vibration module is used to notify the warning by vibration. The final output is a real-time warning notification to the user.

[0285] Step 7:

[0286] The user (worker) receives a warning from the robot and responds immediately. The worker receives a warning via voice and vibration and responds to the surrounding risks.

[0287] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0288] MODE FOR CARRYING OUT THE INVENTION

[0289] This invention is a system to prevent traffic accidents and robberies that elementary school children and the elderly may encounter in their daily lives. Specifically, it uses a portable device equipped with a front and rear camera, and utilizes AI technology and an emotion engine to detect surrounding risks and issue warnings according to the user's emotional state.

[0290] System configuration

[0291] This system consists of the following main components:

[0292] 1. Handheld devices

[0293] Front and rear cameras: These are necessary elements for capturing images in real time.

[0294] Communication module: Has the function to send captured images to a server.

[0295] Emotion engine: Has the function to analyze the user's emotional state in real time.

[0296] Warning notification function: This module notifies the user of warning messages received from the server. Specifically, it has audio output and vibration functions.

[0297] 2. Server

[0298] Video receiving module: Has the function to receive video data transmitted from a portable device.

[0299] AI analysis module: Analyzes the received video data and executes algorithms to detect sudden jumps in and suspicious individuals.

[0300] Alert generation module: Based on the detection results, it generates an alert message and sends it to the mobile device.

[0301] System Operation

[0302] The operation of this system will be explained below.

[0303] 1. Video acquisition and transmission

[0304] Terminal: The mobile device captures video in real time using front and rear cameras and transmits the data to the server at regular intervals.

[0305] 2. Video Analysis

[0306] Server: The server receives the video data sent from the mobile device and inputs it into the AI ​​analysis module, which uses deep learning to detect objects in the video (e.g., vehicles, pedestrians, suspicious individuals).

[0307] 3. Analysis of the user's emotional state

[0308] Terminal: The emotion engine installed in the portable device uses a camera and audio sensors to analyze the user's facial expressions and tone of voice, and determine their emotional state in real time.

[0309] 4. Risk detection and warning

[0310] Server: Using AI analysis, it assesses the risk of pedestrians running out into the road and the approach of suspicious individuals. For example, it detects when a pedestrian is about to run out into the road or when a suspicious individual is approaching an elderly person.

[0311] Server: When a risk is detected, it generates an appropriate warning message according to the user's emotional state and sends it to the mobile device. Depending on the emotional state, it adjusts the content of the warning and the notification method.

[0312] 5. Warning Notification

[0313] Terminal: The terminal notifies the user of warning messages received from the server by voice or vibration. Specifically, it uses voice synthesis to issue warnings such as "Danger! Watch out for people jumping out into the street!" or "There's a suspicious person behind you, so be careful."

[0314] Specific examples

[0315] Detecting elementary school students jumping out

[0316] Device: As elementary school students walk to school, the front-facing camera captures footage of the road.

[0317] Terminal: Video data is sent to the server.

[0318] Server: The video received by the server is analyzed by AI, and any movement of the elementary school student trying to jump out is detected.

[0319] Device: The emotion engine analyzes the emotional state of elementary school students and generates stronger warning messages if they are not concentrating.

[0320] Server: Generates a "Watch out for people jumping out!" warning message and sends it to the handheld device.

[0321] Device: Elementary school students will hear a voice warning saying, "Danger! Watch out for people jumping out!"

[0322] User: Elementary school student listens to the warning and stops to prevent a traffic accident.

[0323] Detecting suspicious elderly people

[0324] Device: The rear camera captures rearward footage as the elderly person walks home.

[0325] Terminal: The acquired video data is sent to the server.

[0326] Server: The server analyzes the video and detects suspicious individuals approaching the elderly.

[0327] Device: Detects the emotional state of the elderly person and generates a soft-toned warning message if they are in a state of tension.

[0328] Server: Generates a "Beware of suspicious person!" warning message and sends it to the mobile device.

[0329] Terminal: A voice warning is issued to the elderly person saying, "There is a suspicious person behind you, so please be careful."

[0330] Users: Elderly people receive a warning and pay attention to their surroundings to prevent harm from suspicious people.

[0331] As described above, this invention improves the safety of elementary school children and the elderly in their daily lives through a system that combines a portable device and a server. The real-time risk detection and warning functions provided by this system prevent traffic accidents and robberies. Furthermore, by combining it with an emotion engine, it is possible to issue warnings according to the user's emotional state, further improving safety.

[0332] The processing flow will be explained below.

[0333] Program processing flow

[0334] Video acquisition and transmission

[0335] Device:

[0336] 1. Step 1: Activate the front and rear cameras on your mobile device.

[0337] 2. Step 2: Capture video in real time using the front and rear cameras. Video frames are acquired at a rate of, for example, 30 frames per second.

[0338] 3. Step 3: Temporarily save the captured video data in memory.

[0339] 4. Step 4: At regular intervals (for example, every second), the video data stored in memory is divided into packets and organized.

[0340] 5. Step 5: The organized data packets are sent to the server via the communication module using HTTP or WebSocket.

[0341] Video analysis

[0342] server:

[0343] 1. Step 1: The server receives the front and rear camera image data.

[0344] 2. Step 2: The received video data is decoded and input into the AI ​​analysis module.

[0345] 3. Step 3: The AI ​​analysis module uses a deep learning model (e.g., YOLO, SSD) to detect objects (e.g., vehicles, pedestrians, suspicious people) in the video frame.

[0346] 4. Step 4: Track the movement of the detected object and analyze its position and velocity.

[0347] 5. Step 5: Evaluate the risk of pedestrians running out into the street or the approach of suspicious people. Specifically, predict whether certain movements pose a risk of traffic accidents or robberies.

[0348] Analyzing the user's emotional state

[0349] Device:

[0350] 1. Step 1: Start the emotion engine on the mobile device.

[0351] 2. Step 2: Use a camera or audio sensor to capture the user's facial expressions and tone of voice.

[0352] 3. Step 3: Analyze the captured data in real time to determine the user's emotional state, for example, identifying states such as tension, anxiety, or concentration.

[0353] 4. Step 4: Send the analysis results to the server.

[0354] Risk detection and warning

[0355] server:

[0356] 1. Step 1: Determine whether there is a risk based on the results of AI analysis. For example, evaluate whether a pedestrian is about to run out into the road or whether a suspicious person is approaching.

[0357] 2. Step 2: If a risk is detected, a warning message is generated based on the analysis of the transmitted emotional state.

[0358] 3. Step 3: Adjust the content and notification method of the warning message according to the user's emotional state.

[0359] 4. Step 4: Send the generated warning message to the terminal.

[0360] Warning Notification

[0361] Device:

[0362] 1. Step 1: Decode the warning message received from the server.

[0363] 2. Step 2: The decoded warning message is notified to the user. Specifically, the voice synthesis module generates a warning voice and outputs it from the speaker or activates the vibration motor.

[0364] User Feedback

[0365] User:

[0366] 1. Step 1: Receive an alert notification from your device, for example, hear a sound notification or feel a vibration.

[0367] 2. Step 2: Follow the warning and take appropriate action. Specifically, if you receive a warning about people jumping out into the street, stop, and if you receive a warning about suspicious people, turn around and check your surroundings.

[0368] Specific examples

[0369] Detecting elementary school students jumping out

[0370] Device:

[0371] 1. Step 1: When an elementary school student is walking to school, the front camera captures images of the road.

[0372] 2. Step 2: The video data is sent to the server.

[0373] server:

[0374] 1. Step 1: The video received by the server is analyzed by AI, and any movement of the elementary school student trying to jump out is detected.

[0375] 2. Step 2: Receive data from the emotion engine to analyze the emotional state of elementary school students.

[0376] 3. Step 3: If the elementary school student is not concentrating, generate a stronger warning message.

[0377] 4. Step 4: Generate a warning message "Watch out for people jumping out!" and send it to the mobile device.

[0378] Device:

[0379] 1. Step 1: Elementary school students will hear a voice warning saying, "Danger! Watch out for people jumping out!"

[0380] User:

[0381] 1. Step 1: Elementary school students listen to the warning and stop to prevent traffic accidents.

[0382] Detecting suspicious elderly people

[0383] Device:

[0384] 1. Step 1: When an elderly person is returning home, the rear camera captures the rear view.

[0385] 2. Step 2: The acquired video data is sent to the server.

[0386] 3. Step 3: The emotion engine captures the emotional state of the elderly person and sends the analysis results to the server.

[0387] server:

[0388] 1. Step 1: The server analyzes the video and detects suspicious individuals approaching the elderly person.

[0389] 2. Step 2: Evaluate the elderly person's emotional state and generate a soft-toned warning message if they are in a tense state.

[0390] 3. Step 3: Generate a warning message saying "Beware of suspicious person!" and send it to the mobile device.

[0391] Device:

[0392] 1. Step 1: The elderly person receives a voice warning saying, "There is a suspicious person behind you, so please be careful."

[0393] User:

[0394] 1. Step 1: The elderly person receives a warning and pays attention to their surroundings to prevent harm from suspicious individuals.

[0395] Example 2

[0396] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0397] Traffic accidents and robberies pose significant threats, especially to elementary school children and the elderly. However, existing safety systems lack the ability to detect risks in real time or to issue warnings based on the user's emotional state. Therefore, a more effective system is needed to prevent these risks.

[0398] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0399] In this invention, the server includes means for analyzing front and rear images using AI to detect pedestrians jumping out and suspicious people in the vicinity, means for analyzing the emotional state of the user and generating a warning message based on the emotional state, and means for transmitting the warning message to the portable device and notifying the user by voice or vibration. This allows the user to receive risk notifications in real time and receive appropriate warning messages tailored to their emotional state.

[0400] "Front camera and rear camera" refers to a photographing device that is attached to a mobile device and captures images of the front and rear in real time.

[0401] A "portable device" is an electronic device that can be carried by a user and has functions such as a camera, a communication module, and an emotion engine.

[0402] A "communication module" is a device that provides wireless communication functionality for transmitting video data from a portable device to a server.

[0403] The "server" is a central processing unit that receives video data sent from the portable device, analyzes it to detect risks, and generates and sends warning messages.

[0404] "AI analysis" is the process of using artificial intelligence technology on a server to analyze video data to detect objects and assess risks.

[0405] A "warning message" is a notification sent to a portable device to alert a user to a danger, and is presented by sound or vibration.

[0406] An "emotion engine" is software or hardware that analyzes a user's facial expressions and tone of voice to determine their emotional state in real time.

[0407] "Speech synthesis" is a technology that generates natural-sounding speech from text.

[0408] "Vibrate" is a feature that allows a mobile device to generate physical vibrations to provide notifications to the user.

[0409] A "deep learning model" is a multi-layered neural network used to learn complex patterns from data.

[0410] "Object detection" is the process of recognizing specific objects in video data and identifying their type and location.

[0411] "Tracking" is the technique of following the movement of a detected object over time and continuously monitoring its behavior.

[0412] This invention is a system for preventing traffic accidents and robberies that elementary school students and the elderly encounter in their daily lives. The system uses a portable device and a server to analyze video data and the user's emotional state, providing real-time risk detection and warnings.

[0413] System configuration

[0414] The system consists of the following main components:

[0415] 1. Handheld devices

[0416] Front and rear cameras: Used to capture footage in real time.

[0417] Communication module: Used to send acquired video data to the server.

[0418] Emotion engine: Used to analyze the user's facial expressions and tone of voice to determine their emotional state in real time.

[0419] Warning notification function: Equipped with voice output and vibration functions to notify the user of warning messages received from the server.

[0420] 2. Server

[0421] Video receiving module: Used to receive video data transmitted from a portable device.

[0422] AI analysis module: Analyzes the received video data and uses deep learning models to detect objects in the video (vehicles, pedestrians, suspicious individuals, etc.).

[0423] Alert generation module: Based on the detection results and the user's emotional state, an alert message is generated and sent to the mobile device.

[0424] Operational Overview

[0425] Video acquisition and transmission

[0426] The device uses front and rear cameras to capture video in real time and transmits the data to a server at regular intervals, with the communication module playing a key role in this process.

[0427] Video analysis

[0428] The server receives the video data sent from the device and inputs it into an AI analysis module, which uses deep learning to detect objects in the video and track their movements.

[0429] Analyzing the user's emotional state

[0430] The device uses an emotion engine to analyze the user's facial expressions and tone of voice in real time to determine their emotional state.

[0431] Risk detection

[0432] The server assesses risks based on the AI ​​analysis results and the user's emotional state, such as detecting pedestrians running out into the street or suspicious people approaching.

[0433] Generate and send warning messages

[0434] The server generates an appropriate warning message depending on the detected risk and the user's emotional state and sends it to the mobile device.

[0435] Warning Notification

[0436] The device notifies the user of the warning message received from the server by voice or vibration. Specific notification content is voice synthesis such as "Danger! Watch out for people jumping out into the street!" or "There is a suspicious person behind you, so be careful."

[0437] Specific examples

[0438] Example of detecting elementary school students jumping out

[0439] 1. Device: When an elementary school student walks to school, the front camera captures images of the road.

[0440] 2. Terminal: The captured video data is sent to the server.

[0441] 3. Server: The received video data is analyzed by an AI analysis module, and any movement of the elementary school student attempting to jump out is detected.

[0442] 4. Device: The emotion engine analyzes the facial expressions of elementary school students and detects when they are not concentrating.

[0443] 5. Server: Generates a warning message saying "Watch out for people jumping out!" and sends it to the mobile device.

[0444] 6. Device: Elementary school students will receive a voice warning and vibration saying, "Danger! Watch out for people jumping out!"

[0445] 7. User: Elementary school students who hear the warning will stop and be able to prevent a traffic accident.

[0446] Example of detecting suspicious elderly people

[0447] 1. Device: When the elderly person is returning home, the rear camera captures the rear view.

[0448] 2. Terminal: The captured video data is sent to the server.

[0449] 3. Server: The received video data is analyzed by an AI analysis module to detect suspicious individuals approaching the elderly.

[0450] 4. Device: The emotion engine analyzes the facial expressions of the elderly person and detects their state of tension.

[0451] 5. Server: Generates a warning message saying "Beware of suspicious person!" and sends it to the mobile device.

[0452] 6. Terminal: The elderly person will receive a voice warning and vibration saying, "There is a suspicious person behind you, so please be careful."

[0453] 7. User: By receiving the warning, elderly people can prevent harm from suspicious people by paying attention to their surroundings.

[0454] Prompt Sentence Examples

[0455] Below are some example prompts to input to the generative AI model:

[0456] "If an elementary school student is walking and is about to run out into the road, how would the system detect the danger and what kind of warning would it give?"

[0457] An example of a generated description:

[0458] The system captures road images in real time using the mobile device's front camera and sends them to a server. An AI analysis module on the server analyzes the video data and detects any movements of the elementary school student attempting to run out into the street. An emotion engine then analyzes the student's emotional state, and if the student is not concentrating, it generates a stronger warning message, "Danger! Watch out for those running out into the street!", and sends it to the mobile device. Finally, the mobile device issues a voice warning message to alert the elementary school student to the danger.

[0459] The above is a specific embodiment for carrying out the present invention. This system is expected to significantly improve the safety of elementary school children and the elderly, and to prevent traffic accidents and robberies.

[0460] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0461] Step 1: Acquire footage

[0462] The device uses the front and rear cameras to capture video in real time, a process that occurs continuously at a constant frame rate.

[0463] Input: The surrounding physical environment.

[0464] Output: Front and rear video data.

[0465] What it does: The camera on your handheld device captures real-time video of what's in front of and behind you, even while you're moving.

[0466] Step 2: Sending video data

[0467] The device transmits the captured video data to the server at regular intervals using a communication module.

[0468] Input: Front and rear video data.

[0469] Output: Video data transferred to the server.

[0470] Specific operation: Video data is compressed and sent in packet format to the server every 5 seconds.

[0471] Step 3: Receiving video data

[0472] The server receives the video data sent from the terminal. This function is performed by the video receiving module.

[0473] Input: Video data sent from the device.

[0474] Output: Video data stored in the server's memory.

[0475] Specific operation: The server's video receiving module receives the data packets and expands them into memory for analysis.

[0476] Step 4: Analyzing the video data

[0477] The server inputs the received video data into an AI analysis module, which uses deep learning algorithms to detect objects in the video (vehicles, pedestrians, suspicious individuals, etc.) and track their movements.

[0478] Input: Video data stored on the server.

[0479] Output: Data about detected objects and their movements.

[0480] How it works: The AI ​​analysis module analyzes the video data frame by frame, identifies objects in each frame, and tracks their position and movement.

[0481] Step 5: Analyzing the user's emotional state

[0482] The device uses an emotion engine to analyze the user's facial expressions and tone of voice in real time to determine their emotional state.

[0483] Input: User facial and voice data from the camera and microphone on the handheld device.

[0484] Output: The user's emotional state (angry, surprised, relaxed, etc.).

[0485] How it works: The device's emotion engine captures the user's face and voice and analyzes them to determine their emotional state.

[0486] Step 6: Detect risks

[0487] The server combines the results of AI analysis with the user's emotional state to detect risks, such as pedestrians running out into the street or suspicious people approaching.

[0488] Input: Detection results from the AI ​​analysis module, and user emotional state data.

[0489] Output: Information about the detected risks.

[0490] Specific behavior: The server integrates the detected object movement with the user's emotional state to identify dangerous situations, such as when the user is about to run out into the road or when a suspicious person is approaching.

[0491] Step 7: Generate a warning message

[0492] The server generates appropriate warning messages based on the detected risks, which are tailored according to the user's emotional state.

[0493] Input: Information on detected risks, user emotional state data.

[0494] Output: A warning message.

[0495] Specific behavior: The server generates a soft tone warning if the user is nervous and a hard tone warning if the user is unaware.

[0496] Step 8: Sending a warning message

[0497] The server sends the generated warning message to the terminal.

[0498] Input: A warning message.

[0499] Output: The warning message sent to the terminal.

[0500] How it works: The server sends real-time alert messages to the device using a low-latency communication protocol.

[0501] Step 9: Notification of warnings

[0502] The terminal notifies the user of the warning message received from the server by voice or vibration.

[0503] Input: The warning message received from the server.

[0504] Output: The warning that was notified to the user.

[0505] What it does: The device's voice synthesis function warns the user, saying "Danger! Watch out for people jumping out!", while also providing physical feedback to the user using the device's vibration function.

[0506] As described above, each step works in conjunction with the others, allowing users to receive risk notifications in real time and receive appropriate warning messages according to their emotional state.

[0507] (Application example 2)

[0508] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0509] Conventional safety systems using portable devices focus only on risk detection and do not optimize warnings based on the user's emotional state. As a result, the effectiveness of warnings is limited, making it difficult to ensure sufficient safety, especially for users with large emotional fluctuations, such as elementary school children and the elderly. Furthermore, the generated warning messages are uniform, making it difficult to provide warnings appropriate to the situation and the user's state.

[0510] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing front and rear images using AI to detect pedestrians jumping out and suspicious people in the vicinity, means for providing the portable device with an emotion engine for analyzing the user's emotional state, means for generating a warning message based on the detection results and the user's emotional state, and means for optimizing the content of the warning message using a generative AI model. This makes it possible to generate and notify the user of an optimal warning message in real time according to their emotional state, thereby providing greater safety than conventional systems.

[0511] "Front camera and rear camera" refers to a photographing device attached to a mobile device for capturing images of the front and rear.

[0512] A "portable device" is a device that can be easily carried around and is equipped with a communication module and a warning notification function.

[0513] "AI analysis" means using artificial intelligence technology to analyze acquired video data and detect specific patterns and risks.

[0514] "Detecting pedestrians running out into the road and suspicious people in the surrounding area" means that AI analyzes video data to find pedestrians who are about to run out into the road or people behaving abnormally.

[0515] An "emotion engine" is an algorithm or software that analyzes information such as a user's facial expressions and tone of voice to determine their emotional state in real time.

[0516] "Generating a warning message" means creating appropriate warning content based on the detected risk and the user's emotional state.

[0517] "Notifying the user of a warning by voice or vibration" means that the generated warning message is conveyed to the user by using a voice output device or a vibration function.

[0518] A "generative AI model" is an artificial intelligence model that creates appropriate responses and outputs based on generated data and input information.

[0519] "Optimization" means making adjustments to obtain the most effective or efficient method or result under given conditions.

[0520] "Generating and notifying in real time" means analyzing data in response to the situation at hand and providing results and information immediately.

[0521] 1. System program generation

[0522] This system uses a portable device equipped with a front and rear camera and works in conjunction with a server to analyze both risk and emotional state. Users can carry the portable device and monitor the risks around them in real time. The system has the following functions:

[0523] 2. Program processing explanation

[0524] Video acquisition and transmission

[0525] The terminal captures video in real time using the front and rear cameras. This video data is sent to the server via a communication module. The terminal has a built-in communication module that supports sending and receiving video data.

[0526] Server-based video analysis

[0527] The server analyzes the received video data. Specifically, the AI ​​analysis module uses a deep learning model to detect objects in the video (e.g., vehicles, pedestrians, suspicious individuals) and track their movements. This makes it possible to detect risks such as pedestrians running out into the street or suspicious individuals approaching in real time.

[0528] Emotion analysis

[0529] The emotion engine installed in the device uses a camera and audio sensors to analyze the user's facial expressions and tone of voice to determine the user's emotional state. A deep learning model is used to analyze the emotional state.

[0530] Generating and notifying warning messages

[0531] The server generates an optimal warning message based on the risk detection results and the user's emotional state. The content of this warning message is optimized using a generative AI model. The generated warning message is then sent back to the device via the communication module.

[0532] The device notifies the user of the received warning message using a voice output device or vibration function. For example, it generates messages such as "Danger! Watch out for people jumping out into the street!" or "Be careful, there is a suspicious person behind you."

[0533] 3. Specific Examples

[0534] Detecting elementary school students jumping out

[0535] When elementary school students use the vehicle on their way home from school, the front camera captures images of the road in real time. If the server detects an approaching vehicle and determines that the child is not concentrating through emotional analysis, it issues a voice warning saying "Danger! Watch out for children running out into the street!" and simultaneously activates a vibration. This allows elementary school students to pay close attention to the warning and prevent them from running out into the street.

[0536] Detecting suspicious elderly people

[0537] When an elderly person is returning home, a rear camera captures video of the area behind them. If analysis on the server detects that a suspicious person is approaching from behind, and emotion analysis determines that the elderly person is nervous, a gentle voice warning is issued saying, "There is a suspicious person behind you, please be careful." This allows the elderly person to deal with the situation calmly and prevent harm from the suspicious person.

[0538] Prompt Sentence Examples

[0539] After-school risk detection prompts:

[0540] Design a system that uses an AI algorithm to detect in real time when a vehicle ahead is approaching, and issues an audio warning saying "Danger! Watch out for vehicles jumping out!" if the user's emotional state is not focused.

[0541] Hardware and Software

[0542] The server requires a data center component with high-performance computing resources and an AI framework (e.g., TensorFlow, PyTorch) for running deep learning models, while the device requires a high-resolution camera, an audio sensor, a communication module (e.g., Wi-Fi, 4G / 5G module), and an on-device AI framework (e.g., TensorFlow Lite) for running a sentiment analysis engine.

[0543] This enables the entire system to work together to detect risks in real time and deliver optimized warning messages to enhance user safety.

[0544] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0545] Step 1:

[0546] The terminal captures video in real time using a front camera and a rear camera. A user carries a portable device and periodically captures video of the front and rear. This allows the terminal to obtain video data of the current surrounding environment. The input is the real-time video data obtained from the camera, and the output is this video data.

[0547] Step 2:

[0548] The terminal transmits the acquired front and rear video data to the server via the communication module. The communication module in the device is activated and uploads the video data to the server via the network. At this stage, the input is the front and rear video data, and the output is the video data received by the server.

[0549] Step 3:

[0550] The server analyzes the received video data using a deep learning model. Specifically, the AI ​​analysis module detects objects in the video data and tracks their movements. The server evaluates risks, such as pedestrians running out into the street or suspicious people approaching, in real time. The input is the video data received by the server, and the output is the risk detection results.

[0551] Step 4:

[0552] The device uses an emotion engine to analyze the user's emotional state in real time. The device's camera and audio sensors capture the user's face and voice data, which are then analyzed by an emotion analysis algorithm. The input is the user's face and voice data, and the output is the analyzed emotional state.

[0553] Step 5:

[0554] The server uses a generative AI model to generate an optimal warning message based on the risk detection results and the user's emotional state. The server determines the appropriate warning message by taking into account the type of risk and the user's current emotional state. The input is the risk detection results and the user's emotional state, and the output is the warning message.

[0555] Step 6:

[0556] The terminal receives the generated warning message via the communication module. The terminal acquires the warning message sent from the server and notifies the user using audio output and vibration functions. The input is the warning message sent from the server, and the output is the warning message notified to the user.

[0557] This allows the entire system to work together to detect risks and provide emotion-based warning notifications to enhance user safety in real time.

[0558] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0559] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0560] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0561] [Second embodiment]

[0562] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0563] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0564] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0565] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0566] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0567] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0568] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0569] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0570] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0571] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0572] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0573] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0574] MODE FOR CARRYING OUT THE INVENTION

[0575] This invention is a system to prevent traffic accidents and robberies that elementary school children and the elderly may encounter in their daily lives. Specifically, it uses a portable device equipped with a front and rear camera and utilizes AI technology to detect surrounding risks and issue a warning.

[0576] System configuration

[0577] This system consists of the following main components:

[0578] 1. Handheld devices

[0579] Front and rear cameras: These are necessary elements for capturing images in real time.

[0580] Communication module: Has the function to send captured images to a server.

[0581] Warning notification function: This module notifies the user of warning messages received from the server. Specifically, it has audio output and vibration functions.

[0582] 2. Server

[0583] Video receiving module: Has the function to receive video data transmitted from a portable device.

[0584] AI analysis module: Analyzes the received video data and executes algorithms to detect sudden jumps in and suspicious individuals.

[0585] Alert generation module: Based on the detection results, it generates an alert message and sends it to the mobile device.

[0586] System Operation

[0587] The operation of this system will be explained below.

[0588] 1. Video acquisition and transmission

[0589] Terminal: The mobile device captures video in real time using front and rear cameras and transmits the data to the server at regular intervals.

[0590] 2. Video Analysis

[0591] Server: The server receives the video data sent from the mobile device and inputs it into the AI ​​analysis module, which uses deep learning to detect objects in the video (e.g., vehicles, pedestrians, suspicious individuals).

[0592] 3. Risk detection and prediction

[0593] Server: Using AI analysis, it assesses the risk of pedestrians running out into the road and the approach of suspicious individuals. For example, it detects when a pedestrian is about to run out into the road or when a suspicious individual is approaching an elderly person.

[0594] 4. Generating and sending warning messages

[0595] Server: Automatically generates warning messages based on each risk and sends them to the mobile device.

[0596] 5. Warning Notification

[0597] Terminal: The terminal notifies the user of warning messages received from the server by voice or vibration. Specifically, it uses voice synthesis to issue warnings such as "Danger! Watch out for people jumping out into the street!" or "There's a suspicious person behind you, so be careful."

[0598] Specific examples

[0599] Detecting elementary school students jumping out

[0600] Device: As elementary school students walk to school, the front-facing camera captures footage of the road.

[0601] Terminal: Video data is sent to the server.

[0602] Server: The video received by the server is analyzed by AI, and any movement of the elementary school student trying to jump out is detected.

[0603] Server: Generates a "Watch out for people jumping out!" warning message and sends it to the handheld device.

[0604] Device: Elementary school students will hear a voice warning saying, "Danger! Watch out for people jumping out!"

[0605] User: Elementary school student listens to the warning and stops to prevent a traffic accident.

[0606] Detecting suspicious elderly people

[0607] Device: The rear camera captures rearward footage as the elderly person walks home.

[0608] Terminal: The acquired video data is sent to the server.

[0609] Server: The server analyzes the video and detects suspicious individuals approaching the elderly.

[0610] Server: Generates a "Beware of suspicious person!" warning message and sends it to the mobile device.

[0611] Terminal: A voice warning is issued to the elderly person saying, "There is a suspicious person behind you, so please be careful."

[0612] Users: Elderly people receive a warning and pay attention to their surroundings to prevent harm from suspicious people.

[0613] As described above, the present invention improves the safety of elementary school children and the elderly in their daily lives through a system that combines a portable device and a server. The real-time risk detection and warning functions provided by this system prevent traffic accidents and robberies.

[0614] The processing flow will be explained below.

[0615] Program processing flow

[0616] Video acquisition and transmission

[0617] Device:

[0618] 1. Step 1: Activate the front and rear cameras on your mobile device.

[0619] 2. Step 2: Capture video in real time using the front and rear cameras. Video frames are acquired at a rate of, for example, 30 frames per second.

[0620] 3. Step 3: Temporarily save the captured video data in memory.

[0621] 4. Step 4: At regular intervals (for example, every second), the video data stored in memory is divided into packets and organized.

[0622] 5. Step 5: The organized data packets are sent to the server via the communication module using HTTP or WebSocket.

[0623] Video analysis

[0624] server:

[0625] 1. Step 1: The server receives the front and rear camera image data.

[0626] 2. Step 2: The received video data is decoded and input into the AI ​​analysis module.

[0627] 3. Step 3: The AI ​​analysis module uses a deep learning model (e.g., YOLO, SSD) to detect objects (e.g., vehicles, pedestrians, suspicious people) in the video frame.

[0628] 4. Step 4: Track the movement of the detected object and analyze its position and velocity.

[0629] 5. Step 5: Evaluate the risk of pedestrians running out into the street or the approach of suspicious people. Specifically, predict whether certain movements pose a risk of traffic accidents or robberies.

[0630] Risk detection and warning

[0631] server:

[0632] 1. Step 1: Determine whether there is a risk based on the results of AI analysis. For example, evaluate whether a pedestrian is about to run out into the road or whether a suspicious person is approaching.

[0633] 2. Step 2: If a risk is detected, generate an appropriate warning message, which includes the warning type (e.g., watch out for a sudden jump, watch out for a suspicious person) and a recommended action (e.g., stop, turn around).

[0634] 3. Step 3: Send the generated warning message to the terminal.

[0635] Warning Notification

[0636] Device:

[0637] 1. Step 1: Decode the warning message received from the server.

[0638] 2. Step 2: The decoded warning message is notified to the user. Specifically, the voice synthesis module generates a warning voice and outputs it from the speaker or activates the vibration motor.

[0639] User Feedback

[0640] User:

[0641] 1. Step 1: Receive an alert notification from your device, for example, hear a sound notification or feel a vibration.

[0642] 2. Step 2: Follow the warning and take appropriate action. Specifically, if you receive a warning about people jumping out into the street, stop, and if you receive a warning about suspicious people, turn around and check your surroundings.

[0643] Specific examples

[0644] Detecting elementary school students jumping out

[0645] Device:

[0646] 1. Step 1: When an elementary school student is walking to school, the front camera captures images of the road.

[0647] 2. Step 2: The video data is sent to the server.

[0648] server:

[0649] 1. Step 1: The video received by the server is analyzed by AI, and any movement of the elementary school student trying to jump out is detected.

[0650] 2. Step 2: Generate a warning message "Watch out for people jumping out!" and send it to the mobile device.

[0651] Device:

[0652] 1. Step 1: Elementary school students will hear a voice warning saying, "Danger! Watch out for people jumping out!"

[0653] User:

[0654] 1. Step 1: Elementary school students listen to the warning and stop to prevent traffic accidents.

[0655] Detecting suspicious elderly people

[0656] Device:

[0657] 1. Step 1: When an elderly person is returning home, the rear camera captures the rear view.

[0658] 2. Step 2: The acquired video data is sent to the server.

[0659] server:

[0660] 1. Step 1: The server analyzes the video and detects suspicious individuals approaching the elderly person.

[0661] 2. Step 2: Generate a warning message saying "Beware of suspicious person!" and send it to the mobile device.

[0662] Device:

[0663] 1. Step 1: The elderly person receives a voice warning saying, "There is a suspicious person behind you, so please be careful."

[0664] User:

[0665] 1. Step 1: The elderly person receives a warning and pays attention to their surroundings to prevent harm from suspicious individuals.

[0666] Example 1

[0667] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0668] This invention relates to a system that prevents traffic accidents and robberies that elementary school students and the elderly may encounter in their daily lives. In particular, it aims to provide a system that can detect surrounding risks in real time and issue accurate warnings. Conventional systems sometimes experience delays in detecting risks and generating warning messages, making them insufficient to avoid actual danger. Therefore, the present invention aims to prevent traffic accidents and robberies by quickly and accurately detecting risks and issuing warnings to users.

[0669] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0670] In this invention, the server includes means for receiving the acquired video data with a video receiving module, analyzing the front and rear video using an AI analysis module, and detecting pedestrians running out into the road and suspicious people in the vicinity, means for generating a warning message with a warning generation module based on the detection results, and means for transmitting the warning message to the mobile terminal device via a communication module and notifying the user of the warning using a warning notification function. This makes it possible to detect risks quickly and accurately and notify the user of appropriate warnings in real time.

[0671] A "portable terminal device" is a device that has a front camera and a rear camera and can be carried by a user.

[0672] A "communication module" is a combination of hardware and software for transmitting data from a mobile terminal device to a server.

[0673] A "server" is a computing system that receives, analyzes, and processes data sent from a mobile terminal device.

[0674] The "video receiving module" is a module having a function for receiving video data transmitted from a portable terminal device at a server.

[0675] The "AI analysis module" is a module that uses artificial intelligence technology to analyze received video data and detect specific patterns and risks.

[0676] The "warning generation module" is a module for generating appropriate warning messages based on the detection results of the AI ​​analysis module.

[0677] The "warning notification function" is a function in which the portable terminal device notifies the user of a warning message by voice or vibration.

[0678] "Pedestrian running out" refers to a pedestrian unintentionally running out into a dangerous place such as the roadway.

[0679] A "suspicious individual" is an individual who deviates from normal patterns of behavior and is perceived as a threat to users.

[0680] A "deep learning model" is an artificial intelligence technology that learns complex patterns from large amounts of data and uses them for analysis.

[0681] This invention is a system for preventing traffic accidents and robberies in the daily lives of elementary school children and the elderly. This system is constructed by combining a portable terminal device equipped with a front camera and a rear camera with a server. The configuration and operation of this system are described in detail below.

[0682] System configuration

[0683] 1. Portable terminal device

[0684] Front and rear cameras:

[0685] The mobile terminal device is equipped with a camera that captures images of the front and rear in real time, allowing the user to monitor the situation around them in detail.

[0686] Communication Module:

[0687] The mobile terminal device includes a communication module for transmitting the acquired video data to a server. Specifically, a 4G LTE module is used.

[0688] Warning notification function:

[0689] The terminal is equipped with a voice output and vibration function as a warning notification function, which allows the terminal to convey the warning message received from the server to the user.

[0690] 2. Server

[0691] Video Receiving Module:

[0692] The server uses a video receiving module, specifically Apache Kafka, to receive video data sent from the mobile terminal device.

[0693] AI analysis module:

[0694] The server is equipped with an AI analysis module that analyzes the received video data and uses deep learning models (e.g., TensorFlow) to detect objects in the video (e.g., vehicles, pedestrians, suspicious individuals).

[0695] Warning generation modules:

[0696] The server includes a warning generation module for generating a warning message based on the detection result by the AI ​​analysis module, and the generated warning message is transmitted to the mobile terminal device again.

[0697] System Operation

[0698] Video acquisition and transmission

[0699] Device: Front and rear cameras capture images of the user's surroundings in real time. The image data is sent to a server at regular intervals.

[0700] Receiving and analyzing video data

[0701] Server: The video receiving module receives the video data and inputs it into the AI ​​analysis module. Analysis detects pedestrians running out into the street and suspicious people approaching.

[0702] Evaluating risks and generating warning messages

[0703] Server: Based on the results from the AI ​​analysis module, a risk assessment is performed and an appropriate warning message is generated in the warning generation module. For example, messages such as "Watch out for people jumping out!" or "Watch out for suspicious people!" are generated.

[0704] Warning message notification

[0705] Terminal: A mobile terminal device receives the warning message and notifies the user through voice or vibration. A warning message is generated using voice synthesis technology (e.g., Google Cloud Text-to-Speech) and output from the speaker.

[0706] Specific examples

[0707] Detecting elementary school students jumping out

[0708] Device: When an elementary school student walks to school, the front camera captures footage of the road.

[0709] Terminal: Video data is sent to the server.

[0710] Server: The video received by the server is analyzed by AI, and any movement of the elementary school student trying to jump out is detected.

[0711] Server: Generates a warning message saying "Watch out for people jumping out!" and sends it to the mobile terminal device.

[0712] Device: Elementary school students will hear a voice warning saying, "Danger! Watch out for people jumping out!"

[0713] User: Elementary school student listens to the warning and stops to prevent a traffic accident.

[0714] Detecting suspicious elderly people

[0715] Device: The rear camera captures rearward footage as the elderly person walks home.

[0716] Terminal: The acquired video data is sent to the server.

[0717] Server: The server analyzes the video and detects suspicious individuals approaching the elderly.

[0718] Server: Generates a warning message saying "Beware of suspicious person!" and sends it to the mobile terminal device.

[0719] Terminal: A voice warning is issued to the elderly person saying, "There is a suspicious person behind you, so please be careful."

[0720] Users: Elderly people receive a warning and pay attention to their surroundings to prevent harm from suspicious people.

[0721] Prompt Sentence Examples

[0722] "While elementary school students are walking to school, please capture road footage with a front camera, use AI to detect the risk of them running out into the street, and use a notification system to generate a warning message. Please also tell us details about the hardware and software required."

[0723] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0724] Step 1: Acquire footage

[0725] Terminal: The mobile terminal device captures video in real time using the front and rear cameras, capturing video at 30 frames per second and storing it in an internal buffer.

[0726] Input: Real-time video of the user's surroundings.

[0727] Output: Front and rear video data stored in internal buffers.

[0728] Step 2: Sending the video

[0729] Terminal: At regular intervals (e.g., every 5 seconds), the video data stored in the buffer is sent to the server via the communications module. Specifically, the data is encoded via the 4G LTE module and transferred to the server via TCP / IP protocol.

[0730] Input: Video data stored in the internal buffer.

[0731] Output: Video data sent to the server.

[0732] Step 3: Receiving video data

[0733] Server: The server-side video receiving module receives the video data sent from the mobile terminal device and stores it in a database. Specifically, it uses Apache Kafka to manage the data stream and store it appropriately.

[0734] Input: Video data transmitted from a mobile terminal device.

[0735] Output: Video data stored in a database.

[0736] Step 4: Analyzing the footage

[0737] Server: The received video data is input into the AI ​​analysis module and analyzed using a deep learning model (e.g., TensorFlow). Specifically, objects in the video (vehicles, pedestrians, suspicious individuals, etc.) are detected and their movements are tracked.

[0738] Input: Video data stored in a database.

[0739] Output: Analysis results (e.g., JSON format) containing information about detected objects.

[0740] Step 5: Risk detection and prediction

[0741] Server: Based on the results of the AI ​​analysis module, it evaluates the risk of pedestrians running out into the street and the approach of suspicious individuals. Specifically, it calculates the speed and distance of objects and executes an algorithm to assess the risk level.

[0742] Input: Analysis results of the AI ​​analysis module.

[0743] Output: Evaluated risk information (e.g., risk of people jumping out, suspicious person warning).

[0744] Step 6: Generate a warning message

[0745] Server: Based on the detected risks, the warning generation module generates warning messages. Specifically, a Python script is used to generate warning messages (e.g., "Watch out for people jumping out!", "Watch out for suspicious people!") as text and audio data.

[0746] Input: Assessed risk information.

[0747] Output: Warning message in text and audio format.

[0748] Step 7: Sending a warning message

[0749] Server: The server transmits the generated warning message to the mobile terminal device. Specifically, the server encodes the warning message via the communication module and transmits it.

[0750] Input: Warning message in text and audio format.

[0751] Output: The warning message sent to the handheld device.

[0752] Step 8: Notification of warnings

[0753] Terminal: The mobile terminal device notifies the user of the received warning message. Specifically, it generates the warning message using voice synthesis technology and outputs it from the speaker. It also issues a physical warning using the vibration function.

[0754] Input: The warning message received.

[0755] Output: Audio and vibration alert notifications.

[0756] (Application example 1)

[0757] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0758] Conventional safety monitoring systems in factories mainly use fixed cameras, making it difficult to effectively monitor the movements of dynamically moving robots and people in real time. Furthermore, advanced analysis technology is required to detect abnormal or suspicious behavior, and it is not possible to issue appropriate warnings immediately. This results in low accuracy in improving safety in factories, and there is a high possibility of delayed response in emergencies.

[0759] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0760] In this invention, the server includes a means for acquiring video from the front and rear, a means for analyzing the acquired video using AI to detect abnormal or suspicious behavior, and a means for generating a warning message based on the detection results and transmitting it to the portable device or robot. This makes it possible to detect abnormal or suspicious behavior with high accuracy in real time and to immediately issue warnings to robots and workers in the factory by voice and vibration.

[0761] "Front camera and rear camera" are imaging devices for capturing images in front of and behind the device.

[0762] A "portable device" is a device that can be carried around, has front and rear cameras, and has the ability to capture and transmit images.

[0763] A "server" is a computer system that receives and analyzes video sent from a portable device.

[0764] "AI analysis" is the process of using artificial intelligence technology to analyze video data and detect specific patterns or anomalies.

[0765] A "warning message" is a message that is generated to warn the user when an abnormality or risk is detected.

[0766] A "factory robot" is a mechanical device that operates autonomously or semi-autonomously in a factory to perform designated tasks.

[0767] An "audio output device" is a device for producing electronically generated speech.

[0768] The "vibration function" is a function that causes a device or robot to vibrate to provide a physical notification to the user.

[0769] "Real-time" refers to the fact that information acquisition, processing, and output of results are carried out almost simultaneously.

[0770] "Abnormal and suspicious behavior" refers to behavior that deviates from normal patterns of behavior or that appears suspicious.

[0771] A "deep learning model" is an artificial intelligence model that uses a multi-layer neural network to learn the characteristics of data and perform advanced predictions and classifications.

[0772] "Tracking" is a technology that continuously tracks a specific object or movement and evaluates its position and status.

[0773] The present invention provides a system for monitoring the safety of robots and workers in a factory, detecting abnormal or suspicious behavior in real time, and issuing a warning. A specific embodiment of this system will be described below.

[0774] Main components of the system

[0775] 1. Handheld devices

[0776] It is equipped with a front-facing camera and a rear-facing camera, which capture images in real time. It also has the function of transmitting the images to a server via a communication module and receiving warning messages.

[0777] 2. Server

[0778] It receives video data sent from a portable device, performs AI analysis, and generates a warning message based on the analysis results, which is then sent back to the portable device.

[0779] 3. Robots in factories

[0780] Equipped with front and rear cameras, it captures video footage to detect abnormal or suspicious activity, and receives warning messages and issues audio and vibration alerts.

[0781] Hardware and Software Used

[0782] Cameras: Front and rear facing cameras

[0783] Communication module: Wi-Fi or Bluetooth

[0784] Server: High-performance server equipped with GPU

[0785] AI analysis software: TensorFlow, PyTorch

[0786] Video capture software: OpenCV

[0787] Warning notification software: paho-mqtt (communication), pyttsx3 (voice synthesis)

[0788] System Operation

[0789] The server receives video data sent from the mobile devices and factory robots. It analyzes this data using a deep learning model to detect abnormal or suspicious behavior in real time. Based on the results, it generates a warning message and sends it back to the mobile devices and robots. The robots then immediately notify them of the received warning message by voice and vibration.

[0790] Specific examples

[0791] Abnormal behavior detection

[0792] 1. Terminal (handheld device): Robots in the factory capture front and rear images in real time and send them to the server.

[0793] 2. Server: Analyzes the received video using a deep learning model to detect abnormal behavior, such as when a robot deviates from its normal behavior pattern.

[0794] 3. Server: Generates a warning message saying "Danger! Caution! Watch the machine!" and sends it to the robots in the factory.

[0795] 4. Terminal (robot in factory): Notifies warnings by voice and vibration.

[0796] Suspicious Activity Detection

[0797] 1. Terminal (portable device): A robot moving around the factory captures rearward-facing images and sends them to a server.

[0798] 2. Server: AI analyzes the received video and detects suspicious movements, such as when a suspicious person approaches the robot.

[0799] 3. Server: Generates a warning message saying "Suspicious behavior detected!" and sends it to the robots in the factory.

[0800] 4. Terminal (robot in factory): Notifies warnings by voice and vibration.

[0801] Prompt Sentence Examples

[0802] Factory robots capture images from cameras in front and behind them in real time and use AI to monitor the safety of the work area. For example, if a person or machine behaves abnormally, it will automatically issue a voice warning saying, "Danger! Watch out for machine operation!" What will happen if an abnormality is detected?

[0803] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0804] Step 1:

[0805] The terminal (a robot in a factory) acquires real-time images from the front and rear. It uses a front camera and a rear camera to capture the images. The input obtained is raw image data.

[0806] Step 2:

[0807] The video data captured by the terminal (a robot in a factory) is sent to the server via a communication module. Using a communication protocol (e.g., MQTT or HTTP), the video data is divided into packets and sent to the server. The server then receives the video data as input.

[0808] Step 3:

[0809] The video data received by the server is input into an AI analysis module, which uses a deep learning model (e.g., TensorFlow or PyTorch) to recognize objects in the video and analyze their behavior. This outputs intermediate data (e.g., the position of each object and its movement vector) that can be used to identify abnormal or suspicious behavior.

[0810] Step 4:

[0811] The server performs risk assessment based on the results of AI analysis. It evaluates identified abnormal or suspicious behavior and runs an algorithm to calculate the risk level. This generates a warning message for any detected risks.

[0812] Step 5:

[0813] The server generates a warning message and sends it back to the terminal (the robot in the factory) via the communication module, which then receives the warning message.

[0814] Step 6:

[0815] Based on the warning message received by the terminal (a robot in the factory), a warning is issued by voice and vibration. The voice output device is used to issue a voice warning, and the vibration module is used to notify the warning by vibration. The final output is a real-time warning notification to the user.

[0816] Step 7:

[0817] The user (worker) receives a warning from the robot and responds immediately. The worker receives a warning via voice and vibration and responds to the surrounding risks.

[0818] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0819] MODE FOR CARRYING OUT THE INVENTION

[0820] This invention is a system to prevent traffic accidents and robberies that elementary school children and the elderly may encounter in their daily lives. Specifically, it uses a portable device equipped with a front and rear camera, and utilizes AI technology and an emotion engine to detect surrounding risks and issue warnings according to the user's emotional state.

[0821] System configuration

[0822] This system consists of the following main components:

[0823] 1. Handheld devices

[0824] Front and rear cameras: These are necessary elements for capturing images in real time.

[0825] Communication module: Has the function to send captured images to a server.

[0826] Emotion engine: Has the function to analyze the user's emotional state in real time.

[0827] Warning notification function: This module notifies the user of warning messages received from the server. Specifically, it has audio output and vibration functions.

[0828] 2. Server

[0829] Video receiving module: Has the function to receive video data transmitted from a portable device.

[0830] AI analysis module: Analyzes the received video data and executes algorithms to detect sudden jumps in and suspicious individuals.

[0831] Alert generation module: Based on the detection results, it generates an alert message and sends it to the mobile device.

[0832] System Operation

[0833] The operation of this system will be explained below.

[0834] 1. Video acquisition and transmission

[0835] Terminal: The mobile device captures video in real time using front and rear cameras and transmits the data to the server at regular intervals.

[0836] 2. Video Analysis

[0837] Server: The server receives the video data sent from the mobile device and inputs it into the AI ​​analysis module, which uses deep learning to detect objects in the video (e.g., vehicles, pedestrians, suspicious individuals).

[0838] 3. Analysis of the user's emotional state

[0839] Terminal: The emotion engine installed in the portable device uses a camera and audio sensors to analyze the user's facial expressions and tone of voice, and determine their emotional state in real time.

[0840] 4. Risk detection and warning

[0841] Server: Using AI analysis, it assesses the risk of pedestrians running out into the road and the approach of suspicious individuals. For example, it detects when a pedestrian is about to run out into the road or when a suspicious individual is approaching an elderly person.

[0842] Server: When a risk is detected, it generates an appropriate warning message according to the user's emotional state and sends it to the mobile device. Depending on the emotional state, it adjusts the content of the warning and the notification method.

[0843] 5. Warning Notification

[0844] Terminal: The terminal notifies the user of warning messages received from the server by voice or vibration. Specifically, it uses voice synthesis to issue warnings such as "Danger! Watch out for people jumping out into the street!" or "There's a suspicious person behind you, so be careful."

[0845] Specific examples

[0846] Detecting elementary school students jumping out

[0847] Device: As elementary school students walk to school, the front-facing camera captures footage of the road.

[0848] Terminal: Video data is sent to the server.

[0849] Server: The video received by the server is analyzed by AI, and any movement of the elementary school student trying to jump out is detected.

[0850] Device: The emotion engine analyzes the emotional state of elementary school students and generates stronger warning messages if they are not concentrating.

[0851] Server: Generates a "Watch out for people jumping out!" warning message and sends it to the handheld device.

[0852] Device: Elementary school students will hear a voice warning saying, "Danger! Watch out for people jumping out!"

[0853] User: Elementary school student listens to the warning and stops to prevent a traffic accident.

[0854] Detecting suspicious elderly people

[0855] Device: The rear camera captures rearward footage as the elderly person walks home.

[0856] Terminal: The acquired video data is sent to the server.

[0857] Server: The server analyzes the video and detects suspicious individuals approaching the elderly.

[0858] Device: Detects the emotional state of the elderly person and generates a soft-toned warning message if they are in a state of tension.

[0859] Server: Generates a "Beware of suspicious person!" warning message and sends it to the mobile device.

[0860] Terminal: A voice warning is issued to the elderly person saying, "There is a suspicious person behind you, so please be careful."

[0861] Users: Elderly people receive a warning and pay attention to their surroundings to prevent harm from suspicious people.

[0862] As described above, this invention improves the safety of elementary school children and the elderly in their daily lives through a system that combines a portable device and a server. The real-time risk detection and warning functions provided by this system prevent traffic accidents and robberies. Furthermore, by combining it with an emotion engine, it is possible to issue warnings according to the user's emotional state, further improving safety.

[0863] The processing flow will be explained below.

[0864] Program processing flow

[0865] Video acquisition and transmission

[0866] Device:

[0867] 1. Step 1: Activate the front and rear cameras on your mobile device.

[0868] 2. Step 2: Capture video in real time using the front and rear cameras. Video frames are acquired at a rate of, for example, 30 frames per second.

[0869] 3. Step 3: Temporarily save the captured video data in memory.

[0870] 4. Step 4: At regular intervals (for example, every second), the video data stored in memory is divided into packets and organized.

[0871] 5. Step 5: The organized data packets are sent to the server via the communication module using HTTP or WebSocket.

[0872] Video analysis

[0873] server:

[0874] 1. Step 1: The server receives the front and rear camera image data.

[0875] 2. Step 2: The received video data is decoded and input into the AI ​​analysis module.

[0876] 3. Step 3: The AI ​​analysis module uses a deep learning model (e.g., YOLO, SSD) to detect objects (e.g., vehicles, pedestrians, suspicious people) in the video frame.

[0877] 4. Step 4: Track the movement of the detected object and analyze its position and velocity.

[0878] 5. Step 5: Evaluate the risk of pedestrians running out into the street or the approach of suspicious people. Specifically, predict whether certain movements pose a risk of traffic accidents or robberies.

[0879] Analyzing the user's emotional state

[0880] Device:

[0881] 1. Step 1: Start the emotion engine on the mobile device.

[0882] 2. Step 2: Use a camera or audio sensor to capture the user's facial expressions and tone of voice.

[0883] 3. Step 3: Analyze the captured data in real time to determine the user's emotional state, for example, identifying states such as tension, anxiety, or concentration.

[0884] 4. Step 4: Send the analysis results to the server.

[0885] Risk detection and warning

[0886] server:

[0887] 1. Step 1: Determine whether there is a risk based on the results of AI analysis. For example, evaluate whether a pedestrian is about to run out into the road or whether a suspicious person is approaching.

[0888] 2. Step 2: If a risk is detected, a warning message is generated based on the analysis of the transmitted emotional state.

[0889] 3. Step 3: Adjust the content and notification method of the warning message according to the user's emotional state.

[0890] 4. Step 4: Send the generated warning message to the terminal.

[0891] Warning Notification

[0892] Device:

[0893] 1. Step 1: Decode the warning message received from the server.

[0894] 2. Step 2: The decoded warning message is notified to the user. Specifically, the voice synthesis module generates a warning voice and outputs it from the speaker or activates the vibration motor.

[0895] User Feedback

[0896] User:

[0897] 1. Step 1: Receive an alert notification from your device, for example, hear a sound notification or feel a vibration.

[0898] 2. Step 2: Follow the warning and take appropriate action. Specifically, if you receive a warning about people jumping out into the street, stop, and if you receive a warning about suspicious people, turn around and check your surroundings.

[0899] Specific examples

[0900] Detecting elementary school students jumping out

[0901] Device:

[0902] 1. Step 1: When an elementary school student is walking to school, the front camera captures images of the road.

[0903] 2. Step 2: The video data is sent to the server.

[0904] server:

[0905] 1. Step 1: The video received by the server is analyzed by AI, and any movement of the elementary school student trying to jump out is detected.

[0906] 2. Step 2: Receive data from the emotion engine to analyze the emotional state of elementary school students.

[0907] 3. Step 3: If the elementary school student is not concentrating, generate a stronger warning message.

[0908] 4. Step 4: Generate a warning message "Watch out for people jumping out!" and send it to the mobile device.

[0909] Device:

[0910] 1. Step 1: Elementary school students will hear a voice warning saying, "Danger! Watch out for people jumping out!"

[0911] User:

[0912] 1. Step 1: Elementary school students listen to the warning and stop to prevent traffic accidents.

[0913] Detecting suspicious elderly people

[0914] Device:

[0915] 1. Step 1: When an elderly person is returning home, the rear camera captures the rear view.

[0916] 2. Step 2: The acquired video data is sent to the server.

[0917] 3. Step 3: The emotion engine captures the emotional state of the elderly person and sends the analysis results to the server.

[0918] server:

[0919] 1. Step 1: The server analyzes the video and detects suspicious individuals approaching the elderly person.

[0920] 2. Step 2: Evaluate the elderly person's emotional state and generate a soft-toned warning message if they are in a tense state.

[0921] 3. Step 3: Generate a warning message saying "Beware of suspicious person!" and send it to the mobile device.

[0922] Device:

[0923] 1. Step 1: The elderly person receives a voice warning saying, "There is a suspicious person behind you, so please be careful."

[0924] User:

[0925] 1. Step 1: The elderly person receives a warning and pays attention to their surroundings to prevent harm from suspicious individuals.

[0926] Example 2

[0927] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0928] Traffic accidents and robberies pose significant threats, especially to elementary school children and the elderly. However, existing safety systems lack the ability to detect risks in real time or to issue warnings based on the user's emotional state. Therefore, a more effective system is needed to prevent these risks.

[0929] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0930] In this invention, the server includes means for analyzing front and rear images using AI to detect pedestrians jumping out and suspicious people in the vicinity, means for analyzing the emotional state of the user and generating a warning message based on the emotional state, and means for transmitting the warning message to the portable device and notifying the user by voice or vibration. This allows the user to receive risk notifications in real time and receive appropriate warning messages tailored to their emotional state.

[0931] "Front camera and rear camera" refers to a photographing device that is attached to a mobile device and captures images of the front and rear in real time.

[0932] A "portable device" is an electronic device that can be carried by a user and has functions such as a camera, a communication module, and an emotion engine.

[0933] A "communication module" is a device that provides wireless communication functionality for transmitting video data from a portable device to a server.

[0934] The "server" is a central processing unit that receives video data sent from the portable device, analyzes it to detect risks, and generates and sends warning messages.

[0935] "AI analysis" is the process of using artificial intelligence technology on a server to analyze video data to detect objects and assess risks.

[0936] A "warning message" is a notification sent to a portable device to alert a user to a danger, and is presented by sound or vibration.

[0937] An "emotion engine" is software or hardware that analyzes a user's facial expressions and tone of voice to determine their emotional state in real time.

[0938] "Speech synthesis" is a technology that generates natural-sounding speech from text.

[0939] "Vibrate" is a feature that allows a mobile device to generate physical vibrations to provide notifications to the user.

[0940] A "deep learning model" is a multi-layered neural network used to learn complex patterns from data.

[0941] "Object detection" is the process of recognizing specific objects in video data and identifying their type and location.

[0942] "Tracking" is the technique of following the movement of a detected object over time and continuously monitoring its behavior.

[0943] This invention is a system for preventing traffic accidents and robberies that elementary school students and the elderly encounter in their daily lives. The system uses a portable device and a server to analyze video data and the user's emotional state, providing real-time risk detection and warnings.

[0944] System configuration

[0945] The system consists of the following main components:

[0946] 1. Handheld devices

[0947] Front and rear cameras: Used to capture footage in real time.

[0948] Communication module: Used to send acquired video data to the server.

[0949] Emotion engine: Used to analyze the user's facial expressions and tone of voice to determine their emotional state in real time.

[0950] Warning notification function: Equipped with voice output and vibration functions to notify the user of warning messages received from the server.

[0951] 2. Server

[0952] Video receiving module: Used to receive video data transmitted from a portable device.

[0953] AI analysis module: Analyzes the received video data and uses deep learning models to detect objects in the video (vehicles, pedestrians, suspicious individuals, etc.).

[0954] Alert generation module: Based on the detection results and the user's emotional state, an alert message is generated and sent to the mobile device.

[0955] Operational Overview

[0956] Video acquisition and transmission

[0957] The device uses front and rear cameras to capture video in real time and transmits the data to a server at regular intervals, with the communication module playing a key role in this process.

[0958] Video analysis

[0959] The server receives the video data sent from the device and inputs it into an AI analysis module, which uses deep learning to detect objects in the video and track their movements.

[0960] Analyzing the user's emotional state

[0961] The device uses an emotion engine to analyze the user's facial expressions and tone of voice in real time to determine their emotional state.

[0962] Risk detection

[0963] The server assesses risks based on the AI ​​analysis results and the user's emotional state, such as detecting pedestrians running out into the street or suspicious people approaching.

[0964] Generate and send warning messages

[0965] The server generates an appropriate warning message depending on the detected risk and the user's emotional state and sends it to the mobile device.

[0966] Warning Notification

[0967] The device notifies the user of the warning message received from the server by voice or vibration. Specific notification content is voice synthesis such as "Danger! Watch out for people jumping out into the street!" or "There is a suspicious person behind you, so be careful."

[0968] Specific examples

[0969] Example of detecting elementary school students jumping out

[0970] 1. Device: When an elementary school student walks to school, the front camera captures images of the road.

[0971] 2. Terminal: The captured video data is sent to the server.

[0972] 3. Server: The received video data is analyzed by an AI analysis module, and any movement of the elementary school student attempting to jump out is detected.

[0973] 4. Device: The emotion engine analyzes the facial expressions of elementary school students and detects when they are not concentrating.

[0974] 5. Server: Generates a warning message saying "Watch out for people jumping out!" and sends it to the mobile device.

[0975] 6. Device: Elementary school students will receive a voice warning and vibration saying, "Danger! Watch out for people jumping out!"

[0976] 7. User: Elementary school students who hear the warning will stop and be able to prevent a traffic accident.

[0977] Example of detecting suspicious elderly people

[0978] 1. Device: When the elderly person is returning home, the rear camera captures the rear view.

[0979] 2. Terminal: The captured video data is sent to the server.

[0980] 3. Server: The received video data is analyzed by an AI analysis module to detect suspicious individuals approaching the elderly.

[0981] 4. Device: The emotion engine analyzes the facial expressions of the elderly person and detects their state of tension.

[0982] 5. Server: Generates a warning message saying "Beware of suspicious person!" and sends it to the mobile device.

[0983] 6. Terminal: The elderly person will receive a voice warning and vibration saying, "There is a suspicious person behind you, so please be careful."

[0984] 7. User: By receiving the warning, elderly people can prevent harm from suspicious people by paying attention to their surroundings.

[0985] Prompt Sentence Examples

[0986] Below are some example prompts to input to the generative AI model:

[0987] "If an elementary school student is walking and is about to run out into the road, how would the system detect the danger and what kind of warning would it give?"

[0988] An example of a generated description:

[0989] The system captures road images in real time using the mobile device's front camera and sends them to a server. An AI analysis module on the server analyzes the video data and detects any movements of the elementary school student attempting to run out into the street. An emotion engine then analyzes the student's emotional state, and if the student is not concentrating, it generates a stronger warning message, "Danger! Watch out for those running out into the street!", and sends it to the mobile device. Finally, the mobile device issues a voice warning message to alert the elementary school student to the danger.

[0990] The above is a specific embodiment for carrying out the present invention. This system is expected to significantly improve the safety of elementary school children and the elderly, and to prevent traffic accidents and robberies.

[0991] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0992] Step 1: Acquire footage

[0993] The device uses the front and rear cameras to capture video in real time, a process that occurs continuously at a constant frame rate.

[0994] Input: The surrounding physical environment.

[0995] Output: Front and rear video data.

[0996] What it does: The camera on your handheld device captures real-time video of what's in front of and behind you, even while you're moving.

[0997] Step 2: Sending video data

[0998] The device transmits the captured video data to the server at regular intervals using a communication module.

[0999] Input: Front and rear video data.

[1000] Output: Video data transferred to the server.

[1001] Specific operation: Video data is compressed and sent in packet format to the server every 5 seconds.

[1002] Step 3: Receiving video data

[1003] The server receives the video data sent from the terminal. This function is performed by the video receiving module.

[1004] Input: Video data sent from the device.

[1005] Output: Video data stored in the server's memory.

[1006] Specific operation: The server's video receiving module receives the data packets and expands them into memory for analysis.

[1007] Step 4: Analyzing the video data

[1008] The server inputs the received video data into an AI analysis module, which uses deep learning algorithms to detect objects in the video (vehicles, pedestrians, suspicious individuals, etc.) and track their movements.

[1009] Input: Video data stored on the server.

[1010] Output: Data about detected objects and their movements.

[1011] How it works: The AI ​​analysis module analyzes the video data frame by frame, identifies objects in each frame, and tracks their position and movement.

[1012] Step 5: Analyzing the user's emotional state

[1013] The device uses an emotion engine to analyze the user's facial expressions and tone of voice in real time to determine their emotional state.

[1014] Input: User facial and voice data from the camera and microphone on the handheld device.

[1015] Output: The user's emotional state (angry, surprised, relaxed, etc.).

[1016] How it works: The device's emotion engine captures the user's face and voice and analyzes them to determine their emotional state.

[1017] Step 6: Detect risks

[1018] The server combines the results of AI analysis with the user's emotional state to detect risks, such as pedestrians running out into the street or suspicious people approaching.

[1019] Input: Detection results from the AI ​​analysis module, and user emotional state data.

[1020] Output: Information about the detected risks.

[1021] Specific behavior: The server integrates the detected object movement with the user's emotional state to identify dangerous situations, such as when the user is about to run out into the road or when a suspicious person is approaching.

[1022] Step 7: Generate a warning message

[1023] The server generates appropriate warning messages based on the detected risks, which are tailored according to the user's emotional state.

[1024] Input: Information on detected risks, user emotional state data.

[1025] Output: A warning message.

[1026] Specific behavior: The server generates a soft tone warning if the user is nervous and a hard tone warning if the user is unaware.

[1027] Step 8: Sending a warning message

[1028] The server sends the generated warning message to the terminal.

[1029] Input: A warning message.

[1030] Output: The warning message sent to the terminal.

[1031] How it works: The server sends real-time alert messages to the device using a low-latency communication protocol.

[1032] Step 9: Notification of warnings

[1033] The terminal notifies the user of the warning message received from the server by voice or vibration.

[1034] Input: The warning message received from the server.

[1035] Output: The warning that was notified to the user.

[1036] What it does: The device's voice synthesis function warns the user, saying "Danger! Watch out for people jumping out!", while also providing physical feedback to the user using the device's vibration function.

[1037] As described above, each step works in conjunction with the others, allowing users to receive risk notifications in real time and receive appropriate warning messages according to their emotional state.

[1038] (Application example 2)

[1039] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1040] Conventional safety systems using portable devices focus only on risk detection and do not optimize warnings based on the user's emotional state. As a result, the effectiveness of warnings is limited, making it difficult to ensure sufficient safety, especially for users with large emotional fluctuations, such as elementary school children and the elderly. Furthermore, the generated warning messages are uniform, making it difficult to provide warnings appropriate to the situation and the user's state.

[1041] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing front and rear images using AI to detect pedestrians jumping out and suspicious people in the vicinity, means for providing the portable device with an emotion engine for analyzing the user's emotional state, means for generating a warning message based on the detection results and the user's emotional state, and means for optimizing the content of the warning message using a generative AI model. This makes it possible to generate and notify the user of an optimal warning message in real time according to their emotional state, thereby providing greater safety than conventional systems.

[1042] "Front camera and rear camera" refers to a photographing device attached to a mobile device for capturing images of the front and rear.

[1043] A "portable device" is a device that can be easily carried around and is equipped with a communication module and a warning notification function.

[1044] "AI analysis" means using artificial intelligence technology to analyze acquired video data and detect specific patterns and risks.

[1045] "Detecting pedestrians running out into the road and suspicious people in the surrounding area" means that AI analyzes video data to find pedestrians who are about to run out into the road or people behaving abnormally.

[1046] An "emotion engine" is an algorithm or software that analyzes information such as a user's facial expressions and tone of voice to determine their emotional state in real time.

[1047] "Generating a warning message" means creating appropriate warning content based on the detected risk and the user's emotional state.

[1048] "Notifying the user of a warning by voice or vibration" means that the generated warning message is conveyed to the user by using a voice output device or a vibration function.

[1049] A "generative AI model" is an artificial intelligence model that creates appropriate responses and outputs based on generated data and input information.

[1050] "Optimization" means making adjustments to obtain the most effective or efficient method or result under given conditions.

[1051] "Generating and notifying in real time" means analyzing data in response to the situation at hand and providing results and information immediately.

[1052] 1. System program generation

[1053] This system uses a portable device equipped with a front and rear camera and works in conjunction with a server to analyze both risk and emotional state. Users can carry the portable device and monitor the risks around them in real time. The system has the following functions:

[1054] 2. Program processing explanation

[1055] Video acquisition and transmission

[1056] The terminal captures video in real time using the front and rear cameras. This video data is sent to the server via a communication module. The terminal has a built-in communication module that supports sending and receiving video data.

[1057] Server-based video analysis

[1058] The server analyzes the received video data. Specifically, the AI ​​analysis module uses a deep learning model to detect objects in the video (e.g., vehicles, pedestrians, suspicious individuals) and track their movements. This makes it possible to detect risks such as pedestrians running out into the street or suspicious individuals approaching in real time.

[1059] Emotion analysis

[1060] The emotion engine installed in the device uses a camera and audio sensors to analyze the user's facial expressions and tone of voice to determine the user's emotional state. A deep learning model is used to analyze the emotional state.

[1061] Generating and notifying warning messages

[1062] The server generates an optimal warning message based on the risk detection results and the user's emotional state. The content of this warning message is optimized using a generative AI model. The generated warning message is then sent back to the device via the communication module.

[1063] The device notifies the user of the received warning message using a voice output device or vibration function. For example, it generates messages such as "Danger! Watch out for people jumping out into the street!" or "Be careful, there is a suspicious person behind you."

[1064] 3. Specific Examples

[1065] Detecting elementary school students jumping out

[1066] When elementary school students use the vehicle on their way home from school, the front camera captures images of the road in real time. If the server detects an approaching vehicle and determines that the child is not concentrating through emotional analysis, it issues a voice warning saying "Danger! Watch out for children running out into the street!" and simultaneously activates a vibration. This allows elementary school students to pay close attention to the warning and prevent them from running out into the street.

[1067] Detecting suspicious elderly people

[1068] When an elderly person is returning home, a rear camera captures video of the area behind them. If analysis on the server detects that a suspicious person is approaching from behind, and emotion analysis determines that the elderly person is nervous, a gentle voice warning is issued saying, "There is a suspicious person behind you, please be careful." This allows the elderly person to deal with the situation calmly and prevent harm from the suspicious person.

[1069] Prompt Sentence Examples

[1070] After-school risk detection prompts:

[1071] Design a system that uses an AI algorithm to detect in real time when a vehicle ahead is approaching, and issues an audio warning saying "Danger! Watch out for vehicles jumping out!" if the user's emotional state is not focused.

[1072] Hardware and Software

[1073] The server requires a data center component with high-performance computing resources and an AI framework (e.g., TensorFlow, PyTorch) for running deep learning models, while the device requires a high-resolution camera, an audio sensor, a communication module (e.g., Wi-Fi, 4G / 5G module), and an on-device AI framework (e.g., TensorFlow Lite) for running a sentiment analysis engine.

[1074] This enables the entire system to work together to detect risks in real time and deliver optimized warning messages to enhance user safety.

[1075] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1076] Step 1:

[1077] The terminal captures video in real time using a front camera and a rear camera. A user carries a portable device and periodically captures video of the front and rear. This allows the terminal to obtain video data of the current surrounding environment. The input is the real-time video data obtained from the camera, and the output is this video data.

[1078] Step 2:

[1079] The terminal transmits the acquired front and rear video data to the server via the communication module. The communication module in the device is activated and uploads the video data to the server via the network. At this stage, the input is the front and rear video data, and the output is the video data received by the server.

[1080] Step 3:

[1081] The server analyzes the received video data using a deep learning model. Specifically, the AI ​​analysis module detects objects in the video data and tracks their movements. The server evaluates risks, such as pedestrians running out into the street or suspicious people approaching, in real time. The input is the video data received by the server, and the output is the risk detection results.

[1082] Step 4:

[1083] The device uses an emotion engine to analyze the user's emotional state in real time. The device's camera and audio sensors capture the user's face and voice data, which are then analyzed by an emotion analysis algorithm. The input is the user's face and voice data, and the output is the analyzed emotional state.

[1084] Step 5:

[1085] The server uses a generative AI model to generate an optimal warning message based on the risk detection results and the user's emotional state. The server determines the appropriate warning message by taking into account the type of risk and the user's current emotional state. The input is the risk detection results and the user's emotional state, and the output is the warning message.

[1086] Step 6:

[1087] The terminal receives the generated warning message via the communication module. The terminal acquires the warning message sent from the server and notifies the user using audio output and vibration functions. The input is the warning message sent from the server, and the output is the warning message notified to the user.

[1088] This allows the entire system to work together to detect risks and provide emotion-based warning notifications to enhance user safety in real time.

[1089] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1090] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1091] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1092] [Third embodiment]

[1093] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1094] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[1095] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1096] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1097] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1098] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1099] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1100] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1101] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1102] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1103] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1104] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1105] MODE FOR CARRYING OUT THE INVENTION

[1106] This invention is a system to prevent traffic accidents and robberies that elementary school children and the elderly may encounter in their daily lives. Specifically, it uses a portable device equipped with a front and rear camera and utilizes AI technology to detect surrounding risks and issue a warning.

[1107] System configuration

[1108] This system consists of the following main components:

[1109] 1. Handheld devices

[1110] Front and rear cameras: These are necessary elements for capturing images in real time.

[1111] Communication module: Has the function to send captured images to a server.

[1112] Warning notification function: This module notifies the user of warning messages received from the server. Specifically, it has audio output and vibration functions.

[1113] 2. Server

[1114] Video receiving module: Has the function to receive video data transmitted from a portable device.

[1115] AI analysis module: Analyzes the received video data and executes algorithms to detect sudden jumps in and suspicious individuals.

[1116] Alert generation module: Based on the detection results, it generates an alert message and sends it to the mobile device.

[1117] System Operation

[1118] The operation of this system will be explained below.

[1119] 1. Video acquisition and transmission

[1120] Terminal: The mobile device captures video in real time using front and rear cameras and transmits the data to the server at regular intervals.

[1121] 2. Video Analysis

[1122] Server: The server receives the video data sent from the mobile device and inputs it into the AI ​​analysis module, which uses deep learning to detect objects in the video (e.g., vehicles, pedestrians, suspicious individuals).

[1123] 3. Risk detection and prediction

[1124] Server: Using AI analysis, it assesses the risk of pedestrians running out into the road and the approach of suspicious individuals. For example, it detects when a pedestrian is about to run out into the road or when a suspicious individual is approaching an elderly person.

[1125] 4. Generating and sending warning messages

[1126] Server: Automatically generates warning messages based on each risk and sends them to the mobile device.

[1127] 5. Warning Notification

[1128] Terminal: The terminal notifies the user of warning messages received from the server by voice or vibration. Specifically, it uses voice synthesis to issue warnings such as "Danger! Watch out for people jumping out into the street!" or "There's a suspicious person behind you, so be careful."

[1129] Specific examples

[1130] Detecting elementary school students jumping out

[1131] Device: As elementary school students walk to school, the front-facing camera captures footage of the road.

[1132] Terminal: Video data is sent to the server.

[1133] Server: The video received by the server is analyzed by AI, and any movement of the elementary school student trying to jump out is detected.

[1134] Server: Generates a "Watch out for people jumping out!" warning message and sends it to the handheld device.

[1135] Device: Elementary school students will hear a voice warning saying, "Danger! Watch out for people jumping out!"

[1136] User: Elementary school student listens to the warning and stops to prevent a traffic accident.

[1137] Detecting suspicious elderly people

[1138] Device: The rear camera captures rearward footage as the elderly person walks home.

[1139] Terminal: The acquired video data is sent to the server.

[1140] Server: The server analyzes the video and detects suspicious individuals approaching the elderly.

[1141] Server: Generates a "Beware of suspicious person!" warning message and sends it to the mobile device.

[1142] Terminal: A voice warning is issued to the elderly person saying, "There is a suspicious person behind you, so please be careful."

[1143] Users: Elderly people receive a warning and pay attention to their surroundings to prevent harm from suspicious people.

[1144] As described above, the present invention improves the safety of elementary school children and the elderly in their daily lives through a system that combines a portable device and a server. The real-time risk detection and warning functions provided by this system prevent traffic accidents and robberies.

[1145] The processing flow will be explained below.

[1146] Program processing flow

[1147] Video acquisition and transmission

[1148] Device:

[1149] 1. Step 1: Activate the front and rear cameras on your mobile device.

[1150] 2. Step 2: Capture video in real time using the front and rear cameras. Video frames are acquired at a rate of, for example, 30 frames per second.

[1151] 3. Step 3: Temporarily save the captured video data in memory.

[1152] 4. Step 4: At regular intervals (for example, every second), the video data stored in memory is divided into packets and organized.

[1153] 5. Step 5: The organized data packets are sent to the server via the communication module using HTTP or WebSocket.

[1154] Video analysis

[1155] server:

[1156] 1. Step 1: The server receives the front and rear camera image data.

[1157] 2. Step 2: The received video data is decoded and input into the AI ​​analysis module.

[1158] 3. Step 3: The AI ​​analysis module uses a deep learning model (e.g., YOLO, SSD) to detect objects (e.g., vehicles, pedestrians, suspicious people) in the video frame.

[1159] 4. Step 4: Track the movement of the detected object and analyze its position and velocity.

[1160] 5. Step 5: Evaluate the risk of pedestrians running out into the street or the approach of suspicious people. Specifically, predict whether certain movements pose a risk of traffic accidents or robberies.

[1161] Risk detection and warning

[1162] server:

[1163] 1. Step 1: Determine whether there is a risk based on the results of AI analysis. For example, evaluate whether a pedestrian is about to run out into the road or whether a suspicious person is approaching.

[1164] 2. Step 2: If a risk is detected, generate an appropriate warning message, which includes the warning type (e.g., watch out for a sudden jump, watch out for a suspicious person) and a recommended action (e.g., stop, turn around).

[1165] 3. Step 3: Send the generated warning message to the terminal.

[1166] Warning Notification

[1167] Device:

[1168] 1. Step 1: Decode the warning message received from the server.

[1169] 2. Step 2: The decoded warning message is notified to the user. Specifically, the voice synthesis module generates a warning voice and outputs it from the speaker or activates the vibration motor.

[1170] User Feedback

[1171] User:

[1172] 1. Step 1: Receive an alert notification from your device, for example, hear a sound notification or feel a vibration.

[1173] 2. Step 2: Follow the warning and take appropriate action. Specifically, if you receive a warning about people jumping out into the street, stop, and if you receive a warning about suspicious people, turn around and check your surroundings.

[1174] Specific examples

[1175] Detecting elementary school students jumping out

[1176] Device:

[1177] 1. Step 1: When an elementary school student is walking to school, the front camera captures images of the road.

[1178] 2. Step 2: The video data is sent to the server.

[1179] server:

[1180] 1. Step 1: The video received by the server is analyzed by AI, and any movement of the elementary school student trying to jump out is detected.

[1181] 2. Step 2: Generate a warning message "Watch out for people jumping out!" and send it to the mobile device.

[1182] Device:

[1183] 1. Step 1: Elementary school students will hear a voice warning saying, "Danger! Watch out for people jumping out!"

[1184] User:

[1185] 1. Step 1: Elementary school students listen to the warning and stop to prevent traffic accidents.

[1186] Detecting suspicious elderly people

[1187] Device:

[1188] 1. Step 1: When an elderly person is returning home, the rear camera captures the rear view.

[1189] 2. Step 2: The acquired video data is sent to the server.

[1190] server:

[1191] 1. Step 1: The server analyzes the video and detects suspicious individuals approaching the elderly person.

[1192] 2. Step 2: Generate a warning message saying "Beware of suspicious person!" and send it to the mobile device.

[1193] Device:

[1194] 1. Step 1: The elderly person receives a voice warning saying, "There is a suspicious person behind you, so please be careful."

[1195] User:

[1196] 1. Step 1: The elderly person receives a warning and pays attention to their surroundings to prevent harm from suspicious individuals.

[1197] Example 1

[1198] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1199] This invention relates to a system that prevents traffic accidents and robberies that elementary school students and the elderly may encounter in their daily lives. In particular, it aims to provide a system that can detect surrounding risks in real time and issue accurate warnings. Conventional systems sometimes experience delays in detecting risks and generating warning messages, making them insufficient to avoid actual danger. Therefore, the present invention aims to prevent traffic accidents and robberies by quickly and accurately detecting risks and issuing warnings to users.

[1200] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1201] In this invention, the server includes means for receiving the acquired video data with a video receiving module, analyzing the front and rear video using an AI analysis module, and detecting pedestrians running out into the road and suspicious people in the vicinity, means for generating a warning message with a warning generation module based on the detection results, and means for transmitting the warning message to the mobile terminal device via a communication module and notifying the user of the warning using a warning notification function. This makes it possible to detect risks quickly and accurately and notify the user of appropriate warnings in real time.

[1202] A "portable terminal device" is a device that has a front camera and a rear camera and can be carried by a user.

[1203] A "communication module" is a combination of hardware and software for transmitting data from a mobile terminal device to a server.

[1204] A "server" is a computing system that receives, analyzes, and processes data sent from a mobile terminal device.

[1205] The "video receiving module" is a module having a function for receiving video data transmitted from a portable terminal device at a server.

[1206] The "AI analysis module" is a module that uses artificial intelligence technology to analyze received video data and detect specific patterns and risks.

[1207] The "warning generation module" is a module for generating appropriate warning messages based on the detection results of the AI ​​analysis module.

[1208] The "warning notification function" is a function in which the portable terminal device notifies the user of a warning message by voice or vibration.

[1209] "Pedestrian running out" refers to a pedestrian unintentionally running out into a dangerous place such as the roadway.

[1210] A "suspicious individual" is an individual who deviates from normal patterns of behavior and is perceived as a threat to users.

[1211] A "deep learning model" is an artificial intelligence technology that learns complex patterns from large amounts of data and uses them for analysis.

[1212] This invention is a system for preventing traffic accidents and robberies in the daily lives of elementary school children and the elderly. This system is constructed by combining a portable terminal device equipped with a front camera and a rear camera with a server. The configuration and operation of this system are described in detail below.

[1213] System configuration

[1214] 1. Portable terminal device

[1215] Front and rear cameras:

[1216] The mobile terminal device is equipped with a camera that captures images of the front and rear in real time, allowing the user to monitor the situation around them in detail.

[1217] Communication Module:

[1218] The mobile terminal device includes a communication module for transmitting the acquired video data to a server. Specifically, a 4G LTE module is used.

[1219] Warning notification function:

[1220] The terminal is equipped with a voice output and vibration function as a warning notification function, which allows the terminal to convey the warning message received from the server to the user.

[1221] 2. Server

[1222] Video Receiving Module:

[1223] The server uses a video receiving module, specifically Apache Kafka, to receive video data sent from the mobile terminal device.

[1224] AI analysis module:

[1225] The server is equipped with an AI analysis module that analyzes the received video data and uses deep learning models (e.g., TensorFlow) to detect objects in the video (e.g., vehicles, pedestrians, suspicious individuals).

[1226] Warning generation modules:

[1227] The server includes a warning generation module for generating a warning message based on the detection result by the AI ​​analysis module, and the generated warning message is transmitted to the mobile terminal device again.

[1228] System Operation

[1229] Video acquisition and transmission

[1230] Device: Front and rear cameras capture images of the user's surroundings in real time. The image data is sent to a server at regular intervals.

[1231] Receiving and analyzing video data

[1232] Server: The video receiving module receives the video data and inputs it into the AI ​​analysis module. Analysis detects pedestrians running out into the street and suspicious people approaching.

[1233] Evaluating risks and generating warning messages

[1234] Server: Based on the results from the AI ​​analysis module, a risk assessment is performed and an appropriate warning message is generated in the warning generation module. For example, messages such as "Watch out for people jumping out!" or "Watch out for suspicious people!" are generated.

[1235] Warning message notification

[1236] Terminal: A mobile terminal device receives the warning message and notifies the user through voice or vibration. A warning message is generated using voice synthesis technology (e.g., Google Cloud Text-to-Speech) and output from the speaker.

[1237] Specific examples

[1238] Detecting elementary school students jumping out

[1239] Device: When an elementary school student walks to school, the front camera captures footage of the road.

[1240] Terminal: Video data is sent to the server.

[1241] Server: The video received by the server is analyzed by AI, and any movement of the elementary school student trying to jump out is detected.

[1242] Server: Generates a warning message saying "Watch out for people jumping out!" and sends it to the mobile terminal device.

[1243] Device: Elementary school students will hear a voice warning saying, "Danger! Watch out for people jumping out!"

[1244] User: Elementary school student listens to the warning and stops to prevent a traffic accident.

[1245] Detecting suspicious elderly people

[1246] Device: The rear camera captures rearward footage as the elderly person walks home.

[1247] Terminal: The acquired video data is sent to the server.

[1248] Server: The server analyzes the video and detects suspicious individuals approaching the elderly.

[1249] Server: Generates a warning message saying "Beware of suspicious person!" and sends it to the mobile terminal device.

[1250] Terminal: A voice warning is issued to the elderly person saying, "There is a suspicious person behind you, so please be careful."

[1251] Users: Elderly people receive a warning and pay attention to their surroundings to prevent harm from suspicious people.

[1252] Prompt Sentence Examples

[1253] "While elementary school students are walking to school, please capture road footage with a front camera, use AI to detect the risk of them running out into the street, and use a notification system to generate a warning message. Please also tell us details about the hardware and software required."

[1254] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1255] Step 1: Acquire footage

[1256] Terminal: The mobile terminal device captures video in real time using the front and rear cameras, capturing video at 30 frames per second and storing it in an internal buffer.

[1257] Input: Real-time video of the user's surroundings.

[1258] Output: Front and rear video data stored in internal buffers.

[1259] Step 2: Sending the video

[1260] Terminal: At regular intervals (e.g., every 5 seconds), the video data stored in the buffer is sent to the server via the communications module. Specifically, the data is encoded via the 4G LTE module and transferred to the server via TCP / IP protocol.

[1261] Input: Video data stored in the internal buffer.

[1262] Output: Video data sent to the server.

[1263] Step 3: Receiving video data

[1264] Server: The server-side video receiving module receives the video data sent from the mobile terminal device and stores it in a database. Specifically, it uses Apache Kafka to manage the data stream and store it appropriately.

[1265] Input: Video data transmitted from a mobile terminal device.

[1266] Output: Video data stored in a database.

[1267] Step 4: Analyzing the footage

[1268] Server: The received video data is input into the AI ​​analysis module and analyzed using a deep learning model (e.g., TensorFlow). Specifically, objects in the video (vehicles, pedestrians, suspicious individuals, etc.) are detected and their movements are tracked.

[1269] Input: Video data stored in a database.

[1270] Output: Analysis results (e.g., JSON format) containing information about detected objects.

[1271] Step 5: Risk detection and prediction

[1272] Server: Based on the results of the AI ​​analysis module, it evaluates the risk of pedestrians running out into the street and the approach of suspicious individuals. Specifically, it calculates the speed and distance of objects and executes an algorithm to assess the risk level.

[1273] Input: Analysis results of the AI ​​analysis module.

[1274] Output: Evaluated risk information (e.g., risk of people jumping out, suspicious person warning).

[1275] Step 6: Generate a warning message

[1276] Server: Based on the detected risks, the warning generation module generates warning messages. Specifically, a Python script is used to generate warning messages (e.g., "Watch out for people jumping out!", "Watch out for suspicious people!") as text and audio data.

[1277] Input: Assessed risk information.

[1278] Output: Warning message in text and audio format.

[1279] Step 7: Sending a warning message

[1280] Server: The server transmits the generated warning message to the mobile terminal device. Specifically, the server encodes the warning message via the communication module and transmits it.

[1281] Input: Warning message in text and audio format.

[1282] Output: The warning message sent to the handheld device.

[1283] Step 8: Notification of warnings

[1284] Terminal: The mobile terminal device notifies the user of the received warning message. Specifically, it generates the warning message using voice synthesis technology and outputs it from the speaker. It also issues a physical warning using the vibration function.

[1285] Input: The warning message received.

[1286] Output: Audio and vibration alert notifications.

[1287] (Application example 1)

[1288] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1289] Conventional safety monitoring systems in factories mainly use fixed cameras, making it difficult to effectively monitor the movements of dynamically moving robots and people in real time. Furthermore, advanced analysis technology is required to detect abnormal or suspicious behavior, and it is not possible to issue appropriate warnings immediately. This results in low accuracy in improving safety in factories, and there is a high possibility of delayed response in emergencies.

[1290] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1291] In this invention, the server includes a means for acquiring video from the front and rear, a means for analyzing the acquired video using AI to detect abnormal or suspicious behavior, and a means for generating a warning message based on the detection results and transmitting it to the portable device or robot. This makes it possible to detect abnormal or suspicious behavior with high accuracy in real time and to immediately issue warnings to robots and workers in the factory by voice and vibration.

[1292] "Front camera and rear camera" are imaging devices for capturing images in front of and behind the device.

[1293] A "portable device" is a device that can be carried around, has front and rear cameras, and has the ability to capture and transmit images.

[1294] A "server" is a computer system that receives and analyzes video sent from a portable device.

[1295] "AI analysis" is the process of using artificial intelligence technology to analyze video data and detect specific patterns or anomalies.

[1296] A "warning message" is a message that is generated to warn the user when an abnormality or risk is detected.

[1297] A "factory robot" is a mechanical device that operates autonomously or semi-autonomously in a factory to perform designated tasks.

[1298] An "audio output device" is a device for producing electronically generated speech.

[1299] The "vibration function" is a function that causes a device or robot to vibrate to provide a physical notification to the user.

[1300] "Real-time" refers to the fact that information acquisition, processing, and output of results are carried out almost simultaneously.

[1301] "Abnormal and suspicious behavior" refers to behavior that deviates from normal patterns of behavior or that appears suspicious.

[1302] A "deep learning model" is an artificial intelligence model that uses a multi-layer neural network to learn the characteristics of data and perform advanced predictions and classifications.

[1303] "Tracking" is a technology that continuously tracks a specific object or movement and evaluates its position and status.

[1304] The present invention provides a system for monitoring the safety of robots and workers in a factory, detecting abnormal or suspicious behavior in real time, and issuing a warning. A specific embodiment of this system will be described below.

[1305] Main components of the system

[1306] 1. Handheld devices

[1307] It is equipped with a front-facing camera and a rear-facing camera, which capture images in real time. It also has the function of transmitting the images to a server via a communication module and receiving warning messages.

[1308] 2. Server

[1309] It receives video data sent from a portable device, performs AI analysis, and generates a warning message based on the analysis results, which is then sent back to the portable device.

[1310] 3. Robots in factories

[1311] Equipped with front and rear cameras, it captures video footage to detect abnormal or suspicious activity, and receives warning messages and issues audio and vibration alerts.

[1312] Hardware and Software Used

[1313] Cameras: Front and rear facing cameras

[1314] Communication module: Wi-Fi or Bluetooth

[1315] Server: High-performance server equipped with GPU

[1316] AI analysis software: TensorFlow, PyTorch

[1317] Video capture software: OpenCV

[1318] Warning notification software: paho-mqtt (communication), pyttsx3 (voice synthesis)

[1319] System Operation

[1320] The server receives video data sent from the mobile devices and factory robots. It analyzes this data using a deep learning model to detect abnormal or suspicious behavior in real time. Based on the results, it generates a warning message and sends it back to the mobile devices and robots. The robots then immediately notify them of the received warning message by voice and vibration.

[1321] Specific examples

[1322] Abnormal behavior detection

[1323] 1. Terminal (handheld device): Robots in the factory capture front and rear images in real time and send them to the server.

[1324] 2. Server: Analyzes the received video using a deep learning model to detect abnormal behavior, such as when a robot deviates from its normal behavior pattern.

[1325] 3. Server: Generates a warning message saying "Danger! Caution! Watch the machine!" and sends it to the robots in the factory.

[1326] 4. Terminal (robot in factory): Notifies warnings by voice and vibration.

[1327] Suspicious Activity Detection

[1328] 1. Terminal (portable device): A robot moving around the factory captures rearward-facing images and sends them to a server.

[1329] 2. Server: AI analyzes the received video and detects suspicious movements, such as when a suspicious person approaches the robot.

[1330] 3. Server: Generates a warning message saying "Suspicious behavior detected!" and sends it to the robots in the factory.

[1331] 4. Terminal (robot in factory): Notifies warnings by voice and vibration.

[1332] Prompt Sentence Examples

[1333] Factory robots capture images from cameras in front and behind them in real time and use AI to monitor the safety of the work area. For example, if a person or machine behaves abnormally, it will automatically issue a voice warning saying, "Danger! Watch out for machine operation!" What will happen if an abnormality is detected?

[1334] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1335] Step 1:

[1336] The terminal (a robot in a factory) acquires real-time images from the front and rear. It uses a front camera and a rear camera to capture the images. The input obtained is raw image data.

[1337] Step 2:

[1338] The video data captured by the terminal (a robot in a factory) is sent to the server via a communication module. Using a communication protocol (e.g., MQTT or HTTP), the video data is divided into packets and sent to the server. The server then receives the video data as input.

[1339] Step 3:

[1340] The video data received by the server is input into an AI analysis module, which uses a deep learning model (e.g., TensorFlow or PyTorch) to recognize objects in the video and analyze their behavior. This outputs intermediate data (e.g., the position of each object and its movement vector) that can be used to identify abnormal or suspicious behavior.

[1341] Step 4:

[1342] The server performs risk assessment based on the results of AI analysis. It evaluates identified abnormal or suspicious behavior and runs an algorithm to calculate the risk level. This generates a warning message for any detected risks.

[1343] Step 5:

[1344] The server generates a warning message and sends it back to the terminal (the robot in the factory) via the communication module, which then receives the warning message.

[1345] Step 6:

[1346] Based on the warning message received by the terminal (a robot in the factory), a warning is issued by voice and vibration. The voice output device is used to issue a voice warning, and the vibration module is used to notify the warning by vibration. The final output is a real-time warning notification to the user.

[1347] Step 7:

[1348] The user (worker) receives a warning from the robot and responds immediately. The worker receives a warning via voice and vibration and responds to the surrounding risks.

[1349] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1350] MODE FOR CARRYING OUT THE INVENTION

[1351] This invention is a system to prevent traffic accidents and robberies that elementary school children and the elderly may encounter in their daily lives. Specifically, it uses a portable device equipped with a front and rear camera, and utilizes AI technology and an emotion engine to detect surrounding risks and issue warnings according to the user's emotional state.

[1352] System configuration

[1353] This system consists of the following main components:

[1354] 1. Handheld devices

[1355] Front and rear cameras: These are necessary elements for capturing images in real time.

[1356] Communication module: Has the function to send captured images to a server.

[1357] Emotion engine: Has the function to analyze the user's emotional state in real time.

[1358] Warning notification function: This module notifies the user of warning messages received from the server. Specifically, it has audio output and vibration functions.

[1359] 2. Server

[1360] Video receiving module: Has the function to receive video data transmitted from a portable device.

[1361] AI analysis module: Analyzes the received video data and executes algorithms to detect sudden jumps in and suspicious individuals.

[1362] Alert generation module: Based on the detection results, it generates an alert message and sends it to the mobile device.

[1363] System Operation

[1364] The operation of this system will be explained below.

[1365] 1. Video acquisition and transmission

[1366] Terminal: The mobile device captures video in real time using front and rear cameras and transmits the data to the server at regular intervals.

[1367] 2. Video Analysis

[1368] Server: The server receives the video data sent from the mobile device and inputs it into the AI ​​analysis module, which uses deep learning to detect objects in the video (e.g., vehicles, pedestrians, suspicious individuals).

[1369] 3. Analysis of the user's emotional state

[1370] Terminal: The emotion engine installed in the portable device uses a camera and audio sensors to analyze the user's facial expressions and tone of voice, and determine their emotional state in real time.

[1371] 4. Risk detection and warning

[1372] Server: Using AI analysis, it assesses the risk of pedestrians running out into the road and the approach of suspicious individuals. For example, it detects when a pedestrian is about to run out into the road or when a suspicious individual is approaching an elderly person.

[1373] Server: When a risk is detected, it generates an appropriate warning message according to the user's emotional state and sends it to the mobile device. Depending on the emotional state, it adjusts the content of the warning and the notification method.

[1374] 5. Warning Notification

[1375] Terminal: The terminal notifies the user of warning messages received from the server by voice or vibration. Specifically, it uses voice synthesis to issue warnings such as "Danger! Watch out for people jumping out into the street!" or "There's a suspicious person behind you, so be careful."

[1376] Specific examples

[1377] Detecting elementary school students jumping out

[1378] Device: As elementary school students walk to school, the front-facing camera captures footage of the road.

[1379] Terminal: Video data is sent to the server.

[1380] Server: The video received by the server is analyzed by AI, and any movement of the elementary school student trying to jump out is detected.

[1381] Device: The emotion engine analyzes the emotional state of elementary school students and generates stronger warning messages if they are not concentrating.

[1382] Server: Generates a "Watch out for people jumping out!" warning message and sends it to the handheld device.

[1383] Device: Elementary school students will hear a voice warning saying, "Danger! Watch out for people jumping out!"

[1384] User: Elementary school student listens to the warning and stops to prevent a traffic accident.

[1385] Detecting suspicious elderly people

[1386] Device: The rear camera captures rearward footage as the elderly person walks home.

[1387] Terminal: The acquired video data is sent to the server.

[1388] Server: The server analyzes the video and detects suspicious individuals approaching the elderly.

[1389] Device: Detects the emotional state of the elderly person and generates a soft-toned warning message if they are in a state of tension.

[1390] Server: Generates a "Beware of suspicious person!" warning message and sends it to the mobile device.

[1391] Terminal: A voice warning is issued to the elderly person saying, "There is a suspicious person behind you, so please be careful."

[1392] Users: Elderly people receive a warning and pay attention to their surroundings to prevent harm from suspicious people.

[1393] As described above, this invention improves the safety of elementary school children and the elderly in their daily lives through a system that combines a portable device and a server. The real-time risk detection and warning functions provided by this system prevent traffic accidents and robberies. Furthermore, by combining it with an emotion engine, it is possible to issue warnings according to the user's emotional state, further improving safety.

[1394] The processing flow will be explained below.

[1395] Program processing flow

[1396] Video acquisition and transmission

[1397] Device:

[1398] 1. Step 1: Activate the front and rear cameras on your mobile device.

[1399] 2. Step 2: Capture video in real time using the front and rear cameras. Video frames are acquired at a rate of, for example, 30 frames per second.

[1400] 3. Step 3: Temporarily save the captured video data in memory.

[1401] 4. Step 4: At regular intervals (for example, every second), the video data stored in memory is divided into packets and organized.

[1402] 5. Step 5: The organized data packets are sent to the server via the communication module using HTTP or WebSocket.

[1403] Video analysis

[1404] server:

[1405] 1. Step 1: The server receives the front and rear camera image data.

[1406] 2. Step 2: The received video data is decoded and input into the AI ​​analysis module.

[1407] 3. Step 3: The AI ​​analysis module uses a deep learning model (e.g., YOLO, SSD) to detect objects (e.g., vehicles, pedestrians, suspicious people) in the video frame.

[1408] 4. Step 4: Track the movement of the detected object and analyze its position and velocity.

[1409] 5. Step 5: Evaluate the risk of pedestrians running out into the street or the approach of suspicious people. Specifically, predict whether certain movements pose a risk of traffic accidents or robberies.

[1410] Analyzing the user's emotional state

[1411] Device:

[1412] 1. Step 1: Start the emotion engine on the mobile device.

[1413] 2. Step 2: Use a camera or audio sensor to capture the user's facial expressions and tone of voice.

[1414] 3. Step 3: Analyze the captured data in real time to determine the user's emotional state, for example, identifying states such as tension, anxiety, or concentration.

[1415] 4. Step 4: Send the analysis results to the server.

[1416] Risk detection and warning

[1417] server:

[1418] 1. Step 1: Determine whether there is a risk based on the results of AI analysis. For example, evaluate whether a pedestrian is about to run out into the road or whether a suspicious person is approaching.

[1419] 2. Step 2: If a risk is detected, a warning message is generated based on the analysis of the transmitted emotional state.

[1420] 3. Step 3: Adjust the content and notification method of the warning message according to the user's emotional state.

[1421] 4. Step 4: Send the generated warning message to the terminal.

[1422] Warning Notification

[1423] Device:

[1424] 1. Step 1: Decode the warning message received from the server.

[1425] 2. Step 2: The decoded warning message is notified to the user. Specifically, the voice synthesis module generates a warning voice and outputs it from the speaker or activates the vibration motor.

[1426] User Feedback

[1427] User:

[1428] 1. Step 1: Receive an alert notification from your device, for example, hear a sound notification or feel a vibration.

[1429] 2. Step 2: Follow the warning and take appropriate action. Specifically, if you receive a warning about people jumping out into the street, stop, and if you receive a warning about suspicious people, turn around and check your surroundings.

[1430] Specific examples

[1431] Detecting elementary school students jumping out

[1432] Device:

[1433] 1. Step 1: When an elementary school student is walking to school, the front camera captures images of the road.

[1434] 2. Step 2: The video data is sent to the server.

[1435] server:

[1436] 1. Step 1: The video received by the server is analyzed by AI, and any movement of the elementary school student trying to jump out is detected.

[1437] 2. Step 2: Receive data from the emotion engine to analyze the emotional state of elementary school students.

[1438] 3. Step 3: If the elementary school student is not concentrating, generate a stronger warning message.

[1439] 4. Step 4: Generate a warning message "Watch out for people jumping out!" and send it to the mobile device.

[1440] Device:

[1441] 1. Step 1: Elementary school students will hear a voice warning saying, "Danger! Watch out for people jumping out!"

[1442] User:

[1443] 1. Step 1: Elementary school students listen to the warning and stop to prevent traffic accidents.

[1444] Detecting suspicious elderly people

[1445] Device:

[1446] 1. Step 1: When an elderly person is returning home, the rear camera captures the rear view.

[1447] 2. Step 2: The acquired video data is sent to the server.

[1448] 3. Step 3: The emotion engine captures the emotional state of the elderly person and sends the analysis results to the server.

[1449] server:

[1450] 1. Step 1: The server analyzes the video and detects suspicious individuals approaching the elderly person.

[1451] 2. Step 2: Evaluate the elderly person's emotional state and generate a soft-toned warning message if they are in a tense state.

[1452] 3. Step 3: Generate a warning message saying "Beware of suspicious person!" and send it to the mobile device.

[1453] Device:

[1454] 1. Step 1: The elderly person receives a voice warning saying, "There is a suspicious person behind you, so please be careful."

[1455] User:

[1456] 1. Step 1: The elderly person receives a warning and pays attention to their surroundings to prevent harm from suspicious individuals.

[1457] Example 2

[1458] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1459] Traffic accidents and robberies pose significant threats, especially to elementary school children and the elderly. However, existing safety systems lack the ability to detect risks in real time or to issue warnings based on the user's emotional state. Therefore, a more effective system is needed to prevent these risks.

[1460] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1461] In this invention, the server includes means for analyzing front and rear images using AI to detect pedestrians jumping out and suspicious people in the vicinity, means for analyzing the emotional state of the user and generating a warning message based on the emotional state, and means for transmitting the warning message to the portable device and notifying the user by voice or vibration. This allows the user to receive risk notifications in real time and receive appropriate warning messages tailored to their emotional state.

[1462] "Front camera and rear camera" refers to a photographing device that is attached to a mobile device and captures images of the front and rear in real time.

[1463] A "portable device" is an electronic device that can be carried by a user and has functions such as a camera, a communication module, and an emotion engine.

[1464] A "communication module" is a device that provides wireless communication functionality for transmitting video data from a portable device to a server.

[1465] The "server" is a central processing unit that receives video data sent from the portable device, analyzes it to detect risks, and generates and sends warning messages.

[1466] "AI analysis" is the process of using artificial intelligence technology on a server to analyze video data to detect objects and assess risks.

[1467] A "warning message" is a notification sent to a portable device to alert a user to a danger, and is presented by sound or vibration.

[1468] An "emotion engine" is software or hardware that analyzes a user's facial expressions and tone of voice to determine their emotional state in real time.

[1469] "Speech synthesis" is a technology that generates natural-sounding speech from text.

[1470] "Vibrate" is a feature that allows a mobile device to generate physical vibrations to provide notifications to the user.

[1471] A "deep learning model" is a multi-layered neural network used to learn complex patterns from data.

[1472] "Object detection" is the process of recognizing specific objects in video data and identifying their type and location.

[1473] "Tracking" is the technique of following the movement of a detected object over time and continuously monitoring its behavior.

[1474] This invention is a system for preventing traffic accidents and robberies that elementary school students and the elderly encounter in their daily lives. The system uses a portable device and a server to analyze video data and the user's emotional state, providing real-time risk detection and warnings.

[1475] System configuration

[1476] The system consists of the following main components:

[1477] 1. Handheld devices

[1478] Front and rear cameras: Used to capture footage in real time.

[1479] Communication module: Used to send acquired video data to the server.

[1480] Emotion engine: Used to analyze the user's facial expressions and tone of voice to determine their emotional state in real time.

[1481] Warning notification function: Equipped with voice output and vibration functions to notify the user of warning messages received from the server.

[1482] 2. Server

[1483] Video receiving module: Used to receive video data transmitted from a portable device.

[1484] AI analysis module: Analyzes the received video data and uses deep learning models to detect objects in the video (vehicles, pedestrians, suspicious individuals, etc.).

[1485] Alert generation module: Based on the detection results and the user's emotional state, an alert message is generated and sent to the mobile device.

[1486] Operational Overview

[1487] Video acquisition and transmission

[1488] The device uses front and rear cameras to capture video in real time and transmits the data to a server at regular intervals, with the communication module playing a key role in this process.

[1489] Video analysis

[1490] The server receives the video data sent from the device and inputs it into an AI analysis module, which uses deep learning to detect objects in the video and track their movements.

[1491] Analyzing the user's emotional state

[1492] The device uses an emotion engine to analyze the user's facial expressions and tone of voice in real time to determine their emotional state.

[1493] Risk detection

[1494] The server assesses risks based on the AI ​​analysis results and the user's emotional state, such as detecting pedestrians running out into the street or suspicious people approaching.

[1495] Generate and send warning messages

[1496] The server generates an appropriate warning message depending on the detected risk and the user's emotional state and sends it to the mobile device.

[1497] Warning Notification

[1498] The device notifies the user of the warning message received from the server by voice or vibration. Specific notification content is voice synthesis such as "Danger! Watch out for people jumping out into the street!" or "There is a suspicious person behind you, so be careful."

[1499] Specific examples

[1500] Example of detecting elementary school students jumping out

[1501] 1. Device: When an elementary school student walks to school, the front camera captures images of the road.

[1502] 2. Terminal: The captured video data is sent to the server.

[1503] 3. Server: The received video data is analyzed by an AI analysis module, and any movement of the elementary school student attempting to jump out is detected.

[1504] 4. Device: The emotion engine analyzes the facial expressions of elementary school students and detects when they are not concentrating.

[1505] 5. Server: Generates a warning message saying "Watch out for people jumping out!" and sends it to the mobile device.

[1506] 6. Device: Elementary school students will receive a voice warning and vibration saying, "Danger! Watch out for people jumping out!"

[1507] 7. User: Elementary school students who hear the warning will stop and be able to prevent a traffic accident.

[1508] Example of detecting suspicious elderly people

[1509] 1. Device: When the elderly person is returning home, the rear camera captures the rear view.

[1510] 2. Terminal: The captured video data is sent to the server.

[1511] 3. Server: The received video data is analyzed by an AI analysis module to detect suspicious individuals approaching the elderly.

[1512] 4. Device: The emotion engine analyzes the facial expressions of the elderly person and detects their state of tension.

[1513] 5. Server: Generates a warning message saying "Beware of suspicious person!" and sends it to the mobile device.

[1514] 6. Terminal: The elderly person will receive a voice warning and vibration saying, "There is a suspicious person behind you, so please be careful."

[1515] 7. User: By receiving the warning, elderly people can prevent harm from suspicious people by paying attention to their surroundings.

[1516] Prompt Sentence Examples

[1517] Below are some example prompts to input to the generative AI model:

[1518] "If an elementary school student is walking and is about to run out into the road, how would the system detect the danger and what kind of warning would it give?"

[1519] An example of a generated description:

[1520] The system captures road images in real time using the mobile device's front camera and sends them to a server. An AI analysis module on the server analyzes the video data and detects any movements of the elementary school student attempting to run out into the street. An emotion engine then analyzes the student's emotional state, and if the student is not concentrating, it generates a stronger warning message, "Danger! Watch out for those running out into the street!", and sends it to the mobile device. Finally, the mobile device issues a voice warning message to alert the elementary school student to the danger.

[1521] The above is a specific embodiment for carrying out the present invention. This system is expected to significantly improve the safety of elementary school children and the elderly, and to prevent traffic accidents and robberies.

[1522] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1523] Step 1: Acquire footage

[1524] The device uses the front and rear cameras to capture video in real time, a process that occurs continuously at a constant frame rate.

[1525] Input: The surrounding physical environment.

[1526] Output: Front and rear video data.

[1527] What it does: The camera on your handheld device captures real-time video of what's in front of and behind you, even while you're moving.

[1528] Step 2: Sending video data

[1529] The device transmits the captured video data to the server at regular intervals using a communication module.

[1530] Input: Front and rear video data.

[1531] Output: Video data transferred to the server.

[1532] Specific operation: Video data is compressed and sent in packet format to the server every 5 seconds.

[1533] Step 3: Receiving video data

[1534] The server receives the video data sent from the terminal. This function is performed by the video receiving module.

[1535] Input: Video data sent from the device.

[1536] Output: Video data stored in the server's memory.

[1537] Specific operation: The server's video receiving module receives the data packets and expands them into memory for analysis.

[1538] Step 4: Analyzing the video data

[1539] The server inputs the received video data into an AI analysis module, which uses deep learning algorithms to detect objects in the video (vehicles, pedestrians, suspicious individuals, etc.) and track their movements.

[1540] Input: Video data stored on the server.

[1541] Output: Data about detected objects and their movements.

[1542] How it works: The AI ​​analysis module analyzes the video data frame by frame, identifies objects in each frame, and tracks their position and movement.

[1543] Step 5: Analyzing the user's emotional state

[1544] The device uses an emotion engine to analyze the user's facial expressions and tone of voice in real time to determine their emotional state.

[1545] Input: User facial and voice data from the camera and microphone on the handheld device.

[1546] Output: The user's emotional state (angry, surprised, relaxed, etc.).

[1547] How it works: The device's emotion engine captures the user's face and voice and analyzes them to determine their emotional state.

[1548] Step 6: Detect risks

[1549] The server combines the results of AI analysis with the user's emotional state to detect risks, such as pedestrians running out into the street or suspicious people approaching.

[1550] Input: Detection results from the AI ​​analysis module, and user emotional state data.

[1551] Output: Information about the detected risks.

[1552] Specific behavior: The server integrates the detected object movement with the user's emotional state to identify dangerous situations, such as when the user is about to run out into the road or when a suspicious person is approaching.

[1553] Step 7: Generate a warning message

[1554] The server generates appropriate warning messages based on the detected risks, which are tailored according to the user's emotional state.

[1555] Input: Information on detected risks, user emotional state data.

[1556] Output: A warning message.

[1557] Specific behavior: The server generates a soft tone warning if the user is nervous and a hard tone warning if the user is unaware.

[1558] Step 8: Sending a warning message

[1559] The server sends the generated warning message to the terminal.

[1560] Input: A warning message.

[1561] Output: The warning message sent to the terminal.

[1562] How it works: The server sends real-time alert messages to the device using a low-latency communication protocol.

[1563] Step 9: Notification of warnings

[1564] The terminal notifies the user of the warning message received from the server by voice or vibration.

[1565] Input: The warning message received from the server.

[1566] Output: The warning that was notified to the user.

[1567] What it does: The device's voice synthesis function warns the user, saying "Danger! Watch out for people jumping out!", while also providing physical feedback to the user using the device's vibration function.

[1568] As described above, each step works in conjunction with the others, allowing users to receive risk notifications in real time and receive appropriate warning messages according to their emotional state.

[1569] (Application example 2)

[1570] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1571] Conventional safety systems using portable devices focus only on risk detection and do not optimize warnings based on the user's emotional state. As a result, the effectiveness of warnings is limited, making it difficult to ensure sufficient safety, especially for users with large emotional fluctuations, such as elementary school children and the elderly. Furthermore, the generated warning messages are uniform, making it difficult to provide warnings appropriate to the situation and the user's state.

[1572] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing front and rear images using AI to detect pedestrians jumping out and suspicious people in the vicinity, means for providing the portable device with an emotion engine for analyzing the user's emotional state, means for generating a warning message based on the detection results and the user's emotional state, and means for optimizing the content of the warning message using a generative AI model. This makes it possible to generate and notify the user of an optimal warning message in real time according to their emotional state, thereby providing greater safety than conventional systems.

[1573] "Front camera and rear camera" refers to a photographing device attached to a mobile device for capturing images of the front and rear.

[1574] A "portable device" is a device that can be easily carried around and is equipped with a communication module and a warning notification function.

[1575] "AI analysis" means using artificial intelligence technology to analyze acquired video data and detect specific patterns and risks.

[1576] "Detecting pedestrians running out into the road and suspicious people in the surrounding area" means that AI analyzes video data to find pedestrians who are about to run out into the road or people behaving abnormally.

[1577] An "emotion engine" is an algorithm or software that analyzes information such as a user's facial expressions and tone of voice to determine their emotional state in real time.

[1578] "Generating a warning message" means creating appropriate warning content based on the detected risk and the user's emotional state.

[1579] "Notifying the user of a warning by voice or vibration" means that the generated warning message is conveyed to the user by using a voice output device or a vibration function.

[1580] A "generative AI model" is an artificial intelligence model that creates appropriate responses and outputs based on generated data and input information.

[1581] "Optimization" means making adjustments to obtain the most effective or efficient method or result under given conditions.

[1582] "Generating and notifying in real time" means analyzing data in response to the situation at hand and providing results and information immediately.

[1583] 1. System program generation

[1584] This system uses a portable device equipped with a front and rear camera and works in conjunction with a server to analyze both risk and emotional state. Users can carry the portable device and monitor the risks around them in real time. The system has the following functions:

[1585] 2. Program processing explanation

[1586] Video acquisition and transmission

[1587] The terminal captures video in real time using the front and rear cameras. This video data is sent to the server via a communication module. The terminal has a built-in communication module that supports sending and receiving video data.

[1588] Server-based video analysis

[1589] The server analyzes the received video data. Specifically, the AI ​​analysis module uses a deep learning model to detect objects in the video (e.g., vehicles, pedestrians, suspicious individuals) and track their movements. This makes it possible to detect risks such as pedestrians running out into the street or suspicious individuals approaching in real time.

[1590] Emotion analysis

[1591] The emotion engine installed in the device uses a camera and audio sensors to analyze the user's facial expressions and tone of voice to determine the user's emotional state. A deep learning model is used to analyze the emotional state.

[1592] Generating and notifying warning messages

[1593] The server generates an optimal warning message based on the risk detection results and the user's emotional state. The content of this warning message is optimized using a generative AI model. The generated warning message is then sent back to the device via the communication module.

[1594] The device notifies the user of the received warning message using a voice output device or vibration function. For example, it generates messages such as "Danger! Watch out for people jumping out into the street!" or "Be careful, there is a suspicious person behind you."

[1595] 3. Specific Examples

[1596] Detecting elementary school students jumping out

[1597] When elementary school students use the vehicle on their way home from school, the front camera captures images of the road in real time. If the server detects an approaching vehicle and determines that the child is not concentrating through emotional analysis, it issues a voice warning saying "Danger! Watch out for children running out into the street!" and simultaneously activates a vibration. This allows elementary school students to pay close attention to the warning and prevent them from running out into the street.

[1598] Detecting suspicious elderly people

[1599] When an elderly person is returning home, a rear camera captures video of the area behind them. If analysis on the server detects that a suspicious person is approaching from behind, and emotion analysis determines that the elderly person is nervous, a gentle voice warning is issued saying, "There is a suspicious person behind you, please be careful." This allows the elderly person to deal with the situation calmly and prevent harm from the suspicious person.

[1600] Prompt Sentence Examples

[1601] After-school risk detection prompts:

[1602] Design a system that uses an AI algorithm to detect in real time when a vehicle ahead is approaching, and issues an audio warning saying "Danger! Watch out for vehicles jumping out!" if the user's emotional state is not focused.

[1603] Hardware and Software

[1604] The server requires a data center component with high-performance computing resources and an AI framework (e.g., TensorFlow, PyTorch) for running deep learning models, while the device requires a high-resolution camera, an audio sensor, a communication module (e.g., Wi-Fi, 4G / 5G module), and an on-device AI framework (e.g., TensorFlow Lite) for running a sentiment analysis engine.

[1605] This enables the entire system to work together to detect risks in real time and deliver optimized warning messages to enhance user safety.

[1606] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1607] Step 1:

[1608] The terminal captures video in real time using a front camera and a rear camera. A user carries a portable device and periodically captures video of the front and rear. This allows the terminal to obtain video data of the current surrounding environment. The input is the real-time video data obtained from the camera, and the output is this video data.

[1609] Step 2:

[1610] The terminal transmits the acquired front and rear video data to the server via the communication module. The communication module in the device is activated and uploads the video data to the server via the network. At this stage, the input is the front and rear video data, and the output is the video data received by the server.

[1611] Step 3:

[1612] The server analyzes the received video data using a deep learning model. Specifically, the AI ​​analysis module detects objects in the video data and tracks their movements. The server evaluates risks, such as pedestrians running out into the street or suspicious people approaching, in real time. The input is the video data received by the server, and the output is the risk detection results.

[1613] Step 4:

[1614] The device uses an emotion engine to analyze the user's emotional state in real time. The device's camera and audio sensors capture the user's face and voice data, which are then analyzed by an emotion analysis algorithm. The input is the user's face and voice data, and the output is the analyzed emotional state.

[1615] Step 5:

[1616] The server uses a generative AI model to generate an optimal warning message based on the risk detection results and the user's emotional state. The server determines the appropriate warning message by taking into account the type of risk and the user's current emotional state. The input is the risk detection results and the user's emotional state, and the output is the warning message.

[1617] Step 6:

[1618] The terminal receives the generated warning message via the communication module. The terminal acquires the warning message sent from the server and notifies the user using audio output and vibration functions. The input is the warning message sent from the server, and the output is the warning message notified to the user.

[1619] This allows the entire system to work together to detect risks and provide emotion-based warning notifications to enhance user safety in real time.

[1620] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1621] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1622] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1623] [Fourth embodiment]

[1624] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1625] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1626] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1627] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1628] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1629] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1630] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1631] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1632] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1633] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1634] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1635] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1636] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1637] MODE FOR CARRYING OUT THE INVENTION

[1638] This invention is a system to prevent traffic accidents and robberies that elementary school children and the elderly may encounter in their daily lives. Specifically, it uses a portable device equipped with a front and rear camera and utilizes AI technology to detect surrounding risks and issue a warning.

[1639] System configuration

[1640] This system consists of the following main components:

[1641] 1. Handheld devices

[1642] Front and rear cameras: These are necessary elements for capturing images in real time.

[1643] Communication module: Has the function to send captured images to a server.

[1644] Warning notification function: This module notifies the user of warning messages received from the server. Specifically, it has audio output and vibration functions.

[1645] 2. Server

[1646] Video receiving module: Has the function to receive video data transmitted from a portable device.

[1647] AI analysis module: Analyzes the received video data and executes algorithms to detect sudden jumps in and suspicious individuals.

[1648] Alert generation module: Based on the detection results, it generates an alert message and sends it to the mobile device.

[1649] System Operation

[1650] The operation of this system will be explained below.

[1651] 1. Video acquisition and transmission

[1652] Terminal: The mobile device captures video in real time using front and rear cameras and transmits the data to the server at regular intervals.

[1653] 2. Video Analysis

[1654] Server: The server receives the video data sent from the mobile device and inputs it into the AI ​​analysis module, which uses deep learning to detect objects in the video (e.g., vehicles, pedestrians, suspicious individuals).

[1655] 3. Risk detection and prediction

[1656] Server: Using AI analysis, it assesses the risk of pedestrians running out into the road and the approach of suspicious individuals. For example, it detects when a pedestrian is about to run out into the road or when a suspicious individual is approaching an elderly person.

[1657] 4. Generating and sending warning messages

[1658] Server: Automatically generates warning messages based on each risk and sends them to the mobile device.

[1659] 5. Warning Notification

[1660] Terminal: The terminal notifies the user of warning messages received from the server by voice or vibration. Specifically, it uses voice synthesis to issue warnings such as "Danger! Watch out for people jumping out into the street!" or "There's a suspicious person behind you, so be careful."

[1661] Specific examples

[1662] Detecting elementary school students jumping out

[1663] Device: As elementary school students walk to school, the front-facing camera captures footage of the road.

[1664] Terminal: Video data is sent to the server.

[1665] Server: The video received by the server is analyzed by AI, and any movement of the elementary school student trying to jump out is detected.

[1666] Server: Generates a "Watch out for people jumping out!" warning message and sends it to the handheld device.

[1667] Device: Elementary school students will hear a voice warning saying, "Danger! Watch out for people jumping out!"

[1668] User: Elementary school student listens to the warning and stops to prevent a traffic accident.

[1669] Detecting suspicious elderly people

[1670] Device: The rear camera captures rearward footage as the elderly person walks home.

[1671] Terminal: The acquired video data is sent to the server.

[1672] Server: The server analyzes the video and detects suspicious individuals approaching the elderly.

[1673] Server: Generates a "Beware of suspicious person!" warning message and sends it to the mobile device.

[1674] Terminal: A voice warning is issued to the elderly person saying, "There is a suspicious person behind you, so please be careful."

[1675] Users: Elderly people receive a warning and pay attention to their surroundings to prevent harm from suspicious people.

[1676] As described above, the present invention improves the safety of elementary school children and the elderly in their daily lives through a system that combines a portable device and a server. The real-time risk detection and warning functions provided by this system prevent traffic accidents and robberies.

[1677] The processing flow will be explained below.

[1678] Program processing flow

[1679] Video acquisition and transmission

[1680] Device:

[1681] 1. Step 1: Activate the front and rear cameras on your mobile device.

[1682] 2. Step 2: Capture video in real time using the front and rear cameras. Video frames are acquired at a rate of, for example, 30 frames per second.

[1683] 3. Step 3: Temporarily save the captured video data in memory.

[1684] 4. Step 4: At regular intervals (for example, every second), the video data stored in memory is divided into packets and organized.

[1685] 5. Step 5: The organized data packets are sent to the server via the communication module using HTTP or WebSocket.

[1686] Video analysis

[1687] server:

[1688] 1. Step 1: The server receives the front and rear camera image data.

[1689] 2. Step 2: The received video data is decoded and input into the AI ​​analysis module.

[1690] 3. Step 3: The AI ​​analysis module uses a deep learning model (e.g., YOLO, SSD) to detect objects (e.g., vehicles, pedestrians, suspicious people) in the video frame.

[1691] 4. Step 4: Track the movement of the detected object and analyze its position and velocity.

[1692] 5. Step 5: Evaluate the risk of pedestrians running out into the street or the approach of suspicious people. Specifically, predict whether certain movements pose a risk of traffic accidents or robberies.

[1693] Risk detection and warning

[1694] server:

[1695] 1. Step 1: Determine whether there is a risk based on the results of AI analysis. For example, evaluate whether a pedestrian is about to run out into the road or whether a suspicious person is approaching.

[1696] 2. Step 2: If a risk is detected, generate an appropriate warning message, which includes the warning type (e.g., watch out for a sudden jump, watch out for a suspicious person) and a recommended action (e.g., stop, turn around).

[1697] 3. Step 3: Send the generated warning message to the terminal.

[1698] Warning Notification

[1699] Device:

[1700] 1. Step 1: Decode the warning message received from the server.

[1701] 2. Step 2: The decoded warning message is notified to the user. Specifically, the voice synthesis module generates a warning voice and outputs it from the speaker or activates the vibration motor.

[1702] User Feedback

[1703] User:

[1704] 1. Step 1: Receive an alert notification from your device, for example, hear a sound notification or feel a vibration.

[1705] 2. Step 2: Follow the warning and take appropriate action. Specifically, if you receive a warning about people jumping out into the street, stop, and if you receive a warning about suspicious people, turn around and check your surroundings.

[1706] Specific examples

[1707] Detecting elementary school students jumping out

[1708] Device:

[1709] 1. Step 1: When an elementary school student is walking to school, the front camera captures images of the road.

[1710] 2. Step 2: The video data is sent to the server.

[1711] server:

[1712] 1. Step 1: The video received by the server is analyzed by AI, and any movement of the elementary school student trying to jump out is detected.

[1713] 2. Step 2: Generate a warning message "Watch out for people jumping out!" and send it to the mobile device.

[1714] Device:

[1715] 1. Step 1: Elementary school students will hear a voice warning saying, "Danger! Watch out for people jumping out!"

[1716] User:

[1717] 1. Step 1: Elementary school students listen to the warning and stop to prevent traffic accidents.

[1718] Detecting suspicious elderly people

[1719] Device:

[1720] 1. Step 1: When an elderly person is returning home, the rear camera captures the rear view.

[1721] 2. Step 2: The acquired video data is sent to the server.

[1722] server:

[1723] 1. Step 1: The server analyzes the video and detects suspicious individuals approaching the elderly person.

[1724] 2. Step 2: Generate a warning message saying "Beware of suspicious person!" and send it to the mobile device.

[1725] Device:

[1726] 1. Step 1: The elderly person receives a voice warning saying, "There is a suspicious person behind you, so please be careful."

[1727] User:

[1728] 1. Step 1: The elderly person receives a warning and pays attention to their surroundings to prevent harm from suspicious individuals.

[1729] Example 1

[1730] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1731] This invention relates to a system that prevents traffic accidents and robberies that elementary school students and the elderly may encounter in their daily lives. In particular, it aims to provide a system that can detect surrounding risks in real time and issue accurate warnings. Conventional systems sometimes experience delays in detecting risks and generating warning messages, making them insufficient to avoid actual danger. Therefore, the present invention aims to prevent traffic accidents and robberies by quickly and accurately detecting risks and issuing warnings to users.

[1732] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1733] In this invention, the server includes means for receiving the acquired video data with a video receiving module, analyzing the front and rear video using an AI analysis module, and detecting pedestrians running out into the road and suspicious people in the vicinity, means for generating a warning message with a warning generation module based on the detection results, and means for transmitting the warning message to the mobile terminal device via a communication module and notifying the user of the warning using a warning notification function. This makes it possible to detect risks quickly and accurately and notify the user of appropriate warnings in real time.

[1734] A "portable terminal device" is a device that has a front camera and a rear camera and can be carried by a user.

[1735] A "communication module" is a combination of hardware and software for transmitting data from a mobile terminal device to a server.

[1736] A "server" is a computing system that receives, analyzes, and processes data sent from a mobile terminal device.

[1737] The "video receiving module" is a module having a function for receiving video data transmitted from a portable terminal device at a server.

[1738] The "AI analysis module" is a module that uses artificial intelligence technology to analyze received video data and detect specific patterns and risks.

[1739] The "warning generation module" is a module for generating appropriate warning messages based on the detection results of the AI ​​analysis module.

[1740] The "warning notification function" is a function in which the portable terminal device notifies the user of a warning message by voice or vibration.

[1741] "Pedestrian running out" refers to a pedestrian unintentionally running out into a dangerous place such as the roadway.

[1742] A "suspicious individual" is an individual who deviates from normal patterns of behavior and is perceived as a threat to users.

[1743] A "deep learning model" is an artificial intelligence technology that learns complex patterns from large amounts of data and uses them for analysis.

[1744] This invention is a system for preventing traffic accidents and robberies in the daily lives of elementary school children and the elderly. This system is constructed by combining a portable terminal device equipped with a front camera and a rear camera with a server. The configuration and operation of this system are described in detail below.

[1745] System configuration

[1746] 1. Portable terminal device

[1747] Front and rear cameras:

[1748] The mobile terminal device is equipped with a camera that captures images of the front and rear in real time, allowing the user to monitor the situation around them in detail.

[1749] Communication Module:

[1750] The mobile terminal device includes a communication module for transmitting the acquired video data to a server. Specifically, a 4G LTE module is used.

[1751] Warning notification function:

[1752] The terminal is equipped with a voice output and vibration function as a warning notification function, which allows the terminal to convey the warning message received from the server to the user.

[1753] 2. Server

[1754] Video Receiving Module:

[1755] The server uses a video receiving module, specifically Apache Kafka, to receive video data sent from the mobile terminal device.

[1756] AI analysis module:

[1757] The server is equipped with an AI analysis module that analyzes the received video data and uses deep learning models (e.g., TensorFlow) to detect objects in the video (e.g., vehicles, pedestrians, suspicious individuals).

[1758] Warning generation modules:

[1759] The server includes a warning generation module for generating a warning message based on the detection result by the AI ​​analysis module, and the generated warning message is transmitted to the mobile terminal device again.

[1760] System Operation

[1761] Video acquisition and transmission

[1762] Device: Front and rear cameras capture images of the user's surroundings in real time. The image data is sent to a server at regular intervals.

[1763] Receiving and analyzing video data

[1764] Server: The video receiving module receives the video data and inputs it into the AI ​​analysis module. Analysis detects pedestrians running out into the street and suspicious people approaching.

[1765] Evaluating risks and generating warning messages

[1766] Server: Based on the results from the AI ​​analysis module, a risk assessment is performed and an appropriate warning message is generated in the warning generation module. For example, messages such as "Watch out for people jumping out!" or "Watch out for suspicious people!" are generated.

[1767] Warning message notification

[1768] Terminal: A mobile terminal device receives the warning message and notifies the user through voice or vibration. A warning message is generated using voice synthesis technology (e.g., Google Cloud Text-to-Speech) and output from the speaker.

[1769] Specific examples

[1770] Detecting elementary school students jumping out

[1771] Device: When an elementary school student walks to school, the front camera captures footage of the road.

[1772] Terminal: Video data is sent to the server.

[1773] Server: The video received by the server is analyzed by AI, and any movement of the elementary school student trying to jump out is detected.

[1774] Server: Generates a warning message saying "Watch out for people jumping out!" and sends it to the mobile terminal device.

[1775] Device: Elementary school students will hear a voice warning saying, "Danger! Watch out for people jumping out!"

[1776] User: Elementary school student listens to the warning and stops to prevent a traffic accident.

[1777] Detecting suspicious elderly people

[1778] Device: The rear camera captures rearward footage as the elderly person walks home.

[1779] Terminal: The acquired video data is sent to the server.

[1780] Server: The server analyzes the video and detects suspicious individuals approaching the elderly.

[1781] Server: Generates a warning message saying "Beware of suspicious person!" and sends it to the mobile terminal device.

[1782] Terminal: A voice warning is issued to the elderly person saying, "There is a suspicious person behind you, so please be careful."

[1783] Users: Elderly people receive a warning and pay attention to their surroundings to prevent harm from suspicious people.

[1784] Prompt Sentence Examples

[1785] "While elementary school students are walking to school, please capture road footage with a front camera, use AI to detect the risk of them running out into the street, and use a notification system to generate a warning message. Please also tell us details about the hardware and software required."

[1786] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1787] Step 1: Acquire footage

[1788] Terminal: The mobile terminal device captures video in real time using the front and rear cameras, capturing video at 30 frames per second and storing it in an internal buffer.

[1789] Input: Real-time video of the user's surroundings.

[1790] Output: Front and rear video data stored in internal buffers.

[1791] Step 2: Sending the video

[1792] Terminal: At regular intervals (e.g., every 5 seconds), the video data stored in the buffer is sent to the server via the communications module. Specifically, the data is encoded via the 4G LTE module and transferred to the server via TCP / IP protocol.

[1793] Input: Video data stored in the internal buffer.

[1794] Output: Video data sent to the server.

[1795] Step 3: Receiving video data

[1796] Server: The server-side video receiving module receives the video data sent from the mobile terminal device and stores it in a database. Specifically, it uses Apache Kafka to manage the data stream and store it appropriately.

[1797] Input: Video data transmitted from a mobile terminal device.

[1798] Output: Video data stored in a database.

[1799] Step 4: Analyzing the footage

[1800] Server: The received video data is input into the AI ​​analysis module and analyzed using a deep learning model (e.g., TensorFlow). Specifically, objects in the video (vehicles, pedestrians, suspicious individuals, etc.) are detected and their movements are tracked.

[1801] Input: Video data stored in a database.

[1802] Output: Analysis results (e.g., JSON format) containing information about detected objects.

[1803] Step 5: Risk detection and prediction

[1804] Server: Based on the results of the AI ​​analysis module, it evaluates the risk of pedestrians running out into the street and the approach of suspicious individuals. Specifically, it calculates the speed and distance of objects and executes an algorithm to assess the risk level.

[1805] Input: Analysis results of the AI ​​analysis module.

[1806] Output: Evaluated risk information (e.g., risk of people jumping out, suspicious person warning).

[1807] Step 6: Generate a warning message

[1808] Server: Based on the detected risks, the warning generation module generates warning messages. Specifically, a Python script is used to generate warning messages (e.g., "Watch out for people jumping out!", "Watch out for suspicious people!") as text and audio data.

[1809] Input: Assessed risk information.

[1810] Output: Warning message in text and audio format.

[1811] Step 7: Sending a warning message

[1812] Server: The server transmits the generated warning message to the mobile terminal device. Specifically, the server encodes the warning message via the communication module and transmits it.

[1813] Input: Warning message in text and audio format.

[1814] Output: The warning message sent to the handheld device.

[1815] Step 8: Notification of warnings

[1816] Terminal: The mobile terminal device notifies the user of the received warning message. Specifically, it generates the warning message using voice synthesis technology and outputs it from the speaker. It also issues a physical warning using the vibration function.

[1817] Input: The warning message received.

[1818] Output: Audio and vibration alert notifications.

[1819] (Application example 1)

[1820] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1821] Conventional safety monitoring systems in factories mainly use fixed cameras, making it difficult to effectively monitor the movements of dynamically moving robots and people in real time. Furthermore, advanced analysis technology is required to detect abnormal or suspicious behavior, and it is not possible to issue appropriate warnings immediately. This results in low accuracy in improving safety in factories, and there is a high possibility of delayed response in emergencies.

[1822] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1823] In this invention, the server includes a means for acquiring video from the front and rear, a means for analyzing the acquired video using AI to detect abnormal or suspicious behavior, and a means for generating a warning message based on the detection results and transmitting it to the portable device or robot. This makes it possible to detect abnormal or suspicious behavior with high accuracy in real time and to immediately issue warnings to robots and workers in the factory by voice and vibration.

[1824] "Front camera and rear camera" are imaging devices for capturing images in front of and behind the device.

[1825] A "portable device" is a device that can be carried around, has front and rear cameras, and has the ability to capture and transmit images.

[1826] A "server" is a computer system that receives and analyzes video sent from a portable device.

[1827] "AI analysis" is the process of using artificial intelligence technology to analyze video data and detect specific patterns or anomalies.

[1828] A "warning message" is a message that is generated to warn the user when an abnormality or risk is detected.

[1829] A "factory robot" is a mechanical device that operates autonomously or semi-autonomously in a factory to perform designated tasks.

[1830] An "audio output device" is a device for producing electronically generated speech.

[1831] The "vibration function" is a function that causes a device or robot to vibrate to provide a physical notification to the user.

[1832] "Real-time" refers to the fact that information acquisition, processing, and output of results are carried out almost simultaneously.

[1833] "Abnormal and suspicious behavior" refers to behavior that deviates from normal patterns of behavior or that appears suspicious.

[1834] A "deep learning model" is an artificial intelligence model that uses a multi-layer neural network to learn the characteristics of data and perform advanced predictions and classifications.

[1835] "Tracking" is a technology that continuously tracks a specific object or movement and evaluates its position and status.

[1836] The present invention provides a system for monitoring the safety of robots and workers in a factory, detecting abnormal or suspicious behavior in real time, and issuing a warning. A specific embodiment of this system will be described below.

[1837] Main components of the system

[1838] 1. Handheld devices

[1839] It is equipped with a front-facing camera and a rear-facing camera, which capture images in real time. It also has the function of transmitting the images to a server via a communication module and receiving warning messages.

[1840] 2. Server

[1841] It receives video data sent from a portable device, performs AI analysis, and generates a warning message based on the analysis results, which is then sent back to the portable device.

[1842] 3. Robots in factories

[1843] Equipped with front and rear cameras, it captures video footage to detect abnormal or suspicious activity, and receives warning messages and issues audio and vibration alerts.

[1844] Hardware and Software Used

[1845] Cameras: Front and rear facing cameras

[1846] Communication module: Wi-Fi or Bluetooth

[1847] Server: High-performance server equipped with GPU

[1848] AI analysis software: TensorFlow, PyTorch

[1849] Video capture software: OpenCV

[1850] Warning notification software: paho-mqtt (communication), pyttsx3 (voice synthesis)

[1851] System Operation

[1852] The server receives video data sent from the mobile devices and factory robots. It analyzes this data using a deep learning model to detect abnormal or suspicious behavior in real time. Based on the results, it generates a warning message and sends it back to the mobile devices and robots. The robots then immediately notify them of the received warning message by voice and vibration.

[1853] Specific examples

[1854] Abnormal behavior detection

[1855] 1. Terminal (handheld device): Robots in the factory capture front and rear images in real time and send them to the server.

[1856] 2. Server: Analyzes the received video using a deep learning model to detect abnormal behavior, such as when a robot deviates from its normal behavior pattern.

[1857] 3. Server: Generates a warning message saying "Danger! Caution! Watch the machine!" and sends it to the robots in the factory.

[1858] 4. Terminal (robot in factory): Notifies warnings by voice and vibration.

[1859] Suspicious Activity Detection

[1860] 1. Terminal (portable device): A robot moving around the factory captures rearward-facing images and sends them to a server.

[1861] 2. Server: AI analyzes the received video and detects suspicious movements, such as when a suspicious person approaches the robot.

[1862] 3. Server: Generates a warning message saying "Suspicious behavior detected!" and sends it to the robots in the factory.

[1863] 4. Terminal (robot in factory): Notifies warnings by voice and vibration.

[1864] Prompt Sentence Examples

[1865] Factory robots capture images from cameras in front and behind them in real time and use AI to monitor the safety of the work area. For example, if a person or machine behaves abnormally, it will automatically issue a voice warning saying, "Danger! Watch out for machine operation!" What will happen if an abnormality is detected?

[1866] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1867] Step 1:

[1868] The terminal (a robot in a factory) acquires real-time images from the front and rear. It uses a front camera and a rear camera to capture the images. The input obtained is raw image data.

[1869] Step 2:

[1870] The video data captured by the terminal (a robot in a factory) is sent to the server via a communication module. Using a communication protocol (e.g., MQTT or HTTP), the video data is divided into packets and sent to the server. The server then receives the video data as input.

[1871] Step 3:

[1872] The video data received by the server is input into an AI analysis module, which uses a deep learning model (e.g., TensorFlow or PyTorch) to recognize objects in the video and analyze their behavior. This outputs intermediate data (e.g., the position of each object and its movement vector) that can be used to identify abnormal or suspicious behavior.

[1873] Step 4:

[1874] The server performs risk assessment based on the results of AI analysis. It evaluates identified abnormal or suspicious behavior and runs an algorithm to calculate the risk level. This generates a warning message for any detected risks.

[1875] Step 5:

[1876] The server generates a warning message and sends it back to the terminal (the robot in the factory) via the communication module, which then receives the warning message.

[1877] Step 6:

[1878] Based on the warning message received by the terminal (a robot in the factory), a warning is issued by voice and vibration. The voice output device is used to issue a voice warning, and the vibration module is used to notify the warning by vibration. The final output is a real-time warning notification to the user.

[1879] Step 7:

[1880] The user (worker) receives a warning from the robot and responds immediately. The worker receives a warning via voice and vibration and responds to the surrounding risks.

[1881] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1882] MODE FOR CARRYING OUT THE INVENTION

[1883] This invention is a system to prevent traffic accidents and robberies that elementary school children and the elderly may encounter in their daily lives. Specifically, it uses a portable device equipped with a front and rear camera, and utilizes AI technology and an emotion engine to detect surrounding risks and issue warnings according to the user's emotional state.

[1884] System configuration

[1885] This system consists of the following main components:

[1886] 1. Handheld devices

[1887] Front and rear cameras: These are necessary elements for capturing images in real time.

[1888] Communication module: Has the function to send captured images to a server.

[1889] Emotion engine: Has the function to analyze the user's emotional state in real time.

[1890] Warning notification function: This module notifies the user of warning messages received from the server. Specifically, it has audio output and vibration functions.

[1891] 2. Server

[1892] Video receiving module: Has the function to receive video data transmitted from a portable device.

[1893] AI analysis module: Analyzes the received video data and executes algorithms to detect sudden jumps in and suspicious individuals.

[1894] Alert generation module: Based on the detection results, it generates an alert message and sends it to the mobile device.

[1895] System Operation

[1896] The operation of this system will be explained below.

[1897] 1. Video acquisition and transmission

[1898] Terminal: The mobile device captures video in real time using front and rear cameras and transmits the data to the server at regular intervals.

[1899] 2. Video Analysis

[1900] Server: The server receives the video data sent from the mobile device and inputs it into the AI ​​analysis module, which uses deep learning to detect objects in the video (e.g., vehicles, pedestrians, suspicious individuals).

[1901] 3. Analysis of the user's emotional state

[1902] Terminal: The emotion engine installed in the portable device uses a camera and audio sensors to analyze the user's facial expressions and tone of voice, and determine their emotional state in real time.

[1903] 4. Risk detection and warning

[1904] Server: Using AI analysis, it assesses the risk of pedestrians running out into the road and the approach of suspicious individuals. For example, it detects when a pedestrian is about to run out into the road or when a suspicious individual is approaching an elderly person.

[1905] Server: When a risk is detected, it generates an appropriate warning message according to the user's emotional state and sends it to the mobile device. Depending on the emotional state, it adjusts the content of the warning and the notification method.

[1906] 5. Warning Notification

[1907] Terminal: The terminal notifies the user of warning messages received from the server by voice or vibration. Specifically, it uses voice synthesis to issue warnings such as "Danger! Watch out for people jumping out into the street!" or "There's a suspicious person behind you, so be careful."

[1908] Specific examples

[1909] Detecting elementary school students jumping out

[1910] Device: As elementary school students walk to school, the front-facing camera captures footage of the road.

[1911] Terminal: Video data is sent to the server.

[1912] Server: The video received by the server is analyzed by AI, and any movement of the elementary school student trying to jump out is detected.

[1913] Device: The emotion engine analyzes the emotional state of elementary school students and generates stronger warning messages if they are not concentrating.

[1914] Server: Generates a "Watch out for people jumping out!" warning message and sends it to the handheld device.

[1915] Device: Elementary school students will hear a voice warning saying, "Danger! Watch out for people jumping out!"

[1916] User: Elementary school student listens to the warning and stops to prevent a traffic accident.

[1917] Detecting suspicious elderly people

[1918] Device: The rear camera captures rearward footage as the elderly person walks home.

[1919] Terminal: The acquired video data is sent to the server.

[1920] Server: The server analyzes the video and detects suspicious individuals approaching the elderly.

[1921] Device: Detects the emotional state of the elderly person and generates a soft-toned warning message if they are in a state of tension.

[1922] Server: Generates a "Beware of suspicious person!" warning message and sends it to the mobile device.

[1923] Terminal: A voice warning is issued to the elderly person saying, "There is a suspicious person behind you, so please be careful."

[1924] Users: Elderly people receive a warning and pay attention to their surroundings to prevent harm from suspicious people.

[1925] As described above, this invention improves the safety of elementary school children and the elderly in their daily lives through a system that combines a portable device and a server. The real-time risk detection and warning functions provided by this system prevent traffic accidents and robberies. Furthermore, by combining it with an emotion engine, it is possible to issue warnings according to the user's emotional state, further improving safety.

[1926] The processing flow will be explained below.

[1927] Program processing flow

[1928] Video acquisition and transmission

[1929] Device:

[1930] 1. Step 1: Activate the front and rear cameras on your mobile device.

[1931] 2. Step 2: Capture video in real time using the front and rear cameras. Video frames are acquired at a rate of, for example, 30 frames per second.

[1932] 3. Step 3: Temporarily save the captured video data in memory.

[1933] 4. Step 4: At regular intervals (for example, every second), the video data stored in memory is divided into packets and organized.

[1934] 5. Step 5: The organized data packets are sent to the server via the communication module using HTTP or WebSocket.

[1935] Video analysis

[1936] server:

[1937] 1. Step 1: The server receives the front and rear camera image data.

[1938] 2. Step 2: The received video data is decoded and input into the AI ​​analysis module.

[1939] 3. Step 3: The AI ​​analysis module uses a deep learning model (e.g., YOLO, SSD) to detect objects (e.g., vehicles, pedestrians, suspicious people) in the video frame.

[1940] 4. Step 4: Track the movement of the detected object and analyze its position and velocity.

[1941] 5. Step 5: Evaluate the risk of pedestrians running out into the street or the approach of suspicious people. Specifically, predict whether certain movements pose a risk of traffic accidents or robberies.

[1942] Analyzing the user's emotional state

[1943] Device:

[1944] 1. Step 1: Start the emotion engine on the mobile device.

[1945] 2. Step 2: Use a camera or audio sensor to capture the user's facial expressions and tone of voice.

[1946] 3. Step 3: Analyze the captured data in real time to determine the user's emotional state, for example, identifying states such as tension, anxiety, or concentration.

[1947] 4. Step 4: Send the analysis results to the server.

[1948] Risk detection and warning

[1949] server:

[1950] 1. Step 1: Determine whether there is a risk based on the results of AI analysis. For example, evaluate whether a pedestrian is about to run out into the road or whether a suspicious person is approaching.

[1951] 2. Step 2: If a risk is detected, a warning message is generated based on the analysis of the transmitted emotional state.

[1952] 3. Step 3: Adjust the content and notification method of the warning message according to the user's emotional state.

[1953] 4. Step 4: Send the generated warning message to the terminal.

[1954] Warning Notification

[1955] Device:

[1956] 1. Step 1: Decode the warning message received from the server.

[1957] 2. Step 2: The decoded warning message is notified to the user. Specifically, the voice synthesis module generates a warning voice and outputs it from the speaker or activates the vibration motor.

[1958] User Feedback

[1959] User:

[1960] 1. Step 1: Receive an alert notification from your device, for example, hear a sound notification or feel a vibration.

[1961] 2. Step 2: Follow the warning and take appropriate action. Specifically, if you receive a warning about people jumping out into the street, stop, and if you receive a warning about suspicious people, turn around and check your surroundings.

[1962] Specific examples

[1963] Detecting elementary school students jumping out

[1964] Device:

[1965] 1. Step 1: When an elementary school student is walking to school, the front camera captures images of the road.

[1966] 2. Step 2: The video data is sent to the server.

[1967] server:

[1968] 1. Step 1: The video received by the server is analyzed by AI, and any movement of the elementary school student trying to jump out is detected.

[1969] 2. Step 2: Receive data from the emotion engine to analyze the emotional state of elementary school students.

[1970] 3. Step 3: If the elementary school student is not concentrating, generate a stronger warning message.

[1971] 4. Step 4: Generate a warning message "Watch out for people jumping out!" and send it to the mobile device.

[1972] Device:

[1973] 1. Step 1: Elementary school students will hear a voice warning saying, "Danger! Watch out for people jumping out!"

[1974] User:

[1975] 1. Step 1: Elementary school students listen to the warning and stop to prevent traffic accidents.

[1976] Detecting suspicious elderly people

[1977] Device:

[1978] 1. Step 1: When an elderly person is returning home, the rear camera captures the rear view.

[1979] 2. Step 2: The acquired video data is sent to the server.

[1980] 3. Step 3: The emotion engine captures the emotional state of the elderly person and sends the analysis results to the server.

[1981] server:

[1982] 1. Step 1: The server analyzes the video and detects suspicious individuals approaching the elderly person.

[1983] 2. Step 2: Evaluate the elderly person's emotional state and generate a soft-toned warning message if they are in a tense state.

[1984] 3. Step 3: Generate a warning message saying "Beware of suspicious person!" and send it to the mobile device.

[1985] Device:

[1986] 1. Step 1: The elderly person receives a voice warning saying, "There is a suspicious person behind you, so please be careful."

[1987] User:

[1988] 1. Step 1: The elderly person receives a warning and pays attention to their surroundings to prevent harm from suspicious individuals.

[1989] Example 2

[1990] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1991] Traffic accidents and robberies pose significant threats, especially to elementary school children and the elderly. However, existing safety systems lack the ability to detect risks in real time or to issue warnings based on the user's emotional state. Therefore, a more effective system is needed to prevent these risks.

[1992] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1993] In this invention, the server includes means for analyzing front and rear images using AI to detect pedestrians jumping out and suspicious people in the vicinity, means for analyzing the emotional state of the user and generating a warning message based on the emotional state, and means for transmitting the warning message to the portable device and notifying the user by voice or vibration. This allows the user to receive risk notifications in real time and receive appropriate warning messages tailored to their emotional state.

[1994] "Front camera and rear camera" refers to a photographing device that is attached to a mobile device and captures images of the front and rear in real time.

[1995] A "portable device" is an electronic device that can be carried by a user and has functions such as a camera, a communication module, and an emotion engine.

[1996] A "communication module" is a device that provides wireless communication functionality for transmitting video data from a portable device to a server.

[1997] The "server" is a central processing unit that receives video data sent from the portable device, analyzes it to detect risks, and generates and sends warning messages.

[1998] "AI analysis" is the process of using artificial intelligence technology on a server to analyze video data to detect objects and assess risks.

[1999] A "warning message" is a notification sent to a portable device to alert a user to a danger, and is presented by sound or vibration.

[2000] An "emotion engine" is software or hardware that analyzes a user's facial expressions and tone of voice to determine their emotional state in real time.

[2001] "Speech synthesis" is a technology that generates natural-sounding speech from text.

[2002] "Vibrate" is a feature that allows a mobile device to generate physical vibrations to provide notifications to the user.

[2003] A "deep learning model" is a multi-layered neural network used to learn complex patterns from data.

[2004] "Object detection" is the process of recognizing specific objects in video data and identifying their type and location.

[2005] "Tracking" is the technique of following the movement of a detected object over time and continuously monitoring its behavior.

[2006] This invention is a system for preventing traffic accidents and robberies that elementary school students and the elderly encounter in their daily lives. The system uses a portable device and a server to analyze video data and the user's emotional state, providing real-time risk detection and warnings.

[2007] System configuration

[2008] The system consists of the following main components:

[2009] 1. Handheld devices

[2010] Front and rear cameras: Used to capture footage in real time.

[2011] Communication module: Used to send acquired video data to the server.

[2012] Emotion engine: Used to analyze the user's facial expressions and tone of voice to determine their emotional state in real time.

[2013] Warning notification function: Equipped with voice output and vibration functions to notify the user of warning messages received from the server.

[2014] 2. Server

[2015] Video receiving module: Used to receive video data transmitted from a portable device.

[2016] AI analysis module: Analyzes the received video data and uses deep learning models to detect objects in the video (vehicles, pedestrians, suspicious individuals, etc.).

[2017] Alert generation module: Based on the detection results and the user's emotional state, an alert message is generated and sent to the mobile device.

[2018] Operational Overview

[2019] Video acquisition and transmission

[2020] The device uses front and rear cameras to capture video in real time and transmits the data to a server at regular intervals, with the communication module playing a key role in this process.

[2021] Video analysis

[2022] The server receives the video data sent from the device and inputs it into an AI analysis module, which uses deep learning to detect objects in the video and track their movements.

[2023] Analyzing the user's emotional state

[2024] The device uses an emotion engine to analyze the user's facial expressions and tone of voice in real time to determine their emotional state.

[2025] Risk detection

[2026] The server assesses risks based on the AI ​​analysis results and the user's emotional state, such as detecting pedestrians running out into the street or suspicious people approaching.

[2027] Generate and send warning messages

[2028] The server generates an appropriate warning message depending on the detected risk and the user's emotional state and sends it to the mobile device.

[2029] Warning Notification

[2030] The device notifies the user of the warning message received from the server by voice or vibration. Specific notification content is voice synthesis such as "Danger! Watch out for people jumping out into the street!" or "There is a suspicious person behind you, so be careful."

[2031] Specific examples

[2032] Example of detecting elementary school students jumping out

[2033] 1. Device: When an elementary school student walks to school, the front camera captures images of the road.

[2034] 2. Terminal: The captured video data is sent to the server.

[2035] 3. Server: The received video data is analyzed by an AI analysis module, and any movement of the elementary school student attempting to jump out is detected.

[2036] 4. Device: The emotion engine analyzes the facial expressions of elementary school students and detects when they are not concentrating.

[2037] 5. Server: Generates a warning message saying "Watch out for people jumping out!" and sends it to the mobile device.

[2038] 6. Device: Elementary school students will receive a voice warning and vibration saying, "Danger! Watch out for people jumping out!"

[2039] 7. User: Elementary school students who hear the warning will stop and be able to prevent a traffic accident.

[2040] Example of detecting suspicious elderly people

[2041] 1. Device: When the elderly person is returning home, the rear camera captures the rear view.

[2042] 2. Terminal: The captured video data is sent to the server.

[2043] 3. Server: The received video data is analyzed by an AI analysis module to detect suspicious individuals approaching the elderly.

[2044] 4. Device: The emotion engine analyzes the facial expressions of the elderly person and detects their state of tension.

[2045] 5. Server: Generates a warning message saying "Beware of suspicious person!" and sends it to the mobile device.

[2046] 6. Terminal: The elderly person will receive a voice warning and vibration saying, "There is a suspicious person behind you, so please be careful."

[2047] 7. User: By receiving the warning, elderly people can prevent harm from suspicious people by paying attention to their surroundings.

[2048] Prompt Sentence Examples

[2049] Below are some example prompts to input to the generative AI model:

[2050] "If an elementary school student is walking and is about to run out into the road, how would the system detect the danger and what kind of warning would it give?"

[2051] An example of a generated description:

[2052] The system captures road images in real time using the mobile device's front camera and sends them to a server. An AI analysis module on the server analyzes the video data and detects any movements of the elementary school student attempting to run out into the street. An emotion engine then analyzes the student's emotional state, and if the student is not concentrating, it generates a stronger warning message, "Danger! Watch out for those running out into the street!", and sends it to the mobile device. Finally, the mobile device issues a voice warning message to alert the elementary school student to the danger.

[2053] The above is a specific embodiment for carrying out the present invention. This system is expected to significantly improve the safety of elementary school children and the elderly, and to prevent traffic accidents and robberies.

[2054] The flow of the identification process in the second embodiment will be described with reference to FIG.

[2055] Step 1: Acquire footage

[2056] The device uses the front and rear cameras to capture video in real time, a process that occurs continuously at a constant frame rate.

[2057] Input: The surrounding physical environment.

[2058] Output: Front and rear video data.

[2059] What it does: The camera on your handheld device captures real-time video of what's in front of and behind you, even while you're moving.

[2060] Step 2: Sending video data

[2061] The device transmits the captured video data to the server at regular intervals using a communication module.

[2062] Input: Front and rear video data.

[2063] Output: Video data transferred to the server.

[2064] Specific operation: Video data is compressed and sent in packet format to the server every 5 seconds.

[2065] Step 3: Receiving video data

[2066] The server receives the video data sent from the terminal. This function is performed by the video receiving module.

[2067] Input: Video data sent from the device.

[2068] Output: Video data stored in the server's memory.

[2069] Specific operation: The server's video receiving module receives the data packets and expands them into memory for analysis.

[2070] Step 4: Analyzing the video data

[2071] The server inputs the received video data into an AI analysis module, which uses deep learning algorithms to detect objects in the video (vehicles, pedestrians, suspicious individuals, etc.) and track their movements.

[2072] Input: Video data stored on the server.

[2073] Output: Data about detected objects and their movements.

[2074] How it works: The AI ​​analysis module analyzes the video data frame by frame, identifies objects in each frame, and tracks their position and movement.

[2075] Step 5: Analyzing the user's emotional state

[2076] The device uses an emotion engine to analyze the user's facial expressions and tone of voice in real time to determine their emotional state.

[2077] Input: User facial and voice data from the camera and microphone on the handheld device.

[2078] Output: The user's emotional state (angry, surprised, relaxed, etc.).

[2079] How it works: The device's emotion engine captures the user's face and voice and analyzes them to determine their emotional state.

[2080] Step 6: Detect risks

[2081] The server combines the results of AI analysis with the user's emotional state to detect risks, such as pedestrians running out into the street or suspicious people approaching.

[2082] Input: Detection results from the AI ​​analysis module, and user emotional state data.

[2083] Output: Information about the detected risks.

[2084] Specific behavior: The server integrates the detected object movement with the user's emotional state to identify dangerous situations, such as when the user is about to run out into the road or when a suspicious person is approaching.

[2085] Step 7: Generate a warning message

[2086] The server generates appropriate warning messages based on the detected risks, which are tailored according to the user's emotional state.

[2087] Input: Information on detected risks, user emotional state data.

[2088] Output: A warning message.

[2089] Specific behavior: The server generates a soft tone warning if the user is nervous and a hard tone warning if the user is unaware.

[2090] Step 8: Sending a warning message

[2091] The server sends the generated warning message to the terminal.

[2092] Input: A warning message.

[2093] Output: The warning message sent to the terminal.

[2094] How it works: The server sends real-time alert messages to the device using a low-latency communication protocol.

[2095] Step 9: Notification of warnings

[2096] The terminal notifies the user of the warning message received from the server by voice or vibration.

[2097] Input: The warning message received from the server.

[2098] Output: The warning that was notified to the user.

[2099] What it does: The device's voice synthesis function warns the user, saying "Danger! Watch out for people jumping out!", while also providing physical feedback to the user using the device's vibration function.

[2100] As described above, each step works in conjunction with the others, allowing users to receive risk notifications in real time and receive appropriate warning messages according to their emotional state.

[2101] (Application example 2)

[2102] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2103] Conventional safety systems using portable devices focus only on risk detection and do not optimize warnings based on the user's emotional state. As a result, the effectiveness of warnings is limited, making it difficult to ensure sufficient safety, especially for users with large emotional fluctuations, such as elementary school children and the elderly. Furthermore, the generated warning messages are uniform, making it difficult to provide warnings appropriate to the situation and the user's state.

[2104] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing front and rear images using AI to detect pedestrians jumping out and suspicious people in the vicinity, means for providing the portable device with an emotion engine for analyzing the user's emotional state, means for generating a warning message based on the detection results and the user's emotional state, and means for optimizing the content of the warning message using a generative AI model. This makes it possible to generate and notify the user of an optimal warning message in real time according to their emotional state, thereby providing greater safety than conventional systems.

[2105] "Front camera and rear camera" refers to a photographing device attached to a mobile device for capturing images of the front and rear.

[2106] A "portable device" is a device that can be easily carried around and is equipped with a communication module and a warning notification function.

[2107] "AI analysis" means using artificial intelligence technology to analyze acquired video data and detect specific patterns and risks.

[2108] "Detecting pedestrians running out into the road and suspicious people in the surrounding area" means that AI analyzes video data to find pedestrians who are about to run out into the road or people behaving abnormally.

[2109] An "emotion engine" is an algorithm or software that analyzes information such as a user's facial expressions and tone of voice to determine their emotional state in real time.

[2110] "Generating a warning message" means creating appropriate warning content based on the detected risk and the user's emotional state.

[2111] "Notifying the user of a warning by voice or vibration" means that the generated warning message is conveyed to the user by using a voice output device or a vibration function.

[2112] A "generative AI model" is an artificial intelligence model that creates appropriate responses and outputs based on generated data and input information.

[2113] "Optimization" means making adjustments to obtain the most effective or efficient method or result under given conditions.

[2114] "Generating and notifying in real time" means analyzing data in response to the situation at hand and providing results and information immediately.

[2115] 1. System program generation

[2116] This system uses a portable device equipped with a front and rear camera and works in conjunction with a server to analyze both risk and emotional state. Users can carry the portable device and monitor the risks around them in real time. The system has the following functions:

[2117] 2. Program processing explanation

[2118] Video acquisition and transmission

[2119] The terminal captures video in real time using the front and rear cameras. This video data is sent to the server via a communication module. The terminal has a built-in communication module that supports sending and receiving video data.

[2120] Server-based video analysis

[2121] The server analyzes the received video data. Specifically, the AI ​​analysis module uses a deep learning model to detect objects in the video (e.g., vehicles, pedestrians, suspicious individuals) and track their movements. This makes it possible to detect risks such as pedestrians running out into the street or suspicious individuals approaching in real time.

[2122] Emotion analysis

[2123] The emotion engine installed in the device uses a camera and audio sensors to analyze the user's facial expressions and tone of voice to determine the user's emotional state. A deep learning model is used to analyze the emotional state.

[2124] Generating and notifying warning messages

[2125] The server generates an optimal warning message based on the risk detection results and the user's emotional state. The content of this warning message is optimized using a generative AI model. The generated warning message is then sent back to the device via the communication module.

[2126] The device notifies the user of the received warning message using a voice output device or vibration function. For example, it generates messages such as "Danger! Watch out for people jumping out into the street!" or "Be careful, there is a suspicious person behind you."

[2127] 3. Specific Examples

[2128] Detecting elementary school students jumping out

[2129] When elementary school students use the vehicle on their way home from school, the front camera captures images of the road in real time. If the server detects an approaching vehicle and determines that the child is not concentrating through emotional analysis, it issues a voice warning saying "Danger! Watch out for children running out into the street!" and simultaneously activates a vibration. This allows elementary school students to pay close attention to the warning and prevent them from running out into the street.

[2130] Detecting suspicious elderly people

[2131] When an elderly person is returning home, a rear camera captures video of the area behind them. If analysis on the server detects that a suspicious person is approaching from behind, and emotion analysis determines that the elderly person is nervous, a gentle voice warning is issued saying, "There is a suspicious person behind you, please be careful." This allows the elderly person to deal with the situation calmly and prevent harm from the suspicious person.

[2132] Prompt Sentence Examples

[2133] After-school risk detection prompts:

[2134] Design a system that uses an AI algorithm to detect in real time when a vehicle ahead is approaching, and issues an audio warning saying "Danger! Watch out for vehicles jumping out!" if the user's emotional state is not focused.

[2135] Hardware and Software

[2136] The server requires a data center component with high-performance computing resources and an AI framework (e.g., TensorFlow, PyTorch) for running deep learning models, while the device requires a high-resolution camera, an audio sensor, a communication module (e.g., Wi-Fi, 4G / 5G module), and an on-device AI framework (e.g., TensorFlow Lite) for running a sentiment analysis engine.

[2137] This enables the entire system to work together to detect risks in real time and deliver optimized warning messages to enhance user safety.

[2138] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[2139] Step 1:

[2140] The terminal captures video in real time using a front camera and a rear camera. A user carries a portable device and periodically captures video of the front and rear. This allows the terminal to obtain video data of the current surrounding environment. The input is the real-time video data obtained from the camera, and the output is this video data.

[2141] Step 2:

[2142] The terminal transmits the acquired front and rear video data to the server via the communication module. The communication module in the device is activated and uploads the video data to the server via the network. At this stage, the input is the front and rear video data, and the output is the video data received by the server.

[2143] Step 3:

[2144] The server analyzes the received video data using a deep learning model. Specifically, the AI ​​analysis module detects objects in the video data and tracks their movements. The server evaluates risks, such as pedestrians running out into the street or suspicious people approaching, in real time. The input is the video data received by the server, and the output is the risk detection results.

[2145] Step 4:

[2146] The device uses an emotion engine to analyze the user's emotional state in real time. The device's camera and audio sensors capture the user's face and voice data, which are then analyzed by an emotion analysis algorithm. The input is the user's face and voice data, and the output is the analyzed emotional state.

[2147] Step 5:

[2148] The server uses a generative AI model to generate an optimal warning message based on the risk detection results and the user's emotional state. The server determines the appropriate warning message by taking into account the type of risk and the user's current emotional state. The input is the risk detection results and the user's emotional state, and the output is the warning message.

[2149] Step 6:

[2150] The terminal receives the generated warning message via the communication module. The terminal acquires the warning message sent from the server and notifies the user using audio output and vibration functions. The input is the warning message sent from the server, and the output is the warning message notified to the user.

[2151] This allows the entire system to work together to detect risks and provide emotion-based warning notifications to enhance user safety in real time.

[2152] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[2153] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[2154] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[2155] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[2156] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[2157] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[2158] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[2159] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[2160] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[2161] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[2162] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[2163] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external dev...

Claims

1. means for capturing front and rear images using a handheld device equipped with a front camera and a rear camera; means for transmitting the captured video from the portable device to a server; A means for analyzing the front and rear images in the server using AI to detect pedestrians jumping out and suspicious people in the surrounding area; means for generating a warning message based on the detection result; means for transmitting the warning message to the portable device and notifying a user of the warning; A system including:

2. A means for acquiring front and rear images in real time using a portable device equipped with a front camera and a rear camera and transmitting the images to a server at regular intervals; In the server, a means for analyzing the transmitted video using AI to predict whether a pedestrian will suddenly appear or whether a suspicious person will appear in the vicinity; means for generating a warning message based on the prediction result and transmitting the warning message to the portable device; a means for notifying a user of a warning message by voice or vibration in the portable device; The system of claim 1 , comprising:

3. The system of claim 1, further comprising means for detecting objects in the video and tracking their movements using a deep learning model when analyzing with the AI.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A