System
The system addresses the limitation of existing cameras by implementing real-time video analysis and generative AI to detect abnormal behavior and issue warnings, ensuring prompt safety measures.
Patent Information
- Application Number
- JP2024126245
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-01
- Publication Date
- 2026-02-13
AI Technical Summary
Commercially available security and home cameras lack the ability to analyze video footage in real time to detect abnormal behavior and issue immediate warnings, making it difficult to prevent crimes and accidents, particularly in environments where children or public safety is a concern.
A system that includes video acquisition, real-time video analysis, abnormal behavior detection, warning message generation, and information sharing, utilizing generative AI models to generate appropriate warnings and convert them into audio, with real-time communication to administrators.
Enables rapid detection of abnormal behavior and issuance of contextually relevant warnings, enhancing safety by allowing immediate responses to potential threats.
Smart Images

Figure 2026023924000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Currently, commercially available security cameras and home cameras are limited to recording video footage and lack the ability to analyze the footage in real time to instantly detect abnormal behavior and issue a warning. As a result, it is difficult to prevent crimes and accidents, and there is a particular need for rapid response to ensure the safety of children at home and in public places. The present invention aims to prevent crimes and accidents and create a safe environment by providing a system that instantly detects abnormal behavior through video analysis and issues appropriate warnings based on the detected abnormal behavior. [Means for solving the problem]
[0005] The present invention is a system that includes the following means. First, it includes a video acquisition means that acquires video in real time from a surveillance camera or the like. Next, it includes a video analysis means that analyzes the acquired video and detects specific abnormal behavior. Furthermore, it includes an abnormal behavior detection means that detects abnormal behavior based on the analysis results. It includes a warning message generation means that generates an appropriate warning message using a generative AI model when abnormal behavior is detected. It also includes a warning message reading means that converts the generated warning message into audio and reads it to the target person. Finally, it includes an information sharing means that shares the detected abnormal behavior and the generated warning message with an administrator. This makes it possible to detect abnormal behavior in real time and issue prompt and appropriate warnings.
[0006] "Video acquisition means" refers to a system that has the function of acquiring video in real time from surveillance cameras, home cameras, etc.
[0007] The "video analysis means" is a system that can analyze acquired video data in real time and detect specific abnormal behavior.
[0008] The "abnormal behavior detection means" is a system that determines and detects specific abnormal behavior based on the analysis results provided by the video analysis means.
[0009] The "warning message generating means" is a system for generating an appropriate warning message according to the situation when abnormal behavior is detected.
[0010] The "means for reading out warning text" is a system that converts the generated warning text into audio and actually reads it out loud to the target person.
[0011] The "information sharing means" is a system for sharing detected abnormal behavior and generated warning messages with administrators and other relevant parties in real time. [Brief explanation of the drawings]
[0012] [Figure 1]1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0013] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0014] First, the terms used in the following description will be explained.
[0015] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0016] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0017] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0018] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0020] [First embodiment]
[0021] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0022] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0023] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0024] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0025] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0027] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0028] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0029] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0030] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0031] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0032] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0033] The present invention is a system that processes a series of steps including video acquisition, video analysis, abnormal behavior detection, warning message generation, warning message reading, and information sharing. An embodiment of this system is mainly composed of the following elements.
[0034] System Configuration
[0035] 1. Video acquisition method
[0036] The server acquires video data from surveillance cameras and home cameras in real time.
[0037] The camera has a real-time streaming function, and the captured images are stored in a buffer on the server.
[0038] 2. Video analysis methods
[0039] The video analysis module in the server processes the acquired video data and performs object detection and motion analysis.
[0040] The server uses existing deep learning algorithms to analyze patterns of abnormal behavior.
[0041] 3. Abnormal Behavior Detection Methods
[0042] The server's abnormal behavior detection module determines and detects specific abnormal behavior in real time based on the video analysis results.
[0043] For example, it identifies a child trying to get into a washing machine or a snatching incident on the street.
[0044] 4. Warning statement generation means
[0045] If the server detects abnormal behavior, it uses a generative AI model to generate an appropriate warning message.
[0046] The generated warning message will vary depending on the abnormal behavior detected.
[0047] 5. Warning message reading method
[0048] The device (home smart speaker, street announcement system) converts the generated warning text into audio using a voice generation module and reads it out to the target person.
[0049] For example, at home, a warning will be played saying, "Playing with the washing machine is dangerous. Stop immediately."
[0050] 6. Information sharing methods
[0051] The server generates a warning message and sends abnormal behavior information to an administrator (parent, police, store manager, etc.) in real time.
[0052] Users (parents or administrators) can receive warning information on their smartphones or dedicated apps and respond quickly.
[0053] Specific examples
[0054] Watching over children at home
[0055] The server acquires real-time video from cameras in the home and transmits the video data to the video analysis module.
[0056] The server's video analysis module analyzes the child's behavior of opening the washing machine door, and the abnormal behavior detection module detects this as abnormal behavior.
[0057] The server generates a warning message saying, "Playing with the washing machine is dangerous. Stop immediately," and sends it to a smart speaker in the home.
[0058] The device (smart speaker) reads out the warning message aloud, and at the same time, the server sends the warning information to the parent's smartphone, allowing the parent to check the information immediately.
[0059] Street crime prevention surveillance
[0060] The server acquires real-time video from surveillance cameras installed on the street and transmits the video data to the video analysis module.
[0061] The server's video analysis module detects snatching attempts, and the abnormal behavior detection module determines this as abnormal behavior.
[0062] The server generates a warning message saying, "Purse snatching has been detected here. Please leave immediately," and sends it to the street announcement system.
[0063] The terminal (street announcement system) reads out the warning message, and at the same time, the server transmits the warning information to the police or monitoring center, enabling immediate action.
[0064] Program processing explanation
[0065] The server acquires video data from the surveillance camera and stores it in a buffer.
[0066] A video analysis module in the server analyzes the video data stored in the buffer and identifies abnormal behavior.
[0067] Based on the analysis results, the server's abnormal behavior detection module detects abnormal behavior.
[0068] If abnormal behavior is detected, the server uses a generative AI model to generate an appropriate warning message.
[0069] The generated warning message is sent to a designated device (smart speaker, announcement system) and read aloud.
[0070] The warning message and abnormal behavior information are sent to the administrator in real time by the server, and the administrator (user) can check the warning information on a smartphone or other device.
[0071] In this way, the form for implementing the invention provides a system in which each step operates in conjunction with other steps to quickly detect abnormal behavior and issue appropriate warnings, thereby preventing crimes and accidents before they occur.
[0072] The processing flow will be explained below.
[0073] Step 1:
[0074] The server acquires video data from security cameras and home cameras in real time. This acquisition process involves capturing video at a specified frame rate and storing it in a buffer.
[0075] Step 2:
[0076] The server sends the acquired video data to the video analysis module, where the video data is converted into a format that allows real-time processing.
[0077] Step 3:
[0078] The server's video analytics module processes the video data and runs object detection and motion analysis algorithms, for example, to identify specific human movements or unusual object behavior.
[0079] Step 4:
[0080] The server's abnormal behavior detection module determines abnormal behavior based on the video analysis results. This module uses a trained deep learning model to detect, for example, a child trying to get into a washing machine as an abnormal behavior.
[0081] Step 5:
[0082] When the server detects abnormal behavior, it activates a generative AI model to generate an appropriate warning message. This generation process creates a contextual warning message based on the nature of the abnormal behavior. For example, a warning message might be generated that reads, "Playing with the washing machine is dangerous. Stop immediately."
[0083] Step 6:
[0084] The server then sends the generated warning message to a designated device, such as a smart speaker in the home or a street announcement system.
[0085] Step 7:
[0086] The device converts the received warning message into audio using a speech generation module and reads it out loud from the speaker. For example, a voice warning a child playing with a washing machine might be played saying, "Playing with a washing machine is dangerous. Stop immediately."
[0087] Step 8:
[0088] At the same time, the server sends the generated warning message and detailed information about abnormal behavior to the administrator in real time. The administrator (user) can receive the warning information on their smartphone or a dedicated application and respond quickly. For example, parents can receive a notification on their smartphone and immediately check on their child's safety.
[0089] In this way, the processing flow from step 1 to step 8 realizes a system in which abnormal behavior is quickly detected and appropriate warnings and information sharing are provided.
[0090] Example 1
[0091] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0092] In recent years, safety monitoring in homes and on the streets has become increasingly important, but conventional monitoring systems have had difficulty detecting abnormal behavior in real time and issuing appropriate warnings. Furthermore, they have been unable to quickly share information with administrators, making it difficult to take immediate action to prevent incidents and accidents. New technological solutions are needed to solve these issues.
[0093] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0094] In this invention, the server includes a video acquisition means, a video analysis means, an abnormal behavior detection means, a warning message generation means, a warning message generation means using a generative AI model, a real-time video processing means, and an information sharing means with enhanced functions. This enables real-time analysis of video data, immediate detection of abnormal behavior, dynamic generation of warning messages using the generative AI model, and rapid voice generation of the warning messages and immediate notification to an administrator.
[0095] "Video acquisition means" refers to a means for acquiring video data in real time from a surveillance camera or a home camera and storing that data.
[0096] The "video analysis means" is a means for processing the acquired video data and performing object detection and motion analysis. Specifically, it analyzes the video data using a deep learning algorithm.
[0097] "Abnormal behavior detection means" is a means for determining and detecting specific abnormal behavior in real time based on the results of video analysis.
[0098] The "warning message generating means" is a means for generating an appropriate warning message based on the detected abnormal behavior.
[0099] "Means for generating warning messages using a generative AI model" refers to means for dynamically generating warning messages in response to abnormal behavior using a generative AI model (e.g., natural language processing technology).
[0100] "Means for processing video in real time" refers to means for storing acquired video data in a buffer in real time and analyzing the data sequentially.
[0101] The "means for reading out a warning message" is a means for converting the generated warning message into voice using a voice generation module (for example, TTS technology) and reading it out to the target person.
[0102] The "information sharing means" is a means for transmitting the generated warning message and abnormal behavior information to a manager (e.g., parent, police, store manager, etc.) in real time.
[0103] A "terminal" is a device that reads out warning messages aloud, such as a smart speaker in the home or a street announcement system.
[0104] This invention is a system that processes a series of steps including video acquisition, video analysis, abnormal behavior detection, warning message generation, warning message reading, and information sharing. The embodiment of the invention will be explained mainly by dividing it into a server, a terminal, and a user.
[0105] Video acquisition method
[0106] The server acquires video data in real time from security cameras and home cameras using HTTP or RTSP protocols. The acquired video data is stored in the server's buffer. Specifically, the server acquires video frames step by step using the FFmpeg library and stores them in a buffer in memory.
[0107] Video analysis methods
[0108] The video analysis module in the server processes the real-time video data stored in the buffer. This analysis uses deep learning algorithms using TensorFlow and PyTorch. Specifically, models such as YOLO (You Only Look Once) and OpenPose are used for object detection and motion analysis. The video analysis module analyzes the person's pose for each frame and identifies their joint points. It also uses an object detection model to recognize objects held by the person.
[0109] Abnormal behavior detection method
[0110] The server's anomalous behavior detection module identifies anomalous behavior in real time from the analyzed video data. This detection module references a rule-based model that defines anomalous behavior patterns. For example, it can detect a child trying to get into a washing machine or a purse snatching attempt. The server identifies frame sequences that meet the conditions for anomalous behavior and marks them as anomalous based on that.
[0111] Warning statement generation means
[0112] If the server detects abnormal behavior, it uses a generative AI model to generate an appropriate warning message. For example, GPT-3 is used as this generative AI model. Based on the abnormal behavior information, the server inputs a prompt message into the generative AI model to generate a warning message. An example of a specific prompt message is, "A child has been detected entering the washing machine. Please generate an appropriate warning message." The generated warning message would be, "Playing with the washing machine is dangerous. Please stop immediately."
[0113] Warning message reading method
[0114] The device (smart speaker or street announcement system) uses a voice generation module to convert the generated warning text into audio and read it to the target person. Specifically, the warning text is converted into audio using a TTS (Text-to-Speech) API such as Amazon Polly or Google Text-to-Speech. The device converts the warning text received from the server into audio and reads it aloud from the speaker.
[0115] Information sharing means
[0116] The server sends the generated warning message and abnormal behavior information to an administrator (such as a parent, police, or store manager) in real time. This communication uses an HTTP-based notification API. The user (administrator) receives the warning information on their smartphone or a dedicated app, and can immediately check the content and take action. Specifically, the user receives the warning information using the real-time notification function within the app and can check the details.
[0117] Specific examples
[0118] Watching over children at home
[0119] The server receives real-time video footage from the home camera and stores the data in the server's buffer. The server's video analysis module analyzes the child's behavior when they open the washing machine door, and the abnormal behavior detection module detects this behavior as abnormal. The server uses a generative AI model to generate a warning message saying, "Playing with the washing machine is dangerous. Please stop immediately," and sends this message to the smart speaker. The device (smart speaker) then reads this warning message aloud. At the same time, the server sends the warning information to the parent's smartphone, where the parent can immediately check it.
[0120] Street crime prevention surveillance
[0121] The server acquires real-time video footage from street surveillance cameras and stores the video data in the server's buffer. The server's video analysis module detects snatching attempts, and the abnormal behavior detection module determines the behavior as abnormal. The server uses a generative AI model to generate a warning message stating, "A snatching attempt has been detected here. Please leave immediately," and sends it to the street announcement system. The device (street announcement system) reads this warning message aloud. At the same time, the server sends the warning information to the police or monitoring center, enabling immediate action.
[0122] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0123] Step 1:
[0124] Real-time video capture
[0125] The server obtains video data from a surveillance camera in real time. The input is the video stream from the surveillance camera. The server receives the video stream using HTTP or RTSP protocol, extracts this data frame by frame using the FFmpeg library, and stores it in a buffer in memory. The output is the video data in the buffer. Specifically, the server reads the IP address and connection settings of the surveillance camera, and obtains the data in real time after establishing a connection.
[0126] Step 2:
[0127] Video data analysis
[0128] The video analysis module in the server processes real-time video data stored in the buffer. The input is the video data in the buffer. The server analyzes the video using deep learning algorithms (such as YOLO or OpenPose) using TensorFlow or PyTorch. GPU acceleration is used to perform object detection and motion analysis at high speed. The output is the analysis results, which are the position information of specific objects or people in each frame. Specifically, the server inputs each frame of data into the deep learning model, obtains and stores the analysis results.
[0129] Step 3:
[0130] Abnormal behavior detection
[0131] The server's abnormal behavior detection module identifies abnormal behavior in real time based on analyzed video data. The input is the output of the video analysis module. The server references a rule-based model that defines abnormal behavior patterns in advance and determines whether it matches each frame of data. The output is information about the frame in which abnormal behavior was detected. Specifically, the server checks the conditions under which abnormal behavior occurs (e.g., a specific behavior occurs consecutively within a specific period of time) and determines that an abnormality exists based on that.
[0132] Step 4:
[0133] Generate a warning message
[0134] When the server detects abnormal behavior, it uses a generative AI model (for example, GPT-3) to generate a warning message. The input is information about the abnormal behavior. Based on this information, the server inputs a prompt message into the generative AI model. The output is the generated warning message. An example of a specific prompt message is, "A snatching attempt has been detected. Please generate an appropriate warning message." The generative model generates a warning message based on this message, and outputs the message, "A snatching attempt has been detected here. Please leave immediately."
[0135] Step 5:
[0136] Reading out warning messages
[0137] The device (such as a home smart speaker or a street announcement system) converts the generated warning text into audio using a voice generation module (for example, Amazon Polly or Google Text-to-Speech) and reads it aloud to the target person. The input is the warning text sent from the server. The device calls the TTS API, generates the warning text as an audio file, and outputs the sound from the speaker. The output is the audio of the warning text. Specifically, the device requests the received warning text from the TTS API and plays the generated audio data.
[0138] Step 6:
[0139] Information sharing
[0140] The server sends the generated warning message and abnormal behavior information to an administrator (parent, police, store manager, etc.) in real time. The input is the abnormal behavior information and the generated warning message. The server uses an HTTP-based notification API to send the warning information to the administrator's smartphone or dedicated app. The output is a notification message to the administrator. The user (administrator) receives the warning information on their smartphone or dedicated app, and can immediately check the content and take action. Specifically, the server reads the administrator's connection information and sends the warning data through the notification API.
[0141] (Application example 1)
[0142] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0143] Conventional crime prevention and surveillance systems have difficulty quickly detecting abnormal behavior and taking appropriate countermeasures immediately. Furthermore, delays in transmitting warning information can lead to problems with preventing crimes and accidents. To address these issues, there is a need for a system that can detect abnormal behavior in real time, issue warnings quickly, and share information with relevant parties immediately.
[0144] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0145] In this invention, the server includes a video acquisition means, a video analysis means, an abnormal behavior detection means, a warning message generation means, a warning message reading means, an information sharing means, a means for displaying a warning message on a display of the smart glasses, a means for issuing a voice warning, and a means for sharing abnormal behavior information with a management center, thereby enabling real-time abnormal behavior detection, prompt issuance of a warning, and prompt information sharing.
[0146] "Video acquisition means" refers to means for acquiring video data from a camera or other video capture device in real time.
[0147] The "video analysis means" is a means for processing the acquired video data and performing object detection and motion analysis.
[0148] "Abnormal behavior detection means" is a means of determining and detecting specific abnormal behavior in real time based on the results of video analysis.
[0149] The "warning message generation means" is a means for generating an appropriate warning message using a generative AI model when abnormal behavior is detected.
[0150] The "means for reading out a warning message" is a means for converting the generated warning message into voice using a voice generation module and reading it out loud to the target person.
[0151] The "information sharing means" is a means for sending the generated warning message and abnormal behavior information to the administrator in real time, enabling a prompt response.
[0152] The "means for displaying a warning message on the display of the smart glasses" refers to a means for visually displaying a warning message on the display of the smart glasses.
[0153] The "means for issuing an audio warning" is a means for issuing an audio warning when abnormal behavior is detected.
[0154] The "means for sharing abnormal behavior information with the management center" refers to a means for transmitting information on detected abnormal behavior to the management center in real time and quickly sharing it with relevant parties.
[0155] The present invention is a security system that processes the acquisition and analysis of camera footage, detection of abnormal behavior, generation and reading of warning messages, and information sharing in a series of steps. This system is mainly composed of a video acquisition means, a video analysis means, an abnormal behavior detection means, a warning message generation means, a warning message reading means, an information sharing means, a means for displaying a warning message on a smart glasses display, a means for issuing an audio warning, and a means for sharing abnormal behavior information with a management center.
[0156] Specifically, various hardware and software are used as follows:
[0157] Hardware:
[0158] Smart glasses: Devices that have a camera, display, and speaker (e.g., Google Glass, Vuzix Blade).
[0159] Camera: Surveillance cameras used on the street or in homes (e.g., network cameras).
[0160] Server: A computer system that processes and analyzes video data, detects abnormal behavior, and generates and distributes warning messages.
[0161] software:
[0162] Video Analytics Module: Object detection and motion analysis software using deep learning algorithms.
[0163] Abnormal behavior detection module: Software that identifies abnormal behavior in real time based on video analysis results.
[0164] Generative AI model: Natural language generation software that generates appropriate warning statements based on abnormal behavior.
[0165] Speech Generation Module: Text-to-Speech (TTS) software for vocalizing the generated warning text.
[0166] Information sharing module: Software that sends warning messages and abnormal behavior information to the management center and relevant parties in real time.
[0167] Data processing and calculation:
[0168] The server processes the acquired video data in real time. The video analysis module analyzes the video data sent from the camera and detects abnormal behavior, using existing deep learning methods. When abnormal behavior is detected, the generative AI model generates an appropriate warning message, which is then converted into audio by the voice generation module and sent to the smart glasses and other devices. The warning message is immediately displayed on the smart glasses' display, and a voice warning is also issued at the same time. In addition, abnormal behavior information is sent to the management center via the information sharing module, and relevant parties are immediately notified.
[0169] Examples:
[0170] For example, consider a police officer wearing smart glasses while patrolling the streets. During the patrol, the camera in the glasses captures real-time footage of the street. The footage is sent to a server, where the video analysis module detects violent acts. The generative AI model then generates a warning message, "Violent acts detected, please leave the scene," which is displayed and read aloud on the smart glasses. At the same time, this warning information is sent to a control center, allowing assistance to be quickly dispatched.
[0171] Example prompt sentence:
[0172] For example, by inputting the following prompt sentence into the generative AI model, an appropriate warning sentence will be generated.
[0173] Generate appropriate warning messages when the following actions are detected:
[0174] Street violence: Warning: "Violence detected, please respond immediately."
[0175] Snatching: Warning: "Snatching detected, please leave the scene."
[0176] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0177] Step 1:
[0178] The server acquires video data from the camera in real time. The input is the video stream from the camera, and the output video data is stored in the server's buffer.
[0179] Step 2:
[0180] The server's video analysis means processes the acquired video data. The input is the video data stored in the buffer, and object detection and motion analysis are performed, and the analysis results are output.
[0181] Step 3:
[0182] The server's abnormal behavior detection means detects specific abnormal behavior in real time based on the video analysis results. The input is the video analysis results, and the output is a determination of whether abnormal behavior has been detected.
[0183] Step 4:
[0184] The server generates a warning message using a generative AI model. The input is information about the detected abnormal behavior, and the generated warning message is output. In this case, the server uses a prompt message to generate an appropriate warning message.
[0185] Step 5:
[0186] The server sends the generated warning text to the terminal. The input is the generated warning text, which is sent to smart glasses or other terminals.
[0187] Step 6:
[0188] The terminal (smart glasses) displays a warning message on the display and issues a voice warning. The input is the warning message received from the server, and the output is a displayed warning message and a voice alert.
[0189] Step 7:
[0190] The server's information sharing means sends abnormal behavior information and warning messages to the management center. The input is the warning message and abnormal behavior information, and the output is shared in real time with the management center and relevant parties.
[0191] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0192] The present invention is a system that combines the functions of video acquisition, video analysis, abnormal behavior detection, warning message generation, warning message reading, information sharing, and emotion recognition, and links these functions together to provide appropriate warnings and share information while taking into account the user's emotions. An embodiment of this system is configured as follows.
[0193] System Configuration
[0194] 1. Video acquisition method
[0195] The server acquires video data from surveillance cameras and home cameras in real time.
[0196] The video is stored in a buffer on the server and is then analyzed.
[0197] 2. Video analysis methods
[0198] The video analysis module in the server processes the acquired video data and performs object detection and motion analysis.
[0199] The server uses deep learning algorithms to analyze patterns of abnormal behavior.
[0200] 3. Abnormal Behavior Detection Methods
[0201] The server's abnormal behavior detection module determines abnormal behavior based on the results of video analysis.
[0202] For example, it can detect a child trying to get into a washing machine or a snatch-and-run theft on the street.
[0203] 4. Emotion recognition means
[0204] The server analyzes the user's facial expressions and tone of voice and uses an emotion recognition module to recognize the user's emotions in real time.
[0205] The recognized emotional information is reflected in the generation of warning messages.
[0206] 5. Warning statement generation means
[0207] The server detects abnormal behavior and generates an appropriate warning message using a generative AI model based on the emotion recognition results.
[0208] The generated warning message will be based on the user's emotions. For example, a gentle warning message will be generated to calm the user.
[0209] 6. Warning message reading method
[0210] The device (home smart speaker, street announcement system) uses a voice generation module to convert the warning text into audio and read it out loud.
[0211] For example, a voice may be heard saying to a child, "Playing with the washing machine is dangerous. Stop immediately."
[0212] 7. Information sharing methods
[0213] The server sends the generated warning message and abnormal behavior information to the administrator in real time.
[0214] Users (parents or administrators) can receive warning information on their smartphones or dedicated apps and respond quickly.
[0215] Specific examples
[0216] Watching over children at home
[0217] The server acquires real-time video from cameras in the home and transmits the video data to the video analysis module.
[0218] The server's video analysis module analyzes the child's behavior of opening the washing machine door, and the abnormal behavior detection module detects this as abnormal behavior.
[0219] The server uses an emotion recognition module to analyze the child's facial expressions and tone of voice to obtain emotional information.
[0220] The server generates a warning message saying, "Playing with the washing machine is dangerous. Stop immediately," and adjusts the tone based on the emotion recognition results.
[0221] The device (smart speaker) reads out the warning message aloud, and at the same time, the server sends the warning information to the parent's smartphone, allowing the parent to check the information immediately.
[0222] Street crime prevention surveillance
[0223] The server acquires real-time video from surveillance cameras installed on the street and transmits the video data to the video analysis module.
[0224] The server's video analysis module detects snatching attempts, and the abnormal behavior detection module determines this as abnormal behavior.
[0225] The server uses an emotion recognition module to analyze the facial expressions and tone of voice of people around it to obtain emotional information.
[0226] The server generates a warning message saying, "Purse snatching has been detected here. Please leave immediately," and adjusts the tone based on emotional information.
[0227] The terminal (street announcement system) reads out the warning message, and at the same time, the server transmits the warning information to the police or monitoring center, enabling immediate action.
[0228] Program processing explanation
[0229] The server acquires video data from the surveillance camera and stores it in a buffer.
[0230] A video analysis module in the server analyzes the video data and identifies abnormal behavior.
[0231] When abnormal behavior is detected, the server activates the emotion recognition module and analyzes the user's emotional information.
[0232] Based on the emotional information, the server uses a generative AI model to generate an appropriate warning message.
[0233] The generated warning message is sent to the specified device and read aloud.
[0234] The warning message and abnormal behavior information are sent to the administrator in real time by the server, and the administrator (user) can check the warning information on a smartphone or other device.
[0235] In this way, it is possible to provide a system that can quickly detect abnormal behavior, recognize emotions, issue appropriate warnings, and share information.
[0236] The processing flow will be explained below.
[0237] Step 1:
[0238] The server acquires video data in real time from home cameras and street surveillance cameras. This video data is broken down into multiple frames and stored in the server's temporary memory (buffer).
[0239] Step 2:
[0240] The server sends the acquired video data to the video analysis module, which converts the video data into a format that can be processed in real time and then provides it for analysis.
[0241] Step 3:
[0242] The video analysis module in the server processes the video data and applies object detection and motion analysis algorithms to analyze the movement, shape, and position of people and objects, preparing to identify abnormal behavior.
[0243] Step 4:
[0244] The server's abnormal behavior detection module determines abnormal behavior based on the video analysis results. For example, it detects a child trying to get into a washing machine or a person moving abnormally in a specific area.
[0245] Step 5:
[0246] After the server detects abnormal behavior, it activates the emotion recognition module, which analyzes the user's facial expressions and tone of voice from the video data to recognize their emotions.
[0247] Step 6:
[0248] The server uses a generative AI model to generate appropriate warning messages based on the emotion recognition results. During this generation process, the tone and content of the warning messages are adjusted to match the user's emotions.
[0249] Step 7:
[0250] The server generates a warning message and sends it to a designated device, such as a smart speaker in the home or an announcement system on the street.
[0251] Step 8:
[0252] The warning message received by the device is converted into voice by a voice generation module and actually read aloud from a speaker. For example, at home, a warning such as "Playing with the washing machine is dangerous. Stop immediately" is played in a soft tone.
[0253] Step 9:
[0254] At the same time, the server sends the generated warning message and detailed information about the abnormal behavior to the administrator in real time. The administrator (user) receives the warning information on their smartphone or a dedicated application, enabling them to respond quickly. For example, a parent could receive a notification on their smartphone and check on their child's safety.
[0255] In this way, the processing flow from step 1 to step 9 provides a system that quickly detects abnormal behavior and provides appropriate warnings and information sharing that take the user's emotions into consideration.
[0256] Example 2
[0257] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0258] Conventional surveillance systems were able to detect abnormal behavior by acquiring and analyzing video data, but they lacked the ability to issue prompt and appropriate warnings and share information while taking into account the user's emotions. As a result, when abnormal behavior occurred, responses were sometimes delayed, making it difficult to implement effective preventative measures.
[0259] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a video acquisition means, a video analysis means, an abnormal behavior detection means, an emotion recognition means, a warning message generation means, a warning message reading means, and an information sharing means. This makes it possible to issue a prompt and accurate warning that takes into account the user's emotions and to share abnormal behavior information in real time.
[0260] The "image acquisition means" is a means for acquiring image data from a camera device in real time.
[0261] The "video analysis means" is a means for processing the acquired video data and performing object detection and motion analysis.
[0262] The "abnormal behavior detection means" is a means for detecting specific abnormal behavior based on the analysis results of the video analysis means.
[0263] The "emotion recognition means" is a means for analyzing the user's facial expressions and voice and recognizing their emotions.
[0264] The "warning message generation means" is a means for generating a warning message according to the user's emotions using a generative AI model.
[0265] The "means for reading out a warning message" is a means for reading out the generated warning message by voice.
[0266] The "information sharing means" is a means for transmitting the generated warning message and abnormal behavior information to the administrator in real time.
[0267] This system is implemented by combining the following elements: video acquisition means, video analysis means, abnormal behavior detection means, emotion recognition means, warning message generation means, warning message reading means, and information sharing means.
[0268] System Configuration
[0269] 1. Video acquisition method:
[0270] The server acquires video data from the camera device in real time. For example, a home IP camera is connected to the server and transmits video using the RTSP protocol. This video data is stored in a buffer on the server.
[0271] 2. Video analysis methods:
[0272] The video analysis module in the server processes the acquired video data and performs object detection and motion analysis. Specifically, it uses deep learning algorithms such as YOLO and SSD to detect objects in the video frame in real time.
[0273] 3. Abnormal behavior detection methods:
[0274] The server's abnormal behavior detection module detects specific abnormal behavior based on the video analysis results and determines whether the behavior is abnormal by comparing the detected behavior pattern with a predefined abnormal behavior database.
[0275] 4. Emotion recognition means:
[0276] The server analyzes the user's facial expressions and voice and recognizes the user's emotions in real time using an emotion recognition module. For example, OpenFace can be used for facial expression recognition and OpenSMILE for voice analysis.
[0277] 5. Warning statement generation means:
[0278] The server detects abnormal behavior and generates an appropriate warning message using a generative AI model based on the emotion recognition results. Specifically, a prompt message is input into a generative model such as GPT-4 to generate the warning message.
[0279] Example prompt sentence:
[0280] "Playing with the washing machine is dangerous. Stop immediately."
[0281] 6. Warning message reading method:
[0282] The device (such as a smart speaker in the home or a public address system) uses a speech generation module to convert the warning text into a voice and read it aloud. A text-to-speech (TTS) engine is used to generate an audio file, which is then played on the device.
[0283] 7. Information sharing methods:
[0284] The server generates a warning message and sends the abnormal behavior information to the administrator in real time. The administrator (e.g., parent or police) receives the warning information on their smartphone or a dedicated app and can respond promptly.
[0285] Specific examples
[0286] Watching over children at home
[0287] The server acquires real-time video from cameras in the home and transmits the video data to the video analysis module.
[0288] The server's video analysis module analyzes the child's behavior of opening the washing machine door, and the abnormal behavior detection module detects this as abnormal behavior.
[0289] The server uses an emotion recognition module to analyze the child's facial expressions and tone of voice to obtain emotional information.
[0290] The server generates a warning message saying, "Playing with the washing machine is dangerous. Stop immediately," and adjusts the tone based on the emotion recognition results.
[0291] The device (smart speaker) reads out the warning message aloud, and at the same time, the server sends the warning information to the parent's smartphone, allowing the parent to check the information immediately.
[0292] Street crime prevention surveillance
[0293] The server acquires real-time video from cameras installed on the street and sends the video data to the video analysis module.
[0294] The server's video analysis module detects snatching attempts, and the abnormal behavior detection module determines this as abnormal behavior.
[0295] The server uses an emotion recognition module to analyze the facial expressions and tone of voice of people around it to obtain emotional information.
[0296] The server generates a warning message saying, "Purse snatching has been detected here. Please leave immediately," and adjusts the tone based on emotional information.
[0297] The terminal (street announcement system) reads out the warning message, and at the same time, the server transmits the warning information to the police or monitoring center, enabling immediate action.
[0298] In this way, a system can be built that can quickly detect abnormal behavior in real time, issue warnings, and share information.
[0299] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0300] Step 1:
[0301] The server receives video data from the camera device in real time. The input is the video stream from the camera, and the output is the video data stored in the server's buffer. Specifically, the server communicates with the IP camera using the RTSP protocol and stores the received data in a buffer.
[0302] Step 2:
[0303] The video analysis module in the server retrieves video data from the buffer and performs object detection and motion analysis. The input is the video data stored in the buffer, and the output is the analyzed object and motion information. Specifically, it uses a deep learning algorithm (such as YOLO or SSD) to detect objects in the video frame in real time.
[0304] Step 3:
[0305] The server's abnormal behavior detection module detects specific abnormal behavior based on the video analysis results. The input is the output of the video analysis module, and the output is the abnormal behavior detection results. Specifically, the analysis results are compared with a predefined abnormal behavior database to determine abnormal behavior. For example, a child trying to open the washing machine door is detected as abnormal behavior.
[0306] Step 4:
[0307] The server uses an emotion recognition module to analyze the user's facial expressions and tone of voice to recognize emotions. The input is the user's face and voice data obtained through video analysis, and the output is the recognized emotional information. Specific operations include running a facial expression recognition model (e.g., OpenFace) and a voice analysis model (e.g., OpenSMILE). It may recognize emotions such as fear or surprise from a child's face.
[0308] Step 5:
[0309] The server detects abnormal behavior and generates a warning message using a generative AI model based on the emotion recognition results. The input is the abnormal behavior detection results and emotion recognition results, and the output is the generated warning message. Specifically, a prompt message is input into a generative AI model such as GPT-4 to generate a warning message. For example, a warning message such as "Playing with the washing machine is dangerous. Stop immediately" is generated.
[0310] Step 6:
[0311] The device (such as a home smart speaker or a street announcement system) uses a voice generation module to read out the warning text. The input is the generated warning text, and the output is a voiced version of the warning text. Specifically, the generated warning text is input into a text-to-speech (TTS) engine, which generates an audio file and plays it back.
[0312] Step 7:
[0313] The server generates a warning message and sends the abnormal behavior information to the administrator in real time. The input is the warning message and abnormal behavior information, and the output is the warning information sent to the administrator's device. Specifically, the abnormal behavior and warning message are sent via push notification to the administrator's smartphone or dedicated app, allowing parents, police, etc. to respond quickly.
[0314] (Application example 2)
[0315] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0316] While conventional video surveillance systems can detect abnormal behavior, they are unable to generate appropriate warnings based on the user's emotions and share information quickly. As a result, appropriate responses to abnormal behavior are delayed, making it difficult to provide effective security and monitoring. To solve this problem, the present invention aims to provide a system that integrates abnormal behavior detection and user emotion recognition to generate appropriate warnings and share information.
[0317] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0318] In this invention, the server includes a video acquisition means, a video analysis means, an abnormal behavior detection means, an emotion recognition means, a warning message generation means, a warning message reading means, and an information sharing means. This allows the server to acquire video from a surveillance camera in real time and analyze the video data to detect abnormal behavior. It can also recognize emotions in real time from a user's facial expression and tone of voice and generate an appropriate warning message based on the emotion information. The generated warning message is read aloud, and the warning information and abnormal behavior information are further transmitted to an administrator in real time, enabling a prompt and effective response.
[0319] "Video acquisition means" refers to a device or function that acquires video data in real time from a surveillance camera, a home camera, or the like.
[0320] "Video analysis means" refers to modules and algorithms for processing acquired video data and performing object detection and motion analysis.
[0321] "Abnormal behavior detection means" refers to a module or device for detecting specific abnormal behavior based on the results of video analysis.
[0322] "Emotion recognition means" refers to modules or algorithms that analyze the user's facial expressions and tone of voice and recognize emotions in real time.
[0323] A "warning message generation means" is a device or software that uses a generative AI model to generate an appropriate warning message based on the results of emotion recognition.
[0324] The "means for reading out a warning message" is a device or function for reading out the generated warning message aloud.
[0325] The "information sharing means" is a device or system for transmitting the generated warning message and abnormal behavior information to the administrator in real time.
[0326] A "generative AI model" is an artificial intelligence model for generating natural language based on specific input data.
[0327] A "prompt sentence" is an input sentence that causes a generative AI model to generate a specific output.
[0328] The present invention relates to a system that integrates video acquisition, video analysis, abnormal behavior detection, emotion recognition, warning message generation, warning message reading, and information sharing. This system can analyze video data acquired from surveillance cameras and home cameras in real time to detect abnormal behavior. Furthermore, it has the ability to recognize user emotions, generate appropriate warning messages based on those emotions, and read them aloud. Furthermore, the generated warning messages and abnormal behavior information are sent to administrators in real time, enabling prompt and effective response.
[0329] System configuration:
[0330] 1. Video acquisition method
[0331] The server receives video data from surveillance cameras and home cameras in real time, stores the video in a buffer on the server, and then analyzes it.
[0332] 2. Video analysis methods
[0333] The server's video analysis module processes the captured video data and performs object detection and motion analysis. The server uses deep learning algorithms to analyze patterns of abnormal behavior.
[0334] 3. Abnormal Behavior Detection Methods
[0335] The server's abnormal behavior detection module determines abnormal behavior based on the video analysis results, such as falls, trespassing in specific areas, and dangerous behavior.
[0336] 4. Emotion recognition means
[0337] The server analyzes the user's facial expressions and tone of voice, and uses an emotion recognition module to recognize the user's emotions in real time. The recognized emotional information is reflected in the generation of warning messages.
[0338] 5. Warning statement generation means
[0339] The server detects abnormal behavior and generates an appropriate warning message based on the emotion recognition results using a generative AI model. The generated warning message is tailored to the user's emotions.
[0340] 6. Warning message reading method
[0341] The device (such as a smart speaker in the home or a street announcement system) uses a voice generation module to convert the warning message into audio and read it aloud. For example, a voice might say, "Falling is dangerous. Please sit down and rest."
[0342] 7. Information sharing methods
[0343] The server generates a warning message and sends abnormal behavior information to the administrator in real time. The user (administrator) receives the warning information on their smartphone or a dedicated app, allowing them to respond quickly.
[0344] Examples:
[0345] Elderly monitoring system:
[0346] When an elderly person falls, abnormal behavior is detected from camera footage and, based on emotion recognition, a gentle warning message is generated, such as "Falling is dangerous. Please sit down and rest." The warning message is read aloud by the smart speaker and this information is sent to family members' smartphones in real time, encouraging them to take prompt action.
[0347] Examples of prompts:
[0348] "If the user is sad and falls, create a warning message in a gentle tone encouraging them to rest."
[0349] "If the user has a neutral emotion, generate a warning message to prevent them from falling."
[0350] Hardware and software used:
[0351] The hardware used is a surveillance camera, a home camera, a cloud server, a smart speaker, and a smartphone, while the software used is a video analysis module, an emotion recognition module, a generative AI model, a warning message generation module, and a notification sending module.
[0352] This system efficiently carries out a series of processes, from real-time acquisition and analysis of video data, detection of abnormal behavior, emotion recognition, generation and voice reading of appropriate warning messages, and information sharing, thereby achieving effective security and monitoring.
[0353] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0354] Step 1:
[0355] The server acquires video data in real time from surveillance cameras and home cameras. The video is stored in the server's buffer. The input is real-time video data from the camera, and the output is video data stored in the server's buffer. Specifically, the server acquires video data frame by frame from the camera stream and stores it in memory.
[0356] Step 2:
[0357] The server's video analysis module processes the acquired video data and performs object detection and motion analysis. The input is the video data stored in the buffer, and the output is the analysis results (e.g., object position and motion information). Specifically, the server uses a deep learning algorithm to detect face and body movements from the video data.
[0358] Step 3:
[0359] The server's abnormal behavior detection module detects specific abnormal behavior based on the video analysis results. The input is the result data of object detection and motion analysis, and the output is the abnormal behavior detection results. Specifically, the server compares motion patterns and determines abnormal behavior such as falls or dangerous behavior.
[0360] Step 4:
[0361] The server's emotion recognition module analyzes the user's facial expressions and tone of voice to recognize emotions in real time. The input is facial and voice data, and the output is recognized emotional information. Specifically, deep learning is used to estimate emotions such as joy, anger, sadness, and happiness from the user's facial and voice data.
[0362] Step 5:
[0363] The server detects abnormal behavior and generates an appropriate warning message using a generative AI model based on the emotion recognition results. The inputs are the abnormal behavior detection results and emotion recognition results, and the output is the generated warning message. Specifically, the server inputs a prompt message into the generative AI model and generates an appropriate warning message according to the emotion. For example, the prompt message for the generative AI model could be, "If the user is sad and falls, please create a warning message in a gentle tone encouraging them to rest."
[0364] Step 6:
[0365] The device (such as a home smart speaker or a street announcement system) uses a voice generation module to convert the warning text into audio and read it out loud. The input is the generated warning text, and the output is a voice warning. Specifically, the device uses text-to-speech conversion technology to play back the generated warning text aloud.
[0366] Step 7:
[0367] The server sends the generated warning message and abnormal behavior information to the administrator in real time. The input is the warning message and abnormal behavior information, and the output is a notification to the administrator's smartphone or a dedicated app. Specifically, the server uses a notification sending module to send the warning information as an email or push notification.
[0368] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0369] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0370] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0371] [Second embodiment]
[0372] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0373] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0374] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0375] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0376] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0377] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0378] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0379] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0380] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0381] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0382] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0383] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0384] The present invention is a system that processes a series of steps including video acquisition, video analysis, abnormal behavior detection, warning message generation, warning message reading, and information sharing. An embodiment of this system is mainly composed of the following elements.
[0385] System Configuration
[0386] 1. Video acquisition method
[0387] The server acquires video data from surveillance cameras and home cameras in real time.
[0388] The camera has a real-time streaming function, and the captured images are stored in a buffer on the server.
[0389] 2. Video analysis methods
[0390] The video analysis module in the server processes the acquired video data and performs object detection and motion analysis.
[0391] The server uses existing deep learning algorithms to analyze patterns of abnormal behavior.
[0392] 3. Abnormal Behavior Detection Methods
[0393] The server's abnormal behavior detection module determines and detects specific abnormal behavior in real time based on the video analysis results.
[0394] For example, it identifies a child trying to get into a washing machine or a snatching incident on the street.
[0395] 4. Warning statement generation means
[0396] If the server detects abnormal behavior, it uses a generative AI model to generate an appropriate warning message.
[0397] The generated warning message will vary depending on the abnormal behavior detected.
[0398] 5. Warning message reading method
[0399] The device (home smart speaker, street announcement system) converts the generated warning text into audio using a voice generation module and reads it out to the target person.
[0400] For example, at home, a warning will be played saying, "Playing with the washing machine is dangerous. Stop immediately."
[0401] 6. Information sharing methods
[0402] The server generates a warning message and sends abnormal behavior information to an administrator (parent, police, store manager, etc.) in real time.
[0403] Users (parents or administrators) can receive warning information on their smartphones or dedicated apps and respond quickly.
[0404] Specific examples
[0405] Watching over children at home
[0406] The server acquires real-time video from cameras in the home and transmits the video data to the video analysis module.
[0407] The server's video analysis module analyzes the child's behavior of opening the washing machine door, and the abnormal behavior detection module detects this as abnormal behavior.
[0408] The server generates a warning message saying, "Playing with the washing machine is dangerous. Stop immediately," and sends it to a smart speaker in the home.
[0409] The device (smart speaker) reads out the warning message aloud, and at the same time, the server sends the warning information to the parent's smartphone, allowing the parent to check the information immediately.
[0410] Street crime prevention surveillance
[0411] The server acquires real-time video from surveillance cameras installed on the street and transmits the video data to the video analysis module.
[0412] The server's video analysis module detects snatching attempts, and the abnormal behavior detection module determines this as abnormal behavior.
[0413] The server generates a warning message saying, "Purse snatching has been detected here. Please leave immediately," and sends it to the street announcement system.
[0414] The terminal (street announcement system) reads out the warning message, and at the same time, the server transmits the warning information to the police or monitoring center, enabling immediate action.
[0415] Program processing explanation
[0416] The server acquires video data from the surveillance camera and stores it in a buffer.
[0417] A video analysis module in the server analyzes the video data stored in the buffer and identifies abnormal behavior.
[0418] Based on the analysis results, the server's abnormal behavior detection module detects abnormal behavior.
[0419] If abnormal behavior is detected, the server uses a generative AI model to generate an appropriate warning message.
[0420] The generated warning message is sent to a designated device (smart speaker, announcement system) and read aloud.
[0421] The warning message and abnormal behavior information are sent to the administrator in real time by the server, and the administrator (user) can check the warning information on a smartphone or other device.
[0422] In this way, the form for implementing the invention provides a system in which each step operates in conjunction with other steps to quickly detect abnormal behavior and issue appropriate warnings, thereby preventing crimes and accidents before they occur.
[0423] The processing flow will be explained below.
[0424] Step 1:
[0425] The server acquires video data from security cameras and home cameras in real time. This acquisition process involves capturing video at a specified frame rate and storing it in a buffer.
[0426] Step 2:
[0427] The server sends the acquired video data to the video analysis module, where the video data is converted into a format that allows real-time processing.
[0428] Step 3:
[0429] The server's video analytics module processes the video data and runs object detection and motion analysis algorithms, for example, to identify specific human movements or unusual object behavior.
[0430] Step 4:
[0431] The server's abnormal behavior detection module determines abnormal behavior based on the video analysis results. This module uses a trained deep learning model to detect, for example, a child trying to get into a washing machine as an abnormal behavior.
[0432] Step 5:
[0433] When the server detects abnormal behavior, it activates a generative AI model to generate an appropriate warning message. This generation process creates a contextual warning message based on the nature of the abnormal behavior. For example, a warning message might be generated that reads, "Playing with the washing machine is dangerous. Stop immediately."
[0434] Step 6:
[0435] The server then sends the generated warning message to a designated device, such as a smart speaker in the home or a street announcement system.
[0436] Step 7:
[0437] The device converts the received warning message into audio using a speech generation module and reads it out loud from the speaker. For example, a voice warning a child playing with a washing machine might be played saying, "Playing with a washing machine is dangerous. Stop immediately."
[0438] Step 8:
[0439] At the same time, the server sends the generated warning message and detailed information about abnormal behavior to the administrator in real time. The administrator (user) can receive the warning information on their smartphone or a dedicated application and respond quickly. For example, parents can receive a notification on their smartphone and immediately check on their child's safety.
[0440] In this way, the processing flow from step 1 to step 8 realizes a system in which abnormal behavior is quickly detected and appropriate warnings and information sharing are provided.
[0441] Example 1
[0442] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0443] In recent years, safety monitoring in homes and on the streets has become increasingly important, but conventional monitoring systems have had difficulty detecting abnormal behavior in real time and issuing appropriate warnings. Furthermore, they have been unable to quickly share information with administrators, making it difficult to take immediate action to prevent incidents and accidents. New technological solutions are needed to solve these issues.
[0444] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0445] In this invention, the server includes a video acquisition means, a video analysis means, an abnormal behavior detection means, a warning message generation means, a warning message generation means using a generative AI model, a real-time video processing means, and an information sharing means with enhanced functions. This enables real-time analysis of video data, immediate detection of abnormal behavior, dynamic generation of warning messages using the generative AI model, and rapid voice generation of the warning messages and immediate notification to an administrator.
[0446] "Video acquisition means" refers to a means for acquiring video data in real time from a surveillance camera or a home camera and storing that data.
[0447] The "video analysis means" is a means for processing the acquired video data and performing object detection and motion analysis. Specifically, it analyzes the video data using a deep learning algorithm.
[0448] "Abnormal behavior detection means" is a means for determining and detecting specific abnormal behavior in real time based on the results of video analysis.
[0449] The "warning message generating means" is a means for generating an appropriate warning message based on the detected abnormal behavior.
[0450] "Means for generating warning messages using a generative AI model" refers to means for dynamically generating warning messages in response to abnormal behavior using a generative AI model (e.g., natural language processing technology).
[0451] "Means for processing video in real time" refers to means for storing acquired video data in a buffer in real time and analyzing the data sequentially.
[0452] The "means for reading out a warning message" is a means for converting the generated warning message into voice using a voice generation module (for example, TTS technology) and reading it out to the target person.
[0453] The "information sharing means" is a means for transmitting the generated warning message and abnormal behavior information to a manager (e.g., parent, police, store manager, etc.) in real time.
[0454] A "terminal" is a device that reads out warning messages aloud, such as a smart speaker in the home or a street announcement system.
[0455] This invention is a system that processes a series of steps including video acquisition, video analysis, abnormal behavior detection, warning message generation, warning message reading, and information sharing. The embodiment of the invention will be explained mainly by dividing it into a server, a terminal, and a user.
[0456] Video acquisition method
[0457] The server acquires video data in real time from security cameras and home cameras using HTTP or RTSP protocols. The acquired video data is stored in the server's buffer. Specifically, the server acquires video frames step by step using the FFmpeg library and stores them in a buffer in memory.
[0458] Video analysis methods
[0459] The video analysis module in the server processes the real-time video data stored in the buffer. This analysis uses deep learning algorithms using TensorFlow and PyTorch. Specifically, models such as YOLO (You Only Look Once) and OpenPose are used for object detection and motion analysis. The video analysis module analyzes the person's pose for each frame and identifies their joint points. It also uses an object detection model to recognize objects held by the person.
[0460] Abnormal behavior detection method
[0461] The server's anomalous behavior detection module identifies anomalous behavior in real time from the analyzed video data. This detection module references a rule-based model that defines anomalous behavior patterns. For example, it can detect a child trying to get into a washing machine or a purse snatching attempt. The server identifies frame sequences that meet the conditions for anomalous behavior and marks them as anomalous based on that.
[0462] Warning statement generation means
[0463] If the server detects abnormal behavior, it uses a generative AI model to generate an appropriate warning message. For example, GPT-3 is used as this generative AI model. Based on the abnormal behavior information, the server inputs a prompt message into the generative AI model to generate a warning message. An example of a specific prompt message is, "A child has been detected entering the washing machine. Please generate an appropriate warning message." The generated warning message would be, "Playing with the washing machine is dangerous. Please stop immediately."
[0464] Warning message reading method
[0465] The device (smart speaker or street announcement system) uses a voice generation module to convert the generated warning text into audio and read it to the target person. Specifically, the warning text is converted into audio using a TTS (Text-to-Speech) API such as Amazon Polly or Google Text-to-Speech. The device converts the warning text received from the server into audio and reads it aloud from the speaker.
[0466] Information sharing means
[0467] The server sends the generated warning message and abnormal behavior information to an administrator (such as a parent, police, or store manager) in real time. This communication uses an HTTP-based notification API. The user (administrator) receives the warning information on their smartphone or a dedicated app, and can immediately check the content and take action. Specifically, the user receives the warning information using the real-time notification function within the app and can check the details.
[0468] Specific examples
[0469] Watching over children at home
[0470] The server receives real-time video footage from the home camera and stores the data in the server's buffer. The server's video analysis module analyzes the child's behavior when they open the washing machine door, and the abnormal behavior detection module detects this behavior as abnormal. The server uses a generative AI model to generate a warning message saying, "Playing with the washing machine is dangerous. Please stop immediately," and sends this message to the smart speaker. The device (smart speaker) then reads this warning message aloud. At the same time, the server sends the warning information to the parent's smartphone, where the parent can immediately check it.
[0471] Street crime prevention surveillance
[0472] The server acquires real-time video footage from street surveillance cameras and stores the video data in the server's buffer. The server's video analysis module detects snatching attempts, and the abnormal behavior detection module determines the behavior as abnormal. The server uses a generative AI model to generate a warning message stating, "A snatching attempt has been detected here. Please leave immediately," and sends it to the street announcement system. The device (street announcement system) reads this warning message aloud. At the same time, the server sends the warning information to the police or monitoring center, enabling immediate action.
[0473] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0474] Step 1:
[0475] Real-time video capture
[0476] The server obtains video data from a surveillance camera in real time. The input is the video stream from the surveillance camera. The server receives the video stream using HTTP or RTSP protocol, extracts this data frame by frame using the FFmpeg library, and stores it in a buffer in memory. The output is the video data in the buffer. Specifically, the server reads the IP address and connection settings of the surveillance camera, and obtains the data in real time after establishing a connection.
[0477] Step 2:
[0478] Video data analysis
[0479] The video analysis module in the server processes real-time video data stored in the buffer. The input is the video data in the buffer. The server analyzes the video using deep learning algorithms (such as YOLO or OpenPose) using TensorFlow or PyTorch. GPU acceleration is used to perform object detection and motion analysis at high speed. The output is the analysis results, which are the position information of specific objects or people in each frame. Specifically, the server inputs each frame of data into the deep learning model, obtains and stores the analysis results.
[0480] Step 3:
[0481] Abnormal behavior detection
[0482] The server's abnormal behavior detection module identifies abnormal behavior in real time based on analyzed video data. The input is the output of the video analysis module. The server references a rule-based model that defines abnormal behavior patterns in advance and determines whether it matches each frame of data. The output is information about the frame in which abnormal behavior was detected. Specifically, the server checks the conditions under which abnormal behavior occurs (e.g., a specific behavior occurs consecutively within a specific period of time) and determines that an abnormality exists based on that.
[0483] Step 4:
[0484] Generate a warning message
[0485] When the server detects abnormal behavior, it uses a generative AI model (for example, GPT-3) to generate a warning message. The input is information about the abnormal behavior. Based on this information, the server inputs a prompt message into the generative AI model. The output is the generated warning message. An example of a specific prompt message is, "A snatching attempt has been detected. Please generate an appropriate warning message." The generative model generates a warning message based on this message, and outputs the message, "A snatching attempt has been detected here. Please leave immediately."
[0486] Step 5:
[0487] Reading out warning messages
[0488] The device (such as a home smart speaker or a street announcement system) converts the generated warning text into audio using a voice generation module (for example, Amazon Polly or Google Text-to-Speech) and reads it aloud to the target person. The input is the warning text sent from the server. The device calls the TTS API, generates the warning text as an audio file, and outputs the sound from the speaker. The output is the audio of the warning text. Specifically, the device requests the received warning text from the TTS API and plays the generated audio data.
[0489] Step 6:
[0490] Information sharing
[0491] The server sends the generated warning message and abnormal behavior information to an administrator (parent, police, store manager, etc.) in real time. The input is the abnormal behavior information and the generated warning message. The server uses an HTTP-based notification API to send the warning information to the administrator's smartphone or dedicated app. The output is a notification message to the administrator. The user (administrator) receives the warning information on their smartphone or dedicated app, and can immediately check the content and take action. Specifically, the server reads the administrator's connection information and sends the warning data through the notification API.
[0492] (Application example 1)
[0493] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0494] Conventional crime prevention and surveillance systems have difficulty quickly detecting abnormal behavior and taking appropriate countermeasures immediately. Furthermore, delays in transmitting warning information can lead to problems with preventing crimes and accidents. To address these issues, there is a need for a system that can detect abnormal behavior in real time, issue warnings quickly, and share information with relevant parties immediately.
[0495] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0496] In this invention, the server includes a video acquisition means, a video analysis means, an abnormal behavior detection means, a warning message generation means, a warning message reading means, an information sharing means, a means for displaying a warning message on a display of the smart glasses, a means for issuing a voice warning, and a means for sharing abnormal behavior information with a management center, thereby enabling real-time abnormal behavior detection, prompt issuance of a warning, and prompt information sharing.
[0497] "Video acquisition means" refers to means for acquiring video data from a camera or other video capture device in real time.
[0498] The "video analysis means" is a means for processing the acquired video data and performing object detection and motion analysis.
[0499] "Abnormal behavior detection means" is a means of determining and detecting specific abnormal behavior in real time based on the results of video analysis.
[0500] The "warning message generation means" is a means for generating an appropriate warning message using a generative AI model when abnormal behavior is detected.
[0501] The "means for reading out a warning message" is a means for converting the generated warning message into voice using a voice generation module and reading it out loud to the target person.
[0502] The "information sharing means" is a means for sending the generated warning message and abnormal behavior information to the administrator in real time, enabling a prompt response.
[0503] The "means for displaying a warning message on the display of the smart glasses" refers to a means for visually displaying a warning message on the display of the smart glasses.
[0504] The "means for issuing an audio warning" is a means for issuing an audio warning when abnormal behavior is detected.
[0505] The "means for sharing abnormal behavior information with the management center" refers to a means for transmitting information on detected abnormal behavior to the management center in real time and quickly sharing it with relevant parties.
[0506] The present invention is a security system that processes the acquisition and analysis of camera footage, detection of abnormal behavior, generation and reading of warning messages, and information sharing in a series of steps. This system is mainly composed of a video acquisition means, a video analysis means, an abnormal behavior detection means, a warning message generation means, a warning message reading means, an information sharing means, a means for displaying a warning message on a smart glasses display, a means for issuing an audio warning, and a means for sharing abnormal behavior information with a management center.
[0507] Specifically, various hardware and software are used as follows:
[0508] Hardware:
[0509] Smart glasses: Devices that have a camera, display, and speaker (e.g., Google Glass, Vuzix Blade).
[0510] Camera: Surveillance cameras used on the street or in homes (e.g., network cameras).
[0511] Server: A computer system that processes and analyzes video data, detects abnormal behavior, and generates and distributes warning messages.
[0512] software:
[0513] Video Analytics Module: Object detection and motion analysis software using deep learning algorithms.
[0514] Abnormal behavior detection module: Software that identifies abnormal behavior in real time based on video analysis results.
[0515] Generative AI model: Natural language generation software that generates appropriate warning statements based on abnormal behavior.
[0516] Speech Generation Module: Text-to-Speech (TTS) software for vocalizing the generated warning text.
[0517] Information sharing module: Software that sends warning messages and abnormal behavior information to the management center and relevant parties in real time.
[0518] Data processing and calculation:
[0519] The server processes the acquired video data in real time. The video analysis module analyzes the video data sent from the camera and detects abnormal behavior, using existing deep learning methods. When abnormal behavior is detected, the generative AI model generates an appropriate warning message, which is then converted into audio by the voice generation module and sent to the smart glasses and other devices. The warning message is immediately displayed on the smart glasses' display, and a voice warning is also issued at the same time. In addition, abnormal behavior information is sent to the management center via the information sharing module, and relevant parties are immediately notified.
[0520] Examples:
[0521] For example, consider a police officer wearing smart glasses while patrolling the streets. During the patrol, the camera in the glasses captures real-time footage of the street. The footage is sent to a server, where the video analysis module detects violent acts. The generative AI model then generates a warning message, "Violent acts detected, please leave the scene," which is displayed and read aloud on the smart glasses. At the same time, this warning information is sent to a control center, allowing assistance to be quickly dispatched.
[0522] Example prompt sentence:
[0523] For example, by inputting the following prompt sentence into the generative AI model, an appropriate warning sentence will be generated.
[0524] Generate appropriate warning messages when the following actions are detected:
[0525] Street violence: Warning: "Violence detected, please respond immediately."
[0526] Snatching: Warning: "Snatching detected, please leave the scene."
[0527] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0528] Step 1:
[0529] The server acquires video data from the camera in real time. The input is the video stream from the camera, and the output video data is stored in the server's buffer.
[0530] Step 2:
[0531] The server's video analysis means processes the acquired video data. The input is the video data stored in the buffer, and object detection and motion analysis are performed, and the analysis results are output.
[0532] Step 3:
[0533] The server's abnormal behavior detection means detects specific abnormal behavior in real time based on the video analysis results. The input is the video analysis results, and the output is a determination of whether abnormal behavior has been detected.
[0534] Step 4:
[0535] The server generates a warning message using a generative AI model. The input is information about the detected abnormal behavior, and the generated warning message is output. In this case, the server uses a prompt message to generate an appropriate warning message.
[0536] Step 5:
[0537] The server sends the generated warning text to the terminal. The input is the generated warning text, which is sent to smart glasses or other terminals.
[0538] Step 6:
[0539] The terminal (smart glasses) displays a warning message on the display and issues a voice warning. The input is the warning message received from the server, and the output is a displayed warning message and a voice alert.
[0540] Step 7:
[0541] The server's information sharing means sends abnormal behavior information and warning messages to the management center. The input is the warning message and abnormal behavior information, and the output is shared in real time with the management center and relevant parties.
[0542] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0543] The present invention is a system that combines the functions of video acquisition, video analysis, abnormal behavior detection, warning message generation, warning message reading, information sharing, and emotion recognition, and links these functions together to provide appropriate warnings and share information while taking into account the user's emotions. An embodiment of this system is configured as follows.
[0544] System Configuration
[0545] 1. Video acquisition method
[0546] The server acquires video data from surveillance cameras and home cameras in real time.
[0547] The video is stored in a buffer on the server and is then analyzed.
[0548] 2. Video analysis methods
[0549] The video analysis module in the server processes the acquired video data and performs object detection and motion analysis.
[0550] The server uses deep learning algorithms to analyze patterns of abnormal behavior.
[0551] 3. Abnormal Behavior Detection Methods
[0552] The server's abnormal behavior detection module determines abnormal behavior based on the results of video analysis.
[0553] For example, it can detect a child trying to get into a washing machine or a snatch-and-run theft on the street.
[0554] 4. Emotion recognition means
[0555] The server analyzes the user's facial expressions and tone of voice and uses an emotion recognition module to recognize the user's emotions in real time.
[0556] The recognized emotional information is reflected in the generation of warning messages.
[0557] 5. Warning statement generation means
[0558] The server detects abnormal behavior and generates an appropriate warning message using a generative AI model based on the emotion recognition results.
[0559] The generated warning message will be based on the user's emotions. For example, a gentle warning message will be generated to calm the user.
[0560] 6. Warning message reading method
[0561] The device (home smart speaker, street announcement system) uses a voice generation module to convert the warning text into audio and read it out loud.
[0562] For example, a voice may be heard saying to a child, "Playing with the washing machine is dangerous. Stop immediately."
[0563] 7. Information sharing methods
[0564] The server sends the generated warning message and abnormal behavior information to the administrator in real time.
[0565] Users (parents or administrators) can receive warning information on their smartphones or dedicated apps and respond quickly.
[0566] Specific examples
[0567] Watching over children at home
[0568] The server acquires real-time video from cameras in the home and transmits the video data to the video analysis module.
[0569] The server's video analysis module analyzes the child's behavior of opening the washing machine door, and the abnormal behavior detection module detects this as abnormal behavior.
[0570] The server uses an emotion recognition module to analyze the child's facial expressions and tone of voice to obtain emotional information.
[0571] The server generates a warning message saying, "Playing with the washing machine is dangerous. Stop immediately," and adjusts the tone based on the emotion recognition results.
[0572] The device (smart speaker) reads out the warning message aloud, and at the same time, the server sends the warning information to the parent's smartphone, allowing the parent to check the information immediately.
[0573] Street crime prevention surveillance
[0574] The server acquires real-time video from surveillance cameras installed on the street and transmits the video data to the video analysis module.
[0575] The server's video analysis module detects snatching attempts, and the abnormal behavior detection module determines this as abnormal behavior.
[0576] The server uses an emotion recognition module to analyze the facial expressions and tone of voice of people around it to obtain emotional information.
[0577] The server generates a warning message saying, "Purse snatching has been detected here. Please leave immediately," and adjusts the tone based on emotional information.
[0578] The terminal (street announcement system) reads out the warning message, and at the same time, the server transmits the warning information to the police or monitoring center, enabling immediate action.
[0579] Program processing explanation
[0580] The server acquires video data from the surveillance camera and stores it in a buffer.
[0581] A video analysis module in the server analyzes the video data and identifies abnormal behavior.
[0582] When abnormal behavior is detected, the server activates the emotion recognition module and analyzes the user's emotional information.
[0583] Based on the emotional information, the server uses a generative AI model to generate an appropriate warning message.
[0584] The generated warning message is sent to the specified device and read aloud.
[0585] The warning message and abnormal behavior information are sent to the administrator in real time by the server, and the administrator (user) can check the warning information on a smartphone or other device.
[0586] In this way, it is possible to provide a system that can quickly detect abnormal behavior, recognize emotions, issue appropriate warnings, and share information.
[0587] The processing flow will be explained below.
[0588] Step 1:
[0589] The server acquires video data in real time from home cameras and street surveillance cameras. This video data is broken down into multiple frames and stored in the server's temporary memory (buffer).
[0590] Step 2:
[0591] The server sends the acquired video data to the video analysis module, which converts the video data into a format that can be processed in real time and then provides it for analysis.
[0592] Step 3:
[0593] The video analysis module in the server processes the video data and applies object detection and motion analysis algorithms to analyze the movement, shape, and position of people and objects, preparing to identify abnormal behavior.
[0594] Step 4:
[0595] The server's abnormal behavior detection module determines abnormal behavior based on the video analysis results. For example, it detects a child trying to get into a washing machine or a person moving abnormally in a specific area.
[0596] Step 5:
[0597] After the server detects abnormal behavior, it activates the emotion recognition module, which analyzes the user's facial expressions and tone of voice from the video data to recognize their emotions.
[0598] Step 6:
[0599] The server uses a generative AI model to generate appropriate warning messages based on the emotion recognition results. During this generation process, the tone and content of the warning messages are adjusted to match the user's emotions.
[0600] Step 7:
[0601] The server generates a warning message and sends it to a designated device, such as a smart speaker in the home or an announcement system on the street.
[0602] Step 8:
[0603] The warning message received by the device is converted into voice by a voice generation module and actually read aloud from a speaker. For example, at home, a warning such as "Playing with the washing machine is dangerous. Stop immediately" is played in a soft tone.
[0604] Step 9:
[0605] At the same time, the server sends the generated warning message and detailed information about the abnormal behavior to the administrator in real time. The administrator (user) receives the warning information on their smartphone or a dedicated application, enabling them to respond quickly. For example, a parent could receive a notification on their smartphone and check on their child's safety.
[0606] In this way, the processing flow from step 1 to step 9 provides a system that quickly detects abnormal behavior and provides appropriate warnings and information sharing that take the user's emotions into consideration.
[0607] Example 2
[0608] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0609] Conventional surveillance systems were able to detect abnormal behavior by acquiring and analyzing video data, but they lacked the ability to issue prompt and appropriate warnings and share information while taking into account the user's emotions. As a result, when abnormal behavior occurred, responses were sometimes delayed, making it difficult to implement effective preventative measures.
[0610] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a video acquisition means, a video analysis means, an abnormal behavior detection means, an emotion recognition means, a warning message generation means, a warning message reading means, and an information sharing means. This makes it possible to issue a prompt and accurate warning that takes into account the user's emotions and to share abnormal behavior information in real time.
[0611] The "image acquisition means" is a means for acquiring image data from a camera device in real time.
[0612] The "video analysis means" is a means for processing the acquired video data and performing object detection and motion analysis.
[0613] The "abnormal behavior detection means" is a means for detecting specific abnormal behavior based on the analysis results of the video analysis means.
[0614] The "emotion recognition means" is a means for analyzing the user's facial expressions and voice and recognizing their emotions.
[0615] The "warning message generation means" is a means for generating a warning message according to the user's emotions using a generative AI model.
[0616] The "means for reading out a warning message" is a means for reading out the generated warning message by voice.
[0617] The "information sharing means" is a means for transmitting the generated warning message and abnormal behavior information to the administrator in real time.
[0618] This system is implemented by combining the following elements: video acquisition means, video analysis means, abnormal behavior detection means, emotion recognition means, warning message generation means, warning message reading means, and information sharing means.
[0619] System Configuration
[0620] 1. Video acquisition method:
[0621] The server acquires video data from the camera device in real time. For example, a home IP camera is connected to the server and transmits video using the RTSP protocol. This video data is stored in a buffer on the server.
[0622] 2. Video analysis methods:
[0623] The video analysis module in the server processes the acquired video data and performs object detection and motion analysis. Specifically, it uses deep learning algorithms such as YOLO and SSD to detect objects in the video frame in real time.
[0624] 3. Abnormal behavior detection methods:
[0625] The server's abnormal behavior detection module detects specific abnormal behavior based on the video analysis results and determines whether the behavior is abnormal by comparing the detected behavior pattern with a predefined abnormal behavior database.
[0626] 4. Emotion recognition means:
[0627] The server analyzes the user's facial expressions and voice and recognizes the user's emotions in real time using an emotion recognition module. For example, OpenFace can be used for facial expression recognition and OpenSMILE for voice analysis.
[0628] 5. Warning statement generation means:
[0629] The server detects abnormal behavior and generates an appropriate warning message using a generative AI model based on the emotion recognition results. Specifically, a prompt message is input into a generative model such as GPT-4 to generate the warning message.
[0630] Example prompt sentence:
[0631] "Playing with the washing machine is dangerous. Stop immediately."
[0632] 6. Warning message reading method:
[0633] The device (such as a smart speaker in the home or a public address system) uses a speech generation module to convert the warning text into a voice and read it aloud. A text-to-speech (TTS) engine is used to generate an audio file, which is then played on the device.
[0634] 7. Information sharing methods:
[0635] The server generates a warning message and sends the abnormal behavior information to the administrator in real time. The administrator (e.g., parent or police) receives the warning information on their smartphone or a dedicated app and can respond promptly.
[0636] Specific examples
[0637] Watching over children at home
[0638] The server acquires real-time video from cameras in the home and transmits the video data to the video analysis module.
[0639] The server's video analysis module analyzes the child's behavior of opening the washing machine door, and the abnormal behavior detection module detects this as abnormal behavior.
[0640] The server uses an emotion recognition module to analyze the child's facial expressions and tone of voice to obtain emotional information.
[0641] The server generates a warning message saying, "Playing with the washing machine is dangerous. Stop immediately," and adjusts the tone based on the emotion recognition results.
[0642] The device (smart speaker) reads out the warning message aloud, and at the same time, the server sends the warning information to the parent's smartphone, allowing the parent to check the information immediately.
[0643] Street crime prevention surveillance
[0644] The server acquires real-time video from cameras installed on the street and sends the video data to the video analysis module.
[0645] The server's video analysis module detects snatching attempts, and the abnormal behavior detection module determines this as abnormal behavior.
[0646] The server uses an emotion recognition module to analyze the facial expressions and tone of voice of people around it to obtain emotional information.
[0647] The server generates a warning message saying, "Purse snatching has been detected here. Please leave immediately," and adjusts the tone based on emotional information.
[0648] The terminal (street announcement system) reads out the warning message, and at the same time, the server transmits the warning information to the police or monitoring center, enabling immediate action.
[0649] In this way, a system can be built that can quickly detect abnormal behavior in real time, issue warnings, and share information.
[0650] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0651] Step 1:
[0652] The server receives video data from the camera device in real time. The input is the video stream from the camera, and the output is the video data stored in the server's buffer. Specifically, the server communicates with the IP camera using the RTSP protocol and stores the received data in a buffer.
[0653] Step 2:
[0654] The video analysis module in the server retrieves video data from the buffer and performs object detection and motion analysis. The input is the video data stored in the buffer, and the output is the analyzed object and motion information. Specifically, it uses a deep learning algorithm (such as YOLO or SSD) to detect objects in the video frame in real time.
[0655] Step 3:
[0656] The server's abnormal behavior detection module detects specific abnormal behavior based on the video analysis results. The input is the output of the video analysis module, and the output is the abnormal behavior detection results. Specifically, the analysis results are compared with a predefined abnormal behavior database to determine abnormal behavior. For example, a child trying to open the washing machine door is detected as abnormal behavior.
[0657] Step 4:
[0658] The server uses an emotion recognition module to analyze the user's facial expressions and tone of voice to recognize emotions. The input is the user's face and voice data obtained through video analysis, and the output is the recognized emotional information. Specific operations include running a facial expression recognition model (e.g., OpenFace) and a voice analysis model (e.g., OpenSMILE). It may recognize emotions such as fear or surprise from a child's face.
[0659] Step 5:
[0660] The server detects abnormal behavior and generates a warning message using a generative AI model based on the emotion recognition results. The input is the abnormal behavior detection results and emotion recognition results, and the output is the generated warning message. Specifically, a prompt message is input into a generative AI model such as GPT-4 to generate a warning message. For example, a warning message such as "Playing with the washing machine is dangerous. Stop immediately" is generated.
[0661] Step 6:
[0662] The device (such as a home smart speaker or a street announcement system) uses a voice generation module to read out the warning text. The input is the generated warning text, and the output is a voiced version of the warning text. Specifically, the generated warning text is input into a text-to-speech (TTS) engine, which generates an audio file and plays it back.
[0663] Step 7:
[0664] The server generates a warning message and sends the abnormal behavior information to the administrator in real time. The input is the warning message and abnormal behavior information, and the output is the warning information sent to the administrator's device. Specifically, the abnormal behavior and warning message are sent via push notification to the administrator's smartphone or dedicated app, allowing parents, police, etc. to respond quickly.
[0665] (Application example 2)
[0666] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0667] While conventional video surveillance systems can detect abnormal behavior, they are unable to generate appropriate warnings based on the user's emotions and share information quickly. As a result, appropriate responses to abnormal behavior are delayed, making it difficult to provide effective security and monitoring. To solve this problem, the present invention aims to provide a system that integrates abnormal behavior detection and user emotion recognition to generate appropriate warnings and share information.
[0668] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0669] In this invention, the server includes a video acquisition means, a video analysis means, an abnormal behavior detection means, an emotion recognition means, a warning message generation means, a warning message reading means, and an information sharing means. This allows the server to acquire video from a surveillance camera in real time and analyze the video data to detect abnormal behavior. It can also recognize emotions in real time from a user's facial expression and tone of voice and generate an appropriate warning message based on the emotion information. The generated warning message is read aloud, and the warning information and abnormal behavior information are further transmitted to an administrator in real time, enabling a prompt and effective response.
[0670] "Video acquisition means" refers to a device or function that acquires video data in real time from a surveillance camera, a home camera, or the like.
[0671] "Video analysis means" refers to modules and algorithms for processing acquired video data and performing object detection and motion analysis.
[0672] "Abnormal behavior detection means" refers to a module or device for detecting specific abnormal behavior based on the results of video analysis.
[0673] "Emotion recognition means" refers to modules or algorithms that analyze the user's facial expressions and tone of voice and recognize emotions in real time.
[0674] A "warning message generation means" is a device or software that uses a generative AI model to generate an appropriate warning message based on the results of emotion recognition.
[0675] The "means for reading out a warning message" is a device or function for reading out the generated warning message aloud.
[0676] The "information sharing means" is a device or system for transmitting the generated warning message and abnormal behavior information to the administrator in real time.
[0677] A "generative AI model" is an artificial intelligence model for generating natural language based on specific input data.
[0678] A "prompt sentence" is an input sentence that causes a generative AI model to generate a specific output.
[0679] The present invention relates to a system that integrates video acquisition, video analysis, abnormal behavior detection, emotion recognition, warning message generation, warning message reading, and information sharing. This system can analyze video data acquired from surveillance cameras and home cameras in real time to detect abnormal behavior. Furthermore, it has the ability to recognize user emotions, generate appropriate warning messages based on those emotions, and read them aloud. Furthermore, the generated warning messages and abnormal behavior information are sent to administrators in real time, enabling prompt and effective response.
[0680] System configuration:
[0681] 1. Video acquisition method
[0682] The server receives video data from surveillance cameras and home cameras in real time, stores the video in a buffer on the server, and then analyzes it.
[0683] 2. Video analysis methods
[0684] The server's video analysis module processes the captured video data and performs object detection and motion analysis. The server uses deep learning algorithms to analyze patterns of abnormal behavior.
[0685] 3. Abnormal Behavior Detection Methods
[0686] The server's abnormal behavior detection module determines abnormal behavior based on the video analysis results, such as falls, trespassing in specific areas, and dangerous behavior.
[0687] 4. Emotion recognition means
[0688] The server analyzes the user's facial expressions and tone of voice, and uses an emotion recognition module to recognize the user's emotions in real time. The recognized emotional information is reflected in the generation of warning messages.
[0689] 5. Warning statement generation means
[0690] The server detects abnormal behavior and generates an appropriate warning message based on the emotion recognition results using a generative AI model. The generated warning message is tailored to the user's emotions.
[0691] 6. Warning message reading method
[0692] The device (such as a smart speaker in the home or a street announcement system) uses a voice generation module to convert the warning message into audio and read it aloud. For example, a voice might say, "Falling is dangerous. Please sit down and rest."
[0693] 7. Information sharing methods
[0694] The server generates a warning message and sends abnormal behavior information to the administrator in real time. The user (administrator) receives the warning information on their smartphone or a dedicated app, allowing them to respond quickly.
[0695] Examples:
[0696] Elderly monitoring system:
[0697] When an elderly person falls, abnormal behavior is detected from camera footage and, based on emotion recognition, a gentle warning message is generated, such as "Falling is dangerous. Please sit down and rest." The warning message is read aloud by the smart speaker and this information is sent to family members' smartphones in real time, encouraging them to take prompt action.
[0698] Examples of prompts:
[0699] "If the user is sad and falls, create a warning message in a gentle tone encouraging them to rest."
[0700] "If the user has a neutral emotion, generate a warning message to prevent them from falling."
[0701] Hardware and software used:
[0702] The hardware used is a surveillance camera, a home camera, a cloud server, a smart speaker, and a smartphone, while the software used is a video analysis module, an emotion recognition module, a generative AI model, a warning message generation module, and a notification sending module.
[0703] This system efficiently carries out a series of processes, from real-time acquisition and analysis of video data, detection of abnormal behavior, emotion recognition, generation and voice reading of appropriate warning messages, and information sharing, thereby achieving effective security and monitoring.
[0704] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0705] Step 1:
[0706] The server acquires video data in real time from surveillance cameras and home cameras. The video is stored in the server's buffer. The input is real-time video data from the camera, and the output is video data stored in the server's buffer. Specifically, the server acquires video data frame by frame from the camera stream and stores it in memory.
[0707] Step 2:
[0708] The server's video analysis module processes the acquired video data and performs object detection and motion analysis. The input is the video data stored in the buffer, and the output is the analysis results (e.g., object position and motion information). Specifically, the server uses a deep learning algorithm to detect face and body movements from the video data.
[0709] Step 3:
[0710] The server's abnormal behavior detection module detects specific abnormal behavior based on the video analysis results. The input is the result data of object detection and motion analysis, and the output is the abnormal behavior detection results. Specifically, the server compares motion patterns and determines abnormal behavior such as falls or dangerous behavior.
[0711] Step 4:
[0712] The server's emotion recognition module analyzes the user's facial expressions and tone of voice to recognize emotions in real time. The input is facial and voice data, and the output is recognized emotional information. Specifically, deep learning is used to estimate emotions such as joy, anger, sadness, and happiness from the user's facial and voice data.
[0713] Step 5:
[0714] The server detects abnormal behavior and generates an appropriate warning message using a generative AI model based on the emotion recognition results. The inputs are the abnormal behavior detection results and emotion recognition results, and the output is the generated warning message. Specifically, the server inputs a prompt message into the generative AI model and generates an appropriate warning message according to the emotion. For example, the prompt message for the generative AI model could be, "If the user is sad and falls, please create a warning message in a gentle tone encouraging them to rest."
[0715] Step 6:
[0716] The device (such as a home smart speaker or a street announcement system) uses a voice generation module to convert the warning text into audio and read it out loud. The input is the generated warning text, and the output is a voice warning. Specifically, the device uses text-to-speech conversion technology to play back the generated warning text aloud.
[0717] Step 7:
[0718] The server sends the generated warning message and abnormal behavior information to the administrator in real time. The input is the warning message and abnormal behavior information, and the output is a notification to the administrator's smartphone or a dedicated app. Specifically, the server uses a notification sending module to send the warning information as an email or push notification.
[0719] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0720] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0721] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0722] [Third embodiment]
[0723] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0724] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0725] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0726] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0727] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0728] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0729] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0730] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0731] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0732] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0733] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0734] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0735] The present invention is a system that processes a series of steps including video acquisition, video analysis, abnormal behavior detection, warning message generation, warning message reading, and information sharing. An embodiment of this system is mainly composed of the following elements.
[0736] System Configuration
[0737] 1. Video acquisition method
[0738] The server acquires video data from surveillance cameras and home cameras in real time.
[0739] The camera has a real-time streaming function, and the captured images are stored in a buffer on the server.
[0740] 2. Video analysis methods
[0741] The video analysis module in the server processes the acquired video data and performs object detection and motion analysis.
[0742] The server uses existing deep learning algorithms to analyze patterns of abnormal behavior.
[0743] 3. Abnormal Behavior Detection Methods
[0744] The server's abnormal behavior detection module determines and detects specific abnormal behavior in real time based on the video analysis results.
[0745] For example, it identifies a child trying to get into a washing machine or a snatching incident on the street.
[0746] 4. Warning statement generation means
[0747] If the server detects abnormal behavior, it uses a generative AI model to generate an appropriate warning message.
[0748] The generated warning message will vary depending on the abnormal behavior detected.
[0749] 5. Warning message reading method
[0750] The device (home smart speaker, street announcement system) converts the generated warning text into audio using a voice generation module and reads it out to the target person.
[0751] For example, at home, a warning will be played saying, "Playing with the washing machine is dangerous. Stop immediately."
[0752] 6. Information sharing methods
[0753] The server generates a warning message and sends abnormal behavior information to an administrator (parent, police, store manager, etc.) in real time.
[0754] Users (parents or administrators) can receive warning information on their smartphones or dedicated apps and respond quickly.
[0755] Specific examples
[0756] Watching over children at home
[0757] The server acquires real-time video from cameras in the home and transmits the video data to the video analysis module.
[0758] The server's video analysis module analyzes the child's behavior of opening the washing machine door, and the abnormal behavior detection module detects this as abnormal behavior.
[0759] The server generates a warning message saying, "Playing with the washing machine is dangerous. Stop immediately," and sends it to a smart speaker in the home.
[0760] The device (smart speaker) reads out the warning message aloud, and at the same time, the server sends the warning information to the parent's smartphone, allowing the parent to check the information immediately.
[0761] Street crime prevention surveillance
[0762] The server acquires real-time video from surveillance cameras installed on the street and transmits the video data to the video analysis module.
[0763] The server's video analysis module detects snatching attempts, and the abnormal behavior detection module determines this as abnormal behavior.
[0764] The server generates a warning message saying, "Purse snatching has been detected here. Please leave immediately," and sends it to the street announcement system.
[0765] The terminal (street announcement system) reads out the warning message, and at the same time, the server transmits the warning information to the police or monitoring center, enabling immediate action.
[0766] Program processing explanation
[0767] The server acquires video data from the surveillance camera and stores it in a buffer.
[0768] A video analysis module in the server analyzes the video data stored in the buffer and identifies abnormal behavior.
[0769] Based on the analysis results, the server's abnormal behavior detection module detects abnormal behavior.
[0770] If abnormal behavior is detected, the server uses a generative AI model to generate an appropriate warning message.
[0771] The generated warning message is sent to a designated device (smart speaker, announcement system) and read aloud.
[0772] The warning message and abnormal behavior information are sent to the administrator in real time by the server, and the administrator (user) can check the warning information on a smartphone or other device.
[0773] In this way, the form for implementing the invention provides a system in which each step operates in conjunction with other steps to quickly detect abnormal behavior and issue appropriate warnings, thereby preventing crimes and accidents before they occur.
[0774] The processing flow will be explained below.
[0775] Step 1:
[0776] The server acquires video data from security cameras and home cameras in real time. This acquisition process involves capturing video at a specified frame rate and storing it in a buffer.
[0777] Step 2:
[0778] The server sends the acquired video data to the video analysis module, where the video data is converted into a format that allows real-time processing.
[0779] Step 3:
[0780] The server's video analytics module processes the video data and runs object detection and motion analysis algorithms, for example, to identify specific human movements or unusual object behavior.
[0781] Step 4:
[0782] The server's abnormal behavior detection module determines abnormal behavior based on the video analysis results. This module uses a trained deep learning model to detect, for example, a child trying to get into a washing machine as an abnormal behavior.
[0783] Step 5:
[0784] When the server detects abnormal behavior, it activates a generative AI model to generate an appropriate warning message. This generation process creates a contextual warning message based on the nature of the abnormal behavior. For example, a warning message might be generated that reads, "Playing with the washing machine is dangerous. Stop immediately."
[0785] Step 6:
[0786] The server then sends the generated warning message to a designated device, such as a smart speaker in the home or a street announcement system.
[0787] Step 7:
[0788] The device converts the received warning message into audio using a speech generation module and reads it out loud from the speaker. For example, a voice warning a child playing with a washing machine might be played saying, "Playing with a washing machine is dangerous. Stop immediately."
[0789] Step 8:
[0790] At the same time, the server sends the generated warning message and detailed information about abnormal behavior to the administrator in real time. The administrator (user) can receive the warning information on their smartphone or a dedicated application and respond quickly. For example, parents can receive a notification on their smartphone and immediately check on their child's safety.
[0791] In this way, the processing flow from step 1 to step 8 realizes a system in which abnormal behavior is quickly detected and appropriate warnings and information sharing are provided.
[0792] Example 1
[0793] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0794] In recent years, safety monitoring in homes and on the streets has become increasingly important, but conventional monitoring systems have had difficulty detecting abnormal behavior in real time and issuing appropriate warnings. Furthermore, they have been unable to quickly share information with administrators, making it difficult to take immediate action to prevent incidents and accidents. New technological solutions are needed to solve these issues.
[0795] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0796] In this invention, the server includes a video acquisition means, a video analysis means, an abnormal behavior detection means, a warning message generation means, a warning message generation means using a generative AI model, a real-time video processing means, and an information sharing means with enhanced functions. This enables real-time analysis of video data, immediate detection of abnormal behavior, dynamic generation of warning messages using the generative AI model, and rapid voice generation of the warning messages and immediate notification to an administrator.
[0797] "Video acquisition means" refers to a means for acquiring video data in real time from a surveillance camera or a home camera and storing that data.
[0798] The "video analysis means" is a means for processing the acquired video data and performing object detection and motion analysis. Specifically, it analyzes the video data using a deep learning algorithm.
[0799] "Abnormal behavior detection means" is a means for determining and detecting specific abnormal behavior in real time based on the results of video analysis.
[0800] The "warning message generating means" is a means for generating an appropriate warning message based on the detected abnormal behavior.
[0801] "Means for generating warning messages using a generative AI model" refers to means for dynamically generating warning messages in response to abnormal behavior using a generative AI model (e.g., natural language processing technology).
[0802] "Means for processing video in real time" refers to means for storing acquired video data in a buffer in real time and analyzing the data sequentially.
[0803] The "means for reading out a warning message" is a means for converting the generated warning message into voice using a voice generation module (for example, TTS technology) and reading it out to the target person.
[0804] The "information sharing means" is a means for transmitting the generated warning message and abnormal behavior information to a manager (e.g., parent, police, store manager, etc.) in real time.
[0805] A "terminal" is a device that reads out warning messages aloud, such as a smart speaker in the home or a street announcement system.
[0806] This invention is a system that processes a series of steps including video acquisition, video analysis, abnormal behavior detection, warning message generation, warning message reading, and information sharing. The embodiment of the invention will be explained mainly by dividing it into a server, a terminal, and a user.
[0807] Video acquisition method
[0808] The server acquires video data in real time from security cameras and home cameras using HTTP or RTSP protocols. The acquired video data is stored in the server's buffer. Specifically, the server acquires video frames step by step using the FFmpeg library and stores them in a buffer in memory.
[0809] Video analysis methods
[0810] The video analysis module in the server processes the real-time video data stored in the buffer. This analysis uses deep learning algorithms using TensorFlow and PyTorch. Specifically, models such as YOLO (You Only Look Once) and OpenPose are used for object detection and motion analysis. The video analysis module analyzes the person's pose for each frame and identifies their joint points. It also uses an object detection model to recognize objects held by the person.
[0811] Abnormal behavior detection method
[0812] The server's anomalous behavior detection module identifies anomalous behavior in real time from the analyzed video data. This detection module references a rule-based model that defines anomalous behavior patterns. For example, it can detect a child trying to get into a washing machine or a purse snatching attempt. The server identifies frame sequences that meet the conditions for anomalous behavior and marks them as anomalous based on that.
[0813] Warning statement generation means
[0814] If the server detects abnormal behavior, it uses a generative AI model to generate an appropriate warning message. For example, GPT-3 is used as this generative AI model. Based on the abnormal behavior information, the server inputs a prompt message into the generative AI model to generate a warning message. An example of a specific prompt message is, "A child has been detected entering the washing machine. Please generate an appropriate warning message." The generated warning message would be, "Playing with the washing machine is dangerous. Please stop immediately."
[0815] Warning message reading method
[0816] The device (smart speaker or street announcement system) uses a voice generation module to convert the generated warning text into audio and read it to the target person. Specifically, the warning text is converted into audio using a TTS (Text-to-Speech) API such as Amazon Polly or Google Text-to-Speech. The device converts the warning text received from the server into audio and reads it aloud from the speaker.
[0817] Information sharing means
[0818] The server sends the generated warning message and abnormal behavior information to an administrator (such as a parent, police, or store manager) in real time. This communication uses an HTTP-based notification API. The user (administrator) receives the warning information on their smartphone or a dedicated app, and can immediately check the content and take action. Specifically, the user receives the warning information using the real-time notification function within the app and can check the details.
[0819] Specific examples
[0820] Watching over children at home
[0821] The server receives real-time video footage from the home camera and stores the data in the server's buffer. The server's video analysis module analyzes the child's behavior when they open the washing machine door, and the abnormal behavior detection module detects this behavior as abnormal. The server uses a generative AI model to generate a warning message saying, "Playing with the washing machine is dangerous. Please stop immediately," and sends this message to the smart speaker. The device (smart speaker) then reads this warning message aloud. At the same time, the server sends the warning information to the parent's smartphone, where the parent can immediately check it.
[0822] Street crime prevention surveillance
[0823] The server acquires real-time video footage from street surveillance cameras and stores the video data in the server's buffer. The server's video analysis module detects snatching attempts, and the abnormal behavior detection module determines the behavior as abnormal. The server uses a generative AI model to generate a warning message stating, "A snatching attempt has been detected here. Please leave immediately," and sends it to the street announcement system. The device (street announcement system) reads this warning message aloud. At the same time, the server sends the warning information to the police or monitoring center, enabling immediate action.
[0824] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0825] Step 1:
[0826] Real-time video capture
[0827] The server obtains video data from a surveillance camera in real time. The input is the video stream from the surveillance camera. The server receives the video stream using HTTP or RTSP protocol, extracts this data frame by frame using the FFmpeg library, and stores it in a buffer in memory. The output is the video data in the buffer. Specifically, the server reads the IP address and connection settings of the surveillance camera, and obtains the data in real time after establishing a connection.
[0828] Step 2:
[0829] Video data analysis
[0830] The video analysis module in the server processes real-time video data stored in the buffer. The input is the video data in the buffer. The server analyzes the video using deep learning algorithms (such as YOLO or OpenPose) using TensorFlow or PyTorch. GPU acceleration is used to perform object detection and motion analysis at high speed. The output is the analysis results, which are the position information of specific objects or people in each frame. Specifically, the server inputs each frame of data into the deep learning model, obtains and stores the analysis results.
[0831] Step 3:
[0832] Abnormal behavior detection
[0833] The server's abnormal behavior detection module identifies abnormal behavior in real time based on analyzed video data. The input is the output of the video analysis module. The server references a rule-based model that defines abnormal behavior patterns in advance and determines whether it matches each frame of data. The output is information about the frame in which abnormal behavior was detected. Specifically, the server checks the conditions under which abnormal behavior occurs (e.g., a specific behavior occurs consecutively within a specific period of time) and determines that an abnormality exists based on that.
[0834] Step 4:
[0835] Generate a warning message
[0836] When the server detects abnormal behavior, it uses a generative AI model (for example, GPT-3) to generate a warning message. The input is information about the abnormal behavior. Based on this information, the server inputs a prompt message into the generative AI model. The output is the generated warning message. An example of a specific prompt message is, "A snatching attempt has been detected. Please generate an appropriate warning message." The generative model generates a warning message based on this message, and outputs the message, "A snatching attempt has been detected here. Please leave immediately."
[0837] Step 5:
[0838] Reading out warning messages
[0839] The device (such as a home smart speaker or a street announcement system) converts the generated warning text into audio using a voice generation module (for example, Amazon Polly or Google Text-to-Speech) and reads it aloud to the target person. The input is the warning text sent from the server. The device calls the TTS API, generates the warning text as an audio file, and outputs the sound from the speaker. The output is the audio of the warning text. Specifically, the device requests the received warning text from the TTS API and plays the generated audio data.
[0840] Step 6:
[0841] Information sharing
[0842] The server sends the generated warning message and abnormal behavior information to an administrator (parent, police, store manager, etc.) in real time. The input is the abnormal behavior information and the generated warning message. The server uses an HTTP-based notification API to send the warning information to the administrator's smartphone or dedicated app. The output is a notification message to the administrator. The user (administrator) receives the warning information on their smartphone or dedicated app, and can immediately check the content and take action. Specifically, the server reads the administrator's connection information and sends the warning data through the notification API.
[0843] (Application example 1)
[0844] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0845] Conventional crime prevention and surveillance systems have difficulty quickly detecting abnormal behavior and taking appropriate countermeasures immediately. Furthermore, delays in transmitting warning information can lead to problems with preventing crimes and accidents. To address these issues, there is a need for a system that can detect abnormal behavior in real time, issue warnings quickly, and share information with relevant parties immediately.
[0846] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0847] In this invention, the server includes a video acquisition means, a video analysis means, an abnormal behavior detection means, a warning message generation means, a warning message reading means, an information sharing means, a means for displaying a warning message on a display of the smart glasses, a means for issuing a voice warning, and a means for sharing abnormal behavior information with a management center, thereby enabling real-time abnormal behavior detection, prompt issuance of a warning, and prompt information sharing.
[0848] "Video acquisition means" refers to means for acquiring video data from a camera or other video capture device in real time.
[0849] The "video analysis means" is a means for processing the acquired video data and performing object detection and motion analysis.
[0850] "Abnormal behavior detection means" is a means of determining and detecting specific abnormal behavior in real time based on the results of video analysis.
[0851] The "warning message generation means" is a means for generating an appropriate warning message using a generative AI model when abnormal behavior is detected.
[0852] The "means for reading out a warning message" is a means for converting the generated warning message into voice using a voice generation module and reading it out loud to the target person.
[0853] The "information sharing means" is a means for sending the generated warning message and abnormal behavior information to the administrator in real time, enabling a prompt response.
[0854] The "means for displaying a warning message on the display of the smart glasses" refers to a means for visually displaying a warning message on the display of the smart glasses.
[0855] The "means for issuing an audio warning" is a means for issuing an audio warning when abnormal behavior is detected.
[0856] The "means for sharing abnormal behavior information with the management center" refers to a means for transmitting information on detected abnormal behavior to the management center in real time and quickly sharing it with relevant parties.
[0857] The present invention is a security system that processes the acquisition and analysis of camera footage, detection of abnormal behavior, generation and reading of warning messages, and information sharing in a series of steps. This system is mainly composed of a video acquisition means, a video analysis means, an abnormal behavior detection means, a warning message generation means, a warning message reading means, an information sharing means, a means for displaying a warning message on a smart glasses display, a means for issuing an audio warning, and a means for sharing abnormal behavior information with a management center.
[0858] Specifically, various hardware and software are used as follows:
[0859] Hardware:
[0860] Smart glasses: Devices that have a camera, display, and speaker (e.g., Google Glass, Vuzix Blade).
[0861] Camera: Surveillance cameras used on the street or in homes (e.g., network cameras).
[0862] Server: A computer system that processes and analyzes video data, detects abnormal behavior, and generates and distributes warning messages.
[0863] software:
[0864] Video Analytics Module: Object detection and motion analysis software using deep learning algorithms.
[0865] Abnormal behavior detection module: Software that identifies abnormal behavior in real time based on video analysis results.
[0866] Generative AI model: Natural language generation software that generates appropriate warning statements based on abnormal behavior.
[0867] Speech Generation Module: Text-to-Speech (TTS) software for vocalizing the generated warning text.
[0868] Information sharing module: Software that sends warning messages and abnormal behavior information to the management center and relevant parties in real time.
[0869] Data processing and calculation:
[0870] The server processes the acquired video data in real time. The video analysis module analyzes the video data sent from the camera and detects abnormal behavior, using existing deep learning methods. When abnormal behavior is detected, the generative AI model generates an appropriate warning message, which is then converted into audio by the voice generation module and sent to the smart glasses and other devices. The warning message is immediately displayed on the smart glasses' display, and a voice warning is also issued at the same time. In addition, abnormal behavior information is sent to the management center via the information sharing module, and relevant parties are immediately notified.
[0871] Examples:
[0872] For example, consider a police officer wearing smart glasses while patrolling the streets. During the patrol, the camera in the glasses captures real-time footage of the street. The footage is sent to a server, where the video analysis module detects violent acts. The generative AI model then generates a warning message, "Violent acts detected, please leave the scene," which is displayed and read aloud on the smart glasses. At the same time, this warning information is sent to a control center, allowing assistance to be quickly dispatched.
[0873] Example prompt sentence:
[0874] For example, by inputting the following prompt sentence into the generative AI model, an appropriate warning sentence will be generated.
[0875] Generate appropriate warning messages when the following actions are detected:
[0876] Street violence: Warning: "Violence detected, please respond immediately."
[0877] Snatching: Warning: "Snatching detected, please leave the scene."
[0878] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0879] Step 1:
[0880] The server acquires video data from the camera in real time. The input is the video stream from the camera, and the output video data is stored in the server's buffer.
[0881] Step 2:
[0882] The server's video analysis means processes the acquired video data. The input is the video data stored in the buffer, and object detection and motion analysis are performed, and the analysis results are output.
[0883] Step 3:
[0884] The server's abnormal behavior detection means detects specific abnormal behavior in real time based on the video analysis results. The input is the video analysis results, and the output is a determination of whether abnormal behavior has been detected.
[0885] Step 4:
[0886] The server generates a warning message using a generative AI model. The input is information about the detected abnormal behavior, and the generated warning message is output. In this case, the server uses a prompt message to generate an appropriate warning message.
[0887] Step 5:
[0888] The server sends the generated warning text to the terminal. The input is the generated warning text, which is sent to smart glasses or other terminals.
[0889] Step 6:
[0890] The terminal (smart glasses) displays a warning message on the display and issues a voice warning. The input is the warning message received from the server, and the output is a displayed warning message and a voice alert.
[0891] Step 7:
[0892] The server's information sharing means sends abnormal behavior information and warning messages to the management center. The input is the warning message and abnormal behavior information, and the output is shared in real time with the management center and relevant parties.
[0893] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0894] The present invention is a system that combines the functions of video acquisition, video analysis, abnormal behavior detection, warning message generation, warning message reading, information sharing, and emotion recognition, and links these functions together to provide appropriate warnings and share information while taking into account the user's emotions. An embodiment of this system is configured as follows.
[0895] System Configuration
[0896] 1. Video acquisition method
[0897] The server acquires video data from surveillance cameras and home cameras in real time.
[0898] The video is stored in a buffer on the server and is then analyzed.
[0899] 2. Video analysis methods
[0900] The video analysis module in the server processes the acquired video data and performs object detection and motion analysis.
[0901] The server uses deep learning algorithms to analyze patterns of abnormal behavior.
[0902] 3. Abnormal Behavior Detection Methods
[0903] The server's abnormal behavior detection module determines abnormal behavior based on the results of video analysis.
[0904] For example, it can detect a child trying to get into a washing machine or a snatch-and-run theft on the street.
[0905] 4. Emotion recognition means
[0906] The server analyzes the user's facial expressions and tone of voice and uses an emotion recognition module to recognize the user's emotions in real time.
[0907] The recognized emotional information is reflected in the generation of warning messages.
[0908] 5. Warning statement generation means
[0909] The server detects abnormal behavior and generates an appropriate warning message using a generative AI model based on the emotion recognition results.
[0910] The generated warning message will be based on the user's emotions. For example, a gentle warning message will be generated to calm the user.
[0911] 6. Warning message reading method
[0912] The device (home smart speaker, street announcement system) uses a voice generation module to convert the warning text into audio and read it out loud.
[0913] For example, a voice may be heard saying to a child, "Playing with the washing machine is dangerous. Stop immediately."
[0914] 7. Information sharing methods
[0915] The server sends the generated warning message and abnormal behavior information to the administrator in real time.
[0916] Users (parents or administrators) can receive warning information on their smartphones or dedicated apps and respond quickly.
[0917] Specific examples
[0918] Watching over children at home
[0919] The server acquires real-time video from cameras in the home and transmits the video data to the video analysis module.
[0920] The server's video analysis module analyzes the child's behavior of opening the washing machine door, and the abnormal behavior detection module detects this as abnormal behavior.
[0921] The server uses an emotion recognition module to analyze the child's facial expressions and tone of voice to obtain emotional information.
[0922] The server generates a warning message saying, "Playing with the washing machine is dangerous. Stop immediately," and adjusts the tone based on the emotion recognition results.
[0923] The device (smart speaker) reads out the warning message aloud, and at the same time, the server sends the warning information to the parent's smartphone, allowing the parent to check the information immediately.
[0924] Street crime prevention surveillance
[0925] The server acquires real-time video from surveillance cameras installed on the street and transmits the video data to the video analysis module.
[0926] The server's video analysis module detects snatching attempts, and the abnormal behavior detection module determines this as abnormal behavior.
[0927] The server uses an emotion recognition module to analyze the facial expressions and tone of voice of people around it to obtain emotional information.
[0928] The server generates a warning message saying, "Purse snatching has been detected here. Please leave immediately," and adjusts the tone based on emotional information.
[0929] The terminal (street announcement system) reads out the warning message, and at the same time, the server transmits the warning information to the police or monitoring center, enabling immediate action.
[0930] Program processing explanation
[0931] The server acquires video data from the surveillance camera and stores it in a buffer.
[0932] A video analysis module in the server analyzes the video data and identifies abnormal behavior.
[0933] When abnormal behavior is detected, the server activates the emotion recognition module and analyzes the user's emotional information.
[0934] Based on the emotional information, the server uses a generative AI model to generate an appropriate warning message.
[0935] The generated warning message is sent to the specified device and read aloud.
[0936] The warning message and abnormal behavior information are sent to the administrator in real time by the server, and the administrator (user) can check the warning information on a smartphone or other device.
[0937] In this way, it is possible to provide a system that can quickly detect abnormal behavior, recognize emotions, issue appropriate warnings, and share information.
[0938] The processing flow will be explained below.
[0939] Step 1:
[0940] The server acquires video data in real time from home cameras and street surveillance cameras. This video data is broken down into multiple frames and stored in the server's temporary memory (buffer).
[0941] Step 2:
[0942] The server sends the acquired video data to the video analysis module, which converts the video data into a format that can be processed in real time and then provides it for analysis.
[0943] Step 3:
[0944] The video analysis module in the server processes the video data and applies object detection and motion analysis algorithms to analyze the movement, shape, and position of people and objects, preparing to identify abnormal behavior.
[0945] Step 4:
[0946] The server's abnormal behavior detection module determines abnormal behavior based on the video analysis results. For example, it detects a child trying to get into a washing machine or a person moving abnormally in a specific area.
[0947] Step 5:
[0948] After the server detects abnormal behavior, it activates the emotion recognition module, which analyzes the user's facial expressions and tone of voice from the video data to recognize their emotions.
[0949] Step 6:
[0950] The server uses a generative AI model to generate appropriate warning messages based on the emotion recognition results. During this generation process, the tone and content of the warning messages are adjusted to match the user's emotions.
[0951] Step 7:
[0952] The server generates a warning message and sends it to a designated device, such as a smart speaker in the home or an announcement system on the street.
[0953] Step 8:
[0954] The warning message received by the device is converted into voice by a voice generation module and actually read aloud from a speaker. For example, at home, a warning such as "Playing with the washing machine is dangerous. Stop immediately" is played in a soft tone.
[0955] Step 9:
[0956] At the same time, the server sends the generated warning message and detailed information about the abnormal behavior to the administrator in real time. The administrator (user) receives the warning information on their smartphone or a dedicated application, enabling them to respond quickly. For example, a parent could receive a notification on their smartphone and check on their child's safety.
[0957] In this way, the processing flow from step 1 to step 9 provides a system that quickly detects abnormal behavior and provides appropriate warnings and information sharing that take the user's emotions into consideration.
[0958] Example 2
[0959] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0960] Conventional surveillance systems were able to detect abnormal behavior by acquiring and analyzing video data, but they lacked the ability to issue prompt and appropriate warnings and share information while taking into account the user's emotions. As a result, when abnormal behavior occurred, responses were sometimes delayed, making it difficult to implement effective preventative measures.
[0961] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a video acquisition means, a video analysis means, an abnormal behavior detection means, an emotion recognition means, a warning message generation means, a warning message reading means, and an information sharing means. This makes it possible to issue a prompt and accurate warning that takes into account the user's emotions and to share abnormal behavior information in real time.
[0962] The "image acquisition means" is a means for acquiring image data from a camera device in real time.
[0963] The "video analysis means" is a means for processing the acquired video data and performing object detection and motion analysis.
[0964] The "abnormal behavior detection means" is a means for detecting specific abnormal behavior based on the analysis results of the video analysis means.
[0965] The "emotion recognition means" is a means for analyzing the user's facial expressions and voice and recognizing their emotions.
[0966] The "warning message generation means" is a means for generating a warning message according to the user's emotions using a generative AI model.
[0967] The "means for reading out a warning message" is a means for reading out the generated warning message by voice.
[0968] The "information sharing means" is a means for transmitting the generated warning message and abnormal behavior information to the administrator in real time.
[0969] This system is implemented by combining the following elements: video acquisition means, video analysis means, abnormal behavior detection means, emotion recognition means, warning message generation means, warning message reading means, and information sharing means.
[0970] System Configuration
[0971] 1. Video acquisition method:
[0972] The server acquires video data from the camera device in real time. For example, a home IP camera is connected to the server and transmits video using the RTSP protocol. This video data is stored in a buffer on the server.
[0973] 2. Video analysis methods:
[0974] The video analysis module in the server processes the acquired video data and performs object detection and motion analysis. Specifically, it uses deep learning algorithms such as YOLO and SSD to detect objects in the video frame in real time.
[0975] 3. Abnormal behavior detection methods:
[0976] The server's abnormal behavior detection module detects specific abnormal behavior based on the video analysis results and determines whether the behavior is abnormal by comparing the detected behavior pattern with a predefined abnormal behavior database.
[0977] 4. Emotion recognition means:
[0978] The server analyzes the user's facial expressions and voice and recognizes the user's emotions in real time using an emotion recognition module. For example, OpenFace can be used for facial expression recognition and OpenSMILE for voice analysis.
[0979] 5. Warning statement generation means:
[0980] The server detects abnormal behavior and generates an appropriate warning message using a generative AI model based on the emotion recognition results. Specifically, a prompt message is input into a generative model such as GPT-4 to generate the warning message.
[0981] Example prompt sentence:
[0982] "Playing with the washing machine is dangerous. Stop immediately."
[0983] 6. Warning message reading method:
[0984] The device (such as a smart speaker in the home or a public address system) uses a speech generation module to convert the warning text into a voice and read it aloud. A text-to-speech (TTS) engine is used to generate an audio file, which is then played on the device.
[0985] 7. Information sharing methods:
[0986] The server generates a warning message and sends the abnormal behavior information to the administrator in real time. The administrator (e.g., parent or police) receives the warning information on their smartphone or a dedicated app and can respond promptly.
[0987] Specific examples
[0988] Watching over children at home
[0989] The server acquires real-time video from cameras in the home and transmits the video data to the video analysis module.
[0990] The server's video analysis module analyzes the child's behavior of opening the washing machine door, and the abnormal behavior detection module detects this as abnormal behavior.
[0991] The server uses an emotion recognition module to analyze the child's facial expressions and tone of voice to obtain emotional information.
[0992] The server generates a warning message saying, "Playing with the washing machine is dangerous. Stop immediately," and adjusts the tone based on the emotion recognition results.
[0993] The device (smart speaker) reads out the warning message aloud, and at the same time, the server sends the warning information to the parent's smartphone, allowing the parent to check the information immediately.
[0994] Street crime prevention surveillance
[0995] The server acquires real-time video from cameras installed on the street and sends the video data to the video analysis module.
[0996] The server's video analysis module detects snatching attempts, and the abnormal behavior detection module determines this as abnormal behavior.
[0997] The server uses an emotion recognition module to analyze the facial expressions and tone of voice of people around it to obtain emotional information.
[0998] The server generates a warning message saying, "Purse snatching has been detected here. Please leave immediately," and adjusts the tone based on emotional information.
[0999] The terminal (street announcement system) reads out the warning message, and at the same time, the server transmits the warning information to the police or monitoring center, enabling immediate action.
[1000] In this way, a system can be built that can quickly detect abnormal behavior in real time, issue warnings, and share information.
[1001] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1002] Step 1:
[1003] The server receives video data from the camera device in real time. The input is the video stream from the camera, and the output is the video data stored in the server's buffer. Specifically, the server communicates with the IP camera using the RTSP protocol and stores the received data in a buffer.
[1004] Step 2:
[1005] The video analysis module in the server retrieves video data from the buffer and performs object detection and motion analysis. The input is the video data stored in the buffer, and the output is the analyzed object and motion information. Specifically, it uses a deep learning algorithm (such as YOLO or SSD) to detect objects in the video frame in real time.
[1006] Step 3:
[1007] The server's abnormal behavior detection module detects specific abnormal behavior based on the video analysis results. The input is the output of the video analysis module, and the output is the abnormal behavior detection results. Specifically, the analysis results are compared with a predefined abnormal behavior database to determine abnormal behavior. For example, a child trying to open the washing machine door is detected as abnormal behavior.
[1008] Step 4:
[1009] The server uses an emotion recognition module to analyze the user's facial expressions and tone of voice to recognize emotions. The input is the user's face and voice data obtained through video analysis, and the output is the recognized emotional information. Specific operations include running a facial expression recognition model (e.g., OpenFace) and a voice analysis model (e.g., OpenSMILE). It may recognize emotions such as fear or surprise from a child's face.
[1010] Step 5:
[1011] The server detects abnormal behavior and generates a warning message using a generative AI model based on the emotion recognition results. The input is the abnormal behavior detection results and emotion recognition results, and the output is the generated warning message. Specifically, a prompt message is input into a generative AI model such as GPT-4 to generate a warning message. For example, a warning message such as "Playing with the washing machine is dangerous. Stop immediately" is generated.
[1012] Step 6:
[1013] The device (such as a home smart speaker or a street announcement system) uses a voice generation module to read out the warning text. The input is the generated warning text, and the output is a voiced version of the warning text. Specifically, the generated warning text is input into a text-to-speech (TTS) engine, which generates an audio file and plays it back.
[1014] Step 7:
[1015] The server generates a warning message and sends the abnormal behavior information to the administrator in real time. The input is the warning message and abnormal behavior information, and the output is the warning information sent to the administrator's device. Specifically, the abnormal behavior and warning message are sent via push notification to the administrator's smartphone or dedicated app, allowing parents, police, etc. to respond quickly.
[1016] (Application example 2)
[1017] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1018] While conventional video surveillance systems can detect abnormal behavior, they are unable to generate appropriate warnings based on the user's emotions and share information quickly. As a result, appropriate responses to abnormal behavior are delayed, making it difficult to provide effective security and monitoring. To solve this problem, the present invention aims to provide a system that integrates abnormal behavior detection and user emotion recognition to generate appropriate warnings and share information.
[1019] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1020] In this invention, the server includes a video acquisition means, a video analysis means, an abnormal behavior detection means, an emotion recognition means, a warning message generation means, a warning message reading means, and an information sharing means. This allows the server to acquire video from a surveillance camera in real time and analyze the video data to detect abnormal behavior. It can also recognize emotions in real time from a user's facial expression and tone of voice and generate an appropriate warning message based on the emotion information. The generated warning message is read aloud, and the warning information and abnormal behavior information are further transmitted to an administrator in real time, enabling a prompt and effective response.
[1021] "Video acquisition means" refers to a device or function that acquires video data in real time from a surveillance camera, a home camera, or the like.
[1022] "Video analysis means" refers to modules and algorithms for processing acquired video data and performing object detection and motion analysis.
[1023] "Abnormal behavior detection means" refers to a module or device for detecting specific abnormal behavior based on the results of video analysis.
[1024] "Emotion recognition means" refers to modules or algorithms that analyze the user's facial expressions and tone of voice and recognize emotions in real time.
[1025] A "warning message generation means" is a device or software that uses a generative AI model to generate an appropriate warning message based on the results of emotion recognition.
[1026] The "means for reading out a warning message" is a device or function for reading out the generated warning message aloud.
[1027] The "information sharing means" is a device or system for transmitting the generated warning message and abnormal behavior information to the administrator in real time.
[1028] A "generative AI model" is an artificial intelligence model for generating natural language based on specific input data.
[1029] A "prompt sentence" is an input sentence that causes a generative AI model to generate a specific output.
[1030] The present invention relates to a system that integrates video acquisition, video analysis, abnormal behavior detection, emotion recognition, warning message generation, warning message reading, and information sharing. This system can analyze video data acquired from surveillance cameras and home cameras in real time to detect abnormal behavior. Furthermore, it has the ability to recognize user emotions, generate appropriate warning messages based on those emotions, and read them aloud. Furthermore, the generated warning messages and abnormal behavior information are sent to administrators in real time, enabling prompt and effective response.
[1031] System configuration:
[1032] 1. Video acquisition method
[1033] The server receives video data from surveillance cameras and home cameras in real time, stores the video in a buffer on the server, and then analyzes it.
[1034] 2. Video analysis methods
[1035] The server's video analysis module processes the captured video data and performs object detection and motion analysis. The server uses deep learning algorithms to analyze patterns of abnormal behavior.
[1036] 3. Abnormal Behavior Detection Methods
[1037] The server's abnormal behavior detection module determines abnormal behavior based on the video analysis results, such as falls, trespassing in specific areas, and dangerous behavior.
[1038] 4. Emotion recognition means
[1039] The server analyzes the user's facial expressions and tone of voice, and uses an emotion recognition module to recognize the user's emotions in real time. The recognized emotional information is reflected in the generation of warning messages.
[1040] 5. Warning statement generation means
[1041] The server detects abnormal behavior and generates an appropriate warning message based on the emotion recognition results using a generative AI model. The generated warning message is tailored to the user's emotions.
[1042] 6. Warning message reading method
[1043] The device (such as a smart speaker in the home or a street announcement system) uses a voice generation module to convert the warning message into audio and read it aloud. For example, a voice might say, "Falling is dangerous. Please sit down and rest."
[1044] 7. Information sharing methods
[1045] The server generates a warning message and sends abnormal behavior information to the administrator in real time. The user (administrator) receives the warning information on their smartphone or a dedicated app, allowing them to respond quickly.
[1046] Examples:
[1047] Elderly monitoring system:
[1048] When an elderly person falls, abnormal behavior is detected from camera footage and, based on emotion recognition, a gentle warning message is generated, such as "Falling is dangerous. Please sit down and rest." The warning message is read aloud by the smart speaker and this information is sent to family members' smartphones in real time, encouraging them to take prompt action.
[1049] Examples of prompts:
[1050] "If the user is sad and falls, create a warning message in a gentle tone encouraging them to rest."
[1051] "If the user has a neutral emotion, generate a warning message to prevent them from falling."
[1052] Hardware and software used:
[1053] The hardware used is a surveillance camera, a home camera, a cloud server, a smart speaker, and a smartphone, while the software used is a video analysis module, an emotion recognition module, a generative AI model, a warning message generation module, and a notification sending module.
[1054] This system efficiently carries out a series of processes, from real-time acquisition and analysis of video data, detection of abnormal behavior, emotion recognition, generation and voice reading of appropriate warning messages, and information sharing, thereby achieving effective security and monitoring.
[1055] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1056] Step 1:
[1057] The server acquires video data in real time from surveillance cameras and home cameras. The video is stored in the server's buffer. The input is real-time video data from the camera, and the output is video data stored in the server's buffer. Specifically, the server acquires video data frame by frame from the camera stream and stores it in memory.
[1058] Step 2:
[1059] The server's video analysis module processes the acquired video data and performs object detection and motion analysis. The input is the video data stored in the buffer, and the output is the analysis results (e.g., object position and motion information). Specifically, the server uses a deep learning algorithm to detect face and body movements from the video data.
[1060] Step 3:
[1061] The server's abnormal behavior detection module detects specific abnormal behavior based on the video analysis results. The input is the result data of object detection and motion analysis, and the output is the abnormal behavior detection results. Specifically, the server compares motion patterns and determines abnormal behavior such as falls or dangerous behavior.
[1062] Step 4:
[1063] The server's emotion recognition module analyzes the user's facial expressions and tone of voice to recognize emotions in real time. The input is facial and voice data, and the output is recognized emotional information. Specifically, deep learning is used to estimate emotions such as joy, anger, sadness, and happiness from the user's facial and voice data.
[1064] Step 5:
[1065] The server detects abnormal behavior and generates an appropriate warning message using a generative AI model based on the emotion recognition results. The inputs are the abnormal behavior detection results and emotion recognition results, and the output is the generated warning message. Specifically, the server inputs a prompt message into the generative AI model and generates an appropriate warning message according to the emotion. For example, the prompt message for the generative AI model could be, "If the user is sad and falls, please create a warning message in a gentle tone encouraging them to rest."
[1066] Step 6:
[1067] The device (such as a home smart speaker or a street announcement system) uses a voice generation module to convert the warning text into audio and read it out loud. The input is the generated warning text, and the output is a voice warning. Specifically, the device uses text-to-speech conversion technology to play back the generated warning text aloud.
[1068] Step 7:
[1069] The server sends the generated warning message and abnormal behavior information to the administrator in real time. The input is the warning message and abnormal behavior information, and the output is a notification to the administrator's smartphone or a dedicated app. Specifically, the server uses a notification sending module to send the warning information as an email or push notification.
[1070] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1071] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1072] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1073] [Fourth embodiment]
[1074] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1075] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1076] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1077] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1078] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1079] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1080] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1081] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1082] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1083] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1084] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1085] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1086] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1087] The present invention is a system that processes a series of steps including video acquisition, video analysis, abnormal behavior detection, warning message generation, warning message reading, and information sharing. An embodiment of this system is mainly composed of the following elements.
[1088] System Configuration
[1089] 1. Video acquisition method
[1090] The server acquires video data from surveillance cameras and home cameras in real time.
[1091] The camera has a real-time streaming function, and the captured images are stored in a buffer on the server.
[1092] 2. Video analysis methods
[1093] The video analysis module in the server processes the acquired video data and performs object detection and motion analysis.
[1094] The server uses existing deep learning algorithms to analyze patterns of abnormal behavior.
[1095] 3. Abnormal Behavior Detection Methods
[1096] The server's abnormal behavior detection module determines and detects specific abnormal behavior in real time based on the video analysis results.
[1097] For example, it identifies a child trying to get into a washing machine or a snatching incident on the street.
[1098] 4. Warning statement generation means
[1099] If the server detects abnormal behavior, it uses a generative AI model to generate an appropriate warning message.
[1100] The generated warning message will vary depending on the abnormal behavior detected.
[1101] 5. Warning message reading method
[1102] The device (home smart speaker, street announcement system) converts the generated warning text into audio using a voice generation module and reads it out to the target person.
[1103] For example, at home, a warning will be played saying, "Playing with the washing machine is dangerous. Stop immediately."
[1104] 6. Information sharing methods
[1105] The server generates a warning message and sends abnormal behavior information to an administrator (parent, police, store manager, etc.) in real time.
[1106] Users (parents or administrators) can receive warning information on their smartphones or dedicated apps and respond quickly.
[1107] Specific examples
[1108] Watching over children at home
[1109] The server acquires real-time video from cameras in the home and transmits the video data to the video analysis module.
[1110] The server's video analysis module analyzes the child's behavior of opening the washing machine door, and the abnormal behavior detection module detects this as abnormal behavior.
[1111] The server generates a warning message saying, "Playing with the washing machine is dangerous. Stop immediately," and sends it to a smart speaker in the home.
[1112] The device (smart speaker) reads out the warning message aloud, and at the same time, the server sends the warning information to the parent's smartphone, allowing the parent to check the information immediately.
[1113] Street crime prevention surveillance
[1114] The server acquires real-time video from surveillance cameras installed on the street and transmits the video data to the video analysis module.
[1115] The server's video analysis module detects snatching attempts, and the abnormal behavior detection module determines this as abnormal behavior.
[1116] The server generates a warning message saying, "Purse snatching has been detected here. Please leave immediately," and sends it to the street announcement system.
[1117] The terminal (street announcement system) reads out the warning message, and at the same time, the server transmits the warning information to the police or monitoring center, enabling immediate action.
[1118] Program processing explanation
[1119] The server acquires video data from the surveillance camera and stores it in a buffer.
[1120] A video analysis module in the server analyzes the video data stored in the buffer and identifies abnormal behavior.
[1121] Based on the analysis results, the server's abnormal behavior detection module detects abnormal behavior.
[1122] If abnormal behavior is detected, the server uses a generative AI model to generate an appropriate warning message.
[1123] The generated warning message is sent to a designated device (smart speaker, announcement system) and read aloud.
[1124] The warning message and abnormal behavior information are sent to the administrator in real time by the server, and the administrator (user) can check the warning information on a smartphone or other device.
[1125] In this way, the form for implementing the invention provides a system in which each step operates in conjunction with other steps to quickly detect abnormal behavior and issue appropriate warnings, thereby preventing crimes and accidents before they occur.
[1126] The processing flow will be explained below.
[1127] Step 1:
[1128] The server acquires video data from security cameras and home cameras in real time. This acquisition process involves capturing video at a specified frame rate and storing it in a buffer.
[1129] Step 2:
[1130] The server sends the acquired video data to the video analysis module, where the video data is converted into a format that allows real-time processing.
[1131] Step 3:
[1132] The server's video analytics module processes the video data and runs object detection and motion analysis algorithms, for example, to identify specific human movements or unusual object behavior.
[1133] Step 4:
[1134] The server's abnormal behavior detection module determines abnormal behavior based on the video analysis results. This module uses a trained deep learning model to detect, for example, a child trying to get into a washing machine as an abnormal behavior.
[1135] Step 5:
[1136] When the server detects abnormal behavior, it activates a generative AI model to generate an appropriate warning message. This generation process creates a contextual warning message based on the nature of the abnormal behavior. For example, a warning message might be generated that reads, "Playing with the washing machine is dangerous. Stop immediately."
[1137] Step 6:
[1138] The server then sends the generated warning message to a designated device, such as a smart speaker in the home or a street announcement system.
[1139] Step 7:
[1140] The device converts the received warning message into audio using a speech generation module and reads it out loud from the speaker. For example, a voice warning a child playing with a washing machine might be played saying, "Playing with a washing machine is dangerous. Stop immediately."
[1141] Step 8:
[1142] At the same time, the server sends the generated warning message and detailed information about abnormal behavior to the administrator in real time. The administrator (user) can receive the warning information on their smartphone or a dedicated application and respond quickly. For example, parents can receive a notification on their smartphone and immediately check on their child's safety.
[1143] In this way, the processing flow from step 1 to step 8 realizes a system in which abnormal behavior is quickly detected and appropriate warnings and information sharing are provided.
[1144] Example 1
[1145] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1146] In recent years, safety monitoring in homes and on the streets has become increasingly important, but conventional monitoring systems have had difficulty detecting abnormal behavior in real time and issuing appropriate warnings. Furthermore, they have been unable to quickly share information with administrators, making it difficult to take immediate action to prevent incidents and accidents. New technological solutions are needed to solve these issues.
[1147] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1148] In this invention, the server includes a video acquisition means, a video analysis means, an abnormal behavior detection means, a warning message generation means, a warning message generation means using a generative AI model, a real-time video processing means, and an information sharing means with enhanced functions. This enables real-time analysis of video data, immediate detection of abnormal behavior, dynamic generation of warning messages using the generative AI model, and rapid voice generation of the warning messages and immediate notification to an administrator.
[1149] "Video acquisition means" refers to a means for acquiring video data in real time from a surveillance camera or a home camera and storing that data.
[1150] The "video analysis means" is a means for processing the acquired video data and performing object detection and motion analysis. Specifically, it analyzes the video data using a deep learning algorithm.
[1151] "Abnormal behavior detection means" is a means for determining and detecting specific abnormal behavior in real time based on the results of video analysis.
[1152] The "warning message generating means" is a means for generating an appropriate warning message based on the detected abnormal behavior.
[1153] "Means for generating warning messages using a generative AI model" refers to means for dynamically generating warning messages in response to abnormal behavior using a generative AI model (e.g., natural language processing technology).
[1154] "Means for processing video in real time" refers to means for storing acquired video data in a buffer in real time and analyzing the data sequentially.
[1155] The "means for reading out a warning message" is a means for converting the generated warning message into voice using a voice generation module (for example, TTS technology) and reading it out to the target person.
[1156] The "information sharing means" is a means for transmitting the generated warning message and abnormal behavior information to a manager (e.g., parent, police, store manager, etc.) in real time.
[1157] A "terminal" is a device that reads out warning messages aloud, such as a smart speaker in the home or a street announcement system.
[1158] This invention is a system that processes a series of steps including video acquisition, video analysis, abnormal behavior detection, warning message generation, warning message reading, and information sharing. The embodiment of the invention will be explained mainly by dividing it into a server, a terminal, and a user.
[1159] Video acquisition method
[1160] The server acquires video data in real time from security cameras and home cameras using HTTP or RTSP protocols. The acquired video data is stored in the server's buffer. Specifically, the server acquires video frames step by step using the FFmpeg library and stores them in a buffer in memory.
[1161] Video analysis methods
[1162] The video analysis module in the server processes the real-time video data stored in the buffer. This analysis uses deep learning algorithms using TensorFlow and PyTorch. Specifically, models such as YOLO (You Only Look Once) and OpenPose are used for object detection and motion analysis. The video analysis module analyzes the person's pose for each frame and identifies their joint points. It also uses an object detection model to recognize objects held by the person.
[1163] Abnormal behavior detection method
[1164] The server's anomalous behavior detection module identifies anomalous behavior in real time from the analyzed video data. This detection module references a rule-based model that defines anomalous behavior patterns. For example, it can detect a child trying to get into a washing machine or a purse snatching attempt. The server identifies frame sequences that meet the conditions for anomalous behavior and marks them as anomalous based on that.
[1165] Warning statement generation means
[1166] If the server detects abnormal behavior, it uses a generative AI model to generate an appropriate warning message. For example, GPT-3 is used as this generative AI model. Based on the abnormal behavior information, the server inputs a prompt message into the generative AI model to generate a warning message. An example of a specific prompt message is, "A child has been detected entering the washing machine. Please generate an appropriate warning message." The generated warning message would be, "Playing with the washing machine is dangerous. Please stop immediately."
[1167] Warning message reading method
[1168] The device (smart speaker or street announcement system) uses a voice generation module to convert the generated warning text into audio and read it to the target person. Specifically, the warning text is converted into audio using a TTS (Text-to-Speech) API such as Amazon Polly or Google Text-to-Speech. The device converts the warning text received from the server into audio and reads it aloud from the speaker.
[1169] Information sharing means
[1170] The server sends the generated warning message and abnormal behavior information to an administrator (such as a parent, police, or store manager) in real time. This communication uses an HTTP-based notification API. The user (administrator) receives the warning information on their smartphone or a dedicated app, and can immediately check the content and take action. Specifically, the user receives the warning information using the real-time notification function within the app and can check the details.
[1171] Specific examples
[1172] Watching over children at home
[1173] The server receives real-time video footage from the home camera and stores the data in the server's buffer. The server's video analysis module analyzes the child's behavior when they open the washing machine door, and the abnormal behavior detection module detects this behavior as abnormal. The server uses a generative AI model to generate a warning message saying, "Playing with the washing machine is dangerous. Please stop immediately," and sends this message to the smart speaker. The device (smart speaker) then reads this warning message aloud. At the same time, the server sends the warning information to the parent's smartphone, where the parent can immediately check it.
[1174] Street crime prevention surveillance
[1175] The server acquires real-time video footage from street surveillance cameras and stores the video data in the server's buffer. The server's video analysis module detects snatching attempts, and the abnormal behavior detection module determines the behavior as abnormal. The server uses a generative AI model to generate a warning message stating, "A snatching attempt has been detected here. Please leave immediately," and sends it to the street announcement system. The device (street announcement system) reads this warning message aloud. At the same time, the server sends the warning information to the police or monitoring center, enabling immediate action.
[1176] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1177] Step 1:
[1178] Real-time video capture
[1179] The server obtains video data from a surveillance camera in real time. The input is the video stream from the surveillance camera. The server receives the video stream using HTTP or RTSP protocol, extracts this data frame by frame using the FFmpeg library, and stores it in a buffer in memory. The output is the video data in the buffer. Specifically, the server reads the IP address and connection settings of the surveillance camera, and obtains the data in real time after establishing a connection.
[1180] Step 2:
[1181] Video data analysis
[1182] The video analysis module in the server processes real-time video data stored in the buffer. The input is the video data in the buffer. The server analyzes the video using deep learning algorithms (such as YOLO or OpenPose) using TensorFlow or PyTorch. GPU acceleration is used to perform object detection and motion analysis at high speed. The output is the analysis results, which are the position information of specific objects or people in each frame. Specifically, the server inputs each frame of data into the deep learning model, obtains and stores the analysis results.
[1183] Step 3:
[1184] Abnormal behavior detection
[1185] The server's abnormal behavior detection module identifies abnormal behavior in real time based on analyzed video data. The input is the output of the video analysis module. The server references a rule-based model that defines abnormal behavior patterns in advance and determines whether it matches each frame of data. The output is information about the frame in which abnormal behavior was detected. Specifically, the server checks the conditions under which abnormal behavior occurs (e.g., a specific behavior occurs consecutively within a specific period of time) and determines that an abnormality exists based on that.
[1186] Step 4:
[1187] Generate a warning message
[1188] When the server detects abnormal behavior, it uses a generative AI model (for example, GPT-3) to generate a warning message. The input is information about the abnormal behavior. Based on this information, the server inputs a prompt message into the generative AI model. The output is the generated warning message. An example of a specific prompt message is, "A snatching attempt has been detected. Please generate an appropriate warning message." The generative model generates a warning message based on this message, and outputs the message, "A snatching attempt has been detected here. Please leave immediately."
[1189] Step 5:
[1190] Reading out warning messages
[1191] The device (such as a home smart speaker or a street announcement system) converts the generated warning text into audio using a voice generation module (for example, Amazon Polly or Google Text-to-Speech) and reads it aloud to the target person. The input is the warning text sent from the server. The device calls the TTS API, generates the warning text as an audio file, and outputs the sound from the speaker. The output is the audio of the warning text. Specifically, the device requests the received warning text from the TTS API and plays the generated audio data.
[1192] Step 6:
[1193] Information sharing
[1194] The server sends the generated warning message and abnormal behavior information to an administrator (parent, police, store manager, etc.) in real time. The input is the abnormal behavior information and the generated warning message. The server uses an HTTP-based notification API to send the warning information to the administrator's smartphone or dedicated app. The output is a notification message to the administrator. The user (administrator) receives the warning information on their smartphone or dedicated app, and can immediately check the content and take action. Specifically, the server reads the administrator's connection information and sends the warning data through the notification API.
[1195] (Application example 1)
[1196] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1197] Conventional crime prevention and surveillance systems have difficulty quickly detecting abnormal behavior and taking appropriate countermeasures immediately. Furthermore, delays in transmitting warning information can lead to problems with preventing crimes and accidents. To address these issues, there is a need for a system that can detect abnormal behavior in real time, issue warnings quickly, and share information with relevant parties immediately.
[1198] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1199] In this invention, the server includes a video acquisition means, a video analysis means, an abnormal behavior detection means, a warning message generation means, a warning message reading means, an information sharing means, a means for displaying a warning message on a display of the smart glasses, a means for issuing a voice warning, and a means for sharing abnormal behavior information with a management center, thereby enabling real-time abnormal behavior detection, prompt issuance of a warning, and prompt information sharing.
[1200] "Video acquisition means" refers to means for acquiring video data from a camera or other video capture device in real time.
[1201] The "video analysis means" is a means for processing the acquired video data and performing object detection and motion analysis.
[1202] "Abnormal behavior detection means" is a means of determining and detecting specific abnormal behavior in real time based on the results of video analysis.
[1203] The "warning message generation means" is a means for generating an appropriate warning message using a generative AI model when abnormal behavior is detected.
[1204] The "means for reading out a warning message" is a means for converting the generated warning message into voice using a voice generation module and reading it out loud to the target person.
[1205] The "information sharing means" is a means for sending the generated warning message and abnormal behavior information to the administrator in real time, enabling a prompt response.
[1206] The "means for displaying a warning message on the display of the smart glasses" refers to a means for visually displaying a warning message on the display of the smart glasses.
[1207] The "means for issuing an audio warning" is a means for issuing an audio warning when abnormal behavior is detected.
[1208] The "means for sharing abnormal behavior information with the management center" refers to a means for transmitting information on detected abnormal behavior to the management center in real time and quickly sharing it with relevant parties.
[1209] The present invention is a security system that processes the acquisition and analysis of camera footage, detection of abnormal behavior, generation and reading of warning messages, and information sharing in a series of steps. This system is mainly composed of a video acquisition means, a video analysis means, an abnormal behavior detection means, a warning message generation means, a warning message reading means, an information sharing means, a means for displaying a warning message on a smart glasses display, a means for issuing an audio warning, and a means for sharing abnormal behavior information with a management center.
[1210] Specifically, various hardware and software are used as follows:
[1211] Hardware:
[1212] Smart glasses: Devices that have a camera, display, and speaker (e.g., Google Glass, Vuzix Blade).
[1213] Camera: Surveillance cameras used on the street or in homes (e.g., network cameras).
[1214] Server: A computer system that processes and analyzes video data, detects abnormal behavior, and generates and distributes warning messages.
[1215] software:
[1216] Video Analytics Module: Object detection and motion analysis software using deep learning algorithms.
[1217] Abnormal behavior detection module: Software that identifies abnormal behavior in real time based on video analysis results.
[1218] Generative AI model: Natural language generation software that generates appropriate warning statements based on abnormal behavior.
[1219] Speech Generation Module: Text-to-Speech (TTS) software for vocalizing the generated warning text.
[1220] Information sharing module: Software that sends warning messages and abnormal behavior information to the management center and relevant parties in real time.
[1221] Data processing and calculation:
[1222] The server processes the acquired video data in real time. The video analysis module analyzes the video data sent from the camera and detects abnormal behavior, using existing deep learning methods. When abnormal behavior is detected, the generative AI model generates an appropriate warning message, which is then converted into audio by the voice generation module and sent to the smart glasses and other devices. The warning message is immediately displayed on the smart glasses' display, and a voice warning is also issued at the same time. In addition, abnormal behavior information is sent to the management center via the information sharing module, and relevant parties are immediately notified.
[1223] Examples:
[1224] For example, consider a police officer wearing smart glasses while patrolling the streets. During the patrol, the camera in the glasses captures real-time footage of the street. The footage is sent to a server, where the video analysis module detects violent acts. The generative AI model then generates a warning message, "Violent acts detected, please leave the scene," which is displayed and read aloud on the smart glasses. At the same time, this warning information is sent to a control center, allowing assistance to be quickly dispatched.
[1225] Example prompt sentence:
[1226] For example, by inputting the following prompt sentence into the generative AI model, an appropriate warning sentence will be generated.
[1227] Generate appropriate warning messages when the following actions are detected:
[1228] Street violence: Warning: "Violence detected, please respond immediately."
[1229] Snatching: Warning: "Snatching detected, please leave the scene."
[1230] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1231] Step 1:
[1232] The server acquires video data from the camera in real time. The input is the video stream from the camera, and the output video data is stored in the server's buffer.
[1233] Step 2:
[1234] The server's video analysis means processes the acquired video data. The input is the video data stored in the buffer, and object detection and motion analysis are performed, and the analysis results are output.
[1235] Step 3:
[1236] The server's abnormal behavior detection means detects specific abnormal behavior in real time based on the video analysis results. The input is the video analysis results, and the output is a determination of whether abnormal behavior has been detected.
[1237] Step 4:
[1238] The server generates a warning message using a generative AI model. The input is information about the detected abnormal behavior, and the generated warning message is output. In this case, the server uses a prompt message to generate an appropriate warning message.
[1239] Step 5:
[1240] The server sends the generated warning text to the terminal. The input is the generated warning text, which is sent to smart glasses or other terminals.
[1241] Step 6:
[1242] The terminal (smart glasses) displays a warning message on the display and issues a voice warning. The input is the warning message received from the server, and the output is a displayed warning message and a voice alert.
[1243] Step 7:
[1244] The server's information sharing means sends abnormal behavior information and warning messages to the management center. The input is the warning message and abnormal behavior information, and the output is shared in real time with the management center and relevant parties.
[1245] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1246] The present invention is a system that combines the functions of video acquisition, video analysis, abnormal behavior detection, warning message generation, warning message reading, information sharing, and emotion recognition, and links these functions together to provide appropriate warnings and share information while taking into account the user's emotions. An embodiment of this system is configured as follows.
[1247] System Configuration
[1248] 1. Video acquisition method
[1249] The server acquires video data from surveillance cameras and home cameras in real time.
[1250] The video is stored in a buffer on the server and is then analyzed.
[1251] 2. Video analysis methods
[1252] The video analysis module in the server processes the acquired video data and performs object detection and motion analysis.
[1253] The server uses deep learning algorithms to analyze patterns of abnormal behavior.
[1254] 3. Abnormal Behavior Detection Methods
[1255] The server's abnormal behavior detection module determines abnormal behavior based on the results of video analysis.
[1256] For example, it can detect a child trying to get into a washing machine or a snatch-and-run theft on the street.
[1257] 4. Emotion recognition means
[1258] The server analyzes the user's facial expressions and tone of voice and uses an emotion recognition module to recognize the user's emotions in real time.
[1259] The recognized emotional information is reflected in the generation of warning messages.
[1260] 5. Warning statement generation means
[1261] The server detects abnormal behavior and generates an appropriate warning message using a generative AI model based on the emotion recognition results.
[1262] The generated warning message will be based on the user's emotions. For example, a gentle warning message will be generated to calm the user.
[1263] 6. Warning message reading method
[1264] The device (home smart speaker, street announcement system) uses a voice generation module to convert the warning text into audio and read it out loud.
[1265] For example, a voice may be heard saying to a child, "Playing with the washing machine is dangerous. Stop immediately."
[1266] 7. Information sharing methods
[1267] The server sends the generated warning message and abnormal behavior information to the administrator in real time.
[1268] Users (parents or administrators) can receive warning information on their smartphones or dedicated apps and respond quickly.
[1269] Specific examples
[1270] Watching over children at home
[1271] The server acquires real-time video from cameras in the home and transmits the video data to the video analysis module.
[1272] The server's video analysis module analyzes the child's behavior of opening the washing machine door, and the abnormal behavior detection module detects this as abnormal behavior.
[1273] The server uses an emotion recognition module to analyze the child's facial expressions and tone of voice to obtain emotional information.
[1274] The server generates a warning message saying, "Playing with the washing machine is dangerous. Stop immediately," and adjusts the tone based on the emotion recognition results.
[1275] The device (smart speaker) reads out the warning message aloud, and at the same time, the server sends the warning information to the parent's smartphone, allowing the parent to check the information immediately.
[1276] Street crime prevention surveillance
[1277] The server acquires real-time video from surveillance cameras installed on the street and transmits the video data to the video analysis module.
[1278] The server's video analysis module detects snatching attempts, and the abnormal behavior detection module determines this as abnormal behavior.
[1279] The server uses an emotion recognition module to analyze the facial expressions and tone of voice of people around it to obtain emotional information.
[1280] The server generates a warning message saying, "Purse snatching has been detected here. Please leave immediately," and adjusts the tone based on emotional information.
[1281] The terminal (street announcement system) reads out the warning message, and at the same time, the server transmits the warning information to the police or monitoring center, enabling immediate action.
[1282] Program processing explanation
[1283] The server acquires video data from the surveillance camera and stores it in a buffer.
[1284] A video analysis module in the server analyzes the video data and identifies abnormal behavior.
[1285] When abnormal behavior is detected, the server activates the emotion recognition module and analyzes the user's emotional information.
[1286] Based on the emotional information, the server uses a generative AI model to generate an appropriate warning message.
[1287] The generated warning message is sent to the specified device and read aloud.
[1288] The warning message and abnormal behavior information are sent to the administrator in real time by the server, and the administrator (user) can check the warning information on a smartphone or other device.
[1289] In this way, it is possible to provide a system that can quickly detect abnormal behavior, recognize emotions, issue appropriate warnings, and share information.
[1290] The processing flow will be explained below.
[1291] Step 1:
[1292] The server acquires video data in real time from home cameras and street surveillance cameras. This video data is broken down into multiple frames and stored in the server's temporary memory (buffer).
[1293] Step 2:
[1294] The server sends the acquired video data to the video analysis module, which converts the video data into a format that can be processed in real time and then provides it for analysis.
[1295] Step 3:
[1296] The video analysis module in the server processes the video data and applies object detection and motion analysis algorithms to analyze the movement, shape, and position of people and objects, preparing to identify abnormal behavior.
[1297] Step 4:
[1298] The server's abnormal behavior detection module determines abnormal behavior based on the video analysis results. For example, it detects a child trying to get into a washing machine or a person moving abnormally in a specific area.
[1299] Step 5:
[1300] After the server detects abnormal behavior, it activates the emotion recognition module, which analyzes the user's facial expressions and tone of voice from the video data to recognize their emotions.
[1301] Step 6:
[1302] The server uses a generative AI model to generate appropriate warning messages based on the emotion recognition results. During this generation process, the tone and content of the warning messages are adjusted to match the user's emotions.
[1303] Step 7:
[1304] The server generates a warning message and sends it to a designated device, such as a smart speaker in the home or an announcement system on the street.
[1305] Step 8:
[1306] The warning message received by the device is converted into voice by a voice generation module and actually read aloud from a speaker. For example, at home, a warning such as "Playing with the washing machine is dangerous. Stop immediately" is played in a soft tone.
[1307] Step 9:
[1308] At the same time, the server sends the generated warning message and detailed information about the abnormal behavior to the administrator in real time. The administrator (user) receives the warning information on their smartphone or a dedicated application, enabling them to respond quickly. For example, a parent could receive a notification on their smartphone and check on their child's safety.
[1309] In this way, the processing flow from step 1 to step 9 provides a system that quickly detects abnormal behavior and provides appropriate warnings and information sharing that take the user's emotions into consideration.
[1310] Example 2
[1311] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1312] Conventional surveillance systems were able to detect abnormal behavior by acquiring and analyzing video data, but they lacked the ability to issue prompt and appropriate warnings and share information while taking into account the user's emotions. As a result, when abnormal behavior occurred, responses were sometimes delayed, making it difficult to implement effective preventative measures.
[1313] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a video acquisition means, a video analysis means, an abnormal behavior detection means, an emotion recognition means, a warning message generation means, a warning message reading means, and an information sharing means. This makes it possible to issue a prompt and accurate warning that takes into account the user's emotions and to share abnormal behavior information in real time.
[1314] The "image acquisition means" is a means for acquiring image data from a camera device in real time.
[1315] The "video analysis means" is a means for processing the acquired video data and performing object detection and motion analysis.
[1316] The "abnormal behavior detection means" is a means for detecting specific abnormal behavior based on the analysis results of the video analysis means.
[1317] The "emotion recognition means" is a means for analyzing the user's facial expressions and voice and recognizing their emotions.
[1318] The "warning message generation means" is a means for generating a warning message according to the user's emotions using a generative AI model.
[1319] The "means for reading out a warning message" is a means for reading out the generated warning message by voice.
[1320] The "information sharing means" is a means for transmitting the generated warning message and abnormal behavior information to the administrator in real time.
[1321] This system is implemented by combining the following elements: video acquisition means, video analysis means, abnormal behavior detection means, emotion recognition means, warning message generation means, warning message reading means, and information sharing means.
[1322] System Configuration
[1323] 1. Video acquisition method:
[1324] The server acquires video data from the camera device in real time. For example, a home IP camera is connected to the server and transmits video using the RTSP protocol. This video data is stored in a buffer on the server.
[1325] 2. Video analysis methods:
[1326] The video analysis module in the server processes the acquired video data and performs object detection and motion analysis. Specifically, it uses deep learning algorithms such as YOLO and SSD to detect objects in the video frame in real time.
[1327] 3. Abnormal behavior detection methods:
[1328] The server's abnormal behavior detection module detects specific abnormal behavior based on the video analysis results and determines whether the behavior is abnormal by comparing the detected behavior pattern with a predefined abnormal behavior database.
[1329] 4. Emotion recognition means:
[1330] The server analyzes the user's facial expressions and voice and recognizes the user's emotions in real time using an emotion recognition module. For example, OpenFace can be used for facial expression recognition and OpenSMILE for voice analysis.
[1331] 5. Warning statement generation means:
[1332] The server detects abnormal behavior and generates an appropriate warning message using a generative AI model based on the emotion recognition results. Specifically, a prompt message is input into a generative model such as GPT-4 to generate the warning message.
[1333] Example prompt sentence:
[1334] "Playing with the washing machine is dangerous. Stop immediately."
[1335] 6. Warning message reading method:
[1336] The device (such as a smart speaker in the home or a public address system) uses a speech generation module to convert the warning text into a voice and read it aloud. A text-to-speech (TTS) engine is used to generate an audio file, which is then played on the device.
[1337] 7. Information sharing methods:
[1338] The server generates a warning message and sends the abnormal behavior information to the administrator in real time. The administrator (e.g., parent or police) receives the warning information on their smartphone or a dedicated app and can respond promptly.
[1339] Specific examples
[1340] Watching over children at home
[1341] The server acquires real-time video from cameras in the home and transmits the video data to the video analysis module.
[1342] The server's video analysis module analyzes the child's behavior of opening the washing machine door, and the abnormal behavior detection module detects this as abnormal behavior.
[1343] The server uses an emotion recognition module to analyze the child's facial expressions and tone of voice to obtain emotional information.
[1344] The server generates a warning message saying, "Playing with the washing machine is dangerous. Stop immediately," and adjusts the tone based on the emotion recognition results.
[1345] The device (smart speaker) reads out the warning message aloud, and at the same time, the server sends the warning information to the parent's smartphone, allowing the parent to check the information immediately.
[1346] Street crime prevention surveillance
[1347] The server acquires real-time video from cameras installed on the street and sends the video data to the video analysis module.
[1348] The server's video analysis module detects snatching attempts, and the abnormal behavior detection module determines this as abnormal behavior.
[1349] The server uses an emotion recognition module to analyze the facial expressions and tone of voice of people around it to obtain emotional information.
[1350] The server generates a warning message saying, "Purse snatching has been detected here. Please leave immediately," and adjusts the tone based on emotional information.
[1351] The terminal (street announcement system) reads out the warning message, and at the same time, the server transmits the warning information to the police or monitoring center, enabling immediate action.
[1352] In this way, a system can be built that can quickly detect abnormal behavior in real time, issue warnings, and share information.
[1353] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1354] Step 1:
[1355] The server receives video data from the camera device in real time. The input is the video stream from the camera, and the output is the video data stored in the server's buffer. Specifically, the server communicates with the IP camera using the RTSP protocol and stores the received data in a buffer.
[1356] Step 2:
[1357] The video analysis module in the server retrieves video data from the buffer and performs object detection and motion analysis. The input is the video data stored in the buffer, and the output is the analyzed object and motion information. Specifically, it uses a deep learning algorithm (such as YOLO or SSD) to detect objects in the video frame in real time.
[1358] Step 3:
[1359] The server's abnormal behavior detection module detects specific abnormal behavior based on the video analysis results. The input is the output of the video analysis module, and the output is the abnormal behavior detection results. Specifically, the analysis results are compared with a predefined abnormal behavior database to determine abnormal behavior. For example, a child trying to open the washing machine door is detected as abnormal behavior.
[1360] Step 4:
[1361] The server uses an emotion recognition module to analyze the user's facial expressions and tone of voice to recognize emotions. The input is the user's face and voice data obtained through video analysis, and the output is the recognized emotional information. Specific operations include running a facial expression recognition model (e.g., OpenFace) and a voice analysis model (e.g., OpenSMILE). It may recognize emotions such as fear or surprise from a child's face.
[1362] Step 5:
[1363] The server detects abnormal behavior and generates a warning message using a generative AI model based on the emotion recognition results. The input is the abnormal behavior detection results and emotion recognition results, and the output is the generated warning message. Specifically, a prompt message is input into a generative AI model such as GPT-4 to generate a warning message. For example, a warning message such as "Playing with the washing machine is dangerous. Stop immediately" is generated.
[1364] Step 6:
[1365] The device (such as a home smart speaker or a street announcement system) uses a voice generation module to read out the warning text. The input is the generated warning text, and the output is a voiced version of the warning text. Specifically, the generated warning text is input into a text-to-speech (TTS) engine, which generates an audio file and plays it back.
[1366] Step 7:
[1367] The server generates a warning message and sends the abnormal behavior information to the administrator in real time. The input is the warning message and abnormal behavior information, and the output is the warning information sent to the administrator's device. Specifically, the abnormal behavior and warning message are sent via push notification to the administrator's smartphone or dedicated app, allowing parents, police, etc. to respond quickly.
[1368] (Application example 2)
[1369] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1370] While conventional video surveillance systems can detect abnormal behavior, they are unable to generate appropriate warnings based on the user's emotions and share information quickly. As a result, appropriate responses to abnormal behavior are delayed, making it difficult to provide effective security and monitoring. To solve this problem, the present invention aims to provide a system that integrates abnormal behavior detection and user emotion recognition to generate appropriate warnings and share information.
[1371] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1372] In this invention, the server includes a video acquisition means, a video analysis means, an abnormal behavior detection means, an emotion recognition means, a warning message generation means, a warning message reading means, and an information sharing means. This allows the server to acquire video from a surveillance camera in real time and analyze the video data to detect abnormal behavior. It can also recognize emotions in real time from a user's facial expression and tone of voice and generate an appropriate warning message based on the emotion information. The generated warning message is read aloud, and the warning information and abnormal behavior information are further transmitted to an administrator in real time, enabling a prompt and effective response.
[1373] "Video acquisition means" refers to a device or function that acquires video data in real time from a surveillance camera, a home camera, or the like.
[1374] "Video analysis means" refers to modules and algorithms for processing acquired video data and performing object detection and motion analysis.
[1375] "Abnormal behavior detection means" refers to a module or device for detecting specific abnormal behavior based on the results of video analysis.
[1376] "Emotion recognition means" refers to modules or algorithms that analyze the user's facial expressions and tone of voice and recognize emotions in real time.
[1377] A "warning message generation means" is a device or software that uses a generative AI model to generate an appropriate warning message based on the results of emotion recognition.
[1378] The "means for reading out a warning message" is a device or function for reading out the generated warning message aloud.
[1379] The "information sharing means" is a device or system for transmitting the generated warning message and abnormal behavior information to the administrator in real time.
[1380] A "generative AI model" is an artificial intelligence model for generating natural language based on specific input data.
[1381] A "prompt sentence" is an input sentence that causes a generative AI model to generate a specific output.
[1382] The present invention relates to a system that integrates video acquisition, video analysis, abnormal behavior detection, emotion recognition, warning message generation, warning message reading, and information sharing. This system can analyze video data acquired from surveillance cameras and home cameras in real time to detect abnormal behavior. Furthermore, it has the ability to recognize user emotions, generate appropriate warning messages based on those emotions, and read them aloud. Furthermore, the generated warning messages and abnormal behavior information are sent to administrators in real time, enabling prompt and effective response.
[1383] System configuration:
[1384] 1. Video acquisition method
[1385] The server receives video data from surveillance cameras and home cameras in real time, stores the video in a buffer on the server, and then analyzes it.
[1386] 2. Video analysis methods
[1387] The server's video analysis module processes the captured video data and performs object detection and motion analysis. The server uses deep learning algorithms to analyze patterns of abnormal behavior.
[1388] 3. Abnormal Behavior Detection Methods
[1389] The server's abnormal behavior detection module determines abnormal behavior based on the video analysis results, such as falls, trespassing in specific areas, and dangerous behavior.
[1390] 4. Emotion recognition means
[1391] The server analyzes the user's facial expressions and tone of voice, and uses an emotion recognition module to recognize the user's emotions in real time. The recognized emotional information is reflected in the generation of warning messages.
[1392] 5. Warning statement generation means
[1393] The server detects abnormal behavior and generates an appropriate warning message based on the emotion recognition results using a generative AI model. The generated warning message is tailored to the user's emotions.
[1394] 6. Warning message reading method
[1395] The device (such as a smart speaker in the home or a street announcement system) uses a voice generation module to convert the warning message into audio and read it aloud. For example, a voice might say, "Falling is dangerous. Please sit down and rest."
[1396] 7. Information sharing methods
[1397] The server generates a warning message and sends abnormal behavior information to the administrator in real time. The user (administrator) receives the warning information on their smartphone or a dedicated app, allowing them to respond quickly.
[1398] Examples:
[1399] Elderly monitoring system:
[1400] When an elderly person falls, abnormal behavior is detected from camera footage and, based on emotion recognition, a gentle warning message is generated, such as "Falling is dangerous. Please sit down and rest." The warning message is read aloud by the smart speaker and this information is sent to family members' smartphones in real time, encouraging them to take prompt action.
[1401] Examples of prompts:
[1402] "If the user is sad and falls, create a warning message in a gentle tone encouraging them to rest."
[1403] "If the user has a neutral emotion, generate a warning message to prevent them from falling."
[1404] Hardware and software used:
[1405] The hardware used is a surveillance camera, a home camera, a cloud server, a smart speaker, and a smartphone, while the software used is a video analysis module, an emotion recognition module, a generative AI model, a warning message generation module, and a notification sending module.
[1406] This system efficiently carries out a series of processes, from real-time acquisition and analysis of video data, detection of abnormal behavior, emotion recognition, generation and voice reading of appropriate warning messages, and information sharing, thereby achieving effective security and monitoring.
[1407] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1408] Step 1:
[1409] The server acquires video data in real time from surveillance cameras and home cameras. The video is stored in the server's buffer. The input is real-time video data from the camera, and the output is video data stored in the server's buffer. Specifically, the server acquires video data frame by frame from the camera stream and stores it in memory.
[1410] Step 2:
[1411] The server's video analysis module processes the acquired video data and performs object detection and motion analysis. The input is the video data stored in the buffer, and the output is the analysis results (e.g., object position and motion information). Specifically, the server uses a deep learning algorithm to detect face and body movements from the video data.
[1412] Step 3:
[1413] The server's abnormal behavior detection module detects specific abnormal behavior based on the video analysis results. The input is the result data of object detection and motion analysis, and the output is the abnormal behavior detection results. Specifically, the server compares motion patterns and determines abnormal behavior such as falls or dangerous behavior.
[1414] Step 4:
[1415] The server's emotion recognition module analyzes the user's facial expressions and tone of voice to recognize emotions in real time. The input is facial and voice data, and the output is recognized emotional information. Specifically, deep learning is used to estimate emotions such as joy, anger, sadness, and happiness from the user's facial and voice data.
[1416] Step 5:
[1417] The server detects abnormal behavior and generates an appropriate warning message using a generative AI model based on the emotion recognition results. The inputs are the abnormal behavior detection results and emotion recognition results, and the output is the generated warning message. Specifically, the server inputs a prompt message into the generative AI model and generates an appropriate warning message according to the emotion. For example, the prompt message for the generative AI model could be, "If the user is sad and falls, please create a warning message in a gentle tone encouraging them to rest."
[1418] Step 6:
[1419] The device (such as a home smart speaker or a street announcement system) uses a voice generation module to convert the warning text into audio and read it out loud. The input is the generated warning text, and the output is a voice warning. Specifically, the device uses text-to-speech conversion technology to play back the generated warning text aloud.
[1420] Step 7:
[1421] The server sends the generated warning message and abnormal behavior information to the administrator in real time. The input is the warning message and abnormal behavior information, and the output is a notification to the administrator's smartphone or a dedicated app. Specifically, the server uses a notification sending module to send the warning information as an email or push notification.
[1422] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1423] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1424] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1425] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1426] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1427] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1428] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1429] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1430] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1431] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1432] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1433] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1434] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1435] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1436] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1437] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1438] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1439] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1440] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1441] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1442] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1443] The following is further disclosed regarding the above embodiment.
[1444] (Claim 1)
[1445] An image acquisition means;
[1446] Video analysis means;
[1447] An abnormal behavior detection means;
[1448] a warning message generating means;
[1449] a warning reading means;
[1450] Information sharing means;
[1451] A system including:
[1452] (Claim 2)
[1453] 2. The system according to claim 1, wherein the image acquisition means is a means for acquiring real-time images from a surveillance camera.
[1454] (Claim 3)
[1455] 2. The system according to claim 1, wherein the abnormal behavior detection means is a means for detecting specific abnormal behavior based on video analysis results.
[1456] "Example 1"
[1457] (Claim 1)
[1458] An image acquisition means;
[1459] Video analysis means;
[1460] An abnormal behavior detection means;
[1461] a warning message generating means;
[1462] a warning reading means;
[1463] Information sharing means;
[1464] A warning message generation means using a generative AI model;
[1465] a means for processing the video in real time;
[1466] a terminal that converts the generated warning text into audio and reads it out loud;
[1467] A system including:
[1468] (Claim 2)
[1469] The system of claim 1, wherein the video acquisition means is a means for acquiring real-time video from a surveillance camera, and the video analysis means analyzes the video data using a deep learning algorithm.
[1470] (Claim 3)
[1471] The system of claim 1, wherein the abnormal behavior detection means is a means for detecting specific abnormal behavior in real time based on video analysis results, and includes an information sharing means for notifying an administrator of abnormal behavior information in real time.
[1472] "Application Example 1"
[1473] (Claim 1)
[1474] An image acquisition means;
[1475] Video analysis means;
[1476] An abnormal behavior detection means;
[1477] a warning message generating means;
[1478] a warning reading means;
[1479] Information sharing means;
[1480] means for displaying a warning statement on a display of the smart glasses;
[1481] means for issuing an audio warning;
[1482] A means of sharing abnormal behavior information with the management center;
[1483] A system including:
[1484] (Claim 2)
[1485] 2. The system according to claim 1, wherein the image acquisition means is a means for acquiring real-time images from a surveillance camera.
[1486] (Claim 3)
[1487] 2. The system according to claim 1, wherein the abnormal behavior detection means is a means for detecting specific abnormal behavior based on video analysis results.
[1488] "Example 2: Combining Emotion Engines"
[1489] (Claim 1)
[1490] An image acquisition means;
[1491] Video analysis means;
[1492] An abnormal behavior detection means;
[1493] An emotion recognition means;
[1494] a warning message generating means;
[1495] a warning reading means;
[1496] Information sharing means;
[1497] A system including:
[1498] (Claim 2)
[1499] 2. The system according to claim 1, wherein the image acquisition means is a means for acquiring real-time images from a camera device.
[1500] (Claim 3)
[1501] 2. The system according to claim 1, wherein the abnormal behavior detection means is a means for detecting specific abnormal behavior based on video analysis results.
[1502] (Claim 4)
[1503] 2. The system according to claim 1, wherein the emotion recognition means is a means for recognizing emotions by analyzing the user's facial expressions and voice.
[1504] (Claim 5)
[1505] 2. The system according to claim 1, wherein the warning message generation means is a means for generating a warning message according to the user's emotions using a generation AI model.
[1506] (Claim 6)
[1507] 2. The system according to claim 1, wherein the warning message reading means is means for reading the generated warning message aloud.
[1508] (Claim 7)
[1509] 2. The system according to claim 1, wherein the information sharing means is means for transmitting the generated warning message and abnormal behavior information to an administrator in real time.
[1510] "Application example 2 when combining emotion engines"
[1511] (Claim 1)
[1512] An image acquisition means;
[1513] Video analysis means;
[1514] An abnormal behavior detection means;
[1515] An emotion recognition means;
[1516] a warning message generating means;
[1517] a warning reading means;
[1518] Information sharing means;
[1519] A system including:
[1520] (Claim 2)
[1521] 2. The system according to claim 1, wherein the image acquisition means is a means for acquiring real-time images from a surveillance camera.
[1522] (Claim 3)
[1523] 2. The system according to claim 1, wherein the abnormal behavior detection means is a means for detecting specific abnormal behavior based on video analysis results.
[1524] (Claim 4)
[1525] 2. The system according to claim 1, wherein the emotion recognition means is a means for analyzing a user's facial expression and tone of voice to recognize emotions in real time.
[1526] (Claim 5)
[1527] 2. The system according to claim 1, wherein the warning message generation means is a means for generating an appropriate warning message based on the results of emotion recognition using a generative AI model.
[1528] (Claim 6)
[1529] 2. The system according to claim 1, wherein the warning message reading means is means for reading the generated warning message aloud.
[1530] (Claim 7)
[1531] 2. The system according to claim 1, wherein the information sharing means is means for transmitting the generated warning message and abnormal behavior information to an administrator in real time.
[1532] (Claim 8)
[1533] The system of claim 1 uses a combination of an emotion recognition means as set forth in claim 4, a warning message generation means using a generative AI model as set forth in claim 5, and a warning message reading means as set forth in claim 6. [Explanation of symbols]
[1534] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. An image acquisition means; Video analysis means; An abnormal behavior detection means; a warning message generating means; a warning reading means; Information sharing means; A system including:
2. 2. The system according to claim 1, wherein the image acquisition means is means for acquiring real-time images from a surveillance camera.
3. 2. The system according to claim 1, wherein the abnormal behavior detection means is means for detecting specific abnormal behavior based on video analysis results.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A