System
The system addresses high costs and slow response times in security systems by using real-time video analysis and AI for anomaly detection and automatic notification, enhancing crime prevention through rapid and accurate risk assessment.
Patent Information
- Application Number
- JP2024121613
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-26
- Publication Date
- 2026-02-05
AI Technical Summary
Current security systems face high costs, frequent malfunctions, and slow response times, with a lack of efficient systems for quickly and accurately assessing risks and notifying external agencies.
A system that includes real-time video analysis, anomaly detection, risk assessment, automatic notification, and behavior inference using surveillance cameras and AI, enabling rapid and accurate detection of abnormalities and automatic notification to external organizations.
The system achieves highly accurate anomaly detection, rapid risk assessment, and effective notification, significantly improving crime prevention by reducing costs and shortening response times.
Smart Images

Figure 2026019865000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Current security systems face several challenges, including high costs, frequent malfunctions, and slow response times. Continuously monitoring surveillance camera footage requires enormous resources, making them practically inefficient. Furthermore, there is a lack of systems that can quickly and accurately assess risks and appropriately notify external agencies. Given these circumstances, there is a demand for lower-cost, more accurate security systems. [Means for solving the problem]
[0005] The present invention provides a system that includes a means for acquiring video footage, a means for analyzing the acquired video footage in real time to detect abnormalities, a means for assessing the risk of an abnormality when it is detected, a means for automatically notifying an external organization of the analysis results when the risk is assessed to be high, and a means for inferring the next course of action using past data. This system can detect abnormal behavior not only from surveillance camera footage but also from other sources with high accuracy, and quickly notify external organizations only when necessary. This reduces costs, prevents malfunctions, and shortens response times.
[0006] "Video" refers to image data and audio data captured by imaging equipment such as a surveillance camera.
[0007] "Means for acquiring" refers to a function or device for acquiring video using a surveillance camera or other imaging device.
[0008] "Means for analysis" refers to programs and algorithms that analyze acquired video data in real time and detect abnormalities.
[0009] An "anomaly" is a phenomenon that includes movements, sounds, or specific behavioral patterns that are different from normal situations and are deemed to pose a security risk.
[0010] A "means for assessing risk" is an algorithm or model for assessing the risk of an abnormal behavior based on the detected behavior.
[0011] "Means of notification" refers to communication functions and protocols for automatically notifying external agencies of information when a risk is assessed as high.
[0012] "External agencies" include, but are not limited to, public agencies such as police and security-related organizations.
[0013] "Past data" refers to records of crimes and accidents that have occurred in the past and data related to these.
[0014] "Means for inferring next actions" are algorithms or models that use past data to predict and infer future actions from the current situation. [Brief explanation of the drawings]
[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0017] First, the terms used in the following description will be explained.
[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0023] [First embodiment]
[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0036] The present invention is a high-precision security system that uses surveillance cameras and combines real-time analysis of video, anomaly detection, risk assessment, automatic notification, behavior inference, etc. Specific embodiments of the present invention are described below.
[0037] System Configuration
[0038] 1. Surveillance camera (terminal)
[0039] The device is equipped with a high-resolution camera for wide-area monitoring and a built-in AI analysis engine for initial analysis.
[0040] 2. Central Server (Server)
[0041] The server receives the video data sent from each camera and performs advanced AI analysis. Notifications to external agencies such as the police are also sent from this server.
[0042] The server stores a large amount of past data, and the next action is inferred based on this data.
[0043] 3. User terminal (user)
[0044] Users, such as system administrators and police, have an interface for monitoring and managing security.
[0045] Program processing
[0046] Real-time video analysis
[0047] The device captures video 24 hours a day and analyzes it in real time using a built-in AI analysis engine. For example, a store's surveillance camera constantly captures human movement and audio, and if it detects abnormal behavior or loud voices, it automatically triggers an event.
[0048] Anomaly detection and risk assessment
[0049] When the device detects an abnormality, it immediately assesses the risk based on a previously trained abnormal behavior model. For example, if it detects a person with a knife in front of a store, it will assess the risk of robbery as high.
[0050] Sending data
[0051] If the anomaly is assessed as high risk, the video data is sent to a server, where it is compressed and quickly transmitted over the network.
[0052] Advanced AI analysis and notifications
[0053] The server analyzes the received data and performs a more detailed risk assessment. At the same time, it compares it with past crime data and infers the next action. For example, based on abnormal behavior data in front of a store, the server may infer that the next robber is likely to enter the store.
[0054] Notification to external agencies
[0055] If the risk is deemed to be extremely high, the server automatically notifies external agencies such as the police, and the notification includes the risk assessment results, video data, and inferences about the next course of action.
[0056] Specific examples
[0057] The device captured suspicious activity in front of the store, including a loud voice and a person holding a knife.
[0058] The device immediately detects an abnormality, assesses the risk, and if it determines that the risk is high, sends several minutes of video footage to the server.
[0059] The server analyzes the video and determines that there is a high risk of robbery. Based on past crime data, it infers that the next move is likely to be to break into a store.
[0060] The server automatically notifies the police and provides relevant information (risk assessment, footage, inference results).
[0061] Police can quickly rush to the scene and prevent incidents from occurring. During this time, the server provides real-time video footage of the scene to the police.
[0062] In this way, the system of the present invention overcomes the challenges posed by conventional security systems through highly accurate anomaly detection, rapid risk assessment, and effective notification, and can significantly improve crime prevention effectiveness.
[0063] The processing flow will be explained below.
[0064] Step 1:
[0065] The device constantly operates the surveillance camera and captures video in real time, and this video data is sent to the camera's internal initial analysis engine.
[0066] Step 2:
[0067] The device's built-in analysis engine analyzes the captured video data in real time, analyzing movement, sound, behavioral patterns, etc. to determine whether or not an abnormality is detected. For example, if a person is holding a knife, it will detect this as an abnormality.
[0068] Step 3:
[0069] When a device detects an anomaly, the anomaly triggers a risk assessment, which calculates the risk of abnormal behavior based on a pre-trained model.
[0070] Step 4:
[0071] If the device assesses the risk as high, it compresses the video data of the anomalous event and prepares it for transmission to the server, including the time, location, and video clip of the anomalous behavior.
[0072] Step 5:
[0073] The server receives the data sent from the terminal, then extracts the data and prepares it for reanalysis.
[0074] Step 6:
[0075] The server then begins detailed AI analysis. Based on the initial analysis results, it compares them with past criminal data and the behavioral patterns of former criminals. If similar patterns are detected, the risk assessment is further refined.
[0076] Step 7:
[0077] The server infers the next course of action based on the results of the risk assessment. For example, it calculates the probability that a person carrying a knife will turn into a robber based on past data.
[0078] Step 8:
[0079] If the server determines that the risk is high, it automatically notifies external agencies such as the police, along with the results of the inference. The notification includes the risk assessment results, video data, and an inference on the next course of action.
[0080] Step 9:
[0081] The user (police officer or security officer) receives notifications from the server and takes action as needed, checks real-time video footage, and prepares for a response to the scene.
[0082] Step 10:
[0083] The server stores the notification and subsequent response as a log, which can be used for future analysis and system improvement.
[0084] Step 11:
[0085] If the device detects a new anomaly, it restarts the process from step 1, enabling continuous monitoring and response.
[0086] Example 1
[0087] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0088] In recent years, there has been a growing emphasis on improving security, but conventional security systems still face serious challenges in anomaly detection and risk assessment. For example, the accuracy of anomaly detection is low, and risk assessment is not performed quickly, often resulting in delayed responses. Furthermore, behavioral inference using past data has not been realized, limiting the effectiveness of crime prevention. Furthermore, despite the overwhelming computational resources of cloud servers and data centers, these resources are often not fully utilized, preventing the system from achieving its full potential. To address these challenges, it is necessary to develop a highly functional security system that can perform anomaly detection, risk assessment, data analysis, and notification in real time.
[0089] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0090] In this invention, the server includes means for acquiring video, means for analyzing the acquired video in real time to detect abnormalities, means for assessing the risk when the abnormality is detected, means for transmitting compressed video data to a central processing unit when the risk is assessed to be high, means for the central processing unit to analyze the received data in detail and compare it with a past database to infer the next action, means for automatically notifying an external institution of the inference results and risk assessment results, and means for performing a series of processes from abnormality detection to risk assessment, analysis, and notification in real time. This enables highly accurate abnormality detection and risk assessment, rapid data analysis and notification, and further inference of the next action using past data.
[0091] A "means for acquiring video" is a combination of hardware and software that uses a monitoring device to collect video data within a specific range.
[0092] "Means for analyzing acquired video in real time to detect abnormalities" refers to means for instantly processing collected video data and using algorithms to recognize abnormalities based on pre-set criteria and patterns.
[0093] The "means for assessing risk" refers to a means including an assessment algorithm and a related database for determining the severity and urgency of an abnormality when the abnormality is detected.
[0094] "Means for transmitting compressed video data to a central processing unit" refers to a protocol and associated means for efficiently compressing collected video data and rapidly transmitting it to a central processing unit via a communications network when an abnormality is detected and assessed as high risk.
[0095] A "central processing unit" is a computer system that receives data from multiple monitoring devices via a network, performs advanced analysis, and stores and manages the results.
[0096] "Means of inferring next actions by conducting detailed analysis and comparing it with a past database" refers to an advanced analytical algorithm and matching system that analyzes received video data in more detail and compares and collates it with data accumulated in the past to predict future actions.
[0097] "Means for automatically notifying external organizations of inference results and risk assessment results" refers to a system and communication protocol that automatically notifies appropriate external organizations (e.g., law enforcement agencies) based on analysis results and risk assessment information.
[0098] "A means for carrying out a series of processes in real time, from detecting an anomaly to risk assessment, analysis, and notification" refers to a system and algorithm that executes all processes, from the moment an anomaly is detected, such as risk assessment, data analysis, and notification to external organizations, in real time without any time delay.
[0099] MODE FOR CARRYING OUT THE INVENTION
[0100] The present invention is a highly accurate security system that uses a monitoring device and combines real-time video analysis, anomaly detection, risk assessment, automatic notification, behavior inference, etc. Specific embodiments of the present invention are described below.
[0101] System Configuration
[0102] Monitoring device (terminal)
[0103] The device is equipped with a high-resolution monitoring device for wide-area monitoring. It also has a built-in AI analysis engine for initial analysis. The device captures video 24 hours a day and analyzes it in real time using the built-in AI analysis engine. For example, a store's surveillance camera constantly captures human movement and audio, and if it detects abnormal behavior or loud voices, it automatically triggers an event.
[0104] Central Processing Unit (Server)
[0105] The server receives the video data sent from each device and performs advanced AI analysis. A large amount of past data is stored on the server, and the next course of action is inferred based on this data. If the risk is determined to be extremely high, the server automatically notifies an external agency, such as a security agency. This notification includes the risk assessment results, video data, and the inferred next course of action.
[0106] User terminal (user)
[0107] Users, such as system administrators and law enforcement agencies, are provided with an interface for monitoring and managing security.
[0108] Specific Examples
[0109] 1. The device captures suspicious activity in front of the store, such as a loud voice or a person holding a knife.
[0110] 2. The device immediately detects the abnormality and assesses the risk. If it determines that the risk is high, it sends several minutes of video footage to the server.
[0111] 3. The server performs advanced analysis of the received video data and compares it with a database of past crimes. If it determines that there is a high risk of robbery, it infers that the next move is likely to be breaking into a store.
[0112] 4. The server automatically notifies security agencies of any detected anomalies and their risk assessment results. The notification includes video data of the abnormal behavior, the risk assessment results, and the inferred next action.
[0113] 5. The security agency (user) receives the notification and quickly rushes to the scene to respond. During this time, the server provides the security agency with real-time video footage of the scene.
[0114] In this way, the system of the present invention overcomes the challenges posed by conventional security systems through highly accurate anomaly detection, rapid risk assessment, and effective notification, and can significantly improve crime prevention effectiveness.
[0115] Example prompts for generative AI models
[0116] "Please explain how an anomaly detection system works in real-time video analysis. Please also explain in detail what processes are performed from the perspectives of the device, server, and user. Please also provide specific examples."
[0117] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0118] Step 1:
[0119] The device captures video 24 hours a day.
[0120] Specific operation: The device uses a high-resolution camera to constantly capture video data of the monitored area. The video data collected by the camera is temporarily stored in the internal memory.
[0121] Input: Video of the monitored area
[0122] Output: High-resolution video data
[0123] Step 2:
[0124] The video data acquired by the device is analyzed in real time to detect abnormalities.
[0125] How it works: The AI analysis engine analyzes the movements, sounds, and types of objects in the video to detect abnormal patterns. For example, if abnormal movements or loud voices are detected, an alert will be automatically generated.
[0126] Input: High-resolution video data
[0127] Output: Anomaly detection results (e.g., abnormal behavior detected)
[0128] Step 3:
[0129] If the device detects an abnormality, it evaluates the risk.
[0130] How it works: The built-in model performs a risk assessment based on the type of anomaly detected and the surrounding circumstances. For example, if a person holding a knife is captured on video, the risk will be assessed as "high."
[0131] Input: Anomaly detection result
[0132] Output: Risk assessment result (e.g., high risk)
[0133] Step 4:
[0134] If the device is assessed as being at high risk, the video data is compressed and sent to a central processing unit (server).
[0135] Specific operation: Using a compression algorithm, the size of the video data is reduced and quickly sent to the server over the network.
[0136] Input: Risk assessment results, high-resolution video data
[0137] Output: Compressed video data
[0138] Step 5:
[0139] The server analyzes the received data in detail and compares it with a past database to infer the next action.
[0140] Specific operation: The server analyzes the received video data and searches a database for similar abnormal behavior. For example, if the data matches past data on robbery, the server infers that the next behavior is likely to be an attempted robbery.
[0141] Input: Compressed video data
[0142] Output: Behavioral inference result (e.g., probability of attempted robbery)
[0143] Step 6:
[0144] The server automatically notifies external agencies based on the inference results and risk assessment results.
[0145] Specific operation: The server's notification system automatically generates risk assessment results and behavioral inference results and notifies security agencies in real time.
[0146] Input: Risk assessment results, behavioral inference results
[0147] Output: Notification to external agencies (e.g. emergency contact with security agencies)
[0148] Step 7:
[0149] The user (security agency or system administrator) receives a notification, checks the situation on site in real time, and takes action.
[0150] Specific operation: The security agency receives the notification and decides to dispatch to the scene based on the inference results and risk assessment. At the same time, the server provides the security agency with real-time video footage of the scene to assist in the on-site response.
[0151] Input: Notification to external agency
[0152] Output: On-site response (e.g., police dispatch)
[0153] In this way, the system of the present invention can achieve highly accurate anomaly detection, rapid risk assessment, detailed data analysis, effective notification, and real-time on-site response support.
[0154] (Application example 1)
[0155] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0156] Conventional surveillance camera systems require manual monitoring, which requires a great deal of effort and cost. Furthermore, if abnormalities are not detected and risks are not assessed promptly, it becomes difficult to prevent serious incidents. In particular, there is a lack of support for security guards to respond quickly on-site. The purpose of this invention is to solve these issues and build an efficient security system by achieving highly accurate abnormality detection, risk assessment, and real-time notifications.
[0157] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0158] In this invention, the server includes a system using a wearable device worn by a security guard, which includes means for acquiring video, means for analyzing the acquired video in real time to detect abnormalities, means for assessing the risk when the abnormality is detected, means for automatically notifying an external institution of the analysis result when the risk is assessed to be high, means for inferring the next action using past data, and means for issuing visual and audio warnings when an abnormality is detected, thereby enabling the security guard to respond quickly and accurately on site.
[0159] "Means for acquiring images" refers to the function of collecting visual information using surveillance cameras, wearable devices, etc.
[0160] "Means for analyzing acquired video in real time and detecting abnormalities" refers to a function that instantly processes collected video data and detects specific abnormal behavior or abnormal conditions.
[0161] "Means for assessing the risk" refers to a function that instantly determines the danger and impact of a detected abnormality and evaluates the level of risk.
[0162] "Means for automatically notifying external organizations of analysis results" refers to a function that, if the detected risk is high, judges the information and notifies relevant organizations such as the police.
[0163] "Means of inferring next actions using past data" is a function that uses previously collected and analyzed data to predict the next possible action or incident based on the current situation.
[0164] "Visual and audio warning means" refers to a function that notifies security guards via a display or audio alert when an abnormality is detected.
[0165] A "wearable device" is a device worn by security guards and has the ability to obtain and notify information in real time.
[0166] This invention is a system that combines real-time video analysis, anomaly detection, risk assessment, automatic notification, and behavioral inference to realize a highly accurate security system. This system consists of the following three main components:
[0167] System Components
[0168] 1. Terminal
[0169] The device is equipped with a high-resolution camera that continuously captures video of the monitored area. The device also includes a built-in AI analytics engine that analyzes the captured video in real time. This analytics engine uses deep learning models to detect anomalies. For example, a surveillance camera can capture human movement and audio and detect loud voices or suspicious behavior.
[0170] 2. Server
[0171] The server receives the video data sent from the device and performs advanced AI analysis. The server stores and references large amounts of past data, and performs risk assessments of abnormal behavior and infers behavior. If a high risk is determined, the server automatically notifies the police and other relevant agencies along with the analysis results. This server-side analysis makes it possible to predict the next likely event based on past crime data.
[0172] 3. Users
[0173] Users are mainly system administrators, security guards, police, etc., who monitor and control the entire system through a management interface. In particular, the wearable devices (e.g., smart glasses) worn by security guards include a function to notify visual and audio warnings when an abnormality is detected, enabling prompt response on site.
[0174] Explanation of program processing
[0175] The server uses the following hardware and software:
[0176] Hardware: high-performance servers, surveillance cameras, wearable devices.
[0177] Software: OpenCV (image processing library), requests (HTTP request sending library), deep learning model.
[0178] The server first receives the video data sent from the device and analyzes it in real time using a deep learning model. If an abnormality is detected, it evaluates the risk of that abnormality and, if it is determined to be high risk, automatically notifies the relevant authorities of the video data. It also infers the next action based on past data and notifies the predicted results.
[0179] The device mainly acquires video in real time and performs initial analysis. If an abnormality is detected, the video data of the abnormal part is sent to the server.
[0180] For example, a security guard wearing smart glasses will receive an alert when an abnormality is detected, allowing the user to recognize the abnormality visually and audibly and respond quickly.
[0181] Examples and prompts
[0182] A concrete example would be:
[0183] "We would like you to develop a security monitoring application for smart glasses worn by security guards. The application must meet the following requirements:
[0184] 1. Images are acquired from the camera in real time and analyzed using an AI model for anomaly detection.
[0185] 2. If an abnormality is detected, an alert will be displayed on the smart glasses' display and an audio notification will be given through the built-in speaker.
[0186] 3. Compress the abnormal video frames and send them to the security server.
[0187] 4. The SDK of the smart glasses used is assumed to be SmartGlassesAPI.
[0188] Please include specific program code and a description of the libraries you use in your response."
[0189] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0190] Step 1:
[0191] The device captures video data in real time using a high-resolution camera. At this time, the device has a built-in AI analysis engine that performs initial analysis. The input is video data collected in real time, and the output is pre-processed video data for immediate anomaly detection.
[0192] Step 2:
[0193] The video data acquired by the device is analyzed in real time using an AI analysis engine. Anomalies are detected from the analyzed data and the risk of the detected anomalies is assessed. The input is preprocessed video data, and the output is an anomaly score and risk assessment results. The specific operation of this is to calculate an anomaly score for each frame using a deep learning model.
[0194] Step 3:
[0195] The device evaluates the anomaly score, and if it is judged to be high risk, it sends the corresponding video data to the server. At this time, the video data is compressed using a data compression algorithm. The input is the risk assessment result and the original video data, and the output is the compressed high-risk video data.
[0196] Step 4:
[0197] The server receives high-risk video data sent from the device and performs advanced AI analysis. This analysis uses a large amount of past data. The input is compressed high-risk video data, and the output is a detailed risk assessment and next action prediction.
[0198] Step 5:
[0199] The server automatically notifies external agencies such as the police based on the detailed risk assessment results and next action predictions. This notification includes the risk assessment results, analyzed video data, and inferred next action results. The input is the detailed risk assessment results and next action predictions, and the output is notification information to the relevant agencies.
[0200] Step 6:
[0201] The user, a security guard, receives an anomaly detection notification through the smart glasses. The notification is visual (display) and audio (built-in speaker). The input is the anomaly detection notification from the terminal, and the output is a real-time notification to the security guard. Specifically, a warning message is displayed on the smart glasses' display and an alarm sounds from the speaker.
[0202] Step 7:
[0203] The user, a security guard, rushes to the scene and responds quickly based on the detailed information provided by the server. The input is the alert from the smart glasses and additional data from the server, and the output is the response action at the scene. This allows the security guard to make quick and appropriate decisions.
[0204] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0205] The present invention combines an emotion engine with a highly accurate security system using surveillance cameras, and combines real-time video analysis, anomaly detection, risk assessment, automatic notification, behavioral inference, and emotion recognition. Specific embodiments of the present invention are described below.
[0206] System Configuration
[0207] 1. Surveillance camera (terminal)
[0208] The device is equipped with a high-resolution camera for wide-area monitoring and a built-in AI analysis engine for initial analysis.
[0209] 2. Emotion Engine
[0210] The device's built-in emotion engine analyzes the user's facial expressions and tone of voice from video and audio data captured by the camera, and has the ability to recognize emotions.
[0211] 3. Central Server (Server)
[0212] The server receives the video data and emotion recognition data sent from each camera and performs advanced AI analysis. It also notifies external agencies such as the police.
[0213] 4. User terminal (user)
[0214] Users, such as system administrators and police, have an interface for monitoring and managing security.
[0215] Program processing
[0216] Real-time video analysis
[0217] The device captures video 24 hours a day and analyzes it in real time using a built-in AI analysis engine. For example, a store's surveillance camera constantly captures human movement and audio, and if it detects abnormal behavior or loud voices, it automatically triggers an event.
[0218] Emotion recognition and additional anomaly detection
[0219] The emotion engine installed in the device analyzes the user's facial expressions and tone of voice from the captured video and audio data to recognize their emotions. For example, if the user is feeling extremely scared or has an aggressive expression, this will be detected as an abnormality.
[0220] Anomaly detection and risk assessment
[0221] When a device detects an anomaly, it immediately assesses the risk based on previously learned abnormal behavior models and emotional data. The risk assessment takes into account not only movement and voice, but also emotional data. For example, if a distorted facial expression or an angry tone of voice is detected, it is deemed to be a high risk.
[0222] Sending data
[0223] If an anomaly is assessed as high risk, the video and emotion data are sent to a server, where the video data is compressed and quickly transmitted over a network.
[0224] Advanced AI analysis and notifications
[0225] The server analyzes the received data and performs a more detailed risk assessment. At the same time, it compares it with past crime data and emotion data to infer the next action. For example, based on abnormal behavior data in front of a store, the server may infer that the next robber is likely to break into the store.
[0226] Notification to external agencies
[0227] If the risk is deemed to be extremely high, the server automatically notifies external agencies such as the police. This notification includes the risk assessment results, video data, emotional data, and inference results for the next action.
[0228] Specific examples
[0229] The device captured suspicious activity in front of the store, including a loud voice and a person holding a knife.
[0230] The device immediately detects abnormalities and detects emotions such as fear and anger from facial expressions, assesses the risk, determines it to be high risk, and sends several minutes of video and emotional data to a server.
[0231] The server analyzes the video and emotional data and determines that there is a very high risk of robbery. Based on past crime data and emotional data, it infers that the next action is likely to be breaking into a store.
[0232] The server automatically notifies the police and provides relevant information (risk assessment, video footage, emotional data, and inference results).
[0233] Police can quickly rush to the scene and prevent incidents from occurring. During this time, the server provides real-time video footage of the scene to the police.
[0234] In this way, the system of the present invention overcomes the challenges of conventional security systems and significantly improves crime prevention effectiveness through highly accurate anomaly detection, rapid risk assessment, and effective notification. The introduction of an emotion engine enables even more accurate anomaly detection and risk assessment, enabling more appropriate responses.
[0235] The processing flow will be explained below.
[0236] Step 1:
[0237] The device constantly operates the surveillance camera, capturing video and audio in real time, which is then sent to the camera's internal analysis engine.
[0238] Step 2:
[0239] The device's built-in analysis engine analyzes the captured video and audio data in real time, analyzing movement, audio, and behavioral patterns to determine whether or not an abnormality has been detected. For example, if a person is holding a knife, it will detect this as an abnormality.
[0240] Step 3:
[0241] The emotion engine installed in the device analyzes the user's facial expressions and tone of voice from video and audio data to recognize their emotions. If the user is feeling extremely scared or has an aggressive expression, this is detected as an abnormality.
[0242] Step 4:
[0243] The device reassesses the anomaly based on the analysis results of the emotion engine and calculates the risk. For example, if a fearful facial expression or an angry tone of voice is detected, the device will determine that the risk is high.
[0244] Step 5:
[0245] If the device assesses the risk as high, it compresses the video and emotion data of the anomalous event and prepares to send it to the server. The data to be sent includes the time, location, video clip, and emotion data of the anomalous behavior.
[0246] Step 6:
[0247] The server receives the data sent from the device, then extracts the data and prepares it for reanalysis.
[0248] Step 7:
[0249] The server then begins detailed AI analysis. Based on the initial analysis results, it compares the data with past criminal records and the behavioral patterns of former criminals. If similar patterns are detected, the risk assessment is further refined.
[0250] Step 8:
[0251] The server analyzes the data, including emotional data, and infers the next action. For example, it calculates the probability that a person with a knife will turn into a robber based on past data.
[0252] Step 9:
[0253] If the server determines that the risk is high, it will automatically notify external agencies such as the police, along with the inference results. The notification will include the risk assessment results, video data, emotional data, and inferences about the next action.
[0254] Step 10:
[0255] The user (police officer or security officer) receives notifications from the server and takes action as needed, checking real-time video and emotion data and preparing for a response at the scene.
[0256] Step 11:
[0257] The content of the notification from the server and the subsequent response are saved as a log, which can be used for future analysis and system improvement.
[0258] Step 12:
[0259] If the device continues to detect new anomalies, the process restarts from step 1, ensuring continuous monitoring and response.
[0260] Example 2
[0261] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0262] Conventional security systems focused on detecting anomalies in video and assessing risk, but the accuracy of anomaly detection using emotion recognition was not improved, and so they were often insufficient. Furthermore, even if an anomaly was detected, it was difficult to immediately notify external agencies appropriately, making it difficult to respond quickly. This limited the effectiveness of crime prevention, and led to a demand for more accurate and rapid anomaly detection and response.
[0263] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0264] In this invention, the server includes means for acquiring video, means for analyzing the acquired video in real time to detect abnormalities, means for recognizing emotions from the acquired video and audio data, means for assessing the risk when the abnormality is detected, means for automatically notifying an external institution of the analysis results when the risk is assessed to be high, and means for inferring the next action using past data. This enables highly accurate anomaly detection using emotion recognition, and enables effective crime prevention through risk assessment and prompt external notification.
[0265] The "means for acquiring video" refers to a means for collecting video data of the area to be monitored using a terminal such as a surveillance camera.
[0266] "Means for analyzing acquired video in real time to detect abnormalities" refers to means for instantly analyzing acquired video data and identifying suspicious or abnormal behavior.
[0267] The "means for recognizing emotions from acquired video and audio data" refers to a means for analyzing video and audio data and identifying emotions from the subject's facial expressions and tone of voice.
[0268] "Means for assessing the risk" refers to means for determining the degree of danger associated with a detected abnormality.
[0269] "Means for automatically notifying external organizations of analysis results" refers to a means for automatically notifying relevant organizations such as the police based on the results of risk assessment.
[0270] "Means for inferring future behavior using past data" refers to a means for predicting future behavior using abnormal behavior and emotional data collected in the past.
[0271] The present invention combines a highly accurate security system using surveillance cameras with an emotion engine, and combines real-time video analysis, anomaly detection, risk assessment, automatic notification, behavioral inference, and emotion recognition. Specific embodiments of the present invention are described below.
[0272] System Configuration
[0273] 1. Surveillance camera (terminal)
[0274] The device is equipped with a high-resolution camera for wide-area monitoring and an AI analysis engine for initial analysis. For example, it captures video data from a surveillance camera and detects abnormal behavior in real time.
[0275] 2. Emotion Engine
[0276] The device's built-in emotion engine analyzes the user's facial expressions and tone of voice from video and audio data captured by the camera, and has the ability to recognize emotions. For example, if the user is feeling extremely frightened or has an aggressive expression, it will detect this as an abnormality.
[0277] 3. Central Server (Server)
[0278] The server receives the video data and emotion recognition data sent from each camera and performs advanced AI analysis. Notifications to external agencies such as the police are also sent from this server. For example, the server compares the data with past crime data to infer the next action and automatically notifies the police.
[0279] 4. User terminal (user)
[0280] Users, such as system administrators and police, have an interface for monitoring and managing security. For example, user devices can view surveillance footage in real time, supporting rapid response if an abnormality is detected.
[0281] Specific examples
[0282] For example, a store's surveillance cameras capture footage 24 hours a day. One day, the cameras capture a suspicious individual and detect abnormal behavior. The individual is shouting and holding a knife. The system analyzes this abnormal behavior and assesses the risk. As a result, it determines that the individual is likely to attempt robbery. Based on this analysis, the system automatically notifies the police.
[0283] Prompt Sentence Examples
[0284] The following prompts can be fed into the generative AI model to help understand what this system is doing:
[0285] Specific prompt examples:
[0286] A store's surveillance cameras capture footage 24 hours a day. One day, the cameras capture a suspicious individual and detect abnormal behavior. The individual is shouting and holding a knife. The system analyzes this abnormal behavior and assesses the risk. As a result, it determines that the individual is likely to attempt robbery. The system automatically notifies the police. Please explain this process in detail.
[0287] In this way, the system of the present invention overcomes the challenges of conventional security systems and significantly improves crime prevention effectiveness through highly accurate anomaly detection, rapid risk assessment, and effective notification. The introduction of an emotion engine enables even more accurate anomaly detection and risk assessment, enabling more appropriate responses.
[0288] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0289] Step 1:
[0290] Video data acquisition and initial analysis
[0291] The device receives video footage from surveillance cameras 24 / 7. As input, high-resolution video data is acquired from the surveillance cameras. This video data is sent in real time to the device's AI analysis engine for initial analysis. Specifically, the device captures video data and uses AI models to detect anomalies in human movement and audio. As output, potentially anomalous behavior and audio data are filtered.
[0292] Step 2:
[0293] Emotion Recognition and Anomaly Detection
[0294] The emotion engine built into the device receives the filtered video and audio data from the previous step as input. It analyzes this data and recognizes emotions based on the user's facial expressions and tone of voice. Specifically, the device uses facial expression recognition algorithms and voice analysis algorithms to identify abnormal emotions such as fear or anger. The output is the recognized emotion information and a judgment on whether it is abnormal.
[0295] Step 3:
[0296] Anomaly detection and risk assessment
[0297] The device receives the emotional information and abnormal behavior data acquired in step 2 as input and performs a risk assessment. Previously learned abnormal behavior models and emotional data are used for the risk assessment. Specifically, the device compares the behavior data and emotional data with the previous abnormal behavior models to assess the risk level. The output generates an assessment result such as high risk, low risk, or medium risk.
[0298] Step 4:
[0299] Sending data to the server
[0300] If the device evaluates the anomaly as high risk, it sends the video data and emotion data to the server. The input includes the risk assessment result, video data, and emotion data. Specifically, the device compresses these data and quickly transmits them to the server via the network. The output is the data received by the server.
[0301] Step 5:
[0302] Advanced AI analysis
[0303] The server receives data from the device as input and performs a more detailed risk assessment. The server compares it with past crime data and emotion data to infer the next course of action. Specifically, the server uses an advanced AI analysis engine to analyze the received data and infer the next course of action and the most likely scenario. The output is the inference result and a more detailed risk assessment result.
[0304] Step 6:
[0305] Notification to external agencies
[0306] If the server determines that the risk is extremely high, it automatically notifies external agencies such as the police. The inputs include risk assessment results, inference results, video data, and emotion data. Specifically, the server activates an automatic notification system, attaches relevant information, and sends it to the police. The output is a notification sent to an external agency. This notification allows for a rapid response.
[0307] (Application example 2)
[0308] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0309] Conventional security systems have focused on detecting anomalies in people's behavior and actions, but do not include the ability to recognize anomalies based on emotions or facial expressions. As a result, emotional risks may be overlooked when abnormal behavior occurs, and the accuracy of anomaly detection may be insufficient. Furthermore, it is often difficult to respond to detected abnormal behavior in real time. This makes it difficult to notify external agencies in a timely manner, making it difficult to take appropriate measures.
[0310] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0311] In this invention, the server includes a means for acquiring video, a means for analyzing the acquired video in real time to detect abnormalities, a means for recognizing emotions from the acquired video, a means for assessing the risk if the abnormality is detected, a means for automatically notifying an external organization of the analysis results and emotion recognition results if the risk is assessed as high, and a means for inferring the next action using past data. This enables anomaly detection based not only on abnormal behavior but also on emotions, enabling highly accurate risk assessment and rapid notification. Furthermore, by taking appropriate measures in real time, a safer environment can be provided.
[0312] Definition of Terms
[0313] "Footage" means visual data captured by surveillance cameras or other imaging devices.
[0314] "Real-time" means that video data is analyzed and processed as soon as it is acquired.
[0315] "Analysis" refers to the act of analyzing acquired video data using algorithms and AI models to detect specific information or anomalies.
[0316] "Abnormal" refers to behavior or situations that deviate from normal behavior or situations, including behavior or facial expressions that the system judges to be high risk.
[0317] "Emotions" refer to sensations and moods extracted from human facial expressions and tone of voice recognized by the system.
[0318] "Emotion recognition" is a technology that analyzes video and audio data captured by surveillance cameras to identify a person's emotional state.
[0319] "Risk assessment" is the process of determining how dangerous a situation is based on detected anomalies and emotions.
[0320] "External agencies" refers to the police, security companies, and other organizations or groups responsible for security.
[0321] "Notification" refers to the act of the system sending analysis results or risk assessment results to an external organization.
[0322] "Past data" refers to historical information such as previously acquired video data, emotion recognition results, and anomaly detection data.
[0323] "Behavioral inference" is a technology that allows a system to predict the next likely action or situation based on past data.
[0324] MODE FOR CARRYING OUT THE INVENTION
[0325] The present invention provides an advanced security system that combines a surveillance camera and an emotion engine. Specific examples thereof are described below.
[0326] 1. Hardware Configuration
[0327] The system includes a server, a terminal, and a user terminal.
[0328] The device is equipped with a high-resolution camera for wide-area monitoring and an AI analysis engine for initial analysis.
[0329] The server has the computing power to perform advanced AI analysis and the ability to notify external agencies such as the police.
[0330] The user terminal is provided with an interface that allows security managers, police, etc. to monitor and manage abnormalities.
[0331] 2. Software Configuration
[0332] The device comes with software to capture and analyze video data in real time, using OpenCV to capture video data and Keras to run emotion recognition models.
[0333] The server runs software that performs advanced analysis of the data it receives, which is then stored in a database and compared with previous data.
[0334] The user terminal is provided with a graphical user interface (GUI) for displaying the results of anomaly detection and risk assessment.
[0335] 3. Processing flow
[0336] The device monitors the video in real time 24 hours a day, and if an abnormality is detected, the emotion engine analyzes the user's emotions from the video data. For example, if a store's surveillance camera detects someone shouting and recognizes anger from their facial expression, it will assess the risk as an abnormality.
[0337] The server receives the abnormal information and emotion recognition results that are judged to be high risk, compares it with past data, and infers the next action. This then notifies external agencies of the next possible action. For example, if it infers that there is a high possibility of a robber breaking into a store, it will automatically notify the police.
[0338] 4. Specific Examples
[0339] The device detects a suspicious person in front of the store and confirms that the person is holding a knife by shouting.
[0340] The device's emotion engine recognizes anger from the person's facial expression and determines that the person is at high risk.
[0341] Several minutes of video and emotional data are sent to a server, which then performs a detailed analysis.
[0342] The server compares the data with past crime data and determines that the person is at high risk of breaking into a store.
[0343] The server automatically notifies the police of the risk assessment results and video data, allowing them to respond quickly to the scene.
[0344] 5. Examples of prompts
[0345] TXT
[0346] Create an AI program that analyzes video data from security cameras in real time and detects suspicious behavior and emotional anomalies. Consider the following points:
[0347] 1. A system for acquiring video data in real time
[0348] 2. Use a pre-trained emotion recognition model to perform emotion recognition.
[0349] 3. A system that notifies you in real time when an abnormality is detected
[0350] 4. Specific examples and explanations of the program
[0351] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0352] Program processing flow
[0353] Step 1:
[0354] The device acquires video data from the surveillance camera in real time.
[0355] Input: Visual data (footage) captured by surveillance cameras.
[0356] Output: Real-time video stream.
[0357] How it works: The camera module in the device captures video and sends the data to the analysis engine. Specifically, it uses libraries such as OpenCV to capture video.
[0358] Step 2:
[0359] The video data acquired by the device is analyzed in real time to detect abnormalities.
[0360] Input: Real-time video stream.
[0361] Output: Anomaly detection results (e.g., unusual behavior, loud audio, etc.).
[0362] How it works: The device's AI analysis engine analyzes the video data and compares it with a pre-trained abnormal behavior model to detect abnormalities. Specifically, it performs object recognition and movement analysis.
[0363] Step 3:
[0364] The device recognizes emotions using video data.
[0365] Input: Real-time video stream and detected anomalies.
[0366] Output: Emotion recognition result (e.g. anger, fear, etc.).
[0367] How it works: The device's emotion recognition engine analyzes human facial expressions from video data captured by the camera and recognizes emotions. Specifically, it runs an emotion recognition model using Keras.
[0368] Step 4:
[0369] The device assesses risk based on anomalies and emotion recognition results.
[0370] Input: Anomaly detection results and emotion recognition results.
[0371] Output: Risk assessment result (e.g. high risk, medium risk, low risk).
[0372] How it works: The device's risk assessment module evaluates how dangerous a situation is based on abnormal behavior and emotional data. For example, if anger and loud voices are detected at the same time, it will be judged as high risk.
[0373] Step 5:
[0374] If a server is assessed as high risk, the analysis results and emotion recognition results will be automatically notified to external agencies.
[0375] Input: Risk assessment results and analysis data, emotion recognition results.
[0376] Output: Notification data to external agencies (e.g. police).
[0377] How it works: The server generates a notification message containing detailed information about the anomaly and sends it to an external agency, such as the police. Specifically, this can be done by sending data via an HTTP request.
[0378] Step 6:
[0379] The server uses past data to infer the next action.
[0380] Input: Detailed anomaly data, historical database information.
[0381] Output: The inferred result of the next possible action.
[0382] Operation: The server's AI inference module analyzes past data, compares it with current anomalies, and predicts the next action. Based on this, it becomes possible to provide more appropriate countermeasures to external organizations. Specifically, it performs behavior analysis using a predictive algorithm.
[0383] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0384] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0385] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0386] [Second embodiment]
[0387] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0388] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0389] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0390] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0391] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0392] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0393] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0394] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0395] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0396] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0397] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0398] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0399] The present invention is a high-precision security system that uses surveillance cameras and combines real-time analysis of video, anomaly detection, risk assessment, automatic notification, behavior inference, etc. Specific embodiments of the present invention are described below.
[0400] System Configuration
[0401] 1. Surveillance camera (terminal)
[0402] The device is equipped with a high-resolution camera for wide-area monitoring and a built-in AI analysis engine for initial analysis.
[0403] 2. Central Server (Server)
[0404] The server receives the video data sent from each camera and performs advanced AI analysis. Notifications to external agencies such as the police are also sent from this server.
[0405] The server stores a large amount of past data, and the next action is inferred based on this data.
[0406] 3. User terminal (user)
[0407] Users, such as system administrators and police, have an interface for monitoring and managing security.
[0408] Program processing
[0409] Real-time video analysis
[0410] The device captures video 24 hours a day and analyzes it in real time using a built-in AI analysis engine. For example, a store's surveillance camera constantly captures human movement and audio, and if it detects abnormal behavior or loud voices, it automatically triggers an event.
[0411] Anomaly detection and risk assessment
[0412] When the device detects an abnormality, it immediately assesses the risk based on a previously trained abnormal behavior model. For example, if it detects a person with a knife in front of a store, it will assess the risk of robbery as high.
[0413] Sending data
[0414] If the anomaly is assessed as high risk, the video data is sent to a server, where it is compressed and quickly transmitted over the network.
[0415] Advanced AI analysis and notifications
[0416] The server analyzes the received data and performs a more detailed risk assessment. At the same time, it compares it with past crime data and infers the next action. For example, based on abnormal behavior data in front of a store, the server may infer that the next robber is likely to enter the store.
[0417] Notification to external agencies
[0418] If the risk is deemed to be extremely high, the server automatically notifies external agencies such as the police, and the notification includes the risk assessment results, video data, and inferences about the next course of action.
[0419] Specific examples
[0420] The device captured suspicious activity in front of the store, including a loud voice and a person holding a knife.
[0421] The device immediately detects an abnormality, assesses the risk, and if it determines that the risk is high, sends several minutes of video footage to the server.
[0422] The server analyzes the video and determines that there is a high risk of robbery. Based on past crime data, it infers that the next move is likely to be to break into a store.
[0423] The server automatically notifies the police and provides relevant information (risk assessment, footage, inference results).
[0424] Police can quickly rush to the scene and prevent incidents from occurring. During this time, the server provides real-time video footage of the scene to the police.
[0425] In this way, the system of the present invention overcomes the challenges posed by conventional security systems through highly accurate anomaly detection, rapid risk assessment, and effective notification, and can significantly improve crime prevention effectiveness.
[0426] The processing flow will be explained below.
[0427] Step 1:
[0428] The device constantly operates the surveillance camera and captures video in real time, and this video data is sent to the camera's internal initial analysis engine.
[0429] Step 2:
[0430] The device's built-in analysis engine analyzes the captured video data in real time, analyzing movement, sound, behavioral patterns, etc. to determine whether or not an abnormality is detected. For example, if a person is holding a knife, it will detect this as an abnormality.
[0431] Step 3:
[0432] When a device detects an anomaly, the anomaly triggers a risk assessment, which calculates the risk of abnormal behavior based on a pre-trained model.
[0433] Step 4:
[0434] If the device assesses the risk as high, it compresses the video data of the anomalous event and prepares it for transmission to the server, including the time, location, and video clip of the anomalous behavior.
[0435] Step 5:
[0436] The server receives the data sent from the terminal, then extracts the data and prepares it for reanalysis.
[0437] Step 6:
[0438] The server then begins detailed AI analysis. Based on the initial analysis results, it compares them with past criminal data and the behavioral patterns of former criminals. If similar patterns are detected, the risk assessment is further refined.
[0439] Step 7:
[0440] The server infers the next course of action based on the results of the risk assessment. For example, it calculates the probability that a person carrying a knife will turn into a robber based on past data.
[0441] Step 8:
[0442] If the server determines that the risk is high, it automatically notifies external agencies such as the police, along with the results of the inference. The notification includes the risk assessment results, video data, and an inference on the next course of action.
[0443] Step 9:
[0444] The user (police officer or security officer) receives notifications from the server and takes action as needed, checks real-time video footage, and prepares for a response to the scene.
[0445] Step 10:
[0446] The server stores the notification and subsequent response as a log, which can be used for future analysis and system improvement.
[0447] Step 11:
[0448] If the device detects a new anomaly, it restarts the process from step 1, enabling continuous monitoring and response.
[0449] Example 1
[0450] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0451] In recent years, there has been a growing emphasis on improving security, but conventional security systems still face serious challenges in anomaly detection and risk assessment. For example, the accuracy of anomaly detection is low, and risk assessment is not performed quickly, often resulting in delayed responses. Furthermore, behavioral inference using past data has not been realized, limiting the effectiveness of crime prevention. Furthermore, despite the overwhelming computational resources of cloud servers and data centers, these resources are often not fully utilized, preventing the system from achieving its full potential. To address these challenges, it is necessary to develop a highly functional security system that can perform anomaly detection, risk assessment, data analysis, and notification in real time.
[0452] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0453] In this invention, the server includes means for acquiring video, means for analyzing the acquired video in real time to detect abnormalities, means for assessing the risk when the abnormality is detected, means for transmitting compressed video data to a central processing unit when the risk is assessed to be high, means for the central processing unit to analyze the received data in detail and compare it with a past database to infer the next action, means for automatically notifying an external institution of the inference results and risk assessment results, and means for performing a series of processes from abnormality detection to risk assessment, analysis, and notification in real time. This enables highly accurate abnormality detection and risk assessment, rapid data analysis and notification, and further inference of the next action using past data.
[0454] A "means for acquiring video" is a combination of hardware and software that uses a monitoring device to collect video data within a specific range.
[0455] "Means for analyzing acquired video in real time to detect abnormalities" refers to means for instantly processing collected video data and using algorithms to recognize abnormalities based on pre-set criteria and patterns.
[0456] The "means for assessing risk" refers to a means including an assessment algorithm and a related database for determining the severity and urgency of an abnormality when the abnormality is detected.
[0457] "Means for transmitting compressed video data to a central processing unit" refers to a protocol and associated means for efficiently compressing collected video data and rapidly transmitting it to a central processing unit via a communications network when an abnormality is detected and assessed as high risk.
[0458] A "central processing unit" is a computer system that receives data from multiple monitoring devices via a network, performs advanced analysis, and stores and manages the results.
[0459] "Means of inferring next actions by conducting detailed analysis and comparing it with a past database" refers to an advanced analytical algorithm and matching system that analyzes received video data in more detail and compares and collates it with data accumulated in the past to predict future actions.
[0460] "Means for automatically notifying external organizations of inference results and risk assessment results" refers to a system and communication protocol that automatically notifies appropriate external organizations (e.g., law enforcement agencies) based on analysis results and risk assessment information.
[0461] "A means for carrying out a series of processes in real time, from detecting an anomaly to risk assessment, analysis, and notification" refers to a system and algorithm that executes all processes, from the moment an anomaly is detected, such as risk assessment, data analysis, and notification to external organizations, in real time without any time delay.
[0462] MODE FOR CARRYING OUT THE INVENTION
[0463] The present invention is a highly accurate security system that uses a monitoring device and combines real-time video analysis, anomaly detection, risk assessment, automatic notification, behavior inference, etc. Specific embodiments of the present invention are described below.
[0464] System Configuration
[0465] Monitoring device (terminal)
[0466] The device is equipped with a high-resolution monitoring device for wide-area monitoring. It also has a built-in AI analysis engine for initial analysis. The device captures video 24 hours a day and analyzes it in real time using the built-in AI analysis engine. For example, a store's surveillance camera constantly captures human movement and audio, and if it detects abnormal behavior or loud voices, it automatically triggers an event.
[0467] Central Processing Unit (Server)
[0468] The server receives the video data sent from each device and performs advanced AI analysis. A large amount of past data is stored on the server, and the next course of action is inferred based on this data. If the risk is determined to be extremely high, the server automatically notifies an external agency, such as a security agency. This notification includes the risk assessment results, video data, and the inferred next course of action.
[0469] User terminal (user)
[0470] Users, such as system administrators and law enforcement agencies, are provided with an interface for monitoring and managing security.
[0471] Specific Examples
[0472] 1. The device captures suspicious activity in front of the store, such as a loud voice or a person holding a knife.
[0473] 2. The device immediately detects the abnormality and assesses the risk. If it determines that the risk is high, it sends several minutes of video footage to the server.
[0474] 3. The server performs advanced analysis of the received video data and compares it with a database of past crimes. If it determines that there is a high risk of robbery, it infers that the next move is likely to be breaking into a store.
[0475] 4. The server automatically notifies security agencies of any detected anomalies and their risk assessment results. The notification includes video data of the abnormal behavior, the risk assessment results, and the inferred next action.
[0476] 5. The security agency (user) receives the notification and quickly rushes to the scene to respond. During this time, the server provides the security agency with real-time video footage of the scene.
[0477] In this way, the system of the present invention overcomes the challenges posed by conventional security systems through highly accurate anomaly detection, rapid risk assessment, and effective notification, and can significantly improve crime prevention effectiveness.
[0478] Example prompts for generative AI models
[0479] Please explain how an anomaly detection system works in real-time video analysis. Please also explain in detail what processes are performed from the perspectives of the device, server, and user. Please also provide specific examples.
[0480] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0481] Step 1:
[0482] The device captures video 24 hours a day.
[0483] Specific operation: The device uses a high-resolution camera to constantly capture video data of the monitored area. The video data collected by the camera is temporarily stored in the internal memory.
[0484] Input: Video of the monitored area
[0485] Output: High-resolution video data
[0486] Step 2:
[0487] The video data acquired by the device is analyzed in real time to detect abnormalities.
[0488] How it works: The AI analysis engine analyzes the movements, sounds, and types of objects in the video to detect abnormal patterns. For example, if abnormal movements or loud voices are detected, an alert will be automatically generated.
[0489] Input: High-resolution video data
[0490] Output: Anomaly detection results (e.g., abnormal behavior detected)
[0491] Step 3:
[0492] If the device detects an abnormality, it evaluates the risk.
[0493] How it works: The built-in model performs a risk assessment based on the type of anomaly detected and the surrounding circumstances. For example, if a person holding a knife is captured on video, the risk will be assessed as "high."
[0494] Input: Anomaly detection result
[0495] Output: Risk assessment result (e.g., high risk)
[0496] Step 4:
[0497] If the device is assessed as being at high risk, the video data is compressed and sent to a central processing unit (server).
[0498] Specific operation: Using a compression algorithm, the size of the video data is reduced and quickly sent to the server over the network.
[0499] Input: Risk assessment results, high-resolution video data
[0500] Output: Compressed video data
[0501] Step 5:
[0502] The server analyzes the received data in detail and compares it with a past database to infer the next action.
[0503] Specific operation: The server analyzes the received video data and searches a database for similar abnormal behavior. For example, if the data matches past data on robbery, the server infers that the next behavior is likely to be an attempted robbery.
[0504] Input: Compressed video data
[0505] Output: Behavioral inference result (e.g., probability of attempted robbery)
[0506] Step 6:
[0507] The server automatically notifies external agencies based on the inference results and risk assessment results.
[0508] Specific operation: The server's notification system automatically generates risk assessment results and behavioral inference results and notifies security agencies in real time.
[0509] Input: Risk assessment results, behavioral inference results
[0510] Output: Notification to external agencies (e.g. emergency contact with security agencies)
[0511] Step 7:
[0512] The user (security agency or system administrator) receives a notification, checks the situation on site in real time, and takes action.
[0513] Specific operation: The security agency receives the notification and decides to dispatch to the scene based on the inference results and risk assessment. At the same time, the server provides the security agency with real-time video footage of the scene to assist in the on-site response.
[0514] Input: Notification to external agency
[0515] Output: On-site response (e.g., police dispatch)
[0516] In this way, the system of the present invention can achieve highly accurate anomaly detection, rapid risk assessment, detailed data analysis, effective notification, and real-time on-site response support.
[0517] (Application example 1)
[0518] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0519] Conventional surveillance camera systems require manual monitoring, which requires a great deal of effort and cost. Furthermore, if abnormalities are not detected and risks are not assessed promptly, it becomes difficult to prevent serious incidents. In particular, there is a lack of support for security guards to respond quickly on-site. The purpose of this invention is to solve these issues and build an efficient security system by achieving highly accurate abnormality detection, risk assessment, and real-time notifications.
[0520] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0521] In this invention, the server includes a system using a wearable device worn by a security guard, which includes means for acquiring video, means for analyzing the acquired video in real time to detect abnormalities, means for assessing the risk when the abnormality is detected, means for automatically notifying an external institution of the analysis result when the risk is assessed to be high, means for inferring the next action using past data, and means for issuing visual and audio warnings when an abnormality is detected, thereby enabling the security guard to respond quickly and accurately on site.
[0522] "Means for acquiring images" refers to the function of collecting visual information using surveillance cameras, wearable devices, etc.
[0523] "Means for analyzing acquired video in real time and detecting abnormalities" refers to a function that instantly processes collected video data and detects specific abnormal behavior or abnormal conditions.
[0524] "Means for assessing the risk" refers to a function that instantly determines the danger and impact of a detected abnormality and evaluates the level of risk.
[0525] "Means for automatically notifying external organizations of analysis results" refers to a function that, if the detected risk is high, judges the information and notifies relevant organizations such as the police.
[0526] "Means of inferring next actions using past data" is a function that uses previously collected and analyzed data to predict the next possible action or incident based on the current situation.
[0527] "Visual and audio warning means" refers to a function that notifies security guards via a display or audio alert when an abnormality is detected.
[0528] A "wearable device" is a device worn by security guards and has the ability to obtain and notify information in real time.
[0529] This invention is a system that combines real-time video analysis, anomaly detection, risk assessment, automatic notification, and behavioral inference to realize a highly accurate security system. This system consists of the following three main components:
[0530] System Components
[0531] 1. Terminal
[0532] The device is equipped with a high-resolution camera that continuously captures video of the monitored area. The device also includes a built-in AI analytics engine that analyzes the captured video in real time. This analytics engine uses deep learning models to detect anomalies. For example, a surveillance camera can capture human movement and audio and detect loud voices or suspicious behavior.
[0533] 2. Server
[0534] The server receives the video data sent from the device and performs advanced AI analysis. The server stores and references large amounts of past data, and performs risk assessments of abnormal behavior and infers behavior. If a high risk is determined, the server automatically notifies the police and other relevant agencies along with the analysis results. This server-side analysis makes it possible to predict the next likely event based on past crime data.
[0535] 3. Users
[0536] Users are mainly system administrators, security guards, police, etc., who monitor and control the entire system through a management interface. In particular, the wearable devices (e.g., smart glasses) worn by security guards include a function to notify visual and audio warnings when an abnormality is detected, enabling prompt response on site.
[0537] Explanation of program processing
[0538] The server uses the following hardware and software:
[0539] Hardware: high-performance servers, surveillance cameras, wearable devices.
[0540] Software: OpenCV (image processing library), requests (HTTP request sending library), deep learning model.
[0541] The server first receives the video data sent from the device and analyzes it in real time using a deep learning model. If an abnormality is detected, it evaluates the risk of that abnormality and, if it is determined to be high risk, automatically notifies the relevant authorities of the video data. It also infers the next action based on past data and notifies the predicted results.
[0542] The device mainly acquires video in real time and performs initial analysis. If an abnormality is detected, the video data of the abnormal part is sent to the server.
[0543] For example, a security guard wearing smart glasses will receive an alert when an abnormality is detected, allowing the user to recognize the abnormality visually and audibly and respond quickly.
[0544] Examples and prompts
[0545] A concrete example would be:
[0546] "We would like you to develop a security monitoring application for smart glasses worn by security guards. The application must meet the following requirements:
[0547] 1. Images are acquired from the camera in real time and analyzed using an AI model for anomaly detection.
[0548] 2. If an abnormality is detected, an alert will be displayed on the smart glasses' display and an audio notification will be given through the built-in speaker.
[0549] 3. Compress the abnormal video frames and send them to the security server.
[0550] 4. The SDK of the smart glasses used is assumed to be SmartGlassesAPI.
[0551] Please include specific program code and a description of the libraries you use in your response."
[0552] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0553] Step 1:
[0554] The device captures video data in real time using a high-resolution camera. At this time, the device has a built-in AI analysis engine that performs initial analysis. The input is video data collected in real time, and the output is pre-processed video data for immediate anomaly detection.
[0555] Step 2:
[0556] The video data acquired by the device is analyzed in real time using an AI analysis engine. Anomalies are detected from the analyzed data and the risk of the detected anomalies is assessed. The input is preprocessed video data, and the output is an anomaly score and risk assessment results. The specific operation of this is to calculate an anomaly score for each frame using a deep learning model.
[0557] Step 3:
[0558] The device evaluates the anomaly score, and if it is judged to be high risk, it sends the corresponding video data to the server. At this time, the video data is compressed using a data compression algorithm. The input is the risk assessment result and the original video data, and the output is the compressed high-risk video data.
[0559] Step 4:
[0560] The server receives high-risk video data sent from the device and performs advanced AI analysis. This analysis uses a large amount of past data. The input is compressed high-risk video data, and the output is a detailed risk assessment and next action prediction.
[0561] Step 5:
[0562] The server automatically notifies external agencies such as the police based on the detailed risk assessment results and next action predictions. This notification includes the risk assessment results, analyzed video data, and inferred next action results. The input is the detailed risk assessment results and next action predictions, and the output is notification information to the relevant agencies.
[0563] Step 6:
[0564] The user, a security guard, receives an anomaly detection notification through the smart glasses. The notification is visual (display) and audio (built-in speaker). The input is the anomaly detection notification from the terminal, and the output is a real-time notification to the security guard. Specifically, a warning message is displayed on the smart glasses' display and an alarm sounds from the speaker.
[0565] Step 7:
[0566] The user, a security guard, rushes to the scene and responds quickly based on the detailed information provided by the server. The input is the alert from the smart glasses and additional data from the server, and the output is the response action at the scene, allowing the security guard to make quick and appropriate decisions.
[0567] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0568] The present invention combines an emotion engine with a highly accurate security system using surveillance cameras, and combines real-time video analysis, anomaly detection, risk assessment, automatic notification, behavioral inference, and emotion recognition. Specific embodiments of the present invention are described below.
[0569] System Configuration
[0570] 1. Surveillance camera (terminal)
[0571] The device is equipped with a high-resolution camera for wide-area monitoring and a built-in AI analysis engine for initial analysis.
[0572] 2. Emotion Engine
[0573] The device's built-in emotion engine analyzes the user's facial expressions and tone of voice from video and audio data captured by the camera, and has the ability to recognize emotions.
[0574] 3. Central Server (Server)
[0575] The server receives the video data and emotion recognition data sent from each camera and performs advanced AI analysis. It also notifies external agencies such as the police.
[0576] 4. User terminal (user)
[0577] Users, such as system administrators and police, have an interface for monitoring and managing security.
[0578] Program processing
[0579] Real-time video analysis
[0580] The device captures video 24 hours a day and analyzes it in real time using a built-in AI analysis engine. For example, a store's surveillance camera constantly captures human movement and audio, and if it detects abnormal behavior or loud voices, it automatically triggers an event.
[0581] Emotion recognition and additional anomaly detection
[0582] The emotion engine installed in the device analyzes the user's facial expressions and tone of voice from the captured video and audio data to recognize their emotions. For example, if the user is feeling extremely scared or has an aggressive expression, this will be detected as an abnormality.
[0583] Anomaly detection and risk assessment
[0584] When a device detects an anomaly, it immediately assesses the risk based on previously learned abnormal behavior models and emotional data. The risk assessment takes into account not only movement and voice, but also emotional data. For example, if a distorted facial expression or an angry tone of voice is detected, it is deemed to be a high risk.
[0585] Sending data
[0586] If an anomaly is assessed as high risk, the video and emotion data are sent to a server, where the video data is compressed and quickly transmitted over a network.
[0587] Advanced AI analysis and notifications
[0588] The server analyzes the received data and performs a more detailed risk assessment. At the same time, it compares it with past crime data and emotion data to infer the next action. For example, based on abnormal behavior data in front of a store, the server may infer that the next robber is likely to break into the store.
[0589] Notification to external agencies
[0590] If the risk is deemed to be extremely high, the server automatically notifies external agencies such as the police. This notification includes the risk assessment results, video data, emotional data, and inference results for the next action.
[0591] Specific examples
[0592] The device captured suspicious activity in front of the store, including a loud voice and a person holding a knife.
[0593] The device immediately detects abnormalities and detects emotions such as fear and anger from facial expressions, assesses the risk, determines it to be high risk, and sends several minutes of video and emotional data to a server.
[0594] The server analyzes the video and emotional data and determines that there is a very high risk of robbery. Based on past crime data and emotional data, it infers that the next action is likely to be breaking into a store.
[0595] The server automatically notifies the police and provides relevant information (risk assessment, video footage, emotional data, and inference results).
[0596] Police can quickly rush to the scene and prevent incidents from occurring. During this time, the server provides real-time video footage of the scene to the police.
[0597] In this way, the system of the present invention overcomes the challenges of conventional security systems and significantly improves crime prevention effectiveness through highly accurate anomaly detection, rapid risk assessment, and effective notification. The introduction of an emotion engine enables even more accurate anomaly detection and risk assessment, enabling more appropriate responses.
[0598] The processing flow will be explained below.
[0599] Step 1:
[0600] The device constantly operates the surveillance camera, capturing video and audio in real time, which is then sent to the camera's internal analysis engine.
[0601] Step 2:
[0602] The device's built-in analysis engine analyzes the captured video and audio data in real time, analyzing movement, audio, and behavioral patterns to determine whether or not an abnormality has been detected. For example, if a person is holding a knife, it will detect this as an abnormality.
[0603] Step 3:
[0604] The emotion engine installed in the device analyzes the user's facial expressions and tone of voice from video and audio data to recognize their emotions. If the user is feeling extremely scared or has an aggressive expression, this is detected as an abnormality.
[0605] Step 4:
[0606] The device reassesses the anomaly based on the analysis results of the emotion engine and calculates the risk. For example, if a fearful facial expression or an angry tone of voice is detected, the device will determine that the risk is high.
[0607] Step 5:
[0608] If the device assesses the risk as high, it compresses the video and emotion data of the anomalous event and prepares to send it to the server. The data to be sent includes the time, location, video clip, and emotion data of the anomalous behavior.
[0609] Step 6:
[0610] The server receives the data sent from the device, then extracts the data and prepares it for reanalysis.
[0611] Step 7:
[0612] The server then begins detailed AI analysis. Based on the initial analysis results, it compares the data with past criminal records and the behavioral patterns of former criminals. If similar patterns are detected, the risk assessment is further refined.
[0613] Step 8:
[0614] The server analyzes the data, including emotional data, and infers the next action. For example, it calculates the probability that a person with a knife will turn into a robber based on past data.
[0615] Step 9:
[0616] If the server determines that the risk is high, it will automatically notify external agencies such as the police, along with the inference results. The notification will include the risk assessment results, video data, emotional data, and inferences about the next action.
[0617] Step 10:
[0618] The user (police officer or security officer) receives notifications from the server and takes action as needed, checking real-time video and emotion data and preparing for a response at the scene.
[0619] Step 11:
[0620] The content of the notification from the server and the subsequent response are saved as a log, which can be used for future analysis and system improvement.
[0621] Step 12:
[0622] If the device continues to detect new anomalies, the process restarts from step 1, ensuring continuous monitoring and response.
[0623] Example 2
[0624] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0625] Conventional security systems focused on detecting anomalies in video and assessing risk, but the accuracy of anomaly detection using emotion recognition was not improved, and so they were often insufficient. Furthermore, even if an anomaly was detected, it was difficult to immediately notify external agencies appropriately, making it difficult to respond quickly. This limited the effectiveness of crime prevention, and led to a demand for more accurate and rapid anomaly detection and response.
[0626] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0627] In this invention, the server includes means for acquiring video, means for analyzing the acquired video in real time to detect abnormalities, means for recognizing emotions from the acquired video and audio data, means for assessing the risk when the abnormality is detected, means for automatically notifying an external institution of the analysis results when the risk is assessed to be high, and means for inferring the next action using past data. This enables highly accurate anomaly detection using emotion recognition, and enables effective crime prevention through risk assessment and prompt external notification.
[0628] The "means for acquiring video" refers to a means for collecting video data of the area to be monitored using a terminal such as a surveillance camera.
[0629] "Means for analyzing acquired video in real time to detect abnormalities" refers to means for instantly analyzing acquired video data and identifying suspicious or abnormal behavior.
[0630] The "means for recognizing emotions from acquired video and audio data" refers to a means for analyzing video and audio data and identifying emotions from the subject's facial expressions and tone of voice.
[0631] "Means for assessing the risk" refers to means for determining the degree of danger associated with a detected abnormality.
[0632] "Means for automatically notifying external organizations of analysis results" refers to a means for automatically notifying relevant organizations such as the police based on the results of risk assessment.
[0633] "Means for inferring future behavior using past data" refers to a means for predicting future behavior using abnormal behavior and emotional data collected in the past.
[0634] The present invention combines a highly accurate security system using surveillance cameras with an emotion engine, and combines real-time video analysis, anomaly detection, risk assessment, automatic notification, behavioral inference, and emotion recognition. Specific embodiments of the present invention are described below.
[0635] System Configuration
[0636] 1. Surveillance camera (terminal)
[0637] The device is equipped with a high-resolution camera for wide-area monitoring and an AI analysis engine for initial analysis. For example, it captures video data from a surveillance camera and detects abnormal behavior in real time.
[0638] 2. Emotion Engine
[0639] The device's built-in emotion engine analyzes the user's facial expressions and tone of voice from video and audio data captured by the camera, and has the ability to recognize emotions. For example, if the user is feeling extremely frightened or has an aggressive expression, it will detect this as an abnormality.
[0640] 3. Central Server (Server)
[0641] The server receives the video data and emotion recognition data sent from each camera and performs advanced AI analysis. Notifications to external agencies such as the police are also sent from this server. For example, the server compares the data with past crime data to infer the next action and automatically notifies the police.
[0642] 4. User terminal (user)
[0643] Users, such as system administrators and police, have an interface for monitoring and managing security. For example, user devices can view surveillance footage in real time, supporting rapid response if an abnormality is detected.
[0644] Specific examples
[0645] For example, a store's surveillance cameras capture footage 24 hours a day. One day, the cameras capture a suspicious individual and detect abnormal behavior. The individual is shouting and holding a knife. The system analyzes this abnormal behavior and assesses the risk. As a result, it determines that the individual is likely to attempt robbery. Based on this analysis, the system automatically notifies the police.
[0646] Prompt Sentence Examples
[0647] The following prompts can be fed into the generative AI model to help understand what this system is doing:
[0648] Specific prompt examples:
[0649] A store's surveillance cameras capture footage 24 hours a day. One day, the cameras capture a suspicious individual and detect abnormal behavior. The individual is shouting and holding a knife. The system analyzes this abnormal behavior and assesses the risk. As a result, it determines that the individual is likely to attempt robbery. The system automatically notifies the police. Please explain this process in detail.
[0650] In this way, the system of the present invention overcomes the challenges of conventional security systems and significantly improves crime prevention effectiveness through highly accurate anomaly detection, rapid risk assessment, and effective notification. The introduction of an emotion engine enables even more accurate anomaly detection and risk assessment, enabling more appropriate responses.
[0651] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0652] Step 1:
[0653] Video data acquisition and initial analysis
[0654] The device receives video footage from surveillance cameras 24 / 7. As input, high-resolution video data is acquired from the surveillance cameras. This video data is sent in real time to the device's AI analysis engine for initial analysis. Specifically, the device captures video data and uses AI models to detect anomalies in human movement and audio. As output, potentially anomalous behavior and audio data are filtered.
[0655] Step 2:
[0656] Emotion Recognition and Anomaly Detection
[0657] The emotion engine built into the device receives the filtered video and audio data from the previous step as input. It analyzes this data and recognizes emotions based on the user's facial expressions and tone of voice. Specifically, the device uses facial expression recognition algorithms and voice analysis algorithms to identify abnormal emotions such as fear or anger. The output is the recognized emotion information and a judgment on whether it is abnormal.
[0658] Step 3:
[0659] Anomaly detection and risk assessment
[0660] The device receives the emotional information and abnormal behavior data acquired in step 2 as input and performs a risk assessment. Previously learned abnormal behavior models and emotional data are used for the risk assessment. Specifically, the device compares the behavior data and emotional data with the previous abnormal behavior models to assess the risk level. The output generates an assessment result such as high risk, low risk, or medium risk.
[0661] Step 4:
[0662] Sending data to the server
[0663] If the device evaluates the anomaly as high risk, it sends the video data and emotion data to the server. The input includes the risk assessment result, video data, and emotion data. Specifically, the device compresses these data and quickly transmits them to the server via the network. The output is the data received by the server.
[0664] Step 5:
[0665] Advanced AI analysis
[0666] The server receives data from the device as input and performs a more detailed risk assessment. The server compares it with past crime data and emotion data to infer the next course of action. Specifically, the server uses an advanced AI analysis engine to analyze the received data and infer the next course of action and the most likely scenario. The output is the inference result and a more detailed risk assessment result.
[0667] Step 6:
[0668] Notification to external agencies
[0669] If the server determines that the risk is extremely high, it automatically notifies external agencies such as the police. The inputs include risk assessment results, inference results, video data, and emotion data. Specifically, the server activates an automatic notification system, attaches relevant information, and sends it to the police. The output is a notification sent to an external agency. This notification allows for a rapid response.
[0670] (Application example 2)
[0671] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0672] Conventional security systems have focused on detecting anomalies in people's behavior and actions, but do not include the ability to recognize anomalies based on emotions or facial expressions. As a result, emotional risks may be overlooked when abnormal behavior occurs, and the accuracy of anomaly detection may be insufficient. Furthermore, it is often difficult to respond to detected abnormal behavior in real time. This makes it difficult to notify external agencies in a timely manner, making it difficult to take appropriate measures.
[0673] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0674] In this invention, the server includes a means for acquiring video, a means for analyzing the acquired video in real time to detect abnormalities, a means for recognizing emotions from the acquired video, a means for assessing the risk if the abnormality is detected, a means for automatically notifying an external organization of the analysis results and emotion recognition results if the risk is assessed as high, and a means for inferring the next action using past data. This enables anomaly detection based not only on abnormal behavior but also on emotions, enabling highly accurate risk assessment and rapid notification. Furthermore, by taking appropriate measures in real time, a safer environment can be provided.
[0675] Definition of Terms
[0676] "Footage" means visual data captured by surveillance cameras or other imaging devices.
[0677] "Real-time" means that video data is analyzed and processed as soon as it is acquired.
[0678] "Analysis" refers to the act of analyzing acquired video data using algorithms and AI models to detect specific information or anomalies.
[0679] "Abnormal" refers to behavior or situations that deviate from normal behavior or situations, including behavior or facial expressions that the system judges to be high risk.
[0680] "Emotions" refer to sensations and moods extracted from human facial expressions and tone of voice recognized by the system.
[0681] "Emotion recognition" is a technology that analyzes video and audio data captured by surveillance cameras to identify a person's emotional state.
[0682] "Risk assessment" is the process of determining how dangerous a situation is based on detected anomalies and emotions.
[0683] "External agencies" refers to the police, security companies, and other organizations or groups responsible for security.
[0684] "Notification" refers to the act of the system sending analysis results or risk assessment results to an external organization.
[0685] "Past data" refers to historical information such as previously acquired video data, emotion recognition results, and anomaly detection data.
[0686] "Behavioral inference" is a technology that allows a system to predict the next likely action or situation based on past data.
[0687] MODE FOR CARRYING OUT THE INVENTION
[0688] The present invention provides an advanced security system that combines a surveillance camera and an emotion engine. Specific examples thereof are described below.
[0689] 1. Hardware Configuration
[0690] The system includes a server, a terminal, and a user terminal.
[0691] The device is equipped with a high-resolution camera for wide-area monitoring and an AI analysis engine for initial analysis.
[0692] The server has the computing power to perform advanced AI analysis and the ability to notify external agencies such as the police.
[0693] The user terminal is provided with an interface that allows security managers, police, etc. to monitor and manage abnormalities.
[0694] 2. Software Configuration
[0695] The device comes with software to capture and analyze video data in real time, using OpenCV to capture video data and Keras to run emotion recognition models.
[0696] The server runs software that performs advanced analysis of the data it receives, which is then stored in a database and compared with previous data.
[0697] The user terminal is provided with a graphical user interface (GUI) for displaying the results of anomaly detection and risk assessment.
[0698] 3. Processing flow
[0699] The device monitors the video in real time 24 hours a day, and if an abnormality is detected, the emotion engine analyzes the user's emotions from the video data. For example, if a store's surveillance camera detects someone shouting and recognizes anger from their facial expression, it will assess the risk as an abnormality.
[0700] The server receives the abnormal information and emotion recognition results that are judged to be high risk, compares it with past data, and infers the next action. This then notifies external agencies of the next possible action. For example, if it infers that there is a high possibility that a robber will break into a store, it will automatically notify the police.
[0701] 4. Specific Examples
[0702] The device detects a suspicious person in front of the store and confirms that the person is holding a knife by shouting.
[0703] The device's emotion engine recognizes anger from the person's facial expression and determines that the person is at high risk.
[0704] Several minutes of video and emotional data are sent to a server, which then performs a detailed analysis.
[0705] The server compares the data with past crime data and determines that the person is at high risk of breaking into a store.
[0706] The server automatically notifies the police of the risk assessment results and video data, allowing them to respond quickly to the scene.
[0707] 5. Examples of prompts
[0708] TXT
[0709] Create an AI program that analyzes video data from security cameras in real time and detects suspicious behavior and emotional anomalies. Consider the following points:
[0710] 1. A system for acquiring video data in real time
[0711] 2. Use a pre-trained emotion recognition model to perform emotion recognition.
[0712] 3. A system that notifies you in real time when an abnormality is detected
[0713] 4. Specific examples and explanations of the program
[0714] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0715] Program processing flow
[0716] Step 1:
[0717] The device acquires video data from the surveillance camera in real time.
[0718] Input: Visual data (footage) captured by surveillance cameras.
[0719] Output: Real-time video stream.
[0720] How it works: The camera module in the device captures video and sends the data to the analysis engine. Specifically, it uses libraries such as OpenCV to capture video.
[0721] Step 2:
[0722] The video data acquired by the device is analyzed in real time to detect abnormalities.
[0723] Input: Real-time video stream.
[0724] Output: Anomaly detection results (e.g., unusual behavior, loud audio, etc.).
[0725] How it works: The device's AI analysis engine analyzes the video data and compares it with a pre-trained abnormal behavior model to detect abnormalities. Specifically, it performs object recognition and movement analysis.
[0726] Step 3:
[0727] The device recognizes emotions using video data.
[0728] Input: Real-time video stream and detected anomalies.
[0729] Output: Emotion recognition result (e.g. anger, fear, etc.).
[0730] How it works: The device's emotion recognition engine analyzes human facial expressions from video data captured by the camera and recognizes emotions. Specifically, it runs an emotion recognition model using Keras.
[0731] Step 4:
[0732] The device assesses risk based on anomalies and emotion recognition results.
[0733] Input: Anomaly detection results and emotion recognition results.
[0734] Output: Risk assessment result (e.g. high risk, medium risk, low risk).
[0735] How it works: The device's risk assessment module evaluates how dangerous a situation is based on abnormal behavior and emotional data. For example, if anger and loud voices are detected at the same time, it will be judged as high risk.
[0736] Step 5:
[0737] If a server is assessed as high risk, the analysis results and emotion recognition results will be automatically notified to external agencies.
[0738] Input: Risk assessment results and analysis data, emotion recognition results.
[0739] Output: Notification data to external agencies (e.g. police).
[0740] How it works: The server generates a notification message containing detailed information about the anomaly and sends it to an external agency, such as the police. Specifically, this can be done by sending data via an HTTP request.
[0741] Step 6:
[0742] The server uses past data to infer the next action.
[0743] Input: Detailed anomaly data, historical database information.
[0744] Output: The inferred result of the next possible action.
[0745] Operation: The server's AI inference module analyzes past data, compares it with current anomalies, and predicts the next action. Based on this, it becomes possible to provide more appropriate countermeasures to external organizations. Specifically, it performs behavior analysis using a predictive algorithm.
[0746] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0747] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0748] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0749] [Third embodiment]
[0750] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0751] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0752] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0753] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0754] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0755] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0756] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0757] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0758] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0759] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0760] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0761] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0762] The present invention is a high-precision security system that uses surveillance cameras and combines real-time analysis of video, anomaly detection, risk assessment, automatic notification, behavior inference, etc. Specific embodiments of the present invention are described below.
[0763] System Configuration
[0764] 1. Surveillance camera (terminal)
[0765] The device is equipped with a high-resolution camera for wide-area monitoring and a built-in AI analysis engine for initial analysis.
[0766] 2. Central Server (Server)
[0767] The server receives the video data sent from each camera and performs advanced AI analysis. Notifications to external agencies such as the police are also sent from this server.
[0768] The server stores a large amount of past data, and the next action is inferred based on this data.
[0769] 3. User terminal (user)
[0770] Users, such as system administrators and police, have an interface for monitoring and managing security.
[0771] Program processing
[0772] Real-time video analysis
[0773] The device captures video 24 hours a day and analyzes it in real time using a built-in AI analysis engine. For example, a store's surveillance camera constantly captures human movement and audio, and if it detects abnormal behavior or loud voices, it automatically triggers an event.
[0774] Anomaly detection and risk assessment
[0775] When the device detects an abnormality, it immediately assesses the risk based on a previously trained abnormal behavior model. For example, if it detects a person with a knife in front of a store, it will assess the risk of robbery as high.
[0776] Sending data
[0777] If the anomaly is assessed as high risk, the video data is sent to a server, where it is compressed and quickly transmitted over the network.
[0778] Advanced AI analysis and notifications
[0779] The server analyzes the received data and performs a more detailed risk assessment. At the same time, it compares it with past crime data and infers the next action. For example, based on abnormal behavior data in front of a store, the server may infer that the next robber is likely to enter the store.
[0780] Notification to external agencies
[0781] If the risk is deemed to be extremely high, the server automatically notifies external agencies such as the police, and the notification includes the risk assessment results, video data, and inferences about the next course of action.
[0782] Specific examples
[0783] The device captured suspicious activity in front of the store, including a loud voice and a person holding a knife.
[0784] The device immediately detects an abnormality, assesses the risk, and if it determines that the risk is high, sends several minutes of video footage to the server.
[0785] The server analyzes the video and determines that there is a high risk of robbery. Based on past crime data, it infers that the next move is likely to be to break into a store.
[0786] The server automatically notifies the police and provides relevant information (risk assessment, footage, inference results).
[0787] Police can quickly rush to the scene and prevent incidents from occurring. During this time, the server provides real-time video footage of the scene to the police.
[0788] In this way, the system of the present invention overcomes the challenges posed by conventional security systems through highly accurate anomaly detection, rapid risk assessment, and effective notification, and can significantly improve crime prevention effectiveness.
[0789] The processing flow will be explained below.
[0790] Step 1:
[0791] The device constantly operates the surveillance camera and captures video in real time, and this video data is sent to the camera's internal initial analysis engine.
[0792] Step 2:
[0793] The device's built-in analysis engine analyzes the captured video data in real time, analyzing movement, sound, behavioral patterns, etc. to determine whether or not an abnormality is detected. For example, if a person is holding a knife, it will detect this as an abnormality.
[0794] Step 3:
[0795] When a device detects an anomaly, the anomaly triggers a risk assessment, which calculates the risk of abnormal behavior based on a pre-trained model.
[0796] Step 4:
[0797] If the device assesses the risk as high, it compresses the video data of the anomalous event and prepares it for transmission to the server, including the time, location, and video clip of the anomalous behavior.
[0798] Step 5:
[0799] The server receives the data sent from the terminal, then extracts the data and prepares it for reanalysis.
[0800] Step 6:
[0801] The server then begins detailed AI analysis. Based on the initial analysis results, it compares them with past criminal data and the behavioral patterns of former criminals. If similar patterns are detected, the risk assessment is further refined.
[0802] Step 7:
[0803] The server infers the next course of action based on the results of the risk assessment. For example, it calculates the probability that a person carrying a knife will turn into a robber based on past data.
[0804] Step 8:
[0805] If the server determines that the risk is high, it automatically notifies external agencies such as the police, along with the results of the inference. The notification includes the risk assessment results, video data, and an inference on the next course of action.
[0806] Step 9:
[0807] The user (police officer or security officer) receives notifications from the server and takes action as needed, checks real-time video footage, and prepares for a response to the scene.
[0808] Step 10:
[0809] The server stores the notification and subsequent response as a log, which can be used for future analysis and system improvement.
[0810] Step 11:
[0811] If the device detects a new anomaly, it restarts the process from step 1, enabling continuous monitoring and response.
[0812] Example 1
[0813] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0814] In recent years, there has been a growing emphasis on improving security, but conventional security systems still face serious challenges in anomaly detection and risk assessment. For example, the accuracy of anomaly detection is low, and risk assessment is not performed quickly, often resulting in delayed responses. Furthermore, behavioral inference using past data has not been realized, limiting the effectiveness of crime prevention. Furthermore, despite the overwhelming computational resources of cloud servers and data centers, these resources are often not fully utilized, preventing the system from achieving its full potential. To address these challenges, it is necessary to develop a highly functional security system that can perform anomaly detection, risk assessment, data analysis, and notification in real time.
[0815] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0816] In this invention, the server includes means for acquiring video, means for analyzing the acquired video in real time to detect abnormalities, means for assessing the risk when the abnormality is detected, means for transmitting compressed video data to a central processing unit when the risk is assessed to be high, means for the central processing unit to analyze the received data in detail and compare it with a past database to infer the next action, means for automatically notifying an external institution of the inference results and risk assessment results, and means for performing a series of processes from abnormality detection to risk assessment, analysis, and notification in real time. This enables highly accurate abnormality detection and risk assessment, rapid data analysis and notification, and further inference of the next action using past data.
[0817] A "means for acquiring video" is a combination of hardware and software that uses a monitoring device to collect video data within a specific range.
[0818] "Means for analyzing acquired video in real time to detect abnormalities" refers to means for instantly processing collected video data and using algorithms to recognize abnormalities based on pre-set criteria and patterns.
[0819] The "means for assessing risk" refers to a means including an assessment algorithm and a related database for determining the severity and urgency of an abnormality when the abnormality is detected.
[0820] "Means for transmitting compressed video data to a central processing unit" refers to a protocol and associated means for efficiently compressing collected video data and rapidly transmitting it to a central processing unit via a communications network when an abnormality is detected and assessed as high risk.
[0821] A "central processing unit" is a computer system that receives data from multiple monitoring devices via a network, performs advanced analysis, and stores and manages the results.
[0822] "Means of inferring next actions by conducting detailed analysis and comparing it with a past database" refers to an advanced analytical algorithm and matching system that analyzes received video data in more detail and compares and collates it with data accumulated in the past to predict future actions.
[0823] "Means for automatically notifying external organizations of inference results and risk assessment results" refers to a system and communication protocol that automatically notifies appropriate external organizations (e.g., law enforcement agencies) based on analysis results and risk assessment information.
[0824] "A means for carrying out a series of processes in real time, from detecting an anomaly to risk assessment, analysis, and notification" refers to a system and algorithm that executes all processes, from the moment an anomaly is detected, such as risk assessment, data analysis, and notification to external organizations, in real time without any time delay.
[0825] MODE FOR CARRYING OUT THE INVENTION
[0826] The present invention is a highly accurate security system that uses a monitoring device and combines real-time video analysis, anomaly detection, risk assessment, automatic notification, behavior inference, etc. Specific embodiments of the present invention are described below.
[0827] System Configuration
[0828] Monitoring device (terminal)
[0829] The device is equipped with a high-resolution monitoring device for wide-area monitoring. It also has a built-in AI analysis engine for initial analysis. The device captures video 24 hours a day and analyzes it in real time using the built-in AI analysis engine. For example, a store's surveillance camera constantly captures human movement and audio, and if it detects abnormal behavior or loud voices, it automatically triggers an event.
[0830] Central Processing Unit (Server)
[0831] The server receives the video data sent from each device and performs advanced AI analysis. A large amount of past data is stored on the server, and the next course of action is inferred based on this data. If the risk is determined to be extremely high, the server automatically notifies an external agency, such as a security agency. This notification includes the risk assessment results, video data, and the inferred next course of action.
[0832] User terminal (user)
[0833] Users, such as system administrators and law enforcement agencies, are provided with an interface for monitoring and managing security.
[0834] Specific Examples
[0835] 1. The device captures suspicious activity in front of the store, such as a loud voice or a person holding a knife.
[0836] 2. The device immediately detects the abnormality and assesses the risk. If it determines that the risk is high, it sends several minutes of video footage to the server.
[0837] 3. The server performs advanced analysis of the received video data and compares it with a database of past crimes. If it determines that there is a high risk of robbery, it infers that the next move is likely to be breaking into a store.
[0838] 4. The server automatically notifies security agencies of any detected anomalies and their risk assessment results. The notification includes video data of the abnormal behavior, the risk assessment results, and the inferred next action.
[0839] 5. The security agency (user) receives the notification and quickly rushes to the scene to respond. During this time, the server provides the security agency with real-time video footage of the scene.
[0840] In this way, the system of the present invention overcomes the challenges posed by conventional security systems through highly accurate anomaly detection, rapid risk assessment, and effective notification, and can significantly improve crime prevention effectiveness.
[0841] Example prompts for generative AI models
[0842] Please explain how an anomaly detection system works in real-time video analysis. Please also explain in detail what processes are performed from the perspectives of the device, server, and user. Please also provide specific examples.
[0843] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0844] Step 1:
[0845] The device captures video 24 hours a day.
[0846] Specific operation: The device uses a high-resolution camera to constantly capture video data of the monitored area. The video data collected by the camera is temporarily stored in the internal memory.
[0847] Input: Video of the monitored area
[0848] Output: High-resolution video data
[0849] Step 2:
[0850] The video data acquired by the device is analyzed in real time to detect abnormalities.
[0851] How it works: The AI analysis engine analyzes the movements, sounds, and types of objects in the video to detect abnormal patterns. For example, if abnormal movements or loud voices are detected, an alert will be automatically generated.
[0852] Input: High-resolution video data
[0853] Output: Anomaly detection results (e.g., abnormal behavior detected)
[0854] Step 3:
[0855] If the device detects an abnormality, it evaluates the risk.
[0856] How it works: The built-in model performs a risk assessment based on the type of anomaly detected and the surrounding circumstances. For example, if a person holding a knife is captured on video, the risk will be assessed as "high."
[0857] Input: Anomaly detection result
[0858] Output: Risk assessment result (e.g., high risk)
[0859] Step 4:
[0860] If the device is assessed as being at high risk, the video data is compressed and sent to a central processing unit (server).
[0861] Specific operation: Using a compression algorithm, the size of the video data is reduced and quickly sent to the server over the network.
[0862] Input: Risk assessment results, high-resolution video data
[0863] Output: Compressed video data
[0864] Step 5:
[0865] The server analyzes the received data in detail and compares it with a past database to infer the next action.
[0866] Specific operation: The server analyzes the received video data and searches a database for similar abnormal behavior. For example, if the data matches past data on robbery, the server infers that the next behavior is likely to be an attempted robbery.
[0867] Input: Compressed video data
[0868] Output: Behavioral inference result (e.g., probability of attempted robbery)
[0869] Step 6:
[0870] The server automatically notifies external agencies based on the inference results and risk assessment results.
[0871] Specific operation: The server's notification system automatically generates risk assessment results and behavioral inference results and notifies security agencies in real time.
[0872] Input: Risk assessment results, behavioral inference results
[0873] Output: Notification to external agencies (e.g. emergency contact with security agencies)
[0874] Step 7:
[0875] The user (security agency or system administrator) receives a notification, checks the situation on site in real time, and takes action.
[0876] Specific operation: The security agency receives the notification and decides to dispatch to the scene based on the inference results and risk assessment. At the same time, the server provides the security agency with real-time video footage of the scene to assist in the on-site response.
[0877] Input: Notification to external agency
[0878] Output: On-site response (e.g., police dispatch)
[0879] In this way, the system of the present invention can achieve highly accurate anomaly detection, rapid risk assessment, detailed data analysis, effective notification, and real-time on-site response support.
[0880] (Application example 1)
[0881] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0882] Conventional surveillance camera systems require manual monitoring, which requires a great deal of effort and cost. Furthermore, if abnormalities are not detected and risks are not assessed promptly, it becomes difficult to prevent serious incidents. In particular, there is a lack of support for security guards to respond quickly on-site. The purpose of this invention is to solve these issues and build an efficient security system by achieving highly accurate abnormality detection, risk assessment, and real-time notifications.
[0883] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0884] In this invention, the server includes a system using a wearable device worn by a security guard, which includes means for acquiring video, means for analyzing the acquired video in real time to detect abnormalities, means for assessing the risk when the abnormality is detected, means for automatically notifying an external institution of the analysis result when the risk is assessed to be high, means for inferring the next action using past data, and means for issuing visual and audio warnings when an abnormality is detected, thereby enabling the security guard to respond quickly and accurately on site.
[0885] "Means for acquiring images" refers to the function of collecting visual information using surveillance cameras, wearable devices, etc.
[0886] "Means for analyzing acquired video in real time and detecting abnormalities" refers to a function that instantly processes collected video data and detects specific abnormal behavior or abnormal conditions.
[0887] "Means for assessing the risk" refers to a function that instantly determines the danger and impact of a detected abnormality and evaluates the level of risk.
[0888] "Means for automatically notifying external organizations of analysis results" refers to a function that, if the detected risk is high, judges the information and notifies relevant organizations such as the police.
[0889] "Means of inferring next actions using past data" is a function that uses previously collected and analyzed data to predict the next possible action or incident based on the current situation.
[0890] "Visual and audio warning means" refers to a function that notifies security guards via a display or audio alert when an abnormality is detected.
[0891] A "wearable device" is a device worn by security guards and has the ability to obtain and notify information in real time.
[0892] This invention is a system that combines real-time video analysis, anomaly detection, risk assessment, automatic notification, and behavioral inference to realize a highly accurate security system. This system consists of the following three main components:
[0893] System Components
[0894] 1. Terminal
[0895] The device is equipped with a high-resolution camera that continuously captures video of the monitored area. The device also includes a built-in AI analytics engine that analyzes the captured video in real time. This analytics engine uses deep learning models to detect anomalies. For example, a surveillance camera can capture human movement and audio and detect loud voices or suspicious behavior.
[0896] 2. Server
[0897] The server receives the video data sent from the device and performs advanced AI analysis. The server stores and references large amounts of past data, and performs risk assessments of abnormal behavior and infers behavior. If a high risk is determined, the server automatically notifies the police and other relevant agencies along with the analysis results. This server-side analysis makes it possible to predict the next likely event based on past crime data.
[0898] 3. Users
[0899] Users are mainly system administrators, security guards, police, etc., who monitor and control the entire system through a management interface. In particular, the wearable devices (e.g., smart glasses) worn by security guards include a function to notify visual and audio warnings when an abnormality is detected, enabling prompt response on site.
[0900] Explanation of program processing
[0901] The server uses the following hardware and software:
[0902] Hardware: high-performance servers, surveillance cameras, wearable devices.
[0903] Software: OpenCV (image processing library), requests (HTTP request sending library), deep learning model.
[0904] The server first receives the video data sent from the device and analyzes it in real time using a deep learning model. If an abnormality is detected, it evaluates the risk of that abnormality and, if it is determined to be high risk, automatically notifies the relevant authorities of the video data. It also infers the next action based on past data and notifies the predicted results.
[0905] The device mainly acquires video in real time and performs initial analysis. If an abnormality is detected, the video data of the abnormal part is sent to the server.
[0906] For example, a security guard wearing smart glasses will receive an alert when an abnormality is detected, allowing the user to recognize the abnormality visually and audibly and respond quickly.
[0907] Examples and prompts
[0908] A concrete example would be:
[0909] "We would like you to develop a security monitoring application for smart glasses worn by security guards. The application must meet the following requirements:
[0910] 1. Images are acquired from the camera in real time and analyzed using an AI model for anomaly detection.
[0911] 2. If an abnormality is detected, an alert will be displayed on the smart glasses' display and an audio notification will be given through the built-in speaker.
[0912] 3. Compress the abnormal video frames and send them to the security server.
[0913] 4. The SDK of the smart glasses used is assumed to be SmartGlassesAPI.
[0914] Please include specific program code and a description of the libraries you use in your response."
[0915] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0916] Step 1:
[0917] The device captures video data in real time using a high-resolution camera. At this time, the device has a built-in AI analysis engine that performs initial analysis. The input is video data collected in real time, and the output is pre-processed video data for immediate anomaly detection.
[0918] Step 2:
[0919] The video data acquired by the device is analyzed in real time using an AI analysis engine. Anomalies are detected from the analyzed data and the risk of the detected anomalies is assessed. The input is preprocessed video data, and the output is an anomaly score and risk assessment results. The specific operation of this is to calculate an anomaly score for each frame using a deep learning model.
[0920] Step 3:
[0921] The device evaluates the anomaly score, and if it is judged to be high risk, it sends the corresponding video data to the server. At this time, the video data is compressed using a data compression algorithm. The input is the risk assessment result and the original video data, and the output is the compressed high-risk video data.
[0922] Step 4:
[0923] The server receives high-risk video data sent from the device and performs advanced AI analysis. This analysis uses a large amount of past data. The input is compressed high-risk video data, and the output is a detailed risk assessment and next action prediction.
[0924] Step 5:
[0925] The server automatically notifies external agencies such as the police based on the detailed risk assessment results and next action predictions. This notification includes the risk assessment results, analyzed video data, and inferred next action results. The input is the detailed risk assessment results and next action predictions, and the output is notification information to the relevant agencies.
[0926] Step 6:
[0927] The user, a security guard, receives an anomaly detection notification through the smart glasses. The notification is visual (display) and audio (built-in speaker). The input is the anomaly detection notification from the terminal, and the output is a real-time notification to the security guard. Specifically, a warning message is displayed on the smart glasses' display and an alarm sounds from the speaker.
[0928] Step 7:
[0929] The user, a security guard, rushes to the scene and responds quickly based on the detailed information provided by the server. The input is the alert from the smart glasses and additional data from the server, and the output is the response action at the scene. This allows the security guard to make quick and appropriate decisions.
[0930] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0931] The present invention combines an emotion engine with a highly accurate security system using surveillance cameras, and combines real-time video analysis, anomaly detection, risk assessment, automatic notification, behavioral inference, and emotion recognition. Specific embodiments of the present invention are described below.
[0932] System Configuration
[0933] 1. Surveillance camera (terminal)
[0934] The device is equipped with a high-resolution camera for wide-area monitoring and a built-in AI analysis engine for initial analysis.
[0935] 2. Emotion Engine
[0936] The device's built-in emotion engine analyzes the user's facial expressions and tone of voice from video and audio data captured by the camera, and has the ability to recognize emotions.
[0937] 3. Central Server (Server)
[0938] The server receives the video data and emotion recognition data sent from each camera and performs advanced AI analysis. It also notifies external agencies such as the police.
[0939] 4. User terminal (user)
[0940] Users, such as system administrators and police, have an interface for monitoring and managing security.
[0941] Program processing
[0942] Real-time video analysis
[0943] The device captures video 24 hours a day and analyzes it in real time using a built-in AI analysis engine. For example, a store's surveillance camera constantly captures human movement and audio, and if it detects abnormal behavior or loud voices, it automatically triggers an event.
[0944] Emotion recognition and additional anomaly detection
[0945] The emotion engine installed in the device analyzes the user's facial expressions and tone of voice from the captured video and audio data to recognize their emotions. For example, if the user is feeling extremely scared or has an aggressive expression, this will be detected as an abnormality.
[0946] Anomaly detection and risk assessment
[0947] When a device detects an anomaly, it immediately assesses the risk based on previously learned abnormal behavior models and emotional data. The risk assessment takes into account not only movement and voice, but also emotional data. For example, if a distorted facial expression or an angry tone of voice is detected, it is deemed to be a high risk.
[0948] Sending data
[0949] If an anomaly is assessed as high risk, the video and emotion data are sent to a server, where the video data is compressed and quickly transmitted over a network.
[0950] Advanced AI analysis and notifications
[0951] The server analyzes the received data and performs a more detailed risk assessment. At the same time, it compares it with past crime data and emotion data to infer the next action. For example, based on abnormal behavior data in front of a store, the server may infer that the next robber is likely to break into the store.
[0952] Notification to external agencies
[0953] If the risk is deemed to be extremely high, the server automatically notifies external agencies such as the police. This notification includes the risk assessment results, video data, emotional data, and inference results for the next action.
[0954] Specific examples
[0955] The device captured suspicious activity in front of the store, including a loud voice and a person holding a knife.
[0956] The device immediately detects abnormalities and detects emotions such as fear and anger from facial expressions, assesses the risk, determines it to be high risk, and sends several minutes of video and emotional data to a server.
[0957] The server analyzes the video and emotional data and determines that there is a very high risk of robbery. Based on past crime data and emotional data, it infers that the next action is likely to be breaking into a store.
[0958] The server automatically notifies the police and provides relevant information (risk assessment, video footage, emotional data, and inference results).
[0959] Police can quickly rush to the scene and prevent incidents from occurring. During this time, the server provides real-time video footage of the scene to the police.
[0960] In this way, the system of the present invention overcomes the challenges of conventional security systems and significantly improves crime prevention effectiveness through highly accurate anomaly detection, rapid risk assessment, and effective notification. The introduction of an emotion engine enables even more accurate anomaly detection and risk assessment, enabling more appropriate responses.
[0961] The processing flow will be explained below.
[0962] Step 1:
[0963] The device constantly operates the surveillance camera, capturing video and audio in real time, which is then sent to the camera's internal analysis engine.
[0964] Step 2:
[0965] The device's built-in analysis engine analyzes the captured video and audio data in real time, analyzing movement, audio, and behavioral patterns to determine whether or not an abnormality has been detected. For example, if a person is holding a knife, it will detect this as an abnormality.
[0966] Step 3:
[0967] The emotion engine installed in the device analyzes the user's facial expressions and tone of voice from video and audio data to recognize their emotions. If the user is feeling extremely scared or has an aggressive expression, this is detected as an abnormality.
[0968] Step 4:
[0969] The device reassesses the anomaly based on the analysis results of the emotion engine and calculates the risk. For example, if a fearful facial expression or an angry tone of voice is detected, the device will determine that the risk is high.
[0970] Step 5:
[0971] If the device assesses the risk as high, it compresses the video and emotion data of the anomalous event and prepares to send it to the server. The data to be sent includes the time, location, video clip, and emotion data of the anomalous behavior.
[0972] Step 6:
[0973] The server receives the data sent from the device, then extracts the data and prepares it for reanalysis.
[0974] Step 7:
[0975] The server then begins detailed AI analysis. Based on the initial analysis results, it compares the data with past criminal records and the behavioral patterns of former criminals. If similar patterns are detected, the risk assessment is further refined.
[0976] Step 8:
[0977] The server analyzes the data, including emotional data, and infers the next action. For example, it calculates the probability that a person with a knife will turn into a robber based on past data.
[0978] Step 9:
[0979] If the server determines that the risk is high, it will automatically notify external agencies such as the police, along with the inference results. The notification will include the risk assessment results, video data, emotional data, and inferences about the next action.
[0980] Step 10:
[0981] The user (police officer or security officer) receives notifications from the server and takes action as needed, checking real-time video and emotion data and preparing for a response at the scene.
[0982] Step 11:
[0983] The content of the notification from the server and the subsequent response are saved as a log, which can be used for future analysis and system improvement.
[0984] Step 12:
[0985] If the device continues to detect new anomalies, the process restarts from step 1, ensuring continuous monitoring and response.
[0986] Example 2
[0987] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0988] Conventional security systems focused on detecting anomalies in video and assessing risk, but the accuracy of anomaly detection using emotion recognition was not improved, and so they were often insufficient. Furthermore, even if an anomaly was detected, it was difficult to immediately notify external agencies appropriately, making it difficult to respond quickly. This limited the effectiveness of crime prevention, and led to a demand for more accurate and rapid anomaly detection and response.
[0989] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0990] In this invention, the server includes means for acquiring video, means for analyzing the acquired video in real time to detect abnormalities, means for recognizing emotions from the acquired video and audio data, means for assessing the risk when the abnormality is detected, means for automatically notifying an external institution of the analysis results when the risk is assessed to be high, and means for inferring the next action using past data. This enables highly accurate anomaly detection using emotion recognition, and enables effective crime prevention through risk assessment and prompt external notification.
[0991] The "means for acquiring video" refers to a means for collecting video data of the area to be monitored using a terminal such as a surveillance camera.
[0992] "Means for analyzing acquired video in real time to detect abnormalities" refers to means for instantly analyzing acquired video data and identifying suspicious or abnormal behavior.
[0993] The "means for recognizing emotions from acquired video and audio data" refers to a means for analyzing video and audio data and identifying emotions from the subject's facial expressions and tone of voice.
[0994] "Means for assessing the risk" refers to means for determining the degree of danger associated with a detected abnormality.
[0995] "Means for automatically notifying external organizations of analysis results" refers to a means for automatically notifying relevant organizations such as the police based on the results of risk assessment.
[0996] "Means for inferring future behavior using past data" refers to a means for predicting future behavior using abnormal behavior and emotional data collected in the past.
[0997] The present invention combines a highly accurate security system using surveillance cameras with an emotion engine, and combines real-time video analysis, anomaly detection, risk assessment, automatic notification, behavioral inference, and emotion recognition. Specific embodiments of the present invention are described below.
[0998] System Configuration
[0999] 1. Surveillance camera (terminal)
[1000] The device is equipped with a high-resolution camera for wide-area monitoring and an AI analysis engine for initial analysis. For example, it captures video data from a surveillance camera and detects abnormal behavior in real time.
[1001] 2. Emotion Engine
[1002] The device's built-in emotion engine analyzes the user's facial expressions and tone of voice from video and audio data captured by the camera, and has the ability to recognize emotions. For example, if the user is feeling extremely frightened or has an aggressive expression, it will detect this as an abnormality.
[1003] 3. Central Server (Server)
[1004] The server receives the video data and emotion recognition data sent from each camera and performs advanced AI analysis. Notifications to external agencies such as the police are also sent from this server. For example, the server compares the data with past crime data to infer the next action and automatically notifies the police.
[1005] 4. User terminal (user)
[1006] Users, such as system administrators and police, have an interface for monitoring and managing security. For example, user devices can view surveillance footage in real time, supporting rapid response if an abnormality is detected.
[1007] Specific examples
[1008] For example, a store's surveillance cameras capture footage 24 hours a day. One day, the cameras capture a suspicious individual and detect abnormal behavior. The individual is shouting and holding a knife. The system analyzes this abnormal behavior and assesses the risk. As a result, it determines that the individual is likely to attempt robbery. Based on this analysis, the system automatically notifies the police.
[1009] Prompt Sentence Examples
[1010] The following prompts can be fed into the generative AI model to help understand what this system is doing:
[1011] Specific prompt examples:
[1012] A store's surveillance cameras capture footage 24 hours a day. One day, the cameras capture a suspicious individual and detect abnormal behavior. The individual is shouting and holding a knife. The system analyzes this abnormal behavior and assesses the risk. As a result, it determines that the individual is likely to attempt robbery. The system automatically notifies the police. Please explain this process in detail.
[1013] In this way, the system of the present invention overcomes the challenges of conventional security systems and significantly improves crime prevention effectiveness through highly accurate anomaly detection, rapid risk assessment, and effective notification. The introduction of an emotion engine enables even more accurate anomaly detection and risk assessment, enabling more appropriate responses.
[1014] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1015] Step 1:
[1016] Video data acquisition and initial analysis
[1017] The device receives video footage from surveillance cameras 24 / 7. As input, high-resolution video data is acquired from the surveillance cameras. This video data is sent in real time to the device's AI analysis engine for initial analysis. Specifically, the device captures video data and uses AI models to detect anomalies in human movement and audio. As output, potentially anomalous behavior and audio data are filtered.
[1018] Step 2:
[1019] Emotion Recognition and Anomaly Detection
[1020] The emotion engine built into the device receives the filtered video and audio data from the previous step as input. It analyzes this data and recognizes emotions based on the user's facial expressions and tone of voice. Specifically, the device uses facial expression recognition algorithms and voice analysis algorithms to identify abnormal emotions such as fear or anger. The output is the recognized emotion information and a judgment on whether it is abnormal.
[1021] Step 3:
[1022] Anomaly detection and risk assessment
[1023] The device receives the emotional information and abnormal behavior data acquired in step 2 as input and performs a risk assessment. Previously learned abnormal behavior models and emotional data are used for the risk assessment. Specifically, the device compares the behavior data and emotional data with the previous abnormal behavior models to assess the risk level. The output generates an assessment result such as high risk, low risk, or medium risk.
[1024] Step 4:
[1025] Sending data to the server
[1026] If the device evaluates the anomaly as high risk, it sends the video data and emotion data to the server. The input includes the risk assessment result, video data, and emotion data. Specifically, the device compresses these data and quickly transmits them to the server via the network. The output is the data received by the server.
[1027] Step 5:
[1028] Advanced AI analysis
[1029] The server receives data from the device as input and performs a more detailed risk assessment. The server compares it with past crime data and emotion data to infer the next course of action. Specifically, the server uses an advanced AI analysis engine to analyze the received data and infer the next course of action and the most likely scenario. The output is the inference result and a more detailed risk assessment result.
[1030] Step 6:
[1031] Notification to external agencies
[1032] If the server determines that the risk is extremely high, it automatically notifies external agencies such as the police. The inputs include risk assessment results, inference results, video data, and emotion data. Specifically, the server activates an automatic notification system, attaches relevant information, and sends it to the police. The output is a notification sent to an external agency. This notification allows for a rapid response.
[1033] (Application example 2)
[1034] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1035] Conventional security systems have focused on detecting anomalies in people's behavior and actions, but do not include the ability to recognize anomalies based on emotions or facial expressions. As a result, emotional risks may be overlooked when abnormal behavior occurs, and the accuracy of anomaly detection may be insufficient. Furthermore, it is often difficult to respond to detected abnormal behavior in real time. This makes it difficult to notify external agencies in a timely manner, making it difficult to take appropriate measures.
[1036] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1037] In this invention, the server includes a means for acquiring video, a means for analyzing the acquired video in real time to detect abnormalities, a means for recognizing emotions from the acquired video, a means for assessing the risk if the abnormality is detected, a means for automatically notifying an external organization of the analysis results and emotion recognition results if the risk is assessed as high, and a means for inferring the next action using past data. This enables anomaly detection based not only on abnormal behavior but also on emotions, enabling highly accurate risk assessment and rapid notification. Furthermore, by taking appropriate measures in real time, a safer environment can be provided.
[1038] Definition of Terms
[1039] "Footage" means visual data captured by surveillance cameras or other imaging devices.
[1040] "Real-time" means that video data is analyzed and processed as soon as it is acquired.
[1041] "Analysis" refers to the act of analyzing acquired video data using algorithms and AI models to detect specific information or anomalies.
[1042] "Abnormal" refers to behavior or situations that deviate from normal behavior or situations, including behavior or facial expressions that the system judges to be high risk.
[1043] "Emotions" refer to sensations and moods extracted from human facial expressions and tone of voice recognized by the system.
[1044] "Emotion recognition" is a technology that analyzes video and audio data captured by surveillance cameras to identify a person's emotional state.
[1045] "Risk assessment" is the process of determining how dangerous a situation is based on detected anomalies and emotions.
[1046] "External agencies" refers to the police, security companies, and other organizations or groups responsible for security.
[1047] "Notification" refers to the act of the system sending analysis results or risk assessment results to an external organization.
[1048] "Past data" refers to historical information such as previously acquired video data, emotion recognition results, and anomaly detection data.
[1049] "Behavioral inference" is a technology that allows a system to predict the next likely action or situation based on past data.
[1050] MODE FOR CARRYING OUT THE INVENTION
[1051] The present invention provides an advanced security system that combines a surveillance camera and an emotion engine. Specific examples thereof are described below.
[1052] 1. Hardware Configuration
[1053] The system includes a server, a terminal, and a user terminal.
[1054] The device is equipped with a high-resolution camera for wide-area monitoring and an AI analysis engine for initial analysis.
[1055] The server has the computing power to perform advanced AI analysis and the ability to notify external agencies such as the police.
[1056] The user terminal is provided with an interface that allows security managers, police, etc. to monitor and manage abnormalities.
[1057] 2. Software Configuration
[1058] The device comes with software to capture and analyze video data in real time, using OpenCV to capture video data and Keras to run emotion recognition models.
[1059] The server runs software that performs advanced analysis of the data it receives, which is then stored in a database and compared with previous data.
[1060] The user terminal is provided with a graphical user interface (GUI) for displaying the results of anomaly detection and risk assessment.
[1061] 3. Processing flow
[1062] The device monitors the video in real time 24 hours a day, and if an abnormality is detected, the emotion engine analyzes the user's emotions from the video data. For example, if a store's surveillance camera detects someone shouting and recognizes anger from their facial expression, it will assess the risk as an abnormality.
[1063] The server receives the abnormal information and emotion recognition results that are judged to be high risk, compares it with past data, and infers the next action. This then notifies external agencies of the next possible action. For example, if it infers that there is a high possibility of a robber breaking into a store, it will automatically notify the police.
[1064] 4. Specific Examples
[1065] The device detects a suspicious person in front of the store and confirms that the person is holding a knife by shouting.
[1066] The device's emotion engine recognizes anger from the person's facial expression and determines that the person is at high risk.
[1067] Several minutes of video and emotional data are sent to a server, which then performs a detailed analysis.
[1068] The server compares the data with past crime data and determines that the person is at high risk of breaking into a store.
[1069] The server automatically notifies the police of the risk assessment results and video data, allowing them to respond quickly to the scene.
[1070] 5. Examples of prompts
[1071] TXT
[1072] Create an AI program that analyzes video data from security cameras in real time and detects suspicious behavior and emotional anomalies. Consider the following points:
[1073] 1. A system for acquiring video data in real time
[1074] 2. Use a pre-trained emotion recognition model to perform emotion recognition.
[1075] 3. A system that notifies you in real time when an abnormality is detected
[1076] 4. Specific examples and explanations of the program
[1077] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1078] Program processing flow
[1079] Step 1:
[1080] The device acquires video data from the surveillance camera in real time.
[1081] Input: Visual data (footage) captured by surveillance cameras.
[1082] Output: Real-time video stream.
[1083] How it works: The camera module in the device captures video and sends the data to the analysis engine. Specifically, it uses libraries such as OpenCV to capture video.
[1084] Step 2:
[1085] The video data acquired by the device is analyzed in real time to detect abnormalities.
[1086] Input: Real-time video stream.
[1087] Output: Anomaly detection results (e.g., unusual behavior, loud audio, etc.).
[1088] How it works: The device's AI analysis engine analyzes the video data and compares it with a pre-trained abnormal behavior model to detect abnormalities. Specifically, it performs object recognition and movement analysis.
[1089] Step 3:
[1090] The device recognizes emotions using video data.
[1091] Input: Real-time video stream and detected anomalies.
[1092] Output: Emotion recognition result (e.g. anger, fear, etc.).
[1093] How it works: The device's emotion recognition engine analyzes human facial expressions from video data captured by the camera and recognizes emotions. Specifically, it runs an emotion recognition model using Keras.
[1094] Step 4:
[1095] The device assesses risk based on anomalies and emotion recognition results.
[1096] Input: Anomaly detection results and emotion recognition results.
[1097] Output: Risk assessment result (e.g. high risk, medium risk, low risk).
[1098] How it works: The device's risk assessment module evaluates how dangerous a situation is based on abnormal behavior and emotional data. For example, if anger and loud voices are detected at the same time, it will be judged as high risk.
[1099] Step 5:
[1100] If a server is assessed as high risk, the analysis results and emotion recognition results will be automatically notified to external agencies.
[1101] Input: Risk assessment results and analysis data, emotion recognition results.
[1102] Output: Notification data to external agencies (e.g. police).
[1103] How it works: The server generates a notification message containing detailed information about the anomaly and sends it to an external agency, such as the police. Specifically, this can be done by sending data via an HTTP request.
[1104] Step 6:
[1105] The server uses past data to infer the next action.
[1106] Input: Detailed anomaly data, historical database information.
[1107] Output: The inferred result of the next possible action.
[1108] Operation: The server's AI inference module analyzes past data, compares it with current anomalies, and predicts the next action. Based on this, it becomes possible to provide more appropriate countermeasures to external organizations. Specifically, it performs behavior analysis using a predictive algorithm.
[1109] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1110] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1111] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1112] [Fourth embodiment]
[1113] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1114] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1115] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1116] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1117] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1118] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1119] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1120] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1121] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1122] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1123] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1124] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1125] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1126] The present invention is a high-precision security system that uses surveillance cameras and combines real-time analysis of video, anomaly detection, risk assessment, automatic notification, behavior inference, etc. Specific embodiments of the present invention are described below.
[1127] System Configuration
[1128] 1. Surveillance camera (terminal)
[1129] The device is equipped with a high-resolution camera for wide-area monitoring and a built-in AI analysis engine for initial analysis.
[1130] 2. Central Server (Server)
[1131] The server receives the video data sent from each camera and performs advanced AI analysis. Notifications to external agencies such as the police are also sent from this server.
[1132] The server stores a large amount of past data, and the next action is inferred based on this data.
[1133] 3. User terminal (user)
[1134] Users, such as system administrators and police, have an interface for monitoring and managing security.
[1135] Program processing
[1136] Real-time video analysis
[1137] The device captures video 24 hours a day and analyzes it in real time using a built-in AI analysis engine. For example, a store's surveillance camera constantly captures human movement and audio, and if it detects abnormal behavior or loud voices, it automatically triggers an event.
[1138] Anomaly detection and risk assessment
[1139] When the device detects an abnormality, it immediately assesses the risk based on a previously trained abnormal behavior model. For example, if it detects a person with a knife in front of a store, it will assess the risk of robbery as high.
[1140] Sending data
[1141] If the anomaly is assessed as high risk, the video data is sent to a server, where it is compressed and quickly transmitted over the network.
[1142] Advanced AI analysis and notifications
[1143] The server analyzes the received data and performs a more detailed risk assessment. At the same time, it compares it with past crime data and infers the next action. For example, based on abnormal behavior data in front of a store, the server may infer that the next robber is likely to enter the store.
[1144] Notification to external agencies
[1145] If the risk is deemed to be extremely high, the server automatically notifies external agencies such as the police, and the notification includes the risk assessment results, video data, and inferences about the next course of action.
[1146] Specific examples
[1147] The device captured suspicious activity in front of the store, including a loud voice and a person holding a knife.
[1148] The device immediately detects an abnormality, assesses the risk, and if it determines that the risk is high, sends several minutes of video footage to the server.
[1149] The server analyzes the video and determines that there is a high risk of robbery. Based on past crime data, it infers that the next move is likely to be to break into a store.
[1150] The server automatically notifies the police and provides relevant information (risk assessment, footage, inference results).
[1151] Police can quickly rush to the scene and prevent incidents from occurring. During this time, the server provides real-time video footage of the scene to the police.
[1152] In this way, the system of the present invention overcomes the challenges posed by conventional security systems through highly accurate anomaly detection, rapid risk assessment, and effective notification, and can significantly improve crime prevention effectiveness.
[1153] The processing flow will be explained below.
[1154] Step 1:
[1155] The device constantly operates the surveillance camera and captures video in real time, and this video data is sent to the camera's internal initial analysis engine.
[1156] Step 2:
[1157] The device's built-in analysis engine analyzes the captured video data in real time, analyzing movement, sound, behavioral patterns, etc. to determine whether or not an abnormality is detected. For example, if a person is holding a knife, it will detect this as an abnormality.
[1158] Step 3:
[1159] When a device detects an anomaly, the anomaly triggers a risk assessment, which calculates the risk of abnormal behavior based on a pre-trained model.
[1160] Step 4:
[1161] If the device assesses the risk as high, it compresses the video data of the anomalous event and prepares it for transmission to the server, including the time, location, and video clip of the anomalous behavior.
[1162] Step 5:
[1163] The server receives the data sent from the terminal, then extracts the data and prepares it for reanalysis.
[1164] Step 6:
[1165] The server then begins detailed AI analysis. Based on the initial analysis results, it compares them with past criminal data and the behavioral patterns of former criminals. If similar patterns are detected, the risk assessment is further refined.
[1166] Step 7:
[1167] The server infers the next course of action based on the results of the risk assessment. For example, it calculates the probability that a person carrying a knife will turn into a robber based on past data.
[1168] Step 8:
[1169] If the server determines that the risk is high, it automatically notifies external agencies such as the police, along with the results of the inference. The notification includes the risk assessment results, video data, and an inference on the next course of action.
[1170] Step 9:
[1171] The user (police officer or security officer) receives notifications from the server and takes action as needed, checks real-time video footage, and prepares for a response to the scene.
[1172] Step 10:
[1173] The server stores the notification and subsequent response as a log, which can be used for future analysis and system improvement.
[1174] Step 11:
[1175] If the device detects a new anomaly, it restarts the process from step 1, enabling continuous monitoring and response.
[1176] Example 1
[1177] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1178] In recent years, there has been a growing emphasis on improving security, but conventional security systems still face serious challenges in anomaly detection and risk assessment. For example, the accuracy of anomaly detection is low, and risk assessment is not performed quickly, often resulting in delayed responses. Furthermore, behavioral inference using past data has not been realized, limiting the effectiveness of crime prevention. Furthermore, despite the overwhelming computational resources of cloud servers and data centers, these resources are often not fully utilized, preventing the system from achieving its full potential. To address these challenges, it is necessary to develop a highly functional security system that can perform anomaly detection, risk assessment, data analysis, and notification in real time.
[1179] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1180] In this invention, the server includes means for acquiring video, means for analyzing the acquired video in real time to detect abnormalities, means for assessing the risk when the abnormality is detected, means for transmitting compressed video data to a central processing unit when the risk is assessed to be high, means for the central processing unit to analyze the received data in detail and compare it with a past database to infer the next action, means for automatically notifying an external institution of the inference results and risk assessment results, and means for performing a series of processes from abnormality detection to risk assessment, analysis, and notification in real time. This enables highly accurate abnormality detection and risk assessment, rapid data analysis and notification, and further inference of the next action using past data.
[1181] A "means for acquiring video" is a combination of hardware and software that uses a monitoring device to collect video data within a specific range.
[1182] "Means for analyzing acquired video in real time to detect abnormalities" refers to means for instantly processing collected video data and using algorithms to recognize abnormalities based on pre-set criteria and patterns.
[1183] The "means for assessing risk" refers to a means including an assessment algorithm and a related database for determining the severity and urgency of an abnormality when the abnormality is detected.
[1184] "Means for transmitting compressed video data to a central processing unit" refers to a protocol and associated means for efficiently compressing collected video data and rapidly transmitting it to a central processing unit via a communications network when an abnormality is detected and assessed as high risk.
[1185] A "central processing unit" is a computer system that receives data from multiple monitoring devices via a network, performs advanced analysis, and stores and manages the results.
[1186] "Means of inferring next actions by conducting detailed analysis and comparing it with a past database" refers to an advanced analytical algorithm and matching system that analyzes received video data in more detail and compares and collates it with data accumulated in the past to predict future actions.
[1187] "Means for automatically notifying external organizations of inference results and risk assessment results" refers to a system and communication protocol that automatically notifies appropriate external organizations (e.g., law enforcement agencies) based on analysis results and risk assessment information.
[1188] "A means for carrying out a series of processes in real time, from detecting an anomaly to risk assessment, analysis, and notification" refers to a system and algorithm that executes all processes, from the moment an anomaly is detected, such as risk assessment, data analysis, and notification to external organizations, in real time without any time delay.
[1189] MODE FOR CARRYING OUT THE INVENTION
[1190] The present invention is a highly accurate security system that uses a monitoring device and combines real-time video analysis, anomaly detection, risk assessment, automatic notification, behavior inference, etc. Specific embodiments of the present invention are described below.
[1191] System Configuration
[1192] Monitoring device (terminal)
[1193] The device is equipped with a high-resolution monitoring device for wide-area monitoring. It also has a built-in AI analysis engine for initial analysis. The device captures video 24 hours a day and analyzes it in real time using the built-in AI analysis engine. For example, a store's surveillance camera constantly captures human movement and audio, and if it detects abnormal behavior or loud voices, it automatically triggers an event.
[1194] Central Processing Unit (Server)
[1195] The server receives the video data sent from each device and performs advanced AI analysis. A large amount of past data is stored on the server, and the next course of action is inferred based on this data. If the risk is determined to be extremely high, the server automatically notifies an external agency, such as a security agency. This notification includes the risk assessment results, video data, and the inferred next course of action.
[1196] User terminal (user)
[1197] Users, such as system administrators and law enforcement agencies, are provided with an interface for monitoring and managing security.
[1198] Specific Examples
[1199] 1. The device captures suspicious activity in front of the store, such as a loud voice or a person holding a knife.
[1200] 2. The device immediately detects the abnormality and assesses the risk. If it determines that the risk is high, it sends several minutes of video footage to the server.
[1201] 3. The server performs advanced analysis of the received video data and compares it with a database of past crimes. If it determines that there is a high risk of robbery, it infers that the next move is likely to be breaking into a store.
[1202] 4. The server automatically notifies security agencies of any detected anomalies and their risk assessment results. The notification includes video data of the abnormal behavior, the risk assessment results, and the inferred next action.
[1203] 5. The security agency (user) receives the notification and quickly rushes to the scene to respond. During this time, the server provides the security agency with real-time video footage of the scene.
[1204] In this way, the system of the present invention overcomes the challenges posed by conventional security systems through highly accurate anomaly detection, rapid risk assessment, and effective notification, and can significantly improve crime prevention effectiveness.
[1205] Example prompts for generative AI models
[1206] "Please explain how an anomaly detection system works in real-time video analysis. Please also explain in detail what processes are performed from the perspectives of the device, server, and user. Please also provide specific examples."
[1207] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1208] Step 1:
[1209] The device captures video 24 hours a day.
[1210] Specific operation: The device uses a high-resolution camera to constantly capture video data of the monitored area. The video data collected by the camera is temporarily stored in the internal memory.
[1211] Input: Video of the monitored area
[1212] Output: High-resolution video data
[1213] Step 2:
[1214] The video data acquired by the device is analyzed in real time to detect abnormalities.
[1215] How it works: The AI analysis engine analyzes the movements, sounds, and types of objects in the video to detect abnormal patterns. For example, if abnormal movements or loud voices are detected, an alert will be automatically generated.
[1216] Input: High-resolution video data
[1217] Output: Anomaly detection results (e.g., abnormal behavior detected)
[1218] Step 3:
[1219] If the device detects an abnormality, it evaluates the risk.
[1220] How it works: The built-in model performs a risk assessment based on the type of anomaly detected and the surrounding circumstances. For example, if a person holding a knife is captured on video, the risk will be assessed as "high."
[1221] Input: Anomaly detection result
[1222] Output: Risk assessment result (e.g., high risk)
[1223] Step 4:
[1224] If the device is assessed as being at high risk, the video data is compressed and sent to a central processing unit (server).
[1225] Specific operation: Using a compression algorithm, the size of the video data is reduced and quickly sent to the server over the network.
[1226] Input: Risk assessment results, high-resolution video data
[1227] Output: Compressed video data
[1228] Step 5:
[1229] The server analyzes the received data in detail and compares it with a past database to infer the next action.
[1230] Specific operation: The server analyzes the received video data and searches a database for similar abnormal behavior. For example, if the data matches past data on robbery, the server infers that the next behavior is likely to be an attempted robbery.
[1231] Input: Compressed video data
[1232] Output: Behavioral inference result (e.g., probability of attempted robbery)
[1233] Step 6:
[1234] The server automatically notifies external agencies based on the inference results and risk assessment results.
[1235] Specific operation: The server's notification system automatically generates risk assessment results and behavioral inference results and notifies security agencies in real time.
[1236] Input: Risk assessment results, behavioral inference results
[1237] Output: Notification to external agencies (e.g. emergency contact with security agencies)
[1238] Step 7:
[1239] The user (security agency or system administrator) receives a notification, checks the situation on site in real time, and takes action.
[1240] Specific operation: The security agency receives the notification and decides to dispatch to the scene based on the inference results and risk assessment. At the same time, the server provides the security agency with real-time video footage of the scene to assist in the on-site response.
[1241] Input: Notification to external agency
[1242] Output: On-site response (e.g., police dispatch)
[1243] In this way, the system of the present invention can achieve highly accurate anomaly detection, rapid risk assessment, detailed data analysis, effective notification, and real-time on-site response support.
[1244] (Application example 1)
[1245] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1246] Conventional surveillance camera systems require manual monitoring, which requires a great deal of effort and cost. Furthermore, if abnormalities are not detected and risks are not assessed promptly, it becomes difficult to prevent serious incidents. In particular, there is a lack of support for security guards to respond quickly on-site. The purpose of this invention is to solve these issues and build an efficient security system by achieving highly accurate abnormality detection, risk assessment, and real-time notifications.
[1247] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1248] In this invention, the server includes a system using a wearable device worn by a security guard, which includes means for acquiring video, means for analyzing the acquired video in real time to detect abnormalities, means for assessing the risk when the abnormality is detected, means for automatically notifying an external institution of the analysis result when the risk is assessed to be high, means for inferring the next action using past data, and means for issuing visual and audio warnings when an abnormality is detected, thereby enabling the security guard to respond quickly and accurately on site.
[1249] "Means for acquiring images" refers to the function of collecting visual information using surveillance cameras, wearable devices, etc.
[1250] "Means for analyzing acquired video in real time and detecting abnormalities" refers to a function that instantly processes collected video data and detects specific abnormal behavior or abnormal conditions.
[1251] "Means for assessing the risk" refers to a function that instantly determines the danger and impact of a detected abnormality and evaluates the level of risk.
[1252] "Means for automatically notifying external organizations of analysis results" refers to a function that, if the detected risk is high, judges the information and notifies relevant organizations such as the police.
[1253] "Means of inferring next actions using past data" is a function that uses previously collected and analyzed data to predict the next possible action or incident based on the current situation.
[1254] "Visual and audio warning means" refers to a function that notifies security guards via a display or audio alert when an abnormality is detected.
[1255] A "wearable device" is a device worn by security guards and has the ability to obtain and notify information in real time.
[1256] This invention is a system that combines real-time video analysis, anomaly detection, risk assessment, automatic notification, and behavioral inference to realize a highly accurate security system. This system consists of the following three main components:
[1257] System Components
[1258] 1. Terminal
[1259] The device is equipped with a high-resolution camera that continuously captures video of the monitored area. The device also includes a built-in AI analytics engine that analyzes the captured video in real time. This analytics engine uses deep learning models to detect anomalies. For example, a surveillance camera can capture human movement and audio and detect loud voices or suspicious behavior.
[1260] 2. Server
[1261] The server receives the video data sent from the device and performs advanced AI analysis. The server stores and references large amounts of past data, and performs risk assessments of abnormal behavior and infers behavior. If a high risk is determined, the server automatically notifies the police and other relevant agencies along with the analysis results. This server-side analysis makes it possible to predict the next likely event based on past crime data.
[1262] 3. Users
[1263] Users are mainly system administrators, security guards, police, etc., who monitor and control the entire system through a management interface. In particular, the wearable devices (e.g., smart glasses) worn by security guards include a function to notify visual and audio warnings when an abnormality is detected, enabling prompt response on site.
[1264] Explanation of program processing
[1265] The server uses the following hardware and software:
[1266] Hardware: high-performance servers, surveillance cameras, wearable devices.
[1267] Software: OpenCV (image processing library), requests (HTTP request sending library), deep learning model.
[1268] The server first receives the video data sent from the device and analyzes it in real time using a deep learning model. If an abnormality is detected, it evaluates the risk of that abnormality and, if it is determined to be high risk, automatically notifies the relevant authorities of the video data. It also infers the next action based on past data and notifies the predicted results.
[1269] The device mainly acquires video in real time and performs initial analysis. If an abnormality is detected, the video data of the abnormal part is sent to the server.
[1270] For example, a security guard wearing smart glasses will receive an alert when an abnormality is detected, allowing the user to recognize the abnormality visually and audibly and respond quickly.
[1271] Examples and prompts
[1272] A concrete example would be:
[1273] "We would like you to develop a security monitoring application for smart glasses worn by security guards. The application must meet the following requirements:
[1274] 1. Images are acquired from the camera in real time and analyzed using an AI model for anomaly detection.
[1275] 2. If an abnormality is detected, an alert will be displayed on the smart glasses' display and an audio notification will be given through the built-in speaker.
[1276] 3. Compress the abnormal video frames and send them to the security server.
[1277] 4. The SDK of the smart glasses used is assumed to be SmartGlassesAPI.
[1278] Please include specific program code and a description of the libraries you use in your response."
[1279] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1280] Step 1:
[1281] The device captures video data in real time using a high-resolution camera. At this time, the device has a built-in AI analysis engine that performs initial analysis. The input is video data collected in real time, and the output is pre-processed video data for immediate anomaly detection.
[1282] Step 2:
[1283] The video data acquired by the device is analyzed in real time using an AI analysis engine. Anomalies are detected from the analyzed data and the risk of the detected anomalies is assessed. The input is preprocessed video data, and the output is an anomaly score and risk assessment results. The specific operation of this is to calculate an anomaly score for each frame using a deep learning model.
[1284] Step 3:
[1285] The device evaluates the anomaly score, and if it is judged to be high risk, it sends the corresponding video data to the server. At this time, the video data is compressed using a data compression algorithm. The input is the risk assessment result and the original video data, and the output is the compressed high-risk video data.
[1286] Step 4:
[1287] The server receives high-risk video data sent from the device and performs advanced AI analysis. This analysis uses a large amount of past data. The input is compressed high-risk video data, and the output is a detailed risk assessment and next action prediction.
[1288] Step 5:
[1289] The server automatically notifies external agencies such as the police based on the detailed risk assessment results and next action predictions. This notification includes the risk assessment results, analyzed video data, and inferred next action results. The input is the detailed risk assessment results and next action predictions, and the output is notification information to the relevant agencies.
[1290] Step 6:
[1291] The user, a security guard, receives an anomaly detection notification through the smart glasses. The notification is visual (display) and audio (built-in speaker). The input is the anomaly detection notification from the terminal, and the output is a real-time notification to the security guard. Specifically, a warning message is displayed on the smart glasses' display and an alarm sounds from the speaker.
[1292] Step 7:
[1293] The user, a security guard, rushes to the scene and responds quickly based on the detailed information provided by the server. The input is the alert from the smart glasses and additional data from the server, and the output is the response action at the scene. This allows the security guard to make quick and appropriate decisions.
[1294] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1295] The present invention combines an emotion engine with a highly accurate security system using surveillance cameras, and combines real-time video analysis, anomaly detection, risk assessment, automatic notification, behavioral inference, and emotion recognition. Specific embodiments of the present invention are described below.
[1296] System Configuration
[1297] 1. Surveillance camera (terminal)
[1298] The device is equipped with a high-resolution camera for wide-area monitoring and a built-in AI analysis engine for initial analysis.
[1299] 2. Emotion Engine
[1300] The device's built-in emotion engine analyzes the user's facial expressions and tone of voice from video and audio data captured by the camera, and has the ability to recognize emotions.
[1301] 3. Central Server (Server)
[1302] The server receives the video data and emotion recognition data sent from each camera and performs advanced AI analysis. It also notifies external agencies such as the police.
[1303] 4. User terminal (user)
[1304] Users, such as system administrators and police, have an interface for monitoring and managing security.
[1305] Program processing
[1306] Real-time video analysis
[1307] The device captures video 24 hours a day and analyzes it in real time using a built-in AI analysis engine. For example, a store's surveillance camera constantly captures human movement and audio, and if it detects abnormal behavior or loud voices, it automatically triggers an event.
[1308] Emotion recognition and additional anomaly detection
[1309] The emotion engine installed in the device analyzes the user's facial expressions and tone of voice from the captured video and audio data to recognize their emotions. For example, if the user is feeling extremely scared or has an aggressive expression, this will be detected as an abnormality.
[1310] Anomaly detection and risk assessment
[1311] When a device detects an anomaly, it immediately assesses the risk based on previously learned abnormal behavior models and emotional data. The risk assessment takes into account not only movement and voice, but also emotional data. For example, if a distorted facial expression or an angry tone of voice is detected, it is deemed to be a high risk.
[1312] Sending data
[1313] If an anomaly is assessed as high risk, the video and emotion data are sent to a server, where the video data is compressed and quickly transmitted over a network.
[1314] Advanced AI analysis and notifications
[1315] The server analyzes the received data and performs a more detailed risk assessment. At the same time, it compares it with past crime data and emotion data to infer the next action. For example, based on abnormal behavior data in front of a store, the server may infer that the next robber is likely to break into the store.
[1316] Notification to external agencies
[1317] If the risk is deemed to be extremely high, the server automatically notifies external agencies such as the police. This notification includes the risk assessment results, video data, emotional data, and inference results for the next action.
[1318] Specific examples
[1319] The device captured suspicious activity in front of the store, including a loud voice and a person holding a knife.
[1320] The device immediately detects abnormalities and detects emotions such as fear and anger from facial expressions, assesses the risk, determines it to be high risk, and sends several minutes of video and emotional data to a server.
[1321] The server analyzes the video and emotional data and determines that there is a very high risk of robbery. Based on past crime data and emotional data, it infers that the next action is likely to be breaking into a store.
[1322] The server automatically notifies the police and provides relevant information (risk assessment, video footage, emotional data, and inference results).
[1323] Police can quickly rush to the scene and prevent incidents from occurring. During this time, the server provides real-time video footage of the scene to the police.
[1324] In this way, the system of the present invention overcomes the challenges of conventional security systems and significantly improves crime prevention effectiveness through highly accurate anomaly detection, rapid risk assessment, and effective notification. The introduction of an emotion engine enables even more accurate anomaly detection and risk assessment, enabling more appropriate responses.
[1325] The processing flow will be explained below.
[1326] Step 1:
[1327] The device constantly operates the surveillance camera, capturing video and audio in real time, which is then sent to the camera's internal analysis engine.
[1328] Step 2:
[1329] The device's built-in analysis engine analyzes the captured video and audio data in real time, analyzing movement, audio, and behavioral patterns to determine whether or not an abnormality has been detected. For example, if a person is holding a knife, it will detect this as an abnormality.
[1330] Step 3:
[1331] The emotion engine installed in the device analyzes the user's facial expressions and tone of voice from video and audio data to recognize their emotions. If the user is feeling extremely scared or has an aggressive expression, this is detected as an abnormality.
[1332] Step 4:
[1333] The device reassesses the anomaly based on the analysis results of the emotion engine and calculates the risk. For example, if a fearful facial expression or an angry tone of voice is detected, the device will determine that the risk is high.
[1334] Step 5:
[1335] If the device assesses the risk as high, it compresses the video and emotion data of the anomalous event and prepares to send it to the server. The data to be sent includes the time, location, video clip, and emotion data of the anomalous behavior.
[1336] Step 6:
[1337] The server receives the data sent from the device, then extracts the data and prepares it for reanalysis.
[1338] Step 7:
[1339] The server then begins detailed AI analysis. Based on the initial analysis results, it compares the data with past criminal records and the behavioral patterns of former criminals. If similar patterns are detected, the risk assessment is further refined.
[1340] Step 8:
[1341] The server analyzes the data, including emotional data, and infers the next action. For example, it calculates the probability that a person with a knife will turn into a robber based on past data.
[1342] Step 9:
[1343] If the server determines that the risk is high, it will automatically notify external agencies such as the police, along with the inference results. The notification will include the risk assessment results, video data, emotional data, and inferences about the next action.
[1344] Step 10:
[1345] The user (police officer or security officer) receives notifications from the server and takes action as needed, checking real-time video and emotion data and preparing for a response at the scene.
[1346] Step 11:
[1347] The content of the notification from the server and the subsequent response are saved as a log, which can be used for future analysis and system improvement.
[1348] Step 12:
[1349] If the device continues to detect new anomalies, the process restarts from step 1, ensuring continuous monitoring and response.
[1350] Example 2
[1351] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1352] Conventional security systems focused on detecting anomalies in video and assessing risk, but the accuracy of anomaly detection using emotion recognition was not improved, and so they were often insufficient. Furthermore, even if an anomaly was detected, it was difficult to immediately notify external agencies appropriately, making it difficult to respond quickly. This limited the effectiveness of crime prevention, and led to a demand for more accurate and rapid anomaly detection and response.
[1353] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1354] In this invention, the server includes means for acquiring video, means for analyzing the acquired video in real time to detect abnormalities, means for recognizing emotions from the acquired video and audio data, means for assessing the risk when the abnormality is detected, means for automatically notifying an external institution of the analysis results when the risk is assessed to be high, and means for inferring the next action using past data. This enables highly accurate anomaly detection using emotion recognition, and enables effective crime prevention through risk assessment and prompt external notification.
[1355] The "means for acquiring video" refers to a means for collecting video data of the area to be monitored using a terminal such as a surveillance camera.
[1356] "Means for analyzing acquired video in real time to detect abnormalities" refers to means for instantly analyzing acquired video data and identifying suspicious or abnormal behavior.
[1357] The "means for recognizing emotions from acquired video and audio data" refers to a means for analyzing video and audio data and identifying emotions from the subject's facial expressions and tone of voice.
[1358] "Means for assessing the risk" refers to means for determining the degree of danger associated with a detected abnormality.
[1359] "Means for automatically notifying external organizations of analysis results" refers to a means for automatically notifying relevant organizations such as the police based on the results of risk assessment.
[1360] "Means for inferring future behavior using past data" refers to a means for predicting future behavior using abnormal behavior and emotional data collected in the past.
[1361] The present invention combines a highly accurate security system using surveillance cameras with an emotion engine, and combines real-time video analysis, anomaly detection, risk assessment, automatic notification, behavioral inference, and emotion recognition. Specific embodiments of the present invention are described below.
[1362] System Configuration
[1363] 1. Surveillance camera (terminal)
[1364] The device is equipped with a high-resolution camera for wide-area monitoring and an AI analysis engine for initial analysis. For example, it captures video data from a surveillance camera and detects abnormal behavior in real time.
[1365] 2. Emotion Engine
[1366] The device's built-in emotion engine analyzes the user's facial expressions and tone of voice from video and audio data captured by the camera, and has the ability to recognize emotions. For example, if the user is feeling extremely scared or has an aggressive expression, it will detect this as an abnormality.
[1367] 3. Central Server (Server)
[1368] The server receives the video data and emotion recognition data sent from each camera and performs advanced AI analysis. Notifications to external agencies such as the police are also sent from this server. For example, the server compares the data with past crime data to infer the next action and automatically notifies the police.
[1369] 4. User terminal (user)
[1370] Users, such as system administrators and police, have an interface for monitoring and managing security. For example, user devices can view surveillance footage in real time, supporting rapid response if an abnormality is detected.
[1371] Specific examples
[1372] For example, a store's surveillance cameras capture footage 24 hours a day. One day, the cameras capture a suspicious individual and detect abnormal behavior. The individual is shouting and holding a knife. The system analyzes this abnormal behavior and assesses the risk. As a result, it determines that the individual is likely to attempt robbery. Based on this analysis, the system automatically notifies the police.
[1373] Prompt Sentence Examples
[1374] The following prompts can be fed into the generative AI model to help understand what this system is doing:
[1375] Specific prompt examples:
[1376] A store's surveillance cameras capture footage 24 hours a day. One day, the cameras capture a suspicious individual and detect abnormal behavior. The individual is shouting and holding a knife. The system analyzes this abnormal behavior and assesses the risk. As a result, it determines that the individual is likely to attempt robbery. The system automatically notifies the police. Please explain this process in detail.
[1377] In this way, the system of the present invention overcomes the challenges of conventional security systems and significantly improves crime prevention effectiveness through highly accurate anomaly detection, rapid risk assessment, and effective notification. The introduction of an emotion engine enables even more accurate anomaly detection and risk assessment, enabling more appropriate responses.
[1378] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1379] Step 1:
[1380] Video data acquisition and initial analysis
[1381] The device receives video footage from surveillance cameras 24 / 7. As input, high-resolution video data is acquired from the surveillance cameras. This video data is sent in real time to the device's AI analysis engine for initial analysis. Specifically, the device captures video data and uses AI models to detect anomalies in human movement and audio. As output, potentially anomalous behavior and audio data are filtered.
[1382] Step 2:
[1383] Emotion Recognition and Anomaly Detection
[1384] The emotion engine built into the device receives the filtered video and audio data from the previous step as input. It analyzes this data and recognizes emotions based on the user's facial expressions and tone of voice. Specifically, the device uses facial expression recognition algorithms and voice analysis algorithms to identify abnormal emotions such as fear or anger. The output is the recognized emotion information and a judgment on whether it is abnormal.
[1385] Step 3:
[1386] Anomaly detection and risk assessment
[1387] The device receives the emotional information and abnormal behavior data acquired in step 2 as input and performs a risk assessment. Previously learned abnormal behavior models and emotional data are used for the risk assessment. Specifically, the device compares the behavior data and emotional data with the previous abnormal behavior models to assess the risk level. The output generates an assessment result such as high risk, low risk, or medium risk.
[1388] Step 4:
[1389] Sending data to the server
[1390] If the device evaluates the anomaly as high risk, it sends the video data and emotion data to the server. The input includes the risk assessment result, video data, and emotion data. Specifically, the device compresses these data and quickly transmits them to the server via the network. The output is the data received by the server.
[1391] Step 5:
[1392] Advanced AI analysis
[1393] The server receives data from the device as input and performs a more detailed risk assessment. The server compares it with past crime data and emotion data to infer the next course of action. Specifically, the server uses an advanced AI analysis engine to analyze the received data and infer the next course of action and the most likely scenario. The output is the inference result and a more detailed risk assessment result.
[1394] Step 6:
[1395] Notification to external agencies
[1396] If the server determines that the risk is extremely high, it automatically notifies external agencies such as the police. The inputs include risk assessment results, inference results, video data, and emotion data. Specifically, the server activates an automatic notification system, attaches relevant information, and sends it to the police. The output is a notification sent to an external agency. This notification allows for a rapid response.
[1397] (Application example 2)
[1398] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1399] Conventional security systems have focused on detecting anomalies in people's behavior and actions, but do not include the ability to recognize anomalies based on emotions or facial expressions. As a result, emotional risks may be overlooked when abnormal behavior occurs, and the accuracy of anomaly detection may be insufficient. Furthermore, it is often difficult to respond to detected abnormal behavior in real time. This makes it difficult to notify external agencies in a timely manner, making it difficult to take appropriate measures.
[1400] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1401] In this invention, the server includes a means for acquiring video, a means for analyzing the acquired video in real time to detect abnormalities, a means for recognizing emotions from the acquired video, a means for assessing the risk if the abnormality is detected, a means for automatically notifying an external organization of the analysis results and emotion recognition results if the risk is assessed as high, and a means for inferring the next action using past data. This enables anomaly detection based not only on abnormal behavior but also on emotions, enabling highly accurate risk assessment and rapid notification. Furthermore, by taking appropriate measures in real time, a safer environment can be provided.
[1402] Definition of Terms
[1403] "Footage" means visual data captured by surveillance cameras or other imaging devices.
[1404] "Real-time" means that video data is analyzed and processed as soon as it is acquired.
[1405] "Analysis" refers to the act of analyzing acquired video data using algorithms and AI models to detect specific information or anomalies.
[1406] "Abnormal" refers to behavior or situations that deviate from normal behavior or situations, including behavior or facial expressions that the system judges to be high risk.
[1407] "Emotions" refer to sensations and moods extracted from human facial expressions and tone of voice recognized by the system.
[1408] "Emotion recognition" is a technology that analyzes video and audio data captured by surveillance cameras to identify a person's emotional state.
[1409] "Risk assessment" is the process of determining how dangerous a situation is based on detected anomalies and emotions.
[1410] "External agencies" refers to the police, security companies, and other organizations or groups responsible for security.
[1411] "Notification" refers to the act of the system sending analysis results or risk assessment results to an external organization.
[1412] "Past data" refers to historical information such as previously acquired video data, emotion recognition results, and anomaly detection data.
[1413] "Behavioral inference" is a technology that allows a system to predict the next likely action or situation based on past data.
[1414] MODE FOR CARRYING OUT THE INVENTION
[1415] The present invention provides an advanced security system that combines a surveillance camera and an emotion engine. Specific examples thereof are described below.
[1416] 1. Hardware Configuration
[1417] The system includes a server, a terminal, and a user terminal.
[1418] The device is equipped with a high-resolution camera for wide-area monitoring and an AI analysis engine for initial analysis.
[1419] The server has the computing power to perform advanced AI analysis and the ability to notify external agencies such as the police.
[1420] The user terminal is provided with an interface that allows security managers, police, etc. to monitor and manage abnormalities.
[1421] 2. Software Configuration
[1422] The device comes with software to capture and analyze video data in real time, using OpenCV to capture video data and Keras to run emotion recognition models.
[1423] The server runs software that performs advanced analysis of the data it receives, which is then stored in a database and compared with previous data.
[1424] The user terminal is provided with a graphical user interface (GUI) for displaying the results of anomaly detection and risk assessment.
[1425] 3. Processing flow
[1426] The device monitors the video in real time 24 hours a day, and if an abnormality is detected, the emotion engine analyzes the user's emotions from the video data. For example, if a store's surveillance camera detects someone shouting and recognizes anger from their facial expression, it will assess the risk as an abnormality.
[1427] The server receives the abnormal information and emotion recognition results that are judged to be high risk, compares it with past data, and infers the next action. This then notifies external agencies of the next possible action. For example, if it infers that there is a high possibility of a robber breaking into a store, it will automatically notify the police.
[1428] 4. Specific Examples
[1429] The device detects a suspicious person in front of the store and confirms that the person is holding a knife by shouting.
[1430] The device's emotion engine recognizes anger from the person's facial expression and determines that the person is at high risk.
[1431] Several minutes of video and emotional data are sent to a server, which then performs a detailed analysis.
[1432] The server compares the data with past crime data and determines that the person is at high risk of breaking into a store.
[1433] The server automatically notifies the police of the risk assessment results and video data, allowing them to respond quickly to the scene.
[1434] 5. Examples of prompts
[1435] TXT
[1436] Create an AI program that analyzes video data from security cameras in real time and detects suspicious behavior and emotional anomalies. Consider the following points:
[1437] 1. A system for acquiring video data in real time
[1438] 2. Use a pre-trained emotion recognition model to perform emotion recognition.
[1439] 3. A system that notifies you in real time when an abnormality is detected
[1440] 4. Specific examples and explanations of the program
[1441] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1442] Program processing flow
[1443] Step 1:
[1444] The device acquires video data from the surveillance camera in real time.
[1445] Input: Visual data (footage) captured by surveillance cameras.
[1446] Output: Real-time video stream.
[1447] How it works: The camera module in the device captures video and sends the data to the analysis engine. Specifically, it uses libraries such as OpenCV to capture video.
[1448] Step 2:
[1449] The video data acquired by the device is analyzed in real time to detect abnormalities.
[1450] Input: Real-time video stream.
[1451] Output: Anomaly detection results (e.g., unusual behavior, loud audio, etc.).
[1452] How it works: The device's AI analysis engine analyzes the video data and compares it with a pre-trained abnormal behavior model to detect abnormalities. Specifically, it performs object recognition and movement analysis.
[1453] Step 3:
[1454] The device recognizes emotions using video data.
[1455] Input: Real-time video stream and detected anomalies.
[1456] Output: Emotion recognition result (e.g. anger, fear, etc.).
[1457] How it works: The device's emotion recognition engine analyzes human facial expressions from video data captured by the camera and recognizes emotions. Specifically, it runs an emotion recognition model using Keras.
[1458] Step 4:
[1459] The device assesses risk based on anomalies and emotion recognition results.
[1460] Input: Anomaly detection results and emotion recognition results.
[1461] Output: Risk assessment result (e.g. high risk, medium risk, low risk).
[1462] How it works: The device's risk assessment module evaluates how dangerous a situation is based on abnormal behavior and emotional data. For example, if anger and loud voices are detected at the same time, it will be judged as high risk.
[1463] Step 5:
[1464] If a server is assessed as high risk, the analysis results and emotion recognition results will be automatically notified to external agencies.
[1465] Input: Risk assessment results and analysis data, emotion recognition results.
[1466] Output: Notification data to external agencies (e.g. police).
[1467] How it works: The server generates a notification message containing detailed information about the anomaly and sends it to an external agency, such as the police. Specifically, this can be done by sending data via an HTTP request.
[1468] Step 6:
[1469] The server uses past data to infer the next action.
[1470] Input: Detailed anomaly data, historical database information.
[1471] Output: The inferred result of the next possible action.
[1472] Operation: The server's AI inference module analyzes past data, compares it with current anomalies, and predicts the next action. Based on this, it becomes possible to provide more appropriate countermeasures to external organizations. Specifically, it performs behavior analysis using a predictive algorithm.
[1473] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1474] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1475] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1476] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1477] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1478] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1479] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1480] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1481] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1482] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1483] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1484] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1485] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1486] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1487] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1488] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1489] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1490] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1491] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1492] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1493] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1494] The following is further disclosed regarding the above embodiment.
[1495] (Claim 1)
[1496] a means for acquiring the image;
[1497] A means of analyzing the acquired video in real time to detect abnormalities,
[1498] a means for evaluating a risk when the abnormality is detected;
[1499] means for automatically notifying an external organization of the analysis results when the risk is assessed to be high;
[1500] A means of inferring next actions using past data;
[1501] A system including:
[1502] (Claim 2)
[1503] The system according to claim 1, wherein the system notifies the user of a predicted next action together with the analysis result.
[1504] (Claim 3)
[1505] 2. The system of claim 1, wherein the external agency is the police.
[1506] "Example 1"
[1507] (Claim 1)
[1508] a means for acquiring the image;
[1509] A means of analyzing the acquired video in real time to detect abnormalities,
[1510] a means for evaluating a risk when the abnormality is detected;
[1511] means for transmitting the compressed video data to a central processing unit if the risk is assessed to be high;
[1512] A means for the central processing unit to analyze the received data in detail and compare it with a past database to infer the next action;
[1513] a means for automatically communicating the inference results and risk assessment results to external parties;
[1514] A means to carry out a series of processes in real time, from anomaly detection to risk assessment, analysis, and notification,
[1515] A system including:
[1516] (Claim 2)
[1517] The system according to claim 1, wherein the system notifies the user of a predicted next action together with the analysis result.
[1518] (Claim 3)
[1519] 2. The system of claim 1, wherein the external agency is a security agency.
[1520] "Application Example 1"
[1521] (Claim 1)
[1522] a means for acquiring the image;
[1523] A means of analyzing the acquired video in real time to detect abnormalities,
[1524] a means for evaluating a risk when the abnormality is detected;
[1525] means for automatically notifying an external organization of the analysis results when the risk is assessed to be high;
[1526] A means of inferring next actions using past data;
[1527] A system that uses a wearable device worn by security guards, which includes means for providing visual and audio warnings when an abnormality is detected.
[1528] (Claim 2)
[1529] The system according to claim 1, wherein the system notifies the user of a predicted next action together with the analysis result.
[1530] (Claim 3)
[1531] 2. The system of claim 1, wherein the external agency is a law enforcement organization.
[1532] "Example 2: Combining Emotion Engines"
[1533] (Claim 1)
[1534] a means for acquiring the image;
[1535] A means of analyzing the acquired video in real time to detect abnormalities,
[1536] means for recognizing emotions from the acquired video and audio data;
[1537] a means for evaluating a risk when the abnormality is detected;
[1538] means for automatically notifying an external organization of the analysis results when the risk is assessed to be high;
[1539] A means of inferring next actions using past data;
[1540] A system including:
[1541] (Claim 2)
[1542] The system according to claim 1, wherein the system notifies the user of a predicted next action together with the analysis result.
[1543] (Claim 3)
[1544] 2. The system of claim 1, wherein the external agency is the police.
[1545] "Application example 2 when combining emotion engines"
[1546] New Claims
[1547] (Claim 1)
[1548] a means for acquiring the image;
[1549] A means of analyzing the acquired video in real time to detect abnormalities,
[1550] means for recognizing emotions from the acquired video;
[1551] a means for evaluating a risk when the abnormality is detected;
[1552] means for automatically notifying an external organization of the analysis results and emotion recognition results when the risk is assessed to be high;
[1553] A means of inferring next actions using past data;
[1554] A system including:
[1555] (Claim 2)
[1556] The system according to claim 1, wherein a predicted next action is notified together with the analysis result and the emotion recognition result.
[1557] (Claim 3)
[1558] 2. The system of claim 1, wherein the external agency is the police. [Explanation of symbols]
[1559] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. a means for acquiring the image; A means of analyzing the acquired video in real time to detect abnormalities, a means for evaluating a risk when the abnormality is detected; means for automatically notifying an external organization of the analysis results when the risk is assessed to be high; A means of inferring next actions using past data; A system including:
2. The system according to claim 1 , wherein a predicted next action is notified together with the analysis result.
3. 2. The system of claim 1, wherein the external agency is a police force.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A