System
The system addresses inefficiencies in crime prevention by preprocessing and analyzing security data with generative models, sending real-time warnings, and building a feedback loop to improve prediction accuracy and response times.
Patent Information
- Application Number
- JP2024130303
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-06
- Publication Date
- 2026-02-19
AI Technical Summary
Existing crime prevention systems face inefficiencies in data analysis, leading to delayed responses and inaccurate predictions due to false alarms and insufficient feedback loops, making them inadequate for real-time crime prediction and response.
A system that collects security data from surveillance devices, preprocesses it, analyzes crime risk using a generative model, sends real-time warnings, and builds a feedback loop to improve prediction accuracy by learning from user interactions.
Enables efficient data processing, accurate crime risk analysis, and immediate response by integrating data preprocessing, generative models, and a feedback mechanism to enhance prediction accuracy.
Smart Images

Figure 2026028005000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In modern society, there are an increasing number of situations where crime prevention and immediate response are required, and simple surveillance cameras and alarm systems alone are not sufficient in many cases. In particular, there is a demand for systems that can achieve highly accurate, real-time crime prediction and automatic response. However, current systems face challenges that make it difficult to achieve complete crime prevention due to inefficient data analysis, delayed responses due to false alarms and delays, and insufficient feedback loops. [Means for solving the problem]
[0005] The present invention provides the following means.
[0006] The system includes a means for collecting security data from a monitoring device, a means for preprocessing the collected data, a means for analyzing crime risk using a generative model based on the preprocessed data, a means for sending a warning notification to a user or relevant organization based on the analysis results, and a means for feeding back data after the warning to the generative model to perform learning.
[0007] This system enables efficient data pre-processing and highly accurate analysis, strengthening crime prevention measures that require immediate response. In addition, by creating a feedback loop, prediction accuracy can be improved based on the latest information, eliminating the problems of false alarms and delays and realizing swift and accurate responses.
[0008] A "surveillance device" is a hardware device used to monitor a physical location or environment and collect data in the form of video, audio, motion sensors, or the like.
[0009] "Security data" is a general term for all security-related data, including video data, audio data, and motion sensor data collected from surveillance devices.
[0010] "Means of collection" refers to the processes and mechanisms used to obtain security data from monitoring devices and incorporate it into the system.
[0011] "Preprocessing" refers to processes that convert collected security data into a format that is easy to analyze with a generative model. Specifically, this includes data standardization, image frame division, and audio chunking.
[0012] A "generative model" is a computer model that uses machine learning algorithms to analyze and make predictions based on collected and preprocessed data.
[0013] A "means for analyzing crime risk" is a process for identifying indicators and risks of crime from pre-processed security data using a generative model.
[0014] The "means for sending warnings or notifications to users or relevant organizations based on the analysis results" refers to the process of receiving the analysis results of the Generative Model and electronically sending warnings or notifications to users or relevant organizations as necessary.
[0015] The "means of feeding back data after issuing a warning to the generative model for learning" refers to a mechanism that collects the reactions and results of users and related organizations in response to the issued warning and uses them as learning data for the generative model, thereby improving the model's predictive accuracy. [Brief explanation of the drawings]
[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10]1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0024] [First embodiment]
[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0037] Overall overview
[0038] This system uses security data collected from surveillance devices to analyze crime risks using generative models and issue real-time warnings. It is composed of a server, terminals, and users, and operates in cooperation with each other.
[0039] Program processing
[0040] Data collection
[0041] The server collects data in real time from multiple surveillance devices (surveillance cameras, audio sensors, motion sensors, etc.) and has a buffer to temporarily store the data pulled from these devices.
[0042] Data Preprocessing
[0043] The server pre-processes the collected security data: video data is split into frames, audio data is split into analyzable chunks based on sample rate, motion sensor data is digitized, and image processing such as noise removal and edge detection is performed.
[0044] Data analysis
[0045] The device inputs the preprocessed data into a generative model and performs analysis. It uses image recognition algorithms to detect suspicious movements and voice recognition algorithms to detect abnormal sounds. It analyzes motion patterns and detects unnatural movements.
[0046] Crime risk assessment
[0047] The server then uses the analysis results to determine the risk of crime. This determination is based on pre-set thresholds and rules. For example, a risk score is calculated based on the identification pattern of a suspicious person or specific frequencies of a voice.
[0048] Notifications and Alerts
[0049] Based on the risk assessment results, the device will send notifications or warnings to the user or relevant authorities via push notifications to smartphones or tablets, email, SMS, etc. If necessary, an API will also be used to automatically notify the police or security companies.
[0050] User Support
[0051] Users can take appropriate action based on the warnings and notifications they receive, such as checking the warning on their smartphone, checking live footage from surveillance cameras, or contacting security companies or police to instruct them to investigate the scene.
[0052] Building a feedback loop
[0053] The server collects the results of the warnings and new crime data, and feeds them back as training data for the generative model. This feedback allows the model to continuously learn and improve its prediction accuracy.
[0054] Specific examples
[0055] Example 1: Detecting and warning suspicious individuals
[0056] 1. The server collects video data from the surveillance cameras.
[0057] 2. The device analyzes this video data using a generative model to detect suspicious activity.
[0058] 3. The server determines that the crime risk is high and raises an alert flag.
[0059] 4. The device sends a warning notification to the user's smartphone.
[0060] 5. The user checks the alert on their smartphone and checks the surveillance camera footage in real time.
[0061] 6. The user will notify the police if necessary.
[0062] Example 2: Abnormal sound detection and warning
[0063] 1. The server collects audio data from intercoms and audio sensors.
[0064] 2. The device analyzes the audio data and detects abnormal sounds such as breaking glass.
[0065] 3. The server determines that the risk of crime is high and sends a warning notice to the user while automatically reporting the incident to the police.
[0066] 4. The user checks the alerts on their smartphone and monitors the situation at the site.
[0067] 5. The police will rush to the scene and take appropriate action.
[0068] This will enable 24-hour security response and immediate crime prevention, improving the safety of the property.
[0069] The processing flow will be explained below.
[0070] Step 1:
[0071] The server collects real-time data from surveillance devices such as surveillance cameras, audio sensors, motion sensors, etc. Specifically, it periodically pulls data from each device (e.g., every second) and temporarily stores this data in a buffer on the server.
[0072] Step 2:
[0073] The server pre-processes the collected security data, splitting the video data into frames and the audio data into analyzable chunks based on sample rate, including applying image processing techniques such as noise removal and edge detection to shape the data.
[0074] Step 3:
[0075] The device inputs the preprocessed data into a generative model, which uses an image recognition algorithm to detect suspicious movements and a voice recognition algorithm to detect abnormal sounds. The device then analyzes the motion patterns to detect unnatural movements.
[0076] Step 4:
[0077] The server determines the crime risk based on the analysis results of the generative model. It calculates a risk score based on pre-set thresholds and rules, and if the risk score exceeds a certain threshold, it sets an alert flag. This result is recorded in a database.
[0078] Step 5:
[0079] If the device determines that there is a high risk of crime, it will send a warning to the user and relevant authorities via push notification to the smartphone or tablet, email or SMS, and, if necessary, automatically notify the police or security company using an API.
[0080] Step 6:
[0081] The user checks the received warnings and notifications and takes appropriate action. For example, they can check the warnings on their smartphones, check real-time footage from surveillance cameras, and, if necessary, contact a security company or police and instruct them to investigate the scene.
[0082] Step 7:
[0083] The server collects the results of warnings and new crime data and feeds it back to the generative model. This feedback allows the generative model to continuously learn and improve its prediction accuracy. The collected data is fed into the model as re-training data, and the parameters are adjusted.
[0084] Example 1
[0085] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0086] In recent years, the spread of security devices has progressed as crime risks have increased. However, conventional security systems have faced challenges in analyzing crime risks in real time and immediately sending appropriate warning notifications. Furthermore, there have been concerns about a decline in prediction accuracy due to insufficient updating of the training data for generative models. The objective of the present invention is to solve these challenges and provide a system that realizes highly accurate crime risk analysis and warning notifications in real time.
[0087] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0088] In this invention, the server includes means for collecting security data from monitoring devices, means for preprocessing the collected data, means for analyzing crime risk based on the preprocessed data using a generative model, means for sending a warning to a user or a relevant organization based on the analysis results, means for feeding back data after the warning to the generative model for learning, means for acquiring data in real time and providing a buffer for temporarily storing it, means for using a voice analysis algorithm to detect abnormal sounds, and means for using a cloud messaging service as a notification method. This makes it possible to analyze the data collected from the monitoring devices with high accuracy, determine crime risk in real time, and promptly send appropriate warnings.
[0089] A "surveillance device" is a device, such as a camera, audio sensor, or motion sensor, that collects information occurring within a specific space or environment.
[0090] "Security data" is a general term for information such as video, audio, and behavior patterns obtained from surveillance devices.
[0091] "Preprocessing" refers to the process of converting collected data into an analyzable form, and specifically includes dividing video data into frames and chunking audio data.
[0092] "Generative model" generally refers to machine learning and artificial intelligence algorithms used to predict and analyze crime risk from data.
[0093] "Crime risk" is a value that assesses the likelihood of suspicious activity or abnormal events occurring based on collected security data using certain criteria.
[0094] "Cloud Messaging Service" means an online service for sending push notifications over the Internet.
[0095] "Feedback" is the process of collecting results and new data after a warning and reusing them as training data for the analytical model.
[0096] "Buffer" refers to a temporary storage area for data, and is used to temporarily store data collected in real time.
[0097] An "audio analysis algorithm" is a mathematical method or technique for analyzing audio data to identify specific sounds or patterns.
[0098] "Notification" refers to the act of sending warnings or information to users or relevant organizations based on the analysis results.
[0099] As an embodiment of the invention, the system consists of three main components: a server, a terminal, and a user. The whole system is based on monitoring devices, where the collected data is sequentially pre-processed, analyzed, notified, and fed back.
[0100] Data collection
[0101] The server collects real-time data from multiple monitoring devices (e.g., cameras, audio sensors, motion sensors). This data is first temporarily stored in a buffer on an AWS EC2 instance. RTSP (Real-Time Streaming Protocol) and HTTP are used to obtain real-time data.
[0102] Data Preprocessing
[0103] The server splits the collected video data into frames using OpenCV, and splits the audio data into analyzable chunks using Librosa. It also performs preprocessing such as noise reduction and normalization. Motion sensor data is stored as numerical data after calibration.
[0104] Data analysis
[0105] The device then inputs the preprocessed data into generative AI models to perform analysis. For example, video data is analyzed using TensorFlow's YOLO model to detect suspicious activity, audio data is analyzed using the Google Cloud Speech-to-Text API to identify abnormal sounds, and motion data is analyzed using the SciPy library to detect unnatural movements.
[0106] Crime risk assessment
[0107] The server then uses the analysis results to determine the crime risk according to pre-set thresholds and rules. If the threshold is exceeded, a risk score is calculated and the person is deemed high risk based on that score.
[0108] Notifications and Alerts
[0109] Based on the server's assessment, the device sends a warning to the user or relevant authorities. Notifications are sent via Firebase Cloud Messaging, and emails and SMS are sent via the Twilio API. The system also includes an automatic notification function to the police and security companies using the API.
[0110] User Support
[0111] Users receive a warning notification sent from the device and check it on their smartphone app. The app provides a function to view live footage from surveillance cameras, allowing users to take appropriate action based on this information. If necessary, they can also report the incident to a security company or the police.
[0112] Building a feedback loop
[0113] The server continuously collects the results of warnings and newly acquired crime data, and uses them to retrain the generative model, thereby continuously improving the prediction accuracy of the generative AI model.
[0114] Specific examples
[0115] Example 1: Detecting and warning suspicious individuals
[0116] 1. The server collects video data from the surveillance cameras.
[0117] 2. The device analyzes this video data using a generative model to detect suspicious activity.
[0118] 3. The server determines that the crime risk is high and raises an alert flag.
[0119] 4. The device sends a warning notification to the user's smartphone.
[0120] 5. The user checks the alert on their smartphone and checks the surveillance camera footage in real time.
[0121] 6. The user will notify the police if necessary.
[0122] Example prompts to input to the generative AI model
[0123] "Please explain in natural language how your AI system analyzes collected data and generates a risk score when it detects suspicious activity or abnormal sounds, including specific libraries and tools."
[0124] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0125] Step 1: Data collection
[0126] The server collects data from surveillance devices. Specifically, the server acquires video data from surveillance cameras in real time using RTSP and temporarily stores it in a buffer on an AWS EC2 instance. It also collects audio data from audio sensors via HTTP and data from motion sensors via Bluetooth. It receives stream data from surveillance devices as input and outputs it to a temporary storage buffer.
[0127] Step 2: Preprocess the data
[0128] The server preprocesses the collected security data. Specifically, it uses OpenCV to split the video data into frames and Librosa to split the audio data into one-second chunks. Motion sensor data is digitized, denoised, and normalized. It receives raw data stored in a buffer as input and outputs data that can be converted into an analyzable format.
[0129] Step 3: Data analysis
[0130] The device inputs the preprocessed data into a generative model to perform analysis. Specifically, the device uses TensorFlow's YOLO model to analyze each frame and detect suspicious motion. It uses the Google Cloud Speech-to-Text API to analyze chunked audio data and identify abnormal sounds. It uses SciPy's signal processing functions for motion pattern analysis. It takes the preprocessed data as input and outputs specific analysis results.
[0131] Step 4: Determine crime risk
[0132] The server determines the crime risk based on the analysis results. Specifically, the server calculates a risk score based on thresholds and rules, and if it exceeds a certain value, it determines the risk as high. For example, it sets a score based on the frequency of suspicious movements or abnormal sounds. It receives the analysis results as input and outputs risk assessment data including a risk score.
[0133] Step 5: Notifications and warnings
[0134] The device sends a warning notification to the user or relevant authorities based on the risk assessment results. Specifically, the device sends a push notification to the smartphone via Firebase Cloud Messaging, sends an emergency email or SMS using the Twilio API, and automatically notifies the police or security company via the API. It receives risk assessment data as input and outputs a notification message.
[0135] Step 6: User interaction
[0136] The user can take appropriate action based on the warning notification they receive. Specifically, the user checks the warning on their smartphone and checks live footage from the surveillance camera within the app. If necessary, they can also notify a security company or the police. The app receives the notification message as input and outputs the response action to be taken.
[0137] Step 7: Building a feedback loop
[0138] The server collects the results of the warning and newly acquired crime data and feeds it back as training data for the generative model. Specifically, the server continuously monitors and collects new data and uses it to retrain the generative model, thereby improving the model's predictive accuracy. It receives the post-warning data as input and outputs it as retraining data.
[0139] (Application example 1)
[0140] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0141] In conventional security systems, when analyzing data collected from monitoring devices, it has been difficult to detect anomalies in real time, provide immediate notification, and automatically report to relevant authorities. This increases the possibility of delays in crime prevention and prompt response, posing a security issue. The present invention aims to solve these problems and provide a security system that can detect anomalies in real time and respond immediately.
[0142] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0143] In this invention, the server includes means for collecting security data from monitoring devices, means for preprocessing the collected data, means for analyzing crime risks using a generative model based on the preprocessed data, means for sending a warning to a user or relevant organizations based on the analysis results, means for feeding back data after the warning to the generative model for learning, means for sending a push notification when an abnormality is detected based on the analysis results, and means for automatically notifying relevant organizations when an abnormality is detected. This makes it possible to detect abnormalities in real time and to immediately notify / report them.
[0144] A "surveillance device" is a device, such as a camera or sensor, used to collect security data.
[0145] "Security data" refers to security-related information such as video, audio, and motion sensor data obtained through surveillance devices.
[0146] "Preprocessing" refers to the process of converting collected security data into a format suitable for analysis, and includes dividing image data into frames and audio data into chunks.
[0147] A "generative model" is an artificial intelligence model used to detect patterns and anomalies in collected data.
[0148] "Crime risk analysis" refers to the use of generative models to assess the likelihood of a crime occurring based on collected data.
[0149] "Sending a warning notification" means issuing a warning to users and relevant organizations when an abnormality is detected based on the analysis results.
[0150] "Feedback" refers to the process of feeding post-alert data and new security data back into the generative model to allow it to continuously learn.
[0151] "Push notification" is a function that immediately sends a warning message to a user's smartphone or other device when an abnormality is detected.
[0152] "Automatic reporting" is a function that automatically reports to relevant authorities (such as the police or security companies) when a crime risk or abnormality is detected.
[0153] System configuration
[0154] An embodiment of the present invention is a system that collects security data from monitoring devices, analyzes it using a generative AI model, and determines crime risk. This system is mainly composed of a server, a terminal, and a user, and each element works in cooperation with each other.
[0155] Program processing overview
[0156] The server collects data in real time from multiple monitoring devices (e.g., surveillance cameras, audio sensors, motion sensors, etc.) and has a buffer that temporarily stores it. The collected data is divided into frames, and the image data is divided into analyzable chunks based on the sample rate. The motion sensor data is also digitized and processed for noise removal, edge detection, etc.
[0157] The device inputs the preprocessed data into a generative model and performs analysis to determine the crime risk. This analysis includes detecting suspicious activity using image recognition algorithms, detecting abnormal sounds using voice recognition algorithms, and analyzing motion patterns. Based on the analysis results, the device evaluates the crime risk and calculates a risk score.
[0158] The server determines the crime risk based on the analysis results, and if the risk exceeds a set threshold, it sends a push notification to the user's smartphone or tablet, and automatically reports the situation to relevant authorities, such as the police or security companies, in real time.
[0159] Additionally, post-alert data and new security data are fed back into the generative model, allowing it to continuously learn and improve its predictions. This feedback loop allows the system to make increasingly accurate predictions over time.
[0160] Hardware and software used
[0161] Hardware:
[0162] Smartphone: Used by users to receive warning notifications and check surveillance footage in the event of an abnormality.
[0163] Surveillance cameras: Used to collect security data.
[0164] Sensors: Used to collect additional security data, such as sound and motion.
[0165] software:
[0166] Keras (keras.models, keras.preprocessing): Used to analyze video data using deep learning models.
[0167] OpenCV (cv2): Used to capture and pre-process video data.
[0168] Gmail SMTP (smtplib, email): Email sending system used for automated notifications.
[0169] Plyer (notification): Used to send push notifications to smartphones.
[0170] Specific examples
[0171] For example, a surveillance camera can detect suspicious activity and analyze the video data using a Keras model. If the result indicates a high risk of crime, the Plryer library can be used to send a push notification to the user's smartphone, and the relevant authorities can be automatically notified via Gmail SMTP. Similarly, if an audio sensor detects the sound of glass breaking, an immediate notification and report can be sent.
[0172] Prompt Sentence Examples
[0173] "Please create an application that analyzes surveillance camera footage in a specified area, detects suspicious individuals and abnormal sounds in real time, and sends push notifications and automatic reporting."
[0174] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0175] Step 1:
[0176] The server collects security data in real time from multiple monitoring devices (such as surveillance cameras, audio sensors, and motion sensors). The collected data is temporarily stored in a buffer. The input is real-time data from the monitoring devices, and the output is the security data stored in the buffer. Specifically, data is pulled from each monitoring device and stored in the server's buffer area.
[0177] Step 2:
[0178] The server preprocesses the collected security data. Image data is divided into frames, and audio data is divided into analyzable chunks based on the sample rate. Motion sensor data is also digitized and processed for noise removal and edge detection. The input is the buffered security data, and the output is the preprocessed data. Specifically, OpenCV is used to divide the video data into frames, and the audio data into chunks based on the sample rate.
[0179] Step 3:
[0180] The device inputs the preprocessed data into a generative AI model to analyze crime risk. This analysis includes detecting suspicious activity using an image recognition algorithm, detecting abnormal sounds using a voice recognition algorithm, and analyzing motion patterns. The input is the preprocessed data, and the output is the analysis results. Specifically, Keras is used to load the generative model, and the preprocessed data is input to the model to perform the analysis.
[0181] Step 4:
[0182] The server determines the crime risk based on the analysis results received from the terminal. If the crime risk is determined to be high, it sets an alert flag based on the set threshold. The input is the analysis result, and the output is the status of the alert flag. Specifically, it determines whether the threshold is exceeded based on the analysis result, and if so, sets an alert flag in the internal data structure.
[0183] Step 5:
[0184] When an alert flag is raised, the server sends a push notification to the user's smartphone or tablet. Furthermore, if necessary, it automatically notifies the police or security company. The input is the alert flag status and analysis results, and the output is the warning notification and notification results. Specifically, it uses the Plryer library to send notifications to smartphones, and uses Gmail SMTP to notify the relevant authorities via email.
[0185] Step 6:
[0186] Based on the received warning notification, the user checks the live footage from the surveillance camera and contacts the police or security company if necessary. The input is the content of the warning notification, and the output is the confirmation result of the abnormal situation and the countermeasures. Specifically, the user checks the received notification on their smartphone and displays the real-time footage through the application.
[0187] Step 7:
[0188] The server collects post-warning data and new security data and feeds it back as training data for the generative model. This allows the generative model to continuously learn and improve its prediction accuracy. The input is the post-warning data and new security data, and the output is an updated generative model. Specifically, the collected data is added to the dataset and the model is retrained using Keras.
[0189] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0190] Overall overview
[0191] This invention combines a system that uses security data collected from surveillance devices to analyze crime risks using generative models and send real-time warnings, with an emotion engine that recognizes user emotions. The system is composed of a server, a terminal, and a user, and operates in cooperation with each other.
[0192] Program processing
[0193] Data collection
[0194] The server collects data in real time from surveillance devices such as surveillance cameras, audio sensors, and motion sensors, and the collected data is temporarily stored in a buffer on the server.
[0195] Data Preprocessing
[0196] The server pre-processes the collected security data: video data is split into frames, and audio data is split into analyzable chunks based on sample rate. Image processing such as noise reduction and edge detection is also performed.
[0197] Data analysis
[0198] The device then inputs the preprocessed data into a generative model to perform a crime risk analysis. Image recognition algorithms are used to detect suspicious activity, and voice recognition algorithms are used to detect abnormal sounds. Motion patterns are also analyzed to detect unnatural movements.
[0199] Crime risk assessment
[0200] The server determines the crime risk based on the analysis results of the generative model. It calculates a risk score based on pre-set thresholds and rules, and sets an alert flag if the risk score exceeds a certain threshold.
[0201] Emotion analysis
[0202] The device's built-in emotion engine analyzes the user's emotions in real time, analyzing the user's facial expressions, tone of voice, and movement patterns from collected video and audio data to identify their emotional state.
[0203] Notifications and Alerts
[0204] The device sends notifications and warnings to the user and relevant authorities based on the crime risk assessment results and the user's emotional analysis results. The content and method of the notification change depending on the specific emotional state (e.g., tension, surprise, anger, etc.). For example, if emergency notification is set, the police and security companies will be automatically notified immediately.
[0205] User Support
[0206] Users can take appropriate action based on the warnings and notifications they receive, such as checking the content of the warning and checking live footage from surveillance cameras. Users can also issue more specific instructions based on the analysis results of the emotion engine.
[0207] Building a feedback loop
[0208] The server collects the results of the alerts and new crime data, and feeds it back into the generative model and emotion engine as training data. This feedback allows the model to continuously learn and improve its prediction accuracy.
[0209] Specific examples
[0210] Example 1: Detecting and warning suspicious individuals
[0211] 1. The server collects video data from the surveillance cameras.
[0212] 2. The device analyzes the video data using a generative model to detect suspicious activity.
[0213] 3. The server determines the risk is high and raises an alert flag.
[0214] 4. The emotion engine analyzes the user's facial expressions and voice to identify states of surprise and tension.
[0215] 5. The device sends an alert to the user's smartphone as an emergency notification.
[0216] 6. The user checks the alert on their smartphone and checks live footage from the security camera.
[0217] 7. The user will notify the police if necessary.
[0218] Example 2: Abnormal sound detection and warning
[0219] 1. The server collects audio data from intercoms and audio sensors.
[0220] 2. The device analyzes the audio data using a generative model to detect abnormal sounds, such as the sound of glass breaking.
[0221] 3. The server determines the risk is high and raises an alert flag.
[0222] 4. The emotion engine analyzes the user's voice tone and identifies their state of excitement.
[0223] 5. The device will immediately send a warning notification to the user and automatically notify the police.
[0224] 6. Users can check alerts on their smartphones and monitor the situation at the site.
[0225] 7. The police will rush to the scene and take appropriate action.
[0226] This will enable quick and accurate crime prevention and response that takes into account the user's emotional state, further improving the safety of properties.
[0227] The processing flow will be explained below.
[0228] Step 1:
[0229] The server collects data in real time from monitoring devices such as surveillance cameras, audio sensors, motion sensors, etc. Specifically, it periodically acquires data from each device (e.g., every second) and temporarily stores this data in a buffer on the server.
[0230] Step 2:
[0231] The server preprocesses the collected security data, splitting the video data into frames and the audio data into analyzable chunks based on sample rate, and applying image processing techniques such as noise removal and edge detection to shape the data.
[0232] Step 3:
[0233] The device inputs the preprocessed data into a generative model to perform crime risk analysis. Specifically, it uses image recognition algorithms to detect suspicious activity and voice recognition algorithms to detect abnormal sounds. It also analyzes motion patterns to identify unnatural movements.
[0234] Step 4:
[0235] The server determines the crime risk based on the analysis results of the generative model. It calculates a risk score based on pre-set thresholds and rules, and if this risk score exceeds a certain threshold, it sets an alert flag. This determination result is recorded in a database.
[0236] Step 5:
[0237] The device's built-in emotion engine analyzes the user's emotions in real time by analyzing the user's facial expressions, tone of voice, and movement patterns from collected video and audio data to identify their emotional state.
[0238] Step 6:
[0239] The device sends notifications and warnings to the user and relevant authorities based on the crime risk assessment results and the user's emotional analysis results. The content and method of notifications can be changed depending on the user's emotional state, such as tension or surprise. For example, in an emergency, the device can immediately and automatically notify the police or security company.
[0240] Step 7:
[0241] The user checks the warnings and notifications they receive and takes appropriate action. Specifically, they check the warnings on their smartphones, check live footage from surveillance cameras, and, based on the analysis results of the emotion engine, contact security companies and police and instruct them to investigate the scene.
[0242] Step 8:
[0243] The server collects the results of warnings and new crime data, and feeds it back as training data for the generative model and emotion engine. This feedback allows the model to continuously learn and improve its prediction accuracy. The collected data is fed into the model as re-training data, and the parameters are adjusted.
[0244] Specific examples
[0245] Example 1: Detecting and warning suspicious individuals
[0246] 1. The server collects video data from the surveillance cameras.
[0247] 2. The server preprocesses the video data and divides it into frames.
[0248] 3. The device analyzes the video data using a generative model to detect suspicious activity.
[0249] 4. The server determines that the crime risk is high and raises an alert flag.
[0250] 5. The emotion engine analyzes the user's facial expressions and voice to identify states of surprise and tension.
[0251] 6. The device sends an alert to the user's smartphone as an emergency notification.
[0252] 7. The user checks the alert on their smartphone and checks live footage from the security camera.
[0253] 8. The user will notify the police if necessary.
[0254] Example 2: Abnormal sound detection and warning
[0255] 1. The server collects audio data from intercoms and audio sensors.
[0256] 2. The server preprocesses the audio data and splits it into parseable chunks.
[0257] 3. The device analyzes the audio data using a generative model to detect abnormal sounds, such as the sound of glass breaking.
[0258] 4. The server determines that the crime risk is high and raises an alert flag.
[0259] 5. The emotion engine analyzes the user's voice tone to identify excitement.
[0260] 6. The device will immediately send a warning notification to the user and automatically notify the police.
[0261] 7. Users can check alerts on their smartphones and monitor the situation at the site.
[0262] 8. The police will rush to the scene and take appropriate action.
[0263] This will enable quick and accurate crime prevention and response that takes into account the user's emotional state, further improving the safety of properties.
[0264] Example 2
[0265] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0266] In recent years, the importance of surveillance systems has increased. However, conventional systems have limited accuracy in analyzing crime risks and real-time warning notifications. Furthermore, they do not take the user's emotional state into account, making it difficult to respond quickly and appropriately. Furthermore, it is difficult to improve the learning accuracy of predictive models, and they may lack long-term reliability. This means that there is a lack of effective means to prevent crime risks before they occur.
[0267] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0268] In this invention, the server includes means for collecting security data from the monitoring devices, means for preprocessing the collected data, means for analyzing crime risk using a generative model based on the preprocessed data, means for analyzing user emotions in real time, means for sending a warning notification to the user or relevant organizations based on the analysis results, and means for feeding back data after the warning to the generative model for learning. This enables highly accurate analysis of crime risk and prompt notification that takes into account the user's emotional state, making it possible to provide an effective monitoring system that prevents crime risks before they occur.
[0269] A "surveillance device" is a hardware device for collecting security data, such as a surveillance camera, audio sensor, or motion sensor.
[0270] "Security data" refers to information used for crime risk analysis, such as video data, audio data, and motion data collected from surveillance devices.
[0271] "Preprocessing" refers to the process of converting security data into an analyzable format, specifically splitting video data into frames and audio data into chunks.
[0272] A "generative model" is a machine learning or deep learning model for analyzing crime risk based on security data.
[0273] "Crime risk" is an indicator that shows the degree to which a particular behavior or environment is likely to lead to criminal activity.
[0274] "Emotion analysis" is the process of analyzing facial expressions, vocal tone, and movement patterns to identify a user's emotional state.
[0275] A "warning notification" is an alert or notification that is sent when a criminal risk is determined to be high or based on a user's particular emotional state.
[0276] "Feedback" is the process of reusing post-warning data or newly collected data as training data for a generative model.
[0277] A "threshold" is a standard value used to determine crime risk, and anything above this value is considered high risk.
[0278] An "alert flag" is a signal or mark that is raised when the crime risk exceeds a threshold, and triggers a warning notification.
[0279] This invention is a system that analyzes security data collected from monitoring devices, analyzes crime risks in real time, and sends warning notifications to users and relevant organizations as needed. It also analyzes users' emotions to encourage more appropriate responses. Specific embodiments are described in detail below.
[0280] Hardware and Software Configuration
[0281] This system is mainly composed of a server, terminals, and users, and uses the following hardware and software.
[0282] server
[0283] The server has the following functions:
[0284] Collect security data in real time from surveillance devices (surveillance cameras, audio sensors, motion sensors).
[0285] Preprocess the collected data.
[0286] Conduct analysis to determine crime risk.
[0287] The generative model is trained by feeding back the results after the warning and new data.
[0288] The software used is OpenCV for preprocessing video data and PyDub for preprocessing audio data, and TensorFlow for inputting data into a generative model and performing analysis.
[0289] Terminal
[0290] The terminal has the following features:
[0291] The data received from the server is input into the generative model to analyze crime risk.
[0292] Analyze user emotions in real time using an emotion engine.
[0293] Send warning notifications to users and / or relevant authorities.
[0294] Specifically, it uses Microsoft Azure's Face API for facial recognition and emotion analysis, IBM Watson's Tone Analyzer for voice emotion analysis, and Twilio's API for SMS and phone notifications.
[0295] User
[0296] The user can:
[0297] Check the warning notification you receive and take appropriate action (e.g., check the warning on your smartphone and check live footage from your security camera).
[0298] Provide further instructions as needed (e.g., to call the police).
[0299] Specific examples
[0300] Example 1: Detecting and warning suspicious individuals
[0301] 1. The server collects video data from the surveillance cameras.
[0302] 2. The device analyzes the video data and detects suspicious activity.
[0303] 3. The server determines the risk is high and raises an alert flag.
[0304] 4. The emotion engine analyzes the user's facial expressions and voice to identify states of surprise and tension.
[0305] 5. The device sends an alert to the user's smartphone as an emergency notification.
[0306] 6. The user checks the alert on their smartphone and checks live footage from the security camera.
[0307] 7. The user will notify the police if necessary.
[0308] Example prompt sentence:
[0309] "Suspicious person detected. Please check live feed and call the police if necessary."
[0310] Example 2: Abnormal sound detection and warning
[0311] 1. The server collects audio data from intercoms and audio sensors.
[0312] 2. The device analyzes the audio data and detects abnormal sounds such as breaking glass.
[0313] 3. The server determines the risk is high and raises an alert flag.
[0314] 4. The emotion engine analyzes the user's voice tone and identifies their state of excitement.
[0315] 5. The device will immediately send a warning notification to the user and automatically notify the police.
[0316] 6. Users can check alerts on their smartphones and monitor the situation at the site.
[0317] 7. The police will rush to the scene and take appropriate action.
[0318] Example prompt sentence:
[0319] "An abnormal sound (sound of glass breaking) has been detected. Please check the situation on site."
[0320] Using the above method, the system can analyze crime risks with high accuracy and provide prompt notifications that take into account the user's emotional state, thereby preventing crime risks and further improving the safety of properties.
[0321] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0322] Step 1: Data collection
[0323] The server collects real-time security data from surveillance devices such as surveillance cameras, audio sensors, and motion sensors, and stores the collected data in a buffer.
[0324] Input: Video data from surveillance cameras, audio data from audio sensors, and motion data from motion sensors.
[0325] Processing: Connect to the monitoring device, acquire the data, and store it in a buffer.
[0326] Output: Buffered raw data (video, audio, motion data).
[0327] Specific operation: The server connects to a surveillance camera with a specific IP address and port number and streams video data using RTSP (Real-Time Streaming Protocol). It also sends an HTTP request to the audio sensor to obtain audio data.
[0328] Step 2: Preprocessing the data
[0329] The server pre-processes the collected data to convert it into a format that is easy to analyze: video data is divided into frames, and audio data is divided into analyzable chunks.
[0330] Input: Buffered raw data (video, audio, motion data).
[0331] Processing: Video data is split into frames and noise removal and edge detection are performed. Audio data is split into chunks based on sample rate and noise filters are applied.
[0332] Output: Video data divided into frames and audio data divided into chunks.
[0333] How it works: The server uses OpenCV to split the video data into frames and applies image processing such as histogram smoothing and edge detection to each frame. For audio data, it uses PyDub to split the data into chunks every second and apply a noise reduction filter.
[0334] Step 3: Data analysis
[0335] The device inputs the preprocessed data into a generative AI model to analyze crime risk and detects anomalies using image and voice recognition algorithms.
[0336] Input: Video data divided into frames and audio data divided into chunks.
[0337] Processing: Apply the YOLOv4 model to video data to identify suspicious individuals and abnormal behavior. Use the Google Cloud Speech-to-Text API on audio data to detect abnormal sounds.
[0338] Output: Crime risk analysis results (e.g., presence of suspicious individuals, identification of abnormal behavior, detection of abnormal sounds).
[0339] How it works: The device uses TensorFlow to feed video frames to a pre-trained YOLOv4 model to detect suspicious people and anomalous behavior, and the Google Cloud Speech-to-Text API to identify anomalous sounds.
[0340] Step 4: Determine crime risk
[0341] The server determines the crime risk based on the analysis results obtained from the generative AI model, calculates the risk score, and compares it with a threshold.
[0342] Input: Crime risk analysis results.
[0343] Processing: Calculate a risk value based on the score of the analysis results, and set an alert flag if the value exceeds the set threshold.
[0344] Output: Risk score and alert flag.
[0345] Specific operation: The server calculates a risk value based on the score of the analysis results, and if it exceeds a threshold (e.g., 0.7), it sets an alert flag.
[0346] Step 5: Sentiment Analysis
[0347] The device's built-in emotion engine analyzes the user's emotions in real time, identifying their emotional state from facial expressions, tone of voice, and movement patterns.
[0348] Input: User's video and audio data.
[0349] Processing: Apply Face API to video data to analyze facial expressions and identify emotions. Apply Tone Analyzer to audio data to analyze voice tone.
[0350] Output: The user's emotional state.
[0351] How it works: The device uses Microsoft Azure's Face API to analyze facial expressions and identify emotional states (e.g., surprise, anger, fear, etc.) and IBM Watson's Tone Analyzer to analyze voice emotions.
[0352] Step 6: Notifications and Alerts
[0353] The device sends notifications and warnings based on the crime risk assessment results and the results of user emotion analysis.
[0354] Inputs: Risk score, alert flag, sentiment analysis results.
[0355] Action: Generate notifications based on the risk score and emotional state and send them to the user or relevant authorities.
[0356] Output: Notifications and alerts (e.g. push notifications to smartphones, SMS, phone calls).
[0357] What it does: The device sends a push notification to the user's smartphone and, in some cases, automatically notifies the police or security company. SMS and phone notifications are also possible using Twilio's API.
[0358] Step 7: User interaction
[0359] Users can take appropriate action based on the warnings and notifications they receive.
[0360] Input: The warning message sent to your smartphone.
[0361] Action: Check live surveillance footage and provide additional instructions as needed.
[0362] Output: Confirmed security footage, further instructions (e.g., to call the police).
[0363] What happens: The user sees the alert on their smartphone, checks the live video feed, and, if necessary, calls the police.
[0364] Step 8: Building a feedback loop
[0365] The server collects the results and new data after the warning and feeds them back as training data for the generative model.
[0366] Input: Results after warning, newly collected data.
[0367] Processing: Store new data in the database and periodically retrain the generative model.
[0368] Output: An updated generative model.
[0369] What it does: The server saves the new data to a database and uses services like Azure Machine Learning to retrain the generative model and emotion engine.
[0370] (Application example 2)
[0371] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0372] Conventional security systems analyze crime risks and provide appropriate warnings, but the notification method and content are rarely flexibly adjusted according to the user's state. Furthermore, because they do not take the user's emotional state into account, the appropriate response that is actually required may be delayed. Furthermore, a feedback loop after the warning is not effectively established, resulting in insufficient continuous learning of the model, and there are issues with improving long-term accuracy.
[0373] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0374] In this invention, the server includes means for collecting security data from monitoring devices, means for preprocessing the collected data, means for analyzing crime risk using a generative model based on the preprocessed data, means for analyzing voice tone, facial expressions, movement patterns, etc. from the analyzed data to identify the user's emotional state, means for automatically adjusting the content and method of the warning based on the user's emotional state, and means for feeding back post-warning data to the generative model for learning. This enables flexible warning notifications that take the user's emotional state into consideration, improving the accuracy of crime prevention and response.
[0375] "Surveillance Devices" refers to devices used to remotely detect suspicious activity or unusual conditions, such as surveillance cameras, audio sensors, and motion sensors.
[0376] "Security data" refers to data that details surrounding conditions, such as video, audio, and motion data collected from surveillance devices.
[0377] "Preprocessing" refers to the process of preparing collected security data in a form that is easy to analyze, such as by dividing it into frames and removing noise.
[0378] A "generative model" refers to a machine learning or deep learning algorithm used to identify specific patterns or anomalies based on collected data.
[0379] "Means for analyzing crime risk" refers to a mechanism that uses a generative model to analyze crime risk, such as suspicious movements or abnormal sounds, and evaluates it as a score.
[0380] "Means for sending warning notifications" refers to a function that notifies users or relevant organizations of information regarding crime risks in real time.
[0381] "Means for identifying emotional states" refers to a system that analyzes a user's facial expressions, voice tone, and movement patterns to determine emotions such as joy, anger, surprise, and fear.
[0382] "Means for automatically adjusting the content and method of warnings" refers to a mechanism that dynamically changes the content and method of warnings sent depending on the user's emotional state.
[0383] "Feedback and learning" refers to the process of collecting post-warning results and new crime data and incorporating them into the generative model to continuously improve its accuracy.
[0384] "User" refers to any individual or entity that uses the System.
[0385] Overall overview
[0386] This invention combines a system that uses security data collected from surveillance devices to analyze crime risks using generative models and send real-time warnings, with an emotion engine that recognizes user emotions. The system is composed of a server, a terminal, and a user, and operates in cooperation with each other.
[0387] Hardware and Software
[0388] The hardware used is as follows:
[0389] High-performance smartphone
[0390] Network Camera
[0391] microphone
[0392] The following software is used:
[0393] Video analysis: OpenCV
[0394] Audio analysis: Librosa
[0395] Sentiment analysis: TensorFlow + Keras (emotion recognition model)
[0396] Notification system: Firebase Cloud Messaging
[0397] Data collection
[0398] The server collects data in real time from surveillance devices such as surveillance cameras, audio sensors, and motion sensors, and the collected data is temporarily stored in a buffer on the server.
[0399] Data Preprocessing
[0400] The server pre-processes the collected security data: video data is split into frames, and audio data is split into analyzable chunks based on sample rate. Image processing such as noise reduction and edge detection is also performed.
[0401] Data analysis
[0402] The device then inputs the preprocessed data into a generative model to perform a crime risk analysis. Image recognition algorithms are used to detect suspicious activity, and voice recognition algorithms are used to detect abnormal sounds. Motion patterns are also analyzed to detect unnatural movements.
[0403] Crime risk assessment
[0404] The server determines the crime risk based on the analysis results of the generative model. It calculates a risk score based on pre-set thresholds and rules, and sets an alert flag if the risk score exceeds a certain threshold.
[0405] Emotion analysis
[0406] The device's built-in emotion engine analyzes the user's emotions in real time, analyzing the user's facial expressions, tone of voice, and movement patterns from collected video and audio data to identify their emotional state.
[0407] Notifications and Alerts
[0408] The device sends notifications and warnings to the user and relevant authorities based on the crime risk assessment results and the user's emotional analysis results. The content and method of the notification change depending on the specific emotional state (e.g., tension, surprise, anger, etc.). For example, if emergency notification is set, the police and security companies will be automatically notified immediately.
[0409] User Support
[0410] Users can take appropriate action based on the warnings and notifications they receive, such as checking the content of the warning and checking live footage from surveillance cameras. Users can also issue more specific instructions based on the analysis results of the emotion engine.
[0411] Building a feedback loop
[0412] The server collects the results of the alerts and new crime data, and feeds it back into the generative model and emotion engine as training data. This feedback allows the model to continuously learn and improve its prediction accuracy.
[0413] Examples of prompt statements
[0414] Prompt: "Analyze surveillance camera footage, detect suspicious movements, and analyze emotions from the user's facial expressions. Also detect abnormal sounds (such as glass breaking). Send warning notifications to the user in real time and assess crime risk."
[0415] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0416] Step 1:
[0417] The server collects security data from surveillance devices in real time. Specifically, this includes video data from surveillance cameras, audio data from microphones, and movement data from motion sensors. The collected data is temporarily stored in a buffer on the server. The input is real-time data from the surveillance devices, and the output is the raw data stored in the buffer.
[0418] Step 2:
[0419] The server preprocesses the collected security data by splitting the video data into frames and the audio data into analyzable chunks based on the sample rate. It also performs image processing such as noise reduction and edge detection. The input is the raw data in the buffer, and the output is the preprocessed image and audio data.
[0420] Step 3:
[0421] The device inputs the preprocessed data into a generative model to perform a crime risk analysis. Specifically, it uses an image recognition algorithm to detect suspicious movements and a voice recognition algorithm to detect abnormal sounds. It also analyzes motion patterns to detect unnatural movements. The input is preprocessed image data and voice data, and the output is a crime risk score and analysis results.
[0422] Step 4:
[0423] The server determines the crime risk based on the analysis results of the generative model. Specifically, it calculates a risk score based on pre-set thresholds and rules, and sets an alert flag if the risk score exceeds a certain threshold. The inputs are the crime risk score and the analysis results, and the outputs are an alert flag and a risk assessment.
[0424] Step 5:
[0425] The device uses an emotion engine to analyze the user's emotions in real time based on the crime risk assessment results and the user's emotion analysis results. Specifically, it analyzes the user's facial expressions, voice tone, and movement patterns to identify their emotional state. The input is the user's facial expression data, voice tone, and movement patterns, and the output is the user's emotional state.
[0426] Step 6:
[0427] The device sends notifications and warnings to the user and relevant authorities based on the crime risk assessment results and the user's emotional analysis results. Specifically, the content and method of the notification are changed depending on the user's emotional state (e.g., tension, surprise, anger, etc.). For example, if emergency notification is set, the device will immediately and automatically notify the police or security company. The input is the crime risk assessment and the user's emotional state, and the output is a warning notification.
[0428] Step 7:
[0429] The user can take appropriate action based on the warnings and notifications they receive. Specifically, they can check the content of the warning and check live video footage from security cameras. The user can also issue more specific instructions based on the analysis results of the emotion engine. The input is the warning notification and live video footage, and the output is the appropriate response action.
[0430] Step 8:
[0431] The server collects the results of the warning and new crime data, and feeds it back as training data for the generative model and emotion engine. This feedback allows the model to continuously learn and improve its prediction accuracy. The input is the result data after the warning notification, and the output is an updated generative model and emotion engine.
[0432] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0433] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0434] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0435] [Second embodiment]
[0436] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0437] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0438] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0439] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0440] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0441] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0442] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0443] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0444] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0445] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0446] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0447] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0448] Overall overview
[0449] This system uses security data collected from surveillance devices to analyze crime risks using generative models and issue real-time warnings. It is composed of a server, terminals, and users, and operates in cooperation with each other.
[0450] Program processing
[0451] Data collection
[0452] The server collects data in real time from multiple surveillance devices (surveillance cameras, audio sensors, motion sensors, etc.) and has a buffer to temporarily store the data pulled from these devices.
[0453] Data Preprocessing
[0454] The server pre-processes the collected security data: video data is split into frames, audio data is split into analyzable chunks based on sample rate, motion sensor data is digitized, and image processing such as noise removal and edge detection is performed.
[0455] Data analysis
[0456] The device inputs the preprocessed data into a generative model and performs analysis. It uses image recognition algorithms to detect suspicious movements and voice recognition algorithms to detect abnormal sounds. It analyzes motion patterns and detects unnatural movements.
[0457] Crime risk assessment
[0458] The server then uses the analysis results to determine the risk of crime. This determination is based on pre-set thresholds and rules. For example, a risk score is calculated based on the identification pattern of a suspicious person or specific frequencies of a voice.
[0459] Notifications and Alerts
[0460] Based on the risk assessment results, the device will send notifications or warnings to the user or relevant authorities via push notifications to smartphones or tablets, email, SMS, etc. If necessary, an API will also be used to automatically notify the police or security companies.
[0461] User Support
[0462] Users can take appropriate action based on the warnings and notifications they receive, such as checking the warning on their smartphone, checking live footage from surveillance cameras, or contacting security companies or police to instruct them to investigate the scene.
[0463] Building a feedback loop
[0464] The server collects the results of the warnings and new crime data, and feeds them back as training data for the generative model. This feedback allows the model to continuously learn and improve its prediction accuracy.
[0465] Specific examples
[0466] Example 1: Detecting and warning suspicious individuals
[0467] 1. The server collects video data from the surveillance cameras.
[0468] 2. The device analyzes this video data using a generative model to detect suspicious activity.
[0469] 3. The server determines that the crime risk is high and raises an alert flag.
[0470] 4. The device sends a warning notification to the user's smartphone.
[0471] 5. The user checks the alert on their smartphone and checks the surveillance camera footage in real time.
[0472] 6. The user will notify the police if necessary.
[0473] Example 2: Abnormal sound detection and warning
[0474] 1. The server collects audio data from intercoms and audio sensors.
[0475] 2. The device analyzes the audio data and detects abnormal sounds such as breaking glass.
[0476] 3. The server determines that the risk of crime is high and sends a warning notice to the user while automatically reporting the incident to the police.
[0477] 4. The user checks the alerts on their smartphone and monitors the situation at the site.
[0478] 5. The police will rush to the scene and take appropriate action.
[0479] This will enable 24-hour security response and immediate crime prevention, improving the safety of the property.
[0480] The processing flow will be explained below.
[0481] Step 1:
[0482] The server collects real-time data from surveillance devices such as surveillance cameras, audio sensors, motion sensors, etc. Specifically, it periodically pulls data from each device (e.g., every second) and temporarily stores this data in a buffer on the server.
[0483] Step 2:
[0484] The server pre-processes the collected security data, splitting the video data into frames and the audio data into analyzable chunks based on sample rate, including applying image processing techniques such as noise removal and edge detection to shape the data.
[0485] Step 3:
[0486] The device inputs the preprocessed data into a generative model, which uses an image recognition algorithm to detect suspicious movements and a voice recognition algorithm to detect abnormal sounds. The device then analyzes the motion patterns to detect unnatural movements.
[0487] Step 4:
[0488] The server determines the crime risk based on the analysis results of the generative model. It calculates a risk score based on pre-set thresholds and rules, and if the risk score exceeds a certain threshold, it sets an alert flag. This result is recorded in a database.
[0489] Step 5:
[0490] If the device determines that there is a high risk of crime, it will send a warning to the user and relevant authorities via push notification to the smartphone or tablet, email or SMS, and, if necessary, automatically notify the police or security company using an API.
[0491] Step 6:
[0492] The user checks the warnings and notifications they receive and takes appropriate action. For example, they can check the warnings on their smartphones, check real-time footage from surveillance cameras, and, if necessary, contact a security company or police and instruct them to investigate the scene.
[0493] Step 7:
[0494] The server collects the results of warnings and new crime data and feeds it back to the generative model. This feedback allows the generative model to continuously learn and improve its prediction accuracy. The collected data is fed into the model as re-training data, and the parameters are adjusted.
[0495] Example 1
[0496] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0497] In recent years, the spread of security devices has progressed as crime risks have increased. However, conventional security systems have faced challenges in analyzing crime risks in real time and immediately sending appropriate warning notifications. Furthermore, there have been concerns about a decline in prediction accuracy due to insufficient updating of the training data for generative models. The objective of the present invention is to solve these challenges and provide a system that realizes highly accurate crime risk analysis and warning notifications in real time.
[0498] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0499] In this invention, the server includes means for collecting security data from monitoring devices, means for preprocessing the collected data, means for analyzing crime risk based on the preprocessed data using a generative model, means for sending a warning to a user or a relevant organization based on the analysis results, means for feeding back data after the warning to the generative model for learning, means for acquiring data in real time and providing a buffer for temporarily storing it, means for using a voice analysis algorithm to detect abnormal sounds, and means for using a cloud messaging service as a notification method. This makes it possible to analyze the data collected from the monitoring devices with high accuracy, determine crime risk in real time, and promptly send appropriate warnings.
[0500] A "surveillance device" is a device, such as a camera, audio sensor, or motion sensor, that collects information occurring within a specific space or environment.
[0501] "Security data" is a general term for information such as video, audio, and behavior patterns obtained from surveillance devices.
[0502] "Preprocessing" refers to the process of converting collected data into an analyzable form, and specifically includes dividing video data into frames and chunking audio data.
[0503] "Generative model" generally refers to machine learning and artificial intelligence algorithms used to predict and analyze crime risk from data.
[0504] "Crime risk" is a value that assesses the likelihood of suspicious activity or abnormal events occurring based on collected security data using certain criteria.
[0505] "Cloud Messaging Service" means an online service for sending push notifications over the Internet.
[0506] "Feedback" is the process of collecting results and new data after a warning and reusing them as training data for the analytical model.
[0507] "Buffer" refers to a temporary storage area for data, and is used to temporarily store data collected in real time.
[0508] An "audio analysis algorithm" is a mathematical method or technique for analyzing audio data to identify specific sounds or patterns.
[0509] "Notification" refers to the act of sending warnings or information to users or relevant organizations based on the analysis results.
[0510] As an embodiment of the invention, the system consists of three main components: a server, a terminal, and a user. The whole system is based on monitoring devices, where the collected data is sequentially pre-processed, analyzed, notified, and fed back.
[0511] Data collection
[0512] The server collects real-time data from multiple monitoring devices (e.g., cameras, audio sensors, motion sensors). This data is first temporarily stored in a buffer on an AWS EC2 instance. RTSP (Real-Time Streaming Protocol) and HTTP are used to obtain real-time data.
[0513] Data Preprocessing
[0514] The server splits the collected video data into frames using OpenCV, and splits the audio data into analyzable chunks using Librosa. It also performs preprocessing such as noise reduction and normalization. Motion sensor data is stored as numerical data after calibration.
[0515] Data analysis
[0516] The device then inputs the preprocessed data into generative AI models to perform analysis. For example, video data is analyzed using TensorFlow's YOLO model to detect suspicious activity, audio data is analyzed using the Google Cloud Speech-to-Text API to identify abnormal sounds, and motion data is analyzed using the SciPy library to detect unnatural movements.
[0517] Crime risk assessment
[0518] The server then uses the analysis results to determine the crime risk according to pre-set thresholds and rules. If the threshold is exceeded, a risk score is calculated and the person is deemed high risk based on that score.
[0519] Notifications and Alerts
[0520] Based on the server's assessment, the device sends a warning to the user or relevant authorities. Notifications are sent via Firebase Cloud Messaging, and emails and SMS are sent via the Twilio API. The system also includes an automatic notification function to the police and security companies using the API.
[0521] User Support
[0522] Users receive a warning notification sent from the device and check it on their smartphone app. The app provides a function to view live footage from surveillance cameras, allowing users to take appropriate action based on this information. If necessary, they can also report the incident to a security company or the police.
[0523] Building a feedback loop
[0524] The server continuously collects the results of warnings and newly acquired crime data, and uses them to retrain the generative model, thereby continuously improving the prediction accuracy of the generative AI model.
[0525] Specific examples
[0526] Example 1: Detecting and warning suspicious individuals
[0527] 1. The server collects video data from the surveillance cameras.
[0528] 2. The device analyzes this video data using a generative model to detect suspicious activity.
[0529] 3. The server determines that the crime risk is high and raises an alert flag.
[0530] 4. The device sends a warning notification to the user's smartphone.
[0531] 5. The user checks the alert on their smartphone and checks the surveillance camera footage in real time.
[0532] 6. The user will notify the police if necessary.
[0533] Example prompts to input to the generative AI model
[0534] "Please explain in natural language how your AI system analyzes collected data and generates a risk score when it detects suspicious activity or abnormal sounds, including specific libraries and tools."
[0535] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0536] Step 1: Data collection
[0537] The server collects data from surveillance devices. Specifically, the server acquires video data from surveillance cameras in real time using RTSP and temporarily stores it in a buffer on an AWS EC2 instance. It also collects audio data from audio sensors via HTTP and data from motion sensors via Bluetooth. It receives stream data from surveillance devices as input and outputs it to a temporary storage buffer.
[0538] Step 2: Preprocess the data
[0539] The server preprocesses the collected security data. Specifically, it uses OpenCV to split the video data into frames and Librosa to split the audio data into one-second chunks. Motion sensor data is digitized, denoised, and normalized. It receives raw data stored in a buffer as input and outputs data that can be converted into an analyzable format.
[0540] Step 3: Data analysis
[0541] The device inputs the preprocessed data into a generative model to perform analysis. Specifically, the device uses TensorFlow's YOLO model to analyze each frame and detect suspicious motion. It uses the Google Cloud Speech-to-Text API to analyze chunked audio data and identify abnormal sounds. It uses SciPy's signal processing functions for motion pattern analysis. It takes the preprocessed data as input and outputs specific analysis results.
[0542] Step 4: Determine crime risk
[0543] The server determines the crime risk based on the analysis results. Specifically, the server calculates a risk score based on thresholds and rules, and if it exceeds a certain value, it determines the risk as high. For example, it sets a score based on the frequency of suspicious movements or abnormal sounds. It receives the analysis results as input and outputs risk assessment data including a risk score.
[0544] Step 5: Notifications and warnings
[0545] The device sends a warning notification to the user or relevant authorities based on the risk assessment results. Specifically, the device sends a push notification to the smartphone via Firebase Cloud Messaging, sends an emergency email or SMS using the Twilio API, and automatically notifies the police or security company via the API. It receives risk assessment data as input and outputs a notification message.
[0546] Step 6: User interaction
[0547] The user can take appropriate action based on the warning notification they receive. Specifically, the user checks the warning on their smartphone and checks live footage from the surveillance camera within the app. If necessary, they can also notify a security company or the police. The app receives the notification message as input and outputs the response action to be taken.
[0548] Step 7: Building a feedback loop
[0549] The server collects the results of the warning and newly acquired crime data and feeds it back as training data for the generative model. Specifically, the server continuously monitors and collects new data and uses it to retrain the generative model, thereby improving the model's predictive accuracy. It receives the post-warning data as input and outputs it as retraining data.
[0550] (Application example 1)
[0551] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0552] In conventional security systems, when analyzing data collected from monitoring devices, it has been difficult to detect anomalies in real time, provide immediate notification, and automatically report to relevant authorities. This increases the possibility of delays in crime prevention and prompt response, posing a security issue. The present invention aims to solve these problems and provide a security system that can detect anomalies in real time and respond immediately.
[0553] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0554] In this invention, the server includes means for collecting security data from monitoring devices, means for preprocessing the collected data, means for analyzing crime risks using a generative model based on the preprocessed data, means for sending a warning to a user or relevant organizations based on the analysis results, means for feeding back data after the warning to the generative model for learning, means for sending a push notification when an abnormality is detected based on the analysis results, and means for automatically notifying relevant organizations when an abnormality is detected. This makes it possible to detect abnormalities in real time and to immediately notify / report them.
[0555] A "surveillance device" is a device, such as a camera or sensor, used to collect security data.
[0556] "Security data" refers to security-related information such as video, audio, and motion sensor data obtained through surveillance devices.
[0557] "Preprocessing" refers to the process of converting collected security data into a format suitable for analysis, and includes dividing image data into frames and audio data into chunks.
[0558] A "generative model" is an artificial intelligence model used to detect patterns and anomalies in collected data.
[0559] "Crime risk analysis" refers to the use of generative models to assess the likelihood of a crime occurring based on collected data.
[0560] "Sending a warning notification" means issuing a warning to users and relevant organizations when an abnormality is detected based on the analysis results.
[0561] "Feedback" refers to the process of feeding post-alert data and new security data back into the generative model to allow it to continuously learn.
[0562] "Push notification" is a function that immediately sends a warning message to a user's smartphone or other device when an abnormality is detected.
[0563] "Automatic reporting" is a function that automatically reports to relevant authorities (such as the police or security companies) when a crime risk or abnormality is detected.
[0564] System configuration
[0565] An embodiment of the present invention is a system that collects security data from monitoring devices, analyzes it using a generative AI model, and determines crime risk. This system is mainly composed of a server, a terminal, and a user, and each element works in cooperation with each other.
[0566] Program processing overview
[0567] The server collects data in real time from multiple monitoring devices (e.g., surveillance cameras, audio sensors, motion sensors, etc.) and has a buffer that temporarily stores it. The collected data is divided into frames, and the image data is divided into analyzable chunks based on the sample rate. The motion sensor data is also digitized and processed for noise removal, edge detection, etc.
[0568] The device inputs the preprocessed data into a generative model and performs analysis to determine the crime risk. This analysis includes detecting suspicious activity using image recognition algorithms, detecting abnormal sounds using voice recognition algorithms, and analyzing motion patterns. Based on the analysis results, the device evaluates the crime risk and calculates a risk score.
[0569] The server determines the crime risk based on the analysis results, and if the risk exceeds a set threshold, it sends a push notification to the user's smartphone or tablet, and automatically reports the situation to relevant authorities, such as the police or security companies, in real time.
[0570] Additionally, post-alert data and new security data are fed back into the generative model, allowing it to continuously learn and improve its predictions. This feedback loop allows the system to make increasingly accurate predictions over time.
[0571] Hardware and software used
[0572] Hardware:
[0573] Smartphone: Used by users to receive warning notifications and check surveillance footage in the event of an abnormality.
[0574] Surveillance cameras: Used to collect security data.
[0575] Sensors: Used to collect additional security data, such as sound and motion.
[0576] software:
[0577] Keras (keras.models, keras.preprocessing): Used to analyze video data using deep learning models.
[0578] OpenCV (cv2): Used to capture and pre-process video data.
[0579] Gmail SMTP (smtplib, email): Email sending system used for automated notifications.
[0580] Plyer (notification): Used to send push notifications to smartphones.
[0581] Specific examples
[0582] For example, a surveillance camera can detect suspicious activity and analyze the video data using a Keras model. If the result indicates a high risk of crime, the Plryer library can be used to send a push notification to the user's smartphone and automatically notify relevant authorities via Gmail SMTP. Similarly, if an audio sensor detects the sound of glass breaking, an immediate notification and report can be sent.
[0583] Prompt Sentence Examples
[0584] "Please create an application that analyzes surveillance camera footage in a specified area, detects suspicious individuals and abnormal sounds in real time, and sends push notifications and automatic reporting."
[0585] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0586] Step 1:
[0587] The server collects security data in real time from multiple monitoring devices (such as surveillance cameras, audio sensors, and motion sensors). The collected data is temporarily stored in a buffer. The input is real-time data from the monitoring devices, and the output is the security data stored in the buffer. Specifically, data is pulled from each monitoring device and stored in the server's buffer area.
[0588] Step 2:
[0589] The server preprocesses the collected security data. Image data is divided into frames, and audio data is divided into analyzable chunks based on the sample rate. Motion sensor data is also digitized and processed for noise removal and edge detection. The input is the buffered security data, and the output is the preprocessed data. Specifically, OpenCV is used to divide the video data into frames, and the audio data into chunks based on the sample rate.
[0590] Step 3:
[0591] The device inputs the preprocessed data into a generative AI model to analyze crime risk. This analysis includes detecting suspicious activity using an image recognition algorithm, detecting abnormal sounds using a voice recognition algorithm, and analyzing motion patterns. The input is the preprocessed data, and the output is the analysis results. Specifically, Keras is used to load the generative model, and the preprocessed data is input to the model to perform the analysis.
[0592] Step 4:
[0593] The server determines the crime risk based on the analysis results received from the terminal. If the crime risk is determined to be high, it sets an alert flag based on the set threshold. The input is the analysis result, and the output is the status of the alert flag. Specifically, it determines whether the threshold is exceeded based on the analysis result, and if so, sets an alert flag in the internal data structure.
[0594] Step 5:
[0595] When an alert flag is raised, the server sends a push notification to the user's smartphone or tablet. Furthermore, if necessary, it automatically notifies the police or security company. The input is the alert flag status and analysis results, and the output is the warning notification and notification results. Specifically, it uses the Plryer library to send notifications to smartphones, and uses Gmail SMTP to notify the relevant authorities via email.
[0596] Step 6:
[0597] Based on the received warning notification, the user checks the live footage from the surveillance camera and contacts the police or security company if necessary. The input is the content of the warning notification, and the output is the confirmation result of the abnormal situation and the countermeasures. Specifically, the user checks the received notification on their smartphone and displays the real-time footage through the application.
[0598] Step 7:
[0599] The server collects post-warning data and new security data and feeds it back as training data for the generative model. This allows the generative model to continuously learn and improve its prediction accuracy. The input is the post-warning data and new security data, and the output is an updated generative model. Specifically, the collected data is added to the dataset and the model is retrained using Keras.
[0600] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0601] Overall overview
[0602] This invention combines a system that uses security data collected from surveillance devices to analyze crime risks using generative models and send real-time warnings, with an emotion engine that recognizes user emotions. The system is composed of a server, a terminal, and a user, and operates in cooperation with each other.
[0603] Program processing
[0604] Data collection
[0605] The server collects data in real time from surveillance devices such as surveillance cameras, audio sensors, and motion sensors, and the collected data is temporarily stored in a buffer on the server.
[0606] Data Preprocessing
[0607] The server pre-processes the collected security data: video data is split into frames, and audio data is split into analyzable chunks based on sample rate. Image processing such as noise reduction and edge detection is also performed.
[0608] Data analysis
[0609] The device then inputs the preprocessed data into a generative model to perform a crime risk analysis. Image recognition algorithms are used to detect suspicious activity, and voice recognition algorithms are used to detect abnormal sounds. Motion patterns are also analyzed to detect unnatural movements.
[0610] Crime risk assessment
[0611] The server determines the crime risk based on the analysis results of the generative model. It calculates a risk score based on pre-set thresholds and rules, and sets an alert flag if the risk score exceeds a certain threshold.
[0612] Emotion analysis
[0613] The device's built-in emotion engine analyzes the user's emotions in real time, analyzing the user's facial expressions, tone of voice, and movement patterns from collected video and audio data to identify their emotional state.
[0614] Notifications and Alerts
[0615] The device sends notifications and warnings to the user and relevant authorities based on the crime risk assessment results and the user's emotional analysis results. The content and method of the notification change depending on the specific emotional state (e.g., tension, surprise, anger, etc.). For example, if emergency notification is set, the police and security companies will be automatically notified immediately.
[0616] User Support
[0617] Users can take appropriate action based on the warnings and notifications they receive, such as checking the content of the warning and checking live footage from surveillance cameras. Users can also issue more specific instructions based on the analysis results of the emotion engine.
[0618] Building a feedback loop
[0619] The server collects the results of the alerts and new crime data, and feeds it back into the generative model and emotion engine as training data. This feedback allows the model to continuously learn and improve its prediction accuracy.
[0620] Specific examples
[0621] Example 1: Detecting and warning suspicious individuals
[0622] 1. The server collects video data from the surveillance cameras.
[0623] 2. The device analyzes the video data using a generative model to detect suspicious activity.
[0624] 3. The server determines the risk is high and raises an alert flag.
[0625] 4. The emotion engine analyzes the user's facial expressions and voice to identify states of surprise and tension.
[0626] 5. The device sends an alert to the user's smartphone as an emergency notification.
[0627] 6. The user checks the alert on their smartphone and checks live footage from the security camera.
[0628] 7. The user will notify the police if necessary.
[0629] Example 2: Abnormal sound detection and warning
[0630] 1. The server collects audio data from intercoms and audio sensors.
[0631] 2. The device analyzes the audio data using a generative model to detect abnormal sounds, such as the sound of glass breaking.
[0632] 3. The server determines the risk is high and raises an alert flag.
[0633] 4. The emotion engine analyzes the user's voice tone and identifies their state of excitement.
[0634] 5. The device will immediately send a warning notification to the user and automatically notify the police.
[0635] 6. Users can check alerts on their smartphones and monitor the situation at the site.
[0636] 7. The police will rush to the scene and take appropriate action.
[0637] This will enable quick and accurate crime prevention and response that takes into account the user's emotional state, further improving the safety of properties.
[0638] The processing flow will be explained below.
[0639] Step 1:
[0640] The server collects data in real time from monitoring devices such as surveillance cameras, audio sensors, motion sensors, etc. Specifically, it periodically acquires data from each device (e.g., every second) and temporarily stores this data in a buffer on the server.
[0641] Step 2:
[0642] The server preprocesses the collected security data, splitting the video data into frames and the audio data into analyzable chunks based on sample rate, and applying image processing techniques such as noise reduction and edge detection to shape the data.
[0643] Step 3:
[0644] The device then inputs the preprocessed data into a generative model to perform a crime risk analysis. Specifically, it uses image recognition algorithms to detect suspicious activity and voice recognition algorithms to detect abnormal sounds. It also analyzes motion patterns to identify unnatural movements.
[0645] Step 4:
[0646] The server determines the crime risk based on the analysis results of the generative model. It calculates a risk score based on pre-set thresholds and rules, and if this risk score exceeds a certain threshold, it sets an alert flag. This determination result is recorded in a database.
[0647] Step 5:
[0648] The device's built-in emotion engine analyzes the user's emotions in real time by analyzing the user's facial expressions, tone of voice, and movement patterns from collected video and audio data to identify their emotional state.
[0649] Step 6:
[0650] The device sends notifications and warnings to the user and relevant authorities based on the crime risk assessment results and the user's emotional analysis results. The content and method of notifications can be changed depending on the user's emotional state, such as tension or surprise. For example, in an emergency, the device can immediately and automatically notify the police or security company.
[0651] Step 7:
[0652] The user checks the warnings and notifications they receive and takes appropriate action. Specifically, they check the warnings on their smartphones, check live footage from surveillance cameras, and, based on the analysis results of the emotion engine, contact security companies and police and instruct them to investigate the scene.
[0653] Step 8:
[0654] The server collects the results of warnings and new crime data, and feeds it back as training data for the generative model and emotion engine. This feedback allows the model to continuously learn and improve its prediction accuracy. The collected data is fed into the model as re-training data, and the parameters are adjusted.
[0655] Specific examples
[0656] Example 1: Detecting and warning suspicious individuals
[0657] 1. The server collects video data from the surveillance cameras.
[0658] 2. The server preprocesses the video data and divides it into frames.
[0659] 3. The device analyzes the video data using a generative model to detect suspicious activity.
[0660] 4. The server determines that the crime risk is high and raises an alert flag.
[0661] 5. The emotion engine analyzes the user's facial expressions and voice to identify states of surprise and tension.
[0662] 6. The device sends an alert to the user's smartphone as an emergency notification.
[0663] 7. The user checks the alert on their smartphone and checks live footage from the security camera.
[0664] 8. The user will notify the police if necessary.
[0665] Example 2: Abnormal sound detection and warning
[0666] 1. The server collects audio data from intercoms and audio sensors.
[0667] 2. The server preprocesses the audio data and splits it into parseable chunks.
[0668] 3. The device analyzes the audio data using a generative model to detect abnormal sounds, such as the sound of glass breaking.
[0669] 4. The server determines that the crime risk is high and raises an alert flag.
[0670] 5. The emotion engine analyzes the user's voice tone to identify excitement.
[0671] 6. The device will immediately send a warning notification to the user and automatically notify the police.
[0672] 7. Users can check alerts on their smartphones and monitor the situation at the site.
[0673] 8. The police will rush to the scene and take appropriate action.
[0674] This will enable quick and accurate crime prevention and response that takes into account the user's emotional state, further improving the safety of properties.
[0675] Example 2
[0676] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0677] In recent years, the importance of surveillance systems has increased. However, conventional systems have limited accuracy in analyzing crime risks and real-time warning notifications. Furthermore, they do not take the user's emotional state into account, making it difficult to respond quickly and appropriately. Furthermore, it is difficult to improve the learning accuracy of predictive models, and they may lack long-term reliability. This means that there is a lack of effective means to prevent crime risks before they occur.
[0678] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0679] In this invention, the server includes means for collecting security data from the monitoring devices, means for preprocessing the collected data, means for analyzing crime risk using a generative model based on the preprocessed data, means for analyzing user emotions in real time, means for sending a warning notification to the user or relevant organizations based on the analysis results, and means for feeding back data after the warning to the generative model for learning. This enables highly accurate analysis of crime risk and prompt notification that takes into account the user's emotional state, making it possible to provide an effective monitoring system that prevents crime risks before they occur.
[0680] A "surveillance device" is a hardware device for collecting security data, such as a surveillance camera, audio sensor, or motion sensor.
[0681] "Security data" refers to information used for crime risk analysis, such as video data, audio data, and motion data collected from surveillance devices.
[0682] "Preprocessing" refers to the process of converting security data into an analyzable format, specifically splitting video data into frames and audio data into chunks.
[0683] A "generative model" is a machine learning or deep learning model for analyzing crime risk based on security data.
[0684] "Crime risk" is an indicator that shows the degree to which a particular behavior or environment is likely to lead to criminal activity.
[0685] "Emotion analysis" is the process of analyzing facial expressions, vocal tone, and movement patterns to identify a user's emotional state.
[0686] A "warning notification" is an alert or notification that is sent when a criminal risk is determined to be high or based on a user's particular emotional state.
[0687] "Feedback" is the process of reusing post-warning data or newly collected data as training data for a generative model.
[0688] A "threshold" is a standard value used to determine crime risk, and anything above this value is considered high risk.
[0689] An "alert flag" is a signal or mark that is raised when the crime risk exceeds a threshold, and triggers a warning notification.
[0690] This invention is a system that analyzes security data collected from monitoring devices, analyzes crime risks in real time, and sends warning notifications to users and relevant organizations as needed. It also analyzes users' emotions to encourage more appropriate responses. Specific embodiments are described in detail below.
[0691] Hardware and Software Configuration
[0692] This system is mainly composed of a server, terminals, and users, and uses the following hardware and software.
[0693] server
[0694] The server has the following functions:
[0695] Collect security data in real time from surveillance devices (surveillance cameras, audio sensors, motion sensors).
[0696] Preprocess the collected data.
[0697] Conduct analysis to determine crime risk.
[0698] The generative model is trained by feeding back the results after the warning and new data.
[0699] The software used is OpenCV for preprocessing video data and PyDub for preprocessing audio data, and TensorFlow for inputting data into a generative model and performing analysis.
[0700] Terminal
[0701] The terminal has the following features:
[0702] The data received from the server is input into the generative model to analyze crime risk.
[0703] Analyze user emotions in real time using an emotion engine.
[0704] Send warning notifications to users and / or relevant authorities.
[0705] Specifically, it uses Microsoft Azure's Face API for facial recognition and emotion analysis, IBM Watson's Tone Analyzer for voice emotion analysis, and Twilio's API for SMS and phone notifications.
[0706] User
[0707] The user can:
[0708] Check the warning notification you receive and take appropriate action (e.g., check the warning on your smartphone and check live footage from your security camera).
[0709] Provide further instructions as needed (e.g., to call the police).
[0710] Specific examples
[0711] Example 1: Detecting and warning suspicious individuals
[0712] 1. The server collects video data from the surveillance cameras.
[0713] 2. The device analyzes the video data and detects suspicious activity.
[0714] 3. The server determines the risk is high and raises an alert flag.
[0715] 4. The emotion engine analyzes the user's facial expressions and voice to identify states of surprise and tension.
[0716] 5. The device sends an alert to the user's smartphone as an emergency notification.
[0717] 6. The user checks the alert on their smartphone and checks live footage from the security camera.
[0718] 7. The user will notify the police if necessary.
[0719] Example prompt sentence:
[0720] "Suspicious person detected. Please check live feed and call the police if necessary."
[0721] Example 2: Abnormal sound detection and warning
[0722] 1. The server collects audio data from intercoms and audio sensors.
[0723] 2. The device analyzes the audio data and detects abnormal sounds such as breaking glass.
[0724] 3. The server determines the risk is high and raises an alert flag.
[0725] 4. The emotion engine analyzes the user's voice tone and identifies their state of excitement.
[0726] 5. The device will immediately send a warning notification to the user and automatically notify the police.
[0727] 6. Users can check alerts on their smartphones and monitor the situation at the site.
[0728] 7. The police will rush to the scene and take appropriate action.
[0729] Example prompt sentence:
[0730] "An abnormal sound (sound of glass breaking) has been detected. Please check the situation on site."
[0731] Using the above method, the system can analyze crime risks with high accuracy and provide prompt notifications that take into account the user's emotional state, thereby preventing crime risks and further improving the safety of properties.
[0732] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0733] Step 1: Data collection
[0734] The server collects real-time security data from surveillance devices such as surveillance cameras, audio sensors, and motion sensors, and stores the collected data in a buffer.
[0735] Input: Video data from surveillance cameras, audio data from audio sensors, and motion data from motion sensors.
[0736] Processing: Connect to the monitoring device, acquire the data, and store it in a buffer.
[0737] Output: Buffered raw data (video, audio, motion data).
[0738] Specific operation: The server connects to a surveillance camera with a specific IP address and port number and streams video data using RTSP (Real-Time Streaming Protocol). It also sends an HTTP request to the audio sensor to obtain audio data.
[0739] Step 2: Preprocessing the data
[0740] The server pre-processes the collected data to convert it into a format that is easy to analyze: video data is divided into frames, and audio data is divided into analyzable chunks.
[0741] Input: Buffered raw data (video, audio, motion data).
[0742] Processing: Video data is split into frames and noise removal and edge detection are performed. Audio data is split into chunks based on sample rate and noise filters are applied.
[0743] Output: Video data divided into frames and audio data divided into chunks.
[0744] How it works: The server uses OpenCV to split the video data into frames and applies image processing such as histogram smoothing and edge detection to each frame. For audio data, it uses PyDub to split the data into chunks every second and apply a noise reduction filter.
[0745] Step 3: Data analysis
[0746] The device inputs the preprocessed data into a generative AI model to analyze crime risk and detects anomalies using image and voice recognition algorithms.
[0747] Input: Video data divided into frames and audio data divided into chunks.
[0748] Processing: Apply the YOLOv4 model to video data to identify suspicious individuals and abnormal behavior. Use the Google Cloud Speech-to-Text API on audio data to detect abnormal sounds.
[0749] Output: Crime risk analysis results (e.g., presence of suspicious individuals, identification of abnormal behavior, detection of abnormal sounds).
[0750] How it works: The device uses TensorFlow to feed video frames to a pre-trained YOLOv4 model to detect suspicious people and anomalous behavior, and the Google Cloud Speech-to-Text API to identify anomalous sounds.
[0751] Step 4: Determine crime risk
[0752] The server determines the crime risk based on the analysis results obtained from the generative AI model, calculates the risk score, and compares it with a threshold.
[0753] Input: Crime risk analysis results.
[0754] Processing: Calculate a risk value based on the score of the analysis results, and set an alert flag if the value exceeds the set threshold.
[0755] Output: Risk score and alert flag.
[0756] Specific operation: The server calculates a risk value based on the score of the analysis results, and if it exceeds a threshold (e.g., 0.7), it sets an alert flag.
[0757] Step 5: Sentiment Analysis
[0758] The device's built-in emotion engine analyzes the user's emotions in real time, identifying their emotional state from facial expressions, tone of voice, and movement patterns.
[0759] Input: User's video and audio data.
[0760] Processing: Apply Face API to video data to analyze facial expressions and identify emotions. Apply Tone Analyzer to audio data to analyze voice tone.
[0761] Output: The user's emotional state.
[0762] How it works: The device uses Microsoft Azure's Face API to analyze facial expressions and identify emotional states (e.g., surprise, anger, fear, etc.) and IBM Watson's Tone Analyzer to analyze voice emotions.
[0763] Step 6: Notifications and Alerts
[0764] The device sends notifications and warnings based on the crime risk assessment results and the results of user emotion analysis.
[0765] Inputs: Risk score, alert flag, sentiment analysis results.
[0766] Action: Generate notifications based on the risk score and emotional state and send them to the user or relevant authorities.
[0767] Output: Notifications and alerts (e.g. push notifications to smartphones, SMS, phone calls).
[0768] What it does: The device sends a push notification to the user's smartphone and, in some cases, automatically notifies the police or security company. SMS and phone notifications are also possible using Twilio's API.
[0769] Step 7: User interaction
[0770] Users can take appropriate action based on the warnings and notifications they receive.
[0771] Input: The warning message sent to your smartphone.
[0772] Action: Check live surveillance footage and provide additional instructions as needed.
[0773] Output: Confirmed security footage, further instructions (e.g., to call the police).
[0774] What happens: The user sees the alert on their smartphone, checks the live video feed, and, if necessary, calls the police.
[0775] Step 8: Building a feedback loop
[0776] The server collects the results and new data after the warning and feeds them back as training data for the generative model.
[0777] Input: Results after warning, newly collected data.
[0778] Processing: Store new data in the database and periodically retrain the generative model.
[0779] Output: An updated generative model.
[0780] What it does: The server saves the new data to a database and uses services like Azure Machine Learning to retrain the generative model and emotion engine.
[0781] (Application example 2)
[0782] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0783] Conventional security systems analyze crime risks and provide appropriate warnings, but the notification method and content are rarely flexibly adjusted according to the user's state. Furthermore, because they do not take the user's emotional state into account, the appropriate response that is actually required may be delayed. Furthermore, a feedback loop after the warning is not effectively established, resulting in insufficient continuous learning of the model, and there are issues with improving long-term accuracy.
[0784] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0785] In this invention, the server includes means for collecting security data from monitoring devices, means for preprocessing the collected data, means for analyzing crime risk using a generative model based on the preprocessed data, means for analyzing voice tone, facial expressions, movement patterns, etc. from the analyzed data to identify the user's emotional state, means for automatically adjusting the content and method of the warning based on the user's emotional state, and means for feeding back post-warning data to the generative model for learning. This enables flexible warning notifications that take the user's emotional state into consideration, improving the accuracy of crime prevention and response.
[0786] "Surveillance Devices" refers to devices used to remotely detect suspicious activity or unusual conditions, such as surveillance cameras, audio sensors, and motion sensors.
[0787] "Security data" refers to data that details surrounding conditions, such as video, audio, and motion data collected from surveillance devices.
[0788] "Preprocessing" refers to the process of preparing collected security data in a form that is easy to analyze, such as by dividing it into frames and removing noise.
[0789] A "generative model" refers to a machine learning or deep learning algorithm used to identify specific patterns or anomalies based on collected data.
[0790] "Means for analyzing crime risk" refers to a mechanism that uses a generative model to analyze crime risk, such as suspicious movements or abnormal sounds, and evaluates it as a score.
[0791] "Means for sending warning notifications" refers to a function that notifies users or relevant organizations of information regarding crime risks in real time.
[0792] "Means for identifying emotional states" refers to a system that analyzes a user's facial expressions, voice tone, and movement patterns to determine emotions such as joy, anger, surprise, and fear.
[0793] "Means for automatically adjusting the content and method of warnings" refers to a mechanism that dynamically changes the content and method of warnings sent depending on the user's emotional state.
[0794] "Feedback and learning" refers to the process of collecting post-warning results and new crime data and incorporating them into the generative model to continuously improve its accuracy.
[0795] "User" refers to any individual or entity that uses the System.
[0796] Overall overview
[0797] This invention combines a system that uses security data collected from surveillance devices to analyze crime risks using generative models and send real-time warnings, with an emotion engine that recognizes user emotions. The system is composed of a server, a terminal, and a user, and operates in cooperation with each other.
[0798] Hardware and Software
[0799] The hardware used is as follows:
[0800] High-performance smartphone
[0801] Network Camera
[0802] microphone
[0803] The following software is used:
[0804] Video analysis: OpenCV
[0805] Audio analysis: Librosa
[0806] Sentiment analysis: TensorFlow + Keras (emotion recognition model)
[0807] Notification system: Firebase Cloud Messaging
[0808] Data collection
[0809] The server collects data in real time from surveillance devices such as surveillance cameras, audio sensors, and motion sensors, and the collected data is temporarily stored in a buffer on the server.
[0810] Data Preprocessing
[0811] The server pre-processes the collected security data: video data is split into frames, and audio data is split into analyzable chunks based on sample rate. Image processing such as noise reduction and edge detection is also performed.
[0812] Data analysis
[0813] The device then inputs the preprocessed data into a generative model to perform a crime risk analysis. Image recognition algorithms are used to detect suspicious activity, and voice recognition algorithms are used to detect abnormal sounds. Motion patterns are also analyzed to detect unnatural movements.
[0814] Crime risk assessment
[0815] The server determines the crime risk based on the analysis results of the generative model. It calculates a risk score based on pre-set thresholds and rules, and sets an alert flag if the risk score exceeds a certain threshold.
[0816] Emotion analysis
[0817] The device's built-in emotion engine analyzes the user's emotions in real time, analyzing the user's facial expressions, tone of voice, and movement patterns from collected video and audio data to identify their emotional state.
[0818] Notifications and Alerts
[0819] The device sends notifications and warnings to the user and relevant authorities based on the crime risk assessment results and the user's emotional analysis results. The content and method of the notification change depending on the specific emotional state (e.g., tension, surprise, anger, etc.). For example, if emergency notification is set, the police and security companies will be automatically notified immediately.
[0820] User Support
[0821] Users can take appropriate action based on the warnings and notifications they receive, such as checking the content of the warning and checking live footage from surveillance cameras. Users can also issue more specific instructions based on the analysis results of the emotion engine.
[0822] Building a feedback loop
[0823] The server collects the results of the alerts and new crime data, and feeds it back into the generative model and emotion engine as training data. This feedback allows the model to continuously learn and improve its prediction accuracy.
[0824] Examples of prompt statements
[0825] Prompt: "Analyze surveillance camera footage, detect suspicious movements, and analyze emotions from the user's facial expressions. Also detect abnormal sounds (such as glass breaking). Send warning notifications to the user in real time and assess crime risk."
[0826] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0827] Step 1:
[0828] The server collects security data from surveillance devices in real time. Specifically, this includes video data from surveillance cameras, audio data from microphones, and movement data from motion sensors. The collected data is temporarily stored in a buffer on the server. The input is real-time data from the surveillance devices, and the output is the raw data stored in the buffer.
[0829] Step 2:
[0830] The server preprocesses the collected security data by splitting the video data into frames and the audio data into analyzable chunks based on the sample rate. It also performs image processing such as noise reduction and edge detection. The input is the raw data in the buffer, and the output is the preprocessed image and audio data.
[0831] Step 3:
[0832] The device inputs the preprocessed data into a generative model to perform a crime risk analysis. Specifically, it uses an image recognition algorithm to detect suspicious movements and a voice recognition algorithm to detect abnormal sounds. It also analyzes motion patterns to detect unnatural movements. The input is preprocessed image data and voice data, and the output is a crime risk score and analysis results.
[0833] Step 4:
[0834] The server determines the crime risk based on the analysis results of the generative model. Specifically, it calculates a risk score based on pre-set thresholds and rules, and sets an alert flag if the risk score exceeds a certain threshold. The inputs are the crime risk score and the analysis results, and the outputs are an alert flag and a risk assessment.
[0835] Step 5:
[0836] The device uses an emotion engine to analyze the user's emotions in real time based on the crime risk assessment results and the user's emotion analysis results. Specifically, it analyzes the user's facial expressions, voice tone, and movement patterns to identify their emotional state. The input is the user's facial expression data, voice tone, and movement patterns, and the output is the user's emotional state.
[0837] Step 6:
[0838] The device sends notifications and warnings to the user and relevant authorities based on the crime risk assessment results and the user's emotional analysis results. Specifically, the content and method of the notification are changed depending on the user's emotional state (e.g., tension, surprise, anger, etc.). For example, if emergency notification is set, the device will immediately and automatically notify the police or security company. The input is the crime risk assessment and the user's emotional state, and the output is a warning notification.
[0839] Step 7:
[0840] The user can take appropriate action based on the warnings and notifications they receive. Specifically, they can check the content of the warning and check live video footage from security cameras. The user can also issue more specific instructions based on the analysis results of the emotion engine. The input is the warning notification and live video footage, and the output is the appropriate response action.
[0841] Step 8:
[0842] The server collects the results of the warning and new crime data, and feeds it back as training data for the generative model and emotion engine. This feedback allows the model to continuously learn and improve its prediction accuracy. The input is the result data after the warning notification, and the output is an updated generative model and emotion engine.
[0843] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0844] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0845] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0846] [Third embodiment]
[0847] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0848] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0849] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0850] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0851] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0852] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0853] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0854] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0855] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0856] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0857] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0858] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0859] Overall overview
[0860] This system uses security data collected from surveillance devices to analyze crime risks using generative models and issue real-time warnings. It is composed of a server, terminals, and users, and operates in cooperation with each other.
[0861] Program processing
[0862] Data collection
[0863] The server collects data in real time from multiple surveillance devices (surveillance cameras, audio sensors, motion sensors, etc.) and has a buffer to temporarily store the data pulled from these devices.
[0864] Data Preprocessing
[0865] The server pre-processes the collected security data: video data is split into frames, audio data is split into analyzable chunks based on sample rate, motion sensor data is digitized, and image processing such as noise removal and edge detection is performed.
[0866] Data analysis
[0867] The device inputs the preprocessed data into a generative model and performs analysis. It uses image recognition algorithms to detect suspicious movements and voice recognition algorithms to detect abnormal sounds. It analyzes motion patterns and detects unnatural movements.
[0868] Crime risk assessment
[0869] The server then uses the analysis results to determine the risk of crime. This determination is based on pre-set thresholds and rules. For example, a risk score is calculated based on the identification pattern of a suspicious person or specific frequencies of a voice.
[0870] Notifications and Alerts
[0871] Based on the risk assessment results, the device will send notifications or warnings to the user or relevant authorities via push notifications to smartphones or tablets, email, SMS, etc. If necessary, an API will also be used to automatically notify the police or security companies.
[0872] User Support
[0873] Users can take appropriate action based on the warnings and notifications they receive, such as checking the warning on their smartphone, checking live footage from surveillance cameras, or contacting security companies or police to instruct them to investigate the scene.
[0874] Building a feedback loop
[0875] The server collects the results of the warnings and new crime data, and feeds them back as training data for the generative model. This feedback allows the model to continuously learn and improve its prediction accuracy.
[0876] Specific examples
[0877] Example 1: Detecting and warning suspicious individuals
[0878] 1. The server collects video data from the surveillance cameras.
[0879] 2. The device analyzes this video data using a generative model to detect suspicious activity.
[0880] 3. The server determines that the crime risk is high and raises an alert flag.
[0881] 4. The device sends a warning notification to the user's smartphone.
[0882] 5. The user checks the alert on their smartphone and checks the surveillance camera footage in real time.
[0883] 6. The user will notify the police if necessary.
[0884] Example 2: Abnormal sound detection and warning
[0885] 1. The server collects audio data from intercoms and audio sensors.
[0886] 2. The device analyzes the audio data and detects abnormal sounds such as breaking glass.
[0887] 3. The server determines that the risk of crime is high and sends a warning notice to the user while automatically reporting the incident to the police.
[0888] 4. The user checks the alerts on their smartphone and monitors the situation at the site.
[0889] 5. The police will rush to the scene and take appropriate action.
[0890] This will enable 24-hour security response and immediate crime prevention, improving the safety of the property.
[0891] The processing flow will be explained below.
[0892] Step 1:
[0893] The server collects real-time data from surveillance devices such as surveillance cameras, audio sensors, motion sensors, etc. Specifically, it periodically pulls data from each device (e.g., every second) and temporarily stores this data in a buffer on the server.
[0894] Step 2:
[0895] The server pre-processes the collected security data, splitting the video data into frames and the audio data into analyzable chunks based on sample rate, including applying image processing techniques such as noise removal and edge detection to shape the data.
[0896] Step 3:
[0897] The device inputs the preprocessed data into a generative model, which uses an image recognition algorithm to detect suspicious movements and a voice recognition algorithm to detect abnormal sounds. The device then analyzes the motion patterns to detect unnatural movements.
[0898] Step 4:
[0899] The server determines the crime risk based on the analysis results of the generative model. It calculates a risk score based on pre-set thresholds and rules, and if the risk score exceeds a certain threshold, it sets an alert flag. This result is recorded in a database.
[0900] Step 5:
[0901] If the device determines that there is a high risk of crime, it will send a warning to the user and relevant authorities via push notification to the smartphone or tablet, email or SMS, and, if necessary, automatically notify the police or security company using an API.
[0902] Step 6:
[0903] The user checks the warnings and notifications they receive and takes appropriate action. For example, they can check the warnings on their smartphones, check real-time footage from surveillance cameras, and, if necessary, contact a security company or police and instruct them to investigate the scene.
[0904] Step 7:
[0905] The server collects the results of warnings and new crime data and feeds it back to the generative model. This feedback allows the generative model to continuously learn and improve its prediction accuracy. The collected data is fed into the model as re-training data, and the parameters are adjusted.
[0906] Example 1
[0907] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0908] In recent years, the spread of security devices has progressed as crime risks have increased. However, conventional security systems have faced challenges in analyzing crime risks in real time and immediately sending appropriate warning notifications. Furthermore, there have been concerns about a decline in prediction accuracy due to insufficient updating of the training data for generative models. The objective of the present invention is to solve these challenges and provide a system that realizes highly accurate crime risk analysis and warning notifications in real time.
[0909] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0910] In this invention, the server includes means for collecting security data from monitoring devices, means for preprocessing the collected data, means for analyzing crime risk based on the preprocessed data using a generative model, means for sending a warning to a user or a relevant organization based on the analysis results, means for feeding back data after the warning to the generative model for learning, means for acquiring data in real time and providing a buffer for temporarily storing it, means for using a voice analysis algorithm to detect abnormal sounds, and means for using a cloud messaging service as a notification method. This makes it possible to analyze the data collected from the monitoring devices with high accuracy, determine crime risk in real time, and promptly send appropriate warnings.
[0911] A "surveillance device" is a device, such as a camera, audio sensor, or motion sensor, that collects information occurring within a specific space or environment.
[0912] "Security data" is a general term for information such as video, audio, and behavior patterns obtained from surveillance devices.
[0913] "Preprocessing" refers to the process of converting collected data into an analyzable form, and specifically includes dividing video data into frames and chunking audio data.
[0914] "Generative model" generally refers to machine learning and artificial intelligence algorithms used to predict and analyze crime risk from data.
[0915] "Crime risk" is a value that assesses the likelihood of suspicious activity or abnormal events occurring based on collected security data using certain criteria.
[0916] "Cloud Messaging Service" means an online service for sending push notifications over the Internet.
[0917] "Feedback" is the process of collecting results and new data after a warning and reusing them as training data for the analytical model.
[0918] "Buffer" refers to a temporary storage area for data, and is used to temporarily store data collected in real time.
[0919] An "audio analysis algorithm" is a mathematical method or technique for analyzing audio data to identify specific sounds or patterns.
[0920] "Notification" refers to the act of sending warnings or information to users or relevant organizations based on the analysis results.
[0921] As an embodiment of the invention, the system consists of three main components: a server, a terminal, and a user. The whole system is based on monitoring devices, where the collected data is sequentially pre-processed, analyzed, notified, and fed back.
[0922] Data collection
[0923] The server collects real-time data from multiple monitoring devices (e.g., cameras, audio sensors, motion sensors). This data is first temporarily stored in a buffer on an AWS EC2 instance. RTSP (Real-Time Streaming Protocol) and HTTP are used to obtain real-time data.
[0924] Data Preprocessing
[0925] The server splits the collected video data into frames using OpenCV, and splits the audio data into analyzable chunks using Librosa. It also performs preprocessing such as noise reduction and normalization. Motion sensor data is stored as numerical data after calibration.
[0926] Data analysis
[0927] The device then inputs the preprocessed data into generative AI models to perform analysis. For example, video data is analyzed using TensorFlow's YOLO model to detect suspicious activity, audio data is analyzed using the Google Cloud Speech-to-Text API to identify abnormal sounds, and motion data is analyzed using the SciPy library to detect unnatural movements.
[0928] Crime risk assessment
[0929] The server then uses the analysis results to determine the crime risk according to pre-set thresholds and rules. If the threshold is exceeded, a risk score is calculated and the person is deemed high risk based on that score.
[0930] Notifications and Alerts
[0931] Based on the server's assessment, the device sends a warning to the user or relevant authorities. Notifications are sent via Firebase Cloud Messaging, and emails and SMS are sent via the Twilio API. The system also includes an automatic notification function to the police and security companies using the API.
[0932] User Support
[0933] Users receive a warning notification sent from the device and check it on their smartphone app. The app provides a function to view live footage from surveillance cameras, allowing users to take appropriate action based on this information. If necessary, they can also report the incident to a security company or the police.
[0934] Building a feedback loop
[0935] The server continuously collects the results of warnings and newly acquired crime data, and uses them to retrain the generative model, thereby continuously improving the prediction accuracy of the generative AI model.
[0936] Specific examples
[0937] Example 1: Detecting and warning suspicious individuals
[0938] 1. The server collects video data from the surveillance cameras.
[0939] 2. The device analyzes this video data using a generative model to detect suspicious activity.
[0940] 3. The server determines that the crime risk is high and raises an alert flag.
[0941] 4. The device sends a warning notification to the user's smartphone.
[0942] 5. The user checks the alert on their smartphone and checks the surveillance camera footage in real time.
[0943] 6. The user will notify the police if necessary.
[0944] Example prompts to input to the generative AI model
[0945] "Please explain in natural language how your AI system analyzes collected data and generates a risk score when it detects suspicious activity or abnormal sounds, including specific libraries and tools."
[0946] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0947] Step 1: Data collection
[0948] The server collects data from surveillance devices. Specifically, the server acquires video data from surveillance cameras in real time using RTSP and temporarily stores it in a buffer on an AWS EC2 instance. It also collects audio data from audio sensors via HTTP and data from motion sensors via Bluetooth. It receives stream data from surveillance devices as input and outputs it to a temporary storage buffer.
[0949] Step 2: Preprocess the data
[0950] The server preprocesses the collected security data. Specifically, it uses OpenCV to split the video data into frames and Librosa to split the audio data into one-second chunks. Motion sensor data is digitized, denoised, and normalized. It receives raw data stored in a buffer as input and outputs data that can be converted into an analyzable format.
[0951] Step 3: Data analysis
[0952] The device inputs the preprocessed data into a generative model to perform analysis. Specifically, the device uses TensorFlow's YOLO model to analyze each frame and detect suspicious motion. It uses the Google Cloud Speech-to-Text API to analyze chunked audio data and identify abnormal sounds. It uses SciPy's signal processing functions for motion pattern analysis. It takes the preprocessed data as input and outputs specific analysis results.
[0953] Step 4: Determine crime risk
[0954] The server determines the crime risk based on the analysis results. Specifically, the server calculates a risk score based on thresholds and rules, and if it exceeds a certain value, it determines the risk as high. For example, it sets a score based on the frequency of suspicious movements or abnormal sounds. It receives the analysis results as input and outputs risk assessment data including a risk score.
[0955] Step 5: Notifications and warnings
[0956] The device sends a warning notification to the user or relevant authorities based on the risk assessment results. Specifically, the device sends a push notification to the smartphone via Firebase Cloud Messaging, sends an emergency email or SMS using the Twilio API, and automatically notifies the police or security company via the API. It receives risk assessment data as input and outputs a notification message.
[0957] Step 6: User interaction
[0958] The user can take appropriate action based on the warning notification they receive. Specifically, the user checks the warning on their smartphone and checks live footage from the surveillance camera within the app. If necessary, they can also notify a security company or the police. The app receives the notification message as input and outputs the response action to be taken.
[0959] Step 7: Building a feedback loop
[0960] The server collects the results of the warning and newly acquired crime data and feeds it back as training data for the generative model. Specifically, the server continuously monitors and collects new data and uses it to retrain the generative model, thereby improving the model's predictive accuracy. It receives the post-warning data as input and outputs it as retraining data.
[0961] (Application example 1)
[0962] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0963] In conventional security systems, when analyzing data collected from monitoring devices, it has been difficult to detect anomalies in real time, provide immediate notification, and automatically report to relevant authorities. This increases the possibility of delays in crime prevention and prompt response, posing a security issue. The present invention aims to solve these problems and provide a security system that can detect anomalies in real time and respond immediately.
[0964] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0965] In this invention, the server includes means for collecting security data from monitoring devices, means for preprocessing the collected data, means for analyzing crime risks using a generative model based on the preprocessed data, means for sending a warning to a user or relevant organizations based on the analysis results, means for feeding back data after the warning to the generative model for learning, means for sending a push notification when an abnormality is detected based on the analysis results, and means for automatically notifying relevant organizations when an abnormality is detected. This makes it possible to detect abnormalities in real time and to immediately notify / report them.
[0966] A "surveillance device" is a device, such as a camera or sensor, used to collect security data.
[0967] "Security data" refers to security-related information such as video, audio, and motion sensor data obtained through surveillance devices.
[0968] "Preprocessing" refers to the process of converting collected security data into a format suitable for analysis, and includes dividing image data into frames and audio data into chunks.
[0969] A "generative model" is an artificial intelligence model used to detect patterns and anomalies in collected data.
[0970] "Crime risk analysis" refers to the use of generative models to assess the likelihood of a crime occurring based on collected data.
[0971] "Sending a warning notification" means issuing a warning to users and relevant organizations when an abnormality is detected based on the analysis results.
[0972] "Feedback" refers to the process of feeding post-alert data and new security data back into the generative model to allow it to continuously learn.
[0973] "Push notification" is a function that immediately sends a warning message to a user's smartphone or other device when an abnormality is detected.
[0974] "Automatic reporting" is a function that automatically reports to relevant authorities (such as the police or security companies) when a crime risk or abnormality is detected.
[0975] System configuration
[0976] An embodiment of the present invention is a system that collects security data from monitoring devices, analyzes it using a generative AI model, and determines crime risk. This system is mainly composed of a server, a terminal, and a user, and each element works in cooperation with each other.
[0977] Program processing overview
[0978] The server collects data in real time from multiple monitoring devices (e.g., surveillance cameras, audio sensors, motion sensors, etc.) and has a buffer that temporarily stores it. The collected data is divided into frames, and the image data is divided into analyzable chunks based on the sample rate. The motion sensor data is also digitized and processed for noise removal, edge detection, etc.
[0979] The device inputs the preprocessed data into a generative model and performs analysis to determine the crime risk. This analysis includes detecting suspicious activity using image recognition algorithms, detecting abnormal sounds using voice recognition algorithms, and analyzing motion patterns. Based on the analysis results, the device evaluates the crime risk and calculates a risk score.
[0980] The server determines the crime risk based on the analysis results, and if the risk exceeds a set threshold, it sends a push notification to the user's smartphone or tablet, and automatically reports the situation to relevant authorities, such as the police or security companies, in real time.
[0981] Additionally, post-alert data and new security data are fed back into the generative model, allowing it to continuously learn and improve its predictions. This feedback loop allows the system to make increasingly accurate predictions over time.
[0982] Hardware and software used
[0983] Hardware:
[0984] Smartphone: Used by users to receive warning notifications and check surveillance footage in the event of an abnormality.
[0985] Surveillance cameras: Used to collect security data.
[0986] Sensors: Used to collect additional security data, such as sound and motion.
[0987] software:
[0988] Keras (keras.models, keras.preprocessing): Used to analyze video data using deep learning models.
[0989] OpenCV (cv2): Used to capture and pre-process video data.
[0990] Gmail SMTP (smtplib, email): Email sending system used for automated notifications.
[0991] Plyer (notification): Used to send push notifications to smartphones.
[0992] Specific examples
[0993] For example, a surveillance camera can detect suspicious activity and analyze the video data using a Keras model. If the result indicates a high risk of crime, the Plryer library can be used to send a push notification to the user's smartphone and automatically notify relevant authorities via Gmail SMTP. Similarly, if an audio sensor detects the sound of glass breaking, an immediate notification and report can be sent.
[0994] Prompt Sentence Examples
[0995] "Please create an application that analyzes surveillance camera footage in a specified area, detects suspicious individuals and abnormal sounds in real time, and sends push notifications and automatic reporting."
[0996] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0997] Step 1:
[0998] The server collects security data in real time from multiple monitoring devices (such as surveillance cameras, audio sensors, and motion sensors). The collected data is temporarily stored in a buffer. The input is real-time data from the monitoring devices, and the output is the security data stored in the buffer. Specifically, data is pulled from each monitoring device and stored in the server's buffer area.
[0999] Step 2:
[1000] The server preprocesses the collected security data. Image data is divided into frames, and audio data is divided into analyzable chunks based on the sample rate. Motion sensor data is also digitized and processed for noise removal and edge detection. The input is the buffered security data, and the output is the preprocessed data. Specifically, OpenCV is used to divide the video data into frames, and the audio data into chunks based on the sample rate.
[1001] Step 3:
[1002] The device inputs the preprocessed data into a generative AI model to analyze crime risk. This analysis includes detecting suspicious activity using an image recognition algorithm, detecting abnormal sounds using a voice recognition algorithm, and analyzing motion patterns. The input is the preprocessed data, and the output is the analysis results. Specifically, Keras is used to load the generative model, and the preprocessed data is input to the model to perform the analysis.
[1003] Step 4:
[1004] The server determines the crime risk based on the analysis results received from the terminal. If the crime risk is determined to be high, it sets an alert flag based on the set threshold. The input is the analysis result, and the output is the status of the alert flag. Specifically, it determines whether the threshold is exceeded based on the analysis result, and if so, sets an alert flag in the internal data structure.
[1005] Step 5:
[1006] When an alert flag is raised, the server sends a push notification to the user's smartphone or tablet. Furthermore, if necessary, it automatically notifies the police or security company. The input is the alert flag status and analysis results, and the output is the warning notification and notification results. Specifically, it uses the Plryer library to send notifications to smartphones, and uses Gmail SMTP to notify the relevant authorities via email.
[1007] Step 6:
[1008] Based on the received warning notification, the user checks the live footage from the surveillance camera and contacts the police or security company if necessary. The input is the content of the warning notification, and the output is the confirmation result of the abnormal situation and the countermeasures. Specifically, the user checks the received notification on their smartphone and displays the real-time footage through the application.
[1009] Step 7:
[1010] The server collects post-warning data and new security data and feeds it back as training data for the generative model. This allows the generative model to continuously learn and improve its prediction accuracy. The input is the post-warning data and new security data, and the output is an updated generative model. Specifically, the collected data is added to the dataset and the model is retrained using Keras.
[1011] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1012] Overall overview
[1013] This invention combines a system that uses security data collected from surveillance devices to analyze crime risks using generative models and send real-time warnings, with an emotion engine that recognizes user emotions. The system is composed of a server, a terminal, and a user, and operates in cooperation with each other.
[1014] Program processing
[1015] Data collection
[1016] The server collects data in real time from surveillance devices such as surveillance cameras, audio sensors, and motion sensors, and the collected data is temporarily stored in a buffer on the server.
[1017] Data Preprocessing
[1018] The server pre-processes the collected security data: video data is split into frames, and audio data is split into analyzable chunks based on sample rate. Image processing such as noise reduction and edge detection is also performed.
[1019] Data analysis
[1020] The device then inputs the preprocessed data into a generative model to perform a crime risk analysis. Image recognition algorithms are used to detect suspicious activity, and voice recognition algorithms are used to detect abnormal sounds. Motion patterns are also analyzed to detect unnatural movements.
[1021] Crime risk assessment
[1022] The server determines the crime risk based on the analysis results of the generative model. It calculates a risk score based on pre-set thresholds and rules, and sets an alert flag if the risk score exceeds a certain threshold.
[1023] Emotion analysis
[1024] The device's built-in emotion engine analyzes the user's emotions in real time, analyzing the user's facial expressions, tone of voice, and movement patterns from collected video and audio data to identify their emotional state.
[1025] Notifications and Alerts
[1026] The device sends notifications and warnings to the user and relevant authorities based on the crime risk assessment results and the user's emotional analysis results. The content and method of the notification change depending on the specific emotional state (e.g., tension, surprise, anger, etc.). For example, if emergency notification is set, the police and security companies will be automatically notified immediately.
[1027] User Support
[1028] Users can take appropriate action based on the warnings and notifications they receive, such as checking the content of the warning and checking live footage from surveillance cameras. Users can also issue more specific instructions based on the analysis results of the emotion engine.
[1029] Building a feedback loop
[1030] The server collects the results of the alerts and new crime data, and feeds it back into the generative model and emotion engine as training data. This feedback allows the model to continuously learn and improve its prediction accuracy.
[1031] Specific examples
[1032] Example 1: Detecting and warning suspicious individuals
[1033] 1. The server collects video data from the surveillance cameras.
[1034] 2. The device analyzes the video data using a generative model to detect suspicious activity.
[1035] 3. The server determines the risk is high and raises an alert flag.
[1036] 4. The emotion engine analyzes the user's facial expressions and voice to identify states of surprise and tension.
[1037] 5. The device sends an alert to the user's smartphone as an emergency notification.
[1038] 6. The user checks the alert on their smartphone and checks live footage from the security camera.
[1039] 7. The user will notify the police if necessary.
[1040] Example 2: Abnormal sound detection and warning
[1041] 1. The server collects audio data from intercoms and audio sensors.
[1042] 2. The device analyzes the audio data using a generative model to detect abnormal sounds, such as the sound of glass breaking.
[1043] 3. The server determines the risk is high and raises an alert flag.
[1044] 4. The emotion engine analyzes the user's voice tone and identifies their state of excitement.
[1045] 5. The device will immediately send a warning notification to the user and automatically notify the police.
[1046] 6. Users can check alerts on their smartphones and monitor the situation at the site.
[1047] 7. The police will rush to the scene and take appropriate action.
[1048] This will enable quick and accurate crime prevention and response that takes into account the user's emotional state, further improving the safety of properties.
[1049] The processing flow will be explained below.
[1050] Step 1:
[1051] The server collects data in real time from monitoring devices such as surveillance cameras, audio sensors, motion sensors, etc. Specifically, it periodically acquires data from each device (e.g., every second) and temporarily stores this data in a buffer on the server.
[1052] Step 2:
[1053] The server preprocesses the collected security data, splitting the video data into frames and the audio data into analyzable chunks based on sample rate, and applying image processing techniques such as noise reduction and edge detection to shape the data.
[1054] Step 3:
[1055] The device then inputs the preprocessed data into a generative model to perform a crime risk analysis. Specifically, it uses image recognition algorithms to detect suspicious activity and voice recognition algorithms to detect abnormal sounds. It also analyzes motion patterns to identify unnatural movements.
[1056] Step 4:
[1057] The server determines the crime risk based on the analysis results of the generative model. It calculates a risk score based on pre-set thresholds and rules, and if this risk score exceeds a certain threshold, it sets an alert flag. This determination result is recorded in a database.
[1058] Step 5:
[1059] The device's built-in emotion engine analyzes the user's emotions in real time by analyzing the user's facial expressions, tone of voice, and movement patterns from collected video and audio data to identify their emotional state.
[1060] Step 6:
[1061] The device sends notifications and warnings to the user and relevant authorities based on the crime risk assessment results and the user's emotional analysis results. The content and method of notifications can be changed depending on the user's emotional state, such as tension or surprise. For example, in an emergency, the device can immediately and automatically notify the police or security company.
[1062] Step 7:
[1063] The user checks the warnings and notifications they receive and takes appropriate action. Specifically, they check the warnings on their smartphones, check live footage from surveillance cameras, and, based on the analysis results of the emotion engine, contact security companies and police and instruct them to investigate the scene.
[1064] Step 8:
[1065] The server collects the results of warnings and new crime data, and feeds it back as training data for the generative model and emotion engine. This feedback allows the model to continuously learn and improve its prediction accuracy. The collected data is fed into the model as re-training data, and the parameters are adjusted.
[1066] Specific examples
[1067] Example 1: Detecting and warning suspicious individuals
[1068] 1. The server collects video data from the surveillance cameras.
[1069] 2. The server preprocesses the video data and divides it into frames.
[1070] 3. The device analyzes the video data using a generative model to detect suspicious activity.
[1071] 4. The server determines that the crime risk is high and raises an alert flag.
[1072] 5. The emotion engine analyzes the user's facial expressions and voice to identify states of surprise and tension.
[1073] 6. The device sends an alert to the user's smartphone as an emergency notification.
[1074] 7. The user checks the alert on their smartphone and checks live footage from the security camera.
[1075] 8. The user will notify the police if necessary.
[1076] Example 2: Abnormal sound detection and warning
[1077] 1. The server collects audio data from intercoms and audio sensors.
[1078] 2. The server preprocesses the audio data and splits it into parseable chunks.
[1079] 3. The device analyzes the audio data using a generative model to detect abnormal sounds, such as the sound of glass breaking.
[1080] 4. The server determines that the crime risk is high and raises an alert flag.
[1081] 5. The emotion engine analyzes the user's voice tone to identify excitement.
[1082] 6. The device will immediately send a warning notification to the user and automatically notify the police.
[1083] 7. Users can check alerts on their smartphones and monitor the situation at the site.
[1084] 8. The police will rush to the scene and take appropriate action.
[1085] This will enable quick and accurate crime prevention and response that takes into account the user's emotional state, further improving the safety of properties.
[1086] Example 2
[1087] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1088] In recent years, the importance of surveillance systems has increased. However, conventional systems have limited accuracy in analyzing crime risks and real-time warning notifications. Furthermore, they do not take the user's emotional state into account, making it difficult to respond quickly and appropriately. Furthermore, it is difficult to improve the learning accuracy of predictive models, and they may lack long-term reliability. This means that there is a lack of effective means to prevent crime risks before they occur.
[1089] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1090] In this invention, the server includes means for collecting security data from the monitoring devices, means for preprocessing the collected data, means for analyzing crime risk using a generative model based on the preprocessed data, means for analyzing user emotions in real time, means for sending a warning notification to the user or relevant organizations based on the analysis results, and means for feeding back data after the warning to the generative model for learning. This enables highly accurate analysis of crime risk and prompt notification that takes into account the user's emotional state, making it possible to provide an effective monitoring system that prevents crime risks before they occur.
[1091] A "surveillance device" is a hardware device for collecting security data, such as a surveillance camera, audio sensor, or motion sensor.
[1092] "Security data" refers to information used for crime risk analysis, such as video data, audio data, and motion data collected from surveillance devices.
[1093] "Preprocessing" refers to the process of converting security data into an analyzable format, specifically splitting video data into frames and audio data into chunks.
[1094] A "generative model" is a machine learning or deep learning model for analyzing crime risk based on security data.
[1095] "Crime risk" is an indicator that shows the degree to which a particular behavior or environment is likely to lead to criminal activity.
[1096] "Emotion analysis" is the process of analyzing facial expressions, vocal tone, and movement patterns to identify a user's emotional state.
[1097] A "warning notification" is an alert or notification that is sent when a criminal risk is determined to be high or based on a user's particular emotional state.
[1098] "Feedback" is the process of reusing post-warning data or newly collected data as training data for a generative model.
[1099] A "threshold" is a standard value used to determine crime risk, and anything above this value is considered high risk.
[1100] An "alert flag" is a signal or mark that is raised when the crime risk exceeds a threshold, and triggers a warning notification.
[1101] This invention is a system that analyzes security data collected from monitoring devices, analyzes crime risks in real time, and sends warning notifications to users and relevant organizations as needed. It also analyzes users' emotions to encourage more appropriate responses. Specific embodiments are described in detail below.
[1102] Hardware and Software Configuration
[1103] This system is mainly composed of a server, terminals, and users, and uses the following hardware and software.
[1104] server
[1105] The server has the following functions:
[1106] Collect security data in real time from surveillance devices (surveillance cameras, audio sensors, motion sensors).
[1107] Preprocess the collected data.
[1108] Conduct analysis to determine crime risk.
[1109] The generative model is trained by feeding back the results after the warning and new data.
[1110] The software used is OpenCV for preprocessing video data and PyDub for preprocessing audio data, and TensorFlow for inputting data into a generative model and performing analysis.
[1111] Terminal
[1112] The terminal has the following features:
[1113] The data received from the server is input into the generative model to analyze crime risk.
[1114] Analyze user emotions in real time using an emotion engine.
[1115] Send warning notifications to users and / or relevant authorities.
[1116] Specifically, it uses Microsoft Azure's Face API for facial recognition and emotion analysis, IBM Watson's Tone Analyzer for voice emotion analysis, and Twilio's API for SMS and phone notifications.
[1117] User
[1118] The user can:
[1119] Check the warning notification you receive and take appropriate action (e.g., check the warning on your smartphone and check live footage from your security camera).
[1120] Provide further instructions as needed (e.g., to call the police).
[1121] Specific examples
[1122] Example 1: Detecting and warning suspicious individuals
[1123] 1. The server collects video data from the surveillance cameras.
[1124] 2. The device analyzes the video data and detects suspicious activity.
[1125] 3. The server determines the risk is high and raises an alert flag.
[1126] 4. The emotion engine analyzes the user's facial expressions and voice to identify states of surprise and tension.
[1127] 5. The device sends an alert to the user's smartphone as an emergency notification.
[1128] 6. The user checks the alert on their smartphone and checks live footage from the security camera.
[1129] 7. The user will notify the police if necessary.
[1130] Example prompt sentence:
[1131] "Suspicious person detected. Please check live feed and call the police if necessary."
[1132] Example 2: Abnormal sound detection and warning
[1133] 1. The server collects audio data from intercoms and audio sensors.
[1134] 2. The device analyzes the audio data and detects abnormal sounds such as breaking glass.
[1135] 3. The server determines the risk is high and raises an alert flag.
[1136] 4. The emotion engine analyzes the user's voice tone and identifies their state of excitement.
[1137] 5. The device will immediately send a warning notification to the user and automatically notify the police.
[1138] 6. Users can check alerts on their smartphones and monitor the situation at the site.
[1139] 7. The police will rush to the scene and take appropriate action.
[1140] Example prompt sentence:
[1141] "An abnormal sound (sound of glass breaking) has been detected. Please check the situation on site."
[1142] Using the above method, the system can analyze crime risks with high accuracy and provide prompt notifications that take into account the user's emotional state, thereby preventing crime risks and further improving the safety of properties.
[1143] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1144] Step 1: Data collection
[1145] The server collects real-time security data from surveillance devices such as surveillance cameras, audio sensors, and motion sensors, and stores the collected data in a buffer.
[1146] Input: Video data from surveillance cameras, audio data from audio sensors, and motion data from motion sensors.
[1147] Processing: Connect to the monitoring device, acquire the data, and store it in a buffer.
[1148] Output: Buffered raw data (video, audio, motion data).
[1149] Specific operation: The server connects to a surveillance camera with a specific IP address and port number and streams video data using RTSP (Real-Time Streaming Protocol). It also sends an HTTP request to the audio sensor to obtain audio data.
[1150] Step 2: Preprocessing the data
[1151] The server pre-processes the collected data to convert it into a format that is easy to analyze: video data is divided into frames, and audio data is divided into analyzable chunks.
[1152] Input: Buffered raw data (video, audio, motion data).
[1153] Processing: Video data is split into frames and noise removal and edge detection are performed. Audio data is split into chunks based on sample rate and noise filters are applied.
[1154] Output: Video data divided into frames and audio data divided into chunks.
[1155] How it works: The server uses OpenCV to split the video data into frames and applies image processing such as histogram smoothing and edge detection to each frame. For audio data, it uses PyDub to split the data into chunks every second and apply a noise reduction filter.
[1156] Step 3: Data analysis
[1157] The device inputs the preprocessed data into a generative AI model to analyze crime risk and detects anomalies using image and voice recognition algorithms.
[1158] Input: Video data divided into frames and audio data divided into chunks.
[1159] Processing: Apply the YOLOv4 model to video data to identify suspicious individuals and abnormal behavior. Use the Google Cloud Speech-to-Text API on audio data to detect abnormal sounds.
[1160] Output: Crime risk analysis results (e.g., presence of suspicious individuals, identification of abnormal behavior, detection of abnormal sounds).
[1161] How it works: The device uses TensorFlow to feed video frames to a pre-trained YOLOv4 model to detect suspicious people and anomalous behavior, and the Google Cloud Speech-to-Text API to identify anomalous sounds.
[1162] Step 4: Determine crime risk
[1163] The server determines the crime risk based on the analysis results obtained from the generative AI model, calculates the risk score, and compares it with a threshold.
[1164] Input: Crime risk analysis results.
[1165] Processing: Calculate a risk value based on the score of the analysis results, and set an alert flag if the value exceeds the set threshold.
[1166] Output: Risk score and alert flag.
[1167] Specific operation: The server calculates a risk value based on the score of the analysis results, and if it exceeds a threshold (e.g., 0.7), it sets an alert flag.
[1168] Step 5: Sentiment Analysis
[1169] The device's built-in emotion engine analyzes the user's emotions in real time, identifying their emotional state from facial expressions, tone of voice, and movement patterns.
[1170] Input: User's video and audio data.
[1171] Processing: Apply Face API to video data to analyze facial expressions and identify emotions. Apply Tone Analyzer to audio data to analyze voice tone.
[1172] Output: The user's emotional state.
[1173] How it works: The device uses Microsoft Azure's Face API to analyze facial expressions and identify emotional states (e.g., surprise, anger, fear, etc.) and IBM Watson's Tone Analyzer to analyze voice emotions.
[1174] Step 6: Notifications and Alerts
[1175] The device sends notifications and warnings based on the crime risk assessment results and the results of user emotion analysis.
[1176] Inputs: Risk score, alert flag, sentiment analysis results.
[1177] Action: Generate notifications based on the risk score and emotional state and send them to the user or relevant authorities.
[1178] Output: Notifications and alerts (e.g. push notifications to smartphones, SMS, phone calls).
[1179] What it does: The device sends a push notification to the user's smartphone and, in some cases, automatically notifies the police or security company. SMS and phone notifications are also possible using Twilio's API.
[1180] Step 7: User interaction
[1181] Users can take appropriate action based on the warnings and notifications they receive.
[1182] Input: The warning message sent to your smartphone.
[1183] Action: Check live surveillance footage and provide additional instructions as needed.
[1184] Output: Confirmed security footage, further instructions (e.g., to call the police).
[1185] What happens: The user sees the alert on their smartphone, checks the live video feed, and, if necessary, calls the police.
[1186] Step 8: Building a feedback loop
[1187] The server collects the results and new data after the warning and feeds them back as training data for the generative model.
[1188] Input: Results after warning, newly collected data.
[1189] Processing: Store new data in the database and periodically retrain the generative model.
[1190] Output: An updated generative model.
[1191] What it does: The server saves the new data to a database and retrains the generative model and emotion engine using services like Azure Machine Learning.
[1192] (Application example 2)
[1193] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1194] Conventional security systems analyze crime risks and provide appropriate warnings, but the notification method and content are rarely flexibly adjusted according to the user's state. Furthermore, because they do not take the user's emotional state into account, appropriate responses may be delayed. Furthermore, a feedback loop after the warning is not effectively established, leading to insufficient continuous learning of the model and challenges in improving long-term accuracy.
[1195] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1196] In this invention, the server includes means for collecting security data from monitoring devices, means for preprocessing the collected data, means for analyzing crime risk using a generative model based on the preprocessed data, means for analyzing voice tone, facial expressions, movement patterns, etc. from the analyzed data to identify the user's emotional state, means for automatically adjusting the content and method of the warning based on the user's emotional state, and means for feeding back post-warning data to the generative model for learning. This enables flexible warning notifications that take the user's emotional state into consideration, improving the accuracy of crime prevention and response.
[1197] "Surveillance Devices" refers to devices used to remotely detect suspicious activity or unusual conditions, such as surveillance cameras, audio sensors, and motion sensors.
[1198] "Security data" refers to data that details surrounding conditions, such as video, audio, and motion data collected from surveillance devices.
[1199] "Preprocessing" refers to the process of preparing collected security data in a form that is easy to analyze, such as by dividing it into frames and removing noise.
[1200] A "generative model" refers to a machine learning or deep learning algorithm used to identify specific patterns or anomalies based on collected data.
[1201] "Means for analyzing crime risk" refers to a mechanism that uses a generative model to analyze crime risk, such as suspicious movements or abnormal sounds, and evaluates it as a score.
[1202] "Means for sending warning notifications" refers to a function that notifies users or relevant organizations of information regarding crime risks in real time.
[1203] "Means for identifying emotional states" refers to a system that analyzes a user's facial expressions, voice tone, and movement patterns to determine emotions such as joy, anger, surprise, and fear.
[1204] "Means for automatically adjusting the content and method of warnings" refers to a mechanism that dynamically changes the content and method of warnings sent depending on the user's emotional state.
[1205] "Feedback and learning" refers to the process of collecting post-warning results and new crime data and incorporating them into the generative model to continuously improve the model's accuracy.
[1206] "User" refers to any individual or entity that uses the System.
[1207] Overall overview
[1208] This invention combines a system that uses security data collected from surveillance devices to analyze crime risks using generative models and send real-time warnings, with an emotion engine that recognizes user emotions. The system is composed of a server, a terminal, and a user, and operates in cooperation with each other.
[1209] Hardware and Software
[1210] The hardware used is as follows:
[1211] High-performance smartphone
[1212] Network Camera
[1213] microphone
[1214] The following software is used:
[1215] Video analysis: OpenCV
[1216] Audio analysis: Librosa
[1217] Sentiment analysis: TensorFlow + Keras (emotion recognition model)
[1218] Notification system: Firebase Cloud Messaging
[1219] Data collection
[1220] The server collects data in real time from surveillance devices such as surveillance cameras, audio sensors, and motion sensors, and the collected data is temporarily stored in a buffer on the server.
[1221] Data Preprocessing
[1222] The server pre-processes the collected security data: video data is split into frames, and audio data is split into analyzable chunks based on sample rate. Image processing such as noise reduction and edge detection is also performed.
[1223] Data analysis
[1224] The device then inputs the preprocessed data into a generative model to perform a crime risk analysis. Image recognition algorithms are used to detect suspicious activity, and voice recognition algorithms are used to detect abnormal sounds. Motion patterns are also analyzed to detect unnatural movements.
[1225] Crime risk assessment
[1226] The server determines the crime risk based on the analysis results of the generative model. It calculates a risk score based on pre-set thresholds and rules, and sets an alert flag if the risk score exceeds a certain threshold.
[1227] Emotion analysis
[1228] The device's built-in emotion engine analyzes the user's emotions in real time, analyzing the user's facial expressions, tone of voice, and movement patterns from collected video and audio data to identify their emotional state.
[1229] Notifications and Alerts
[1230] The device sends notifications and warnings to the user and relevant authorities based on the crime risk assessment results and the user's emotional analysis results. The content and method of the notification change depending on the specific emotional state (e.g., tension, surprise, anger, etc.). For example, if emergency notification is set, the police and security companies will be automatically notified immediately.
[1231] User Support
[1232] Users can take appropriate action based on the warnings and notifications they receive, such as checking the content of the warning and checking live footage from surveillance cameras. Users can also issue more specific instructions based on the analysis results of the emotion engine.
[1233] Building a feedback loop
[1234] The server collects the results of the alerts and new crime data, and feeds it back into the generative model and emotion engine as training data. This feedback allows the model to continuously learn and improve its prediction accuracy.
[1235] Examples of prompt statements
[1236] Prompt: "Analyze surveillance camera footage, detect suspicious activity, and analyze the user's emotions from their facial expressions. Detect abnormal sounds (such as glass breaking). Send warning notifications to the user in real time and assess crime risk."
[1237] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1238] Step 1:
[1239] The server collects security data from surveillance devices in real time. Specifically, this includes video data from surveillance cameras, audio data from microphones, and movement data from motion sensors. The collected data is temporarily stored in a buffer on the server. The input is real-time data from the surveillance devices, and the output is the raw data stored in the buffer.
[1240] Step 2:
[1241] The server preprocesses the collected security data by splitting the video data into frames and the audio data into analyzable chunks based on the sample rate. It also performs image processing such as noise reduction and edge detection. The input is the raw data in the buffer, and the output is the preprocessed image and audio data.
[1242] Step 3:
[1243] The device inputs the preprocessed data into a generative model to perform a crime risk analysis. Specifically, it uses an image recognition algorithm to detect suspicious movements and a voice recognition algorithm to detect abnormal sounds. It also analyzes motion patterns to detect unnatural movements. The input is preprocessed image data and voice data, and the output is a crime risk score and analysis results.
[1244] Step 4:
[1245] The server determines the crime risk based on the analysis results of the generative model. Specifically, it calculates a risk score based on pre-set thresholds and rules, and sets an alert flag if the risk score exceeds a certain threshold. The inputs are the crime risk score and the analysis results, and the outputs are an alert flag and a risk assessment.
[1246] Step 5:
[1247] The device uses an emotion engine to analyze the user's emotions in real time based on the crime risk assessment results and the user's emotion analysis results. Specifically, it analyzes the user's facial expressions, voice tone, and movement patterns to identify their emotional state. The input is the user's facial expression data, voice tone, and movement patterns, and the output is the user's emotional state.
[1248] Step 6:
[1249] The device sends notifications and warnings to the user and relevant authorities based on the crime risk assessment results and the user's emotional analysis results. Specifically, the content and method of the notification are changed depending on the user's emotional state (e.g., tension, surprise, anger, etc.). For example, if emergency notification is set, the device will immediately and automatically notify the police or security company. The input is the crime risk assessment and the user's emotional state, and the output is a warning notification.
[1250] Step 7:
[1251] The user can take appropriate action based on the warnings and notifications they receive. Specifically, they can check the content of the warning and check live video footage from security cameras. The user can also issue more specific instructions based on the analysis results of the emotion engine. The input is the warning notification and live video footage, and the output is the appropriate response action.
[1252] Step 8:
[1253] The server collects the results of the warning and new crime data, and feeds it back as training data for the generative model and emotion engine. This feedback allows the model to continuously learn and improve its prediction accuracy. The input is the result data after the warning notification, and the output is an updated generative model and emotion engine.
[1254] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1255] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1256] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1257] [Fourth embodiment]
[1258] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1259] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1260] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1261] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1262] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1263] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1264] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1265] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1266] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1267] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1268] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1269] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1270] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1271] Overall overview
[1272] This system uses security data collected from surveillance devices to analyze crime risks using generative models and issue real-time warnings. It is composed of a server, terminals, and users, and operates in cooperation with each other.
[1273] Program processing
[1274] Data collection
[1275] The server collects data in real time from multiple surveillance devices (surveillance cameras, audio sensors, motion sensors, etc.) and has a buffer to temporarily store the data pulled from these devices.
[1276] Data Preprocessing
[1277] The server pre-processes the collected security data: video data is split into frames, audio data is split into analyzable chunks based on sample rate, motion sensor data is digitized, and image processing such as noise removal and edge detection is performed.
[1278] Data analysis
[1279] The device inputs the preprocessed data into a generative model and performs analysis. It uses image recognition algorithms to detect suspicious movements and voice recognition algorithms to detect abnormal sounds. It analyzes motion patterns and detects unnatural movements.
[1280] Crime risk assessment
[1281] The server then uses the analysis results to determine the risk of crime. This determination is based on pre-set thresholds and rules. For example, a risk score is calculated based on the identification pattern of a suspicious person or specific frequencies of a voice.
[1282] Notifications and Alerts
[1283] Based on the risk assessment results, the device will send notifications or warnings to the user or relevant authorities via push notifications to smartphones or tablets, email, SMS, etc. If necessary, an API will also be used to automatically notify the police or security companies.
[1284] User Support
[1285] Users can take appropriate action based on the warnings and notifications they receive, such as checking the warning on their smartphone, checking live footage from surveillance cameras, or contacting security companies or police to instruct them to investigate the scene.
[1286] Building a feedback loop
[1287] The server collects the results of the warnings and new crime data, and feeds them back as training data for the generative model. This feedback allows the model to continuously learn and improve its prediction accuracy.
[1288] Specific examples
[1289] Example 1: Detecting and warning suspicious individuals
[1290] 1. The server collects video data from the surveillance cameras.
[1291] 2. The device analyzes this video data using a generative model to detect suspicious activity.
[1292] 3. The server determines that the crime risk is high and raises an alert flag.
[1293] 4. The device sends a warning notification to the user's smartphone.
[1294] 5. The user checks the alert on their smartphone and checks the surveillance camera footage in real time.
[1295] 6. The user will notify the police if necessary.
[1296] Example 2: Abnormal sound detection and warning
[1297] 1. The server collects audio data from intercoms and audio sensors.
[1298] 2. The device analyzes the audio data and detects abnormal sounds such as breaking glass.
[1299] 3. The server determines that the risk of crime is high and sends a warning notice to the user while automatically reporting the incident to the police.
[1300] 4. The user checks the alerts on their smartphone and monitors the situation at the site.
[1301] 5. The police will rush to the scene and take appropriate action.
[1302] This will enable 24-hour security response and immediate crime prevention, improving the safety of the property.
[1303] The processing flow will be explained below.
[1304] Step 1:
[1305] The server collects real-time data from surveillance devices such as surveillance cameras, audio sensors, motion sensors, etc. Specifically, it periodically pulls data from each device (e.g., every second) and temporarily stores this data in a buffer on the server.
[1306] Step 2:
[1307] The server pre-processes the collected security data, splitting the video data into frames and the audio data into analyzable chunks based on sample rate, including applying image processing techniques such as noise removal and edge detection to shape the data.
[1308] Step 3:
[1309] The device inputs the preprocessed data into a generative model, which uses an image recognition algorithm to detect suspicious movements and a voice recognition algorithm to detect abnormal sounds. The device then analyzes the motion patterns to detect unnatural movements.
[1310] Step 4:
[1311] The server determines the crime risk based on the analysis results of the generative model. It calculates a risk score based on pre-set thresholds and rules, and if the risk score exceeds a certain threshold, it sets an alert flag. This result is recorded in a database.
[1312] Step 5:
[1313] If the device determines that there is a high risk of crime, it will send a warning to the user and relevant authorities via push notification to the smartphone or tablet, email or SMS, and, if necessary, automatically notify the police or security company using an API.
[1314] Step 6:
[1315] The user checks the warnings and notifications they receive and takes appropriate action. For example, they can check the warnings on their smartphones, check real-time footage from surveillance cameras, and, if necessary, contact a security company or police and instruct them to investigate the scene.
[1316] Step 7:
[1317] The server collects the results of warnings and new crime data and feeds it back to the generative model. This feedback allows the generative model to continuously learn and improve its prediction accuracy. The collected data is fed into the model as re-training data, and the parameters are adjusted.
[1318] Example 1
[1319] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1320] In recent years, the spread of security devices has progressed as crime risks have increased. However, conventional security systems have faced challenges in analyzing crime risks in real time and immediately sending appropriate warning notifications. Furthermore, there have been concerns about a decline in prediction accuracy due to insufficient updating of the training data for generative models. The objective of the present invention is to solve these challenges and provide a system that realizes highly accurate crime risk analysis and warning notifications in real time.
[1321] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1322] In this invention, the server includes means for collecting security data from monitoring devices, means for preprocessing the collected data, means for analyzing crime risk based on the preprocessed data using a generative model, means for sending a warning to a user or a relevant organization based on the analysis results, means for feeding back data after the warning to the generative model for learning, means for acquiring data in real time and providing a buffer for temporarily storing it, means for using a voice analysis algorithm to detect abnormal sounds, and means for using a cloud messaging service as a notification method. This makes it possible to analyze the data collected from the monitoring devices with high accuracy, determine crime risk in real time, and promptly send appropriate warnings.
[1323] A "surveillance device" is a device, such as a camera, audio sensor, or motion sensor, that collects information occurring within a specific space or environment.
[1324] "Security data" is a general term for information such as video, audio, and behavior patterns obtained from surveillance devices.
[1325] "Preprocessing" refers to the process of converting collected data into an analyzable form, and specifically includes dividing video data into frames and chunking audio data.
[1326] "Generative model" generally refers to machine learning and artificial intelligence algorithms used to predict and analyze crime risk from data.
[1327] "Crime risk" is a value that assesses the likelihood of suspicious activity or abnormal events occurring based on collected security data using certain criteria.
[1328] "Cloud Messaging Service" means an online service for sending push notifications over the Internet.
[1329] "Feedback" is the process of collecting results and new data after a warning and reusing them as training data for the analytical model.
[1330] "Buffer" refers to a temporary storage area for data, and is used to temporarily store data collected in real time.
[1331] An "audio analysis algorithm" is a mathematical method or technique for analyzing audio data to identify specific sounds or patterns.
[1332] "Notification" refers to the act of sending warnings or information to users or relevant organizations based on the analysis results.
[1333] As an embodiment of the invention, the system consists of three main components: a server, a terminal, and a user. The whole system is based on monitoring devices, where the collected data is sequentially pre-processed, analyzed, notified, and fed back.
[1334] Data collection
[1335] The server collects real-time data from multiple monitoring devices (e.g., cameras, audio sensors, motion sensors). This data is first temporarily stored in a buffer on an AWS EC2 instance. RTSP (Real-Time Streaming Protocol) and HTTP are used to obtain real-time data.
[1336] Data Preprocessing
[1337] The server splits the collected video data into frames using OpenCV, and splits the audio data into analyzable chunks using Librosa. It also performs preprocessing such as noise reduction and normalization. Motion sensor data is stored as numerical data after calibration.
[1338] Data analysis
[1339] The device then inputs the preprocessed data into generative AI models to perform analysis. For example, video data is analyzed using TensorFlow's YOLO model to detect suspicious activity, audio data is analyzed using the Google Cloud Speech-to-Text API to identify abnormal sounds, and motion data is analyzed using the SciPy library to detect unnatural movements.
[1340] Crime risk assessment
[1341] The server then uses the analysis results to determine the crime risk according to pre-set thresholds and rules. If the threshold is exceeded, a risk score is calculated and the person is deemed high risk based on that score.
[1342] Notifications and Alerts
[1343] Based on the server's assessment, the device sends a warning to the user or relevant authorities. Notifications are sent via Firebase Cloud Messaging, and emails and SMS are sent via the Twilio API. The system also includes an automatic notification function to the police and security companies using the API.
[1344] User Support
[1345] Users receive a warning notification sent from the device and check it on their smartphone app. The app provides a function to view live footage from surveillance cameras, allowing users to take appropriate action based on this information. If necessary, they can also report the incident to a security company or the police.
[1346] Building a feedback loop
[1347] The server continuously collects the results of warnings and newly acquired crime data, and uses them to retrain the generative model, thereby continuously improving the prediction accuracy of the generative AI model.
[1348] Specific examples
[1349] Example 1: Detecting and warning suspicious individuals
[1350] 1. The server collects video data from the surveillance cameras.
[1351] 2. The device analyzes this video data using a generative model to detect suspicious activity.
[1352] 3. The server determines that the crime risk is high and raises an alert flag.
[1353] 4. The device sends a warning notification to the user's smartphone.
[1354] 5. The user checks the alert on their smartphone and checks the surveillance camera footage in real time.
[1355] 6. The user will notify the police if necessary.
[1356] Example prompts to input to the generative AI model
[1357] "Please explain in natural language how your AI system analyzes collected data and generates a risk score when it detects suspicious activity or abnormal sounds, including specific libraries and tools."
[1358] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1359] Step 1: Data collection
[1360] The server collects data from surveillance devices. Specifically, the server acquires video data from surveillance cameras in real time using RTSP and temporarily stores it in a buffer on an AWS EC2 instance. It also collects audio data from audio sensors via HTTP and data from motion sensors via Bluetooth. It receives stream data from surveillance devices as input and outputs it to a temporary storage buffer.
[1361] Step 2: Preprocess the data
[1362] The server preprocesses the collected security data. Specifically, it uses OpenCV to split the video data into frames and Librosa to split the audio data into one-second chunks. Motion sensor data is digitized, denoised, and normalized. It receives raw data stored in a buffer as input and outputs data that can be converted into an analyzable format.
[1363] Step 3: Data analysis
[1364] The device inputs the preprocessed data into a generative model to perform analysis. Specifically, the device uses TensorFlow's YOLO model to analyze each frame and detect suspicious motion. It uses the Google Cloud Speech-to-Text API to analyze chunked audio data and identify abnormal sounds. It uses SciPy's signal processing functions for motion pattern analysis. It takes the preprocessed data as input and outputs specific analysis results.
[1365] Step 4: Determine crime risk
[1366] The server determines the crime risk based on the analysis results. Specifically, the server calculates a risk score based on thresholds and rules, and if it exceeds a certain value, it determines the risk as high. For example, it sets a score based on the frequency of suspicious movements or abnormal sounds. It receives the analysis results as input and outputs risk assessment data including a risk score.
[1367] Step 5: Notifications and warnings
[1368] The device sends a warning notification to the user or relevant authorities based on the risk assessment results. Specifically, the device sends a push notification to the smartphone via Firebase Cloud Messaging, sends an emergency email or SMS using the Twilio API, and automatically notifies the police or security company via the API. It receives risk assessment data as input and outputs a notification message.
[1369] Step 6: User interaction
[1370] The user can take appropriate action based on the warning notification they receive. Specifically, the user checks the warning on their smartphone and checks live footage from the surveillance camera within the app. If necessary, they can also notify a security company or the police. The app receives the notification message as input and outputs the response action to be taken.
[1371] Step 7: Building a feedback loop
[1372] The server collects the results of the warning and newly acquired crime data and feeds it back as training data for the generative model. Specifically, the server continuously monitors and collects new data and uses it to retrain the generative model, thereby improving the model's predictive accuracy. It receives the post-warning data as input and outputs it as retraining data.
[1373] (Application example 1)
[1374] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1375] In conventional security systems, when analyzing data collected from monitoring devices, it has been difficult to detect anomalies in real time, provide immediate notification, and automatically report to relevant authorities. This increases the possibility of delays in crime prevention and prompt response, posing a security issue. The present invention aims to solve these problems and provide a security system that can detect anomalies in real time and respond immediately.
[1376] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1377] In this invention, the server includes means for collecting security data from monitoring devices, means for preprocessing the collected data, means for analyzing crime risks using a generative model based on the preprocessed data, means for sending a warning to a user or relevant organizations based on the analysis results, means for feeding back data after the warning to the generative model for learning, means for sending a push notification when an abnormality is detected based on the analysis results, and means for automatically notifying relevant organizations when an abnormality is detected. This makes it possible to detect abnormalities in real time and to immediately notify / report them.
[1378] A "surveillance device" is a device, such as a camera or sensor, used to collect security data.
[1379] "Security data" refers to security-related information such as video, audio, and motion sensor data obtained through surveillance devices.
[1380] "Preprocessing" refers to the process of converting collected security data into a format suitable for analysis, and includes dividing image data into frames and audio data into chunks.
[1381] A "generative model" is an artificial intelligence model used to detect patterns and anomalies in collected data.
[1382] "Crime risk analysis" refers to the use of generative models to assess the likelihood of a crime occurring based on collected data.
[1383] "Sending a warning notification" means issuing a warning to users and relevant organizations when an abnormality is detected based on the analysis results.
[1384] "Feedback" refers to the process of feeding post-alert data and new security data back into the generative model to allow it to continuously learn.
[1385] "Push notification" is a function that immediately sends a warning message to a user's smartphone or other device when an abnormality is detected.
[1386] "Automatic reporting" is a function that automatically reports to relevant authorities (such as the police or security companies) when a crime risk or abnormality is detected.
[1387] System configuration
[1388] An embodiment of the present invention is a system that collects security data from monitoring devices, analyzes it using a generative AI model, and determines crime risk. This system is mainly composed of a server, a terminal, and a user, and each element works in cooperation with each other.
[1389] Program processing overview
[1390] The server collects data in real time from multiple monitoring devices (e.g., surveillance cameras, audio sensors, motion sensors, etc.) and has a buffer that temporarily stores it. The collected data is divided into frames, and the image data is divided into analyzable chunks based on the sample rate. The motion sensor data is also digitized and processed for noise removal, edge detection, etc.
[1391] The device inputs the preprocessed data into a generative model and performs analysis to determine the crime risk. This analysis includes detecting suspicious activity using image recognition algorithms, detecting abnormal sounds using voice recognition algorithms, and analyzing motion patterns. Based on the analysis results, the device evaluates the crime risk and calculates a risk score.
[1392] The server determines the crime risk based on the analysis results, and if the risk exceeds a set threshold, it sends a push notification to the user's smartphone or tablet, and automatically reports the situation to relevant authorities, such as the police or security companies, in real time.
[1393] Additionally, post-alert data and new security data are fed back into the generative model, allowing it to continuously learn and improve its predictions. This feedback loop allows the system to make increasingly accurate predictions over time.
[1394] Hardware and software used
[1395] Hardware:
[1396] Smartphone: Used by users to receive warning notifications and check surveillance footage in the event of an abnormality.
[1397] Surveillance cameras: Used to collect security data.
[1398] Sensors: Used to collect additional security data, such as sound and motion.
[1399] software:
[1400] Keras (keras.models, keras.preprocessing): Used to analyze video data using deep learning models.
[1401] OpenCV (cv2): Used to capture and pre-process video data.
[1402] Gmail SMTP (smtplib, email): Email sending system used for automated notifications.
[1403] Plyer (notification): Used to send push notifications to smartphones.
[1404] Specific examples
[1405] For example, a surveillance camera can detect suspicious activity and analyze the video data using a Keras model. If the result indicates a high risk of crime, the Plryer library can be used to send a push notification to the user's smartphone and automatically notify relevant authorities via Gmail SMTP. Similarly, if an audio sensor detects the sound of glass breaking, an immediate notification and report can be sent.
[1406] Prompt Sentence Examples
[1407] "Please create an application that analyzes surveillance camera footage in a specified area, detects suspicious individuals and abnormal sounds in real time, and sends push notifications and automatic reporting."
[1408] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1409] Step 1:
[1410] The server collects security data in real time from multiple monitoring devices (such as surveillance cameras, audio sensors, and motion sensors). The collected data is temporarily stored in a buffer. The input is real-time data from the monitoring devices, and the output is the security data stored in the buffer. Specifically, data is pulled from each monitoring device and stored in the server's buffer area.
[1411] Step 2:
[1412] The server preprocesses the collected security data. Image data is divided into frames, and audio data is divided into analyzable chunks based on the sample rate. Motion sensor data is also digitized and processed for noise removal and edge detection. The input is the buffered security data, and the output is the preprocessed data. Specifically, OpenCV is used to divide the video data into frames, and the audio data into chunks based on the sample rate.
[1413] Step 3:
[1414] The device inputs the preprocessed data into a generative AI model to analyze crime risk. This analysis includes detecting suspicious activity using an image recognition algorithm, detecting abnormal sounds using a voice recognition algorithm, and analyzing motion patterns. The input is the preprocessed data, and the output is the analysis results. Specifically, Keras is used to load the generative model, and the preprocessed data is input to the model to perform the analysis.
[1415] Step 4:
[1416] The server determines the crime risk based on the analysis results received from the terminal. If the crime risk is determined to be high, it sets an alert flag based on the set threshold. The input is the analysis result, and the output is the status of the alert flag. Specifically, it determines whether the threshold is exceeded based on the analysis result, and if so, sets an alert flag in the internal data structure.
[1417] Step 5:
[1418] When an alert flag is raised, the server sends a push notification to the user's smartphone or tablet. Furthermore, if necessary, it automatically notifies the police or security company. The input is the alert flag status and analysis results, and the output is the warning notification and notification results. Specifically, it uses the Plryer library to send notifications to smartphones, and uses Gmail SMTP to notify the relevant authorities via email.
[1419] Step 6:
[1420] Based on the received warning notification, the user checks the live footage from the surveillance camera and contacts the police or security company if necessary. The input is the content of the warning notification, and the output is the confirmation result of the abnormal situation and the countermeasures. Specifically, the user checks the received notification on their smartphone and displays the real-time footage through the application.
[1421] Step 7:
[1422] The server collects post-warning data and new security data and feeds it back as training data for the generative model. This allows the generative model to continuously learn and improve its prediction accuracy. The input is the post-warning data and new security data, and the output is an updated generative model. Specifically, the collected data is added to the dataset and the model is retrained using Keras.
[1423] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1424] Overall overview
[1425] This invention combines a system that uses security data collected from surveillance devices to analyze crime risks using generative models and send real-time warnings, with an emotion engine that recognizes user emotions. The system is composed of a server, a terminal, and a user, and operates in cooperation with each other.
[1426] Program processing
[1427] Data collection
[1428] The server collects data in real time from surveillance devices such as surveillance cameras, audio sensors, and motion sensors, and the collected data is temporarily stored in a buffer on the server.
[1429] Data Preprocessing
[1430] The server pre-processes the collected security data: video data is split into frames, and audio data is split into analyzable chunks based on sample rate. Image processing such as noise reduction and edge detection is also performed.
[1431] Data analysis
[1432] The device then inputs the preprocessed data into a generative model to perform a crime risk analysis. Image recognition algorithms are used to detect suspicious activity, and voice recognition algorithms are used to detect abnormal sounds. Motion patterns are also analyzed to detect unnatural movements.
[1433] Crime risk assessment
[1434] The server determines the crime risk based on the analysis results of the generative model. It calculates a risk score based on pre-set thresholds and rules, and sets an alert flag if the risk score exceeds a certain threshold.
[1435] Emotion analysis
[1436] The device's built-in emotion engine analyzes the user's emotions in real time, analyzing the user's facial expressions, tone of voice, and movement patterns from collected video and audio data to identify their emotional state.
[1437] Notifications and Alerts
[1438] The device sends notifications and warnings to the user and relevant authorities based on the crime risk assessment results and the user's emotional analysis results. The content and method of the notification change depending on the specific emotional state (e.g., tension, surprise, anger, etc.). For example, if emergency notification is set, the police and security companies will be automatically notified immediately.
[1439] User Support
[1440] Users can take appropriate action based on the warnings and notifications they receive, such as checking the content of the warning and checking live footage from surveillance cameras. Users can also issue more specific instructions based on the analysis results of the emotion engine.
[1441] Building a feedback loop
[1442] The server collects the results of the alerts and new crime data, and feeds it back into the generative model and emotion engine as training data. This feedback allows the model to continuously learn and improve its prediction accuracy.
[1443] Specific examples
[1444] Example 1: Detecting and warning suspicious individuals
[1445] 1. The server collects video data from the surveillance cameras.
[1446] 2. The device analyzes the video data using a generative model to detect suspicious activity.
[1447] 3. The server determines the risk is high and raises an alert flag.
[1448] 4. The emotion engine analyzes the user's facial expressions and voice to identify states of surprise and tension.
[1449] 5. The device sends an alert to the user's smartphone as an emergency notification.
[1450] 6. The user checks the alert on their smartphone and checks live footage from the security camera.
[1451] 7. The user will notify the police if necessary.
[1452] Example 2: Abnormal sound detection and warning
[1453] 1. The server collects audio data from intercoms and audio sensors.
[1454] 2. The device analyzes the audio data using a generative model to detect abnormal sounds, such as the sound of glass breaking.
[1455] 3. The server determines the risk is high and raises an alert flag.
[1456] 4. The emotion engine analyzes the user's voice tone and identifies their state of excitement.
[1457] 5. The device will immediately send a warning notification to the user and automatically notify the police.
[1458] 6. Users can check alerts on their smartphones and monitor the situation at the site.
[1459] 7. The police will rush to the scene and take appropriate action.
[1460] This will enable quick and accurate crime prevention and response that takes into account the user's emotional state, further improving the safety of properties.
[1461] The processing flow will be explained below.
[1462] Step 1:
[1463] The server collects data in real time from monitoring devices such as surveillance cameras, audio sensors, motion sensors, etc. Specifically, it periodically acquires data from each device (e.g., every second) and temporarily stores this data in a buffer on the server.
[1464] Step 2:
[1465] The server preprocesses the collected security data, splitting the video data into frames and the audio data into analyzable chunks based on sample rate, and applying image processing techniques such as noise reduction and edge detection to shape the data.
[1466] Step 3:
[1467] The device then inputs the preprocessed data into a generative model to perform a crime risk analysis. Specifically, it uses image recognition algorithms to detect suspicious activity and voice recognition algorithms to detect abnormal sounds. It also analyzes motion patterns to identify unnatural movements.
[1468] Step 4:
[1469] The server determines the crime risk based on the analysis results of the generative model. It calculates a risk score based on pre-set thresholds and rules, and if this risk score exceeds a certain threshold, it sets an alert flag. This determination result is recorded in a database.
[1470] Step 5:
[1471] The device's built-in emotion engine analyzes the user's emotions in real time by analyzing the user's facial expressions, tone of voice, and movement patterns from collected video and audio data to identify their emotional state.
[1472] Step 6:
[1473] The device sends notifications and warnings to the user and relevant authorities based on the crime risk assessment results and the user's emotional analysis results. The content and method of notifications can be changed depending on the user's emotional state, such as tension or surprise. For example, in an emergency, the device can immediately and automatically notify the police or security company.
[1474] Step 7:
[1475] The user checks the warnings and notifications they receive and takes appropriate action. Specifically, they check the warnings on their smartphones, check live footage from surveillance cameras, and, based on the analysis results of the emotion engine, contact security companies and police and instruct them to investigate the scene.
[1476] Step 8:
[1477] The server collects the results of warnings and new crime data, and feeds it back as training data for the generative model and emotion engine. This feedback allows the model to continuously learn and improve its prediction accuracy. The collected data is fed into the model as re-training data, and the parameters are adjusted.
[1478] Specific examples
[1479] Example 1: Detecting and warning suspicious individuals
[1480] 1. The server collects video data from the surveillance cameras.
[1481] 2. The server preprocesses the video data and divides it into frames.
[1482] 3. The device analyzes the video data using a generative model to detect suspicious activity.
[1483] 4. The server determines that the crime risk is high and raises an alert flag.
[1484] 5. The emotion engine analyzes the user's facial expressions and voice to identify states of surprise and tension.
[1485] 6. The device sends an alert to the user's smartphone as an emergency notification.
[1486] 7. The user checks the alert on their smartphone and checks live footage from the security camera.
[1487] 8. The user will notify the police if necessary.
[1488] Example 2: Abnormal sound detection and warning
[1489] 1. The server collects audio data from intercoms and audio sensors.
[1490] 2. The server preprocesses the audio data and splits it into parseable chunks.
[1491] 3. The device analyzes the audio data using a generative model to detect abnormal sounds, such as the sound of glass breaking.
[1492] 4. The server determines that the crime risk is high and raises an alert flag.
[1493] 5. The emotion engine analyzes the user's voice tone to identify excitement.
[1494] 6. The device will immediately send a warning notification to the user and automatically notify the police.
[1495] 7. Users can check alerts on their smartphones and monitor the situation at the site.
[1496] 8. The police will rush to the scene and take appropriate action.
[1497] This will enable quick and accurate crime prevention and response that takes into account the user's emotional state, further improving the safety of properties.
[1498] Example 2
[1499] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1500] In recent years, the importance of surveillance systems has increased. However, conventional systems have limited accuracy in analyzing crime risks and real-time warning notifications. Furthermore, they do not take the user's emotional state into account, making it difficult to respond quickly and appropriately. Furthermore, it is difficult to improve the learning accuracy of predictive models, and they may lack long-term reliability. This means that there is a lack of effective means to prevent crime risks before they occur.
[1501] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1502] In this invention, the server includes means for collecting security data from the monitoring devices, means for preprocessing the collected data, means for analyzing crime risk using a generative model based on the preprocessed data, means for analyzing user emotions in real time, means for sending a warning notification to the user or relevant organizations based on the analysis results, and means for feeding back data after the warning to the generative model for learning. This enables highly accurate analysis of crime risk and prompt notification that takes into account the user's emotional state, making it possible to provide an effective monitoring system that prevents crime risks before they occur.
[1503] A "surveillance device" is a hardware device for collecting security data, such as a surveillance camera, audio sensor, or motion sensor.
[1504] "Security data" refers to information used for crime risk analysis, such as video data, audio data, and motion data collected from surveillance devices.
[1505] "Preprocessing" refers to the process of converting security data into an analyzable format, specifically splitting video data into frames and audio data into chunks.
[1506] A "generative model" is a machine learning or deep learning model for analyzing crime risk based on security data.
[1507] "Crime risk" is an indicator that shows the degree to which a particular behavior or environment is likely to lead to criminal activity.
[1508] "Emotion analysis" is the process of analyzing facial expressions, vocal tone, and movement patterns to identify a user's emotional state.
[1509] A "warning notification" is an alert or notification that is sent when a criminal risk is determined to be high or based on a user's particular emotional state.
[1510] "Feedback" is the process of reusing post-warning data or newly collected data as training data for a generative model.
[1511] A "threshold" is a standard value used to determine crime risk, and anything above this value is considered high risk.
[1512] An "alert flag" is a signal or mark that is raised when the crime risk exceeds a threshold, and triggers a warning notification.
[1513] This invention is a system that analyzes security data collected from monitoring devices, analyzes crime risks in real time, and sends warning notifications to users and relevant organizations as needed. It also analyzes users' emotions to encourage more appropriate responses. Specific embodiments are described in detail below.
[1514] Hardware and Software Configuration
[1515] This system is mainly composed of a server, terminals, and users, and uses the following hardware and software.
[1516] server
[1517] The server has the following functions:
[1518] Collect security data in real time from surveillance devices (surveillance cameras, audio sensors, motion sensors).
[1519] Preprocess the collected data.
[1520] Conduct analysis to determine crime risk.
[1521] The generative model is trained by feeding back the results after the warning and new data.
[1522] The software used is OpenCV for preprocessing video data and PyDub for preprocessing audio data, and TensorFlow for inputting data into a generative model and performing analysis.
[1523] Terminal
[1524] The terminal has the following features:
[1525] The data received from the server is input into the generative model to analyze crime risk.
[1526] Analyze user emotions in real time using an emotion engine.
[1527] Send warning notifications to users and / or relevant authorities.
[1528] Specifically, it uses Microsoft Azure's Face API for facial recognition and emotion analysis, IBM Watson's Tone Analyzer for voice emotion analysis, and Twilio's API for SMS and phone notifications.
[1529] User
[1530] The user can:
[1531] Check the warning notification you receive and take appropriate action (e.g., check the warning on your smartphone and check live footage from your security camera).
[1532] Provide further instructions as needed (e.g., to call the police).
[1533] Specific examples
[1534] Example 1: Detecting and warning suspicious individuals
[1535] 1. The server collects video data from the surveillance cameras.
[1536] 2. The device analyzes the video data and detects suspicious activity.
[1537] 3. The server determines the risk is high and raises an alert flag.
[1538] 4. The emotion engine analyzes the user's facial expressions and voice to identify states of surprise and tension.
[1539] 5. The device sends an alert to the user's smartphone as an emergency notification.
[1540] 6. The user checks the alert on their smartphone and checks live footage from the security camera.
[1541] 7. The user will notify the police if necessary.
[1542] Example prompt sentence:
[1543] "Suspicious person detected. Please check live feed and call the police if necessary."
[1544] Example 2: Abnormal sound detection and warning
[1545] 1. The server collects audio data from intercoms and audio sensors.
[1546] 2. The device analyzes the audio data and detects abnormal sounds such as breaking glass.
[1547] 3. The server determines the risk is high and raises an alert flag.
[1548] 4. The emotion engine analyzes the user's voice tone and identifies their state of excitement.
[1549] 5. The device will immediately send a warning notification to the user and automatically notify the police.
[1550] 6. Users can check alerts on their smartphones and monitor the situation at the site.
[1551] 7. The police will rush to the scene and take appropriate action.
[1552] Example prompt sentence:
[1553] "An abnormal sound (sound of glass breaking) has been detected. Please check the situation on site."
[1554] Using the above method, the system can analyze crime risks with high accuracy and provide prompt notifications that take into account the user's emotional state, thereby preventing crime risks and further improving the safety of properties.
[1555] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1556] Step 1: Data collection
[1557] The server collects real-time security data from surveillance devices such as surveillance cameras, audio sensors, and motion sensors, and stores the collected data in a buffer.
[1558] Input: Video data from surveillance cameras, audio data from audio sensors, and motion data from motion sensors.
[1559] Processing: Connect to the monitoring device, acquire the data, and store it in a buffer.
[1560] Output: Buffered raw data (video, audio, motion data).
[1561] Specific operation: The server connects to a surveillance camera with a specific IP address and port number and streams video data using RTSP (Real-Time Streaming Protocol). It also sends an HTTP request to the audio sensor to obtain audio data.
[1562] Step 2: Preprocessing the data
[1563] The server pre-processes the collected data to convert it into a format that is easy to analyze: video data is divided into frames, and audio data is divided into analyzable chunks.
[1564] Input: Buffered raw data (video, audio, motion data).
[1565] Processing: Video data is split into frames and noise removal and edge detection are performed. Audio data is split into chunks based on sample rate and noise filters are applied.
[1566] Output: Video data divided into frames and audio data divided into chunks.
[1567] How it works: The server uses OpenCV to split the video data into frames and applies image processing such as histogram smoothing and edge detection to each frame. For audio data, it uses PyDub to split the data into chunks every second and apply a noise reduction filter.
[1568] Step 3: Data analysis
[1569] The device inputs the preprocessed data into a generative AI model to analyze crime risk and detects anomalies using image and voice recognition algorithms.
[1570] Input: Video data divided into frames and audio data divided into chunks.
[1571] Processing: Apply the YOLOv4 model to video data to identify suspicious individuals and abnormal behavior. Use the Google Cloud Speech-to-Text API on audio data to detect abnormal sounds.
[1572] Output: Crime risk analysis results (e.g., presence of suspicious individuals, identification of abnormal behavior, detection of abnormal sounds).
[1573] How it works: The device uses TensorFlow to feed video frames to a pre-trained YOLOv4 model to detect suspicious people and anomalous behavior, and the Google Cloud Speech-to-Text API to identify anomalous sounds.
[1574] Step 4: Determine crime risk
[1575] The server determines the crime risk based on the analysis results obtained from the generative AI model, calculates the risk score, and compares it with a threshold.
[1576] Input: Crime risk analysis results.
[1577] Processing: Calculate a risk value based on the score of the analysis results, and set an alert flag if the value exceeds the set threshold.
[1578] Output: Risk score and alert flag.
[1579] Specific operation: The server calculates a risk value based on the score of the analysis results, and if it exceeds a threshold (e.g., 0.7), it sets an alert flag.
[1580] Step 5: Sentiment Analysis
[1581] The device's built-in emotion engine analyzes the user's emotions in real time, identifying their emotional state from facial expressions, tone of voice, and movement patterns.
[1582] Input: User's video and audio data.
[1583] Processing: Apply Face API to video data to analyze facial expressions and identify emotions. Apply Tone Analyzer to audio data to analyze voice tone.
[1584] Output: The user's emotional state.
[1585] How it works: The device uses Microsoft Azure's Face API to analyze facial expressions and identify emotional states (e.g., surprise, anger, fear, etc.) and IBM Watson's Tone Analyzer to analyze voice emotions.
[1586] Step 6: Notifications and Alerts
[1587] The device sends notifications and warnings based on the crime risk assessment results and the results of user emotion analysis.
[1588] Inputs: Risk score, alert flag, sentiment analysis results.
[1589] Action: Generate notifications based on the risk score and emotional state and send them to the user or relevant authorities.
[1590] Output: Notifications and alerts (e.g. push notifications to smartphones, SMS, phone calls).
[1591] What it does: The device sends a push notification to the user's smartphone and, in some cases, automatically notifies the police or security company. SMS and phone notifications are also possible using Twilio's API.
[1592] Step 7: User interaction
[1593] Users can take appropriate action based on the warnings and notifications they receive.
[1594] Input: The warning message sent to your smartphone.
[1595] Action: Check live surveillance footage and provide additional instructions as needed.
[1596] Output: Confirmed security footage, further instructions (e.g., to call the police).
[1597] What happens: The user sees the alert on their smartphone, checks the live video feed, and, if necessary, calls the police.
[1598] Step 8: Building a feedback loop
[1599] The server collects the results and new data after the warning and feeds them back as training data for the generative model.
[1600] Input: Results after warning, newly collected data.
[1601] Processing: Store new data in the database and periodically retrain the generative model.
[1602] Output: An updated generative model.
[1603] What it does: The server saves the new data to a database and retrains the generative model and emotion engine using services like Azure Machine Learning.
[1604] (Application example 2)
[1605] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1606] Conventional security systems analyze crime risks and provide appropriate warnings, but the notification method and content are rarely flexibly adjusted according to the user's state. Furthermore, because they do not take the user's emotional state into account, appropriate responses may be delayed. Furthermore, a feedback loop after the warning is not effectively established, leading to insufficient continuous learning of the model and challenges in improving long-term accuracy.
[1607] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1608] In this invention, the server includes means for collecting security data from monitoring devices, means for preprocessing the collected data, means for analyzing crime risk using a generative model based on the preprocessed data, means for analyzing voice tone, facial expressions, movement patterns, etc. from the analyzed data to identify the user's emotional state, means for automatically adjusting the content and method of the warning based on the user's emotional state, and means for feeding back post-warning data to the generative model for learning. This enables flexible warning notifications that take the user's emotional state into consideration, improving the accuracy of crime prevention and response.
[1609] "Surveillance Devices" refers to devices used to remotely detect suspicious activity or unusual conditions, such as surveillance cameras, audio sensors, and motion sensors.
[1610] "Security data" refers to data that details surrounding conditions, such as video, audio, and motion data collected from surveillance devices.
[1611] "Preprocessing" refers to the process of preparing collected security data in a form that is easy to analyze, such as by dividing it into frames and removing noise.
[1612] A "generative model" refers to a machine learning or deep learning algorithm used to identify specific patterns or anomalies based on collected data.
[1613] "Means for analyzing crime risk" refers to a mechanism that uses a generative model to analyze crime risk, such as suspicious movements or abnormal sounds, and evaluates it as a score.
[1614] "Means for sending warning notifications" refers to a function that notifies users or relevant organizations of information regarding crime risks in real time.
[1615] "Means for identifying emotional states" refers to a system that analyzes a user's facial expressions, voice tone, and movement patterns to determine emotions such as joy, anger, surprise, and fear.
[1616] "Means for automatically adjusting the content and method of warnings" refers to a mechanism that dynamically changes the content and method of warnings sent depending on the user's emotional state.
[1617] "Feedback and learning" refers to the process of collecting post-warning results and new crime data and incorporating them into the generative model to continuously improve the model's accuracy.
[1618] "User" refers to any individual or entity that uses the System.
[1619] Overall overview
[1620] This invention combines a system that uses security data collected from surveillance devices to analyze crime risks using generative models and send real-time warnings, with an emotion engine that recognizes user emotions. The system is composed of a server, a terminal, and a user, and operates in cooperation with each other.
[1621] Hardware and Software
[1622] The hardware used is as follows:
[1623] High-performance smartphone
[1624] Network Camera
[1625] microphone
[1626] The following software is used:
[1627] Video analysis: OpenCV
[1628] Audio analysis: Librosa
[1629] Sentiment analysis: TensorFlow + Keras (emotion recognition model)
[1630] Notification system: Firebase Cloud Messaging
[1631] Data collection
[1632] The server collects data in real time from surveillance devices such as surveillance cameras, audio sensors, and motion sensors, and the collected data is temporarily stored in a buffer on the server.
[1633] Data Preprocessing
[1634] The server pre-processes the collected security data: video data is split into frames, and audio data is split into analyzable chunks based on sample rate. Image processing such as noise reduction and edge detection is also performed.
[1635] Data analysis
[1636] The device then inputs the preprocessed data into a generative model to perform a crime risk analysis. Image recognition algorithms are used to detect suspicious activity, and voice recognition algorithms are used to detect abnormal sounds. Motion patterns are also analyzed to detect unnatural movements.
[1637] Crime risk assessment
[1638] The server determines the crime risk based on the analysis results of the generative model. It calculates a risk score based on pre-set thresholds and rules, and sets an alert flag if the risk score exceeds a certain threshold.
[1639] Emotion analysis
[1640] The device's built-in emotion engine analyzes the user's emotions in real time, analyzing the user's facial expressions, tone of voice, and movement patterns from collected video and audio data to identify their emotional state.
[1641] Notifications and Alerts
[1642] The device sends notifications and warnings to the user and relevant authorities based on the crime risk assessment results and the user's emotional analysis results. The content and method of the notification change depending on the specific emotional state (e.g., tension, surprise, anger, etc.). For example, if emergency notification is set, the police and security companies will be automatically notified immediately.
[1643] User Support
[1644] Users can take appropriate action based on the warnings and notifications they receive, such as checking the content of the warning and checking live footage from surveillance cameras. Users can also issue more specific instructions based on the analysis results of the emotion engine.
[1645] Building a feedback loop
[1646] The server collects the results of the alerts and new crime data, and feeds it back into the generative model and emotion engine as training data. This feedback allows the model to continuously learn and improve its prediction accuracy.
[1647] Examples of prompt statements
[1648] Prompt: "Analyze surveillance camera footage, detect suspicious activity, and analyze the user's emotions from their facial expressions. Detect abnormal sounds (such as glass breaking). Send warning notifications to the user in real time and assess crime risk."
[1649] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1650] Step 1:
[1651] The server collects security data from surveillance devices in real time. Specifically, this includes video data from surveillance cameras, audio data from microphones, and movement data from motion sensors. The collected data is temporarily stored in a buffer on the server. The input is real-time data from the surveillance devices, and the output is the raw data stored in the buffer.
[1652] Step 2:
[1653] The server preprocesses the collected security data by splitting the video data into frames and the audio data into analyzable chunks based on the sample rate. It also performs image processing such as noise reduction and edge detection. The input is the raw data in the buffer, and the output is the preprocessed image and audio data.
[1654] Step 3:
[1655] The device inputs the preprocessed data into a generative model to perform a crime risk analysis. Specifically, it uses an image recognition algorithm to detect suspicious movements and a voice recognition algorithm to detect abnormal sounds. It also analyzes motion patterns to detect unnatural movements. The input is preprocessed image data and voice data, and the output is a crime risk score and analysis results.
[1656] Step 4:
[1657] The server determines the crime risk based on the analysis results of the generative model. Specifically, it calculates a risk score based on pre-set thresholds and rules, and sets an alert flag if the risk score exceeds a certain threshold. The inputs are the crime risk score and the analysis results, and the outputs are an alert flag and a risk assessment.
[1658] Step 5:
[1659] The device uses an emotion engine to analyze the user's emotions in real time based on the crime risk assessment results and the user's emotion analysis results. Specifically, it analyzes the user's facial expressions, voice tone, and movement patterns to identify their emotional state. The input is the user's facial expression data, voice tone, and movement patterns, and the output is the user's emotional state.
[1660] Step 6:
[1661] The device sends notifications and warnings to the user and relevant authorities based on the crime risk assessment results and the user's emotional analysis results. Specifically, the content and method of the notification are changed depending on the user's emotional state (e.g., tension, surprise, anger, etc.). For example, if emergency notification is set, the device will immediately and automatically notify the police or security company. The input is the crime risk assessment and the user's emotional state, and the output is a warning notification.
[1662] Step 7:
[1663] The user can take appropriate action based on the warnings and notifications they receive. Specifically, they can check the content of the warning and check live video footage from security cameras. The user can also issue more specific instructions based on the analysis results of the emotion engine. The input is the warning notification and live video footage, and the output is the appropriate response action.
[1664] Step 8:
[1665] The server collects the results of the warning and new crime data, and feeds it back as training data for the generative model and emotion engine. This feedback allows the model to continuously learn and improve its prediction accuracy. The input is the result data after the warning notification, and the output is an updated generative model and emotion engine.
[1666] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1667] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1668] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1669] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1670] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1671] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1672] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1673] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1674] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1675] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1676] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1677] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1678] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1679] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1680] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1681] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1682] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1683] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1684] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1685] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1686] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1687] The following is further disclosed regarding the above embodiment.
[1688] (Claim 1)
[1689] means for collecting security data from the monitoring device;
[1690] means for pre-processing the collected data;
[1691] A means for analyzing crime risk using a generative model based on the preprocessed data;
[1692] a means for sending a warning notice to a user or a relevant organization based on the analysis result;
[1693] The system includes a means for feeding post-warning data back into the generative model for learning.
[1694] (Claim 2)
[1695] 10. The system of claim 1, wherein the means for pre-processing the collected security data divides image data into frames and audio data into analyzable chunks based on sample rate.
[1696] (Claim 3)
[1697] The system of claim 1, further comprising means for setting an alert flag when a threshold set based on the analysis result is exceeded.
[1698] "Example 1"
[1699] (Claim 1)
[1700] means for collecting security data from the monitoring device;
[1701] means for pre-processing the collected data;
[1702] A means for analyzing crime risk using a generative model based on the preprocessed data;
[1703] a means for sending a warning notice to a user or a relevant organization based on the analysis result;
[1704] A means for feeding back data after the warning to the generative model for learning;
[1705] a means for acquiring data in real time and providing a buffer for temporarily storing the data;
[1706] means for using a sound analysis algorithm to detect abnormal sounds;
[1707] A system including means for utilizing a cloud messaging service as a notification method.
[1708] (Claim 2)
[1709] 10. The system of claim 1, wherein the means for pre-processing the collected security data divides image data into frames and audio data into analyzable chunks based on sample rate.
[1710] (Claim 3)
[1711] The system of claim 1, further comprising means for setting an alert flag when a threshold set based on the analysis result is exceeded.
[1712] "Application Example 1"
[1713] (Claim 1)
[1714] means for collecting security data from the monitoring device;
[1715] means for pre-processing the collected data;
[1716] A means for analyzing crime risk using a generative model based on the preprocessed data;
[1717] a means for sending a warning notice to a user or a relevant organization based on the analysis result;
[1718] A means for feeding back data after the warning to the generative model for learning;
[1719] a means for sending a push notification when an anomaly is detected based on the analysis results;
[1720] A system that includes a means to automatically notify relevant authorities when an abnormality is detected.
[1721] (Claim 2)
[1722] 10. The system of claim 1, wherein the means for pre-processing the collected security data divides image data into frames and audio data into analyzable chunks based on sample rate.
[1723] (Claim 3)
[1724] The system of claim 1, further comprising: means for setting an alert flag when a threshold set based on the analysis result is exceeded; and means for enabling live video to be viewed in real time when an abnormality is detected.
[1725] "Example 2: Combining Emotion Engines"
[1726] (Claim 1)
[1727] means for collecting security data from the monitoring device;
[1728] means for pre-processing the collected data;
[1729] A means for analyzing crime risk using a generative model based on the preprocessed data;
[1730] A means for analyzing user emotions in real time;
[1731] a means for sending a warning notice to a user or a relevant organization based on the analysis result;
[1732] The system includes a means for feeding post-warning data back into the generative model for learning.
[1733] (Claim 2)
[1734] 2. The system of claim 1, wherein the means for pre-processing the collected security data divides video data into frames and audio data into analyzable chunks.
[1735] (Claim 3)
[1736] The system of claim 1, further comprising means for setting an alert flag when a threshold set based on the analysis result is exceeded.
[1737] "Application example 2 when combining emotion engines"
[1738] (Claim 1)
[1739] means for collecting security data from the monitoring device;
[1740] means for pre-processing the collected data;
[1741] A means for analyzing crime risk using a generative model based on the preprocessed data;
[1742] a means for sending a warning notice to a user or a relevant organization based on the analysis result;
[1743] means for analyzing voice tones, facial expressions, movement patterns, etc. from the analyzed data to identify the emotional state of the user;
[1744] means for automatically adjusting the content and manner of warnings based on the user's emotional state;
[1745] The system includes a means for feeding post-warning data back into the generative model for learning.
[1746] (Claim 2)
[1747] 10. The system of claim 1, wherein the means for pre-processing the collected security data divides image data into frames and audio data into analyzable chunks based on sample rate.
[1748] (Claim 3)
[1749] The system of claim 1, further comprising means for setting an alert flag when a threshold set based on the analysis result is exceeded. [Explanation of symbols]
[1750] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for collecting security data from the monitoring device; means for pre-processing the collected data; A means for analyzing crime risk using a generative model based on the preprocessed data; a means for sending a warning notice to a user or a relevant organization based on the analysis result; The system includes a means for feeding post-warning data back into the generative model for learning.
2. 2. The system of claim 1, wherein the means for pre-processing the collected security data divides image data into frames and audio data into analyzable chunks based on sample rate.
3. The system according to claim 1, further comprising means for setting an alert flag when a threshold value set based on the analysis result is exceeded.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A