System
The system addresses the labor shortage and monitoring inefficiencies in surveillance by using real-time video data processing and generative AI to detect and alert security personnel to suspicious or lost items, improving public safety.
Patent Information
- Application Number
- JP2024130364
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-06
- Publication Date
- 2026-02-19
AI Technical Summary
The security industry faces a labor shortage and insufficient real-time monitoring of surveillance camera footage, leading to delayed detection of lost items and suspicious objects, which hampers effective incident prevention and reduces public safety.
A system that acquires video data from surveillance cameras in real-time, preprocesses it, and uses a generative AI model to detect objects, generating alerts for relevant personnel and recording information for analysis.
Enables real-time, accurate detection and response to suspicious or lost items, enhancing security by improving detection efficiency and accuracy.
Smart Images

Figure 2026028066000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] The security and security industry faces a serious labor shortage, and real-time monitoring of surveillance camera footage is often insufficient. As a result, lost items and suspicious objects are discovered late, making it difficult to prevent incidents and minimize damage. A decline in public safety and a drop in arrest rates are also becoming a problem, and effective measures to address these issues are needed. In conventional surveillance systems, video analysis and alert generation are often performed manually, resulting in a lack of efficiency and accuracy. [Means for solving the problem]
[0005] This invention provides a system that acquires video data from surveillance cameras in real time and detects objects using a generative AI model. Specifically, the system includes a means for acquiring video data from surveillance cameras, a means for preprocessing the acquired video data and inputting it into a generative AI model, a means for detecting objects using the generative AI model and identifying suspicious or lost items, and a means for generating and notifying relevant personnel when a suspicious or lost item is detected. This system enables real-time detection of suspicious or lost items, significantly improving the effectiveness of security and guarding. Furthermore, the system also includes a means for identifying the occupant or suspicious individual of an object using past video data, and a means for recording information based on the detection of a suspicious or lost item for use in subsequent analysis and improvement, thereby achieving sustained security enhancements.
[0006] 1. Surveillance cameras
[0007] A "surveillance camera" is a device for monitoring a specific area as video.
[0008] 2. Video data
[0009] "Video data" refers to information about images and videos captured by surveillance cameras.
[0010] 3. Pretreatment
[0011] "Preprocessing" refers to the processing and conversion of video data so that it can be efficiently analyzed by a generative AI model.
[0012] 4. Generative AI Models
[0013] A "generative AI model" is an algorithm or software that uses machine learning or deep learning to detect objects.
[0014] 5. Object
[0015] "Object" refers to a person, object, or abnormal condition detected in surveillance camera footage.
[0016] 6. Suspicious Objects
[0017] A "suspicious object" refers to an object in video data that is determined to be abnormal based on specific criteria.
[0018] 7. Lost and Found
[0019] "Lost property" refers to an object that was originally owned by someone else but has now been left behind in its current location.
[0020] 8. Alerts
[0021] An "alert" refers to a warning message or mechanism used to notify relevant parties of information about detected suspicious or lost items.
[0022] 9. Stakeholders
[0023] "Stakeholders" refers to the people and organizations that receive and respond to alerts.
[0024] 10. Logs
[0025] "Log" refers to data that keeps records of events and alerts that occur within a system.
[0026] 11. Notification
[0027] "Notification" refers to the communication methods and messages used to convey detection results and alert information to relevant parties.
[0028] 12. Past video data
[0029] "Past video data" refers to previously acquired video data that has been stored for comparison with the current video.
[0030] 13. Occupant
[0031] "Possessor" refers to a person who is in or was in possession of a particular object.
[0032] 14. Records
[0033] "Recording" refers to the act of preserving information or events observed within a system, or that information.
[0034] 15. Analysis
[0035] "Analysis" refers to the act of examining and analyzing acquired data in detail to understand its meaning and characteristics. [Brief explanation of the drawings]
[0036] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10]1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0037] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0038] First, the terms used in the following description will be explained.
[0039] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0040] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0041] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0042] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0043] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0044] [First embodiment]
[0045] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0046] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0047] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0048] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0049] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0050] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0051] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0052] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0053] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0054] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0055] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0056] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0057] The system for implementing this invention consists of a surveillance camera, a server, a terminal, and a user. The server acquires video data from the surveillance camera in real time and analyzes the video data using a generative AI model. The terminal receives notifications from the server, and the user takes action based on the notifications.
[0058] Program processing explanation
[0059] 1. Obtaining surveillance camera footage
[0060] Server: Acquires video data from surveillance cameras in real time. Specifically, it receives video data in stream format via the surveillance camera's URL or IP address.
[0061] 2. Preprocessing of video data
[0062] Server: Preprocesses the acquired video data to input it into the generative AI model. Preprocessing includes resizing and normalizing the video frames.
[0063] 3. Object detection using AI models
[0064] Server: The generative AI model detects objects in preprocessed video frames, and calculates a label and its confidence score for each object.
[0065] 4. Analysis of detection results
[0066] Server: Identifies suspicious or lost objects based on the confidence score of the detected object. Criteria for suspicious or lost objects are set in advance, and judgments are made based on those criteria.
[0067] 5. Comparison with past data
[0068] Server: Detected objects are compared with past video data to identify the object's occupant or suspicious individuals. This improves the accuracy of determining whether an object's status is suspicious.
[0069] 6. Alert Generation and Notification
[0070] Server: If an item is determined to be suspicious or lost, it automatically generates an alert and notifies the relevant parties via email, SMS, or a dedicated application.
[0071] Terminal: The relevant person's terminal receives the notification and displays the alert content. The user checks the notification and takes appropriate action.
[0072] 7. Record of correspondence
[0073] Server: Records the actions taken in response to generated alerts, which are used for later analysis and system improvement.
[0074] Specific examples
[0075] A specific example from a shopping mall is shown below.
[0076] Surveillance camera: Monitoring the food court in the shopping mall.
[0077] Server: Captures surveillance camera footage in real time and analyzes it using a generative AI model. The model detects unattended bags as suspicious objects. It checks that unattended bags in the footage have been left unattended for a certain period of time. It compares this with past data to identify the culprit or owner. Based on these results, it generates an alert and notifies security guards.
[0078] Terminal: The security guard's smartphone receives an alert and displays a notification saying, "There is an unattended bag in the food court on the first floor. Please check."
[0079] User (security guard): The security guard checks the notification, goes to the scene, checks the safety of the bag, and enters the details of the action into the terminal and records it on the server.
[0080] This will enable early detection of suspicious objects and lost items in real time and prompt response, improving security within shopping malls.
[0081] The processing flow will be explained below.
[0082] Step 1:
[0083] The server acquires video data from the surveillance camera in real time, receiving the video stream via the surveillance camera's URL or IP address and capturing it as video data.
[0084] Step 2:
[0085] The server converts the acquired video data into a format that is easy for the generative AI model to process, specifically by resizing the video frames to 224x224 pixels and normalizing them to match the model's input format.
[0086] Step 3:
[0087] The server inputs the preprocessed video data into a generative AI model to detect objects, which then calculates object labels (e.g., bag, person, etc.) and their confidence scores.
[0088] Step 4:
[0089] The server receives the detection results from the generative AI model and evaluates whether the confidence score exceeds a certain threshold. If the detection result exceeds the threshold, it is listed as a candidate for suspicious or lost items.
[0090] Step 5:
[0091] The server compares the listed suspicious or lost items with past video data. By using past data, it can identify the occupant of the object or suspicious individuals, improving the accuracy of the judgment.
[0092] Step 6:
[0093] The server automatically generates an alert if an object is detected as suspicious or lost, including the object's label, location, and time stamp.
[0094] Step 7:
[0095] The server notifies the relevant parties of the generated alerts via email, SMS or a dedicated application.
[0096] Step 8:
[0097] The device receives the notification sent from the server and displays the alert content. The relevant parties can check this notification and take appropriate action.
[0098] Step 9:
[0099] The user (security guard) checks the alert displayed on the terminal, heads to the scene to check the safety of the object, and then inputs the details of the response into the terminal.
[0100] Step 10:
[0101] The server records the responses from users (security guards), which are used for future analysis and system improvement.
[0102] Example 1
[0103] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0104] Although there are systems that analyze surveillance footage in real time and quickly identify and notify suspicious or lost items, it is difficult to respond with high accuracy and efficiency.In addition, there is a lack of cross-checking with past video data and management of response records, so there is a need for stronger crime prevention measures.
[0105] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0106] In this invention, the server includes means for acquiring video data from a monitoring device in real time, means for preprocessing the acquired video data and inputting it into an artificial intelligence model for detecting objects, means for detecting objects using the artificial intelligence model and identifying suspicious objects or lost items, means for generating and notifying relevant parties when a suspicious object or lost item is detected, means for identifying the occupant of the object or a suspicious person using past video data, and means for recording detected information based on the detection of a suspicious object or lost item and using it for later analysis and improvement, thereby enabling highly accurate detection in real time and rapid response.
[0107] "Monitoring device" refers to a device for acquiring video data in real time.
[0108] "Video data" refers to real-time video information obtained from surveillance equipment.
[0109] "Preprocessing" refers to a series of steps that are performed on the captured video data before it can be input into an AI model. Specifically, this includes processes such as resizing and normalizing the video frames.
[0110] "Artificial intelligence model" refers to the algorithms and learning models used to analyze acquired video data and detect objects.
[0111] "Objects" refer to objects that the AI model detects in the video data, including suspicious objects and lost items.
[0112] A "suspicious object" is an object that is determined to be abnormal or suspicious based on certain criteria.
[0113] "Lost property" refers to an object that has been determined to be unclaimed or abandoned.
[0114] An "alert" refers to warning information used to notify relevant parties when a suspicious object or lost item is detected.
[0115] "Interested parties" refers to people or organizations that should receive the alert, including security guards and administrators.
[0116] "Past video data" refers to video information that was recorded in the past among the acquired video data.
[0117] "Matching" refers to the process of comparing currently detected objects with past video data to identify the occupant or suspicious person of a matching object.
[0118] "Response record" refers to data used to store the history of the detection of suspicious or lost items and the response taken thereto.
[0119] The system for implementing this invention consists of a monitoring device, a server, a terminal, and a user. The server acquires video data from the monitoring device in real time, preprocesses the data, and inputs it into a generative AI model for analysis. The terminal receives notifications from the server, and the user responds based on the notifications.
[0120] The server is equipped with hardware and software for acquiring, preprocessing, and analyzing video data. Specifically, libraries such as OpenCV and FFmpeg are used to perform preprocessing such as resizing and normalizing video frames. Generative AI models such as YOLO (You Only Look Once) and SSD (Single Shot Multibox Detector) are used to detect objects in the preprocessed video data.
[0121] As an example, consider surveillance in a shopping mall. Video data is acquired from a surveillance device installed in the mall's food court and received by a server. The server resizes and normalizes the acquired video data (for example, to 640x480 pixels) and inputs it into the YOLO v4 model. The model detects suspicious objects, such as unattended bags, with high accuracy. An object is deemed suspicious if the confidence score is 0.85 or higher.
[0122] The server then compares the detected suspicious object with video data from the past 24 hours to identify the occupant or suspicious person. This reduces false positives and enables more accurate identification of suspicious objects. It automatically generates an alert and sends a notification to the device of the relevant person (e.g., security guard). This notification can be sent via email, SMS, or a dedicated application.
[0123] A specific example of the notification content is a message such as "There is an unattended bag in the food court on the first floor. Please check." Security personnel receive this notification and rush to the scene to respond. The response at that time is recorded on the server via the terminal. The recorded information is used for future analysis and system improvements.
[0124] Prompt Sentence Examples
[0125] Example of using security systems in shopping malls:
[0126] The surveillance cameras capture real-time footage of the food court and send it to the server. The server uses a generative AI model (YOLO v4) to detect unattended bags and compares the detection results with video data from the past 24 hours to identify suspicious individuals. If a suspicious object is determined based on the specified criteria, the security guard's smartphone should be notified, "There is an unattended bag in the 1st floor food court. Please take a look." The server should record the response results.
[0127] This invention enables highly accurate real-time detection of suspicious objects and rapid response, improving security in shopping malls and other public places.
[0128] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0129] Step 1: Obtaining security camera footage
[0130] Specific behavior:
[0131] Server: Obtain video data in real time using the IP address or URL of the surveillance device. For example, use OpenCV to capture the camera stream with cv2.VideoCapture.
[0132] Input: Video data from a surveillance device (stream format)
[0133] Output: Captured video frames
[0134] Example: If the URL of the monitoring device is rtsp: / / 192.168.1.10 / stream, use this URL to get the video.
[0135] Step 2: Preprocessing the video data
[0136] Specific behavior:
[0137] Server: Resize and normalize the captured video frames. To resize, use OpenCV's cv2.resize function, and to normalize, scale pixel values to the range 0 to 1.
[0138] Input: Captured video frame
[0139] Output: Preprocessed video frame (resized image data)
[0140] Example: Resize and normalize the frame to 640x480 pixels.
[0141] Step 3: Object detection with an AI model
[0142] Specific behavior:
[0143] Server: Inputs preprocessed video frames into a generative AI model (such as YOLO) to perform object detection, and outputs object labels and confidence scores.
[0144] Input: Preprocessed video frames
[0145] Output: Detected objects (including labels and confidence scores)
[0146] Example: Detect unattended bags using YOLO v4 model. If the confidence score is 0.85 or higher, it is considered suspicious.
[0147] Step 4: Analyzing the detection results
[0148] Specific behavior:
[0149] Server: Identifies suspicious or lost objects based on the confidence score and label of the detected object. If the confidence score exceeds a certain threshold, the object is deemed suspicious.
[0150] Input: Detected objects (labels and confidence scores)
[0151] Output: Suspicious or lost item determination result
[0152] Example: If the "bag" label is detected with a confidence score of 0.85 or higher, identify it as suspicious.
[0153] Step 5: Check against historical data
[0154] Specific behavior:
[0155] Server: For detected suspicious objects, retrieves video data from the past 24 hours from a database and compares it to identify occupants or suspicious individuals. Past frames are input into the AI model again to confirm the object.
[0156] Input: Information on detected suspicious objects, past video data
[0157] Output: Information on identified occupants or suspicious individuals
[0158] Example: Matching data from the past 24 hours to determine who the bag belongs to or was involved in it.
[0159] Step 6: Alerting and Notification
[0160] Specific behavior:
[0161] Server: If an item is determined to be suspicious or lost, an alert is automatically generated and relevant parties are notified via email, SMS, or a dedicated application.
[0162] Terminal: The terminal of the person involved (security guard) receives the notification and displays a pop-up notification.
[0163] Input: Suspicious object judgment result
[0164] Output: Alert notification
[0165] Example: A security guard's smartphone displays a notification such as, "There is an unattended bag in the food court on the first floor. Please check it."
[0166] Step 7: Record your response
[0167] Specific behavior:
[0168] Server: Records the actions taken in response to generated alerts. This includes the date and time of the action, the person who took the action, and the action content. Records this in a database.
[0169] Input: Detailed information about the response (date, time, responder, content)
[0170] Output: Corresponding record
[0171] Example: A security guard inspects the scene, enters the details into a terminal, and the information is recorded on the server.
[0172] The above steps create a system that can detect suspicious objects with high accuracy in real time and respond quickly.
[0173] (Application example 1)
[0174] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0175] In security systems that use surveillance cameras, the detection of suspicious objects and lost items often relies on manual monitoring, which can easily lead to oversights. Existing systems also lack the ability to notify relevant parties of detection results in real time, resulting in information delays and making it difficult to respond quickly. Furthermore, the ability to identify suspicious individuals and related activities using past video data is also insufficient, creating a need for enhanced security.
[0176] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0177] In this invention, the server
[0178] A means for acquiring video data from a surveillance camera in real time;
[0179] A means for pre-processing the acquired video data and inputting it into a generative AI model for object detection;
[0180] A means for detecting objects using a generative AI model and identifying suspicious or lost items;
[0181] A means for generating alerts and notifying relevant parties via email or a dedicated application about detected suspicious or lost items;
[0182] a means for displaying detection results and alerts to relevant personnel in real time using a head-mounted display;
[0183] This enables highly accurate detection and rapid response in real time.
[0184] A "surveillance camera" is a device that acquires video data in real time and monitors a surveillance area.
[0185] "Video data" refers to a video stream captured by a surveillance camera, and is a set of image information that is subject to analysis and monitoring.
[0186] "Preprocessing" refers to the process of converting video data into a format suitable for the generative AI model, and specifically includes resizing, normalization, and noise removal.
[0187] A "generative AI model" is a technology that uses deep learning and machine learning algorithms to generate models, learn patterns from input data, and extract features.
[0188] A "target" is an object that exists in a surveillance camera image and that needs to be identified or detected.
[0189] A "suspicious object" is an object within a surveillance area that is deemed abnormal or dangerous.
[0190] "Lost property" is an object that has been left behind with no known owner.
[0191] An "alert" is a notification that warns or calls attention to detected suspicious objects or lost items.
[0192] "Email" is a means of communication for sending and receiving messages over the Internet.
[0193] A "dedicated application" is software developed for a specific purpose, and is a system designed to provide notifications and information.
[0194] A "head-mounted display" is a wearable display device that displays information within the user's field of vision when worn by the user.
[0195] "Interested parties" are those who should be notified by the system, and are typically security personnel such as security guards or administrators.
[0196] The system for implementing this invention consists of a surveillance camera, a server, a head-mounted display terminal, and a user (mainly a security guard). The server acquires video data from the surveillance camera in real time and analyzes the video data using a generative AI model. The head-mounted display (HMD) terminal worn by the security guard receives notifications from the server and displays the notification contents to the security guard in real time.
[0197] Hardware and software used
[0198] Hardware: surveillance cameras, servers, head-mounted displays (HMDs)
[0199] Software: TensorFlow (generative AI model), OpenCV (image processing)
[0200] Details of data processing and calculation
[0201] The server receives video data from the surveillance cameras in real time. This data is first preprocessed using OpenCV. Specifically, the preprocessing involves resizing and normalizing the video frames.
[0202] The pre-processed video data is then fed into a generative AI model for object detection. The generative AI model, built using TensorFlow, detects objects in the video and calculates a label and confidence score.
[0203] If any of the detected objects are deemed suspicious or lost, the server compares them with past video data to identify the object's occupant or suspicious person.
[0204] The server generates an alert based on the results and automatically notifies the relevant parties (security guards) via email or a dedicated application. The head-mounted display terminal worn by the guard also displays the detection results and alert details in real time.
[0205] Specific examples
[0206] For example, if a suspicious object (an unattended bag) is left unattended for a certain period of time in a food court in a shopping mall, the system will operate as follows:
[0207] Surveillance cameras capture footage of the food court.
[0208] The server acquires the video data in real time, preprocesses it, and inputs it into the generative AI model.
[0209] The generative AI model detects unattended bags, analyzes the information and compares it with historical data.
[0210] The server checks for the presence of suspicious objects and generates an alert.
[0211] The security guard's head-mounted display displays a real-time message saying, "There is an unattended bag in the food court on the first floor. Please check."
[0212] Prompt Sentence Examples
[0213] For example, use the following prompt:
[0214] "Monitoring the food court for 45 minutes. Detecting an unattended bag and confirming that it had been left there for 20 minutes. Checking against past data to identify the owner. Notifying a security officer wearing an HMD that 'There is an unattended bag in the food court on the first floor. Please check.'"
[0215] This configuration enables highly accurate detection of suspicious objects and lost items within the surveillance area, enabling rapid response in real time.
[0216] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0217] Step 1:
[0218] The server acquires video data from the surveillance camera in real time. The input is the video stream sent from the surveillance camera, and the output is the raw video data stored in the server. The video data is received in stream format via the URL or IP address of the surveillance camera.
[0219] Step 2:
[0220] The server performs preprocessing on the acquired video data before inputting it into the generative AI model. The input is the raw video data acquired in step 1, and the output is preprocessed video frames. Preprocessing includes resizing, normalizing, and denoising the video frames. Specifically, OpenCV is used to resize the frames to the required size and normalize the pixel values.
[0221] Step 3:
[0222] The server inputs the preprocessed video data into a generative AI model to detect objects. The input is the video frames preprocessed in step 2, and the output is the labels and confidence scores of the detected objects. A generative AI model using TensorFlow is used to identify objects in the video frames. Each object is assigned a label and its confidence score.
[0223] Step 4:
[0224] The server identifies suspicious or lost objects based on the confidence score of the detected object. The input is the object label and confidence score generated in step 3, and the output is the result of the suspicious or lost object judgment. Criteria for suspicious or lost objects are set in advance, and the judgment is made according to those criteria, and a flag is raised if action is required.
[0225] Step 5:
[0226] The server compares the detected object with past video data to identify the object's occupant or suspicious person. The input is the judgment result from step 4 and past video data, and the output is the identification of the owner or suspicious person of the suspicious object or lost item. This improves the accuracy of the judgment by comparing with past behavioral history.
[0227] Step 6:
[0228] If the server determines that an item is suspicious or lost, it automatically generates an alert and notifies the relevant parties. The input is the determination result and identification result, and the output is the generated alert and notification message. The alert is sent via email or a dedicated application.
[0229] Step 7:
[0230] The terminal (head-mounted display) receives an alert notification from the server and displays the notification content to the security guard. The input is the alert notification sent from the server, and the output is the warning message displayed on the head-mounted display. The notification content is specific, such as "There is an unattended bag in the food court on the first floor. Please check it."
[0231] Step 8:
[0232] The user (security guard) checks the notification displayed on the head-mounted display, heads to the scene, and takes action. The input is the alert content displayed on the head-mounted display, and the output is the execution and reporting of the response. The report content is entered into the terminal and recorded on the server.
[0233] These steps enable highly accurate detection of suspicious or lost items in real time and rapid response.
[0234] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0235] The system embodying this invention consists of a surveillance camera, a server, a terminal, a user, and an emotion engine. The server acquires video data from the surveillance camera in real time and analyzes the video data using a generative AI model. The emotion engine recognizes the user's emotions and incorporates that information into the analysis. The terminal receives notifications from the server, and the user responds based on those notifications.
[0236] Program processing explanation
[0237] 1. Obtaining surveillance camera footage
[0238] Server: Acquires video data from surveillance cameras in real time. Receives video streams via the surveillance camera's URL or IP address and acquires them as video data.
[0239] 2. Preprocessing of video data
[0240] Server: Converts acquired video data into a format that is easy for the generative AI model to process. Preprocessing includes resizing and normalizing video frames.
[0241] 3. Object detection using AI models
[0242] Server: The generative AI model detects objects in pre-processed video frames, and calculates a label for each object (e.g., bag, person, etc.) and its confidence score.
[0243] 4. Analysis of detection results
[0244] Server: Identifies suspicious or lost objects based on the confidence score of the detected object. Criteria for suspicious or lost objects are set in advance, and judgments are made based on those criteria.
[0245] 5. Comparison with past data
[0246] Server: Detected objects are compared with past video data to identify the object's occupant or suspicious individuals. This improves the accuracy of determining whether an object's status is suspicious.
[0247] 6. Emotion Recognition by Emotion Engine
[0248] Server: Uses an emotion engine that recognizes the user's emotions using video and audio data, thereby obtaining emotional information such as whether the person involved is feeling stressed.
[0249] 7. Alert Generation and Coordination
[0250] Server: If an item is determined to be suspicious or lost, an alert is automatically generated. The content of the alert and notification method are adjusted based on the recognition results of the emotion engine. For example, if the person involved is under high stress, the content of the notification may be softened.
[0251] Terminal: A notification is sent to the relevant person's terminal and the alert details are displayed. The user checks this notification and takes appropriate action.
[0252] 8. Record of correspondence
[0253] Server: Records the actions taken in response to generated alerts, which are used for later analysis and system improvement.
[0254] Specific examples
[0255] Here is a specific example from inside a shopping mall.
[0256] Surveillance camera: Monitoring the food court in the shopping mall.
[0257] Server: Captures surveillance camera footage in real time and analyzes it using a generative AI model. The model detects unattended bags as suspicious objects. It checks that unattended bags in the footage have been left unattended for a certain period of time and compares them with past data to identify the owner. It also uses an emotion engine to evaluate the stress level of security guards. Based on the detection results, it generates tailored alerts and notifies security guards.
[0258] Device: The security guard's smartphone receives an alert and displays a notification saying, "There is an unattended bag in the food court on the first floor. Please check." If the security guard is in a high-stress state, the notification content is changed to, "There is an unattended bag in the food court on the first floor. Please check."
[0259] User (security guard): The security guard checks the notification, heads to the scene, checks the safety of the bag, and then enters the details of the action into the terminal.
[0260] Server: Records the responses from security guards and uses them for later analysis.
[0261] This system will improve security within shopping malls by enabling early detection and rapid response of suspicious or lost items in real time. In addition, by combining it with an emotion engine, it will be possible to respond taking into account the emotional state of those involved.
[0262] The processing flow will be explained below.
[0263] Step 1:
[0264] The server acquires video data from the surveillance camera in real time, receives the video stream via the surveillance camera's URL or IP address, and acquires the video data frame by frame.
[0265] Step 2:
[0266] The server converts the acquired video data into a format that is easy for the generative AI model to process, specifically by resizing the video frames to 224x224 pixels and normalizing them to match the model's input format.
[0267] Step 3:
[0268] The server inputs the preprocessed video data into a generative AI model for object detection, which identifies each object in the frame and outputs a label (e.g., bag, person, etc.) and its confidence score.
[0269] Step 4:
[0270] The server analyzes the output of the generative AI model and evaluates whether the confidence score exceeds a certain threshold. Detections that exceed the threshold are listed as suspicious or lost items.
[0271] Step 5:
[0272] The server compares the listed suspicious or lost items with past video data, and based on the past data, identifies the occupant or suspicious person of the object, and confirms that it is a suspicious or lost item.
[0273] Step 6:
[0274] The server uses the video and audio data to recognize the user's emotions through an emotion engine, analyzing the user's facial expressions and tone of voice to evaluate their emotional state, such as their stress level.
[0275] Step 7:
[0276] The server automatically generates an alert if the item is determined to be a suspicious or lost item. The content of the alert and notification method are adjusted based on the results of the emotion engine. For example, if the user is in a state of high stress, the notification content will be changed to a softer expression.
[0277] Step 8:
[0278] The server notifies the relevant parties of the generated alerts via email, SMS, or a dedicated application sent to the device.
[0279] Step 9:
[0280] The terminal receives the notification from the server and displays the alert content. The user can check this notification and take the necessary action.
[0281] Step 10:
[0282] The user (security guard) checks the alert displayed on the terminal, goes to the scene to check the object, and then inputs the response details into the terminal.
[0283] Step 11:
[0284] The server records the responses entered by the user (security guard), which will be used for future analysis and system improvement.
[0285] As a concrete example, consider the case where an unattended bag is left in the food court of a shopping mall. The server detects the unattended bag from surveillance camera footage. It compares this with past data to identify the bag's owner. It uses an emotion engine to check the stress level of the security guard and generates an alert in appropriate language to notify the guard. The security guard then checks the notification, checks the safety of the bag on-site, and enters the results into the server. This series of processes enables effective and human-friendly security operations.
[0286] Example 2
[0287] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0288] Conventional surveillance systems detect suspicious objects or lost items, but do not take into account the emotional state of those involved, which can lead to stress when dealing with the incident. Furthermore, detection accuracy can be low because the system does not compare the detected items with past video data. Furthermore, the recording and subsequent analysis of detected information is insufficient, which can hinder system improvement.
[0289] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for acquiring video data from a surveillance camera in real time, means for preprocessing the acquired video data and inputting it into a generative AI model for detecting objects, means for detecting objects using the generative AI model and identifying suspicious or lost objects, means for evaluating the emotional state of relevant parties using an emotion engine and adjusting the content of an alert when a suspicious or lost object is detected, and means for generating and notifying relevant parties of an alert. This enables highly accurate detection of suspicious or lost objects and appropriate responses that take into account the emotional state of relevant parties. In addition, detection accuracy is improved by comparing with past video data, and continuous system improvement is possible through recording and subsequent analysis of detection information.
[0290] A "surveillance camera" is a device that captures images of a pre-set location in real time.
[0291] A "server" is a computer system that obtains, processes, and manages data from other devices over a network.
[0292] "Real-time" refers to processing immediately without delay, and refers to a state in which appropriate action can be taken the moment a specific event occurs.
[0293] "Video data" means a collection of visual information captured from a surveillance camera or other device and stored as video frames.
[0294] "Preprocessing" refers to a series of operations that transform data into a suitable format before it is sent to the main processing stage, including, for example, resizing and normalization.
[0295] A "generative AI model" is an artificial intelligence algorithm that is pre-trained and used to perform a specific task (in this case, object detection).
[0296] "Subject" refers to an entity with specific characteristics, such as a person or object, that exists in the video.
[0297] A "suspicious object" is an object that is out of place or has been left unattended for a period of time and may pose a threat to normal activity.
[0298] "Lost property" refers to property that has been unintentionally abandoned or forgotten.
[0299] An "emotion engine" is software or algorithms that analyze the emotional state (e.g., stress, anxiety) of participants based on video and audio data.
[0300] An "alert" refers to a warning or notification message in response to a specific event (e.g., the discovery of a suspicious object).
[0301] "Notification" means the act of sending a message to inform interested parties of an alert or other important information.
[0302] The system for implementing this invention includes a surveillance camera, a server, a terminal, a user, and an emotion engine. Each component functions as follows.
[0303] surveillance cameras
[0304] Surveillance cameras capture video in real time from pre-defined locations. For example, if a surveillance camera is used in a food court in a shopping mall, it will monitor each area of the food court and capture video. The video from the camera is sent to a server via a specified URL or IP address.
[0305] server
[0306] The server acquires video data from the surveillance cameras in real time and processes it to detect suspicious objects and lost items. Specifically, it performs the following processes.
[0307] 1. Video data acquisition and preprocessing:
[0308] The server receives the video stream from the surveillance camera using the RTSP protocol, resizes the frame (e.g., from 1920x1080 to 640x480) and normalizes the pixel values using the OpenCV library.
[0309] 2. Object detection using AI models:
[0310] The preprocessed video frames are input into a generative AI model (e.g., YOLO, SSD) to detect objects (e.g., bags, people, etc.) in each frame, which includes a label for each object and its confidence score.
[0311] 3. Analysis of detection results:
[0312] The system analyzes detected object information to identify suspicious or lost items. In particular, if an unattended bag is left unattended for a certain period of time, it will be identified as a suspicious object. It will also compare the object with past video data to identify the occupant or suspicious person.
[0313] 4. Emotion Recognition with Emotion Engine:
[0314] Analyze the emotional state of the person involved (e.g., stress level, anxiety) using video and audio data, using an emotion recognition model (e.g., Microsoft Azure's Emotion API), and adjust the alert content if the person involved is in a high-stress state.
[0315] 5. Alert Generation and Notification:
[0316] If a suspicious or lost item is identified, an alert is automatically generated, tailored based on the results of the emotion engine, and sent to the relevant device (e.g., smartphone).
[0317] Terminal
[0318] The terminal (e.g., a security guard's smartphone) receives the alert sent from the server and displays it on the screen. For example, a notification such as "There is an unattended bag in the food court on the first floor. Please check." The terminal provides an interface that allows the security guard to input the details of the response after checking the scene.
[0319] User
[0320] The user (e.g., a security guard) checks the notification displayed on the terminal and takes action at the scene. For example, if the owner of an unattended bag is confirmed, the user enters the details of the action into the terminal.
[0321] Specific examples
[0322] The server receives real-time footage from surveillance cameras installed in the food court of a shopping mall and uses a generative AI model to detect unattended bags. If a bag is detected left in a specific area for a certain period of time, the server uses an emotion engine to evaluate the emotional state of the security guard and adjust the alert content. A notification is sent to the security guard's smartphone saying, "There is an unattended bag in the food court on the first floor. Please check." The security guard then heads to the scene, checks the safety of the bag, and enters the details into the terminal.
[0323] Example prompt sentence:
[0324] "Detect unattended bags from surveillance camera footage in a food court and calculate the amount of time they have been left there. Also, recognize the emotional state (stress level) of the people present."
[0325] This system enables highly accurate detection of suspicious objects and lost items in real time and rapid response, improving environmental security. In addition, by combining it with an emotion engine, flexible responses can be made taking into account the emotional state of those involved.
[0326] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0327] Step 1:
[0328] Acquiring video data
[0329] Server: Specify the URL or IP address of the surveillance camera and receive the video stream in real time using the RTSP protocol. The input is the video data from the surveillance camera, and the output is the captured real-time video frames. The server captures the video at a rate of 30 frames per second and stores it in memory.
[0330] Step 2:
[0331] Video data preprocessing
[0332] Server: Converts captured video frames into a format that is easy for the generative AI model to process. Specifically, it uses the OpenCV library to resize the frames (e.g., from 1920x1080 to 640x480) and normalizes pixel values to the range of 0-255. The input is a video frame captured in real time, and the output is a preprocessed video frame.
[0333] Step 3:
[0334] Object detection with AI models
[0335] Server: Inputs preprocessed video frames into a generative AI model (e.g., YOLO, SSD) to detect objects in each frame. The input is the preprocessed video frames, and the output is each object's label (e.g., bag, person), its coordinates, and a confidence score. The server updates the list of detected objects for each frame.
[0336] Step 4:
[0337] Analysis of detection results
[0338] Server: Analyzes the detected object information and identifies suspicious or lost objects. For example, if an unattended bag is left unattended for a certain period of time, it will be recorded as a suspicious object. The input is a list of detected objects and a confidence score, and the output is a list of suspicious or lost objects. The server analyzes the object status based on specific criteria and stores the results in a database.
[0339] Step 5:
[0340] Comparison with past data
[0341] Server: Compares newly detected suspicious or lost objects with past video data. Specifically, it compares the object with past records to identify the occupant or suspicious person. The input is a list of suspicious or lost objects and past video data, and the output is the identification information of the occupant or suspicious person. The server compares the object with past video frames stored in a database and analyzes the object's history.
[0342] Step 6:
[0343] Emotion recognition by emotion engine
[0344] Server: An emotion engine uses video and audio data to recognize the emotional state of the participants. This uses an emotion recognition model (e.g., Microsoft Azure's Emotion API). The input is video and audio data captured in real time, and the output is the participants' emotional information (e.g., stress level, anxiety). The server analyzes the participants' facial expressions and tone of voice to evaluate their emotional state.
[0345] Step 7:
[0346] Alerting and Tuning
[0347] Server: Detects suspicious or lost items and adjusts the alert content based on the results of the emotion engine. The input is a list of suspicious or lost items and emotional information, and the output is an alert message sent to the relevant person. For example, a person in a high stress state might be notified in a gentler way, such as "There is a bag in the food court on the first floor. Please check it." The server adjusts the alert content and sends it to the relevant person's device.
[0348] Step 8:
[0349] Record of correspondence
[0350] Server: Records the responses of relevant parties to generated alerts. The input is the response details from the relevant parties, and the output is the recorded response data. The server stores the response results entered by the relevant parties into a database (e.g., confirming the owner of the bag, confirming there is no threat) for later analysis and system improvement.
[0351] (Application example 2)
[0352] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0353] Conventional surveillance systems have limitations in accuracy and real-time detection of suspicious objects and lost items, making it difficult to respond quickly. Furthermore, because the emotional state of the user (such as a security guard) is not taken into consideration, the content of notifications and response methods cannot be optimized, resulting in stress for those involved and efficiency issues. The objective of this invention is to enable accurate real-time detection of suspicious objects and lost items while also enabling appropriate notifications and responses according to the user's emotional state.
[0354] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0355] In this invention, the server includes means for acquiring video data from a monitoring device in real time, means for preprocessing the acquired video data and inputting it into a generative AI model for detecting objects, means for detecting objects using the generative AI model and identifying suspicious objects or lost items, means for recognizing a user's emotional state using an emotion engine and reflecting that information in the analysis, and means for generating an alert and notifying the user according to the user's emotional state when a suspicious object or lost item is detected. This enables highly accurate detection of suspicious objects and lost items in real time, and also enables detailed notifications and responses according to the user's emotional state.
[0356] A "monitoring device" is a device that acquires video data in real time and transmits it to a server.
[0357] "Means for acquiring" refers to a method or device for collecting video data from a surveillance device.
[0358] "Preprocessing" is the process of converting the acquired video data into a format that is easy for the generative AI model to process.
[0359] A "generative AI model" is a program that uses machine learning and deep learning techniques to detect and analyze objects in video data.
[0360] The term "object" refers to an object that is the target of detection within the video data.
[0361] The "emotion engine" is a program that analyzes video and audio data to recognize the user's emotional state.
[0362] "Reflecting in the analysis" means incorporating the emotional information recognized by the emotion engine into the analysis results of the generative AI model.
[0363] "Alert" means a warning or notification generated when a suspicious or lost item is detected.
[0364] "Means for notifying" refers to a method or device for communicating a generated alert to a user.
[0365] The system for implementing this invention is composed of a monitoring device, a server, a terminal, a user, and an emotion engine. Each element will be described in detail below.
[0366] The server acquires video data from the surveillance equipment in real time. The acquired video data undergoes preprocessing before being input into the generative AI model. This preprocessing includes data processing such as resizing and normalizing the video frames. The preprocessed video data is then input into the generative AI model for object detection. The generative AI model achieves highly accurate object detection using traditional machine learning and deep learning techniques. Detected objects are analyzed to identify suspicious objects or lost items. By comparing the data with past video data, it is also possible to identify the occupant of the object or a suspicious person.
[0367] The emotion engine recognizes the user's emotions from video and audio data. The emotion information recognized by the emotion engine is reflected in the analysis on the server. This emotion information is used when generating alerts, and the alert content and notification method are adjusted according to the user's emotional state.
[0368] The terminal is a device that is primarily used by users (such as security guards). Notifications from the server are sent to the terminal, and an alert is displayed to the user. The content of the alert is adjusted based on emotional information recognized by the emotion engine. For example, if a security guard is in a state of high stress, the content of the notification can be softened to reduce the stress.
[0369] The user is the person who primarily operates the terminal and takes appropriate action based on notifications from the system. The user checks the notification and takes action on-site. After that, they provide feedback to the server by entering the details of their response into the terminal. This feedback information can be used for subsequent analysis and system improvement.
[0370] As a concrete example, consider the case of monitoring a food court in a shopping mall. Surveillance cameras send video data from inside the food court to a server in real time. The server preprocesses the video data and uses a generative AI model to detect unattended bags as suspicious objects. It confirms that the detected bag has been left unattended for a certain period of time and compares it with past video data to identify the owner. It also uses an emotion engine to evaluate the stress level of the security guard. Based on the detection results, an adjusted alert is generated and notified to the security guard. The security guard's terminal displays the message, "There is a bag in the food court on the first floor. Please check."
[0371] An example of a prompt sentence can be expressed as follows:
[0372] "Design a system that uses a generative AI model to detect suspicious objects and lost items from video data from surveillance equipment, and recognizes the user's emotional state using an emotion engine. Based on the recognition results, notify the user and encourage them to take the appropriate action."
[0373] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0374] Step 1:
[0375] The server acquires video data from the monitoring device in real time. To acquire video data, the server receives the video stream via the IP address or URL of the monitoring device. The video data is then passed to the server as input. The acquired video data is then saved in the server and formatted for subsequent processing.
[0376] Step 2:
[0377] The server preprocesses the acquired video data, specifically by resizing and normalizing the video frames. The input to this step is the raw video frames, and the output is preprocessed video frames in a format that is easy for the generative AI model to analyze.
[0378] Step 3:
[0379] The server inputs the preprocessed video data into a generative AI model to detect objects. Specifically, the generative AI model detects objects in the video frames and calculates their labels (e.g., bag, person, etc.) and confidence scores. The input for this step is the preprocessed video frames, and the output is a list of detected objects.
[0380] Step 4:
[0381] The server analyzes the confidence scores of the objects detected by the generative AI model and identifies them as suspicious or lost. Specifically, the server compares each object's label and confidence score with the set criteria to determine whether it is suspicious or lost. The input for this step is a list of detected objects, and the output is the identification of suspicious or lost objects.
[0382] Step 5:
[0383] The server compares the detected object with past video data and identifies the occupant or suspicious person of the object. Specifically, it searches for related video data from a past database and confirms the object's presence time and past owners. The input to this step is the identification result of the suspicious object or lost item, and the output is detailed information about the object (the identification result of the occupant or suspicious person).
[0384] Step 6:
[0385] The server uses an emotion engine to recognize the user's emotional state from video and audio data. Specifically, it inputs video frames and audio clips into the emotion engine and outputs an emotion label (e.g., high stress, low stress). The input for this step is the current video frame or audio data, and the output is the user's emotional state.
[0386] Step 7:
[0387] When a suspicious object or lost property is detected, the server generates an alert based on the recognition results of the emotion engine and notifies the user. Specifically, it creates an alert message (e.g., "Caution: A highly suspicious object has been detected. Please act with caution") according to the security guard's emotional state and sends it to the terminal. The input of this step is the suspicious object identification result and the emotional state, and the output is a notification alert.
[0388] Step 8:
[0389] The terminal displays the alert sent from the server and notifies the user. Specifically, the notification alert pops up on the screen, prompting the user to confirm it. The user then takes appropriate action based on the notification and inputs the response results into the terminal. The input for this step is the notification alert, and the output is the response results.
[0390] Step 9:
[0391] The server records the response results received from the terminal and uses them for later analysis and system improvement. Specifically, it saves the recorded response results in a database and uses them later for system evaluation and improvement. The input of this step is the response results, and the output is the recorded response data.
[0392] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0393] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0394] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0395] [Second embodiment]
[0396] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0397] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0398] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0399] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0400] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0401] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0402] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0403] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0404] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0405] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0406] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0407] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0408] The system for implementing this invention consists of a surveillance camera, a server, a terminal, and a user. The server acquires video data from the surveillance camera in real time and analyzes the video data using a generative AI model. The terminal receives notifications from the server, and the user takes action based on the notifications.
[0409] Program processing explanation
[0410] 1. Obtaining surveillance camera footage
[0411] Server: Acquires video data from surveillance cameras in real time. Specifically, it receives video data in stream format via the surveillance camera's URL or IP address.
[0412] 2. Preprocessing of video data
[0413] Server: Preprocesses the acquired video data to input it into the generative AI model. Preprocessing includes resizing and normalizing the video frames.
[0414] 3. Object detection using AI models
[0415] Server: The generative AI model detects objects in preprocessed video frames, and calculates a label and its confidence score for each object.
[0416] 4. Analysis of detection results
[0417] Server: Identifies suspicious or lost objects based on the confidence score of the detected object. Criteria for suspicious or lost objects are set in advance, and judgments are made based on those criteria.
[0418] 5. Comparison with past data
[0419] Server: Detected objects are compared with past video data to identify the object's occupant or suspicious individuals. This improves the accuracy of determining whether an object's status is suspicious.
[0420] 6. Alert Generation and Notification
[0421] Server: If an item is determined to be suspicious or lost, it automatically generates an alert and notifies the relevant parties via email, SMS, or a dedicated application.
[0422] Terminal: The relevant person's terminal receives the notification and displays the alert content. The user checks the notification and takes appropriate action.
[0423] 7. Record of correspondence
[0424] Server: Records the actions taken in response to generated alerts, which are used for later analysis and system improvement.
[0425] Specific examples
[0426] A specific example from a shopping mall is shown below.
[0427] Surveillance camera: Monitoring the food court in the shopping mall.
[0428] Server: Captures surveillance camera footage in real time and analyzes it using a generative AI model. The model detects unattended bags as suspicious objects. It checks that unattended bags in the footage have been left unattended for a certain period of time. It compares this with past data to identify the culprit or owner. Based on these results, it generates an alert and notifies security guards.
[0429] Terminal: The security guard's smartphone receives an alert and displays a notification saying, "There is an unattended bag in the food court on the first floor. Please check."
[0430] User (security guard): The security guard checks the notification, goes to the scene, checks the safety of the bag, and enters the details of the action into the terminal and records it on the server.
[0431] This will enable early detection of suspicious objects and lost items in real time and prompt response, improving security within shopping malls.
[0432] The processing flow will be explained below.
[0433] Step 1:
[0434] The server acquires video data from the surveillance camera in real time, receiving the video stream via the surveillance camera's URL or IP address and capturing it as video data.
[0435] Step 2:
[0436] The server converts the acquired video data into a format that is easy for the generative AI model to process, specifically by resizing the video frames to 224x224 pixels and normalizing them to match the model's input format.
[0437] Step 3:
[0438] The server inputs the preprocessed video data into a generative AI model to detect objects, which then calculates object labels (e.g., bag, person, etc.) and their confidence scores.
[0439] Step 4:
[0440] The server receives the detection results from the generative AI model and evaluates whether the confidence score exceeds a certain threshold. If the detection result exceeds the threshold, it is listed as a candidate for suspicious or lost items.
[0441] Step 5:
[0442] The server compares the listed suspicious or lost items with past video data. By using past data, it can identify the occupant of the object or suspicious individuals, improving the accuracy of the judgment.
[0443] Step 6:
[0444] The server automatically generates an alert if an object is detected as suspicious or lost, including the object's label, location, and time stamp.
[0445] Step 7:
[0446] The server notifies the relevant parties of the generated alerts via email, SMS or a dedicated application.
[0447] Step 8:
[0448] The device receives the notification sent from the server and displays the alert content. The relevant parties can check this notification and take appropriate action.
[0449] Step 9:
[0450] The user (security guard) checks the alert displayed on the terminal, heads to the scene to check the safety of the object, and then inputs the details of the response into the terminal.
[0451] Step 10:
[0452] The server records the responses from users (security guards), which are used for future analysis and system improvement.
[0453] Example 1
[0454] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0455] Although there are systems that analyze surveillance footage in real time and quickly identify and notify suspicious or lost items, it is difficult to respond with high accuracy and efficiency.In addition, there is a lack of cross-checking with past video data and management of response records, so there is a need for stronger crime prevention measures.
[0456] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0457] In this invention, the server includes means for acquiring video data from a monitoring device in real time, means for preprocessing the acquired video data and inputting it into an artificial intelligence model for detecting objects, means for detecting objects using the artificial intelligence model and identifying suspicious objects or lost items, means for generating and notifying relevant parties when a suspicious object or lost item is detected, means for identifying the occupant of the object or a suspicious person using past video data, and means for recording detected information based on the detection of a suspicious object or lost item and using it for later analysis and improvement, thereby enabling highly accurate detection in real time and rapid response.
[0458] "Monitoring device" refers to a device for acquiring video data in real time.
[0459] "Video data" refers to real-time video information obtained from surveillance equipment.
[0460] "Preprocessing" refers to a series of steps that are performed on the captured video data before it can be input into an AI model. Specifically, this includes processes such as resizing and normalizing the video frames.
[0461] "Artificial intelligence model" refers to the algorithms and learning models used to analyze acquired video data and detect objects.
[0462] "Objects" refer to objects that the AI model detects in the video data, including suspicious objects and lost items.
[0463] A "suspicious object" is an object that is determined to be abnormal or suspicious based on certain criteria.
[0464] "Lost property" refers to an object that has been determined to be unclaimed or abandoned.
[0465] An "alert" refers to warning information used to notify relevant parties when a suspicious object or lost item is detected.
[0466] "Interested parties" refers to people or organizations that should receive the alert, including security guards and administrators.
[0467] "Past video data" refers to video information that was recorded in the past among the acquired video data.
[0468] "Matching" refers to the process of comparing currently detected objects with past video data to identify the occupant or suspicious person of a matching object.
[0469] "Response record" refers to data used to store the history of the detection of suspicious or lost items and the response taken thereto.
[0470] The system for implementing this invention consists of a monitoring device, a server, a terminal, and a user. The server acquires video data from the monitoring device in real time, preprocesses the data, and inputs it into a generative AI model for analysis. The terminal receives notifications from the server, and the user responds based on the notifications.
[0471] The server is equipped with hardware and software for acquiring, preprocessing, and analyzing video data. Specifically, libraries such as OpenCV and FFmpeg are used to perform preprocessing such as resizing and normalizing video frames. Generative AI models such as YOLO (You Only Look Once) and SSD (Single Shot Multibox Detector) are used to detect objects in the preprocessed video data.
[0472] As an example, consider surveillance in a shopping mall. Video data is acquired from a surveillance device installed in the mall's food court and received by a server. The server resizes and normalizes the acquired video data (for example, to 640x480 pixels) and inputs it into the YOLO v4 model. The model detects suspicious objects, such as unattended bags, with high accuracy. An object is deemed suspicious if the confidence score is 0.85 or higher.
[0473] The server then compares the detected suspicious object with video data from the past 24 hours to identify the occupant or suspicious person. This reduces false positives and enables more accurate identification of suspicious objects. It automatically generates an alert and sends a notification to the device of the relevant person (e.g., security guard). This notification can be sent via email, SMS, or a dedicated application.
[0474] A specific example of the notification content is a message such as "There is an unattended bag in the food court on the first floor. Please check." Security personnel receive this notification and rush to the scene to respond. The response at that time is recorded on the server via the terminal. The recorded information is used for future analysis and system improvements.
[0475] Prompt Sentence Examples
[0476] Example of using security systems in shopping malls:
[0477] The surveillance cameras capture real-time footage of the food court and send it to the server. The server uses a generative AI model (YOLO v4) to detect unattended bags and compares the detection results with video data from the past 24 hours to identify suspicious individuals. If a suspicious object is determined based on the specified criteria, the security guard's smartphone should be notified, "There is an unattended bag in the 1st floor food court. Please take a look." The server should record the response results.
[0478] This invention enables highly accurate real-time detection of suspicious objects and rapid response, improving security in shopping malls and other public places.
[0479] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0480] Step 1: Obtaining security camera footage
[0481] Specific behavior:
[0482] Server: Obtain video data in real time using the IP address or URL of the surveillance device. For example, use OpenCV to capture the camera stream with cv2.VideoCapture.
[0483] Input: Video data from a surveillance device (stream format)
[0484] Output: Captured video frames
[0485] Example: If the URL of the monitoring device is rtsp: / / 192.168.1.10 / stream, use this URL to get the video.
[0486] Step 2: Preprocessing the video data
[0487] Specific behavior:
[0488] Server: Resize and normalize the captured video frames. To resize, use OpenCV's cv2.resize function, and to normalize, scale pixel values to the range 0 to 1.
[0489] Input: Captured video frame
[0490] Output: Preprocessed video frame (resized image data)
[0491] Example: Resize and normalize the frame to 640x480 pixels.
[0492] Step 3: Object detection with an AI model
[0493] Specific behavior:
[0494] Server: Inputs preprocessed video frames into a generative AI model (such as YOLO) to perform object detection, and outputs object labels and confidence scores.
[0495] Input: Preprocessed video frames
[0496] Output: Detected objects (including labels and confidence scores)
[0497] Example: Detect unattended bags using YOLO v4 model. If the confidence score is 0.85 or higher, it is considered suspicious.
[0498] Step 4: Analyzing the detection results
[0499] Specific behavior:
[0500] Server: Identifies suspicious or lost objects based on the confidence score and label of the detected object. If the confidence score exceeds a certain threshold, the object is deemed suspicious.
[0501] Input: Detected objects (labels and confidence scores)
[0502] Output: Suspicious or lost item determination result
[0503] Example: If the "bag" label is detected with a confidence score of 0.85 or higher, identify it as suspicious.
[0504] Step 5: Check against historical data
[0505] Specific behavior:
[0506] Server: For detected suspicious objects, retrieves video data from the past 24 hours from a database and compares it to identify occupants or suspicious individuals. Past frames are input into the AI model again to confirm the object.
[0507] Input: Information on detected suspicious objects, past video data
[0508] Output: Information on identified occupants or suspicious individuals
[0509] Example: Matching data from the past 24 hours to determine who the bag belongs to or was involved in it.
[0510] Step 6: Alerting and Notification
[0511] Specific behavior:
[0512] Server: If an item is determined to be suspicious or lost, an alert is automatically generated and relevant parties are notified via email, SMS, or a dedicated application.
[0513] Terminal: The terminal of the person involved (security guard) receives the notification and displays a pop-up notification.
[0514] Input: Suspicious object judgment result
[0515] Output: Alert notification
[0516] Example: A security guard's smartphone displays a notification such as, "There is an unattended bag in the food court on the first floor. Please check it."
[0517] Step 7: Record your response
[0518] Specific behavior:
[0519] Server: Records the actions taken in response to generated alerts. This includes the date and time of the action, the person who took the action, and the action itself. Records this in a database.
[0520] Input: Detailed information about the response (date, time, responder, content)
[0521] Output: Corresponding record
[0522] Example: A security guard inspects the scene, enters the details into a terminal, and the information is recorded on the server.
[0523] The above steps create a system that can detect suspicious objects with high accuracy in real time and respond quickly.
[0524] (Application example 1)
[0525] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0526] In security systems that use surveillance cameras, the detection of suspicious objects and lost items often relies on manual monitoring, which can easily lead to oversights. Existing systems also lack the ability to notify relevant parties of detection results in real time, resulting in information delays and making it difficult to respond quickly. Furthermore, the system lacks the ability to identify suspicious individuals and related activities using past video data, creating a need for enhanced security.
[0527] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0528] In this invention, the server
[0529] A means for acquiring video data from a surveillance camera in real time;
[0530] A means for pre-processing the acquired video data and inputting it into a generative AI model for object detection;
[0531] A means for detecting objects using a generative AI model and identifying suspicious or lost items;
[0532] A means for generating alerts and notifying relevant parties about detected suspicious or lost items via email or a dedicated application;
[0533] a means for displaying detection results and alerts to relevant personnel in real time using a head-mounted display;
[0534] This enables highly accurate detection and rapid response in real time.
[0535] A "surveillance camera" is a device that acquires video data in real time and monitors a surveillance area.
[0536] "Video data" refers to a video stream captured by a surveillance camera, and is a set of image information that is subject to analysis and monitoring.
[0537] "Preprocessing" refers to the process of converting video data into a format suitable for the generative AI model, and specifically includes resizing, normalization, noise removal, etc.
[0538] A "generative AI model" is a technology that uses deep learning and machine learning algorithms to generate models, learn patterns from input data, and extract features.
[0539] A "target" is an object that exists in a surveillance camera image and that needs to be identified or detected.
[0540] A "suspicious object" is an object within a surveillance area that is deemed abnormal or dangerous.
[0541] "Lost property" is an object that has been left behind with no known owner.
[0542] An "alert" is a notification that warns or calls attention to detected suspicious objects or lost items.
[0543] "Email" is a means of communication for sending and receiving messages over the Internet.
[0544] A "dedicated application" is software developed for a specific purpose, and is a system designed to provide notifications and information.
[0545] A "head-mounted display" is a wearable display device that displays information within the user's field of vision when worn by the user.
[0546] "Interested parties" are those who should be notified by the system, and are typically security personnel such as security guards or administrators.
[0547] The system for implementing this invention consists of a surveillance camera, a server, a head-mounted display terminal, and a user (mainly a security guard). The server acquires video data from the surveillance camera in real time and analyzes the video data using a generative AI model. The head-mounted display (HMD) terminal worn by the security guard receives notifications from the server and displays the notification contents to the security guard in real time.
[0548] Hardware and software used
[0549] Hardware: surveillance cameras, servers, head-mounted displays (HMDs)
[0550] Software: TensorFlow (generative AI model), OpenCV (image processing)
[0551] Details of data processing and calculation
[0552] The server receives video data from the surveillance cameras in real time. This data is first preprocessed using OpenCV. Specifically, the preprocessing involves resizing and normalizing the video frames.
[0553] The pre-processed video data is then fed into a generative AI model for object detection. The generative AI model, built using TensorFlow, detects objects in the video and calculates a label and confidence score.
[0554] If any of the detected objects are deemed suspicious or lost, the server compares them with past video data to identify the object's occupant or suspicious person.
[0555] The server generates an alert based on the results and automatically notifies the relevant parties (security guards) via email or a dedicated application. The head-mounted display terminal worn by the guard also displays the detection results and alert details in real time.
[0556] Specific examples
[0557] For example, if a suspicious object (an unattended bag) is left unattended for a certain period of time in a food court in a shopping mall, the system will operate as follows:
[0558] Surveillance cameras capture footage of the food court.
[0559] The server acquires the video data in real time, preprocesses it, and inputs it into the generative AI model.
[0560] The generative AI model detects unattended bags, analyzes the information and compares it with historical data.
[0561] The server checks for the presence of suspicious objects and generates an alert.
[0562] The security guard's head-mounted display displays a real-time message saying, "There is an unattended bag in the food court on the first floor. Please check."
[0563] Prompt Sentence Examples
[0564] For example, use the following prompt:
[0565] "Monitoring the food court for 45 minutes. Detecting an unattended bag and confirming that it had been left there for 20 minutes. Checking against past data to identify the owner. Notifying a security officer wearing an HMD that 'There is an unattended bag in the food court on the first floor. Please check.'"
[0566] This configuration enables highly accurate detection of suspicious objects and lost items within the surveillance area, enabling rapid response in real time.
[0567] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0568] Step 1:
[0569] The server acquires video data from the surveillance camera in real time. The input is the video stream sent from the surveillance camera, and the output is the raw video data stored in the server. The video data is received in stream format via the URL or IP address of the surveillance camera.
[0570] Step 2:
[0571] The server performs preprocessing on the acquired video data before inputting it into the generative AI model. The input is the raw video data acquired in step 1, and the output is preprocessed video frames. Preprocessing includes resizing, normalizing, and denoising the video frames. Specifically, OpenCV is used to resize the frames to the required size and normalize the pixel values.
[0572] Step 3:
[0573] The server inputs the preprocessed video data into a generative AI model to detect objects. The input is the video frames preprocessed in step 2, and the output is the labels and confidence scores of the detected objects. A generative AI model using TensorFlow is used to identify objects in the video frames. Each object is assigned a label and its confidence score.
[0574] Step 4:
[0575] The server identifies suspicious or lost objects based on the confidence score of the detected object. The input is the object label and confidence score generated in step 3, and the output is the result of the suspicious or lost object judgment. Criteria for suspicious or lost objects are set in advance, and the judgment is made according to those criteria, and a flag is raised if action is required.
[0576] Step 5:
[0577] The server compares the detected object with past video data to identify the object's occupant or suspicious person. The input is the judgment result from step 4 and past video data, and the output is the identification of the owner or suspicious person of the suspicious object or lost item. This improves the accuracy of the judgment by comparing with past behavioral history.
[0578] Step 6:
[0579] If the server determines that an item is suspicious or lost, it automatically generates an alert and notifies the relevant parties. The input is the determination result and identification result, and the output is the generated alert and notification message. The alert is sent via email or a dedicated application.
[0580] Step 7:
[0581] The terminal (head-mounted display) receives an alert notification from the server and displays the notification content to the security guard. The input is the alert notification sent from the server, and the output is the warning message displayed on the head-mounted display. The notification content is specific, such as "There is an unattended bag in the food court on the first floor. Please check it."
[0582] Step 8:
[0583] The user (security guard) checks the notification displayed on the head-mounted display, heads to the scene, and takes action. The input is the alert content displayed on the head-mounted display, and the output is the execution and reporting of the response. The report content is entered into the terminal and recorded on the server.
[0584] These steps enable highly accurate detection of suspicious or lost items in real time and rapid response.
[0585] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0586] The system embodying this invention consists of a surveillance camera, a server, a terminal, a user, and an emotion engine. The server acquires video data from the surveillance camera in real time and analyzes the video data using a generative AI model. The emotion engine recognizes the user's emotions and incorporates that information into the analysis. The terminal receives notifications from the server, and the user responds based on those notifications.
[0587] Program processing explanation
[0588] 1. Obtaining surveillance camera footage
[0589] Server: Acquires video data from surveillance cameras in real time. Receives video streams via the surveillance camera's URL or IP address and acquires them as video data.
[0590] 2. Preprocessing of video data
[0591] Server: Converts acquired video data into a format that is easy for the generative AI model to process. Preprocessing includes resizing and normalizing video frames.
[0592] 3. Object detection using AI models
[0593] Server: The generative AI model detects objects in pre-processed video frames, and calculates a label for each object (e.g., bag, person, etc.) and its confidence score.
[0594] 4. Analysis of detection results
[0595] Server: Identifies suspicious or lost objects based on the confidence score of the detected object. Criteria for suspicious or lost objects are set in advance, and judgments are made based on those criteria.
[0596] 5. Comparison with past data
[0597] Server: Detected objects are compared with past video data to identify the object's occupant or suspicious individuals. This improves the accuracy of determining whether an object's status is suspicious.
[0598] 6. Emotion Recognition by Emotion Engine
[0599] Server: Utilizes an emotion engine that uses video and audio data to recognize the user's emotions, thereby obtaining emotional information such as whether the person involved is feeling stressed.
[0600] 7. Alert Generation and Coordination
[0601] Server: If an item is determined to be suspicious or lost, an alert is automatically generated. The content of the alert and notification method are adjusted based on the recognition results of the emotion engine. For example, if the person involved is under high stress, the content of the notification may be softened.
[0602] Terminal: A notification is sent to the relevant person's terminal and the alert details are displayed. The user checks this notification and takes appropriate action.
[0603] 8. Record of correspondence
[0604] Server: Records the actions taken in response to generated alerts, which are used for later analysis and system improvement.
[0605] Specific examples
[0606] Here is a specific example from inside a shopping mall.
[0607] Surveillance camera: Monitoring the food court in the shopping mall.
[0608] Server: Captures surveillance camera footage in real time and analyzes it using a generative AI model. The model detects unattended bags as suspicious objects. It checks that unattended bags in the footage have been left unattended for a certain period of time and compares them with past data to identify the owner. It also uses an emotion engine to evaluate the stress level of security guards. Based on the detection results, it generates tailored alerts and notifies security guards.
[0609] Device: The security guard's smartphone receives an alert and displays a notification saying, "There is an unattended bag in the food court on the first floor. Please check." If the security guard is in a high-stress state, the notification content is changed to, "There is an unattended bag in the food court on the first floor. Please check."
[0610] User (security guard): The security guard checks the notification, heads to the scene, checks the safety of the bag, and then enters the details of the action into the terminal.
[0611] Server: Records the responses from security guards and uses them for later analysis.
[0612] This system will improve security within shopping malls by enabling early detection and rapid response of suspicious or lost items in real time. In addition, by combining it with an emotion engine, it will be possible to respond taking into account the emotional state of those involved.
[0613] The processing flow will be explained below.
[0614] Step 1:
[0615] The server acquires video data from the surveillance camera in real time, receives the video stream via the surveillance camera's URL or IP address, and acquires the video data frame by frame.
[0616] Step 2:
[0617] The server converts the acquired video data into a format that is easy for the generative AI model to process, specifically by resizing the video frames to 224x224 pixels and normalizing them to match the model's input format.
[0618] Step 3:
[0619] The server inputs the preprocessed video data into a generative AI model for object detection, which identifies each object in the frame and outputs a label (e.g., bag, person, etc.) and its confidence score.
[0620] Step 4:
[0621] The server analyzes the output of the generative AI model and evaluates whether the confidence score exceeds a certain threshold. Detections that exceed the threshold are listed as suspicious or lost items.
[0622] Step 5:
[0623] The server compares the listed suspicious or lost items with past video data, and based on the past data, identifies the occupant of the object or the suspicious person, and confirms that the object is a suspicious or lost item.
[0624] Step 6:
[0625] The server uses the video and audio data to recognize the user's emotions through an emotion engine, analyzing the user's facial expressions and tone of voice to assess their emotional state, such as their stress level.
[0626] Step 7:
[0627] The server automatically generates an alert if the item is determined to be a suspicious or lost item. The content of the alert and notification method are adjusted based on the results of the emotion engine. For example, if the user is in a state of high stress, the notification content will be changed to a softer expression.
[0628] Step 8:
[0629] The server notifies the relevant parties of the generated alerts via email, SMS, or a dedicated application sent to the device.
[0630] Step 9:
[0631] The terminal receives the notification from the server and displays the alert content. The user can check this notification and take the necessary action.
[0632] Step 10:
[0633] The user (security guard) checks the alert displayed on the terminal, goes to the scene to check the object, and then inputs the response details into the terminal.
[0634] Step 11:
[0635] The server records the responses entered by the user (security guard), which will be used for future analysis and system improvement.
[0636] As a concrete example, consider the case where an unattended bag is left in the food court of a shopping mall. The server detects the unattended bag from surveillance camera footage. It compares this with past data to identify the bag's owner. It uses an emotion engine to check the stress level of the security guard and generates an alert in appropriate language to notify the guard. The security guard then checks the notification, checks the safety of the bag on-site, and enters the results into the server. This series of processes enables effective and human-friendly security operations.
[0637] Example 2
[0638] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0639] Conventional surveillance systems detect suspicious objects or lost items, but do not take into account the emotional state of those involved, which can lead to stress when dealing with the incident. Furthermore, detection accuracy can be low because the system does not compare the detected items with past video data. Furthermore, the recording and subsequent analysis of detected information is insufficient, which can hinder system improvement.
[0640] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for acquiring video data from a surveillance camera in real time, means for preprocessing the acquired video data and inputting it into a generative AI model for detecting objects, means for detecting objects using the generative AI model and identifying suspicious or lost objects, means for evaluating the emotional state of relevant parties using an emotion engine and adjusting the content of an alert when a suspicious or lost object is detected, and means for generating and notifying relevant parties of an alert. This enables highly accurate detection of suspicious or lost objects and appropriate responses that take into account the emotional state of relevant parties. In addition, detection accuracy is improved by comparing with past video data, and continuous system improvement is possible through recording and subsequent analysis of detection information.
[0641] A "surveillance camera" is a device that captures images of a pre-set location in real time.
[0642] A "server" is a computer system that obtains, processes, and manages data from other devices over a network.
[0643] "Real-time" refers to processing immediately without delay, and refers to a state in which appropriate action can be taken the moment a specific event occurs.
[0644] "Video data" means a collection of visual information captured from a surveillance camera or other device and stored as video frames.
[0645] "Preprocessing" refers to a series of operations that transform data into a suitable format before it is sent to the main processing stage, including, for example, resizing and normalization.
[0646] A "generative AI model" is an artificial intelligence algorithm that is pre-trained and used to perform a specific task (in this case, object detection).
[0647] "Subject" refers to an entity with specific characteristics, such as a person or object, that exists in the video.
[0648] A "suspicious object" is an object that is out of place or has been left unattended for a period of time and may pose a threat to normal activity.
[0649] "Lost property" refers to property that has been unintentionally abandoned or forgotten.
[0650] An "emotion engine" is software or algorithms that analyze the emotional state (e.g., stress, anxiety) of participants based on video and audio data.
[0651] An "alert" refers to a warning or notification message in response to a specific event (e.g., the discovery of a suspicious object).
[0652] "Notification" means the act of sending a message to inform interested parties of an alert or other important information.
[0653] The system for implementing this invention includes a surveillance camera, a server, a terminal, a user, and an emotion engine. Each component functions as follows.
[0654] surveillance cameras
[0655] Surveillance cameras capture video in real time from pre-defined locations. For example, if a surveillance camera is used in a food court in a shopping mall, it will monitor each area of the food court and capture video. The video from the camera is sent to a server via a specified URL or IP address.
[0656] server
[0657] The server acquires video data from the surveillance cameras in real time and processes it to detect suspicious objects and lost items. Specifically, it performs the following processes.
[0658] 1. Video data acquisition and preprocessing:
[0659] The server receives the video stream from the surveillance camera using the RTSP protocol, resizes the frame (e.g., from 1920x1080 to 640x480) and normalizes the pixel values using the OpenCV library.
[0660] 2. Object detection using AI models:
[0661] The preprocessed video frames are input into a generative AI model (e.g., YOLO, SSD) to detect objects (e.g., bags, people, etc.) in each frame, which includes a label for each object and its confidence score.
[0662] 3. Analysis of detection results:
[0663] The system analyzes detected object information to identify suspicious or lost items. In particular, if an unattended bag is left unattended for a certain period of time, it will be identified as a suspicious object. It will also compare the object with past video data to identify the occupant or suspicious person.
[0664] 4. Emotion Recognition with Emotion Engine:
[0665] Analyze the emotional state of the person involved (e.g., stress level, anxiety) using video and audio data, using an emotion recognition model (e.g., Microsoft Azure's Emotion API), and adjust the alert content if the person involved is in a high-stress state.
[0666] 5. Alert Generation and Notification:
[0667] If a suspicious or lost item is identified, an alert is automatically generated, tailored based on the results of the emotion engine, and sent to the relevant device (e.g., smartphone).
[0668] Terminal
[0669] The terminal (e.g., a security guard's smartphone) receives the alert sent from the server and displays it on the screen. For example, a notification such as "There is an unattended bag in the food court on the first floor. Please check." The terminal provides an interface that allows the security guard to input the details of the response after checking the scene.
[0670] User
[0671] The user (e.g., a security guard) checks the notification displayed on the terminal and takes action at the scene. For example, if the owner of an unattended bag is confirmed, the user enters the details of the action into the terminal.
[0672] Specific examples
[0673] The server receives real-time footage from surveillance cameras installed in the food court of a shopping mall and uses a generative AI model to detect unattended bags. If a bag is detected left in a specific area for a certain period of time, the server uses an emotion engine to evaluate the emotional state of the security guard and adjust the alert content. A notification is sent to the security guard's smartphone saying, "There is an unattended bag in the food court on the first floor. Please check." The security guard then heads to the scene, checks the safety of the bag, and enters the details into the terminal.
[0674] Example prompt sentence:
[0675] "Detect unattended bags from surveillance camera footage in a food court and calculate the amount of time they have been left there. Also, recognize the emotional state (stress level) of the people present."
[0676] This system enables highly accurate detection of suspicious objects and lost items in real time and rapid response, improving environmental security. In addition, by combining it with an emotion engine, flexible responses can be made taking into account the emotional state of those involved.
[0677] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0678] Step 1:
[0679] Acquiring video data
[0680] Server: Specify the URL or IP address of the surveillance camera and receive the video stream in real time using the RTSP protocol. The input is the video data from the surveillance camera, and the output is the captured real-time video frames. The server captures the video at a rate of 30 frames per second and stores it in memory.
[0681] Step 2:
[0682] Video data preprocessing
[0683] Server: Converts captured video frames into a format that is easy for the generative AI model to process. Specifically, it uses the OpenCV library to resize the frames (e.g., from 1920x1080 to 640x480) and normalizes pixel values to the range of 0-255. The input is a video frame captured in real time, and the output is a preprocessed video frame.
[0684] Step 3:
[0685] Object detection with AI models
[0686] Server: Inputs preprocessed video frames into a generative AI model (e.g., YOLO, SSD) to detect objects in each frame. The input is the preprocessed video frames, and the output is each object's label (e.g., bag, person), its coordinates, and a confidence score. The server updates the list of detected objects for each frame.
[0687] Step 4:
[0688] Analysis of detection results
[0689] Server: Analyzes the detected object information and identifies suspicious or lost objects. For example, if an unattended bag is left unattended for a certain period of time, it will be recorded as a suspicious object. The input is a list of detected objects and a confidence score, and the output is a list of suspicious or lost objects. The server analyzes the object status based on specific criteria and stores the results in a database.
[0690] Step 5:
[0691] Comparison with past data
[0692] Server: Compares newly detected suspicious or lost objects with past video data. Specifically, it compares the object with past records to identify the occupant or suspicious person. The input is a list of suspicious or lost objects and past video data, and the output is the identification information of the occupant or suspicious person. The server compares the object with past video frames stored in a database and analyzes the object's history.
[0693] Step 6:
[0694] Emotion recognition by emotion engine
[0695] Server: An emotion engine uses video and audio data to recognize the emotional state of the participants. This uses an emotion recognition model (e.g., Microsoft Azure's Emotion API). The input is video and audio data captured in real time, and the output is the participants' emotional information (e.g., stress level, anxiety). The server analyzes the participants' facial expressions and tone of voice to evaluate their emotional state.
[0696] Step 7:
[0697] Alerting and Tuning
[0698] Server: Detects suspicious or lost items and adjusts the alert content based on the results of the emotion engine. The input is a list of suspicious or lost items and emotional information, and the output is an alert message sent to the relevant person. For example, a person in a high stress state might be notified in a gentler way, such as "There is a bag in the food court on the first floor. Please check it." The server adjusts the alert content and sends it to the relevant person's device.
[0699] Step 8:
[0700] Record of correspondence
[0701] Server: Records the responses of relevant parties to generated alerts. The input is the response details from the relevant parties, and the output is the recorded response data. The server stores the response results entered by the relevant parties into a database (e.g., confirming the owner of the bag, confirming there is no threat) for later analysis and system improvement.
[0702] (Application example 2)
[0703] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0704] Conventional surveillance systems have limitations in accuracy and real-time detection of suspicious objects and lost items, making it difficult to respond quickly. Furthermore, because the emotional state of the user (such as a security guard) is not taken into consideration, the content of notifications and response methods cannot be optimized, resulting in stress for those involved and efficiency issues. The objective of this invention is to enable accurate real-time detection of suspicious objects and lost items while also enabling appropriate notifications and responses according to the user's emotional state.
[0705] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0706] In this invention, the server includes means for acquiring video data from a monitoring device in real time, means for preprocessing the acquired video data and inputting it into a generative AI model for detecting objects, means for detecting objects using the generative AI model and identifying suspicious objects or lost items, means for recognizing a user's emotional state using an emotion engine and reflecting that information in the analysis, and means for generating an alert and notifying the user according to the user's emotional state when a suspicious object or lost item is detected. This enables highly accurate detection of suspicious objects and lost items in real time, and also enables detailed notifications and responses according to the user's emotional state.
[0707] A "monitoring device" is a device that acquires video data in real time and transmits it to a server.
[0708] "Means for acquiring" refers to a method or device for collecting video data from a surveillance device.
[0709] "Preprocessing" is the process of converting the acquired video data into a format that is easy for the generative AI model to process.
[0710] A "generative AI model" is a program that uses machine learning and deep learning techniques to detect and analyze objects in video data.
[0711] The term "object" refers to an object that is the target of detection within the video data.
[0712] The "emotion engine" is a program that analyzes video and audio data to recognize the user's emotional state.
[0713] "Reflecting in the analysis" means incorporating the emotional information recognized by the emotion engine into the analysis results of the generative AI model.
[0714] "Alert" means a warning or notification generated when a suspicious or lost item is detected.
[0715] "Means for notifying" refers to a method or device for communicating a generated alert to a user.
[0716] The system for implementing this invention is composed of a monitoring device, a server, a terminal, a user, and an emotion engine. Each element will be described in detail below.
[0717] The server acquires video data from the surveillance equipment in real time. The acquired video data undergoes preprocessing before being input into the generative AI model. This preprocessing includes data processing such as resizing and normalizing the video frames. The preprocessed video data is then input into the generative AI model for object detection. The generative AI model achieves highly accurate object detection using traditional machine learning and deep learning techniques. Detected objects are analyzed to identify suspicious objects or lost items. By comparing the data with past video data, it is also possible to identify the occupant of the object or a suspicious person.
[0718] The emotion engine recognizes the user's emotions from video and audio data. The emotion information recognized by the emotion engine is reflected in the analysis on the server. This emotion information is used when generating alerts, and the alert content and notification method are adjusted according to the user's emotional state.
[0719] The terminal is a device that is primarily used by users (such as security guards). Notifications from the server are sent to the terminal, and an alert is displayed to the user. The content of the alert is adjusted based on emotional information recognized by the emotion engine. For example, if a security guard is in a state of high stress, the content of the notification can be softened to reduce the stress.
[0720] The user is the person who primarily operates the terminal and takes appropriate action based on notifications from the system. The user checks the notification and takes action on-site. After that, they provide feedback to the server by entering the details of their response into the terminal. This feedback information can be used for subsequent analysis and system improvement.
[0721] As a concrete example, consider the case of monitoring a food court in a shopping mall. Surveillance cameras send video data from inside the food court to a server in real time. The server preprocesses the video data and uses a generative AI model to detect unattended bags as suspicious objects. It confirms that the detected bag has been left unattended for a certain period of time and compares it with past video data to identify the owner. It also uses an emotion engine to evaluate the stress level of the security guard. Based on the detection results, an adjusted alert is generated and notified to the security guard. The security guard's terminal displays the message, "There is a bag in the food court on the first floor. Please check."
[0722] An example of a prompt sentence can be expressed as follows:
[0723] "Design a system that uses a generative AI model to detect suspicious objects and lost items from video data from surveillance equipment, and recognizes the user's emotional state using an emotion engine. Based on the recognition results, notify the user and encourage them to take the appropriate action."
[0724] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0725] Step 1:
[0726] The server acquires video data from the monitoring device in real time. To acquire video data, the server receives the video stream via the IP address or URL of the monitoring device. The video data is then passed to the server as input. The acquired video data is then saved in the server and formatted for subsequent processing.
[0727] Step 2:
[0728] The server preprocesses the acquired video data, specifically by resizing and normalizing the video frames. The input to this step is the raw video frames, and the output is preprocessed video frames in a format that is easy for the generative AI model to analyze.
[0729] Step 3:
[0730] The server inputs the preprocessed video data into a generative AI model to detect objects. Specifically, the generative AI model detects objects in the video frames and calculates their labels (e.g., bag, person, etc.) and confidence scores. The input for this step is the preprocessed video frames, and the output is a list of detected objects.
[0731] Step 4:
[0732] The server analyzes the confidence scores of the objects detected by the generative AI model and identifies them as suspicious or lost. Specifically, the server compares each object's label and confidence score with the set criteria to determine whether it is suspicious or lost. The input for this step is a list of detected objects, and the output is the identification of suspicious or lost objects.
[0733] Step 5:
[0734] The server compares the detected object with past video data and identifies the occupant or suspicious person of the object. Specifically, it searches for related video data from a past database and confirms the object's presence time and past owners. The input to this step is the identification result of the suspicious object or lost item, and the output is detailed information about the object (the identification result of the occupant or suspicious person).
[0735] Step 6:
[0736] The server uses an emotion engine to recognize the user's emotional state from video and audio data. Specifically, it inputs video frames and audio clips into the emotion engine and outputs an emotion label (e.g., high stress, low stress). The input for this step is the current video frame or audio data, and the output is the user's emotional state.
[0737] Step 7:
[0738] When a suspicious object or lost property is detected, the server generates an alert based on the recognition results of the emotion engine and notifies the user. Specifically, it creates an alert message (e.g., "Caution: A highly suspicious object has been detected. Please act with caution") according to the security guard's emotional state and sends it to the terminal. The input of this step is the suspicious object identification result and the emotional state, and the output is a notification alert.
[0739] Step 8:
[0740] The terminal displays the alert sent from the server and notifies the user. Specifically, the notification alert pops up on the screen, prompting the user to confirm it. The user then takes appropriate action based on the notification and inputs the response results into the terminal. The input for this step is the notification alert, and the output is the response results.
[0741] Step 9:
[0742] The server records the response results received from the terminal and uses them for later analysis and system improvement. Specifically, it saves the recorded response results in a database and uses them later for system evaluation and improvement. The input of this step is the response results, and the output is the recorded response data.
[0743] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0744] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0745] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0746] [Third embodiment]
[0747] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0748] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0749] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0750] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0751] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0752] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0753] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0754] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0755] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0756] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0757] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0758] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0759] The system for implementing this invention consists of a surveillance camera, a server, a terminal, and a user. The server acquires video data from the surveillance camera in real time and analyzes the video data using a generative AI model. The terminal receives notifications from the server, and the user takes action based on the notifications.
[0760] Program processing explanation
[0761] 1. Obtaining surveillance camera footage
[0762] Server: Acquires video data from surveillance cameras in real time. Specifically, it receives video data in stream format via the surveillance camera's URL or IP address.
[0763] 2. Preprocessing of video data
[0764] Server: Preprocesses the acquired video data to input it into the generative AI model. Preprocessing includes resizing and normalizing the video frames.
[0765] 3. Object detection using AI models
[0766] Server: The generative AI model detects objects in preprocessed video frames, and calculates a label and its confidence score for each object.
[0767] 4. Analysis of detection results
[0768] Server: Identifies suspicious or lost objects based on the confidence score of the detected object. Criteria for suspicious or lost objects are set in advance, and judgments are made based on those criteria.
[0769] 5. Comparison with past data
[0770] Server: Detected objects are compared with past video data to identify the object's occupant or suspicious individuals. This improves the accuracy of determining whether an object's status is suspicious.
[0771] 6. Alert Generation and Notification
[0772] Server: If an item is determined to be suspicious or lost, it automatically generates an alert and notifies the relevant parties via email, SMS, or a dedicated application.
[0773] Terminal: The relevant person's terminal receives the notification and displays the alert content. The user checks the notification and takes appropriate action.
[0774] 7. Record of correspondence
[0775] Server: Records the actions taken in response to generated alerts, which are used for later analysis and system improvement.
[0776] Specific examples
[0777] A specific example from a shopping mall is shown below.
[0778] Surveillance camera: Monitoring the food court in the shopping mall.
[0779] Server: Captures surveillance camera footage in real time and analyzes it using a generative AI model. The model detects unattended bags as suspicious objects. It checks that unattended bags in the footage have been left unattended for a certain period of time. It compares this with past data to identify the culprit or owner. Based on these results, it generates an alert and notifies security guards.
[0780] Terminal: The security guard's smartphone receives an alert and displays a notification saying, "There is an unattended bag in the food court on the first floor. Please check."
[0781] User (security guard): The security guard checks the notification, goes to the scene, checks the safety of the bag, and enters the details of the action into the terminal and records it on the server.
[0782] This will enable early detection of suspicious objects and lost items in real time and prompt response, improving security within shopping malls.
[0783] The processing flow will be explained below.
[0784] Step 1:
[0785] The server acquires video data from the surveillance camera in real time, receiving the video stream via the surveillance camera's URL or IP address and capturing it as video data.
[0786] Step 2:
[0787] The server converts the acquired video data into a format that is easy for the generative AI model to process, specifically by resizing the video frames to 224x224 pixels and normalizing them to match the model's input format.
[0788] Step 3:
[0789] The server inputs the preprocessed video data into a generative AI model to detect objects, which then calculates object labels (e.g., bag, person, etc.) and their confidence scores.
[0790] Step 4:
[0791] The server receives the detection results from the generative AI model and evaluates whether the confidence score exceeds a certain threshold. If the detection result exceeds the threshold, it is listed as a candidate for suspicious or lost items.
[0792] Step 5:
[0793] The server compares the listed suspicious or lost items with past video data. By using past data, it can identify the occupant of the object or suspicious individuals, improving the accuracy of the judgment.
[0794] Step 6:
[0795] The server automatically generates an alert if an object is detected as suspicious or lost, including the object's label, location, and time stamp.
[0796] Step 7:
[0797] The server notifies the relevant parties of the generated alerts via email, SMS or a dedicated application.
[0798] Step 8:
[0799] The device receives the notification sent from the server and displays the alert content. The relevant parties can check this notification and take appropriate action.
[0800] Step 9:
[0801] The user (security guard) checks the alert displayed on the terminal, heads to the scene to check the safety of the object, and then inputs the details of the response into the terminal.
[0802] Step 10:
[0803] The server records the responses from users (security guards), which are used for future analysis and system improvement.
[0804] Example 1
[0805] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0806] Although there are systems that analyze surveillance footage in real time and quickly identify and notify suspicious or lost items, it is difficult to respond with high accuracy and efficiency.In addition, there is a lack of cross-checking with past video data and management of response records, so there is a need for stronger crime prevention measures.
[0807] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0808] In this invention, the server includes means for acquiring video data from a monitoring device in real time, means for preprocessing the acquired video data and inputting it into an artificial intelligence model for detecting objects, means for detecting objects using the artificial intelligence model and identifying suspicious objects or lost items, means for generating and notifying relevant parties when a suspicious object or lost item is detected, means for identifying the occupant of the object or a suspicious person using past video data, and means for recording detected information based on the detection of a suspicious object or lost item and using it for later analysis and improvement, thereby enabling highly accurate detection in real time and rapid response.
[0809] "Monitoring device" refers to a device for acquiring video data in real time.
[0810] "Video data" refers to real-time video information obtained from surveillance equipment.
[0811] "Preprocessing" refers to a series of steps that are performed on the captured video data before it can be input into an AI model. Specifically, this includes processes such as resizing and normalizing the video frames.
[0812] "Artificial intelligence model" refers to the algorithms and learning models used to analyze acquired video data and detect objects.
[0813] "Objects" refer to objects that the AI model detects in the video data, including suspicious objects and lost items.
[0814] A "suspicious object" is an object that is determined to be abnormal or suspicious based on certain criteria.
[0815] "Lost property" refers to an object that has been determined to be unclaimed or abandoned.
[0816] An "alert" refers to warning information used to notify relevant parties when a suspicious object or lost item is detected.
[0817] "Interested parties" refers to people or organizations that should receive the alert, including security guards and administrators.
[0818] "Past video data" refers to video information that was recorded in the past among the acquired video data.
[0819] "Matching" refers to the process of comparing currently detected objects with past video data to identify the occupant or suspicious person of a matching object.
[0820] "Response record" refers to data used to store the history of the detection of suspicious or lost items and the response taken thereto.
[0821] The system for implementing this invention consists of a monitoring device, a server, a terminal, and a user. The server acquires video data from the monitoring device in real time, preprocesses the data, and inputs it into a generative AI model for analysis. The terminal receives notifications from the server, and the user responds based on the notifications.
[0822] The server is equipped with hardware and software for acquiring, preprocessing, and analyzing video data. Specifically, libraries such as OpenCV and FFmpeg are used to perform preprocessing such as resizing and normalizing video frames. Generative AI models such as YOLO (You Only Look Once) and SSD (Single Shot Multibox Detector) are used to detect objects in the preprocessed video data.
[0823] As an example, consider surveillance in a shopping mall. Video data is acquired from a surveillance device installed in the mall's food court and received by a server. The server resizes and normalizes the acquired video data (for example, to 640x480 pixels) and inputs it into the YOLO v4 model. The model detects suspicious objects, such as unattended bags, with high accuracy. An object is deemed suspicious if the confidence score is 0.85 or higher.
[0824] The server then compares the detected suspicious object with video data from the past 24 hours to identify the occupant or suspicious person. This reduces false positives and enables more accurate identification of suspicious objects. It automatically generates an alert and sends a notification to the device of the relevant person (e.g., security guard). This notification can be sent via email, SMS, or a dedicated application.
[0825] A specific example of the notification content is a message such as "There is an unattended bag in the food court on the first floor. Please check." Security personnel receive this notification and rush to the scene to respond. The response at that time is recorded on the server via the terminal. The recorded information is used for future analysis and system improvements.
[0826] Prompt Sentence Examples
[0827] Example of using security systems in shopping malls:
[0828] The surveillance cameras capture real-time footage of the food court and send it to the server. The server uses a generative AI model (YOLO v4) to detect unattended bags and compares the detection results with video data from the past 24 hours to identify suspicious individuals. If a suspicious object is determined based on the specified criteria, the security guard's smartphone should be notified, "There is an unattended bag in the 1st floor food court. Please take a look." The server should record the response results.
[0829] This invention enables highly accurate real-time detection of suspicious objects and rapid response, improving security in shopping malls and other public places.
[0830] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0831] Step 1: Obtaining security camera footage
[0832] Specific behavior:
[0833] Server: Obtain video data in real time using the IP address or URL of the surveillance device. For example, use OpenCV to capture the camera stream with cv2.VideoCapture.
[0834] Input: Video data from a surveillance device (stream format)
[0835] Output: Captured video frames
[0836] Example: If the URL of the monitoring device is rtsp: / / 192.168.1.10 / stream, use this URL to get the video.
[0837] Step 2: Preprocessing the video data
[0838] Specific behavior:
[0839] Server: Resize and normalize the captured video frames. To resize, use OpenCV's cv2.resize function, and to normalize, scale pixel values to the range 0 to 1.
[0840] Input: Captured video frame
[0841] Output: Preprocessed video frame (resized image data)
[0842] Example: Resize and normalize the frame to 640x480 pixels.
[0843] Step 3: Object detection with an AI model
[0844] Specific behavior:
[0845] Server: Inputs preprocessed video frames into a generative AI model (such as YOLO) to perform object detection, and outputs object labels and confidence scores.
[0846] Input: Preprocessed video frames
[0847] Output: Detected objects (including labels and confidence scores)
[0848] Example: Detect unattended bags using YOLO v4 model. If the confidence score is 0.85 or higher, it is considered suspicious.
[0849] Step 4: Analyzing the detection results
[0850] Specific behavior:
[0851] Server: Identifies suspicious or lost objects based on the confidence score and label of the detected object. If the confidence score exceeds a certain threshold, the object is deemed suspicious.
[0852] Input: Detected objects (labels and confidence scores)
[0853] Output: Suspicious or lost item determination result
[0854] Example: If the "bag" label is detected with a confidence score of 0.85 or higher, identify it as suspicious.
[0855] Step 5: Check against historical data
[0856] Specific behavior:
[0857] Server: For detected suspicious objects, retrieves video data from the past 24 hours from a database and compares it to identify occupants or suspicious individuals. Past frames are input into the AI model again to confirm the object.
[0858] Input: Information on detected suspicious objects, past video data
[0859] Output: Information on identified occupants or suspicious individuals
[0860] Example: Matching data from the past 24 hours to determine who the bag belongs to or was involved in it.
[0861] Step 6: Alerting and Notification
[0862] Specific behavior:
[0863] Server: If an item is determined to be suspicious or lost, an alert is automatically generated and relevant parties are notified via email, SMS, or a dedicated application.
[0864] Terminal: The terminal of the person involved (security guard) receives the notification and displays a pop-up notification.
[0865] Input: Suspicious object judgment result
[0866] Output: Alert notification
[0867] Example: A security guard's smartphone displays a notification such as, "There is an unattended bag in the food court on the first floor. Please check it."
[0868] Step 7: Record your response
[0869] Specific behavior:
[0870] Server: Records the actions taken in response to generated alerts. This includes the date and time of the action, the person who took the action, and the action itself. Records this in a database.
[0871] Input: Detailed information about the response (date, time, responder, content)
[0872] Output: Corresponding record
[0873] Example: A security guard inspects the scene, enters the details into a terminal, and the information is recorded on the server.
[0874] The above steps create a system that can detect suspicious objects with high accuracy in real time and respond quickly.
[0875] (Application example 1)
[0876] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0877] In security systems that use surveillance cameras, the detection of suspicious objects and lost items often relies on manual monitoring, which can easily lead to oversights. Existing systems also lack the ability to notify relevant parties of detection results in real time, resulting in information delays and making it difficult to respond quickly. Furthermore, the system lacks the ability to identify suspicious individuals and related activities using past video data, creating a need for enhanced security.
[0878] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0879] In this invention, the server
[0880] A means for acquiring video data from a surveillance camera in real time;
[0881] A means for pre-processing the acquired video data and inputting it into a generative AI model for object detection;
[0882] A means for detecting objects using a generative AI model and identifying suspicious or lost items;
[0883] A means for generating alerts and notifying relevant parties about detected suspicious or lost items via email or a dedicated application;
[0884] a means for displaying detection results and alerts to relevant personnel in real time using a head-mounted display;
[0885] This enables highly accurate detection and rapid response in real time.
[0886] A "surveillance camera" is a device that acquires video data in real time and monitors a surveillance area.
[0887] "Video data" refers to a video stream captured by a surveillance camera, and is a set of image information that is subject to analysis and monitoring.
[0888] "Preprocessing" refers to the process of converting video data into a format suitable for the generative AI model, and specifically includes resizing, normalization, noise removal, etc.
[0889] A "generative AI model" is a technology that uses deep learning and machine learning algorithms to generate models, learn patterns from input data, and extract features.
[0890] A "target" is an object that exists in a surveillance camera image and that needs to be identified or detected.
[0891] A "suspicious object" is an object within a surveillance area that is deemed abnormal or dangerous.
[0892] "Lost property" is an object that has been left behind with no known owner.
[0893] An "alert" is a notification that warns or calls attention to detected suspicious objects or lost items.
[0894] "Email" is a means of communication for sending and receiving messages over the Internet.
[0895] A "dedicated application" is software developed for a specific purpose, and is a system designed to provide notifications and information.
[0896] A "head-mounted display" is a wearable display device that displays information within the user's field of vision when worn by the user.
[0897] "Interested parties" are those who should be notified by the system, and are typically security personnel such as security guards or administrators.
[0898] The system for implementing this invention consists of a surveillance camera, a server, a head-mounted display terminal, and a user (mainly a security guard). The server acquires video data from the surveillance camera in real time and analyzes the video data using a generative AI model. The head-mounted display (HMD) terminal worn by the security guard receives notifications from the server and displays the notification contents to the security guard in real time.
[0899] Hardware and software used
[0900] Hardware: surveillance cameras, servers, head-mounted displays (HMDs)
[0901] Software: TensorFlow (generative AI model), OpenCV (image processing)
[0902] Details of data processing and calculation
[0903] The server receives video data from the surveillance cameras in real time. This data is first preprocessed using OpenCV. Specifically, the preprocessing involves resizing and normalizing the video frames.
[0904] The pre-processed video data is then fed into a generative AI model for object detection. The generative AI model, built using TensorFlow, detects objects in the video and calculates a label and confidence score.
[0905] If any of the detected objects are deemed suspicious or lost, the server compares them with past video data to identify the object's occupant or suspicious person.
[0906] The server generates an alert based on the results and automatically notifies the relevant parties (security guards) via email or a dedicated application. The head-mounted display terminal worn by the guard also displays the detection results and alert details in real time.
[0907] Specific examples
[0908] For example, if a suspicious object (an unattended bag) is left unattended for a certain period of time in a food court in a shopping mall, the system will operate as follows:
[0909] Surveillance cameras capture footage of the food court.
[0910] The server acquires the video data in real time, preprocesses it, and inputs it into the generative AI model.
[0911] The generative AI model detects unattended bags, analyzes the information and compares it with historical data.
[0912] The server checks for the presence of suspicious objects and generates an alert.
[0913] The security guard's head-mounted display displays a real-time message saying, "There is an unattended bag in the food court on the first floor. Please check."
[0914] Prompt Sentence Examples
[0915] For example, use the following prompt:
[0916] "Monitoring the food court for 45 minutes. Detecting an unattended bag and confirming that it had been left there for 20 minutes. Checking against past data to identify the owner. Notifying a security officer wearing an HMD that 'There is an unattended bag in the food court on the first floor. Please check.'"
[0917] This configuration enables highly accurate detection of suspicious objects and lost items within the surveillance area, enabling rapid response in real time.
[0918] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0919] Step 1:
[0920] The server acquires video data from the surveillance camera in real time. The input is the video stream sent from the surveillance camera, and the output is the raw video data stored in the server. The video data is received in stream format via the URL or IP address of the surveillance camera.
[0921] Step 2:
[0922] The server performs preprocessing on the acquired video data before inputting it into the generative AI model. The input is the raw video data acquired in step 1, and the output is preprocessed video frames. Preprocessing includes resizing, normalizing, and denoising the video frames. Specifically, OpenCV is used to resize the frames to the required size and normalize the pixel values.
[0923] Step 3:
[0924] The server inputs the preprocessed video data into a generative AI model to detect objects. The input is the video frames preprocessed in step 2, and the output is the labels and confidence scores of the detected objects. A generative AI model using TensorFlow is used to identify objects in the video frames. Each object is assigned a label and its confidence score.
[0925] Step 4:
[0926] The server identifies suspicious or lost objects based on the confidence score of the detected object. The input is the object label and confidence score generated in step 3, and the output is the result of the suspicious or lost object judgment. Criteria for suspicious or lost objects are set in advance, and the judgment is made according to those criteria, and a flag is raised if action is required.
[0927] Step 5:
[0928] The server compares the detected object with past video data to identify the object's occupant or suspicious person. The input is the judgment result from step 4 and past video data, and the output is the identification of the owner or suspicious person of the suspicious object or lost item. This improves the accuracy of the judgment by comparing with past behavioral history.
[0929] Step 6:
[0930] If the server determines that an item is suspicious or lost, it automatically generates an alert and notifies the relevant parties. The input is the determination result and identification result, and the output is the generated alert and notification message. The alert is sent via email or a dedicated application.
[0931] Step 7:
[0932] The terminal (head-mounted display) receives an alert notification from the server and displays the notification content to the security guard. The input is the alert notification sent from the server, and the output is the warning message displayed on the head-mounted display. The notification content is specific, such as "There is an unattended bag in the food court on the first floor. Please check it."
[0933] Step 8:
[0934] The user (security guard) checks the notification displayed on the head-mounted display, heads to the scene, and takes action. The input is the alert content displayed on the head-mounted display, and the output is the execution and reporting of the response. The report content is entered into the terminal and recorded on the server.
[0935] These steps enable highly accurate detection of suspicious or lost items in real time and rapid response.
[0936] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0937] The system embodying this invention consists of a surveillance camera, a server, a terminal, a user, and an emotion engine. The server acquires video data from the surveillance camera in real time and analyzes the video data using a generative AI model. The emotion engine recognizes the user's emotions and incorporates that information into the analysis. The terminal receives notifications from the server, and the user responds based on those notifications.
[0938] Program processing explanation
[0939] 1. Obtaining surveillance camera footage
[0940] Server: Acquires video data from surveillance cameras in real time. Receives video streams via the surveillance camera's URL or IP address and acquires them as video data.
[0941] 2. Preprocessing of video data
[0942] Server: Converts acquired video data into a format that is easy for the generative AI model to process. Preprocessing includes resizing and normalizing video frames.
[0943] 3. Object detection using AI models
[0944] Server: The generative AI model detects objects in pre-processed video frames, and calculates a label for each object (e.g., bag, person, etc.) and its confidence score.
[0945] 4. Analysis of detection results
[0946] Server: Identifies suspicious or lost objects based on the confidence score of the detected object. Criteria for suspicious or lost objects are set in advance, and judgments are made based on those criteria.
[0947] 5. Comparison with past data
[0948] Server: Detected objects are compared with past video data to identify the object's occupant or suspicious individuals. This improves the accuracy of determining whether an object's status is suspicious.
[0949] 6. Emotion Recognition by Emotion Engine
[0950] Server: Utilizes an emotion engine that uses video and audio data to recognize the user's emotions, thereby obtaining emotional information such as whether the person involved is feeling stressed.
[0951] 7. Alert Generation and Coordination
[0952] Server: If an item is determined to be suspicious or lost, an alert is automatically generated. The content of the alert and notification method are adjusted based on the recognition results of the emotion engine. For example, if the person involved is under high stress, the content of the notification may be softened.
[0953] Terminal: A notification is sent to the relevant person's terminal and the alert details are displayed. The user checks this notification and takes appropriate action.
[0954] 8. Record of correspondence
[0955] Server: Records the actions taken in response to generated alerts, which are used for later analysis and system improvement.
[0956] Specific examples
[0957] Here is a specific example from inside a shopping mall.
[0958] Surveillance camera: Monitoring the food court in the shopping mall.
[0959] Server: Captures surveillance camera footage in real time and analyzes it using a generative AI model. The model detects unattended bags as suspicious objects. It checks that unattended bags in the footage have been left unattended for a certain period of time and compares them with past data to identify the owner. It also uses an emotion engine to evaluate the stress level of security guards. Based on the detection results, it generates tailored alerts and notifies security guards.
[0960] Device: The security guard's smartphone receives an alert and displays a notification saying, "There is an unattended bag in the food court on the first floor. Please check." If the security guard is in a high-stress state, the notification content is changed to, "There is an unattended bag in the food court on the first floor. Please check."
[0961] User (security guard): The security guard checks the notification, heads to the scene, checks the safety of the bag, and then enters the details of the action into the terminal.
[0962] Server: Records the responses from security guards and uses them for later analysis.
[0963] This system will improve security within shopping malls by enabling early detection and rapid response of suspicious or lost items in real time. In addition, by combining it with an emotion engine, it will be possible to respond taking into account the emotional state of those involved.
[0964] The processing flow will be explained below.
[0965] Step 1:
[0966] The server acquires video data from the surveillance camera in real time, receives the video stream via the surveillance camera's URL or IP address, and acquires the video data frame by frame.
[0967] Step 2:
[0968] The server converts the acquired video data into a format that is easy for the generative AI model to process, specifically by resizing the video frames to 224x224 pixels and normalizing them to match the model's input format.
[0969] Step 3:
[0970] The server inputs the preprocessed video data into a generative AI model for object detection, which identifies each object in the frame and outputs a label (e.g., bag, person, etc.) and its confidence score.
[0971] Step 4:
[0972] The server analyzes the output of the generative AI model and evaluates whether the confidence score exceeds a certain threshold. Detections that exceed the threshold are listed as suspicious or lost items.
[0973] Step 5:
[0974] The server compares the listed suspicious or lost items with past video data, and based on the past data, identifies the occupant of the object or the suspicious person, and confirms that the object is a suspicious or lost item.
[0975] Step 6:
[0976] The server uses the video and audio data to recognize the user's emotions through an emotion engine, analyzing the user's facial expressions and tone of voice to assess their emotional state, such as their stress level.
[0977] Step 7:
[0978] The server automatically generates an alert if the item is determined to be a suspicious or lost item. The content of the alert and notification method are adjusted based on the results of the emotion engine. For example, if the user is in a state of high stress, the notification content will be changed to a softer expression.
[0979] Step 8:
[0980] The server notifies the relevant parties of the generated alerts via email, SMS, or a dedicated application sent to the device.
[0981] Step 9:
[0982] The terminal receives the notification from the server and displays the alert content. The user can check this notification and take the necessary action.
[0983] Step 10:
[0984] The user (security guard) checks the alert displayed on the terminal, goes to the scene to check the object, and then inputs the response details into the terminal.
[0985] Step 11:
[0986] The server records the responses entered by the user (security guard), which will be used for future analysis and system improvement.
[0987] As a concrete example, consider the case where an unattended bag is left in the food court of a shopping mall. The server detects the unattended bag from surveillance camera footage. It compares this with past data to identify the bag's owner. It uses an emotion engine to check the stress level of the security guard and generates an alert in appropriate language to notify the guard. The security guard then checks the notification, checks the safety of the bag on-site, and enters the results into the server. This series of processes enables effective and human-friendly security operations.
[0988] Example 2
[0989] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0990] Conventional surveillance systems detect suspicious objects or lost items, but do not take into account the emotional state of those involved, which can lead to stress when dealing with the incident. Furthermore, detection accuracy can be low because the system does not compare the detected items with past video data. Furthermore, the recording and subsequent analysis of detected information is insufficient, which can hinder system improvement.
[0991] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for acquiring video data from a surveillance camera in real time, means for preprocessing the acquired video data and inputting it into a generative AI model for detecting objects, means for detecting objects using the generative AI model and identifying suspicious or lost objects, means for evaluating the emotional state of relevant parties using an emotion engine and adjusting the content of an alert when a suspicious or lost object is detected, and means for generating and notifying relevant parties of an alert. This enables highly accurate detection of suspicious or lost objects and appropriate responses that take into account the emotional state of relevant parties. In addition, detection accuracy is improved by comparing with past video data, and continuous system improvement is possible through recording and subsequent analysis of detection information.
[0992] A "surveillance camera" is a device that captures images of a pre-set location in real time.
[0993] A "server" is a computer system that obtains, processes, and manages data from other devices over a network.
[0994] "Real-time" refers to processing immediately without delay, and refers to a state in which appropriate action can be taken the moment a specific event occurs.
[0995] "Video data" means a collection of visual information captured from a surveillance camera or other device and stored as video frames.
[0996] "Preprocessing" refers to a series of operations that transform data into a suitable format before it is sent to the main processing stage, including, for example, resizing and normalization.
[0997] A "generative AI model" is an artificial intelligence algorithm that is pre-trained and used to perform a specific task (in this case, object detection).
[0998] "Subject" refers to an entity with specific characteristics, such as a person or object, that exists in the video.
[0999] A "suspicious object" is an object that is out of place or has been left unattended for a period of time and may pose a threat to normal activity.
[1000] "Lost property" refers to property that has been unintentionally abandoned or forgotten.
[1001] An "emotion engine" is software or algorithms that analyze the emotional state (e.g., stress, anxiety) of participants based on video and audio data.
[1002] An "alert" refers to a warning or notification message in response to a specific event (e.g., the discovery of a suspicious object).
[1003] "Notification" means the act of sending a message to inform interested parties of an alert or other important information.
[1004] The system for implementing this invention includes a surveillance camera, a server, a terminal, a user, and an emotion engine. Each component functions as follows.
[1005] surveillance cameras
[1006] Surveillance cameras capture video in real time from pre-defined locations. For example, if a surveillance camera is used in a food court in a shopping mall, it will monitor each area of the food court and capture video. The video from the camera is sent to a server via a specified URL or IP address.
[1007] server
[1008] The server acquires video data from the surveillance cameras in real time and processes it to detect suspicious objects and lost items. Specifically, it performs the following processes.
[1009] 1. Video data acquisition and preprocessing:
[1010] The server receives the video stream from the surveillance camera using the RTSP protocol, resizes the frame (e.g., from 1920x1080 to 640x480) and normalizes the pixel values using the OpenCV library.
[1011] 2. Object detection using AI models:
[1012] The preprocessed video frames are input into a generative AI model (e.g., YOLO, SSD) to detect objects (e.g., bags, people, etc.) in each frame, which includes a label for each object and its confidence score.
[1013] 3. Analysis of detection results:
[1014] The system analyzes detected object information to identify suspicious or lost items. In particular, if an unattended bag is left unattended for a certain period of time, it will be identified as a suspicious object. It will also compare the object with past video data to identify the occupant or suspicious person.
[1015] 4. Emotion Recognition with Emotion Engine:
[1016] Analyze the emotional state of the person involved (e.g., stress level, anxiety) using video and audio data, using an emotion recognition model (e.g., Microsoft Azure's Emotion API), and adjust the alert content if the person involved is in a high-stress state.
[1017] 5. Alert Generation and Notification:
[1018] If a suspicious or lost item is identified, an alert is automatically generated, tailored based on the results of the emotion engine, and sent to the relevant device (e.g., smartphone).
[1019] Terminal
[1020] The terminal (e.g., a security guard's smartphone) receives the alert sent from the server and displays it on the screen. For example, a notification such as "There is an unattended bag in the food court on the first floor. Please check." The terminal provides an interface that allows the security guard to input the details of the response after checking the scene.
[1021] User
[1022] The user (e.g., a security guard) checks the notification displayed on the terminal and takes action at the scene. For example, if the owner of an unattended bag is confirmed, the user enters the details of the action into the terminal.
[1023] Specific examples
[1024] The server receives real-time footage from surveillance cameras installed in the food court of a shopping mall and uses a generative AI model to detect unattended bags. If a bag is detected left in a specific area for a certain period of time, the server uses an emotion engine to evaluate the emotional state of the security guard and adjust the alert content. A notification is sent to the security guard's smartphone saying, "There is an unattended bag in the food court on the first floor. Please check." The security guard then heads to the scene, checks the safety of the bag, and enters the details into the terminal.
[1025] Example prompt sentence:
[1026] "Detect unattended bags from surveillance camera footage in a food court and calculate the time they have been left there. Also, recognize the emotional state (stress level) of people in the area."
[1027] This system enables highly accurate detection of suspicious objects and lost items in real time and rapid response, improving environmental security. In addition, by combining it with an emotion engine, flexible responses can be made taking into account the emotional state of those involved.
[1028] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1029] Step 1:
[1030] Acquiring video data
[1031] Server: Specify the URL or IP address of the surveillance camera and receive the video stream in real time using the RTSP protocol. The input is the video data from the surveillance camera, and the output is the captured real-time video frames. The server captures the video at a rate of 30 frames per second and stores it in memory.
[1032] Step 2:
[1033] Video data preprocessing
[1034] Server: Converts captured video frames into a format that is easy for the generative AI model to process. Specifically, it uses the OpenCV library to resize the frames (e.g., from 1920x1080 to 640x480) and normalizes pixel values to the range of 0-255. The input is a video frame captured in real time, and the output is a preprocessed video frame.
[1035] Step 3:
[1036] Object detection with AI models
[1037] Server: Inputs preprocessed video frames into a generative AI model (e.g., YOLO, SSD) to detect objects in each frame. The input is the preprocessed video frames, and the output is each object's label (e.g., bag, person), its coordinates, and a confidence score. The server updates the list of detected objects for each frame.
[1038] Step 4:
[1039] Analysis of detection results
[1040] Server: Analyzes the detected object information and identifies suspicious or lost objects. For example, if an unattended bag is left unattended for a certain period of time, it will be recorded as a suspicious object. The input is a list of detected objects and a confidence score, and the output is a list of suspicious or lost objects. The server analyzes the object status based on specific criteria and stores the results in a database.
[1041] Step 5:
[1042] Comparison with past data
[1043] Server: Compares newly detected suspicious or lost objects with past video data. Specifically, it compares the object with past records to identify the occupant or suspicious person. The input is a list of suspicious or lost objects and past video data, and the output is the identification information of the occupant or suspicious person. The server compares the object with past video frames stored in a database and analyzes the object's history.
[1044] Step 6:
[1045] Emotion recognition by emotion engine
[1046] Server: An emotion engine uses video and audio data to recognize the emotional state of the participants. This uses an emotion recognition model (e.g., Microsoft Azure's Emotion API). The input is video and audio data captured in real time, and the output is the participants' emotional information (e.g., stress level, anxiety). The server analyzes the participants' facial expressions and tone of voice to evaluate their emotional state.
[1047] Step 7:
[1048] Alerting and Tuning
[1049] Server: Detects suspicious or lost items and adjusts the alert content based on the results of the emotion engine. The input is a list of suspicious or lost items and emotional information, and the output is an alert message sent to the relevant person. For example, a person in a high stress state might be notified in a gentler way, such as "There is a bag in the food court on the first floor. Please check it." The server adjusts the alert content and sends it to the relevant person's device.
[1050] Step 8:
[1051] Record of correspondence
[1052] Server: Records the responses of relevant parties to generated alerts. The input is the response details from the relevant parties, and the output is the recorded response data. The server stores the response results entered by the relevant parties into a database (e.g., confirming the owner of the bag, confirming there is no threat) for later analysis and system improvement.
[1053] (Application example 2)
[1054] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1055] Conventional surveillance systems have limitations in accuracy and real-time detection of suspicious objects and lost items, making it difficult to respond quickly. Furthermore, because the emotional state of the user (such as a security guard) is not taken into consideration, the content of notifications and response methods cannot be optimized, resulting in stress for those involved and efficiency issues. The objective of this invention is to enable accurate real-time detection of suspicious objects and lost items while also enabling appropriate notifications and responses according to the user's emotional state.
[1056] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1057] In this invention, the server includes means for acquiring video data from a monitoring device in real time, means for preprocessing the acquired video data and inputting it into a generative AI model for detecting objects, means for detecting objects using the generative AI model and identifying suspicious objects or lost items, means for recognizing a user's emotional state using an emotion engine and reflecting that information in the analysis, and means for generating an alert and notifying the user according to the user's emotional state when a suspicious object or lost item is detected. This enables highly accurate detection of suspicious objects and lost items in real time, and also enables detailed notifications and responses according to the user's emotional state.
[1058] A "monitoring device" is a device that acquires video data in real time and transmits it to a server.
[1059] "Means for acquiring" refers to a method or device for collecting video data from a surveillance device.
[1060] "Preprocessing" is the process of converting the acquired video data into a format that is easy for the generative AI model to process.
[1061] A "generative AI model" is a program that uses machine learning and deep learning techniques to detect and analyze objects in video data.
[1062] The term "object" refers to an object that is the target of detection within the video data.
[1063] The "emotion engine" is a program that analyzes video and audio data to recognize the user's emotional state.
[1064] "Reflecting in the analysis" means incorporating the emotional information recognized by the emotion engine into the analysis results of the generative AI model.
[1065] "Alert" means a warning or notification generated when a suspicious or lost item is detected.
[1066] "Means for notifying" refers to a method or device for communicating a generated alert to a user.
[1067] The system for implementing this invention is composed of a monitoring device, a server, a terminal, a user, and an emotion engine. Each element will be described in detail below.
[1068] The server acquires video data from the surveillance equipment in real time. The acquired video data undergoes preprocessing before being input into the generative AI model. This preprocessing includes data processing such as resizing and normalizing the video frames. The preprocessed video data is then input into the generative AI model for object detection. The generative AI model achieves highly accurate object detection using traditional machine learning and deep learning techniques. Detected objects are analyzed to identify suspicious objects or lost items. By comparing the data with past video data, it is also possible to identify the occupant of the object or a suspicious person.
[1069] The emotion engine recognizes the user's emotions from video and audio data. The emotion information recognized by the emotion engine is reflected in the analysis on the server. This emotion information is used when generating alerts, and the alert content and notification method are adjusted according to the user's emotional state.
[1070] The terminal is a device that is primarily used by users (such as security guards). Notifications from the server are sent to the terminal, and an alert is displayed to the user. The content of the alert is adjusted based on emotional information recognized by the emotion engine. For example, if a security guard is in a state of high stress, the content of the notification can be softened to reduce the stress.
[1071] The user is the person who primarily operates the terminal and takes appropriate action based on notifications from the system. The user checks the notification and takes action on-site. After that, they provide feedback to the server by entering the details of their response into the terminal. This feedback information can be used for subsequent analysis and system improvement.
[1072] As a concrete example, consider the case of monitoring a food court in a shopping mall. Surveillance cameras send video data from inside the food court to a server in real time. The server preprocesses the video data and uses a generative AI model to detect unattended bags as suspicious objects. It confirms that the detected bag has been left unattended for a certain period of time and compares it with past video data to identify the owner. It also uses an emotion engine to evaluate the stress level of the security guard. Based on the detection results, an adjusted alert is generated and notified to the security guard. The security guard's terminal displays the message, "There is a bag in the food court on the first floor. Please check."
[1073] An example of a prompt sentence can be expressed as follows:
[1074] "Design a system that uses a generative AI model to detect suspicious objects and lost items from video data from surveillance equipment, and recognizes the user's emotional state using an emotion engine. Based on the recognition results, notify the user and encourage them to take the appropriate action."
[1075] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1076] Step 1:
[1077] The server acquires video data from the monitoring device in real time. To acquire video data, the server receives the video stream via the IP address or URL of the monitoring device. The video data is then passed to the server as input. The acquired video data is then saved in the server and formatted for subsequent processing.
[1078] Step 2:
[1079] The server preprocesses the acquired video data, specifically by resizing and normalizing the video frames. The input to this step is the raw video frames, and the output is preprocessed video frames in a format that is easy for the generative AI model to analyze.
[1080] Step 3:
[1081] The server inputs the preprocessed video data into a generative AI model to detect objects. Specifically, the generative AI model detects objects in the video frames and calculates their labels (e.g., bag, person, etc.) and confidence scores. The input for this step is the preprocessed video frames, and the output is a list of detected objects.
[1082] Step 4:
[1083] The server analyzes the confidence scores of the objects detected by the generative AI model and identifies them as suspicious or lost. Specifically, the server compares each object's label and confidence score with the set criteria to determine whether it is suspicious or lost. The input for this step is a list of detected objects, and the output is the identification of suspicious or lost objects.
[1084] Step 5:
[1085] The server compares the detected object with past video data and identifies the occupant or suspicious person of the object. Specifically, it searches for related video data from a past database and confirms the object's presence time and past owners. The input to this step is the identification result of the suspicious object or lost item, and the output is detailed information about the object (the identification result of the occupant or suspicious person).
[1086] Step 6:
[1087] The server uses an emotion engine to recognize the user's emotional state from video and audio data. Specifically, it inputs video frames and audio clips into the emotion engine and outputs an emotion label (e.g., high stress, low stress). The input for this step is the current video frame or audio data, and the output is the user's emotional state.
[1088] Step 7:
[1089] When a suspicious object or lost property is detected, the server generates an alert based on the recognition results of the emotion engine and notifies the user. Specifically, it creates an alert message (e.g., "Caution: A highly suspicious object has been detected. Please act with caution") according to the security guard's emotional state and sends it to the terminal. The input of this step is the suspicious object identification result and the emotional state, and the output is a notification alert.
[1090] Step 8:
[1091] The terminal displays the alert sent from the server and notifies the user. Specifically, the notification alert pops up on the screen, prompting the user to confirm it. The user then takes appropriate action based on the notification and inputs the response results into the terminal. The input for this step is the notification alert, and the output is the response results.
[1092] Step 9:
[1093] The server records the response results received from the terminal and uses them for later analysis and system improvement. Specifically, it saves the recorded response results in a database and uses them later for system evaluation and improvement. The input of this step is the response results, and the output is the recorded response data.
[1094] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1095] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1096] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1097] [Fourth embodiment]
[1098] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1099] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1100] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1101] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1102] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1103] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1104] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1105] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1106] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1107] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1108] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1109] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1110] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1111] The system for implementing this invention consists of a surveillance camera, a server, a terminal, and a user. The server acquires video data from the surveillance camera in real time and analyzes the video data using a generative AI model. The terminal receives notifications from the server, and the user takes action based on the notifications.
[1112] Program processing explanation
[1113] 1. Obtaining surveillance camera footage
[1114] Server: Acquires video data from surveillance cameras in real time. Specifically, it receives video data in stream format via the surveillance camera's URL or IP address.
[1115] 2. Preprocessing of video data
[1116] Server: Preprocesses the acquired video data to input it into the generative AI model. Preprocessing includes resizing and normalizing the video frames.
[1117] 3. Object detection using AI models
[1118] Server: The generative AI model detects objects in preprocessed video frames, and calculates a label and its confidence score for each object.
[1119] 4. Analysis of detection results
[1120] Server: Identifies suspicious or lost objects based on the confidence score of the detected object. Criteria for suspicious or lost objects are set in advance, and judgments are made based on those criteria.
[1121] 5. Comparison with past data
[1122] Server: Detected objects are compared with past video data to identify the object's occupant or suspicious individuals. This improves the accuracy of determining whether an object's status is suspicious.
[1123] 6. Alert Generation and Notification
[1124] Server: If an item is determined to be suspicious or lost, it automatically generates an alert and notifies the relevant parties via email, SMS, or a dedicated application.
[1125] Terminal: The relevant person's terminal receives the notification and displays the alert content. The user checks the notification and takes appropriate action.
[1126] 7. Record of correspondence
[1127] Server: Records the actions taken in response to generated alerts, which are used for later analysis and system improvement.
[1128] Specific examples
[1129] A specific example from a shopping mall is shown below.
[1130] Surveillance camera: Monitoring the food court in the shopping mall.
[1131] Server: Captures surveillance camera footage in real time and analyzes it using a generative AI model. The model detects unattended bags as suspicious objects. It checks that unattended bags in the footage have been left unattended for a certain period of time. It compares this with past data to identify the culprit or owner. Based on these results, it generates an alert and notifies security guards.
[1132] Terminal: The security guard's smartphone receives an alert and displays a notification saying, "There is an unattended bag in the food court on the first floor. Please check."
[1133] User (security guard): The security guard checks the notification, goes to the scene, checks the safety of the bag, and enters the details of the action into the terminal and records it on the server.
[1134] This will enable early detection of suspicious objects and lost items in real time and prompt response, improving security within shopping malls.
[1135] The processing flow will be explained below.
[1136] Step 1:
[1137] The server acquires video data from the surveillance camera in real time, receiving the video stream via the surveillance camera's URL or IP address and capturing it as video data.
[1138] Step 2:
[1139] The server converts the acquired video data into a format that is easy for the generative AI model to process, specifically by resizing the video frames to 224x224 pixels and normalizing them to match the model's input format.
[1140] Step 3:
[1141] The server inputs the preprocessed video data into a generative AI model to detect objects, which then calculates object labels (e.g., bag, person, etc.) and their confidence scores.
[1142] Step 4:
[1143] The server receives the detection results from the generative AI model and evaluates whether the confidence score exceeds a certain threshold. If the detection result exceeds the threshold, it is listed as a candidate for suspicious or lost items.
[1144] Step 5:
[1145] The server compares the listed suspicious or lost items with past video data. By using past data, it can identify the occupant of the object or suspicious individuals, improving the accuracy of the judgment.
[1146] Step 6:
[1147] The server automatically generates an alert if an object is detected as suspicious or lost, including the object's label, location, and time stamp.
[1148] Step 7:
[1149] The server notifies the relevant parties of the generated alerts via email, SMS or a dedicated application.
[1150] Step 8:
[1151] The device receives the notification sent from the server and displays the alert content. The relevant parties can check this notification and take appropriate action.
[1152] Step 9:
[1153] The user (security guard) checks the alert displayed on the terminal, heads to the scene to check the safety of the object, and then inputs the details of the response into the terminal.
[1154] Step 10:
[1155] The server records the responses from users (security guards), which are used for future analysis and system improvement.
[1156] Example 1
[1157] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1158] Although there are systems that analyze surveillance footage in real time and quickly identify and notify suspicious or lost items, it is difficult to respond accurately and efficiently.In addition, there is a lack of cross-checking with past video data and management of response records, so there is a need for stronger crime prevention measures.
[1159] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1160] In this invention, the server includes means for acquiring video data from a monitoring device in real time, means for preprocessing the acquired video data and inputting it into an artificial intelligence model for detecting objects, means for detecting objects using the artificial intelligence model and identifying suspicious objects or lost items, means for generating and notifying relevant parties when a suspicious object or lost item is detected, means for identifying the occupant of the object or a suspicious person using past video data, and means for recording detected information based on the detection of a suspicious object or lost item and using it for later analysis and improvement, thereby enabling highly accurate detection in real time and rapid response.
[1161] "Monitoring device" refers to a device for acquiring video data in real time.
[1162] "Video data" refers to real-time video information obtained from surveillance equipment.
[1163] "Preprocessing" refers to a series of steps that are performed on the captured video data before it can be input into an AI model. Specifically, this includes processes such as resizing and normalizing the video frames.
[1164] "Artificial intelligence model" refers to the algorithms and learning models used to analyze acquired video data and detect objects.
[1165] "Objects" refer to objects that the AI model detects in the video data, including suspicious objects and lost items.
[1166] A "suspicious object" is an object that is determined to be abnormal or suspicious based on certain criteria.
[1167] "Lost property" refers to an object that has been determined to be unclaimed or abandoned.
[1168] An "alert" refers to warning information used to notify relevant parties when a suspicious object or lost item is detected.
[1169] "Interested parties" refers to people or organizations that should receive the alert, including, for example, security guards and administrators.
[1170] "Past video data" refers to video information recorded in the past among the acquired video data.
[1171] "Matching" refers to the process of comparing currently detected objects with past video data to identify the occupant or suspicious person of a matching object.
[1172] "Response record" refers to data used to store the history of the detection of suspicious or lost items and the response taken thereto.
[1173] The system for implementing this invention consists of a monitoring device, a server, a terminal, and a user. The server acquires video data from the monitoring device in real time, preprocesses the data, and inputs it into a generative AI model for analysis. The terminal receives notifications from the server, and the user responds based on the notifications.
[1174] The server is equipped with hardware and software for acquiring, preprocessing, and analyzing video data. Specifically, libraries such as OpenCV and FFmpeg are used to perform preprocessing such as resizing and normalizing video frames. Generative AI models such as YOLO (You Only Look Once) and SSD (Single Shot Multibox Detector) are used to detect objects in the preprocessed video data.
[1175] As an example, consider surveillance in a shopping mall. Video data is acquired from a surveillance device installed in the mall's food court and received by a server. The server resizes and normalizes the acquired video data (for example, to 640x480 pixels) and inputs it into the YOLO v4 model. The model detects suspicious objects, such as unattended bags, with high accuracy. An object is deemed suspicious if the confidence score is 0.85 or higher.
[1176] The server then compares the detected suspicious object with video data from the past 24 hours to identify the occupant or suspicious person. This reduces false positives and enables more accurate identification of suspicious objects. It automatically generates an alert and sends a notification to the device of the relevant person (e.g., security guard). This notification can be sent via email, SMS, or a dedicated application.
[1177] A specific example of the notification content is a message such as "There is an unattended bag in the food court on the first floor. Please check." Security personnel receive this notification and rush to the scene to respond. The response at that time is recorded on the server via the terminal. The recorded information is used for future analysis and system improvements.
[1178] Prompt Sentence Examples
[1179] Example of using security systems in shopping malls:
[1180] The surveillance cameras capture real-time footage of the food court and send it to the server. The server uses a generative AI model (YOLO v4) to detect unattended bags and compares the detection results with video data from the past 24 hours to identify suspicious individuals. If a suspicious object is determined based on the specified criteria, the security guard's smartphone should be notified, "There is an unattended bag in the 1st floor food court. Please take a look." The server should record the response results.
[1181] This invention enables highly accurate real-time detection of suspicious objects and rapid response, improving security in shopping malls and other public places.
[1182] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1183] Step 1: Obtaining security camera footage
[1184] Specific behavior:
[1185] Server: Obtain video data in real time using the IP address or URL of the surveillance device. For example, use OpenCV to capture the camera stream with cv2.VideoCapture.
[1186] Input: Video data from a surveillance device (stream format)
[1187] Output: Captured video frames
[1188] Example: If the URL of the monitoring device is rtsp: / / 192.168.1.10 / stream, use this URL to get the video.
[1189] Step 2: Preprocessing the video data
[1190] Specific behavior:
[1191] Server: Resize and normalize the captured video frames. To resize, use OpenCV's cv2.resize function, and to normalize, scale pixel values to the range 0 to 1.
[1192] Input: Captured video frame
[1193] Output: Preprocessed video frame (resized image data)
[1194] Example: Resize and normalize the frame to 640x480 pixels.
[1195] Step 3: Object detection with an AI model
[1196] Specific behavior:
[1197] Server: Inputs preprocessed video frames into a generative AI model (such as YOLO) to perform object detection, and outputs object labels and confidence scores.
[1198] Input: Preprocessed video frames
[1199] Output: Detected objects (including labels and confidence scores)
[1200] Example: Detect unattended bags using YOLO v4 model. If the confidence score is 0.85 or higher, it is considered suspicious.
[1201] Step 4: Analyzing the detection results
[1202] Specific behavior:
[1203] Server: Identifies suspicious or lost objects based on the confidence score and label of the detected object. If the confidence score exceeds a certain threshold, the object is deemed suspicious.
[1204] Input: Detected objects (labels and confidence scores)
[1205] Output: Suspicious or lost item determination result
[1206] Example: If the "bag" label is detected with a confidence score of 0.85 or higher, identify it as suspicious.
[1207] Step 5: Check against historical data
[1208] Specific behavior:
[1209] Server: For detected suspicious objects, retrieves video data from the past 24 hours from a database and compares it to identify occupants or suspicious individuals. Past frames are input into the AI model again to confirm the object.
[1210] Input: Information on detected suspicious objects, past video data
[1211] Output: Information on identified occupants or suspicious individuals
[1212] Example: Matching data from the past 24 hours to determine who the bag belongs to or was involved in it.
[1213] Step 6: Alerting and Notification
[1214] Specific behavior:
[1215] Server: If an item is determined to be suspicious or lost, an alert is automatically generated and relevant parties are notified via email, SMS, or a dedicated application.
[1216] Terminal: The terminal of the person involved (security guard) receives the notification and displays a pop-up notification.
[1217] Input: Suspicious object judgment result
[1218] Output: Alert notification
[1219] Example: A security guard's smartphone displays a notification such as, "There is an unattended bag in the food court on the first floor. Please check it."
[1220] Step 7: Record your response
[1221] Specific behavior:
[1222] Server: Records the actions taken in response to generated alerts. This includes the date and time of the action, the person who took the action, and the action content. Records this in a database.
[1223] Input: Detailed information about the response (date, time, responder, content)
[1224] Output: Corresponding record
[1225] Example: A security guard inspects the scene, enters the details into a terminal, and the information is recorded on the server.
[1226] The above steps create a system that can detect suspicious objects with high accuracy in real time and respond quickly.
[1227] (Application example 1)
[1228] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1229] In security systems that use surveillance cameras, the detection of suspicious objects and lost items often relies on manual monitoring, which can easily lead to oversights. Existing systems also lack the ability to notify relevant parties of detection results in real time, resulting in information delays and making it difficult to respond quickly. Furthermore, the ability to identify suspicious individuals and related activities using past video data is also insufficient, creating a need for enhanced security.
[1230] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1231] In this invention, the server
[1232] A means for acquiring video data from a surveillance camera in real time;
[1233] A means for pre-processing the acquired video data and inputting it into a generative AI model for object detection;
[1234] A means for detecting objects using a generative AI model and identifying suspicious or lost items;
[1235] A means for generating alerts to relevant parties about detected suspicious or lost items and notifying them via email or a dedicated application;
[1236] a means for displaying detection results and alerts to relevant personnel in real time using a head-mounted display;
[1237] This enables highly accurate detection and rapid response in real time.
[1238] A "surveillance camera" is a device that acquires video data in real time and monitors a surveillance area.
[1239] "Video data" refers to a video stream captured by a surveillance camera, and is a set of image information that is subject to analysis and monitoring.
[1240] "Preprocessing" refers to the process of converting video data into a format suitable for the generative AI model, and specifically includes resizing, normalization, and noise removal.
[1241] A "generative AI model" is a technology that uses deep learning and machine learning algorithms to generate models, learn patterns from input data, and extract features.
[1242] A "target" is an object that exists in a surveillance camera image and that needs to be identified or detected.
[1243] A "suspicious object" is an object within a surveillance area that is deemed abnormal or dangerous.
[1244] "Lost property" is an object that has been left behind with no known owner.
[1245] An "alert" is a notification that warns or calls attention to detected suspicious objects or lost items.
[1246] "Email" is a means of communication for sending and receiving messages over the Internet.
[1247] A "dedicated application" is software developed for a specific purpose, and is a system designed to provide notifications and information.
[1248] A "head-mounted display" is a wearable display device that displays information within the user's field of vision when worn by the user.
[1249] "Interested parties" are those who should be notified by the system, and are typically security personnel such as security guards or administrators.
[1250] The system for implementing this invention consists of a surveillance camera, a server, a head-mounted display terminal, and a user (mainly a security guard). The server acquires video data from the surveillance camera in real time and analyzes the video data using a generative AI model. The head-mounted display (HMD) terminal worn by the security guard receives notifications from the server and displays the notification contents to the security guard in real time.
[1251] Hardware and software used
[1252] Hardware: surveillance cameras, servers, head-mounted displays (HMDs)
[1253] Software: TensorFlow (generative AI model), OpenCV (image processing)
[1254] Details of data processing and calculation
[1255] The server receives video data from the surveillance cameras in real time. This data is first preprocessed using OpenCV. Specifically, the preprocessing involves resizing and normalizing the video frames.
[1256] The pre-processed video data is then fed into a generative AI model for object detection. The generative AI model, built using TensorFlow, detects objects in the video and calculates a label and confidence score.
[1257] If any of the detected objects are deemed suspicious or lost, the server compares them with past video data to identify the object's occupant or suspicious person.
[1258] The server generates an alert based on the results and automatically notifies the relevant parties (security guards) via email or a dedicated application. The head-mounted display terminal worn by the guard also displays the detection results and alert details in real time.
[1259] Specific examples
[1260] For example, if a suspicious object (an unattended bag) is left unattended for a certain period of time in a food court in a shopping mall, the system will operate as follows:
[1261] Surveillance cameras capture footage of the food court.
[1262] The server acquires the video data in real time, preprocesses it, and inputs it into the generative AI model.
[1263] The generative AI model detects unattended bags, analyzes the information and compares it with historical data.
[1264] The server checks for the presence of suspicious objects and generates an alert.
[1265] The security guard's head-mounted display displays a real-time message saying, "There is an unattended bag in the food court on the first floor. Please check."
[1266] Prompt Sentence Examples
[1267] For example, use the following prompt:
[1268] "Monitoring the food court for 45 minutes. Detecting an unattended bag and confirming that it had been left there for 20 minutes. Checking against past data to identify the owner. Notifying a security officer wearing an HMD that 'There is an unattended bag in the food court on the first floor. Please check.'"
[1269] This configuration enables highly accurate detection of suspicious objects and lost items within the surveillance area, enabling rapid response in real time.
[1270] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1271] Step 1:
[1272] The server acquires video data from the surveillance camera in real time. The input is the video stream sent from the surveillance camera, and the output is the raw video data stored in the server. The video data is received in stream format via the URL or IP address of the surveillance camera.
[1273] Step 2:
[1274] The server performs preprocessing on the acquired video data before inputting it into the generative AI model. The input is the raw video data acquired in step 1, and the output is preprocessed video frames. Preprocessing includes resizing, normalizing, and denoising the video frames. Specifically, OpenCV is used to resize the frames to the required size and normalize the pixel values.
[1275] Step 3:
[1276] The server inputs the preprocessed video data into a generative AI model to detect objects. The input is the video frames preprocessed in step 2, and the output is the labels and confidence scores of the detected objects. A generative AI model using TensorFlow is used to identify objects in the video frames. Each object is assigned a label and its confidence score.
[1277] Step 4:
[1278] The server identifies suspicious or lost objects based on the confidence score of the detected object. The input is the object label and confidence score generated in step 3, and the output is the result of the suspicious or lost object judgment. Criteria for suspicious or lost objects are set in advance, and the judgment is made according to those criteria, and a flag is raised if action is required.
[1279] Step 5:
[1280] The server compares the detected object with past video data to identify the object's occupant or suspicious person. The input is the judgment result from step 4 and past video data, and the output is the identification of the owner or suspicious person of the suspicious object or lost item. This improves the accuracy of the judgment by comparing with past behavioral history.
[1281] Step 6:
[1282] If the server determines that an item is suspicious or lost, it automatically generates an alert and notifies the relevant parties. The input is the determination result and identification result, and the output is the generated alert and notification message. The alert is sent via email or a dedicated application.
[1283] Step 7:
[1284] The terminal (head-mounted display) receives an alert notification from the server and displays the notification content to the security guard. The input is the alert notification sent from the server, and the output is the warning message displayed on the head-mounted display. The notification content is specific, such as "There is an unattended bag in the food court on the first floor. Please check it."
[1285] Step 8:
[1286] The user (security guard) checks the notification displayed on the head-mounted display, heads to the scene, and takes action. The input is the alert content displayed on the head-mounted display, and the output is the execution and reporting of the response. The report content is entered into the terminal and recorded on the server.
[1287] These steps enable highly accurate detection of suspicious or lost items in real time and rapid response.
[1288] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1289] The system embodying this invention consists of a surveillance camera, a server, a terminal, a user, and an emotion engine. The server acquires video data from the surveillance camera in real time and analyzes the video data using a generative AI model. The emotion engine recognizes the user's emotions and incorporates that information into the analysis. The terminal receives notifications from the server, and the user responds based on those notifications.
[1290] Program processing explanation
[1291] 1. Obtaining surveillance camera footage
[1292] Server: Acquires video data from surveillance cameras in real time. Receives video streams via the surveillance camera's URL or IP address and acquires them as video data.
[1293] 2. Preprocessing of video data
[1294] Server: Converts acquired video data into a format that is easy for the generative AI model to process. Preprocessing includes resizing and normalizing video frames.
[1295] 3. Object detection using AI models
[1296] Server: The generative AI model detects objects in pre-processed video frames, and calculates a label for each object (e.g., bag, person, etc.) and its confidence score.
[1297] 4. Analysis of detection results
[1298] Server: Identifies suspicious or lost objects based on the confidence score of the detected object. Criteria for suspicious or lost objects are set in advance, and judgments are made based on those criteria.
[1299] 5. Comparison with past data
[1300] Server: Detected objects are compared with past video data to identify the object's occupant or suspicious individuals. This improves the accuracy of determining whether an object's status is suspicious.
[1301] 6. Emotion Recognition by Emotion Engine
[1302] Server: Utilizes an emotion engine that uses video and audio data to recognize the user's emotions, thereby obtaining emotional information such as whether the person involved is feeling stressed.
[1303] 7. Alert Generation and Coordination
[1304] Server: If an item is determined to be suspicious or lost, an alert is automatically generated. The content of the alert and notification method are adjusted based on the recognition results of the emotion engine. For example, if the person involved is under high stress, the content of the notification may be softened.
[1305] Terminal: A notification is sent to the relevant person's terminal and the alert details are displayed. The user checks this notification and takes appropriate action.
[1306] 8. Record of correspondence
[1307] Server: Records the actions taken in response to generated alerts, which are used for later analysis and system improvement.
[1308] Specific examples
[1309] Here is a specific example from inside a shopping mall.
[1310] Surveillance camera: Monitoring the food court in the shopping mall.
[1311] Server: Captures surveillance camera footage in real time and analyzes it using a generative AI model. The model detects unattended bags as suspicious objects. It checks that unattended bags in the footage have been left unattended for a certain period of time and compares them with past data to identify the owner. It also uses an emotion engine to evaluate the stress level of security guards. Based on the detection results, it generates tailored alerts and notifies security guards.
[1312] Device: The security guard's smartphone receives an alert and displays a notification saying, "There is an unattended bag in the food court on the first floor. Please check." If the security guard is in a high-stress state, the notification content is changed to, "There is an unattended bag in the food court on the first floor. Please check."
[1313] User (security guard): The security guard checks the notification, heads to the scene, checks the safety of the bag, and then enters the details of the action into the terminal.
[1314] Server: Records the responses from security guards and uses them for later analysis.
[1315] This system will improve security within shopping malls by enabling early detection and rapid response of suspicious or lost items in real time. In addition, by combining it with an emotion engine, it will be possible to respond taking into account the emotional state of those involved.
[1316] The processing flow will be explained below.
[1317] Step 1:
[1318] The server acquires video data from the surveillance camera in real time, receives the video stream via the surveillance camera's URL or IP address, and acquires the video data frame by frame.
[1319] Step 2:
[1320] The server converts the acquired video data into a format that is easy for the generative AI model to process, specifically by resizing the video frames to 224x224 pixels and normalizing them to match the model's input format.
[1321] Step 3:
[1322] The server inputs the preprocessed video data into a generative AI model for object detection, which identifies each object in the frame and outputs a label (e.g., bag, person, etc.) and its confidence score.
[1323] Step 4:
[1324] The server analyzes the output of the generative AI model and evaluates whether the confidence score exceeds a certain threshold. Detections that exceed the threshold are listed as suspicious or lost items.
[1325] Step 5:
[1326] The server compares the listed suspicious or lost items with past video data, and based on the past data, identifies the occupant of the object or the suspicious person, and confirms that the object is a suspicious or lost item.
[1327] Step 6:
[1328] The server uses the video and audio data to recognize the user's emotions through an emotion engine, analyzing the user's facial expressions and tone of voice to assess their emotional state, such as their stress level.
[1329] Step 7:
[1330] The server automatically generates an alert if the item is determined to be a suspicious or lost item. The content of the alert and notification method are adjusted based on the results of the emotion engine. For example, if the user is in a state of high stress, the notification content will be changed to a softer expression.
[1331] Step 8:
[1332] The server notifies the relevant parties of the generated alerts via email, SMS, or a dedicated application sent to the device.
[1333] Step 9:
[1334] The terminal receives the notification from the server and displays the alert content. The user can check this notification and take the necessary action.
[1335] Step 10:
[1336] The user (security guard) checks the alert displayed on the terminal, goes to the scene to check the object, and then inputs the response details into the terminal.
[1337] Step 11:
[1338] The server records the responses entered by the user (security guard), which will be used for future analysis and system improvement.
[1339] As a concrete example, consider the case where an unattended bag is left in the food court of a shopping mall. The server detects the unattended bag from surveillance camera footage. It compares this with past data to identify the bag's owner. It uses an emotion engine to check the stress level of the security guard and generates an alert in appropriate language to notify the guard. The security guard then checks the notification, checks the safety of the bag on-site, and enters the results into the server. This series of processes enables effective and human-friendly security operations.
[1340] Example 2
[1341] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1342] Conventional surveillance systems detect suspicious objects or lost items, but do not take into account the emotional state of those involved, which can lead to stress when dealing with the incident. Furthermore, detection accuracy can be low because the system does not compare the detected items with past video data. Furthermore, the recording and subsequent analysis of detected information is insufficient, which can hinder system improvement.
[1343] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for acquiring video data from a surveillance camera in real time, means for preprocessing the acquired video data and inputting it into a generative AI model for detecting objects, means for detecting objects using the generative AI model and identifying suspicious or lost objects, means for evaluating the emotional state of relevant parties using an emotion engine and adjusting the content of an alert when a suspicious or lost object is detected, and means for generating and notifying relevant parties of an alert. This enables highly accurate detection of suspicious or lost objects and appropriate responses that take into account the emotional state of relevant parties. In addition, detection accuracy is improved by comparing with past video data, and continuous system improvement is possible through recording and subsequent analysis of detection information.
[1344] A "surveillance camera" is a device that captures images of a pre-set location in real time.
[1345] A "server" is a computer system that obtains, processes, and manages data from other devices over a network.
[1346] "Real-time" refers to processing immediately without delay, and refers to a state in which appropriate action can be taken the moment a specific event occurs.
[1347] "Video data" means a collection of visual information captured from a surveillance camera or other device and stored as video frames.
[1348] "Preprocessing" refers to a series of operations that transform data into a suitable format before it is sent to the main processing stage, including, for example, resizing and normalization.
[1349] A "generative AI model" is an artificial intelligence algorithm that is pre-trained and used to perform a specific task (in this case, object detection).
[1350] "Subject" refers to an entity with specific characteristics, such as a person or object, that exists in the video.
[1351] A "suspicious object" is an object that is out of place or has been left unattended for a period of time and may pose a threat to normal activity.
[1352] "Lost property" refers to property that has been unintentionally abandoned or forgotten.
[1353] An "emotion engine" is software or algorithms that analyze the emotional state (e.g., stress, anxiety) of participants based on video and audio data.
[1354] An "alert" refers to a warning or notification message in response to a specific event (e.g., the discovery of a suspicious object).
[1355] "Notification" means the act of sending a message to inform interested parties of an alert or other important information.
[1356] The system for implementing this invention includes a surveillance camera, a server, a terminal, a user, and an emotion engine. Each component functions as follows.
[1357] surveillance cameras
[1358] Surveillance cameras capture video in real time from pre-defined locations. For example, if a surveillance camera is used in a food court in a shopping mall, it will monitor each area of the food court and capture video. The video from the camera is sent to a server via a specified URL or IP address.
[1359] server
[1360] The server acquires video data from the surveillance cameras in real time and processes it to detect suspicious objects and lost items. Specifically, it performs the following processes.
[1361] 1. Video data acquisition and preprocessing:
[1362] The server receives the video stream from the surveillance camera using the RTSP protocol, resizes the frame (e.g., from 1920x1080 to 640x480) and normalizes the pixel values using the OpenCV library.
[1363] 2. Object detection using AI models:
[1364] The preprocessed video frames are input into a generative AI model (e.g., YOLO, SSD) to detect objects (e.g., bags, people, etc.) in each frame, which includes a label for each object and its confidence score.
[1365] 3. Analysis of detection results:
[1366] The system analyzes detected object information to identify suspicious or lost items. In particular, if an unattended bag is left unattended for a certain period of time, it will be identified as a suspicious object. It will also compare the object with past video data to identify the occupant or suspicious person.
[1367] 4. Emotion Recognition with Emotion Engine:
[1368] Analyze the emotional state of the person involved (e.g., stress level, anxiety) using video and audio data, using an emotion recognition model (e.g., Microsoft Azure's Emotion API), and adjust the alert content if the person involved is in a high-stress state.
[1369] 5. Alert Generation and Notification:
[1370] If a suspicious or lost item is identified, an alert is automatically generated, tailored based on the results of the emotion engine, and sent to the relevant device (e.g., smartphone).
[1371] Terminal
[1372] The terminal (e.g., a security guard's smartphone) receives the alert sent from the server and displays it on the screen. For example, a notification such as "There is an unattended bag in the food court on the first floor. Please check." The terminal provides an interface that allows the security guard to input the details of the response after checking the scene.
[1373] User
[1374] The user (e.g., a security guard) checks the notification displayed on the terminal and takes action at the scene. For example, if the owner of an unattended bag is confirmed, the user enters the details of the action into the terminal.
[1375] Specific examples
[1376] The server receives real-time footage from surveillance cameras installed in the food court of a shopping mall and uses a generative AI model to detect unattended bags. If a bag is detected left in a specific area for a certain period of time, the server uses an emotion engine to evaluate the emotional state of the security guard and adjust the alert content. A notification is sent to the security guard's smartphone saying, "There is an unattended bag in the food court on the first floor. Please check." The security guard then heads to the scene, checks the safety of the bag, and enters the details into the terminal.
[1377] Example prompt sentence:
[1378] "Detect unattended bags from surveillance camera footage in a food court and calculate the time they have been left there. Also, recognize the emotional state (stress level) of people in the area."
[1379] This system enables highly accurate detection of suspicious objects and lost items in real time and rapid response, improving environmental security. In addition, by combining it with an emotion engine, flexible responses can be made taking into account the emotional state of those involved.
[1380] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1381] Step 1:
[1382] Acquiring video data
[1383] Server: Specify the URL or IP address of the surveillance camera and receive the video stream in real time using the RTSP protocol. The input is the video data from the surveillance camera, and the output is the captured real-time video frames. The server captures the video at a rate of 30 frames per second and stores it in memory.
[1384] Step 2:
[1385] Video data preprocessing
[1386] Server: Converts captured video frames into a format that is easy for the generative AI model to process. Specifically, it uses the OpenCV library to resize the frames (e.g., from 1920x1080 to 640x480) and normalizes pixel values to the range of 0-255. The input is a video frame captured in real time, and the output is a preprocessed video frame.
[1387] Step 3:
[1388] Object detection with AI models
[1389] Server: Inputs preprocessed video frames into a generative AI model (e.g., YOLO, SSD) to detect objects in each frame. The input is the preprocessed video frames, and the output is each object's label (e.g., bag, person), its coordinates, and a confidence score. The server updates the list of detected objects for each frame.
[1390] Step 4:
[1391] Analysis of detection results
[1392] Server: Analyzes the detected object information and identifies suspicious or lost objects. For example, if an unattended bag is left unattended for a certain period of time, it will be recorded as a suspicious object. The input is a list of detected objects and a confidence score, and the output is a list of suspicious or lost objects. The server analyzes the object status based on specific criteria and stores the results in a database.
[1393] Step 5:
[1394] Comparison with past data
[1395] Server: Compares newly detected suspicious or lost objects with past video data. Specifically, it compares the object with past records to identify the occupant or suspicious person. The input is a list of suspicious or lost objects and past video data, and the output is the identification information of the occupant or suspicious person. The server compares the object with past video frames stored in a database and analyzes the object's history.
[1396] Step 6:
[1397] Emotion recognition by emotion engine
[1398] Server: An emotion engine uses video and audio data to recognize the emotional state of the participants. This uses an emotion recognition model (e.g., Microsoft Azure's Emotion API). The input is video and audio data captured in real time, and the output is the participants' emotional information (e.g., stress level, anxiety). The server analyzes the participants' facial expressions and tone of voice to evaluate their emotional state.
[1399] Step 7:
[1400] Alerting and Tuning
[1401] Server: Detects suspicious or lost items and adjusts the alert content based on the results of the emotion engine. The input is a list of suspicious or lost items and emotional information, and the output is an alert message sent to the relevant person. For example, a person in a high stress state might be notified in a gentler way, such as "There is a bag in the food court on the first floor. Please check it." The server adjusts the alert content and sends it to the relevant person's device.
[1402] Step 8:
[1403] Record of correspondence
[1404] Server: Records the responses of relevant parties to generated alerts. The input is the response details from the relevant parties, and the output is the recorded response data. The server stores the response results entered by the relevant parties into a database (e.g., confirming the owner of the bag, confirming there is no threat) for later analysis and system improvement.
[1405] (Application example 2)
[1406] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1407] Conventional surveillance systems have limitations in accuracy and real-time detection of suspicious objects and lost items, making it difficult to respond quickly. Furthermore, because the emotional state of the user (such as a security guard) is not taken into consideration, the content of notifications and response methods cannot be optimized, resulting in stress for those involved and efficiency issues. The objective of this invention is to enable accurate real-time detection of suspicious objects and lost items while also enabling appropriate notifications and responses according to the user's emotional state.
[1408] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1409] In this invention, the server includes means for acquiring video data from a monitoring device in real time, means for preprocessing the acquired video data and inputting it into a generative AI model for detecting objects, means for detecting objects using the generative AI model and identifying suspicious objects or lost items, means for recognizing a user's emotional state using an emotion engine and reflecting that information in the analysis, and means for generating an alert and notifying the user according to the user's emotional state when a suspicious object or lost item is detected. This enables highly accurate detection of suspicious objects and lost items in real time, and also enables detailed notifications and responses according to the user's emotional state.
[1410] A "monitoring device" is a device that acquires video data in real time and transmits it to a server.
[1411] "Means for acquiring" refers to a method or device for collecting video data from a surveillance device.
[1412] "Preprocessing" is the process of converting the acquired video data into a format that is easy for the generative AI model to process.
[1413] A "generative AI model" is a program that uses machine learning and deep learning techniques to detect and analyze objects in video data.
[1414] The term "object" refers to an object that is the target of detection within the video data.
[1415] The "emotion engine" is a program that analyzes video and audio data to recognize the user's emotional state.
[1416] "Reflecting in the analysis" means incorporating the emotional information recognized by the emotion engine into the analysis results of the generative AI model.
[1417] "Alert" means a warning or notification generated when a suspicious or lost item is detected.
[1418] "Means for notifying" refers to a method or device for communicating a generated alert to a user.
[1419] The system for implementing this invention is composed of a monitoring device, a server, a terminal, a user, and an emotion engine. Each element will be described in detail below.
[1420] The server acquires video data from the surveillance equipment in real time. The acquired video data undergoes preprocessing before being input into the generative AI model. This preprocessing includes data processing such as resizing and normalizing the video frames. The preprocessed video data is then input into the generative AI model for object detection. The generative AI model achieves highly accurate object detection using traditional machine learning and deep learning techniques. Detected objects are analyzed to identify suspicious objects or lost items. By comparing the data with past video data, it is also possible to identify the occupant of the object or a suspicious person.
[1421] The emotion engine recognizes the user's emotions from video and audio data. The emotion information recognized by the emotion engine is reflected in the analysis on the server. This emotion information is used when generating alerts, and the alert content and notification method are adjusted according to the user's emotional state.
[1422] The terminal is a device that is primarily used by users (such as security guards). Notifications from the server are sent to the terminal, and an alert is displayed to the user. The content of the alert is adjusted based on emotional information recognized by the emotion engine. For example, if a security guard is in a state of high stress, the content of the notification can be softened to reduce the stress.
[1423] The user is the person who primarily operates the terminal and takes appropriate action based on notifications from the system. The user checks the notification and takes action on-site. After that, they provide feedback to the server by entering the details of their response into the terminal. This feedback information can be used for subsequent analysis and system improvement.
[1424] As a concrete example, consider the case of monitoring a food court in a shopping mall. Surveillance cameras send video data from inside the food court to a server in real time. The server preprocesses the video data and uses a generative AI model to detect unattended bags as suspicious objects. It confirms that the detected bag has been left unattended for a certain period of time and compares it with past video data to identify the owner. It also uses an emotion engine to evaluate the stress level of the security guard. Based on the detection results, an adjusted alert is generated and notified to the security guard. The security guard's terminal displays the message, "There is a bag in the food court on the first floor. Please check."
[1425] An example of a prompt sentence can be expressed as follows:
[1426] "Design a system that uses a generative AI model to detect suspicious objects and lost items from video data from surveillance equipment, and recognizes the user's emotional state using an emotion engine. Based on the recognition results, notify the user and encourage them to take the appropriate action."
[1427] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1428] Step 1:
[1429] The server acquires video data from the monitoring device in real time. To acquire video data, the server receives the video stream via the IP address or URL of the monitoring device. The video data is then passed to the server as input. The acquired video data is then saved in the server and formatted for subsequent processing.
[1430] Step 2:
[1431] The server preprocesses the acquired video data, specifically by resizing and normalizing the video frames. The input to this step is the raw video frames, and the output is preprocessed video frames in a format that is easy for the generative AI model to analyze.
[1432] Step 3:
[1433] The server inputs the preprocessed video data into a generative AI model to detect objects. Specifically, the generative AI model detects objects in the video frames and calculates their labels (e.g., bag, person, etc.) and confidence scores. The input for this step is the preprocessed video frames, and the output is a list of detected objects.
[1434] Step 4:
[1435] The server analyzes the confidence scores of the objects detected by the generative AI model and identifies them as suspicious or lost. Specifically, the server compares each object's label and confidence score with the set criteria to determine whether it is suspicious or lost. The input for this step is a list of detected objects, and the output is the identification of suspicious or lost objects.
[1436] Step 5:
[1437] The server compares the detected object with past video data and identifies the occupant or suspicious person of the object. Specifically, it searches for related video data from a past database and confirms the object's presence time and past owners. The input to this step is the identification result of the suspicious object or lost item, and the output is detailed information about the object (the identification result of the occupant or suspicious person).
[1438] Step 6:
[1439] The server uses an emotion engine to recognize the user's emotional state from video and audio data. Specifically, it inputs video frames and audio clips into the emotion engine and outputs an emotion label (e.g., high stress, low stress). The input for this step is the current video frame or audio data, and the output is the user's emotional state.
[1440] Step 7:
[1441] When a suspicious object or lost property is detected, the server generates an alert based on the recognition results of the emotion engine and notifies the user. Specifically, it creates an alert message (e.g., "Caution: A highly suspicious object has been detected. Please act with caution") according to the security guard's emotional state and sends it to the terminal. The input of this step is the suspicious object identification result and the emotional state, and the output is a notification alert.
[1442] Step 8:
[1443] The terminal displays the alert sent from the server and notifies the user. Specifically, the notification alert pops up on the screen, prompting the user to confirm it. The user then takes appropriate action based on the notification and inputs the response results into the terminal. The input for this step is the notification alert, and the output is the response results.
[1444] Step 9:
[1445] The server records the response results received from the terminal and uses them for later analysis and system improvement. Specifically, it saves the recorded response results in a database and uses them later for system evaluation and improvement. The input of this step is the response results, and the output is the recorded response data.
[1446] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1447] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1448] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1449] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1450] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1451] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1452] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1453] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1454] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1455] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1456] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1457] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1458] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1459] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1460] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1461] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1462] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1463] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1464] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1465] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1466] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1467] The following is further disclosed regarding the above embodiment.
[1468] (Claim 1)
[1469] A means for acquiring video data from a surveillance camera in real time;
[1470] A means for pre-processing the acquired video data and inputting it into a generative AI model for object detection;
[1471] A means for detecting objects using a generative AI model and identifying suspicious or lost items;
[1472] a means for generating an alert and notifying relevant parties when a suspicious or lost item is detected;
[1473] A system including:
[1474] (Claim 2)
[1475] 2. The system according to claim 1, further comprising means for identifying an occupant of an object or a suspicious person using past video data.
[1476] (Claim 3)
[1477] 2. The system according to claim 1, further comprising means for recording detected information based on the detection of a suspicious or lost item, for use in subsequent analysis and improvement.
[1478] "Example 1"
[1479] (Claim 1)
[1480] means for acquiring video data from a monitoring device in real time;
[1481] means for pre-processing the acquired video data and inputting it into an artificial intelligence model for detecting objects;
[1482] means for detecting objects and identifying suspicious or lost items using an artificial intelligence model;
[1483] a means for generating an alert and notifying relevant parties when a suspicious or lost item is detected;
[1484] A system including:
[1485] (Claim 2)
[1486] 2. The system according to claim 1, further comprising means for identifying an occupant of an object or a suspicious person using past video data.
[1487] (Claim 3)
[1488] 10. The system of claim 1, further comprising means for recording detected information based on the detection of a suspicious or lost item, for use in subsequent analysis and improvement.
[1489] "Application Example 1"
[1490] (Claim 1)
[1491] A means for acquiring video data from a surveillance camera in real time;
[1492] A means for pre-processing the acquired video data and inputting it into a generative AI model for object detection;
[1493] A means for detecting objects using a generative AI model and identifying suspicious or lost items;
[1494] A means for generating alerts to relevant parties about detected suspicious or lost items and notifying them via email or a dedicated application;
[1495] a means for displaying detection results and alerts to relevant personnel in real time using a head-mounted display;
[1496] A system including:
[1497] (Claim 2)
[1498] 2. The system according to claim 1, further comprising means for identifying an occupant of an object or a suspicious person using past video data.
[1499] (Claim 3)
[1500] 2. The system according to claim 1, further comprising means for recording detected information based on the detection of a suspicious or lost item, for use in subsequent analysis and improvement.
[1501] "Example 2: Combining Emotion Engines"
[1502] (Claim 1)
[1503] A means for acquiring video data from a surveillance camera in real time;
[1504] A means for pre-processing the acquired video data and inputting it into a generative AI model for object detection;
[1505] A means for detecting objects using a generative AI model and identifying suspicious or lost items;
[1506] a means for evaluating the emotional state of the relevant parties using an emotion engine and adjusting the alert content when a suspicious or lost item is detected;
[1507] a means for generating and notifying relevant parties of alerts;
[1508] A system including:
[1509] (Claim 2)
[1510] 2. The system according to claim 1, further comprising means for identifying an occupant of an object or a suspicious person using past video data.
[1511] (Claim 3)
[1512] 10. The system of claim 1, further comprising means for recording detected information based on the detection of a suspicious or lost item, for use in subsequent analysis and improvement.
[1513] "Application example 2 when combining emotion engines"
[1514] (Claim 1)
[1515] means for acquiring video data from a monitoring device in real time;
[1516] A means for pre-processing the acquired video data and inputting it into a generative AI model for object detection;
[1517] A means for detecting objects using a generative AI model and identifying suspicious or lost items;
[1518] A means for recognizing the user's emotional state using an emotion engine and incorporating that information into analysis;
[1519] means for generating an alert and notifying the user in response to the user's emotional state when a suspicious or lost item is detected;
[1520] A system including:
[1521] (Claim 2)
[1522] 2. The system according to claim 1, further comprising means for identifying an occupant of an object or a suspicious person using past video data.
[1523] (Claim 3)
[1524] 2. The system according to claim 1, further comprising means for recording detected information based on the detection of a suspicious or lost item, for use in subsequent analysis and improvement. [Explanation of symbols]
[1525] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for acquiring video data from a surveillance camera in real time; A means for pre-processing the acquired video data and inputting it into a generative AI model for object detection; A means for detecting objects using a generative AI model and identifying suspicious or lost items; a means for generating an alert and notifying relevant parties when a suspicious or lost item is detected; A system including:
2. 2. The system according to claim 1, further comprising means for identifying an occupant of an object or a suspicious person using past video data.
3. 2. The system of claim 1, further comprising means for recording detected information based on the detection of a suspicious or lost item, for use in subsequent analysis and improvement.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A