System and method for increasing situational awareness of alerts generated by a video surveillance system

By applying video analytics algorithms, Gen AI, and LLM to generate text summaries and context for alarms in video surveillance systems, the problem of alarms lacking context awareness is solved, improving operator decision-making efficiency and system automation capabilities.

CN122200459APending Publication Date: 2026-06-12HONEYWELL INTERNATIONAL INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511845790.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-12-11
Filing Date
2025-12-09
Publication Date
2026-06-12

AI Technical Summary

Technical Problem

Existing video surveillance systems generate alarms that lack context awareness, making it difficult for operators to understand the context of the alarms, leading to misoperations and inefficiency.

Method used

By applying video analytics algorithms to detect events and generate alerts in a video surveillance system, and combining generative artificial intelligence (Gen AI) and large language models (LLM) to generate text summaries and context for the alerts, the system outputs enhanced alerts to provide context awareness.

Benefits of technology

It improves operators' contextual awareness of alarms, reduces misoperations, and enhances the system's ability to make automated decisions and predict future alarms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122200459A_ABST
    Figure CN122200459A_ABST
Patent Text Reader

Abstract

Methods and systems are provided for increasing the situational awareness of alerts from a video surveillance system. Video analytics algorithms detect conditions in video streams and generate alerts. For each alert, a video clip containing frames before and / or after the alert is extracted. A generative AI video-to-text summarization model generates text summaries of the video frames, which are processed by a large language model to generate a context for each alert. Enhanced alerts containing both the alert type and the generated context are output to provide increased situational awareness. The system can store a history of alerts with timestamps for pattern analysis and prediction of future alerts. Additional features include multi-event correlation, root cause analysis, and detection of various conditions such as intrusion, loitering, and crowd formation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates generally to video surveillance systems, and more specifically to adding context awareness to alarms generated by video surveillance systems. Background Technology

[0002] Video surveillance systems typically employ various video analytics algorithms associated with different use cases, such as intrusion detection, loitering, people counting, and violence detection. These algorithms can operate at the edge or within the video surveillance system. Each video analytics algorithm identifies certain conditions or events occurring in the video stream of the video surveillance system. When an event is identified, an alert can be issued to the operator of the video surveillance system. In response, the operator must typically identify and subsequently review the video stream from the camera that captured the identified event to determine if the identified event is indeed of concern. If the event is not of concern, the operator can simply acknowledge the alert and proceed. If the event is of concern, the operator can execute a series of pre-defined standard operating procedures (SOPs) to resolve the alert. Alerts typically correspond to a specific event detected in the video by the video analytics algorithm. Alerts typically do not provide any “context” to the security operator, such as the situation or environment that caused the alert and / or what happens after the alert. Systems and methods that automatically determine the context for each alert and provide both the alert and the context to the security operator will be desirable to help increase the operator’s contextual awareness. The context for each alert can provide additional information about movement and / or behavior in the video before and / or after the alert. In some cases, the context can be used to preempt an upcoming alert. Summary of the Invention

[0003] This disclosure relates generally to video surveillance systems, and more specifically to increasing context awareness of alarms generated by video surveillance systems. An example may exist in a method for increasing operator context awareness of alarms generated by a video surveillance system. This exemplary method includes applying one or more video analytics algorithms to a video stream captured by the video surveillance system. Each of the one or more video analytics algorithms is configured to detect a corresponding condition occurring in the video stream, and in response to detecting the corresponding condition in the video stream, the corresponding video analytics algorithm is configured to provide an alarm with alarm metadata, wherein the alarm metadata may include an alarm type and one or more attributes of one or more objects detected in the video stream.

[0004] For each alert, a video clip is extracted from the video stream, comprising one or more video frames from the video stream preceding and / or following the corresponding alert. In some cases, the video clip may also include one or more video frames from the video stream that capture the corresponding alert. For each alert, a generative artificial intelligence (Gen AI) video-to-text summarization model is applied to the corresponding video clip to generate a text summary of each video frame from the one or more video frames preceding and / or following the corresponding alert. A text summary may also be generated for each video frame from the one or more video frames from the video stream that capture the corresponding alert. For each alert, a large language model (LLM) is applied to the text summary of one or more video frames from the corresponding video clip, and in some cases to at least some alert metadata in the alert metadata, to generate context for the corresponding alert. Example contexts may include “clustering of people,” “crowd formation,” “more movement in the area,” “sudden influx of vehicles,” “loud, continuous honking,” etc. For at least some of the alarms triggered by one or more video analytics algorithms, enhanced alarms are output, wherein the enhanced alarms include the alarm type of the corresponding alarm and the generated context for the corresponding alarm, and the enhanced alarms provide increased context awareness for the corresponding alarms. In some cases, the history of alarms and their corresponding contexts, each with a timestamp, are stored for subsequent pattern analysis and prediction. The occurrence of one or more future alarms within future time frames can be predicted based on the history of alarms and their corresponding contexts.

[0005] Another example may exist in a system for increasing operator context awareness of alarms generated by a video surveillance system. This exemplary system includes: an input unit for receiving a video stream captured by the video surveillance system; and a controller operatively coupled to the input unit. The controller is configured to apply one or more video analytics algorithms to the video stream, wherein each of the one or more video analytics algorithms is configured to detect a corresponding event or condition occurring in the video stream, and in response to detecting a corresponding condition in the video stream, the corresponding video analytics algorithm is configured to provide an alarm with an alarm type. The controller is configured to apply a video-to-text summarization model to the video stream to generate a text summary of one or more video frames of the video stream, the video stream including one or more video frames preceding each alarm provided by the one or more video analytics algorithms and / or one or more video frames following each alarm provided by the one or more video analytics algorithms. The controller is configured to apply a large language model (LLM) to text summaries of one or more video frames in a video stream to generate context for each alarm provided by one or more video analytics algorithms, the video stream including one or more video frames preceding and / or following each alarm provided by one or more video analytics algorithms. Text summaries can also be generated for each of the one or more video frames corresponding to the alarms captured in the video stream. The controller is configured to output enhanced alarms for at least some of the alarms provided by one or more video analytics algorithms, wherein the enhanced alarms include the alarm type of the corresponding alarm and the generated context for the corresponding alarm, wherein the enhanced alarms provide increased context awareness for the corresponding alarm.

[0006] Another example may exist in a non-transitory computer-readable medium storing instructions. When executed by one or more processors, the instructions cause one or more processors to apply one or more video analytics algorithms to a video stream, wherein each of the one or more video analytics algorithms is configured to detect a corresponding event or condition occurring in the video stream, and in response to detecting a corresponding condition in the video stream, the corresponding video analytics algorithm is configured to provide an alarm with an alarm type. The instructions also cause one or more processors to apply a video-to-text summarization model to the video stream to generate a text summary of one or more video frames of the video stream, including one or more video frames preceding each alarm provided by the one or more video analytics algorithms and / or one or more video frames following each alarm provided by the one or more video analytics algorithms. A text summary may also be generated for each of the one or more video frames corresponding to the capture of an alarm in the video stream. One or more processors are made to apply a large language model (LLM) to text summaries of one or more video frames in a video stream to generate context for each alarm in an alert provided by one or more video analytics algorithms, the video stream including one or more video frames preceding each alarm in the alerts, one or more video frames following each alarm in the alerts, and / or capturing one or more video frames for each alarm in the alerts provided by one or more video analytics algorithms. One or more processors are made to output enhanced alarms for at least some of the alarms in the alerts provided by one or more video analytics algorithms, wherein the enhanced alarms include the alarm type of the corresponding alarm and the generated context for the corresponding alarm, wherein the enhanced alarms provide increased context awareness for the corresponding alarm.

[0007] The foregoing description is provided to facilitate understanding of the innovative features unique to this disclosure and is not intended as a complete description. A full understanding of this disclosure can be obtained by considering the entire specification, claims, drawings, and abstract as a whole. Attached Figure Description

[0008] This disclosure can be more fully understood by taking into account the following description of various examples in conjunction with the accompanying drawings, in which:

[0009] Figure 1 This is a schematic block diagram illustrating an exemplary system for increasing operator contextual awareness of alarms generated by a video surveillance system;

[0010] Figure 2A and Figure 2B This is a flowchart illustrating an exemplary method for increasing operator contextual awareness of alarms generated by a video surveillance system;

[0011] Figure 3It is a flowchart illustrating a series of exemplary steps that can be executed by one or more processors that execute instructions stored on a non-transitory computer-readable medium;

[0012] Figure 4 This is a flowchart illustrating the overview;

[0013] Figure 5 This is a flowchart showing the details of the data aggregation module;

[0014] Figure 6 This is a flowchart showing the details of the context aggregation module;

[0015] Figure 7 This is a flowchart showing the details of the time data and context aggregation module;

[0016] Figure 8 This is a flowchart illustrating methods for enhancing context awareness; and

[0017] Figure 9 This is a flowchart showing an example output.

[0018] While this disclosure is subject to various modifications and alternatives, its details have been shown by way of example in the accompanying drawings and will be described in detail. However, it should be understood that this disclosure is not intended to limit it to the specific examples described. Rather, it is intended to cover all modifications, equivalents, and alternatives that fall within the substance and scope of this disclosure. Detailed Implementation

[0019] The following description should be read with reference to the accompanying drawings, in which similar elements in different drawings are numbered in a similar manner. The drawings are not necessarily drawn to scale and depict examples that are not intended to limit the scope of this disclosure. While examples of various elements are illustrated, those skilled in the art will recognize that many of the examples provided have suitable alternatives that can be utilized.

[0020] This document assumes that all numbers are modified by the term “about” unless otherwise explicitly stated. Expressions of numerical ranges using endpoints include all numbers contained within that range (e.g., 1 to 5 includes 1, 1.5, 2, 2.75, 3, 3.80, 4, and 5).

[0021] As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” include plural references unless otherwise expressly stated. As used in this specification and the appended claims, the term “or” is generally used in its meaning to include “and / or” unless otherwise expressly stated.

[0022] It should be noted that references to "implementation schemes," "some implementation schemes," or "other implementation schemes" in the specification indicate that the described implementation schemes may include specific features, structures, or characteristics; however, each implementation scheme need not necessarily include that specific feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same implementation scheme. Additionally, when a specific feature, structure, or characteristic is described in conjunction with an implementation scheme, it is conceivable that, whether explicitly described or not, that feature, structure, or characteristic may be applied to other implementation schemes, unless otherwise expressly stated to the contrary.

[0023] Figure 1 This is a schematic block diagram illustrating an exemplary system 10 for increasing operator contextual awareness of alarms generated by a video surveillance system 12. The exemplary system 10 includes an input unit 14 for receiving video streams captured by the video surveillance system 12. A controller 16 is operatively coupled to the input unit 14. In some cases, the controller 16 includes or has access to one or more video analytics algorithms 18. In some cases, the controller 16 includes or has access to a video-to-text summarization model 20. In some cases, the controller 16 includes or has access to a large language model (LLM) 22.

[0024] Controller 16 is configured to apply one or more video analysis algorithms from video analysis algorithms 18 to a video stream. Each of the one or more video analysis algorithms 18 can be configured to detect a corresponding event or condition occurring in the video stream, and in response to detecting a corresponding event or condition in the video stream, the corresponding video analysis algorithm 18 can be configured to provide an alarm with an alarm type. An event can be considered a condition. For each alarm, controller 16 is configured to apply a video-to-text summarization model 20 to the video stream to generate a text summary of one or more video frames of the video stream, which includes one or more video frames preceding each alarm provided by the one or more video analysis algorithms 18 and / or one or more video frames following each alarm provided by the one or more video analysis algorithms. A text summary can also be generated for each of the one or more video frames of the video stream that captures the condition that caused the alarm. The following shows a text-based example summary of three consecutive frames N, N+1, and N+2 of an example video stream:

[0025] Frame N

[0026] A middle-aged man, approximately 5 feet 10 inches tall, was crossing the parking lot. He wore a red hat that cast a shadow on part of his nose and forehead. His facial expression was neutral, his eyebrows relaxed, and his eyes slightly squinted due to the sunlight. His lips were slightly pursed, conveying a calm, focused demeanor. He wore a blue jacket, zipped up to the center of his chest, the fabric ruffling slightly around his elbows and shoulders as he swung his arms. His faded jeans were wrinkled near the knees, and he wore brown leather shoes. His right foot was firmly planted at (x:230, y:400), while his left foot was in the middle of a stride, hovering at (x:245, y:380). The gray asphalt beneath his feet was rough and cracked, with small fissures extending diagonally from (x:100, y:450) to (x:600, y:300). Yellow parking lines appear on either side, approximately 100 pixels apart, and are slightly worn from use. A red sedan is parked about 15 feet away, its front bumper visible at (x:500, y:590), and its windshield reflects bright sunlight. Glare on the windshield creates a bright spot at (x:510, y:580), and the car's body has small dust spots visible along its sides. The car's shadow extends eastward for approximately 120 pixels from (x:480, y:590) to (x:360, y:600). To the left of the scene, a row of bushes sways gently in the breeze, their green leaves casting intricate shadows on the ground from (x:10, y:20) to (x:100, y:100). In the distance, a concrete wall forms the boundary of the parking lot, extending horizontally across the top of the frame.

[0027] Frame N+1

[0028] The man continued walking, his left foot now lowered to the ground at (x:240, y:385), while his right foot began to lift slightly at (x:225, y:395). A faint expression of concentration graced his face; his lips, still tightly closed, were now slightly pursed, as if lost in thought. His red hat sat upright on his head, the shadow beneath the brim becoming more pronounced as the sun moved. His blue jacket swayed with his movement, creating more noticeable creases at the elbows. The jacket material shimmered slightly in the sunlight, especially around his left shoulder, where the sun shone at a particular angle. His jeans had gained several more creases near the knees, and his brown shoes were now lightly scraping the asphalt. The ground beneath his feet was more clearly visible, with cracks in the asphalt appearing more prominent near (x:110, y:470). The yellow stop line remained in place, but near his left foot, faint tire tracks were now visible, likely left by recent traffic. The red sedan remained parked, but the glare on its windshield had shifted slightly, now reflecting more sunlight at (x:515, y:585). Some dust particles were stirred up by a light breeze and floated behind the car at (x:510, y:610). The car's shadow had shortened slightly to 115 pixels, from (x:485, y:595) to (x:370, y:600). The man's shadow, also slightly shortened by the overhead sunlight, was now stretched by 115 pixels from (x:230, y:400) to (x:115, y:460). The bushes on the left swayed more violently, their leaves refracting sunlight and casting intricate shadows on the asphalt. The concrete wall in the background is now partially obscured by swaying leaves, with dappled sunlight filtering through them.

[0029] Frame N+2

[0030] The man's expression was slightly serious, his brows furrowed as if he were deep in thought. His left foot was now fully on the ground at (x:245, y:390), while his right foot was in mid-air at (x:225, y:400), indicating that he was walking with a purpose. As he turned his head slightly, the red hat on his head tilted slightly to the right, casting a longer shadow on the left side of his face. His blue jacket swayed gently, though new creases appeared on his back due to the movement. His right hand was in his jacket pocket, causing the jacket to tug slightly at his waist. His jeans were more wrinkled at the knees, especially his left leg, which stretched out more as he walked. A breeze stirred up some dust from the ground, visible near his left shoe at (x:250, y:415). The red sedan remained parked, but the sunlight reflected from its windshield had intensified, creating a larger glare at (x:520, y:590). As the sun moved slightly, the car's shadow continued to shift, now only 110 pixels long, from (x:480, y:595) to (x:365, y:600). The man's shadow also changed slightly, now stretching from (x:225, y:400) to (x:110, y:460). The bushes swayed more noticeably, and a few leaves fell, drifting across the parking lot, some landing at (x:150, y:600). Sunlight filtering through the bushes cast dappled shadows on the concrete wall behind them.

[0031] Controller 16 is configured to apply LLM 22 to text summaries of one or more video frames in a video stream to generate context for each alarm provided by one or more video analytics algorithms 18, the video stream including one or more video frames capturing the situation that caused the alarm, one or more video frames preceding the alarm, and / or one or more video frames following the alarm. In some cases, when applying LLM model 22, controller 16 may be configured to apply LLM model 22 to text summaries of one or more video frames capturing the situation that caused the alarm, one or more video frames preceding the alarm, and / or one or more video frames following the alarm, along with the corresponding alarm type, to generate context for each alarm provided by one or more video analytics algorithms. Example contexts may include “people clustering,” “crowd formation,” “more movement in an area,” “sudden influx of vehicles,” “loud, continuous honking,” etc. In some cases, controller 16 may be configured to apply LLM 22 to the text digest of each video frame in the video frame that captures the condition that caused the alarm, one or more video frames before the alarm, and / or one or more video frames after the alarm, to generate a frame context for each of the one or more video frames before, during, and / or after each alarm in the alarm, thereby generating a context for a specific alarm by applying LLM 22 to the frame context for one or more video frames before, during, and / or after the corresponding alarm.

[0032] Controller 16 is configured to output enhanced alarms for at least some of the alarms provided by one or more video analytics algorithms 18. The enhanced alarms include the alarm type of the corresponding alarm and the generated context for that alarm. The enhanced alarms provide increased context awareness for the corresponding alarms. In some cases, controller 16 may be configured to store the history of alarms and their corresponding contexts, each with a timestamp, for subsequent pattern analysis and prediction. The controller may be configured to predict the occurrence of one or more future alarms within future time frames based on the history of alarms and their corresponding contexts.

[0033] Figure 2A and Figure 2BThis is a flowchart illustrating an exemplary method 24 for increasing operator context awareness of alarms generated by a video surveillance system, such as video surveillance system 12. Method 24 includes applying one or more video analytics algorithms to a video stream captured by the video surveillance system, wherein each of the one or more video analytics algorithms is configured to detect a corresponding event or condition occurring in the video stream, and in response to detecting the corresponding event or condition in the video stream, the corresponding video analytics algorithm is configured to provide an alarm with alarm metadata, as indicated at box 26. As an example, each alarm may include metadata provided by the corresponding video analytics algorithm, wherein the metadata includes one or more of the following: alarm type, timestamp of the alarm, attributes of one or more objects and / or actors associated with the alarm, location of the camera of the video surveillance system capturing the video stream, and camera ID of the camera of the video surveillance system capturing the video stream. In some cases, one or more video analytics algorithms may be configured to detect one or more objects and / or actors in the video stream. In some cases, the conditions to be detected by one or more video analytics algorithms may include one or more of the following: people detected in the video stream, loitering detected in the video stream, intrusion detected in the video stream, predetermined behavior detected in the video stream, crowds detected in the video stream, specific faces detected in the video stream, specific vehicles detected in the video stream, object abandonment detected in the video stream, and violence detected in the video stream. These are just examples.

[0034] For each alert, a video clip is extracted from the video stream capturing the situation associated with the alert. The video clip includes one or more video frames from the video stream preceding and / or following the corresponding alert, as indicated in box 28. In some cases, the video clip may also include one or more video frames from the video stream capturing the corresponding alert. For each alert, a generative artificial intelligence (Gen AI) video-to-text summarization model is applied to the corresponding video clip to generate a text summary of each video frame in the one or more video frames preceding, during, and / or following the corresponding alert, as indicated in box 30. In some cases, applying the Gen AI model to the corresponding video clip may generate a text summary of each video frame in the one or more video frames preceding, during, and following the corresponding alert. For each alert, an LLM is applied to the text summary of one or more video frames in the corresponding video clip to generate context for the corresponding alert, as indicated in box 32. In some cases, applying an LLM model may include: for each alert, applying the LLM and the corresponding alert type to a text summary of one or more video frames of the corresponding video clip to generate context for that alert. In some cases, metadata provided by the corresponding video analytics algorithm may also be provided to the LLM model.

[0035] For at least some of the alerts generated by one or more video analytics algorithms, enhanced alerts are output, wherein the enhanced alerts include the alert type of the corresponding alert and the generated context for the corresponding alert, wherein the enhanced alerts provide increased context awareness for the corresponding alerts, as indicated in box 38. In some cases, the context of at least some of the alerts may include alert subject, alert object, and alert connecting preposition. As an example, alert connecting prepositions may include one or more of time, location, movement, manner, source, size, and occupancy. In some cases, enhanced alerts may identify one or more events, warnings, or alerts occurring at a threshold distance and a threshold time relative to the corresponding alert, and provide a multi-event correlation tree for increased context awareness. In some cases, method 24 may include storing the history of alerts and their corresponding contexts, each timestamped, for subsequent pattern analysis and prediction, as indicated in box 36. The occurrence of one or more future alerts within future time frames may be predicted based on the history of alerts and their corresponding contexts, as indicated in box 38. In some cases, activity patterns prior to and / or following at least some alarm types can be determined, at least in part, based on the history of alarms and their corresponding context, and the determined activity patterns can be reported to the operator.

[0036] continue Figure 2BMethod 24 may include, for each alarm, applying an LLM to a text summary of each video frame in one or more video frames preceding the corresponding alarm, one or more video frames during the corresponding alarm, and / or one or more video frames following the corresponding alarm, to generate a frame context for each video frame in the corresponding video frames, as indicated at box 40. Method 24 may also include generating a context for the corresponding alarm by applying an LLM to the frame context associated with one or more video frames preceding the corresponding alarm and / or one or more video frames following the corresponding alarm, as indicated at box 42.

[0037] In some cases, method 24 may include storing multiple historical alarms and / or historical enhanced alarms, as indicated at box 44. Multiple historical alarms and / or historical enhanced alarms for a video surveillance system may be used in combination with enhanced alarms to perform pattern analysis to provide additional context and situation awareness for the enhanced alarms. Pattern analysis may include analyzing the history of alarms, including associated objects, object actions, and / or object movement patterns, as indicated at box 46. In some cases, method 24 may include predicting future alarms based at least in part on pattern analysis, as indicated at box 48. In some cases, method 24 may include determining the root cause of one or more enhanced alarms among the enhanced alarms, based at least in part on pattern analysis, as indicated at box 50.

[0038] Figure 3 This is a flowchart illustrating a series of exemplary steps 52 that can be executed by one or more processors when they execute instructions stored on a non-transitory computer-readable medium. In some cases, the one or more processors may be a controller 16. Figure 1As part of ), one or more processors apply one or more video analytics algorithms to a video stream, wherein each of the one or more video analytics algorithms is configured to detect a corresponding event or condition occurring in the video stream, and in response to detecting the corresponding event or condition in the video stream, the corresponding video analytics algorithm is configured to provide an alarm with an alarm type, as indicated in box 54. One or more processors apply a video-to-text summarization model to the video stream to generate a text summary of one or more video frames of the video stream, the video stream including one or more video frames preceding each alarm provided by the one or more video analytics algorithms, one or more video frames during each alarm provided by the one or more video analytics algorithms, and / or one or more video frames following each alarm provided by the one or more video analytics algorithms, as indicated in box 56. One or more processors are instructed to apply a large language model (LLM) to text summaries of one or more video frames in a video stream to generate context for each alert provided by one or more video analytics algorithms, the video stream including one or more video frames preceding each alert, one or more video frames during each alert, and / or one or more video frames following each alert, as indicated in box 58. One or more processors are instructed to output enhanced alerts for at least some of the alerts provided by one or more video analytics algorithms, wherein the enhanced alerts include an alert type for the corresponding alert and a generated context for the corresponding alert, wherein the enhanced alerts provide increased context awareness for the corresponding alert, as indicated in box 60.

[0039] In some cases, one or more processors may be able to store the history of alarms and their corresponding contexts, each with a timestamp, for subsequent pattern analysis and prediction, as indicated in box 62. In some cases, one or more processors may be able to predict the occurrence of one or more future alarms within future time frames based on the history of alarms and their corresponding contexts, as indicated in box 64.

[0040] Figure 4This is a flowchart illustrating the overview of 66. A video stream is provided, as indicated at box 68. Video-to-Image Input 70 receives the video stream. Video-to-Image Input 70 communicates with Video Analysis Module 72 and with box 74, which processes image-to-text conversion using the GenAI tool. The output from Video Analysis Module 72 and box 74 includes events / data / metadata extracted / detected by Video Analysis Module 72 and context generated by the GenAI model in box 74. The context is typically represented in text form, such as "cat on the table," while Video Analysis Module 72 provides metadata output in addition to detected alerts / events, such as cat, table, etc. The output from boxes 72 and 74 is provided to box 75, which includes both Data Aggregation Module 76 and Context Aggregation Module 78. Box 75 outputs to Temporal Context Data Aggregation Module 80. Temporal Context Data Aggregation Module 80 appropriately concatenates the image data to include temporal variations of the video as input. The output of the Time Context Data Aggregation Module 80 contains summary information on metadata and context along the timeline, providing additional and richer information beyond warnings / events and metadata, resulting in enhanced context awareness. The output from the Time Context Data Aggregation Module 80 is provided to the Enhanced SA Module 82, which in turn provides output to the Refinement Box 84. The refinement of context / actions / SOPs, etc., is then performed by the operator or an artificial intelligence (AI) method. The Refinement Box 84 outputs to Box 75. The Enhanced SA Module 82 provides reports and warnings, as indicated at Box 86.

[0041] Figure 5 It is shown Figure 4 A flowchart detailing the data aggregation module 76 is provided. The data aggregation module 76 runs AI models or traditional computer vision techniques to detect different objects, such as people, vehicles, objects, and subcategories, as well as metadata. It also runs various video analytics modules 92 to obtain event or use case alerts, such as loitering, intrusion, behavior analysis, people counting, crowd counting, etc. These alerts / events and metadata are extracted and stored along with timestamps. The data aggregation module 76 primarily processes and stores different object or actor data, as well as alert / event data, forming the context-aware "data" portion. (Reference) Figure 5The system receives video, as indicated in box 88, and provides the video to the image input, as indicated in box 90. Images are sent from box 90 in several directions. Video analysis algorithm box 92 includes, for example, intrusion algorithm 92a, loitering algorithm 92b, behavior analysis algorithm 92c, people counting algorithm 92d, abandoned object detection algorithm 92e, and violence detection algorithm 92f. Video analysis algorithm box 92 outputs alarms (warnings) and metadata to box 96, which counts the alarms (warnings). Image input 90 is also provided to box 96. Image input 90 also outputs to object detection module 94. Both box 94 and box 96 output to box 98, which counts the detected objects. Box 96 outputs to aggregated object and alarm (warning) data box 100.

[0042] Figure 6 It is shown Figure 4 A flowchart detailing the context aggregation module 78 is provided. The context aggregation module 78 focuses on generating the "context" of the actor by automatically detecting prepositions, including prepositions for the actor's time, location, movement, manner, source, measurement, possession, and agency (detected in the data aggregation module 76). The context aggregation module 78 outputs the "context" of the actor / metadata / warning detected by the data aggregation module 76. In some cases, context generation is performed for each image (i.e., video frame) within + / - N minutes, seconds, or hours before / after the corresponding event / warning.

[0043] The system receives video, as indicated at box 88, and provides the video as image input, as indicated at box 90. Image input box 90 outputs to a Gen-AI model, as indicated at box 102. This could include, for example, GPT 4, LLaVA, and / or any other suitable Gen-AI model. Gen-AI model 102 outputs to image-to-text conversion box 104, which outputs to box 106. At box 106, the context of the image is extracted. This could include automatically detected prepositions, including prepositions for the time, place, movement, manner, source, measurement, possession, and agency of the actor. The extracted context is provided to box 108, where key context is extracted, and then provided to box 110. At box 110, the accumulated information is tabulated.

[0044] Figure 7This is a flowchart illustrating method 112, which details how the context generated for each image within a boundary of + / - N hours / minutes / seconds is combined into a single context that contributes to the final context awareness. Multiple video clips 114 include clips from before, during, and after the warning or event. Input images surrounding the warning or event are extracted, as indicated in box 116. Method 112 includes two paths. One path includes providing data to the data aggregation module 118 (which can represent...) Figure 4 The data aggregation module 76 outputs information to the data summarization module 120. The data summarization module 120 outputs to box 122, which stores relevant information before and after the alarm (warning) or event. A second path includes information provided to aggregation box 124. Aggregation box 124 outputs to global context extraction box 126. Global context extraction box 126 also outputs to box 122.

[0045] As can be seen in this exemplary implementation, the two paths of data 118 and context 124 are combined. In the data aggregation module 120, the number of actors, their classification and location, and actions such as running and walking are derived from the image set 114. Actors, their tracks in the video clip, and their actions related to movement can be captured. Tracking methods are used for this. The average number of actors and their movement behavior can also be captured in the data aggregation module 120. This is converted to text using any LLM-LLaVA for ease of understanding.

[0046] The second path 124 of the context is summarized in the global context generation module 126. In the global context generation module 126, each text representing the context of each image is combined again using Gen AI LLM or traditional Natural Language Processing (NLP) techniques to obtain a summary of different contextual interpretations. Typically, the context might look like "seeing a man enter the room, a cat suddenly jumping up from the table," etc. The final module 122 simultaneously stores the exported data in a database or any other storage mechanism, tagged with warnings / events and the timestamps considered. When the alarm (warning) or event is analyzed near real-time or during an investigation, the exported context and the generated warnings / events are presented to better understand the situation, and reports can be generated and sent via a messaging system or as audio input.

[0047] The data exported so far can be combined with the history and context of event / warning data to examine whether the pattern of the warning is the same as or similar to an earlier event. This can provide clearer guidance to any operator / facility manager or authority to better understand the cause and context of the warning than using video analytics for specific use cases, including behavioral analytics for violence, etc. Figure 8 This is a flowchart illustrating an enhanced context-aware method 128. Method 128 includes several inputs, including historical details as indicated at box 130, data and contextual data as indicated at box 132, and warning data as indicated at box 134. Data from boxes 132 and 134 is provided to box 126, where warnings and contextual data can be reported for each event. Outputs from boxes 130 and 136 are provided to a historical context summarization box 138. Both historical context summarization box 138 and box 140, which summarizes nearby events and warnings, output to a pattern analysis box 142. Output from pattern analysis box 142 is provided to a detection and prediction AI model, as indicated at box 144. At box 146, the context is cascaded and subsequently reported, as indicated at box 148.

[0048] Past and post-accident analyses may result in a variety of reports. Figure 9 This is a flowchart illustrating method 150 and the corresponding output from method 150. Box 152, including input video 154 and system status 156, outputs in several directions. Box 152 outputs to existing analysis, as indicated at box 158. Box 158 then outputs to metadata and warning details box 160. From there, data flows to enhanced SA box 162. Box 152 outputs to contextual analysis, as indicated at box 162, and outputs to a text report, as indicated at box 164. Both boxes 162 and 164 output to enhanced SA box 162. From there, data flows to box 168, where past incidents are analyzed, and to box 170, where current incidents are analyzed. Box 168 outputs the cause, as indicated at box 172, and outputs any repeating patterns, as indicated at box 174. From box 162, a summary or abstract is output, as indicated at box 176. From box 170, output the action as indicated in box 178, and output the analysis of the pattern as indicated in box 180.

[0049] Although several exemplary embodiments of this disclosure have been described thus, those skilled in the art will readily understand that other embodiments can be made and used within the scope of the appended claims. However, it should be understood that this disclosure is merely illustrative in many respects. Changes may be made to details, particularly those relating to shape, size, arrangement of parts, and exclusion and sequence of steps, without departing from the scope of this disclosure. The scope of this disclosure is, of course, defined by the language expressed in the appended claims.

Claims

1. A method for increasing operator context awareness of alarms generated by a video surveillance system, the method comprising: One or more video analysis algorithms are applied to a video stream captured by the video surveillance system, wherein each of the one or more video analysis algorithms is configured to detect a corresponding condition occurring in the video stream, and in response to detecting the corresponding condition in the video stream, the corresponding video analysis algorithm is configured to provide an alarm with alarm metadata, wherein the alarm metadata includes an alarm type and one or more attributes of one or more objects detected in the video stream; For each alarm, a video clip is extracted from the video stream, the video clip comprising one or more video frames of the video stream preceding the corresponding alarm and / or one or more video frames following the corresponding alarm; For each alert, a generative artificial intelligence (Gen AI) video-to-text summarization model is applied to the corresponding video clip to generate a text summary of each video frame in one or more video frames preceding the corresponding alert and / or one or more video frames following the corresponding alert. For each alert, a large language model (LLM) is applied to the text summary of the one or more video frames of the corresponding video clip and at least some of the alert metadata to generate a context for the corresponding alert; as well as For at least some of the alarms triggered by the one or more video analytics algorithms, an enhanced alarm is output, wherein the enhanced alarm includes the alarm type of the corresponding alarm and a generated context for the corresponding alarm, wherein the enhanced alarm provides increased context awareness for the corresponding alarm.

2. The method according to claim 1, wherein the method comprises: The history of the alarms and their corresponding contexts are stored, each with a timestamp, for subsequent pattern analysis and prediction; Perform one or more of the following: The activity patterns preceding and / or following at least some alarm types are determined based at least in part on the history of the alarms and their corresponding context, and the determined activity patterns are reported to the operator. as well as Based on the history of alarms and their corresponding context, predict the occurrence of one or more future alarms within a future time frame, and report the predicted future alarms to the operator.

3. The method according to any one of claims 1 or 2, wherein applying the LLM model comprises: For each alert, the large language model (LLM) and the corresponding alert type are applied to the text summary of one or more video frames of the corresponding video clip to generate the context for the corresponding alert; Optionally, the alarm metadata includes one or more of the following: the alarm type, timestamp, attributes of one or more objects and / or actors associated with the alarm, the location of the camera of the video surveillance system that captured the video stream, and the camera ID of the camera of the video surveillance system that captured the video stream.

4. The method according to any one of claims 1 or 2, wherein one or more video analysis algorithms are configured to detect one or more objects and / or actors in the video stream; Optionally, the situation to be detected by one or more video analysis algorithms includes one or more of the following: people detected in the video stream, loitering detected in the video stream, intrusion detected in the video stream, predetermined behavior detected in the video stream, crowds detected in the video stream, specific faces detected in the video stream, specific vehicles detected in the video stream, object abandonment detected in the video stream, and violence detected in the video stream.

5. The method of any one of claims 1 or 2, wherein the generative artificial intelligence (Gen AI) video-to-text summarization model is applied to the corresponding video clip to generate a text summary of each of one or more video frames preceding the corresponding alarm and one or more video frames following the corresponding alarm.

6. The method according to any one of claims 1 or 2, wherein the context for at least some of the alarms includes an alarm subject, an alarm object, and an alarm connection preposition; Optionally, the alarm conjunction is one or more of time, location, movement, manner, source, size, and possession.

7. The method according to any one of claims 1 or 2, wherein the method comprises: For each alert, the Large Language Model (LLM) is applied to the text summary of each video frame in one or more video frames preceding the corresponding alert and / or one or more video frames following the corresponding alert to generate a frame context for each video frame in the corresponding video frame; as well as The context for the corresponding alarm is generated by applying the large language model (LLM) to the frame context associated with the one or more video frames preceding the corresponding alarm and / or the one or more video frames following the corresponding alarm.

8. The method according to any one of claims 1 or 2, wherein: The enhanced alert identifies one or more events, warnings, or alarms that occur relative to a threshold distance and a threshold time period of the corresponding alarm; as well as Provides a multi-event relational tree for enhanced context awareness.

9. The method according to any preceding claim, further comprising: Receive multiple historical alerts and / or enhanced historical alerts; as well as Pattern analysis is performed using the multiple historical alarms and / or historical enhanced alarms for the video surveillance system in combination with the enhanced alarms to provide additional context and context awareness for the enhanced alarms. The pattern analysis includes analyzing the history of the alarms, including associated objects, object actions, and / or object movement patterns. Optionally, the method further includes one or more of the following: Predicting future alerts based at least in part on the aforementioned pattern analysis; and / or The root cause of one or more of the enhanced alerts is determined, at least in part, based on the pattern analysis.

10. A system for increasing operator context awareness of alarms generated by a video surveillance system, the system comprising: An input unit is configured to receive a video stream captured by the video surveillance system. A controller, operatively coupled to the input unit, is configured to: One or more video analysis algorithms are applied to the video stream, wherein each of the one or more video analysis algorithms is configured to detect a corresponding condition occurring in the video stream, and in response to detecting the corresponding condition in the video stream, the corresponding video analysis algorithm is configured to provide an alarm with an alarm type; The video-to-text summarization model is applied to the video stream to generate a text summary of one or more video frames of the video stream, the video stream including one or more video frames preceding each alarm in the alarms provided by the one or more video analysis algorithms and / or one or more video frames following each alarm in the alarms provided by the one or more video analysis algorithms; A large language model (LLM) is applied to the text summaries of the one or more video frames of the video stream to generate context for each of the alerts provided by the one or more video analysis algorithms, the video stream including the one or more video frames preceding each of the alerts provided by the one or more video analysis algorithms and / or the one or more video frames following each of the alerts provided by the one or more video analysis algorithms; as well as For at least some of the alarms provided by the one or more video analytics algorithms, an enhanced alarm is output, wherein the enhanced alarm includes the alarm type of the corresponding alarm and a generated context for the corresponding alarm, wherein the enhanced alarm provides the corresponding alarm with increased context awareness.