Geo-localized, crowd-sourced, video surveillance

The geo-localized, crowd-sourced video surveillance system addresses vulnerabilities in fixed cameras by using AI for real-time threat detection and adaptive monitoring, ensuring efficient and secure security response.

WO2025191566A1PCT designated stage Publication Date: 2025-09-18EYEIN AI LTD

Patent Information

Application Number
PCT/IL2025/050237
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-14
Filing Date
2025-03-13
Publication Date
2025-09-18

AI Technical Summary

Technical Problem

Fixed video cameras in surveillance systems are vulnerable to attackers, costly, and often understaffed, leading to inadequate real-time security monitoring.

Method used

A geo-localized, crowd-sourced video surveillance system using AI analysis with deep learning algorithms for real-time threat detection and classification, integrated with mapping for efficient monitoring and adaptable inference plans to optimize accuracy and speed.

Benefits of technology

Provides rapid threat alerts and adaptive monitoring across large and local areas, enhancing security response capabilities with scalable and secure management of video streams.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IL2025050237_18092025_PF_FP_ABST
    Figure IL2025050237_18092025_PF_FP_ABST
Patent Text Reader

Abstract

A system and methods are provided for area surveillance. The system is configured to acquire multiple video streams, each video stream associated with a geo-location indicator of the given device and to apply machine recognition by multiple specialized agents to images of the multiple respective video streams to generate indications of potential threats. Each Al agent is further configured with alternative early inference options. The one or more indications are then processed by a large language model (ELM) configured to determine an inference plan for processing the multiple respective video streams by an alternate set of Al agents, including alternate early inference options, to improve accuracy and efficiency of the processing by the Al agents, which are then triggered to operate according to the inference plan to generate a revised indication of a potential threat and to issue a corresponding alert.
Need to check novelty before this filing date? Find Prior Art

Description

GEO-LOCALIZED, CROWD SOURCED, VIDEO SURVEILLANCEFIELD OF THE INVENTION

[0001] The present invention relates to the field of video recognition using pattern recognition or machine learning.BACKGROUND

[0002] Threats to public safety are a growing problem in cities around the world. Local authorities and municipalities are finding it difficult to provide adequate security and monitoring solutions for their citizens. Increasingly, local authorities and municipalities are relying on fixed camera systems to monitor their cities. The global video surveillance market size is projected to grow at a CAGR of over - 12.5% from 2023 to 2035. The market revenue is expected to grow to USD 122 billion by the end of 2035, up from a revenue of -USD 50 billion in 2022.

[0003] However, the fixed video cameras of such surveillance systems can easily be mapped, targeted, and disabled by an attacker. Such systems, generally based on CCTV cameras are also expensive to install and maintain, and they are not always available in the right place at the right time. Control rooms designated for monitoring such cameras are often understaffed, and, consequently, operators are not always able to respond in real-time.SUMMARY

[0004] Embodiments of the present invention provide systems and methods for monitoring geographical regions for security threats, based on artificial intelligence (Al) analysis of crowd-sourced video streams. A real-time, Al, threat detection and classification system is provided with a threat analysis system that applies deep learning algorithms to generate event-based output of real-time threat alerts, thereby facilitating rapid analysis and response. The Al video analysis is integrated with mapping, allowing for easy monitoring, both of large geographical regions as well as local areas. The system also modifies the analysis in real time, by means of inference plans, that modify the types of Al agents run and levels of early inference applied to optimize the trade-off between accuracy and execution time.BRIEF DESCRIPTION OF DRAWINGS

[0005] For a better understanding of various embodiments of the invention and to show how the same may be carried into effect, reference is made, by way of example, to the accompanying drawings. Structural details of the invention are shown to provide a fundamental understanding of the invention, the description, taken with the drawings, making apparent to those skilled in the art how the several forms of the invention may be embodied in practice. In the accompanying drawings:

[0006] Fig. 1 is a schematic diagram of elements of a video surveillance system for processing geo-localized, crowd-sourced videos to identify potential threats, according to some embodiments of the present invention;

[0007] Fig. 2 is a schematic diagram of an artificial intelligence-based threat analysis system of the video surveillance system, providing aggregation, prioritization and scoring of threats, to generate human-readable alerts, according to some embodiments of the present invention;

[0008] Figs. 3 and 4 are schematic diagrams of elements of Al modules that form the threat surveillance system, according to some embodiments of the present invention; and

[0009] Fig. 5 is a flow diagram of a process for video surveillance performed by the video surveillance system, according to some embodiments of the present invention.DETAILED DESCRIPTION

[0010] In the following description, various aspects of the present invention are described. For purposes of explanation, specific configurations and details are set forth in order to provide a thorough understanding of the present invention. However, it will also be apparent to one skilled in the art that well known features may have been omitted or simplified in order not to obscure the present invention. With specific reference to the drawings, the particulars shown are by way of example and for purposes of illustrative discussion of the present invention in order to convey the principles and conceptual aspects of the invention. In this regard, structural details of the invention are only shown with the level of detail necessary for enabling the invention, the description taken with the drawings making apparent to those skilled in the art how the several forms of the invention may be implemented in practice.

[0011] Fig. 1 is a schematic diagram of elements of a system 20 for processing geolocalized, crowd-sourced videos to provide threat surveillance. The disclosed system may be implemented as a Video Surveillance as a Service (VSaaS) platform to provide secure, scalable, and efficient management of video streams, user interfaces, and administrative operations. The system generally includes real-time alerting, and options for user interaction across multiple interfaces, including cameras and monitors.

[0012] System 20 includes camera interfaces 30 (generally video streaming interfaces), which may be client or web applications operating on client devices 32, such as cameras, drones, smart phones, wearable devices and web cams. Of these devices, any that are not fixed in place are referred to herein as mobile devices, such devices typically being operated by camera operators 34. Client devices 32 may also include fixed cameras, such as CCTV cameras. Each camera has a system ID.

[0013] The camera interface serves as the entry point for video streams, enabling loT- enabled cameras to upload live or recorded video footage securely. Video streams are transmitted over encrypted channels and routed through a secure ingestion layer, ensuring data integrity and privacy.

[0014] Each camera device is typically registered with the system as an authenticated user with unique credentials, with permissions assigned for video upload and processing in a user’s region and / or network. Ingestion layers of the neural networks described below are configured to support real-time streaming, as well as supporting periodic uploads suitable for environments with intermittent connectivity and / or to support resource efficiency.

[0015] The camera interfaces may also allow users to send and receive alerts and text messages from the human monitors described below.

[0016] The streaming interfaces of each device are typically configured to stream live video to a data center interface 40, which in turn transmits the video streams to a threat analysis system 50. The data center interface 40 may be a cloud server that may use an authentication paradigm for connecting to clients, such as OAuth2.0 and JWT.

[0017] The threat analysis system 50 produces video analysis results, indicating the possible existence and estimated location of threats. Threats that may be detected by the threat analysis system, may include weapons (e.g., firearms, knives, etc.), as well as physical actions suggesting threats or violence (e.g., hitting, threatening, aiming and / orfiring a weapon, aggressive stance / posture). Profiling of possible assailants may also be performed (age, gender, etc.), as well as identification of individuals. In some examples, clothing may also be a parameter triggering threat events, such as presence of army personnel, police officers, paramedics, etc., or outfits indicative of terrorist affiliation, such as headbands, flags, and known symbols. Anomalous group size and motion may also be a factor in determining potential threats, as well as environmental events, such as fires, car accidents, traffic, unusual vehicle activity, etc. In general, any anomalous crowd event may indicate a threat. Possible threats may be categorized by severity level and likelihood, as described further hereinbelow. As described below, tasks of identification are performed by specialized Al agents, whose operation is modified in real time to provide better accuracy in evaluating possible threats.

[0018] Alerts are triggered by the threat analysis system 50 when video streams are determined to include indications of threats. The alerts may be sent to either or both of a control room interface 60 and an alert center 70. The control room interface 60 may be configured to present a “monitoring” client or web application, which may be viewed by control room operators 62. The control room interface 60 may also receive video streams from the data center interface 40, as the threats reference particular video stream channels and the control room interface is configured to the display of a threat with the referenced video streams, enabling operators to view a potential threat live. Results indicating a possible threat may also be issued to the alert center 70, to notify security authorities 72 in real-time, such as police or other armed forces. Threats may include action categories, such as terror attacks, violent actions, intrusion into secure areas, fires and hazards, etc. Threats may also include weapons, categorized by type.

[0019] The monitor interface allows authorized users to view live streams, replay recorded footage, and interact with Al-generated annotations. Users are notified of critical events, such as unauthorized access or unusual activity, through real-time alerts.

[0020] The threat analysis system, typically operated as a cloud application, may include real-time threat detection and classification at several levels of processing (depending on the available resources), including a deep learning architecture allowing computational-heavy inference to provide outputs for real-time alerting. The analysis may include computation of a real-time “global threat score” for a set of cameras or per individual camera, for example, as a number between 0 and 1. Different artificialintelligence (Al) agents may run to process each video stream, each producing a threat score and a connected alert. Alerts (also referred to herein as “alert events”) may be categorized into several levels of priority, depending on the threat score and the type of threat detected. The level of threat may be estimated with the application of higher level decision-making units, including Al or rule -based reasoning and / or similar methodologies. The level of threat may also be affected by user-initiated alerts and their volume and details, by the level of confidence of the Al, by geolocation, by the fact the stream or device was defined as a critical stream or device, or by subsequent analysis corroborating or refuting previously raised alerts.

[0021] The scoring method may also be configured to detect multiple threats and to weight the combinations to form a threat score (e.g., three suspicious people and two possible weapons gives an aggregate score that is the sum of the individual threat scores, or may be calculated by an alternative function to give a higher threat score).

[0022] Typically, each video stream is received together with a geographic indication, i.e., a “geo-location” indicator, of the camera recording the video, as well as an attached audio stream. The machine learning (ML) model of the threat analysis system is typically trained not only to recognize a potential threat (indicated by a threat-associated object and / or a threat-associated action) but to provide an estimated location of the potential threat, according to the geo-location indicator the one or more video streams in which the threat is indicated. When a threat (i.e., a potential / possible threat) is identified, an alert is issued that typically includes an estimated location of the threat, which may be presented on a map of the monitoring (“control room”) interface.

[0023] In further examples of the present invention, the threat analysis system may include audio analysis agents for threat detection and classification from audio streams. The threat analysis system may also include the ability to track and monitor moving people and objects across multiple video streams and to geo-localize their movements on a map. The Al agents may also connect to existing databases of known threats and suspects, and may perform facial and object recognition.

[0024] The “monitoring” client or web application of the control room interface 60 provides an interface for the operators that typically shows both a map of a given area under surveillance, as well as live video streams, streamed from the data center interface, ordirectly from the video upload interfaces. The interface may include real-time prioritization of video streams, real-time prioritization of alerts, and display of alert locations, together with corresponding video streams.

[0025] Human operators may define parameters of the system, such as a "minimal event distance," defining a minimal time gap between any two adjacent threat events. Any two events with a lower distance may be merged into a single, joint event. The motivation for joining the events is to eliminate noise that would be caused by misidentifying a single event as multiple events. Adaptive, region-of-interest definitions may also be provided for camera-specific setup parameters, based on motion-detection and statistical methods. For example, Al analysis may be performed, on the cloud or on the user’s device, to define and adapt a region of interest in the field of view of the camera and to produce a weighted map providing at pixel level the relevancy of a given part of the field of view. The relevancy map may be further refined by monitors and streamers in the camera settings.

[0026] To summarize, the monitoring interface allows authorized users to view live streams, replay recorded footage, and interact with Al-generated annotations. Users are notified of critical events, such as unauthorized access or unusual activity, through realtime alerts. The interface supports low -latency streaming optimized for desktop and mobile devices. Annotated videos include metadata such as event timestamps, detected objects, and system-generated tags and metrics, making it easy for users to review incidents efficiently. The interface provides users with the ability to elevate alerts to their superiors (e.g., admin users) and spread messages to all or specific camera users in their region / network.

[0027] Notifications can be delivered via multiple channels, including SMS, email, or synthesized voice messages, or through the camera interface to ensure timely responses to potential security threats.

[0028] The control center typically also has an administrator interface, which may be a centralized control panel for system administrators to manage the entire VSaaS platform. It provides tools to:• Overview the current state in the monitored area and review the latest alerts. Overview the cameras on the map, in tables, and overview the streams visuals.Overview the status and activity of the monitoring workforce.Onboard and configure cameras.• Manage user roles and permissions for monitors and other stakeholders.• Generate analytics and reports on system usage, event frequency, and performance metrics.• Monitor system health and resource utilization.• Define Al workflows to tailor the platform to specific operational needs (e.g. adapted Al for fence cameras and for street cameras). This could be done through a template selection and by specifically configuring rules and conditions through the application. For instance, the admin could select to raise alerts whenever a specific type of vehicle is detected, at specific times and / or locations.

[0029] This configuration also allows the admin to choose which Al processing to apply, and to which camera, e.g. apply uniform detection when a weapon is detected and apply face recognition on the detected police officers. Administrators can set parameters to:• Customize the alert-raising signals and rules (e.g. when a weapon is detected or when a bag is left unattended for 2 minutes, passage at specific hours, etc.).• Apply template and custom Al workflows to the different cameras• Customize the alerting protocols and communication (e.g., when and how to be alerted).• Add specific details about the expected type of image from the device (e.g. thermic drone camera) allowing to apply a specific preprocessing to the video.

[0030] In some embodiments, the admin interface can be part of a monitoring interface.

[0031] Fig. 2 is a schematic diagram of the threat analysis system 50, which includes an ensemble of Al agents for identifying threats in geo-localized, crowd-sourced videos.

[0032] As described above, video sources 32 transmit video streams (which may also include geo-location and audio data) to a data center interface 40. The data center interface provides the video streams to both the control room interface 60, as described above, and to an Al portal 200 of the threat analysis system 50.

[0033] The Al portal, which may be implemented as a RESTful API service, isconstantly called by the client video upload interfaces of the video sources. In addition, the Al portal receives inference plans that are associated with Al agents 206 that are called to analyze the incoming video streams. An inference plan may include instructions as to which Al agents should be applied, such as motion, person and / or weapon detection agents, as well as the processing level and accuracy to be applied by the agents, as described further hereinbelow. The Al portal, at each call, generates atomic tasks to be performed on specific video streams by one or more Al agents 206, each agent typically being configured as a neural network, with the set of agents forming an agent storage “ensemble” 204. Each task typically includes a priority level, a video stream reference, a set of assigned agents, and execution parameters including the early inference parameter, as described further hereinbelow. The Al portal adds the tasks to an inference queue 202, queuing the tasks for processing by the Al agents. The queue may be implemented with a job queue framework, e.g., well-known frameworks such as RabbitMQ™ or Kafka™, or implement a custom- designed queue to manage the inference of Al agents with input parameters and possibly the preferred allocated resources.

[0034] The Al agents 206 may each include one or more of various types and architectures of Al models such as: Convolutional Neural Networks (CNNs), Deep Neural Networks (DNNs), Multi-Layer Perceptrons (MLPs), Long Short-Term Memory / Recurrent Neural Networks (LSTM / RNN), Temporal Convolutional Networks (TCNs), Large Language Models (LLMs), and multimodal models, including rules-based models. The agents are trained on data representing their target areas of expertise (semantic role) and may be configured to receive multiple forms of data input, including videos and images. The set of agents conforming to each task establishes processing resources, generates agent inferences, and aggregates results. Each Al agent may also apply Al scenarios, pre-defined by a user, or based on default templates.

[0035] The Al agents typically have a common core set of neural network layers, followed by layers conforming to their area of specialty. Early Inference Subnetworks (EISNs) branch from the core agent to provide options for faster but potentially less accurate inference, by exiting the processing pipeline at intermediate layers. Agents typically can issue notifications if pre-defined criteria are met (e.g. one person was detected). The system may also be configured to direct agent alerts directly to users even if additional processing is still required to meet task parameters.

[0036] Agents may also incorporate additional pre-processing and / or post-processing for manipulation of the input and output at inference time. For example, agents may receive parameter information on an image, video stream, and device and apply an adapted preprocessing flow, for example, a type of image processing, such as infrared, drone, or image tiling.

[0037] Agents may use a mask on an image and / or video to avoid performing analysis on certain regions. The mask may be obtained directly from the user (e.g. by painting manually in the application the dead zone of a fixed camera) or from Al and algorithms (e.g. motionless regions). Agents may receive audio input solely or both sound and images as input.

[0038] Typically, Al agents specified by an inference plan perform inference on the input data and write their output results into an agent detection database 210.

[0039] These results are read at intervals (e.g., 1 read / sec) by an LLM Request Composer module of an Alert Aggregator 212. The database 210 typically provides direct write access by the Al agents and read access by the Request Composer. For instance, Al agents may access the database through an asynchronous HTTP call to the Engine Portal API, to store the threat detections, locations, certainty probabilities, etc.

[0040] The composer generates LLM text requests based on context and “insights” from previous calls. These “insights” refer to the results of previous analyses, which have been threats that triggered alerts, having scores above the threshold, or threats that did not trigger alerts, as they generated lower scores. For example, a suspicious act by an individual may not trigger a score above the threshold, but continued suspicious activity would eventually trigger an alert.

[0041] Each text request is sent to an LLM 224 that generates structured text output with specifications of a threat. The specification may indicate the type of threat (e.g., gun, knife, violence, suspicious person), and a threat score (e.g., a measure of confidence level and urgency of threat), as well as the geolocation and channel (i.e., camera supplying video stream). The structured text is then parsed (i.e., “decomposed”) by an LLM Post-Processing module 226. The elements of the parsed text include the “insights” described above, which are written to LLM memory storage 222 to be including in the next LLM analysis of the next set of threats (as read by a new call by the LLM Request Composer to the agentdetection DB 210). The LLM post-processing module 226 also parses from the LLM text a new inference plan, and sends the new inference plan to the data center interface, to be used for the next task executed (“inferred”) by the agents, as described above.

[0042] The LLM 224, aware of the available agents and cameras, outputs an inference plan specifying which agents should be employed, their priority, and for which camera. The LLM may specify which resolution of early inference to employ, and whether to apply quantized or distilled networks. The LLM may also be informed of the available computational power and instructed to specify on which resource to infer the network, allowing for optimized management of cloud and / or edge computing. The LLM may be instructed to provide in its response a specific early inference subnetwork to perform the inference, as well as an early inference threshold to use and the specific camera (e.g. "use person detection, Exit at X% accuracy, on Device Y. For example, all agents may be set to apply a certain level of early inference as a default. If there appears to be a threat, such as a knife, an inference plan may specify that a given Al agent that is trained to detect knives will skip its earlier early inference subnetworks, or perhaps skip early inference entirely, in order to achieve a higher level of confidence (i.e., accuracy) than is provided by the initial threat detection.

[0043] As a further example, if three suspicious people are detected, the Alert Aggregator may generate an inference plan to add processes of person recognition and motion tracking of the people. The motion tracking may be set to a threshold confidence of 0.3, so that a first early inference subnetwork is first applied, and, if the confidence results are too low, the agent is run again with a deeper EISN exit, the process repeating until producing a confidence level meeting the desired confidence level (or reaching the full network limit).

[0044] Similarly, the Alert Aggregator may specify a specific layer number for the early inference subnetwork to stop processing at within the core network. For example, the inference plan may indicate that processing should stop at layer 123 of the core network (even if the predefined threshold of 0.3 was not reached). This approach allows for more fine-grained control over the inference process and can be used together with the early inference threshold. The Alert Aggregator can also specify a starting layer for early inference, instructing the system to skip initial early inference subnetworks due to their low accuracy.

[0045] Similarly, the Alert Aggregator may be instructed to output and the early inference Network make take as input a layer number to start early inference, that is, to skip initial early inference Sub-networks due to their low accuracy.

[0046] Early inference networks are multi-headed neural networks that can be used to derive thinner networks, by taking the first layers of the core network up to a specific layer with the early inference subnetwork associated at that layer. This allows for creating multiple versions of the same network with different depths. Fewer layers generally means less accuracy, but faster inference times. This could also be implemented with quantized and distilled networks adapted to the edge devices, i.e., to run on mobile devices. The LLM request may include details on the available edge and cloud devices, with constant updates at each call to provide the most up-to-date information.

[0047] The context and the notifications of the Al agents are typically provided in a structured way to the LLM, in a human-readable format. A context model is typically provided as a dictionary (JSON, YAML, etc.) and conveys information regarding the role and goal of the aggregator in the context of the Al agents and the video surveillance application, with respect to threat scenarios in focus, the availability of Al agents and their capabilities, threat scores with numerical examples, and the expected output, its format, and with at least one output example. This enables the aggregator to act in an architecture-aware manner, to produce inference plans for the Al agents for specific cameras / devices, like a selective attention mechanism, with a deeper analysis applied dynamically where and when required.

[0048] In some embodiments, the LLM is resource-aware, generating an inference plan that indicates the agent deployment (edge computing or on the cloud) to use for the processing, as well as the inference speed, designating the speed or accuracy of the model to use (e.g. quantized, distilled, or early inference) to set the optimal trade-off between speed and accuracy.

[0049] A memory storage can be used to store part of the LLM's output, such as the threat score, a concise summary of the situation, and the insights and thoughts of the LLM on the situation. This short-term memory information is then conveyed in the next LLM request, for a more coherent, structured, and traceable chain of thoughts and decisions. For example, "Threat Score: 81% - Camera 123: 3 people detected. Suggesting to check forweapons and violent actions.

[0050] To summarize, the alert aggregator provides continuous alert integration with the video surveillance system and provides indications of the additional processing to apply, and which agent to apply in the multi-agent Al system, prioritizing Al agents and cameras for a deeper dynamic analysis, and conveying insights and reasoning to the request for the next periodic call of the LLM (typically in ~l-2 seconds).

[0051] The object detections are conveyed in a structured way, for example, with the following information:{"name": label e.g. person, "probability": float in [0,1], "camera_id": "Cameral23", "timestamp": "2022-01-01T00:00:00Z", "box": [xl, yl, x2, y2] in relative coordinates or [xl, yl, x2, y2, w, h] in absolute coordinates, "threat_score": float in [0,1] interval, "context": "person detected", "agent": "person_detection", "confidence": float in [0,1] }

[0052] The threat score is a way to reduce the complexity of the situation to a single number, that can also be interpreted as the required level of attention from a human monitor. The camera threat score can be used to dynamically reorder the cameras in the monitoring interface.

[0053] The Al agents and their detected labels may be configured as following:• SSDX: people, cars, dogs, cats, bicycles, motorcycles, trucks, buses, vans• SSDX-Mini: a lighter version of SSDX with people, vehicles and animals• Weaponsl: weapons, knives, guns, rifles, pistols• Weapons2: rods, bats, sticks, molotov, grenades• Weapons3: M16, Kalashnikov, shotguns, sniper rifles, RPGs• FacesDet: faces, masks, helmets• Violence: fighting, shooting, stabbing, hitting, kicking• Posture: lying, standing, sitting, running, walking, crouching• Wreckage: Fire, smoke, wreckage, debris, explosion

[0054] Additional agents may include: Face Recognition, LPR, Motion Detection, andSuper Resolution.

[0055] Default Agents may be set to: Motion Detection, SSDX-Mini.

[0056] Examples of agent details may include:Weapons 1 Agent (Sens / Spec: 0.9 / 0.6), produces a lot of false positives (FPs) for knives, but is very accurate for guns.Weapons3 is slower and should be used only when Weaponsl is triggered.Threat Events and Reference Threat Scores may be :- Violent actions: 1- Explosions, fires, wreckage: 0.9- Errant vehicles / people: 0.8- Car stopping at restricted areas: 0.8- Person lying on the ground: 0.7- Unattended bags: 0.7- 10 or more people at night: 0.8- 3 or more people at night: 0.5- Errant dogs: 0.2

[0057] Each detected event may increase a camera threat score. A global threat score is calculated from aggregating the individual camera threat scores.

[0058] A first exemplary output alert may include the following:{ "Alert": "possible person","Summary": A concise description of the overall situation, justifying the threat score, strategy and inference plan."Alert level:": A string indicating the overall alert level, such as "LOW", "MEDIUM", "HIGH", "CRITICAL"."Global Threat Score": A score between 0 and 1, indicating the overall threat level where 1 is the highest threat level."Threat location": {lat, Ion], # a guesstimation of the threat geolocation, ifrelevant"Confidence": float in [0,1],"Timestamp": string with "YYYY-MM-DDTHH:MM:SSZ" format (ISO 8601)"Inference Plan": A diet by cam_id: {cam id: {agent: priority, reason, timestamp, sequence, compute, alert_if } ... }, indicating the agents to apply and the strategic reasons. Timestamps indicate which frame to process, or future frames if negative (eg -Is, -2m, -3h). compute is "edge", or "cloud" by default, sequence is freetext tag for defining processing chains. alert_if is a logical sentence "<label> COUNT / SCORE is EQUAL TO / ABOVE / BELOW value", eg "masks COUNT is ABOVE 0”}

[0059] A second exemplary output alert may include the following:{ "Alert": "Personnel check in restricted area","Summary": "A group of 3 people has been detected in a restricted area. Compare to known personnel, check masks, weapons and wreckage."Alert Level": "MEDIUM","Global Threat Score": 0.6,"Threat location": {"lat": 30.001, "Ion": -30.001],"Confidence": 0.8,"Timestamp": "2022-01-01T01:02:05Z","Inference Plan": {"caml21": [{"FacesDet": 1, "priority": 5", "timestamp": "2022-01- 01T01:02:03Z", "reason": "restricted area access", "sequence" / 'Personnel check" }{"Super Resolution": 1, "sequence" / 'Personnel check"},{"Face Recognition": 1, "sequence" / 'Personnel check", alert_if: "masks COUNT is ABOVE 0"}{"Weaponsl": 1, "priority": 4", "timestamp": "-Is", "reason": "checking access", "compute"= "edge"} },{"Wreckage": 1, "priority": 3", "timestamp": "-Is", "reason": "checking access", "compute"= "edge"}{"Violence": 1, "priority": 2", "timestamp": "-Is", "reason": "checking access", "compute"= "edge"}]},"Details": "Edge computing is recommended on the parallel agents for faster processing. An alert will be triggered if a face mask is detected."}}

[0060] The results of the detections are then conveyed as text as specified in the above context, to compose the whole LLM request.

[0061] Fig. 3 is a schematic diagram of the architecture of each Al agent 206, the multiple Al agents forming the ensemble of the agent ensemble 204 described above. A deep neural network 302 (DNN) of the Al agent receives input data 300, including video streams with geo-located references, and inference plans. As describe above, each Al is specialized, trained to focus on a particular detection task such as weapon detection, person identification, or motion tracking.

[0062] Each agent generates output data 304 indicating possible threats that typically including both “global threat scores,” and “agent threat scores.” The agent threat scores may be based on the inclusion of several Al agents, each outputting events in addition to the score, such as threat detection and classification based on face recognition, objects recognition, car detection, audio analysis, etc. Output is typically generated together with a confidence score.

[0063] The system may also apply parameters of prioritization of cameras being monitored, for real-time tracking of events. Threats (i.e., objects or individuals) may be tracked across multiple, geolocation-enabled video streams. The architecture may be used to adapt existing models to produce faster (i.e., earlier) outputs.

[0064] To efficiently gain inference speed, each agent may cut processing short at internal layers that connect to early inference subnetworks (EISNs), indicated, by way ofexample, by EISNs 306a and 306b, which are shown as generating respective outputs 310a and 310b. These EISNs branch from different depths within each agent's specialized neural network layers, allowing for trade-offs between processing speed and accuracy. Regardless of the EISN, the output results are written to the agent detection database 210, described above. The EISN and the full core network share the same layer architecture but differ in their sizes and in their weights after training. As shown, the inference time is shorter for EISNs linked to earlier layers of the core DNN.

[0065] Fig. 4 is a schematic diagram of the architecture of an early inference subnetwork (EISN). The architecture includes two parts: a core neural network 402, having multiple layers, as described above, and a network ensemble 306, connected to an intermediate layer of the core network. The early inference sub-networks (EISNs) are neural networks aimed at solving the same task and providing an output that is the same or similar to that of the core DNNs.

[0066] Each EISN has three main elements: a Layer Connector 410, a Model Estimator 420, and a Predictor 430.

[0067] The Layer Connector performs dimensionality reduction and preprocessing of an output tensor (a “raw feature tensor”) 404 from a given layer of the Core DNN. Its output is a one-dimensional feature vector 412 of limited size that can be rapidly processed by the model estimator. This can be attained with a Flattening layer ensuring the one- dimensionality of the output, followed by a Pooling layer for data reduction (e.g. Averaging, Sub-sampling, etc.) to reduce the size of the feature vector.

[0068] The Layer Connector may be adapted, in case of multimodal input (e.g. video with sound, or image and text) or in case of parallel processing branches in the Core DNN, to independently process the input for each mode and to fuse result vectors in a concatenation layer before the Model Estimator and Predictor are applied.

[0069] The Model Estimator is a generic artificial intelligence agent composed of DNN layers, aimed at building a basic world representation and understanding from fitting the ID output of the intermediate core layer directly to the task at hand. Compared to the core DNN, the goal is to significantly reduce computation in this step and have a network with fewer parameters than the remainder of the core net, and a rapid inference time.

[0070] The number of parameters (i.e. network weights, e.g. MLP layer size) in theModel Estimator may be set at each layer proportionally to the number of parameters in the remainder layers of the core net, i.e., from this layer to the core output layer, e.g. 2%. In some embodiments, this relationship may be defined as a logarithmic, polynomial or custom function.

[0071] The Model Estimator can be implemented with more sophisticated model estimators such as Mixture of Experts (MoE) to dynamically utilize specialized sub-models, attention mechanisms to focus on the most relevant input features, and autoencoders or variational autoencoders (VAEs) to perform dimensionality reduction or feature extraction for more informative representations.

[0072] The Predictor 430 provides the threat output 310, exemplified by outputs 310a and 310b described above. The predictor may perform tasks of regression, classification. In specific embodiments, the early inference subnetwork may include additional layers including pooling, dropout, CNN, concatenation, and other neural network processing mechanisms known in the art.

[0073] A simple implementation of this architecture may be obtained with a classification task having one flattening layer, one multilayer perceptron (MLP) layer and one Softmax layer.

[0074] Fig. 5 is a flow diagram of a process 500 for video surveillance performed by the video surveillance system, using geo-localized, crowd-sourced video streams. The diagram illustrates a continuous surveillance process whereby data is constantly received and processed.

[0075] At step 502, the system continuously acquires multiple video streams from various cameras, each associated with geo-location data indicating the position of the device capturing the footage.

[0076] In parallel with this ongoing data acquisition, at step 504, the system applies an inference plan to detect potential threats within these video streams. This inference plan specifies which Al agents should analyze the streams and which early inference options to utilize, allowing for efficient resource allocation. As the system processes the continuous video feeds, step 506 generates threat indications with estimated locations based on the geo-location data attached to the source streams.

[0077] When any threat is detected, even those with low probability, the process flowsto step 508, where the inference plan is revised based on the detected threat. This revision may include prioritizing certain Al agents or adjusting early inference thresholds to focus computational resources on areas of potential concern. The revised plan is then fed back to the Al agents, to iterate, at step 504, the reassessment of threats in the video stream, a process that operates in real-time as new video data arrives.

[0078] If the threat indications exceed a preset accuracy threshold during this ongoing analysis cycle, the process advances to step 510, where the system issues alerts containing geo-location data and threat details. These alerts provide security personnel with actionable information about the nature and location of the potential threat, enabling rapid response. The feedback loop between steps 504 and 508 operates continuously and simultaneously with the constant stream of incoming video data, ensuring the system remains adaptive to changing conditions and consistently refines its threat assessment capabilities without interruption.

[0079] The goal of this Al architecture is to accelerate inference time at the expense of the level of confidence and to allow for a gradual response to the task at hand. For that reason, a confidence level estimation mechanism is employed to output estimated uncertainty associated with each prediction.

[0080] The confidence level that is generated may be based on the accuracy of the EISN predictor that is measured during a testing period of training. More advanced schemes for estimating the confidence level may also be employed, such as MC Dropout, Variational Inference, or ensemble methods.

[0081] Protocols for training the EISNs could include the following steps.

[0082] Before training, the following steps may be performed:• Training of the core network or downloading of a pre-trained network and optional fine-tuning.• Definition of the EISN layer architecture, generic to all early inference subnetworks• Definition of the split points in the core network, to which to attach the EISNs. This may be at every or at selected layers of the core net.Definition of the number of parameters of each EISN Model Estimator, automatically or manually.Optionally, hyperparameter pre-tuning, e.g. adapting learning and decay rates of the EISNs at certain layers in consideration of the core architecture. If skipped, the same parameters are used for all subnetworks.• Loss definition, as an aggregation of the EISN losses, each similarly computed as the original loss of the core net. Alternatively, the core network’s accuracy could be set as the benchmark instead of or in lack of ground truth in the dataset, and the loss could be defined based upon the difference between the core and sub networks’ original loss functions.• Uncertainty estimation scheme selection & optional adaptations to the core network.• Setting of the number of epochs• Definition of a pruning scheme, i.e. when should EISNs be discarded, how much at least should be kept, at which layers. . .• Freezing of the core network layers

[0083] Subsequent training may include, until the desired accuracy is attained and the minimal number of EISNs has not been reached, repeating the following:• Training of the whole multi-head network• Automatic Pruning of the EISNs with individual accuracies smaller than a defined threshold, with poor convergence rates or with small discrepancies with EISNs attached to neighboring core layers.• Storage of each EISN head if its accuracy was improved.• Optional hyperparameter tuning.• Optionally, gradual unfreezing and fine-tuning of the core network layers, starting from the layers close to output. This step, when data is available, allows adapting the core network to the dataset and task at hand, and refining its layers to produce more informative data, for the EISNs to better “shortcut” the rest of the network.• A "derating factor" may be used, that is, a multiplication hook for decaying an LSTM / RNN weight, and thereby increasing an overall weight of a rule-based engine. A default of 1.0 deprecates this parameter, as no derating is applied at thissetting.• For Smoothing-Window-Size and the Rolling-Window-Size parameters, higher values typically increase the accuracy but also increase memory / power consumption, as more samples have to be locally stored during the processing.

[0084] It is to be understood that a shallow (e.g., rules-based) model and a deep (e.g., LSTM / RNN) model are orthogonal and can be developed by independent training. Other “deep” models that could also be trained using the data used to train the LSTM / RNN include “temporal convolutional networks” (TCNs). A “shallow” machine learning model that could be trained instead of the rules-based model could be the XGBoost model and its variants.

[0085] The system is typically based on a scalable cloud architecture for distributed surveillance. Typically, the system is configured to accept a wide range of live video stream standards. The monitoring system permits convenient monitoring, both of large geographical regions, as well as local areas.

[0086] The monitor screen may be a split screen, with multiple windows. For example, one window may show a dynamic, geographic map, with icons indicating positions of potential threats, for example, icons indicating positions of individuals that display threatening characteristics (e.g., weapons, types of clothing, actions, or identified by facial recognition). Other windows may include one or more live video streams, with options for manual and / or automatic switching between live streams, for example, based on threats associated with certain video streams. Additional features available on the screen may include: alerts and logs, an alert button and the possibility to provide details and / or alert codes to the system, easy transitions from map view to grid view by the use of the mouse, responsiveness tests for the monitor user may be added to ensure that the human user is always alert, and may be configured to contact a user and / or his supervisor.

[0087] Screens of the video upload interface, described above, may include options for manually adding textual or verbal details that may be processed by the Al model to assist in identifying threats. Such details may include altitude (in meters or just the floor number), and / or the camera's 3D orientation and similar positioning data.

[0088] The Admin interface may provide an overview of the immediate threats on the map, alerts, videos and reports. Admins can add, remove or configure users, regions or usergroups, devices, workforce, etc. Admins also may customize Al behaviors per device, and alerting protocols and settings.

[0089] Users of the system may include: security services at state and city levels, as well as campus-level authorities, such as private companies, schools, universities, hospitals, etc. Users may include: citizens, security guards, police, security companies, private companies, schools, universities, hospitals, etc. The system can be used for any type of threat detection and classification, that is, not only for security and defense, but also for health, safety, and environmental monitoring. Further fields may include transportation, traffic, retail, equipment and environmental monitoring, as well as automotive, robotics, and industrial automation.

[0090] A reverse proxy may be used to manage API requests and video stream routing. This improves system security by obfuscating backend services and ensuring only authorized requests are processed. It also enables load balancing, ensuring even distribution of traffic across backend servers.

[0091] The system may integrate with external messaging services to deliver alerts and notifications to monitors and administrators. Alerts are triggered by events detected in the video streams or system anomalies and can include text messages, emails, synthesized voice messages, and / or as messages in the Camera UI.

[0092] A database may store all system metadata, including user accounts, video metadata, event logs, and audit trails. This enables quick querying for both operational and reporting purposes. The database may also include all information about user roles, communities, Al workflows, and additional user-custom settings for each.

[0093] Once a user has logged into the interface, the backend servers are periodically polled by the interface, e.g. every second, for new alerts. Alternatively, the status update to the UIs could be triggered by an event-based mechanism on the backend servers (e.g. as soon as the early inference nFraetwork has a more pertinent alert). Both approaches (triggering alerts or from the backend) may be adopted in the same embodiment to best reduce the time-to-alert. For instance, the alert raising could be initiated by the trigger, while the alert details are constantly divulged to the user second by second from the periodic calls from the interface to the backend.

[0094] EXAMPLES

[0095] A first example of the present invention is a method EXAMPLES

[0096] A first example of a method for area surveillance includes several steps. The method includes acquiring, by multiple cameras, multiple respective video streams, with each video stream associated with a geo-location indicator of the given device. The method includes applying machine recognition to images of the multiple respective video streams to detect potential threats, wherein the machine recognition is implemented by an artificial intelligence (Al) portal queuing the streams for processing by multiple Al agents. Each Al agent is trained to implement machine recognition of a different aspect of a potential threat, and each Al agent is configured with alternative early inference options. The method includes processing each video stream by the multiple Al agents to generate one or more indications of potential threats, wherein each indication is associated with one or more cameras and an estimated location of each potential threat. The method includes processing the one or more indications by a large language model (LLM) configured to determine an inference plan for processing the multiple respective video streams by an alternate set of Al agents, including alternate early inference options. The method includes applying the inference plan to process the multiple respective video streams to generate a revised indication of a potential threat. The method includes responsively issuing an alert including the revised indication.

[0097] An example 2 of the method includes features of the first example, and issuing the alert includes displaying a map including an estimated location of a potential threat provided by the revised indication.

[0098] An example 3 of the method includes features of some or all of the above examples, and issuing the alert includes providing a message indicating a description of a potential threat, containing one or more of event details, timestamps, a source and type of the alert, and additional images, videos, GIFs.

[0099] An example 4 of the method includes features of some or all of the above examples, and issuing the alert includes providing a score of potential severity of a potential threat provided by the revised indication. The alert is issued after determining that the score exceeds a threshold for issuing alerts.

[0100] An example 5 of the method includes features of some or all of the above examples, and the score is represented as a colored set of partitions on a screen of amonitoring application. Detection of weapons, people, and / or aggressive actions are displayed as small icons at times corresponding in real time to threats.

[0101] An example 6 of the method includes features of some or all of the above examples, and generating the revised indication includes merging threat estimates derived from each of the multiple video streams.

[0102] An example 7 of the method includes features of some or all of the above examples, and further includes transmitting the video streams, together with the associated geo-location indicators for each video stream, to a central computing platform configured to apply the Al agents.

[0103] An example 8 of the method includes features of some or all of the above examples, and the cameras are mobile computing devices. Each camera includes one of the Al agents, dedicated to processing the video stream of the given camera.

[0104] An example 9 of the method includes features of some or all of the above examples, and further includes acquiring additional video streams from fixed cameras. Processing the multiple video streams includes processing the video streams acquired by the mobile devices together with the streams acquired from the fixed cameras.

[0105] An example 10 of the method includes features of some or all of the above examples, and acquiring the video streams further includes acquiring additional details of camera position and / or orientation entered manually by users of the mobile devices. The method includes processing the additional details together with the multiple video streams to generate the indication of the potential threat and the estimated location of the potential threat.

[0106] An example 11 describes a system for area surveillance that includes one or more processors with memory storage with instructions including an artificial intelligence (Al) threat analysis system and a large language model (LLM). When executed, these instructions implement steps described by one or more of the above method examples 1- 10.

Claims

CLAIMS1. A method for area surveillance, comprising: acquiring, by multiple cameras, multiple respective video streams, each video stream associated with a geo-location indicator of the given device; applying machine recognition to images of the multiple respective video streams to detect potential threats, wherein the machine recognition is implemented by an artificial intelligence (Al) portal queuing the streams for processing by multiple Al agents, each Al agent trained to implement machine recognition of a different aspect of a potential threat, each Al agent configured with alternative early inference options; processing each video stream by the multiple Al agents to generate one or more indications of potential threats, wherein each indication is associated with one or more cameras and an estimated location of each potential threat; processing the one or more indications by a large language model (LLM) configured to determine an inference plan for processing the multiple respective video streams by an alternate set of Al agents, including alternate early inference options; applying the inference plan to process the multiple respective video streams to generate a revised indication of a potential threat; and responsively issuing an alert including the revised indication.

2. The method of claim 1, wherein issuing the alert comprises displaying a map including an estimated location of a potential threat provided by the revised indication.

3. The method of claim 1, wherein issuing the alert comprises providing a message indicating a description of a potential threat, containing one or more of event details, timestamps, a source and type of the alert, and additional images, videos, GIFs.

4. The method of claim 1, wherein issuing the alert comprises providing a score of potential severity of a potential threat provided by the revised indication, wherein the alert is issued after determining that the score exceeds a threshold for issuing alerts.

5. The method of claim 4, wherein the score is represented as a colored set of partitions on a screen of a monitoring application and detection of weapons, people, and / or aggressive actions are displayed as small icons at times corresponding in real time to threats.

6. The method of claim 1, wherein generating the revised indication comprises merging threat estimates derived from each of the multiple video streams.

7. The method of claim 1, further comprising transmitting the video streams, together with the associated geo-location indicators for each video stream, to a central computing platform configured to apply the Al agents.

8. The method of claim 1, wherein the cameras are mobile computing devices, wherein each camera comprises one of the Al agents, dedicated to processing the video stream of the given camera.

9. The method of claim 8, further comprising acquiring additional video streams from fixed cameras and wherein processing the multiple video streams comprises processing the video streams acquired by the mobile devices together with the streams acquired from the fixed cameras.

10. The method of claim 8, wherein acquiring the video streams further comprises acquiring additional details of camera position and / or orientation entered manually by users of the mobile devices, and processing the additional details together with the multiple video streams to generate the indication of the potential threat and the estimated location of the potential threat.

11. A system for area surveillance, comprising: one or more processors comprising memory storage with instructions including an artificial intelligence (Al) threat analysis system and a large language model (LLM) that when executed implement steps of: acquiring, by multiple cameras, multiple respective video streams, each video stream associated with a geo-location indicator of the given device; applying machine recognition to images of the multiple respective video streams to detect potential threats, by queuing the streams for processing by multiple Al agents, each Al agent trained to implement machine recognition of a different aspect of a potential threat, each Al agent configured with alternative early inference options; processing each video stream by the multiple Al agents to generate one or more indications of potential threats, wherein each indication is associated with one or more cameras and an estimated location of each potential threat;processing the one or more indications by a large language model (LLM) configured to determine an inference plan for processing the multiple respective video streams by an alternate set of Al agents, including alternate early inference options; applying the inference plan to process the multiple respective video streams to generate a revised indication of a potential threat; and responsively issuing an alert including the revised indication.

Citation Information

Patent Citations

  • Security and protection system risk early warning method, device and equipment and storage medium

    CN117350835A

  • Security analysis agents

    US11916767B1

Cited By

  • System and method for detecting threat events and generating responses and metrics autonomously

    US20250342759A1

  • Adaptive data retrieval using agentic prompt processing units

    US20260236463A1