System and method for multi-camera person re-identification and tracking
Patent Information
- Application Number
- PCT/IN2026/050364
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-27
- Filing Date
- 2026-02-27
- Publication Date
- 2026-09-03
Smart Images

Figure IN2026050364_03092026_PF_FP_ABST
Abstract
Description
SYSTEM AND METHOD FOR MULTI-CAMERA PERSON RE-IDENTIFICATION AND TRACKINGFIELD OF INVENTION
[0001] The present invention relates to surveillance systems, more particularly to a multi-camera-based person re-identification and tracking system that ensures identity continuity across multiple camera views.BACKGROUND OF THE INVENTION
[0002] Surveillance systems play a crucial role in security, public safety, and operational monitoring across various domains, including retail, transportation, smart cities, and critical infrastructure. Traditional surveillance solutions often rely on singlecamera monitoring or facial recognition techniques, which face significant challenges in real-world scenarios. Factors such as occlusions, varying lighting conditions, changes in appearance, and camera blind spots can hinder accurate identification and tracking of individuals.
[0003] Several solutions exist to address these challenges. For instance, the Indian patent application 202141036938, titled "PERSON RE-IDENTIFICATION USING MACHINE LEARNING APPROACHES IN VIDEO SURVEILLANCE," discloses that person recognition is a long-standing challenge in computer vision, particularly for identifying the same person across different video segments. The complexity increases with variations in lighting, resolution, pose, and background, especially in multi-camera environments like airports and city streets. A machine learning-based approach is proposed to enhance person re-identification, along with a discussion of challenges and advancements in this field.
[0004] A United States patent application 2023058106, titled "Multi -camera system to perform movement pattern anomaly detection," describes a method for detecting movement anomalies using multi -camera video analysis. It generates a critical analysis matrix from multiple computer vision algorithms, assigns fusion values based on a fusion map, and triggers alerts when values exceed a set threshold.
[0005] Another Indian patent application, 202411017036, titled "CAMERAS WITH FACIAL RECOGNITION," describes an intelligent security camera system with facial recognition that records faces, gestures, and emotions. It improves accuracy through deep learning, generates security alerts based on recognized faces, and enhances access control for authorized individuals. A user-friendly interface and privacy protection ensure compliance with data regulations.
[0006] While the existing systems and methods introduce significant advancements in surveillance, they exhibit several limitations, highlighting the need for a more advanced and comprehensive solution. Many of these prior works primarily focus on singlecamera analysis, anomaly detection, or predefined threshold-based alerts, limiting their ability to track individuals seamlessly across multiple cameras. They often rely heavily on facial recognition, which becomes ineffective in cases of occlusion, poor lighting, or camera angles that obscure facial features.
[0007] Additionally, movement pattern anomaly detection, while useful for identifying suspicious behavior, does not ensure robust person re-identification across different environments, making it unsuitable for applications requiring continuous identity tracking. These systems also lack multi-sensor fusion, relying solely on video analysis without integrating additional biometric cues such as gait, posture, or body shape, which are crucial for reliable re-identification.
[0008] Furthermore, most existing methods do not incorporate adaptive learning mechanisms to refine tracking accuracy over time, leading to potential errors and inefficiencies. Given these limitations, there is a critical need for a multi-camera, AI-powered person re-identification system that ensures identity continuity, leverages biometric feature extraction beyond facial recognition, integrates multi-sensor fusion, and employs adaptive learning for enhanced performance. This need is addressed by the present invention.OBJECTIVE OF THE INVENTION
[0009] The primary objective of the present invention is to provide a multi-camera person tracking and re-identification system that enhances surveillance accuracy and reliability. This system strengthens security measures by detecting suspicious behaviorand automating person re-identification across multiple cameras, significantly improving loss prevention and operational safety.
[0010] Another objective is to utilize machine learning and adaptive algorithms to enhance identification accuracy over time. The system dynamically refines feature extraction and matching, ensuring improved recognition of individuals across different camera views and environments, even as conditions change.
[0011] Another objective is to integrate diverse sensor data for comprehensive situational awareness. By leveraging behavioral analysis, heatmaps, and movement tracking, the system supports industries such as manufacturing and healthcare in optimizing safety, workflow efficiency, and threat detection.
[0012] Another objective of the invention is to leverage cloud and edge computing for real-time processing and decision-making. This capability ensures faster response times, enhances resource efficiency, and minimizes latency in surveillance and security operations.SUMMARY OF THE INVENTION
[0013] The system in the present invention is designed to enhance security monitoring, behavioral analysis, and retail optimization by leveraging multi-camera person tracking and re-identification. It ensures accurate individual identification by extracting and storing biometric features, integrating real-time tracking algorithms, and utilizing machine learning for adaptive improvements. The system also incorporates multisensor fusion to enhance situational awareness and proactively detect suspicious behavior, making it valuable for both security and commercial applications.
[0014] In one aspect of the invention, the system utilizes biometric feature extraction, including facial recognition and gait analysis, to maintain identity continuity across multiple camera views. It employs real-time tracking and feature-matching algorithms to improve operational efficiency in security environments. A machine learning -based adaptive learning module continuously refines feature extraction and matching accuracy, ensuring reliability across varied conditions.
[0015] In another aspect of the invention, the system integrates a multi-sensor fusion capability that combines data from different sources, enhancing situational awareness and improving the detection of suspicious behavior. This fusion of data from cameras, motion sensors, and other surveillance tools allows for more accurate identification and tracking of individuals across multiple locations.
[0016] In yet another aspect of the invention, a behavioral pattern recognition module leverages machine learning algorithms to detect and classify anomalous activities, enabling proactive security responses. By analyzing movement patterns and behavioral anomalies, the system can identify potential threats in real-time, allowing security personnel to take preventive action.BRIEF DESCRIPTION OF THE DIAGRAM
[0017] The foregoing summary, as well as the following detailed description of the invention will be better understood when read in conjunction with the appended drawings. For the purpose of assisting in the explanation of the invention, there are shown in the drawings, embodiments which are presently preferred and considered illustrative. It should be understood, however, that the invention is not limited to the precise arrangements and instrumentalities shown therein. In the drawings:
[0018] Figure. 1 depicts the process flow of the system
[0019] Figure. 2 illustrates the architecture diagram of the system, depicting its key components and their interactions.
[0020] Figure. 3 illustrates the component diagram of the system, detailing its key functional modules and their interactions.REFERENCE NUMERALS100 depicts the process flow of the system200 illustrates the architecture diagram of the system, depicting its key components and their interactions.300 illustrates the component diagram of the system, detailing its key functional modules and their interactions.DETAILED DESCRIPTION OF THE INVENTION
[0021] The present invention will now be described more fully hereinafter. For the purposes of the following detailed description, it is to be understood that the invention may assume various alternative variations and step sequences, except where expressly specified to the contrary. Thus, before describing the present invention in detail, it is to be understood that this invention is not limited to particularly exemplified systems or embodiments that may of course, vary. The use of examples anywhere in this specification including examples of any terms discussed herein is illustrative only, and in no way limits the scope and meaning of the invention or of any exemplified term. Likewise, the invention is not limited to various embodiments given in this specification.
[0022] As used herein, the singular forms "a," "an," and "the" include plural reference unless the context clearly dictates otherwise. The term "and / or" means one or all of the listed elements or a combination of any two or more of the listed elements.
[0023] The terms “preferred” and “preferably” refer to embodiments of the invention that may afford certain benefits, under certain circumstances. However, other embodiments may also be preferred, under the same or other circumstances. Furthermore, the recitation of one or more preferred embodiments does not imply that other embodiments are not useful, and is not intended to exclude other embodiments from the scope of the invention.
[0024] When the term “about” is used in describing a value or an endpoint of a range, the disclosure should be understood to include both the specific value and endpoint referred to.
[0025] As used herein the terms “comprises”, “comprising”, “includes”, “including”, “containing”, “characterized by”, “having” or any other variation thereof, are intended to cover a non-exclusive inclusion
[0026] The integrated multi -camera person tracking and re-identification system marks a significant leap forward in surveillance technology, delivering unmatched accuracy and reliability in identifying individuals across diverse camera networks. Thisinnovation not only strengthens security measures but also drives operational efficiency across a wide range of applications, including loss prevention, by detecting and tracking suspicious behaviors. By automating person re-identification, our solution sets a new industry standard in surveillance, offering immense potential for both commercialization and further advancement.
[0027] In one embodiment, as illustrated in Figure 1, the system begins by capturing a live video stream from an IP or web camera. This video feed serves as the primary input for further analysis, where individuals and objects within the frame are detected and tracked. Once the video is acquired, an inference model is applied to detect persons and objects. A region of interest (ROI) is assigned to each detected entity, and a tracking algorithm is implemented to continuously monitor movements within the scene. This ensures that objects and individuals are properly identified throughout the analysis.
[0028] To analyze human activity, a behavior classification model is used to predict the actions performed by a person or object. Simultaneously, a classification model infers demographic details such as age and gender, which help in understanding crowd behavior and individual profiling. The system then determines whether a person is engaging in suspicious activity. If suspicious behavior is detected, an alert or notification is triggered, allowing for appropriate intervention. If no unusual behavior is observed, the system marks the journey as normal. Additionally, demographic data of identified individuals is stored for future reference and analysis.
[0029] In another embodiment, queue management and heatmap generation algorithms enhance surveillance operations. A queue management algorithm monitors congestion and ensures smooth movement. If necessary, queue-related alerts are triggered to manage overcrowding. Similarly, a heatmap generation algorithm tracks foot traffic, providing insights into movement patterns and high-traffic areas, which can be utilized for operational enhancements.
[0030] For loss prevention, a dedicated algorithm is implemented at self-checkout points to detect fraudulent activities. If any suspicious behavior is observed during the checkout process, an immediate notification is triggered to prevent potential losses. All collected data and generated insights are processed through a microservice-based rule engine, which applies predefined rules to analyze behavior, manage queues, generateheatmaps, and detect loss prevention scenarios, ensuring seamless automation and decision-making. The system integrates with multiple platforms to deliver notifications and insights. Mobile applications provide real-time alerts and analytics to users, while web applications offer a more detailed dashboard for monitoring detected activities. Additionally, notification services ensure that alerts are communicated promptly through push notifications, emails, or other channels.
[0031] As depicted in Figure 2, the system infrastructure consists of both edge and cloud pipelines. The system begins with cameras capturing video footage, which is then transmitted to a Network Video Recorder (NVR) or Digital Video Recorder (DVR). The recorded video is further sent through a router, allowing connectivity between local and cloud-based processing pipelines via the internet. The edge pipeline is responsible for processing the video stream locally. It integrates an object detection model to identify and track objects in real time. Additionally, a custom video classification model is applied to categorize and analyze detected objects based on predefined criteria. This pipeline ensures rapid processing without relying on cloud resources, reducing latency and improving efficiency.
[0032] Simultaneously, the cloud pipeline processes the video feed remotely. Similar to the edge pipeline, it incorporates an object detection model and a custom classification model for further analysis. The cloud pipeline allows for more extensive data processing, leveraging computational resources to handle complex tasks that may not be feasible at the edge. The processed data is transmitted through a firewall or gateway for security and further analysis. Various services enhance the data, including a metadata parsing and persistence service that structures and stores extracted information. A rule engine evaluates predefined conditions and triggers appropriate actions based on the analyzed video content.
[0033] To ensure real-time monitoring and response, an alert and notification service generates immediate alerts for relevant events detected in the video. A prediction and analytical service provides insights based on historical data, enabling proactive decision-making. Additionally, external hardware components may be integrated into the system for advanced processing or interaction with physical security measures. Finally, all processed data, including detected objects, classifications, alerts, and analytical insights, is stored in a database. This structured storage allows for historicalreference, compliance tracking, and further data analysis to improve system performance and accuracy.
[0034] As described in Figure 3, the system is designed to optimize both local and cloud-based processing. The system begins with a camera capturing video footage, which is then transmitted to a Network Video Recorder (NVR) or Digital Video Recorder (DVR). The NVR / DVR serves as an intermediary storage unit, allowing the video to be processed and streamed further. The recorded video feed is sent through a router, enabling network connectivity for subsequent processing. An edge server receives the video feed from the router. The edge server is responsible for local processing, performing initial computations such as object detection, video classification, and basic analytics. This reduces the need for constant cloud communication, minimizing latency and optimizing response times for real-time applications.
[0035] For more advanced processing, a cloud GPU server is integrated into the system. This server provides high-performance computing capabilities, enabling deep learning models to analyze the video stream with high accuracy. Tasks such as behavior recognition, anomaly detection, and complex visual processing are handled efficiently in the cloud. The processed data is then sent to a cloud server that manages microservices. These microservices handle different functionalities, such as event detection, alert generation, and data management. The cloud server ensures scalability, allowing multiple users and applications to access the processed information seamlessly. All relevant data, including analyzed video insights, object classifications, and detected events, are stored in a structured database. This database serves as a central repository for historical data, enabling analytics, reporting, and compliance tracking.
[0036] Finally, the processed information is made accessible through web and mobile applications. These applications provide real-time monitoring, notifications, and insights to users, allowing them to interact with the system from any location. The integration of cloud computing, edge processing, and mobile accessibility ensures a robust and efficient video analysis framework. This sophisticated surveillance system represents the future of intelligent security solutions, revolutionizing how environments such as retail stores, airports, public transportation hubs, and corporate facilities manage safety and operational efficiency. Through real-time monitoring, automated re-identification, and adaptive intelligence, this system not only minimizes human intervention but also ensures a proactive security approach that can dynamically evolve to address emerging threats and new surveillance challenges.
[0037] The integrated multi -camera person tracking and re-identification system offers several advantages over existing technologies, making it a highly efficient and scalable surveillance solution. By automating tracking and identification processes, the system reduces reliance on manual monitoring, lowering operational costs and eliminating the need for expensive legacy systems. The system delivers live insights, allowing security personnel to make swift, data-driven decisions based on behavioral analysis and situational awareness. It is adaptable to various industries, including retail, manufacturing, public security, and healthcare, ensuring a versatile application across multiple sectors. The system not only enhances security but also provides actionable analytics for optimizing workforce efficiency, compliance monitoring, and customer behavior tracking.
[0038] Advanced encryption and privacy controls ensure that biometric data and tracking information adhere to stringent data protection regulations. Automation minimizes inaccuracies in surveillance operations, ensuring a higher degree of precision in person tracking and suspicious behavior detection. The system is designed to work alongside current security infrastructures, making it a seamless and cost-effective upgrade rather than requiring complete replacements. The technology is suitable for diverse environments, including smart cities, airports, educational institutions, and public venues, enhancing security across various domains.
[0039] With reliable person tracking and re-identification, the system helps prevent unauthorized access and detect anomalies with a high degree of accuracy. Machine learning-powered behavioral pattern recognition ensures faster and more accurate detection of suspicious activities, allowing preemptive measures to be taken. The system dynamically responds to potential threats in real time, providing actionable intelligence that enhances decision-making. Automating identity tracking and reidentification significantly reduces manual workload while improving overall surveillance effectiveness. Whether deployed in small-scale setups or large public infrastructures, the system ensures seamless performance, making it future-ready and adaptable to evolving security challenges. This system represents a next-generationsurveillance framework, combining automation, advanced analytics, and Al-driven decision-making to improve both security and operational efficiency.
[0040] The present invention discloses a distributed edge-cloud video analytics system configured to perform coordinated multi-model artificial intelligence inference and centralized rule-based event orchestration. The system operates across geographically distributed deployments and is designed to provide real-time and near-real-time intelligent analysis of video streams while optimizing latency, scalability, and computational efficiency. The architecture combines edge processing infrastructure positioned proximate to video capture devices with scalable cloud computing infrastructure, thereby enabling dynamic distribution of inference workloads based on system conditions and performance requirements.
[0041] In one embodiment, the system comprises one or more video capture devices configured to continuously or selectively capture video streams of a monitored physical environment. The video capture devices may include fixed cameras, pan-tilt-zoom cameras, or other imaging sensors capable of generating digital video data. Video streams are optionally aggregated or buffered through a network video recorder or digital video recorder and are transmitted through a secure network gateway to an edge processing server. The edge processing server is positioned within the local network environment of the video capture devices to minimize transmission latency and bandwidth utilization.
[0042] The edge processing server is configured to decode incoming video streams and execute one or more artificial intelligence inference models. In a preferred embodiment, an object detection model based on a convolutional neural network architecture performs real-time detection of objects within individual video frames. Each frame undergoes preprocessing operations including resizing, normalization, and format conversion before being processed by the detection network. The detection model generates bounding box coordinates, object class labels, and associated confidence scores. Non-maximum suppression is applied to eliminate redundant detections. For each detected object, a unique entity identifier is generated to enable persistent tracking across successive frames.
[0043] Following detection, a multi-object tracking module associates detected objects across frames using spatial proximity, motion vectors, and appearance embeddings to maintain temporal continuity. The tracking module produces trajectory data representing movement paths, dwell durations, and directionality of each identified entity. Region-of-interest definitions are configurable within the monitored scene and may represent operational zones such as entry points, restricted areas, checkout counters, queue areas, or any user-defined spatial boundary. As tracked entities enter, exit, or remain within such regions, the system computes spatial-temporal metrics including dwell time, transition frequency, and movement direction. These metrics are structured as metadata and stored or transmitted for downstream processing.
[0044] The system further comprises a behavioral classification module configured to analyze temporal motion features derived from tracking data. Rather than operating directly on raw video frames, the behavioral model processes structured feature sequences comprising bounding box coordinates, velocity vectors, acceleration patterns, trajectory curvature, and region transition events over defined time windows. Temporal deep learning architectures, including recurrent neural networks, long shortterm memory networks, gated recurrent units, temporal convolutional networks, or transformer-based sequence models, may be utilized to classify behavioral patterns. The module outputs behavior labels and associated confidence values, which are linked to the corresponding entity identifiers and transmitted as structured metadata.
[0045] In some embodiments, the system includes a demographic inference module configured to estimate demographic attributes of detected persons without performing identity recognition. Cropped person regions are extracted from detection outputs and subjected to preprocessing including alignment, normalization, and illumination adjustment. A trained convolutional neural network predicts probabilistic demographic categories such as age group and gender. The demographic outputs are anonymized and aggregated to preserve privacy and are not associated with personally identifiable information. The inference may be performed at the edge for low latency or selectively offloaded to cloud infrastructure depending on computational load.
[0046] A queue management module may be configured to monitor defined queue regions and calculate queue metrics including queue length, waiting time, congestion levels, and throughput rates. The module utilizes object detection and tracking outputsto identify individuals within queue-designated regions and applies threshold-based or learned models to determine congestion events. Similarly, a heatmap and path analytics module aggregates historical trajectory coordinates to generate spatial density representations and dominant movement flows. Trajectory clustering, smoothing, and temporal filtering techniques may be applied to identify recurrent patterns and high-activity zones within the monitored environment.
[0047] In retail or transactional environments, the system may incorporate a loss prevention module configured to detect anomalous object handling behaviours in selfcheckout regions. The module correlates object detection, tracking, and behavioural signals to identify irregular movement sequences, such as item transfers that bypass seaming zones. Temporal association logic evaluates whether object trajectories correspond to expected scanning workflows. An anomaly score or classification label is generated and forwarded as structured metadata without requiring direct integration with point-of-sale systems.
[0048] As depicted in Figure 100, the system begins with capturing a video stream from one or more video capture devices. The captured video is processed at an edge processing server where object detection is performed to generate bounding box coordinates and classification data. Persistent entity identifiers are assigned to detected objects, and multi -object tracking is performed across successive frames.
[0049] Structured metadata comprising entity identifiers, timestamps, spatial coordinates, and classification labels is generated. The structured metadata is processed using a plurality of artificial intelligence modules including behavioral classification, demographic inference, queue analytics, heatmap analytics, and anomaly detection.
[0050] The processed metadata is transmitted to a centralized microservice-based rule engine which normalizes the metadata into a unified schema and evaluates configurable logical rules including spatial conditions, temporal window constraints, and multisignal correlations. Upon rule satisfaction, event records are generated and stored in a database.
[0051] As illustrated in Figure 200, the system architecture comprises video capture devices, a network video recorder or digital video recorder, a router or secure gateway,an edge processing server, a cloud GPU processing server, a cloud application server hosting microservices, and a database.
[0052] The edge processing server performs local inference including object detection, tracking, and metadata generation. The cloud processing server selectively executes computationally intensive artificial intelligence inference tasks offloaded from the edge processing server. Structured metadata is transmitted to the cloud application server where the microservice-based rule engine evaluates predefined logical rules and generates event records. All structured metadata and generated event records are stored in the database.
[0053] As shown in Figure 300, the system comprises functional modules including an object detection module, a multi-object tracking module, a metadata generation module, a plurality of artificial intelligence inference modules, a cloud processing module, a centralized rule engine, and a database module.
[0054] The object detection module performs frame-level detection and classification. The multi-object tracking module associates detected objects across frames and generates trajectory data. The metadata generation module structures detection and tracking outputs into a standardized metadata format.
[0055] The plurality of artificial intelligence modules process the structured metadata to generate higher-level inference metadata. The centralized rule engine receives the structured metadata, performs schema normalization, evaluates configurable spatial-temporal rules, and generates event records stored in the database module.
[0056] The invention further comprises a cloud processing server including scalable computing resources such as graphics processing units configured to execute computationally intensive inference models. The edge server may dynamically offload selected video frames, cropped regions, or inference tasks to the cloud server based on workload, network conditions, or predefined system policies. The cloud layer enables horizontal scalability across multiple locations and camera networks while preserving low-latency decision capability at the edge.
[0057] All inference modules generate structured metadata rather than transmitting raw video for centralized analysis. Structured metadata includes entity identifiers, spatialcoordinates, timestamps, classification labels, behavioral indicators, confidence scores, and derived metrics. This metadata is transmitted to a centralized backend comprising a microservice-based rule engine and associated database layer. The rule engine normalizes incoming metadata into a unified schema and performs rule-based evaluation across spatial and temporal dimensions.
[0058] Rules are configurable logical expressions that may include threshold comparisons, region-based conditions, time-window constraints, multi-signal correlations, and chained dependencies. The rule engine supports stateful evaluation across defined temporal intervals and enables rule prioritization and dynamic activation or deactivation without interrupting ongoing inference processes. Upon satisfaction of rule conditions, the engine generates event records containing event type, severity level, confidence value, associated entities, timestamps, and contextual information. Event records are stored within the centralized database and made available for notification and analytics services.
[0059] The database layer maintains persistent storage of structured metadata, event logs, historical analytics records, rule definitions, and configuration parameters. By storing structured metadata rather than continuous raw video, the system optimizes storage efficiency while enabling retrospective analytics and trend analysis.
[0060] User interaction with the system is facilitated through web-based and mobile client applications communicating with backend services via secure application programming interfaces. The user interface presents real-time alerts, event feeds, video overlays, analytical dashboards, heatmap visualizations, queue statistics, and historical reports. Administrative functions enable configuration of cameras, region-of-interest definitions, rule parameters, alert thresholds, and user access controls. Notifications generated by the rule engine may be delivered through in-application alerts, push notifications, electronic messages, or integration with external systems such as alarm controllers or enterprise monitoring platforms.
[0061] The artificial intelligence models revealed in this paper were trained, tested, and implemented in various real-time projects of operation, each of which was a live implementation of the distributed edge-cloud video analytics architecture and algorithmic flow as shown in the figures below. As indicated in the model results table,the training was done under between 100 and 200 epochs with overall training time between about 5,555 seconds and 28,821 seconds, depending on the data size and the complexity of deployment. The loss value of training shows consistent converging results on projects and bounding box loss is between about 0.4169 and 0.6704, classification loss is between about 0.2584 and 0.4069 and distribution focal loss is between about 0.8195 to 0.8851. Tire metrics of detection performance show strong real-time accuracy, with a precision value of about 0.7949 to 0.9039, an additional value of recall between about 0.7132 to 0.8350, a mean Average Precision at a 0.5 intersection unit (mAP50) is about 0.7287 to 0.8776 and a mean Average Precision at a 0.95 intersection unit (mAP50-95) is about 0.5587 to Validation losses are close to training losses, validation box loss values are between 0.5540 and 0.7616, validation classification loss values are between 0.3829 and 0.5084 and validation distribution focal loss values are between 0.8594 to 0.9342, which implies that it achieves good generalization in a live condition. The trained models were incorporated into the edge and cloud inference pipelines in the system architecture diagrams, in which live video streams were operated by object detections, object classifications, custom video inference, and region-of-interest based trackings. The results of such models were then used by the behavioral classification, demographic inference, queue management, heatmap and path analytics, and loss prevention modules, indicated in the overall system workflow diagram. Metadata created at every stage was sent to the centralized microservice-based rule engine, which allows producing real-time alerts, computations of analytics, and sending notifications to web and mobile applications. The presented findings thus affirm that the reported system has reported accurate, stable, and scalable inference performance, when implemented in a real-world environment.
[0062] Table 1 illustrates aggregate performance characteristics observed across multiple trained models and deployment scenarios within the disclosed distributed edge-cloud video analytics system.
[0063] Along with the object detection and tracking models, the presented system has an action recognition model, which is trained to categorize behavior patterns as normal and suspicious according to the temporal video aspects. The action recognition model was conditioned on a newly generated set of normal videos to include more variability and minimize the bias on the dataset whereas the suspicious dataset and test dataset were held constant so that the evaluation could be achieved in conditions that are close to production. The training procedure was implemented on 100 epochs and the model performance was measured by using a held-out test set of 100 normal video and 93 suspicious video which were not included in the training set. Recall values of both normal and suspicious classes were tracked during the epoch checkpoints during training. The recorded data show that, in general, normal-class recall was higher than suspicious-class recall during training with recall scores of normal behavior falling between the range of about 40 percent and 80 percent and the suspicious behavior between about 30 percent and 75 percent among epochs. The difference between normal and suspicious recall did not follow a monotonic increasing or decreasing trend1during training, which means that there is no learning behavior change and no overfitting and mode collapse. The model did not show progressive degradation or instability but was consistent between epochs. The results of prediction further exhibited equal separation between truly suspicious and falsely suspicious classes during the training checkpoints. At the epoch 13 and epoch 75, the model correctly classified 55 out of 100 normal videos and 58 out of 93 suspicious videos and 48 out of 100 normal videos and 63 out of 93 suspicious videos, respectively. These findings show that the action recognition model does not lose its discrimination ability with longer training periods. Tire developed action recognition model was incorporated into the distributed edge-cloud model inference system as shown in the system workflow diagrams, where the action recognition system outputs were fused with object detection, tracking, and region-of-interest analytics. The results of behavioral classification were sent to the centralized microservice-based rule engine where they were used via structured metadata to be utilized by rules to perform rule-based correlation, alert generation, and downstream analytics. The obtained findings confirm that the action recognition model can be applied in the real-time in the disclosed system architecture and that it can be used to infer the behaviors with high reliability under production.
[0064] Table 2 illustrates training and evaluation characteristics of the action recognition model across multiple epoch checkpoints under production-representative test conditions.
[0065] Through the integration of distributed edge inference, scalable cloud computation, structured metadata pipelines, and centralized rule-based orchestration, the disclosed invention provides a modular, scalable, and low-latency video analytics system capable of coordinated multi-model artificial intelligence execution and automated event intelligence delivery across diverse deployment environments.
[0066] The above-described embodiment of this patent is detailed; however, it is not limited to the mentioned embodiment. One skilled in the relevant art can make various changes within the scope of this patent, provided they do not deviate from its intended purpose
Claims
I / We Claim:
1. A multi-camera person re-identification and tracking system (200, 300), comprising:one or more video capture devices configured to capture video streams of a monitored environment;an edge processing server operatively coupled to the video capture devices and configured to:receive the video streams,execute an object detection model to detect persons within video frames, generate bounding box coordinates and classification labels corresponding to detected persons,assign persistent entity identifiers to detected persons, and perform multi-object tracking across successive frames to generate trajectory data including movement paths and dwell durations; one or more artificial intelligence modules operatively coupled to the edge processing server and configured to process detection and tracking outputs, the artificial intelligence modules comprising:a behavioral classification module configured to classify behavioral patterns based on temporal motion features derived from the trajectory data; anda demographic inference module configured to infer demographic attributes from cropped image regions corresponding to detected persons;a cloud processing server configured to selectively execute at least one artificial intelligence inference task offloaded from the edge processing server;a centralized microservice-based rule engine configured to:receive structured metadata generated from the object detection model, multi -object tracking, and artificial intelligence modules, normalize the structured metadata into a unified schema,evaluate predefined logical rules across spatial and temporal dimensions, andgenerate event records based on rule evaluation; anda database configured to store the structured metadata and the generated event records.
2. The system as claimed in claim 1, wherein the object detection model comprises a convolutional neural network-based detection architecture.
3. The system as claimed in claim 1, wherein the multi -object tracking associates detected persons across frames using spatial proximity metrics and appearance feature embeddings to maintain persistent entity identifiers.
4. The system as claimed in claim 1, wherein the behavioral classification module processes sequential temporal motion features derived from trajectory data using a sequence learning model.
5. The system as claimed in claim 1, wherein the demographic inference module processes cropped person image regions extracted from detected bounding boxes using a trained classification network.
6. The system as claimed in claim 1, wherein the rule engine performs statefill rule evaluation using configurable temporal windows, multi-signal correlations, and spatial context constraints.
7. A computer-implemented method for multi -camera person re-identification and tracking (100), comprising:capturing video streams from one or more video capture devices; processing the video streams at an edge processing server to:detect persons within video frames,generate bounding box coordinates and classification labels, assign persistent entity identifiers to detected persons, and perform multi-object tracking across successive frames to generate trajectory data;processing detection and tracking outputs using one or more artificial intelligence modules including:behavioral classification; anddemographic inference;selectively offloading at least one artificial intelligence inference task to a cloud processing server;transmitting structured metadata to a centralized microservice-based rule engine; normalizing the structured metadata into a unified schema;evaluating predefined logical rules across spatial and temporal dimensions; and generating event records based on rule evaluation.
8. The method of claim 7, wherein person detection is performed using a convolutional neural network-based detection model.
9. The method of claim 7, wherein behavioral classification is performed using temporal feature sequences derived from tracking data.
10. The method of claim 7, further comprising storing structured metadata and generated event records in a database.