Building cross-camera target tracking and identity verification method and system

By constructing a hierarchical topology and dynamic permissions for building cameras, and combining multimodal feature extraction and spatiotemporal correlation analysis, the problems of poor continuity of cross-camera target tracking and low accuracy of identity verification in building monitoring systems have been solved. This has enabled intelligent identification and early warning of abnormal behavior, thereby improving the intelligence level of building security systems.

CN121865104APending Publication Date: 2026-04-14ZHEJIANG HAISHI ELECTROMECHANICAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG HAISHI ELECTROMECHANICAL TECH CO LTD
Filing Date
2026-01-07
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing building monitoring systems suffer from poor continuity in cross-camera target tracking, low accuracy in identity verification, and a lack of integration of dynamic permissions and topology, making it difficult to provide intelligent early warnings for abnormal behavior.

Method used

A hierarchical topology of cameras is constructed, and dynamic permission is built by combining building space layout and personnel permission information to achieve cross-camera target collaborative tracking. Identity verification is carried out through multimodal feature extraction and spatiotemporal correlation analysis, and abnormal behavior is evaluated by combining motion trajectory and permission status.

Benefits of technology

It improves the efficiency of path planning and trajectory continuity for cross-camera target tracking, enhances the accuracy and robustness of identity recognition, and enables intelligent identification and hierarchical early warning of behaviors such as unauthorized access and abnormal loitering, thereby improving the intelligence level and security control precision of building security systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121865104A_ABST
    Figure CN121865104A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent buildings, in particular to a building cross-camera target tracking and identity verification method and system, and the method comprises the steps: obtaining camera network connection and space deployment information, constructing a layered topological structure, and achieving the hierarchical organization of monitoring resources; building space layout and personnel authority information are combined, a topological structure is fused for dynamic authority modeling, and target tracking authority is generated. Based on the hierarchical topology and the tracking authority, cross-camera cooperative tracking is realized, and a target motion track is acquired. And carrying out space-time correlation analysis and multi-modal feature extraction on the trajectory, combining with a preset identity feature library for dynamic verification, and outputting the target identity credibility. And finally, the motion trail, the identity credibility and the authority matching state are integrated to carry out abnormal behavior intelligent evaluation and generate linkage early warning information, so that the continuity of cross-camera tracking and the identity recognition accuracy are improved, and the intelligentization, refinement and active prevention and control capabilities of the building security system are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of smart building technology, and in particular to a method and system for cross-camera target tracking and identity verification in buildings. Background Technology

[0002] With the rapid development of smart buildings and intelligent security systems, cross-camera target tracking and identity recognition technology based on video surveillance has become a core means to ensure building security and improve management efficiency. Traditional monitoring systems typically treat each camera as an independent monitoring unit, lacking effective modeling of the overall topology of the camera network. This makes it difficult to achieve efficient and continuous trajectory tracking and collaborative analysis when targets move across areas. Furthermore, existing technologies rely heavily on single-modal features (such as facial recognition) for target identity verification, which is susceptible to factors such as changes in lighting, occlusion, and angular shifts, leading to decreased recognition accuracy. Additionally, they lack a dynamic evaluation mechanism for the correlation between target behavior trajectories and access control rules.

[0003] Furthermore, personnel permissions within a building are dynamic and hierarchical, with different personnel having varying access rights to different areas at different times. However, most current systems separate monitoring and tracking from access management, failing to integrate real-time personnel permission status into the target tracking and behavior analysis process. This results in the system's inability to intelligently determine whether target behavior is unauthorized or abnormal. For example, when an unauthorized person enters a sensitive area, or an authorized person appears in a restricted area outside of working hours, the system often struggles to issue timely and accurate warnings.

[0004] The above content is only used to help understand the technical solution of the present invention and does not represent an admission that the above content is prior art. Summary of the Invention

[0005] The main objective of this invention is to provide a method and system for cross-camera target tracking and identity verification in buildings, aiming to solve the technical problems of existing building monitoring systems, such as poor continuity of cross-camera target tracking, low accuracy of identity verification, and lack of intelligent early warning of abnormal behavior by integrating dynamic permissions and topology structure.

[0006] To achieve the above objectives, the present invention provides a method for cross-camera target tracking and identity verification in buildings, the method comprising:

[0007] Obtain network connection data and spatial deployment information of building cameras, construct a layered topology, and obtain the camera layered topology structure.

[0008] The spatial layout and personnel permission information of the building are obtained, and dynamic permission is constructed by combining the hierarchical topology of the camera to obtain target tracking permission;

[0009] Based on the hierarchical topology of the cameras and the target tracking permissions, cross-camera target collaborative tracking is performed to obtain the target motion trajectory;

[0010] Spatiotemporal correlation analysis and multimodal feature extraction are performed on the target's motion trajectory, and dynamic identity verification is performed against a preset identity feature database to obtain the target's identity credibility.

[0011] Based on the target's motion trajectory, the target's identity credibility, and the target's tracking permissions, abnormal behavior is evaluated to obtain linked early warning information.

[0012] Optionally, the step of acquiring network connection data and spatial deployment information of building cameras, and constructing a hierarchical topology to obtain a hierarchical camera topology includes:

[0013] Based on preset camera identifiers, the building cameras are discovered to obtain network connection data;

[0014] The building is divided into spatial zones, and spatial deployment information is obtained by combining the camera installation locations;

[0015] The network connection data is subjected to link quality assessment and hierarchical division to obtain the physical topology hierarchy;

[0016] Logical topology partitions are obtained by logically associating regions based on the spatial deployment information.

[0017] The physical topology hierarchy is merged with the logical topology partition to obtain an initial hierarchical topology;

[0018] The initial hierarchical topology is subjected to connectivity verification and path weight assignment to obtain the camera hierarchical topology structure.

[0019] Optionally, the step of obtaining the building's spatial layout and personnel access information, and combining it with the camera's hierarchical topology to dynamically construct access permissions and obtain target tracking permissions, includes:

[0020] The building is modeled in 3D and functionally partitioned to obtain the spatial layout;

[0021] Based on the spatial layout and the preset personnel management database, the permission levels of building personnel are divided to obtain the personnel permission information;

[0022] The spatial layout is subjected to access area identification and priority encoding to obtain the area security level;

[0023] Based on the personnel permission information, an access control list is generated for the area security level to obtain an initial permission matrix;

[0024] The camera coverage area and historical access records are extracted through the camera hierarchical topology, and permission mapping is performed in combination with the initial permission matrix to obtain the camera permission table;

[0025] Based on the area security level, the camera permission table and the initial permission matrix are dynamically adapted to obtain the target tracking permission.

[0026] Optionally, the step of performing cross-camera target collaborative tracking based on the camera hierarchical topology and the target tracking permission to obtain the target motion trajectory includes:

[0027] Receive real-time video streams from each camera in the hierarchical camera topology, selectively access and decode them according to the target tracking permissions, and obtain authorized video data;

[0028] Multi-scale target detection and feature extraction are performed on the authorized video data to obtain multi-modal target features;

[0029] Based on the physical topology hierarchy and logical topology partitioning in the camera hierarchical topology structure, a cross-camera target matching cost matrix is ​​constructed;

[0030] The matching cost matrix is ​​weighted based on the target tracking permission, and a hierarchical progressive association algorithm is used for target matching and trajectory stitching to obtain the initial motion trajectory.

[0031] The initial motion trajectory is subjected to spatiotemporal continuity verification and outlier correction to obtain the target motion trajectory.

[0032] Optionally, the step of performing spatiotemporal correlation analysis and multimodal feature extraction on the target motion trajectory, and dynamically verifying its identity against a preset identity feature database to obtain the target's identity credibility, includes:

[0033] Key nodes of the target motion trajectory are sampled to extract the trajectory's spatiotemporal features;

[0034] At the key nodes, multimodal biometrics of the target are extracted from the corresponding camera video. These multimodal biometrics include face, body shape, and clothing.

[0035] Based on the spatiotemporal characteristics of the trajectory, the multimodal biometrics are timestamped and their quality is assessed to obtain a feature set to be verified.

[0036] The set of features to be verified is compared with a preset identity feature database by hierarchical retrieval and similarity calculation to obtain preliminary matching results;

[0037] The confidence level of the target identity is obtained by fusing the confidence level of the preliminary matching results based on the historical behavior patterns of the target's motion trajectory.

[0038] Optionally, the step of evaluating abnormal behavior based on the target's motion trajectory, the target's identity credibility, and the target's tracking permissions to obtain linked early warning information includes:

[0039] The path deviation between the target's motion trajectory and the spatial layout is calculated to obtain the trajectory anomaly value.

[0040] Based on the target identity credibility and the target tracking permissions, an access matching degree analysis is performed to obtain permission anomaly values;

[0041] The trajectory anomalies and permission anomalies are weighted and fused to obtain a comprehensive risk score;

[0042] Based on the comparison between the comprehensive risk score and the preset risk threshold, a graded early warning instruction is generated;

[0043] Based on the tiered early warning instructions, cameras in the corresponding areas are linked for focused monitoring, and the linked early warning information is pushed to the management terminal.

[0044] Optionally, the step of linking cameras in the corresponding area for focused monitoring according to the tiered early warning instruction and pushing the linked early warning information to the management terminal includes:

[0045] The tiered early warning instructions are parsed to determine the early warning level, the affected area, and the associated cameras;

[0046] Based on the warning level, resource scheduling is performed on the associated cameras, including increasing the bit rate, adjusting the focal length, and enabling intelligent analysis;

[0047] Predict the target's next position based on the target's movement trajectory, and activate the corresponding area's camera deployment in advance;

[0048] The system integrates information such as warning level, target identity credibility, current location, historical trajectory fragments, and related camera footage.

[0049] Based on the preset terminal permission list, the integrated linkage warning information is pushed to the corresponding level of management terminal.

[0050] Furthermore, to achieve the above objectives, the present invention also provides a building cross-camera target tracking and identity verification system, the system comprising:

[0051] The topology building module is used to acquire network connection data and spatial deployment information of building cameras, perform hierarchical topology building, and obtain the hierarchical topology structure of the cameras;

[0052] The permission modeling module is used to obtain the spatial layout and personnel permission information of the building, and to dynamically construct permissions in combination with the hierarchical topology of the camera to obtain target tracking permissions;

[0053] The collaborative tracking module is used to perform cross-camera target collaborative tracking based on the camera hierarchical topology and the target tracking permission to obtain the target motion trajectory;

[0054] The identity verification module is used to perform spatiotemporal correlation analysis and multimodal feature extraction on the target's motion trajectory, and to perform dynamic identity verification with a preset identity feature library to obtain the target's identity credibility.

[0055] The early warning linkage module is used to evaluate abnormal behavior based on the target's movement trajectory, the target's identity credibility, and the target's tracking permissions, and obtain linkage early warning information.

[0056] In addition, to achieve the above objectives, the present invention also provides a building cross-camera target tracking and identity verification device, the device comprising: a memory, a processor, and a building cross-camera target tracking and identity verification program stored in the memory and executable on the processor, the building cross-camera target tracking and identity verification program being configured to implement the steps of the building cross-camera target tracking and identity verification method as described above.

[0057] In addition, to achieve the above objectives, the present invention also provides a medium storing a building cross-camera target tracking and identity verification program, wherein when the building cross-camera target tracking and identity verification program is executed by a processor, it implements the steps of the building cross-camera target tracking and identity verification method as described above.

[0058] This invention provides a method for cross-camera target tracking and identity verification in buildings. The method constructs a hierarchical camera topology that integrates network connectivity and spatial deployment, achieving systematic organization of monitoring resources and improving the path planning efficiency and trajectory continuity of cross-camera target collaborative tracking. It generates dynamic target tracking permissions by combining building spatial layout and personnel access information, enabling the monitoring process to have context awareness and access control capabilities, effectively constraining illegal tracking behavior and enhancing privacy compliance. During tracking, multimodal feature extraction and spatiotemporal correlation analysis are introduced, and dynamic verification is performed against an identity feature database, significantly improving the accuracy and robustness of identity recognition in complex scenarios. Finally, abnormal behavior assessment is conducted by comprehensively considering motion trajectory, identity credibility, and permission matching status, achieving intelligent identification and graded early warning of risky behaviors such as unauthorized access, abnormal loitering, and trajectory deviation. A linkage response mechanism enhances the system's proactive defense capabilities. The overall solution achieves closed-loop management of topology-guided tracking, permission-constrained behavior, multimodal identity verification, and intelligent risk assessment, significantly improving the intelligence level, security control accuracy, and emergency response efficiency of building security systems. Attached Figure Description

[0059] Figure 1 This is a flowchart illustrating an embodiment of the building cross-camera target tracking and identity verification method of the present invention.

[0060] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0061] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0062] Reference Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the building cross-camera target tracking and identity verification method of the present invention, which presents an embodiment of the building cross-camera target tracking and identity verification method of the present invention.

[0063] In one embodiment, the building cross-camera target tracking and identity verification method includes:

[0064] Step S100: Obtain network connection data and spatial deployment information of building cameras, construct a layered topology, and obtain the camera layered topology structure.

[0065] The network connection data of the building cameras can be data describing the communication link attributes between cameras, including IP address, bandwidth, latency, topological adjacency, etc., which can be used as the logical connection basis for constructing the hierarchical topology of the cameras. Spatial deployment information can be the installation location, orientation, field of view, and coverage area information of the cameras in the building's physical space, which can be used as the spatial geometric basis for constructing the hierarchical topology of the cameras. The hierarchical topology of the cameras can be a graph structure model with hierarchical organization characteristics constructed by integrating the building camera network and physical spatial deployment locations. It can be used to provide structured path guidance for cross-camera target tracking, improving trajectory stitching efficiency and continuity. In this embodiment, the operating principle of the hierarchical topology of the cameras can be explained in context: graph modeling is performed based on the network reachability and spatial adjacency between cameras, and the layers are divided into core layer, convergence layer, edge layer, etc., according to function or coverage. Furthermore, the hierarchical topology of the cameras can: collaborate with cross-camera target collaborative tracking, providing candidate camera jump paths; and collaborate with target tracking permissions, serving as a spatial carrier for permission mapping. For example, the hierarchical topology of the cameras can include, but is not limited to, one or more of the following: logical topology layer, physical deployment layer, and field of view association layer.

[0066] Obtaining network connection data and spatial deployment information from building cameras can be achieved by reading network parameters and installation metadata from the building equipment management system or camera configuration interface, thus providing raw input for hierarchical topology construction. Hierarchical topology construction, resulting in a hierarchical camera topology structure, can be achieved by fusing network connection data and spatial deployment information to construct a graph structure with hierarchical labels. In an exemplary embodiment, hierarchical topology construction can be achieved by clustering camera nodes based on a graph neural network and then injecting spatial adjacency constraints; or by first dividing spatial clusters according to physical floors, and then dividing logical levels within each cluster according to network centrality indicators, thereby enabling the systematic organization of monitoring resources and supporting efficient path planning.

[0067] Step S200: Obtain the building's spatial layout and personnel permission information, and combine the camera's layered topology to construct dynamic permissions, thereby obtaining target tracking permissions.

[0068] The building's spatial layout can be the building's internal structural information, including spatial units such as floors, rooms, passageways, and sensitive areas, and their topological relationships. This can be used to provide physical area semantics for dynamic permission construction. Personnel permission information can be data describing a person's access authorization status to a specific area at a specific time, including roles, validity periods, and area whitelists. This can be used as input for the identity policy in dynamic permission construction. Target tracking permissions can be access control policies dynamically generated based on the person's real-time identity status and building area access rules. These policies constrain whether a target can be tracked in a specific camera or area, enabling the tracking process to have context-aware capabilities, preventing tracking of unauthorized areas or personnel, and enhancing compliance. In one specific embodiment, target tracking permissions can be described in context as follows: mapping personnel permission information and the building's spatial layout to nodes or edges in the camera's hierarchical topology to form dynamically executable tracking permission rules. Furthermore, target tracking permissions can be: coordinated with cross-camera target collaborative tracking to restrict the tracking path to only within authorized areas; and coordinated with abnormal behavior assessment as a basis for judging permission matching status. For example, target tracking permissions may include, but are not limited to, one or more of the following: area access permission, time-limited permission, role-bound permission, etc.

[0069] Obtaining building spatial layout and personnel access information can be achieved by synchronizing building floor plans and personnel authorization policies from a building BIM system or access control platform, thus providing contextual data for dynamic access control. Combining camera hierarchical topology with dynamic access control to obtain target tracking permissions can involve mapping personnel permissions to camera topology nodes, generating an executable set of tracking permission rules. In an exemplary embodiment, this operation can be achieved by converting the permission area into a topology subgraph, allowing target tracking only within the subgraph; or by attaching an Access Control List (ACL) to each camera node to filter unauthorized targets in real time, thereby enabling context-aware and access-control capabilities for tracking.

[0070] Step S300: Based on the camera hierarchical topology and target tracking permissions, perform cross-camera target collaborative tracking to obtain the target motion trajectory.

[0071] The target motion trajectory can be a spatiotemporal path sequence formed by the continuous movement of the target across multiple camera fields of view within a building. It can be used to characterize the target's movement behavior in physical space, providing basic data for abnormal behavior identification. In one specific embodiment, the acquisition method of the target motion trajectory can be described in context: through a cross-camera target collaborative tracking algorithm, guided by a hierarchical camera topology, the detection boxes of the target in each camera are associated and interpolated to generate the trajectory. For example, the target motion trajectory can include, but is not limited to, one or more of the following: short-term local trajectory, cross-regional global trajectory, and dwell point sequence.

[0072] Cross-camera collaborative target tracking, based on a hierarchical camera topology and target tracking permissions, yields the target's motion trajectory. This can be achieved by propagating target features along authorized paths under topology guidance and correlating cross-camera detection results. Furthermore, this operation can employ graph traversal algorithms (such as A*) to search for the most likely transition path within the topology, performing feature matching only at authorized nodes; or by using a graph attention mechanism to dynamically weight candidate cameras and suppress tracking branches in unauthorized areas. This improves trajectory continuity and tracking efficiency while avoiding tracking in unauthorized areas.

[0073] Step S400: Perform spatiotemporal correlation analysis and multimodal feature extraction on the target's motion trajectory, and perform dynamic identity verification with the preset identity feature database to obtain the target's identity credibility.

[0074] Multimodal features can be a collection of various heterogeneous perceptual features extracted from the target video stream to enhance the robustness of identity recognition. In one specific embodiment, the acquisition method of multimodal features can be described in context: semantic feature vectors of visual modalities (such as face, gait, clothing, and body shape) are extracted separately through parallel or cascaded deep neural networks. For example, multimodal features can include, but are not limited to, one or more of face representation features, gait dynamic features, and clothing texture features. The preset identity feature library can be a structured database that stores known multimodal identity features of individuals (such as face embedding vectors, gait templates, and clothing color distribution) and their corresponding permission information. It can be used as a comparison benchmark for identity verification and support dynamic identity credibility calculation. In one exemplary embodiment, the acquisition method of the preset identity feature library can be described in context: multimodal biometric features of individuals are collected through the registration process and encoded by a feature extraction model before being stored, supporting dynamic updates and permission binding. For example, the preset identity feature library can include, but is not limited to, one or more of face feature sub-libraries, gait feature sub-libraries, and appearance feature sub-libraries. The target identity credibility can be a value calculated based on the comparison results of multimodal features and identity feature database. It can be used to quantify identity verification results and for the evaluation of abnormal behavior.

[0075] Spatiotemporal correlation analysis and multimodal feature extraction of the target's motion trajectory can be performed by extracting multimodal features from each frame of the target image along the trajectory and establishing spatiotemporal consistency constraints. Further, this operation can be achieved by using a 3D convolutional network to jointly model the spatiotemporal appearance changes of the trajectory segments; or by extracting independent features from each camera and then fusing them using a temporal alignment module (such as DTW), thereby enhancing feature robustness and overcoming the limitations of a single viewpoint. Dynamic identity verification with a pre-set identity feature library to obtain the target's identity credibility can be achieved by comparing the multimodal features with records in the library for similarity and then weighted and fusing the scores of each modality. In a specific embodiment, this operation can be achieved by employing an adaptive weight fusion strategy to dynamically adjust modal weights based on the current environmental quality (such as illumination and occlusion); or by introducing an online learning mechanism to temporarily cache and incrementally compare frequently occurring but unregistered targets, thereby improving the accuracy and robustness of identity recognition in complex scenarios.

[0076] Step S500: Based on the target's motion trajectory, the target's identity credibility, and the target's tracking permissions, an abnormal behavior assessment is performed to obtain a linkage warning information.

[0077] The abnormal behavior assessment module can be a logical unit that performs multi-dimensional risk judgment based on the target's movement trajectory, identity credibility, and tracking permission status. It can be used to output tiered early warning information and trigger a linkage response mechanism. In one specific embodiment, the abnormal behavior assessment module's acquisition method can be described in context: based on a rule engine or lightweight classification model, it matches or scores patterns such as unauthorized access, abnormal loitering, and trajectory deviation. For example, the abnormal behavior assessment module can include, but is not limited to, one or more of the following: an unauthorized access detector, a loitering behavior analyzer, and a trajectory deviation discriminator. The linkage early warning information can be a structured alarm message containing the abnormal behavior type, risk level, target location, and suggested response measures. It can be used to drive the security system to perform linkage operations such as access control locking, audible and visual alarms, and manual verification.

[0078] Anomaly assessment based on target movement trajectory, target identity credibility, and target tracking permissions yields linked early warning information. This can be achieved by using these three factors as input variables and determining the risk level of the behavior through preset rules or a lightweight model. Furthermore, this operation can be implemented by defining rules such as triggering a level-one warning if the identity credibility is below a threshold and the target enters a high-privilege area; or by training a binary classification model to predict the probability of anomalies using trajectory deviation, number of permission violations, and identity uncertainty as features. This allows for intelligent identification and tiered early warning of behaviors such as unauthorized access, abnormal loitering, and trajectory deviation.

[0079] Taking nighttime security monitoring of an office building as an example, the building cross-camera target tracking and identity verification method in this embodiment can be as follows: Employee A's normal working hours are 9:00-18:00, and their access is limited to the 3rd floor office area. At 22:00, the system detects a target entering the 3rd floor corridor. The camera layered topology guides the tracking path from the lobby camera to the 3rd floor east side camera. The dynamic permission construction module determines that the 3rd floor is a restricted area during this period, but because the target's identity initially matches employee A, limited tracking is still allowed. The multimodal feature extraction module fuses the blurred face and gait features, compares them with the preset identity feature database, and obtains an identity credibility of 0.72. The abnormal behavior assessment module combines "non-working hours + restricted area + medium credibility" to determine "suspicious loitering," generates a secondary linkage warning message, triggers local audio and visual prompts, and notifies the on-duty security personnel for verification.

[0080] This embodiment constructs a hierarchical topology for cameras by integrating network connection data and spatial deployment information, providing hierarchical path guidance for cross-camera target tracking and improving trajectory continuity and planning efficiency. It maps building spatial layout and dynamic personnel permission information to the topology, generating target tracking permissions and enabling context awareness and access control during the tracking process, constraining unauthorized area tracking and enhancing compliance. During tracking, spatiotemporal correlation analysis and multimodal feature extraction are performed simultaneously, and dynamically compared with a preset identity feature database. Modal complementarity overcomes the failure of single features under conditions such as occlusion and lighting changes, improving the robustness of identity verification. Finally, by combining target movement trajectory, identity credibility, and real-time permission status, multidimensional abnormal behavior assessment is conducted to identify risky behaviors such as unauthorized access, abnormal loitering, and trajectory deviation, generating tiered warnings. This achieves a closed-loop management mechanism throughout the entire process, encompassing topology guidance, permission constraints, multimodal verification, and intelligent judgment, thereby improving the intelligence level, security control accuracy, and emergency response efficiency of building security at the system level.

[0081] In one embodiment, network connection data and spatial deployment information of building cameras are acquired, and a hierarchical topology is constructed to obtain a hierarchical topology structure for the cameras, including:

[0082] Based on preset camera identifiers, the system discovers building cameras and obtains network connection data.

[0083] The preset camera identifier can be a pre-configured identifier used to uniquely identify each camera within a building, such as a MAC address, device ID, or IP range prefix. This identifier can support automated device discovery and ensure the integrity and accuracy of network connection data collection. In this embodiment, the preset camera identifier serves as the identification basis for the device discovery process, enabling the system to accurately match the target device. For example, the preset camera identifier can adopt the device ID format from standard network video protocols.

[0084] Device discovery can be the process of actively detecting and identifying all online cameras within a building based on preset identifiers. It can be used to automatically collect network connection data, avoiding errors from manual data entry. Furthermore, device discovery can obtain camera network parameters through methods such as ARP scanning, standard network video protocol queries, or device registry matching. In an exemplary embodiment, device discovery can initiate probe requests in conjunction with standard network video protocols to collect connection parameters such as the responding device's IP address, port, and protocol type, thereby achieving the technical effect of automated and complete collection of camera network connection data.

[0085] The building is divided into spatial zones, and spatial deployment information is obtained by combining the camera installation locations.

[0086] The spatial area division can be a semantic segmentation of physical space based on building function or security level, providing a semantic basis for spatial deployment information construction. In one specific embodiment, the spatial area division can be based on functional area labeling on building BIM or CAD drawings. For example, the spatial area division can include functional areas such as office areas, computer rooms, and corridors. The camera installation location can be the camera's three-dimensional installation coordinates in the building coordinate system, along with its orientation and field of view parameters, which can be combined with the spatial area division to determine the logical zone to which the camera belongs. Furthermore, the camera installation location can be obtained through importing from the building information model or through on-site surveying.

[0087] The building is divided into spatial zones, and spatial deployment information is obtained by combining this with the camera installation locations. This can be achieved by importing building BIM or CAD drawings, marking functional areas, and mapping the camera coordinates to the corresponding areas. Furthermore, this operation can be accomplished by automatically parsing building floor plans and dividing zones based on a semantic segmentation model, or by having administrators manually define zone boundaries and bind cameras in a visual interface. This achieves the technical effect of constructing a camera deployment description with spatial semantics.

[0088] The network connection data is subjected to link quality assessment and hierarchical division to obtain the physical topology hierarchy.

[0089] Link quality assessment can be a quantitative analysis of the performance of network communication links between cameras, providing an objective basis for physical topology hierarchical division. In this embodiment, link quality assessment can collect indicators such as latency, bandwidth, and jitter using ping, iperf, or SNMP, and then assign a weighted score. Hierarchical division can be the operation of classifying cameras into different communication layers based on the link quality assessment results, forming a physical topology hierarchical structure.

[0090] The physical topology hierarchy can be a communication capability hierarchy structure based on the network link quality assessment results between cameras. It reflects the logical location and data transmission performance of devices in the network and can be used to provide communication efficiency-oriented path priority for cross-camera data scheduling and target handover. Furthermore, the physical topology hierarchy can quantify and score network connection data (such as latency, bandwidth, and packet loss rate) and divide cameras into core, aggregation, and edge layers based on centrality or clustering algorithms. In an exemplary embodiment, the physical topology hierarchy can collaborate with logical topology partitioning to jointly form the initial hierarchical topology during the fusion phase; it can also collaborate with path weight assignment, with its hierarchical attributes serving as one of the weight calculation factors. For example, the physical topology hierarchy can include, but is not limited to, one or more of the following: core communication layer, regional aggregation layer, and terminal access layer.

[0091] Link quality assessment and hierarchical classification of network connection data yields a physical topology hierarchy. This can be achieved by quantifying link metrics and assigning hierarchical labels based on thresholds or clustering results. Furthermore, this operation can be implemented by automatically generating a three-layer structure using K-means clustering of link quality vectors, or by rigidly dividing the hierarchy by setting fixed thresholds (e.g., latency <50ms as the core layer), thereby achieving the technical effect of forming a physical topology hierarchy that reflects differences in communication performance.

[0092] Logical topology partitions are obtained by logically associating regions based on spatial deployment information.

[0093] Logical region association can be a process of establishing a mapping relationship between a camera and its associated spatial region and deriving the connectivity between regions, which can be used to generate logical topology partitions. Logical topology partitions can be functional region subgraphs constructed based on building spatial area divisions and camera installation locations, reflecting physical adjacency and field-of-view coverage relationships. They can be used to provide spatial semantic guidance for tracking, ensuring that targets are tracked sequentially along physically continuous paths. In a specific embodiment, logical topology partitioning can be achieved by aggregating cameras within the same region into a subgraph and establishing cross-regional connection edges between adjacent regions. Furthermore, logical topology partitioning can collaborate with physical topology hierarchies, with the two merging to form an initial hierarchical topology; simultaneously, it can collaborate with cross-camera target collaborative tracking, providing reasonable constraints for region jumps. For example, logical topology partitions can include, but are not limited to, one or more of the following: floor functional areas, security-sensitive areas, and public passageways.

[0094] Logical topological partitioning is achieved by associating logical regions based on spatial deployment information. This can be done by aggregating cameras within the same region into a subgraph and establishing cross-regional connection edges between adjacent regions. Furthermore, this operation can be implemented by automatically generating cross-regional edges based on a region adjacency matrix or by dynamically establishing logical connections based on camera field-of-view overlap detection results, thereby achieving the technical effect of constructing logical topological partitions with spatial continuity.

[0095] The physical topology hierarchy and logical topology partitioning are merged to obtain the initial hierarchical topology.

[0096] The initial hierarchical topology can be a camera topology structure that has not yet been verified or optimized, formed by the preliminary fusion of physical topology levels and logical topology partitions. It can serve as an intermediate product for generating the final camera hierarchical topology structure, carrying dual-dimensional structural information. In this embodiment, the initial hierarchical topology can integrate physical level labels and logical partition labels into a unified graph model through graph fusion operations (such as node attribute overlay and edge set union). For example, the initial hierarchical topology can include, but is not limited to, one or more of the following: unverified fused graph, dual-label topology sketch, and initial structural model.

[0097] By fusing physical topology hierarchy with logical topology partitioning to obtain an initial hierarchical topology, each camera node can be labeled with both physical hierarchy and logical partition in a unified graph structure. Furthermore, this operation can be achieved by constructing a dual-attribute graph (nodes containing (physical layer, logical partition) tuples), or by constructing physical and logical graphs separately and aligning edges using node IDs, thus generating an initial topology that integrates both communication and spatial semantics.

[0098] Connectivity verification and path weight assignment are performed on the initial hierarchical topology to obtain the camera hierarchical topology structure.

[0099] Connectivity verification can check whether there is valid communication or view relay path between any two camera nodes in the initial hierarchical topology, ensuring the feasibility of the final topology in actual tracking. In this embodiment, connectivity verification can use a graph traversal algorithm (such as BFS) to detect the connectivity of the entire graph and mark isolated subgraphs or breakpoints. Path weight assignment can assign numerical weights reflecting the tracking quality to each edge in the topology (i.e., the transfer path between cameras), supporting optimal path selection in subsequent tracking processes. Furthermore, path weight assignment can calculate composite weights by considering factors such as bandwidth, view overlap, occlusion probability, and illumination consistency.

[0100] The initial hierarchical topology is validated for connectivity and path weights are assigned to obtain the camera hierarchical topology. This can be achieved by first validating graph connectivity, then calculating and updating the composite weight for each edge. Furthermore, this operation can be implemented by inserting virtual relay nodes into disconnected regions to ensure global reachability, or by using reinforcement learning strategies to adjust path weights online to adapt to real-time environmental changes. This achieves the technical goal of outputting a final camera hierarchical topology that is both engineering-feasible and tracking-oriented.

[0101] Taking cross-regional tracking in a large commercial complex as an example, the building cross-camera target tracking and identity verification method in this embodiment can be as follows: The system first discovers 200 cameras through preset standard network video protocol device IDs and obtains their IP and RTSP ports; at the same time, the mall is divided into three main stores (A / B / C) and a public corridor area, and the area to which each camera belongs is determined by combining the installation coordinates. Link evaluation shows that the underground equipment room cameras have high bandwidth and low latency, and are classified as the core layer; the cameras in the shops are classified as the edge layer. Logically, a strong connection is established between area A and the corridor due to the overlap of the field of view. After fusion, an initial topology is formed. After connectivity verification, it is found that a dead-angle camera in area C is isolated, and the system automatically suggests adding a relay. The final path weight is assigned based on bandwidth (0.4), view overlap (0.4), and occlusion rate (-0.2). When the target enters the corridor from area A, the system prioritizes the high-weight path and switches to the high-definition camera in the corridor to ensure seamless trajectory connection.

[0102] This embodiment achieves the technical effect of generating a hierarchical camera topology structure that combines network communication characteristics and spatial semantics by characterizing communication performance differences from a physical dimension, modeling spatial adjacency relationships from a logical dimension, fusing dual semantics to form an initial structural state, verifying connectivity to ensure engineering feasibility, and introducing multi-factor composite weights to optimize path selection. This provides optimal path guidance for cross-camera target tracking that combines communication efficiency and spatial continuity, effectively solving the tracking breakage and resource silo problems caused by the lack of global topology modeling in traditional systems, significantly improving trajectory continuity and collaborative tracking efficiency, and providing reliable underlying structural support for permission constraints and intelligent judgment.

[0103] In one embodiment, the spatial layout and personnel access information of the building are obtained, and dynamic access control is constructed by combining the hierarchical topology of the cameras to obtain target tracking permissions, including:

[0104] The building is modeled in 3D and its functions are divided to obtain the spatial layout.

[0105] This operation can reconstruct a 3D geometric model of a building based on BIM data or laser point clouds, and divide semantic areas according to business uses (such as offices, equipment rooms, and entrances / exits). Furthermore, this operation can be achieved by introducing building information modeling or high-precision point cloud scanning technology, thereby generating a structured layout with spatial semantics, providing a geographic basis for access control and permission binding.

[0106] Based on the spatial layout and a pre-set personnel management database, the access levels of building personnel are divided to obtain personnel access information.

[0107] This operation can associate fields such as roles, departments, and visitor types in the personnel management database with area attributes in the spatial layout to generate fine-grained permission labels. In an exemplary embodiment, this operation can be achieved by mapping organizational identity information to spatial semantic areas, thus transforming personnel permissions from organizational identity to spatial access capabilities.

[0108] The spatial layout is used to identify and prioritize access areas to obtain the area security level.

[0109] The regional security level can be a security sensitivity level identifier assigned to each spatial area based on building functional zoning and access attributes. This can be used to provide regional-level security context for access control and tracking policies, supporting differentiated permission configuration. In this embodiment, the regional security level can be generated by identifying access areas (such as main passageways, office areas, and computer rooms) in the spatial layout and prioritizing them (such as public, restricted, and high-security) according to business rules, in conjunction with contextual description. For example, the regional security level can include, but is not limited to, one or more of public access level, restricted access level, and high-security control level. This operation can involve analyzing spatial connectivity, pedestrian density, and business sensitivity to classify areas and assign security priority codes. Furthermore, this operation can identify key path nodes (such as access control points and elevator lobbies) using graph theory methods, combine this with manually set security levels, or use historical access control card swipe data to cluster high-frequency / low-frequency access areas and automatically deduce security sensitivity levels. This allows the establishment of a computable regional security attribute system to support differentiated access control.

[0110] Based on personnel permission information, an access control list is generated for the area security level to obtain an initial permission matrix.

[0111] The initial permission matrix can be a structured access control relationship table built based on personnel permission information and area security levels. It represents the static authorization status of personnel to each area and can be used to formally express personnel-area access relationships, serving as the logical basis for subsequent camera-level permission mapping. In a specific embodiment, the initial permission matrix can be described in context by performing a Cartesian product matching between the permission level classification results and the area security level, generating a Boolean or hierarchical access permission matrix according to a preset strategy. For example, the initial permission matrix can include, but is not limited to, role-area permission sub-matrices, time-time-area permission sub-matrices, and temporary visitor permission sub-matrices. This operation can involve matching personnel permission levels with area security levels using a strategy (e.g., "ordinary employees cannot enter high-security control areas") to generate a matrix-style access permission table. Furthermore, this operation can generate a static permission matrix using a role-based access control (RBAC) model or introduce an attribute-based access control (ABAC) mechanism to support joint determination of multi-dimensional attributes (time, task, accompanying status), thereby transforming unstructured permission rules into programmable structured data.

[0112] The camera coverage area and historical access records are extracted by using the camera hierarchical topology structure, and permission mapping is performed by combining the initial permission matrix to obtain the camera permission table.

[0113] The camera permission table can be a device-level access control table formed by mapping abstract region permissions in the initial permission matrix to specific camera nodes. It can be used to implement permission policies on physical monitoring devices, enabling tracking decisions to be made in real-time at the camera level to determine whether to allow continued tracking. In this embodiment, the camera permission table, in conjunction with the context, can project and refine the initial permission matrix using the coverage area (spatial mapping) and historical access records (behavioral statistics) of each node in the camera hierarchical topology. Furthermore, the camera permission table can collaborate with the "camera hierarchical topology": relying on its provided coverage and hierarchical information to complete permission mapping; and collaborate with "target tracking permissions": serving as one of the inputs for dynamic adaptation. For example, the camera permission table can include, but is not limited to, core layer camera permission items, edge layer camera permission items, and cross-regional handover camera permission items. This operation can project region permissions from the initial permission matrix to camera nodes covering the region and adjust the permission granularity by integrating behavioral data such as historical access frequency. In one exemplary embodiment, this operation can be implemented by using the intersection of the camera's view frustum and the three-dimensional spatial layout to determine the precise coverage area and then map permissions, or by increasing the tracking priority of camera-area combinations that frequently have legitimate access and conversely limiting the tracking depth. This enables the implementation of permissions from the area level to the device level, making tracking decisions controllable at the camera granularity.

[0114] Dynamically adapt the camera permission table and initial permission matrix based on the regional security level to obtain target tracking permissions.

[0115] This operation can be achieved by contextually adjusting the static permission table at runtime, taking into account the current regional security level, target identity status, and camera context. Furthermore, this operation can improve management efficiency by temporarily tightening the tracking permission scope in the camera permission table when triggered in a high-security area, or by appropriately relaxing cross-regional tracking permissions for highly trusted targets during low-risk periods. This allows for the generation of target tracking permissions with spatiotemporal context awareness, enabling dynamic compliance constraints on tracking behavior.

[0116] Taking data center building security as an example, the cross-camera target tracking and identity verification method in this embodiment can be as follows: The system performs 3D modeling of the data center, identifying the core area of ​​the server room (high security control level), the operation and maintenance channel (restricted access level), and the reception hall (public access level). IT administrators are classified as high-privilege personnel, while visitors are temporarily classified as low-privilege personnel. The initial permission matrix stipulates that visitors can only enter the reception area. The camera hierarchical topology shows that camera 3 covers the boundary between the reception area and the channel. The system maps the permissions of this area to camera 3, forming a camera permission table: only when the target is a high-privilege personnel are they allowed to cross the channel. One day, a visitor attempts to follow into the channel. The dynamic permission adaptation module detects that the visitor's identity credibility is low and the area's security level jumps, immediately rejecting the tracking request from subsequent cameras and triggering an abnormal behavior assessment, generating an unauthorized warning.

[0117] This embodiment transforms static, fragmented access control into a dynamic tracking and constraint mechanism linked to the monitoring topology. This ensures that the system only performs tracking within the authorized scope, preventing excessive monitoring of unauthorized areas or personnel to protect privacy and compliance. It also provides accurate access control benchmarks for assessing abnormal behavior, effectively supporting the closed-loop management of access control behaviors.

[0118] In one embodiment, cross-camera target collaborative tracking is performed based on the camera hierarchical topology and target tracking permissions to obtain the target motion trajectory, including:

[0119] It receives real-time video streams from each camera in the hierarchical camera topology, selectively accesses and decodes them according to target tracking permissions, and obtains authorized video data.

[0120] The real-time video stream can be the raw video data stream continuously output by each node camera in the hierarchical camera topology, which can be used as the original sensory input source for the tracking system. Authorized video data can be a subset of video streams that are allowed to be accessed and decoded after being filtered according to target tracking permissions. This can be used to ensure that only video content from compliant areas is processed, improving efficiency and ensuring compliance.

[0121] Receiving real-time video streams from each camera in a hierarchical camera topology can be achieved by subscribing to the RTSP or ONVIF streams of all cameras in the topology through a video access service, thereby obtaining raw monitoring data input. Selective access and decoding are performed based on target tracking permissions to obtain authorized video data. This can be achieved by filtering unauthorized camera streams at the video stream access layer based on target tracking permissions, and only decoding authorized streams. Furthermore, this operation can be implemented by deploying lightweight ACL modules on edge nodes to prevent unauthorized streams from being uploaded to the central server, or by dynamically allocating decoding resources in the central decoding pool to prioritize streams from high-privilege areas, thereby reducing unnecessary computational load and enhancing compliance.

[0122] Multi-scale target detection and feature extraction are performed on authorized video data to obtain multi-modal target features.

[0123] Multimodal target features can be a set of target representation vectors extracted from authorized video data during cross-camera tracking, fusing multiple perceptual dimensions (such as appearance, motion, and semantics). These features can enhance the identifiability of targets under complex conditions such as occlusion, low light, and changing viewpoints, thereby improving the robustness of cross-camera matching. In this embodiment, after locating the target using a multi-scale target detector, multimodal target features are extracted using parallel or cascaded deep networks, including heterogeneous features such as face, gait, clothing color, contour, and movement speed, followed by normalization and alignment. For example, multimodal target features can include, but are not limited to, one or more of appearance texture features, dynamic motion features, and semantic context features. Furthermore, multimodal target features can serve as one of the core inputs for constructing the matching cost matrix; their quality directly affects the accuracy of the initial motion trajectory.

[0124] Multi-scale target detection and feature extraction are performed on authorized video data to obtain multimodal target features. This can be achieved by using a multi-scale detection network (such as FPN) to locate targets of different sizes and extract their multimodal representations. In an exemplary embodiment, this operation can be performed by using a shared backbone network and a multi-head output structure to simultaneously extract face, gait, and clothing features, or by processing in stages: first detecting and then cropping, and then feeding each stage into a dedicated feature extraction model. This can improve the richness and discriminativeness of target representations in complex scenes.

[0125] Based on the physical topology hierarchy and logical topology partitioning in the camera hierarchical topology structure, a cross-camera target matching cost matrix is ​​constructed.

[0126] The physical topology hierarchy can be a hierarchical division reflecting the physical spatial organization of a building (such as floors, rooms, and passageways) within the camera's layered topology structure. This hierarchy can provide spatial proximity priors for the matching cost matrix, constraining reasonable transfer paths. Logical topology partitions can be logical regions within the camera's layered topology structure, divided based on function or permission semantics (such as office areas, server rooms, and visitor areas). These partitions can support the organization of tracking processes according to permissions or business rules, enhancing context awareness. The matching cost matrix can be a structured numerical matrix used to quantify the probability of association between detected targets in different cameras. Its elements represent the matching cost of two targets belonging to the same entity. This matrix can provide a unified optimization objective for cross-camera target matching, guiding the trajectory stitching process. In a specific embodiment, the matching cost matrix can calculate the initial cost using multimodal target feature similarity, spatiotemporal distance, and topological adjacency, and then dynamically adjust the weights based on target tracking permissions. For example, the matching cost matrix can include, but is not limited to, feature similarity submatrices, spatiotemporal constraint submatrices, and permission-weighted submatrices.

[0127] Based on the physical topology hierarchy and logical topology partitions in the camera hierarchical topology structure, a cross-camera target matching cost matrix is ​​constructed. This can be achieved by encoding physical adjacency (like floor level) and logical consistency (like permission area) as matching priors and incorporating them into the cost calculation. Furthermore, this operation can be further improved by mapping the topology hierarchy to graph distance as a sparse mask for the cost matrix, or by setting matching thresholds for different logical partitions. Cross-region matching requires higher feature similarity, thereby introducing spatial and semantic structural constraints to suppress unreasonable cross-camera associations.

[0128] The matching cost matrix is ​​weighted based on target tracking permissions, and a hierarchical progressive association algorithm is used for target matching and trajectory stitching to obtain the initial motion trajectory.

[0129] The hierarchical progressive association algorithm can be a tracking algorithm that, guided by a hierarchical camera topology, performs target matching and trajectory aggregation step by step from the bottom local area to the top global scope. It can balance local matching accuracy and global trajectory consistency while reducing computational complexity. In an exemplary embodiment, the hierarchical progressive association algorithm can first complete high-confidence matching within the lowest-level adjacent camera group in the physical topology to generate local sub-trajectories; then merge the sub-trajectories at the logical topology partition level; and finally complete global trajectory stitching at the core layer. Exemplarily, the hierarchical progressive association algorithm may include, but is not limited to, local neighborhood matchers, partitioned trajectory fusion machines, and global trajectory optimizers. The initial motion trajectory can be the unprocessed raw target trajectory sequence output by the hierarchical progressive association algorithm, which may contain noise or abnormal jump points. It can be used as input for the trajectory refinement stage and requires further verification and correction to obtain the final usable trajectory. Furthermore, the initial motion trajectory may include, but is not limited to, local sub-trajectories, cross-regional coarse trajectories, and unverified trajectory segments.

[0130] Adjusting the weights of the matching cost matrix based on target tracking permissions can involve imposing penalty weights on matching items involving unauthorized regions or time periods, thus increasing their matching cost. In one specific embodiment, this operation can be achieved by quantifying the degree of permission violation as a multiplicative factor to amplify the cost of illegal matching, or by directly setting the cost of matching between unauthorized regions to infinite, prohibiting connections, thereby ensuring that trajectory stitching strictly follows access control policies. Using a hierarchical progressive association algorithm for target matching and trajectory stitching to obtain the initial motion trajectory can be achieved by completing local matching at the bottom layer of the topology and then aggregating trajectory fragments layer by layer upwards. Furthermore, this operation can be achieved by using the Hungarian algorithm to solve for the optimal match within each subgraph layer and then achieving cross-layer association through trajectory ID propagation, or by constructing a hierarchical graph model to perform matching in parallel at different granularities and finally fusing the results, thereby balancing computational efficiency and global trajectory consistency.

[0131] The initial trajectory is checked for spatiotemporal continuity and anomaly correction is performed to obtain the target trajectory.

[0132] Spatiotemporal continuity verification can be a process of verifying the physical rationality of the initial trajectory, checking whether the velocity, direction, acceleration, etc., conform to the laws of motion, and can be used to identify and mark abnormal jumps or breaks in the trajectory. Anomaly correction can be a post-processing operation that smooths, interpolates, or removes abnormal trajectory points found in the spatiotemporal continuity verification, and can be used to improve the integrity and reliability of the final target trajectory.

[0133] The initial trajectory undergoes spatiotemporal continuity verification and outlier correction to obtain the target trajectory. This can be achieved by applying a kinematic model (such as the uniform velocity assumption) to detect outliers and then repairing them through filtering or interpolation. Furthermore, this operation can be implemented by using Kalman filtering to estimate and smooth the trajectory's state, or by using linear interpolation based on the positions of preceding and following frames to fill in briefly lost segments, thereby improving the physical plausibility and completeness of the final trajectory.

[0134] For example, in a scenario where personnel are tracking data center personnel during nighttime inspections, the building cross-camera target tracking and identity verification method in this embodiment could involve an authorized maintenance personnel entering the B2 level server room area at 1:00 AM. The system receives video streams from all cameras, but only decodes authorized video data from the B2 level and adjacent channels based on target tracking permissions. A multi-scale detector accurately captures the full-body contour and gait even in low-light conditions, extracting multimodal target features. The matching cost matrix assigns a low cost based on the physical topology level (belonging to the same B2 level), while the logical topology partition (the core area of ​​the server room) further strengthens the internal matching priority. Since the personnel have B2 server room permissions, the matching cost is not penalized. A hierarchical progressive association algorithm first completes local matching within the three cameras on the east side of B2, then stitches it with the trajectory of the camera at the west entrance to form an initial motion trajectory. Spatiotemporal verification detects a positional jump caused by a brief occlusion; after correction using Kalman filtering, a smooth target motion trajectory is output for subsequent identity verification.

[0135] In one embodiment, spatiotemporal correlation analysis and multimodal feature extraction are performed on the target's motion trajectory, and dynamic identity verification is conducted against a preset identity feature database to obtain the target's identity credibility, including:

[0136] Key nodes of the target's motion trajectory are sampled to extract its spatiotemporal features.

[0137] Key nodes can be spatiotemporal location points with high discriminative power or high observation quality within the target's trajectory. They can serve as anchor points for multimodal biometric feature extraction and spatiotemporal alignment, reducing redundant computation and preserving crucial context. In this embodiment, key nodes can be selected by combining spatial transfer events and temporal distribution information provided by the trajectory. For example, key nodes can include, but are not limited to, one or more of the following: field-of-view switching nodes, behavior-residing nodes, and high-quality feature nodes. The trajectory's spatiotemporal features can be spatiotemporal attributes extracted from the target's trajectory that describe its movement pattern. These can be used to guide key node selection, multimodal feature alignment, and historical behavior pattern modeling. Furthermore, trajectory spatiotemporal features can include, but are not limited to, instantaneous motion features, regional transfer features, and temporal distribution features.

[0138] Sampling key nodes of the target's motion trajectory and extracting its spatiotemporal features can be achieved by analyzing the time intervals, spatial change rates, and camera switching events of trajectory points, selecting nodes with high information content, and calculating their kinematic and topological properties. In an exemplary embodiment, this operation can be implemented by detecting and sampling key points based on trajectory curvature and velocity abrupt changes, or by forcibly setting them as key nodes at each cross-camera transfer and adding dwell timeout detection, thereby reducing the amount of data for subsequent feature processing while preserving highly discriminative spatiotemporal context.

[0139] At key nodes, multimodal biometrics of the target are extracted from the corresponding camera video. These multimodal biometrics include face, body shape, and clothing.

[0140] Multimodal biometrics can be various human-related perceptual features extracted from videos that can be used for identity recognition. These features can provide complementary identity cues and improve robustness in complex scenarios. In one specific embodiment, multimodal biometrics may include, but are not limited to, facial representation features, body structure features, and clothing appearance features. At key nodes, extracting multimodal biometrics of the target from the corresponding camera video can involve locating the video frame and target detection box corresponding to the key node, and then calling a dedicated feature extraction network to obtain face, body posture, and clothing embedding vectors respectively. Furthermore, this operation can be achieved by using a unified multi-task network to simultaneously output the three modalities, or by calling independently optimized sub-models (such as ArcFace for face and GaitSet for body posture) for feature extraction, thereby obtaining complementary identity cues and mitigating the problem of single-modality failure.

[0141] Based on the spatiotemporal characteristics of the trajectory, timestamp alignment and quality assessment are performed on multimodal biometrics to obtain a set of features to be verified.

[0142] The feature set to be verified can be a high-quality, temporally consistent set of multimodal biometric features selected after timestamp alignment and quality assessment. This set can be used as input for comparison with an identity feature database, improving matching reliability. In this embodiment, the feature set to be verified can be combined with trajectory spatiotemporal features to synchronously calibrate the original multimodal features, and low-quality samples can be filtered based on indicators such as clarity, completeness, and occlusion rate. For example, the feature set to be verified may include, but is not limited to, aligned facial subsets, effective body posture sequences, and stable clothing feature groups.

[0143] Based on the spatiotemporal characteristics of the trajectory, timestamp alignment and quality assessment are performed on multimodal biometrics to obtain a feature set to be verified. This can be achieved by utilizing the precise temporal and spatial mapping relationship provided by the trajectory to calibrate the temporal reference of each modality feature, and then selecting valid samples based on image quality indicators. Furthermore, this operation can be implemented by using trajectory interpolation to complete feature estimation for missing moments and labeling confidence levels, or by using an occlusion detection model and a blur estimator to score each frame's features and retain only samples above a threshold. This ensures the consistency and reliability of features across cameras and eliminates low-quality interference items.

[0144] The feature set to be verified is compared with the preset identity feature database through hierarchical retrieval and similarity calculation to obtain preliminary matching results.

[0145] Hierarchical retrieval can be a retrieval mechanism that employs a multi-stage strategy for candidate matching in an identity feature database, which can reduce computational overhead while ensuring retrieval accuracy. In a specific embodiment, hierarchical retrieval can combine a fast approximation algorithm with a high-precision similarity model for collaborative operation. For example, hierarchical retrieval may include, but is not limited to, a coarse-grained initial screening layer, a fine-grained precision comparison layer, and a modality-weighted fusion layer. The preliminary matching result can be the candidate identities and their modality similarity scores output after comparing the feature set to be verified with a preset identity feature database, which can be used as the input basis for confidence fusion. Further, the preliminary matching result may include, but is not limited to, face matching candidates, body shape matching candidates, and clothing matching candidates.

[0146] The feature set to be verified is compared with a preset identity feature database through hierarchical retrieval and similarity calculation to obtain preliminary matching results. This can be achieved by first narrowing down the candidate identity set through rapid retrieval, and then performing high-precision multimodal similarity calculation on the candidates. In an exemplary embodiment, this operation can be implemented by using Local Sensitive Hash (LSH) to quickly filter 90% of irrelevant identities in the first layer, calculating cosine similarity and weighting by modality in the second layer, or by pre-filtering the identity database by region permissions before performing fine-grained matching, thereby balancing retrieval efficiency and matching accuracy.

[0147] The confidence level of the target's identity is obtained by fusing the confidence level of the preliminary matching results based on the historical behavior patterns of the target's movement trajectory.

[0148] Historical behavior patterns can be regular behavioral characteristics formed by a target's long-term activities within a building, and can be used to verify consistency and weight confidence of preliminary matching results. In this embodiment, historical behavior patterns can be combined with historical trajectory data clustering or statistical modeling to generate individual behavioral profiles. For example, historical behavior patterns may include, but are not limited to, regional preference patterns, time-activity patterns, and path habit patterns. Confidence fusion can be a logical process of weighted integration of multimodal matching scores and historical behavior consistency scores to generate the final identity credibility, and can be used to achieve joint identity credibility assessment of static features and dynamic behaviors. Furthermore, confidence fusion can dynamically adjust the weights of each dimension based on the current matching quality and behavioral rationality. For example, confidence fusion may include, but is not limited to, modal score fusion, behavioral consistency weighting, and temporal stability correction.

[0149] The confidence level of the target identity is obtained by fusing the confidence scores of the preliminary matching results based on the historical behavior patterns of the target's movement trajectory. This can be achieved by evaluating the consistency between the preliminary matching score and the candidate identity's historical behavior patterns (such as whether it frequently appears in the current area) and fusing them into the final confidence level. In a specific embodiment, this operation can be achieved by significantly reducing the confidence score of the matched identity if it has never entered the area in the current time period, or by using a Bayesian update mechanism to combine the current matching likelihood with the prior historical behavior to calculate the posterior confidence level. This introduces a behavioral corroboration mechanism, enhancing the anti-deception capability and contextual rationality of identity judgment.

[0150] Taking the identity verification of a sensitive laboratory area in a research and development building as an example, the building cross-camera target tracking and identity verification method in this embodiment can be as follows: The system tracks a target from the lobby on the 1st floor to the west corridor on the 3rd floor. The last clear frame before entering the 302 access control is selected as the key node sample. At this node, the system extracts the target's face (partial profile), body shape (medium height, steady gait), and clothing (blue overalls) from the corresponding camera video. The trajectory spatiotemporal features show that the target stayed on the 3rd floor for more than 5 minutes and had no access control card swipe record. The system performs time alignment on the three modal features, removes low-quality facial samples caused by reflection, and retains body shape and clothing features to form a feature set to be verified. The hierarchical search first excludes personnel without access to the 3rd floor, and then performs a detailed comparison on the remaining 20 people, initially matching a certain employee (similarity 0.68). However, the historical behavior pattern shows that employee B has never appeared on the 3rd floor after 20:00 in the past 30 days. Based on this, the confidence fusion module lowers the final identity confidence to 0.41, triggering a "questionable identity + unauthorized stay" composite warning.

[0151] This embodiment reduces the burden of redundant data processing and preserves discriminative spatiotemporal context by sampling key nodes of the target's motion trajectory. It simultaneously extracts multimodal biometric features such as face, body posture, and clothing at key nodes to compensate for the failure of single modalities in complex scenarios. It uses trajectory spatiotemporal features to align multimodal features with timestamps and removes low-quality samples to improve comparison reliability. A hierarchical retrieval strategy balances efficiency and accuracy. Finally, it introduces historical behavior patterns to perform confidence-weighted fusion of preliminary matching results to incorporate behavioral consistency logic. This achieves a dynamic identity credibility assessment mechanism that moves from "passive comparison" to "active verification + behavioral corroboration." It enables high-accuracy and high-reliability cross-camera identity verification in complex building environments, providing a solid foundation for subsequent intelligent early warning of abnormal behavior.

[0152] In one embodiment, abnormal behavior is assessed based on the target's motion trajectory, the target's identity credibility, and the target's tracking permissions to obtain linked early warning information, including:

[0153] The path deviation is calculated based on the target's motion trajectory and spatial layout to obtain trajectory anomaly values;

[0154] The trajectory anomaly value can be a quantitative indicator of path deviation calculated based on a comparison of the target's movement trajectory and the building's spatial layout. It can be used to characterize whether the target exhibits abnormal trajectory behaviors such as detouring into restricted areas, traveling in the wrong direction, or loitering abnormally. In this embodiment, the trajectory anomaly value can be obtained by measuring the geometric or semantic differences between the target's actual movement path and a preset compliant path (such as a passageway or authorized passage), outputting a numerical anomaly score. For example, the trajectory anomaly value may include, but is not limited to, one or more of the following: geometric deviation, semantic violation, and temporal inconsistency.

[0155] Path deviation calculation can be an algorithmic process that measures the difference between a target's actual movement trajectory and a compliant passage path. Furthermore, path deviation calculation can be achieved by using trajectory similarity calculation methods to calculate the geometric similarity between the trajectory and the compliant path and converting it into outliers, or by decomposing the trajectory into a sequence of regions based on a semantic map and comparing the edit distance with the authorized passage area sequence. This allows for the quantification of whether the target exhibits abnormal movement patterns, such as detouring, turning back, or entering restricted areas.

[0156] Based on the target's identity credibility and target tracking permissions, anomaly values ​​in permissions are obtained through permission matching analysis.

[0157] The permission anomaly value can be a quantitative indicator reflecting the degree of matching between the target identity's credibility and its current region's tracking permissions. It can be used to identify permission violations such as unauthorized access, lingering after authorization expires, and impersonation. In an exemplary embodiment, the permission anomaly value can be calculated by comparing the target identity's credibility (e.g., 0.65) with the minimum authorization threshold for the current identity category in that region. Exemplarily, permission anomalies may include, but are not limited to, one or more of the following: identity-region mismatch, time-permission conflict, and role-region overstepping authority.

[0158] Permission matching analysis can be a logical judgment process to assess whether the credibility of a target identity meets the current region's tracking permission requirements. Furthermore, permission matching analysis can be implemented by setting the permission anomaly value to the region threshold minus the credibility if the identity credibility is lower than the region's access threshold, otherwise setting it to zero, or by constructing a permission matching function f(credibility, permission level) to output continuous anomaly scores. This allows for the accurate identification of unauthorized access or permission violations in cases of questionable identity.

[0159] The trajectory anomalies and permission anomalies are weighted and fused to obtain a comprehensive risk score;

[0160] The comprehensive risk score can be a unified risk metric generated by fusing trajectory anomalies and permission anomalies. It can be used to provide a unified criterion for tiered early warning, taking into account both behavioral pattern anomalies and permission status anomalies. In a specific embodiment, the comprehensive risk score can be formed by linearly or non-linearly weighting the two types of anomalies using configurable weights to create a single scalar risk output. For example, the comprehensive risk score can employ weighted linear fusion scoring, logical rule fusion scoring, machine learning fusion scoring, etc.

[0161] Weighted fusion can be a mathematical operation that combines trajectory anomalies and permission anomalies into a comprehensive risk score according to a specific strategy. Furthermore, weighted fusion can be achieved by using fixed weights (such as 0.6:0.4) for linear weighting, or by dynamically adjusting the weight ratio according to the context of the scenario (such as time period, regional sensitivity), thereby achieving unified quantification of multi-dimensional anomaly signals and avoiding misjudgment dominated by a single indicator.

[0162] Based on a comparison between a comprehensive risk score and a preset risk threshold, a tiered early warning instruction is generated.

[0163] Among them, the hierarchical early warning instruction can be a structured response control signal generated based on the comparison result of the comprehensive risk score and the preset risk threshold, which can be used to drive differentiated security response actions and achieve precise grading of early warnings. In an exemplary embodiment, the hierarchical early warning instruction can trigger the response strategy encoding of the corresponding level by mapping the comprehensive risk score to a predefined risk level interval (such as low / medium / high). Exemplarily, the hierarchical early warning instruction can include one or more of, but is not limited to, a first-level emergency early warning instruction, a second-level suspicious early warning instruction, a third-level observation early warning instruction, etc. The preset risk threshold can be a set of predefined comprehensive risk score boundary values for dividing risk levels, which can be used as a judgment benchmark for hierarchical early warnings and support multi-level threshold configuration to adapt to different security strategies.

[0164] Based on the comparison of the comprehensive risk score and the preset risk threshold, the operation of generating the hierarchical early warning instruction can be to match the comprehensive risk score with multiple preset threshold intervals and map it to the corresponding early warning level. Further, this operation can be achieved by triggering a first-level early warning if the score is greater than the high threshold, triggering a second-level early warning if it is between medium and high, and having no early warning if it is lower than the low threshold, or using fuzzy logic to map the score to multi-level membership degrees to generate a composite early warning instruction, so as to achieve refined grading of risks and support differentiated response strategies.

[0165] According to the hierarchical early warning instruction, link the cameras in the corresponding area for key monitoring and push the linked early warning information to the management terminal.

[0166] Among them, the cameras in the corresponding area can be a set of monitoring devices associated with the target's current location or high-risk areas and having physical coverage capabilities, which can be used as the execution terminal for linked responses. After receiving the instruction, they can increase the acquisition frequency, resolution or enable intelligent analysis functions. In a specific embodiment, the cameras in the corresponding area can cooperate with the hierarchical early warning instruction: adjust the working mode according to the instruction content; cooperate with the linked early warning information: its status and video stream are used as supplementary content of the early warning information.

[0167] Key monitoring can be an operation mode of enabling enhanced video acquisition or analysis strategies for specific cameras during specific periods, which can be used to improve the perception granularity and response speed for high-risk targets. Exemplarily, key monitoring can adopt high-frame-rate recording mode, intelligent focusing and tracking mode, multi-modal feature continuous extraction mode, etc. The management terminal can be a human-computer interaction interface or a security duty system that receives and displays the linked early warning information, which can be used for security personnel to view the details of the early warning, confirm the event and perform manual intervention. Exemplarily, the management terminal can include one or more of, but is not limited to, desktop monitoring workstations, mobile emergency APPs, command center large-screen systems, etc.

[0168] The operation of triggering cameras in corresponding areas for focused monitoring based on tiered early warning commands and pushing linked early warning information to the management terminal can involve parsing the level and location information in the early warning command, issuing enhanced monitoring commands to designated cameras, and pushing structured early warning data to the management terminal. Furthermore, this operation can be achieved by triggering cameras to switch to 4K@30fps+AI tracking mode and automatically displaying a pop-up window on the command center's large screen for a level 1 early warning, or by simply increasing the recording bitrate and pushing a notification to security personnel's mobile app for a level 2 early warning. This forms a closed loop from risk assessment to proactive response, improving the system's defense timeliness and accuracy.

[0169] For example, in a scenario of anomaly detection during nighttime inspections of a data center, the building cross-camera target tracking and identity verification method in this embodiment could be as follows: Visitor B is granted access to the reception area on the first floor at 20:00, but at 22:00, their movement trajectory shows them crossing an unauthorized corridor in the equipment room on the second floor. The path deviation calculation module compares their trajectory with the visitor access rules, resulting in a trajectory anomaly value of 0.85; simultaneously, their identity credibility is 0.60, while the second-floor equipment room requires a credibility ≥ 0.90, and the permission matching analysis outputs a permission anomaly value of 0.75. Weighted fusion (weight 0.5:0.5) yields a comprehensive risk score of 0.80, exceeding the high-risk threshold of 0.75, generating a first-level warning instruction. The system automatically activates high-frame-rate recording and face re-recognition for all cameras on the second floor, and pushes the linked warning information, including the target location, risk level, and trajectory screenshot, to the mobile terminal of the on-duty engineer and the command center's large screen.

[0170] This embodiment calculates path deviation by comparing the target's movement trajectory with the building's spatial layout to identify abnormal movements such as detouring into restricted areas and abnormal passage. Simultaneously, it combines the target's identity credibility and real-time tracking permissions to perform permission matching analysis to capture unauthorized access or unauthorized lingering behavior. The two are weighted and fused to form a comprehensive risk score, taking into account both behavioral patterns and permission status to avoid misjudgment by a single indicator. By comparing with a preset risk threshold, risk level classification is achieved to trigger differentiated early warning strategies. Finally, it links designated area cameras to enhance monitoring and pushes structured early warning information to the management terminal to form a closed loop from perception, assessment to response. This can significantly improve the ability to identify complex abnormal behaviors (such as fake authorization + trajectory abnormality), and enhance the accuracy, timeliness, and proactive defense level of early warnings.

[0171] In one embodiment, based on the tiered early warning instruction, cameras in the corresponding areas are linked for focused monitoring, and linked early warning information is pushed to the management terminal, including:

[0172] The tiered early warning instructions are analyzed to determine the warning level, the affected area, and the associated cameras;

[0173] The warning level can be a discrete identifier representing the severity of risk in a tiered warning instruction, which can be used to guide the intensity of camera resource scheduling and the scope of information push. In an exemplary embodiment, the warning level can include, but is not limited to, one or more of Level 1 Emergency, Level 2 Suspicious, and Level 3 Observation. The affected area can be a set of building physical space units where abnormal behavior has occurred or is predicted to occur, which can be used to define the geographical range requiring coordinated response and serve as a basis for camera selection. Associated cameras can be a set of monitoring devices that currently cover the affected area or its adjacent paths and are reachable in the camera hierarchical topology, which can be used as execution units for resource scheduling and image acquisition. Furthermore, associated cameras can coordinate with tiered warning instructions: determined by instruction parsing; and coordinate with resource scheduling: receiving scheduling parameters to adjust their working mode.

[0174] Parsing tiered early warning commands to determine the warning level, affected area, and associated cameras can be achieved by extracting risk level, spatial location, and a list of camera IDs from the structured fields of the commands. For example, this operation can be performed by extracting the 'level', 'zone', and 'cameras' fields from the command using a JSON parser, or by using a rule engine to convert natural language commands into machine-executable area-device mappings. This allows for a structured understanding of early warning events and precise location of response resources.

[0175] Based on the warning level, resources are allocated to associated cameras, including increasing the bit rate, adjusting the focus, and enabling intelligent analysis.

[0176] Resource scheduling can be a control process that dynamically adjusts the operating parameters and functional status of associated cameras based on the warning level. This can be used to prioritize the perception quality of high-risk areas under limited bandwidth and computing power. Bitrate can be the data transmission rate of the video stream per unit time. Increasing the bitrate can enhance image clarity and support subsequent high-precision analysis. Focal length can be an optical parameter representing the focusing distance of the camera lens. Adjusting the focal length can magnify the target in close-up, improving detail recognition. Intelligent analysis can be a real-time video understanding function deployed at the camera end or edge nodes, such as face enhancement, behavior recognition, and target re-identification. This can be enabled during key monitoring phases to improve local perception and feature extraction capabilities. In a specific embodiment, intelligent analysis may include, but is not limited to, one or more of the following: face super-resolution, online gait feature extraction, and abnormal action detection.

[0177] Resource scheduling for associated cameras is implemented based on the warning level, including increasing bitrate, adjusting focus, and enabling intelligent analysis. This can be achieved by mapping preset camera configuration templates to the warning level and issuing parameter adjustment commands. Furthermore, this operation can be implemented by increasing the bitrate to 8Mbps, automatically focusing on the target, and initiating face super-resolution and gait extraction for a Level 1 warning, and only increasing the bitrate to 4Mbps and enabling moving target tracking for a Level 2 warning, or by sending PTZ control and media configuration commands to IP cameras via the ONVIF protocol. This allows for optimization of perception quality and analysis capabilities in key areas under resource-constrained conditions.

[0178] Predict the target's next position based on its movement trajectory and activate the corresponding area's camera deployment in advance;

[0179] The target's next position can be the future short-term spatial coordinates predicted based on the target's motion trajectory using a time-series model. This can be used to pre-activate downstream cameras and eliminate tracking blind spots. In an exemplary embodiment, the target's next position can include, but is not limited to, one or more of the following: short-term position prediction, turning intention estimation, and dwell probability distribution. Camera deployment in the corresponding area can be an operation that pre-sets cameras in the predicted target's entry area to standby or high-sensitivity status. This can be used to achieve seamless tracking and improve the system's proactive responsiveness.

[0180] Predicting the target's next position based on its motion trajectory and pre-activating camera deployment in the corresponding area can be achieved by using a trajectory time series model to predict the position 1-3 seconds in the future, querying the camera hierarchical topology to determine and pre-warm the cameras covering that position. For example, this operation can use the coordinates output by an LSTM trajectory prediction model, match the nearest camera, and pre-load an AI analysis model, or it can be implemented by extrapolating the motion direction and velocity vector to activate the high frame rate mode of the next-hop camera in the topology. This can eliminate tracking blind spots and improve system responsiveness and trajectory continuity.

[0181] The system integrates information such as warning level, target identity credibility, current location, historical trajectory fragments, and related camera footage.

[0182] The current location can be the target's physical coordinates or area identifier within the building space at the current moment, serving as a core spatiotemporal element for early warning information. Historical trajectory fragments can be subsequences of the target's movement path over a period prior to the warning trigger, providing behavioral context and assisting in manual analysis. Associated camera footage can be real-time or cached video frames or keyframe images from associated cameras, serving as visual evidence for early warning information. Information integration is the process of encapsulating multi-source heterogeneous early warning elements (numerical values, text, images, trajectories) into structured data packets, which can help avoid information fragmentation and improve the processing efficiency of management terminals.

[0183] Integrating information such as warning level, target identity credibility, current location, historical trajectory fragments, and related camera footage can be achieved by encapsulating each element into a unified data structure (such as Protobuf or JSON) according to a predefined schema. Furthermore, this operation can be implemented by generating a composite message containing metadata headers, trajectory GeoJSON, identity credibility, and image Base64 encoding, or by constructing a lightweight web interface snapshot with embedded trajectory animation and multi-view stitched images. This avoids information fragmentation and improves the efficiency of manual analysis.

[0184] Based on the preset terminal permission list, the integrated linkage early warning information is pushed to the corresponding level of management terminal.

[0185] The terminal permission list can be a preset mapping table of management terminals and the sensitivity levels of the warning information they can receive. This can be used to control the information distribution scope and ensure that highly sensitive data is only disclosed to authorized terminals. In one specific embodiment, the terminal permission list may include, but is not limited to, one or more of the following: command center terminals with full access, regional security restricted terminals, and inspection personnel read-only terminals. The integrated linked warning information can be a structured message containing elements such as warning level, identity credibility, current location, historical trajectory fragments, and associated camera footage. This can be used as a standardized alarm payload pushed to the management terminal. The corresponding level of management terminal can be a terminal device matched according to the terminal permission list that has the permission to receive the current warning information. This can be used to achieve differentiated information distribution, balancing response efficiency and privacy compliance.

[0186] Based on a pre-defined list of terminal permissions, the integrated alert information is pushed to the corresponding management terminals. This can be achieved by comparing the sensitivity level of the alert information with the terminal permission list, selecting authorized receiving terminals, and then pushing the alerts. For example, a Level 1 alert can be pushed only to the command center's large screen and the duty manager's app, while a Level 3 alert can be broadcast to all regional security terminals. Alternatively, it can be implemented through MQTT topic-based hierarchical publishing, with terminals receiving corresponding messages according to their subscription permissions. This ensures that highly sensitive information is disclosed on demand, balancing response efficiency with data privacy compliance.

[0187] Taking the response to an abnormal intrusion in the VIP area of ​​a financial building as an example, the building cross-camera target tracking and identity verification method in this embodiment can be as follows: The system determines that a target's comprehensive risk score is 0.88 and generates a level-one warning instruction. The parsing module extracts the warning level as "Level 1", the area involved as "5th floor vault passage", and the associated cameras as C501-C503. The resource scheduling module immediately issues instructions to the three cameras: increase the bit rate to 8Mbps, lock the focus on the target, and enable face super-resolution and clothing feature extraction. At the same time, the trajectory prediction module, based on the target's westward movement trend, predicts that it will enter the west stairwell on the 5th floor in 2 seconds, and activates camera C504 in advance to enter a high-sensitivity deployment state. The information integration module packages the warning level, identity credibility of 0.55, current location (5F-passage B), trajectory of the last 30 seconds, and real-time image of C502. After the terminal permission list is matched, the integrated information is only pushed to the security command center's large screen and the vault supervisor's encrypted mobile APP; ordinary floor security terminals do not receive notifications.

[0188] This embodiment achieves a structured understanding of early warning events to accurately locate risk areas and the necessary monitoring resources. It dynamically schedules video parameters of associated cameras based on the early warning level and enables intelligent analysis to optimize perception quality in key areas under resource constraints. It predicts the future location of targets based on their movement trajectories to pre-activate downstream camera deployment, eliminating tracking blind spots. It integrates multi-dimensional information into structured, linked early warning information to avoid information fragmentation. Differential push notifications based on terminal permission lists ensure that highly sensitive information is only disclosed to authorized personnel. This approach transforms passive alarms into a proactive, precise, tiered, and collaborative emergency response system, significantly enhancing the situational awareness, resource scheduling efficiency, and closed-loop handling capabilities of building security systems.

[0189] Furthermore, to achieve the above objectives, the present invention also provides a building cross-camera target tracking and identity verification system, the system comprising:

[0190] The topology building module is used to acquire network connection data and spatial deployment information of building cameras, perform hierarchical topology building, and obtain the hierarchical topology structure of the cameras;

[0191] The permission modeling module is used to obtain the spatial layout and personnel permission information of the building, and to dynamically construct permissions in combination with the hierarchical topology of the camera to obtain target tracking permissions;

[0192] The collaborative tracking module is used to perform cross-camera target collaborative tracking based on the camera hierarchical topology and the target tracking permission to obtain the target motion trajectory;

[0193] The identity verification module is used to perform spatiotemporal correlation analysis and multimodal feature extraction on the target's motion trajectory, and to perform dynamic identity verification with a preset identity feature library to obtain the target's identity credibility.

[0194] The early warning linkage module is used to evaluate abnormal behavior based on the target's movement trajectory, the target's identity credibility, and the target's tracking permissions, and obtain linkage early warning information.

[0195] Other embodiments or specific implementations of the building cross-camera target tracking and identity verification system of the present invention can be referred to the above-described method embodiments, and will not be repeated here.

[0196] In addition, to achieve the above objectives, the present invention also provides a building cross-camera target tracking and identity verification device, the device comprising: a memory, a processor, and a building cross-camera target tracking and identity verification program stored in the memory and executable on the processor, the building cross-camera target tracking and identity verification program being configured to implement the steps of the building cross-camera target tracking and identity verification method as described above.

[0197] In addition, to achieve the above objectives, the present invention also provides a medium storing a building cross-camera target tracking and identity verification program, wherein when the building cross-camera target tracking and identity verification program is executed by a processor, it implements the steps of the building cross-camera target tracking and identity verification method as described above.

[0198] The above are merely preferred embodiments of the present invention and do not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A method for cross-camera target tracking and identity verification in buildings, characterized in that, The method includes: Obtain network connection data and spatial deployment information of building cameras, construct a layered topology, and obtain the camera layered topology structure. The spatial layout and personnel permission information of the building are obtained, and dynamic permission is constructed by combining the hierarchical topology of the camera to obtain target tracking permission; Based on the hierarchical topology of the cameras and the target tracking permissions, cross-camera target collaborative tracking is performed to obtain the target motion trajectory; Spatiotemporal correlation analysis and multimodal feature extraction are performed on the target's motion trajectory, and dynamic identity verification is performed against a preset identity feature database to obtain the target's identity credibility. Based on the target's motion trajectory, the target's identity credibility, and the target's tracking permissions, abnormal behavior is evaluated to obtain linked early warning information.

2. The building cross-camera target tracking and identity verification method as described in claim 1, characterized in that, The process of acquiring network connection data and spatial deployment information of building cameras, constructing a layered topology, and obtaining a layered topology structure for the cameras includes: Based on preset camera identifiers, the building cameras are discovered to obtain network connection data; The building is divided into spatial zones, and spatial deployment information is obtained by combining the camera installation locations; The network connection data is subjected to link quality assessment and hierarchical division to obtain the physical topology hierarchy; Logical topology partitions are obtained by logically associating regions based on the spatial deployment information. The physical topology hierarchy is merged with the logical topology partition to obtain an initial hierarchical topology; The initial hierarchical topology is subjected to connectivity verification and path weight assignment to obtain the camera hierarchical topology structure.

3. The building cross-camera target tracking and identity verification method as described in claim 1, characterized in that, The process of acquiring the building's spatial layout and personnel access information, and combining this with the camera's hierarchical topology to dynamically construct access permissions, thereby obtaining target tracking permissions, includes: The building is modeled in 3D and functionally partitioned to obtain the spatial layout; Based on the spatial layout and the preset personnel management database, the permission levels of building personnel are divided to obtain the personnel permission information; The spatial layout is subjected to access area identification and priority encoding to obtain the area security level; Based on the personnel permission information, an access control list is generated for the area security level to obtain an initial permission matrix; The camera coverage area and historical access records are extracted through the camera hierarchical topology, and permission mapping is performed in combination with the initial permission matrix to obtain the camera permission table; Based on the area security level, the camera permission table and the initial permission matrix are dynamically adapted to obtain the target tracking permission.

4. The building cross-camera target tracking and identity verification method as described in claim 1, characterized in that, The method of cross-camera target collaborative tracking based on the camera hierarchical topology and the target tracking permission to obtain the target motion trajectory includes: Receive real-time video streams from each camera in the hierarchical camera topology, selectively access and decode them according to the target tracking permissions, and obtain authorized video data; Multi-scale target detection and feature extraction are performed on the authorized video data to obtain multi-modal target features; Based on the physical topology hierarchy and logical topology partitioning in the camera hierarchical topology structure, a cross-camera target matching cost matrix is ​​constructed; The matching cost matrix is ​​weighted based on the target tracking permission, and a hierarchical progressive association algorithm is used for target matching and trajectory stitching to obtain the initial motion trajectory. The initial motion trajectory is subjected to spatiotemporal continuity verification and outlier correction to obtain the target motion trajectory.

5. The building cross-camera target tracking and identity verification method as described in claim 1, characterized in that, The process of performing spatiotemporal correlation analysis and multimodal feature extraction on the target's motion trajectory, and dynamically verifying its identity against a preset identity feature database to obtain the target's identity credibility includes: Key nodes of the target motion trajectory are sampled to extract the trajectory's spatiotemporal features; At the key nodes, multimodal biometrics of the target are extracted from the corresponding camera video. These multimodal biometrics include face, body shape, and clothing. Based on the spatiotemporal characteristics of the trajectory, the multimodal biometrics are timestamped and their quality is assessed to obtain a feature set to be verified. The set of features to be verified is compared with a preset identity feature database by hierarchical retrieval and similarity calculation to obtain preliminary matching results; The confidence level of the target identity is obtained by fusing the confidence level of the preliminary matching results based on the historical behavior patterns of the target's motion trajectory.

6. The building cross-camera target tracking and identity verification method as described in claim 1, characterized in that, The abnormal behavior assessment based on the target's motion trajectory, the target's identity credibility, and the target's tracking permissions, to obtain linked early warning information, includes: The path deviation between the target's motion trajectory and the spatial layout is calculated to obtain the trajectory anomaly value. Based on the target identity credibility and the target tracking permissions, an access matching degree analysis is performed to obtain permission anomaly values; The trajectory anomalies and permission anomalies are weighted and fused to obtain a comprehensive risk score; Based on the comparison between the comprehensive risk score and the preset risk threshold, a graded early warning instruction is generated; Based on the tiered early warning instructions, cameras in the corresponding areas are linked for focused monitoring, and the linked early warning information is pushed to the management terminal.

7. The building cross-camera target tracking and identity verification method as described in claim 6, characterized in that, The step of linking cameras in the corresponding areas for focused monitoring based on the tiered early warning instruction and pushing the linked early warning information to the management terminal includes: The tiered early warning instructions are parsed to determine the early warning level, the affected area, and the associated cameras; Based on the warning level, resource scheduling is performed on the associated cameras, including increasing the bit rate, adjusting the focal length, and enabling intelligent analysis; Predict the target's next position based on the target's movement trajectory, and activate the corresponding area's camera deployment in advance; The system integrates information such as warning level, target identity credibility, current location, historical trajectory fragments, and related camera footage. Based on the preset terminal permission list, the integrated linkage warning information is pushed to the corresponding level of management terminal.

8. A building-wide cross-camera target tracking and identity verification system, characterized in that, The system includes: The topology building module is used to acquire network connection data and spatial deployment information of building cameras, perform hierarchical topology building, and obtain the hierarchical topology structure of the cameras; The permission modeling module is used to obtain the spatial layout and personnel permission information of the building, and to dynamically construct permissions in combination with the hierarchical topology of the camera to obtain target tracking permissions; The collaborative tracking module is used to perform cross-camera target collaborative tracking based on the camera hierarchical topology and the target tracking permission to obtain the target motion trajectory; The identity verification module is used to perform spatiotemporal correlation analysis and multimodal feature extraction on the target's motion trajectory, and to perform dynamic identity verification with a preset identity feature library to obtain the target's identity credibility. The early warning linkage module is used to evaluate abnormal behavior based on the target's movement trajectory, the target's identity credibility, and the target's tracking permissions, and obtain linkage early warning information.

9. A building-based cross-camera target tracking and identity verification device, characterized in that, The device includes: a memory, a processor, and a building cross-camera target tracking and identity verification program stored in the memory and executable on the processor, the building cross-camera target tracking and identity verification program being configured to implement the steps of the building cross-camera target tracking and identity verification method as described in any one of claims 1 to 7.

10. A medium, characterized in that, The medium stores a building cross-camera target tracking and identity verification program, which, when executed by a processor, implements the steps of the building cross-camera target tracking and identity verification method as described in any one of claims 1 to 7.