Multi-camera cooperative non-blind area intelligent monitoring method
By employing a multi-camera collaborative blind-spot-free intelligent monitoring method, the problems of blind spots and difficulty in target re-identification in traditional monitoring systems are solved. This method enables seamless collaboration and intelligent analysis of multi-camera systems, improving monitoring efficiency and accuracy, and dynamically optimizing sensing resources and identifying abnormal behavior.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-20
- Publication Date
- 2026-04-03
AI Technical Summary
Traditional surveillance systems with multiple cameras suffer from blind spots, lack a unified spatiotemporal coordinate system, and have difficulty in re-identifying targets, resulting in low monitoring efficiency and difficulty in quickly locating target positions, thus failing to meet the panoramic monitoring needs in large-scale and complex environments.
The system establishes a mapping relationship between the image coordinates of each camera and a unified world coordinate system through system calibration, performs time synchronization and preprocessing, performs continuous tracking based on cross-camera target detection and re-identification algorithms, reconstructs the 3D scene structure in real time using multi-view geometry principles, selects the best observation viewpoint by combining the viewpoint quality assessment model and controls the gimbal camera to perform active tracking, and performs anomaly detection based on spatiotemporal graph convolutional networks.
It enables blind-spot-free collaborative perception of multi-camera systems, ensuring the continuity of target trajectories and consistency of identity, improving the real-time performance and accuracy of the monitoring system, dynamically optimizing perception resources, identifying abnormal behavior, and providing early warnings.
Smart Images

Figure CN121792702A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of intelligent monitoring and security technology, specifically relating to a multi-camera collaborative blind-spot-free intelligent monitoring method. Background Technology
[0002] With the increasing demand for security monitoring, multi-camera systems are widely deployed in large-scale space monitoring scenarios (such as squares, airports, and shopping malls). However, traditional monitoring systems typically use single fixed-view cameras or multiple independent cameras, which have many limitations. For example, a single camera is limited by its field of view and cannot cover the entire monitoring area, resulting in blind spots. Multiple cameras lack effective coordination mechanisms; each camera collects data independently without a unified spatiotemporal coordinate system, making it easy for targets to become confused or their trajectory to be interrupted when switching between different cameras. Furthermore, due to differences in perspective, scale, and lighting conditions, target re-identification is difficult, requiring monitoring personnel to manually switch between multiple views to obtain the target's continuous trajectory, which is inefficient. In response to emergencies, it is difficult to quickly locate the target's specific position in the panoramic view, affecting the real-time performance and accuracy of the security system. While some multi-camera tracking solutions exist in existing technologies, they often fail to achieve truly seamless collaboration and intelligent analysis, and cannot meet the panoramic monitoring needs of large-scale, complex environments. Summary of the Invention
[0003] This application provides a multi-camera collaborative blind-spot-free intelligent monitoring method to solve one of the aforementioned technical problems.
[0004] The technical solution adopted in this application is as follows: This application provides a multi-camera collaborative blind-spot-free intelligent monitoring method, including: Perform system calibration on multiple cameras and establish the mapping relationship between the image coordinates of each camera and a unified world coordinate system; The data collected by the multiple cameras are synchronized in time and preprocessed. Based on cross-camera target detection and re-identification algorithms, the target is continuously tracked; Based on the principle of multi-view geometry, the three-dimensional scene structure of the monitored area is reconstructed in real time. Based on the viewpoint quality assessment model, the best viewing angle is automatically selected and the gimbal camera is controlled to actively track the view. Anomaly detection is achieved by modeling cross-camera behavior using spatiotemporal graph convolutional networks.
[0005] According to one embodiment of this application, system calibration of multiple cameras includes: An automatic calibration algorithm based on feature points is used to establish the mapping relationship between the image coordinates of each camera and the unified world coordinate system; Design an online calibration compensation mechanism to automatically detect changes in camera position and recalibrate.
[0006] According to one embodiment of this application, time synchronization and preprocessing of data collected by the plurality of cameras include: The PTP precision clock protocol is used to achieve microsecond-level time synchronization, ensuring the time consistency of data acquisition from multiple cameras; Establish an image quality assessment module to automatically identify quality issues such as occlusion, overexposure, and blurring, and trigger corresponding camera parameter adjustments. Design an illumination equalization algorithm to eliminate color inconsistencies caused by differences in exposure parameters between different cameras.
[0007] According to one embodiment of this application, continuous tracking of a target based on a cross-camera target detection and re-identification algorithm includes: Construct a cascaded target association network, where: The first level performs coarse correlation based on appearance and motion features; The second level is based on graph neural networks to model the spatiotemporal relationships between targets; The third level utilizes three-dimensional position information for geometric verification; An incremental learning mechanism is designed to update the re-identification model online, adapting to changes in the target's appearance.
[0008] According to one embodiment of this application, real-time reconstruction of the three-dimensional scene structure of a monitoring area based on multi-view geometry principles includes: Design a lightweight SLAM algorithm to achieve real-time scene reconstruction and target localization on edge computing nodes; Establish a fast mapping relationship between two-dimensional image coordinates and three-dimensional world coordinates, and support the ability to locate the target's position in actual space by clicking on the image.
[0009] According to one embodiment of this application, automatically selecting the optimal viewing angle based on a viewing angle quality assessment model includes: Construct a viewpoint quality assessment model to evaluate the target integrity, sharpness, and occlusion level in real time for each camera's image; Design a camera scheduling strategy based on reinforcement learning to automatically select the best viewing angle and control the gimbal camera for active tracking; When the target enters the blind spot or leaves the current camera's field of view, it automatically switches to the camera that can continue tracking.
[0010] According to one embodiment of this application, anomaly detection is achieved by modeling cross-camera behavior based on a spatiotemporal graph convolutional network, including: Design a panoramic anomaly detection algorithm to identify cross-regional anomaly patterns such as crowd gathering, abnormal loitering, and rapid movement; Establish a behavioral semantic map to organize discrete monitoring events into a coherent behavioral narrative.
[0011] According to one embodiment of this application, it also includes: It adopts an edge-cloud collaborative computing architecture, which completes real-time data processing at the edge nodes and performs in-depth analysis in the cloud; Design an intelligent video storage strategy that adaptively adjusts storage quality and duration based on event importance; By employing video summarization technology, timelines of key events are automatically generated, reducing storage space and retrieval time.
[0012] A second aspect of this application provides a computer-readable storage medium having a program stored thereon that, when executed by a processor, implements the steps described in the method.
[0013] A third aspect of this application provides an electronic device including a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the method as described.
[0014] Due to the adoption of the above technical solution, the beneficial effects achieved by this application are as follows: This application fundamentally solves the problem of inconsistent coordinate systems in traditional multi-camera systems by systematically calibrating multiple cameras and establishing a mapping relationship between the image coordinates of each camera and a unified world coordinate system. This method provides a unified spatial reference for the entire system, eliminates positioning errors and blind spots caused by different spatial reference systems, and lays a precise geometric foundation for all subsequent processing procedures.
[0015] By rigorously synchronizing and preprocessing the data acquired by multiple cameras, the inconsistency caused by differences in acquisition time and imaging conditions (such as lighting and exposure) is effectively overcome. This solution ensures the time and quality synergy of the data acquired by subsequent processing modules, providing a reliable data foundation for accurate cross-camera analysis and correlation.
[0016] Based on advanced cross-camera target detection and re-identification algorithms, this approach effectively overcomes tracking interruptions or identity confusion caused by factors such as changes in target appearance, scale scaling, differences in viewpoint, and temporary occlusion. This scheme ensures the continuity of the target's trajectory and the consistency of its identity when crossing different camera fields of view, significantly improving the robustness and reliability of the tracking system.
[0017] By reconstructing the 3D scene structure of the monitored area in real time based on multi-view geometric principles, the system leaps from 2D image perception to 3D spatial perception. This method enables targets not only to be identified and tracked, but also to be accurately located in the real-world coordinate system, greatly improving the accuracy of location-related analysis and applications.
[0018] The system automatically selects the optimal viewing angle based on a viewpoint quality assessment model and controls the gimbal camera for active tracking, enabling the system to dynamically optimize its sensing resources. This approach effectively reduces tracking loss caused by limitations of fixed viewpoints or target movement, achieving uninterrupted, high-quality continuous monitoring of key targets and reducing the need for manual intervention.
[0019] By modeling cross-camera behavior using spatiotemporal graph convolutional networks, the system can understand the complex activity patterns of targets in both spatial and temporal dimensions. Combined with a panoramic anomaly detection algorithm, this approach can efficiently identify abnormal behavior patterns scattered across the fields of view of different cameras, thereby enabling early detection and rapid warning of security threats. Attached Figure Description
[0020] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A flowchart illustrating a multi-camera collaborative blind-spot-free intelligent monitoring method provided in an embodiment of this application; Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0021] Figure label: 810, Processor; 820, Communication interface; 830, Memory; 840, Communication bus. Detailed Implementation
[0022] To more clearly illustrate the overall concept of this application, a detailed explanation is provided below with reference to the accompanying drawings.
[0023] Many specific details are set forth in the following description to provide a thorough understanding of this application. However, this application may also be implemented in other ways different from those described herein. Therefore, the scope of protection of this application is not limited to the specific embodiments disclosed below. It should be noted that, unless otherwise specified, the embodiments of this application and the features thereof can be combined with each other.
[0024] In this application, unless otherwise expressly specified and limited, the "above" or "below" of the second feature can mean that the first and second features are in direct contact, or that the first and second features are in indirect contact through an intermediate medium. In the description of this specification, references to terms such as "an embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described can be combined in any suitable manner in one or more embodiments or examples.
[0025] Example 1 like Figure 1 As shown, a multi-camera collaborative blind-spot-free intelligent monitoring method includes: The system calibrates multiple cameras and establishes a mapping relationship between the image coordinates of each camera and a unified world coordinate system.
[0026] As mentioned above, this step is the core foundation and prerequisite for achieving multi-camera collaborative perception in this scheme. Its purpose is to integrate multiple cameras, which may have different positions, orientations, and focal lengths within a physical space, into a unified spatial reference system. Specifically, "system calibration" refers to solving for the internal parameters (such as focal length, principal point coordinates, distortion coefficients, etc.) and external parameters (i.e., the position and orientation of each camera relative to a common world coordinate system, also known as pose) of each camera through specific algorithms and processes. "Establishing mapping relationships" refers to forming a mathematical transformation model based on the calibration parameters. This model can calculate, through camera models and geometric transformations, the possible ray or specific 3D point coordinates of a pixel coordinate (i.e., a two-dimensional position in the image coordinate system) in the real three-dimensional space (i.e., the world coordinate system); conversely, it can also accurately project any 3D point in the world coordinate system onto the image planes of different cameras, predicting its position in different frames. This process unifies the originally isolated two-dimensional image information into a three-dimensional space with actual physical meaning.
[0027] For example, consider a surveillance system in a large train station waiting hall. The hall is equipped with a fixed camera A at the entrance providing a panoramic view, a rotatable, zoomable pan-tilt camera B in the center of the hall, and a fisheye camera C in a corner providing a wide-angle view. During system initialization, calibration is required. Operators can place a calibration board with known precise dimensions on the waiting hall floor, or directly utilize inherent, known static feature points in the scene (such as the corners of pillars or fixed seats). By having cameras A, B, and C observe these common features from different angles, the system runs a calibration algorithm. The algorithm extracts the pixel coordinates of these feature points from each camera image and combines them with their real-world coordinates to calculate: the optical center of camera A is at (0, 0, 5) meters in the world coordinate system, facing directly downwards; the optical center of camera B is at (10, 5, 3) meters, facing at a certain angle; and the optical center of camera C is at (20, 0, 2.5) meters. Simultaneously, the distortion parameters of its fisheye lens are also obtained. At this point, the mapping relationship is established. For example, when the system detects a passenger at pixel coordinates (320, 240) in the image of camera A, it can calculate the passenger's actual world coordinates on the waiting hall floor (e.g., (15, 8, 0) meters) using this mapping relationship. Based on these coordinates, the system can immediately guide the pan-tilt camera B to rotate and zoom, aligning its image center with the world coordinates (15, 8, 0), thus achieving relay tracking of the target.
[0028] It should be noted that, in specific implementation scenarios, based on the above solution, the system can continuously detect and utilize the intersection points of the trajectories of moving targets (such as pedestrians and vehicles) in different camera views as virtual calibration objects during operation, dynamically optimize and update the camera calibration parameters, thereby automatically compensating for pose changes caused by camera bracket loosening, wind blowing, or manual adjustment, and ensuring the long-term accuracy of the mapping relationship.
[0029] In specific implementation scenarios, the unified world coordinate system and mapping relationships established based on the above scheme can be applied beyond target localization. It can also serve 3D scene reconstruction. By observing the same area from multiple perspectives using multiple cameras and using the 3D information calculated from the mapping relationships, sparse or dense 3D point cloud models of the monitored scene can be constructed in real time, providing richer spatial context for subsequent behavior analysis, virtual fencing, and other applications.
[0030] In specific implementation scenarios, based on the above solutions, the system can also use the corresponding camera imaging models (such as pinhole models, spherical models, etc.) to solve for parameters and perform coordinate transformations for ordinary pinhole cameras, wide-angle cameras, fisheye cameras, and even panoramic cameras. This means that in the process of establishing a unified world coordinate system, the system can properly handle the different distortion characteristics and projection methods brought about by different lens types, achieving true all-round, distortion-free spatial unification.
[0031] The data collected by the multiple cameras are synchronized in time and preprocessed.
[0032] As mentioned above, "time synchronization" refers to providing a unified time reference for all cameras in the network through a precise clock protocol, enabling each frame of image data to be marked with a high-precision timestamp. This ensures that the same moving target seen in different camera views occurs at the same absolute moment, thus providing accurate time basis for cross-camera motion trajectory correlation and target handover judgment. "Preprocessing" refers to a series of quality optimization and standardization operations performed before the data enters the core analysis module. It mainly includes two aspects: first, "image quality assessment," which uses algorithms to automatically detect whether there are quality problems in the image caused by environmental or device-specific factors, such as motion blur, overexposure, underexposure, or image defects caused by obstructions; second, "image quality enhancement," which addresses the above problems or imaging differences between different cameras by performing processes such as illumination equalization and color correction to reduce apparent differences caused by factors other than the target itself.
[0033] For example, consider a traffic monitoring system at a city intersection. The system deploys cameras at four intersections. First, using a precision clock protocol, the system synchronizes the internal clocks of the four cameras to microsecond-level accuracy. When a car crosses the intersection, the moment it appears in the images of all four cameras is recorded with the same timestamp. This allows the system to accurately calculate the precise time sequence of the vehicle's passage through different locations within the intersection, thus determining whether it is speeding or running a red light. During the preprocessing stage, the system detects that the camera facing east is overexposed due to direct sunlight in the morning, making license plates difficult to recognize. Therefore, it automatically triggers an adjustment to the exposure parameters of this camera. Simultaneously, the system finds that the cameras facing north and south are darker due to being in shadow. It then uses a lighting equalization algorithm to adjust the brightness and color style of these cameras to a level similar to the other cameras, ensuring that subsequent vehicle re-identification algorithms do not misidentify the same vehicle as different vehicles due to color distortion.
[0034] It should be noted that, in specific implementation scenarios, the high-quality data obtained after synchronization and preprocessing, based on the above solutions, can directly serve "cross-camera target trajectory association," because accurate timestamps are the only reliable basis for associating the motion trajectory of the same target from different viewpoints. Simultaneously, the consistency of illumination and color achieved through preprocessing significantly improves the accuracy and robustness of the "target re-identification based on appearance features" model. Furthermore, the output of the image quality assessment module (such as occlusion degree and sharpness score) can itself serve as direct input to the "viewpoint quality assessment model," providing real-time decision-making basis for the subsequent "camera scheduling strategy based on reinforcement learning" to dynamically select the optimal viewpoint. The entire preprocessing process can be flexibly completed collaboratively on edge computing nodes or in the cloud, depending on the distribution of computing resources, thereby achieving an optimal balance between efficiency and effectiveness.
[0035] The target is continuously tracked based on a cross-camera target detection and re-identification algorithm.
[0036] As described above, the "cross-camera target detection and re-identification algorithm" is a complex technical process. It first detects and locates all targets of interest (such as pedestrians and vehicles) in real time within each independent camera video stream, generating an independent tracking segment for each target. Then, when a target moves from one camera's field of view to another, the algorithm comprehensively compares the similarity of the targets in a multi-dimensional feature space to determine whether targets appearing in different cameras at different times are the same entity, thus achieving "re-identification." To achieve high-precision re-identification, this scheme employs a cascaded association strategy: first, it performs rapid coarse association based on the target's appearance features (such as color, texture, and clothing) and instantaneous motion features to filter candidate targets; then, it uses graph neural networks to model the complex spatiotemporal relationships between targets to distinguish targets that appear similar but are actually different; finally, it introduces a unified world coordinate system obtained during system calibration and 3D reconstruction to perform geometric consistency verification on candidate associations, i.e., determining whether the path for a target to move from one camera's field of view to another is physically feasible, thereby eliminating obviously erroneous associations. In addition, the algorithm integrates an incremental learning mechanism, which can dynamically update its feature model based on the slow changes in the target's appearance during tracking (such as the effects of lighting and changes in viewpoint), ensuring that the re-identification model can adapt to the evolution of the target's appearance.
[0037] For example, in a surveillance scenario of a large shopping mall, a woman wearing a red top (the target) walks out of the field of view of camera A at the entrance. The system continuously tracks her in camera A and extracts her physical features. When she enters the central hall, this area is covered by cameras B, C, and D. The system needs to determine which of the three cameras showing the woman in the red top is the one from the entrance. Based solely on color (coarse correlation), there may be multiple candidates. At this point, the algorithm activates a graph neural network to analyze the woman's reasonable walking path from the entrance to the hall and the time constraints, inferring that camera C is the most likely candidate. Furthermore, the system uses 3D world coordinates for verification: mapping the target's last position in camera A and the positions of the candidate targets in camera C onto a unified 3D map, determining whether the straight-line distance between the two points is reachable within a reasonable walking time. Through this series of comprehensive judgments, the system ultimately confirms that the target in camera C and the target in camera A are the same person, thus achieving seamless identity transfer and continuous trajectory stitching.
[0038] It should be noted that, in specific implementation scenarios, based on the above solutions, the continuous, cross-camera trajectory output by the cross-camera target detection and re-identification algorithm serves as the core data foundation for multiple upper-layer application functions. This precise trajectory data is directly fed into the "intelligent perspective switching and active tracking" module, serving as the decision-making basis for controlling the PTZ camera to perform relay tracking. Simultaneously, the continuous trajectories of all targets are recorded and aggregated in a unified world coordinate system, collectively forming the raw data for "panoramic behavior analysis and anomaly detection." The spatiotemporal graph convolutional network is able to learn and identify complex group or individual behavioral patterns, such as crowd gatherings and abnormal loitering, based on this extensive, cross-regional trajectory data. Furthermore, the incremental learning mechanism in the re-identification algorithm ensures that the system maintains robustness in tracking when facing long-term changes in the monitoring scenario (such as clothing changes caused by seasonal transitions), thereby improving the stability and reliability throughout the entire system lifecycle.
[0039] Based on the principle of multi-view geometry, the three-dimensional scene structure of the monitored area is reconstructed in real time.
[0040] As mentioned above, "based on multi-view geometry" refers to a series of computer vision methods that use two-dimensional images of the same static scene observed from different viewpoints (i.e., different camera positions) to reconstruct the three-dimensional structure of the scene. Its core lies in finding pixels in different images that correspond to the same physical point in the real world (i.e., feature matching), and combining known camera parameters (obtained through the aforementioned system calibration) and their relative positional relationships, using triangulation to calculate the precise coordinates of that physical point in three-dimensional space. After completing this calculation for a large number of feature points in the scene, it is possible to "reconstruct the three-dimensional scene structure of the monitored area in real time," generating a point cloud model or a more advanced mesh model composed of these three-dimensional points that reflects the geometric contours of the environment. To achieve "real-time performance," this solution preferably employs lightweight Simultaneous Localization and Mapping (SLAM) algorithms or Structure of Motion (SfM) algorithms. These algorithms are optimized to run efficiently on edge computing nodes, continuously process video stream data, and dynamically update the three-dimensional scene model.
[0041] For example, in a museum lobby, cameras deployed at the four corners capture real-time images of the lobby from different angles. The system first uses calibrated camera parameters to perform cross-view matching of feature points detected in each camera's image (such as corner points of picture frames, edges of display stands, and textures of pillars). Then, through multi-view geometric triangulation, the system calculates the 3D coordinates of each matched feature point in the actual space of the lobby. For instance, a chandelier hanging in the center appears as different pixel positions in the four camera images, but using this algorithm, the system can accurately calculate the actual hanging position (X, Y, Z) coordinates of this chandelier in the museum's 3D space. By calculating thousands of such feature points, the system constructs and maintains a 3D digital scene in real-time on its internal computer that is consistent with the geometry of the actual lobby. If the system detects a visitor in a 2D image from one of the cameras, it can immediately determine the visitor's precise location in the 3D digital scene through coordinate mapping.
[0042] It should be noted that, in specific implementation scenarios, the real-time reconstructed 3D scene structure, based on the above solution, enables several advanced functions. First, it directly achieves "rapid mapping between 2D image coordinates and 3D world coordinates," allowing users to directly click on a target (such as a tourist) in any 2D monitoring screen, and the system immediately highlights its precise spatial location in the 3D model, realizing intuitive interaction and "target localization" from image to physical space. Second, this 3D structure provides the ultimate basis for "geometric verification of cross-camera target re-identification," because target candidates from two different cameras are considered valid only if they can be connected in 3D space by a reasonable physical path. Furthermore, this 3D scene model provides spatial context for "intelligent perspective switching and active tracking." Based on the target's position in the 3D model, the system can predict which camera's theoretical field of view it will soon enter, thus allocating resources in advance. Finally, the motion trajectories of all targets are recorded and analyzed in a unified three-dimensional space, which provides the most accurate spatiotemporal data foundation for subsequent "panoramic behavior analysis," making advanced analyses such as calculating the precise moving speed of targets and the aggregation density of analysts in different three-dimensional regions possible.
[0043] Based on the perspective quality assessment model, the system automatically selects the best viewing angle and controls the gimbal camera to actively track the subject.
[0044] As mentioned above, the "viewpoint quality assessment model" is a core decision-making component built into the system. It continuously performs multi-dimensional and quantitative quality assessments of the current images from every available camera in the system (including fixed cameras and pan-tilt cameras). This model comprehensively considers multiple factors affecting observation quality, such as: target integrity (whether the target is of appropriate size in the image and whether it is fully visible), sharpness (whether the target image is clearly focused and free from motion blur), occlusion (whether the target is partially or completely obscured by other objects or pedestrians), and observation angle (whether the camera is in a favorable frontal view or a less favorable side or back view relative to the target). Based on these real-time assessment results, the system generates a global viewpoint quality score. "Automatic selection of the best observation viewpoint" refers to the system dynamically selecting, based on this score and a camera scheduling strategy based on reinforcement learning, the camera view that provides the most comprehensive and clearest target information at the current moment from all available viewpoints. When the selected viewpoint comes from a gimbal camera, the system will further "control the gimbal camera to actively track", that is, send control commands to the gimbal camera to adjust its Pan (horizontal rotation), Tilt (vertical pitch) and Zoom (zoom) parameters, so that the target is always stably in the best observation position in the center of its image, thereby achieving uninterrupted, high-quality continuous tracking.
[0045] For example, in the departure hall of a large airport, the system is tracking a flagged suspicious person. Initially, a fixed camera A, positioned high up with a panoramic view, captures the target, but the distant view results in blurred facial details. At this point, a nearby PTZ camera B receives a higher evaluation score because it can zoom in by adjusting its angle. The system then switches the primary view to camera B and controls its rotation and zoom to clearly capture the person's face. Subsequently, the target walks behind a pillar, creating a brief complete obstruction in camera B's view. The view quality assessment model immediately detects this obstruction event. Almost simultaneously, the system discovers another fixed camera C, located on the other side of the pillar, although with a slightly off-center view, has just captured the target as they step out of the obstruction area. The system then seamlessly switches the primary view to camera C, ensuring the continuity of tracking. Once the target has completely left the pillar area, the system can again adjust the position of PTZ camera B to regain a better frontal view.
[0046] It's important to note that in specific implementation scenarios, the intelligent perspective switching and active tracking functions, building upon the above approach, do not operate in isolation but rather collaborate deeply with other system modules. This function heavily relies on the unified world coordinate system established in the preceding steps. Only when all cameras and targets are on the same 3D map can the system accurately determine which camera "can see" and "how to adjust" to see the target. Simultaneously, cross-camera target re-identification is the fundamental guarantee for maintaining consistency during perspective switching, ensuring that the target is not lost or mistracked when switching cameras. Furthermore, the continuous, high-quality target close-up video stream generated by this module provides richer and more reliable raw data for subsequent panoramic behavior analysis and anomaly detection, especially crucial for recognizing subtle abnormal behaviors such as facial expressions and gesture interactions that require fine observation. Ultimately, this active perception capability enables the entire system to form a closed loop from perception and decision-making to execution, significantly improving the automation and intelligence level of the monitoring system.
[0047] Anomaly detection is achieved by modeling cross-camera behavior using spatiotemporal graph convolutional networks.
[0048] As mentioned above, "modeling cross-camera behavior based on spatiotemporal graph convolutional networks" refers to using a deep learning model specifically designed for processing graph-structured data—graph convolutional networks—and extending it to the spatiotemporal dimension to analyze and understand the complex interactions and dynamic evolution between multiple targets in a surveillance scene. Specifically, the system abstracts the entire surveillance scene as a dynamic graph structure: nodes in the graph represent tracked targets (such as pedestrians and vehicles), and node attributes can include the target's position, speed, and appearance features; edges in the graph represent spatial relationships (such as distance and relative orientation) or social interactions (such as gathering and following) between targets. This graph structure is not static but evolves continuously over time. The "spatiotemporal graph convolutional network" performs convolution operations on this graph simultaneously in both spatial and temporal dimensions: spatial convolution is used to learn the interaction patterns between target nodes and their neighboring nodes at the same moment; temporal convolution is used to learn the motion patterns of individual or group targets over continuous time series. By learning from massive amounts of normal behavior data, this network can construct a baseline model of "normal behavior." The process of "achieving anomaly detection" involves inputting a dynamic graph constructed from real-time monitoring data into the trained model and calculating its deviation from the baseline model of "normal behavior". When the behavior pattern of a target or group (such as abnormal wandering in sensitive areas, sudden reverse movement, rapid gathering in specific areas, etc.) deviates significantly from the learned normal pattern, the system will identify it as an abnormal event and trigger an alarm.
[0049] For example, in a large train station plaza monitoring system, the system continuously tracks hundreds of passengers through the aforementioned steps. A spatiotemporal graph convolutional network constructs and analyzes a dynamic graph composed of these passengers as nodes in real time. Under normal circumstances, the network learns a stable "passenger flow pattern" where passengers move from the entrance to the ticket office, and then to the waiting room or ticket gate. At a certain moment, the system detects the following abnormal graph pattern: in a normally passable area in the center of the plaza, the movement trajectories of five nodes (five people) converge from different directions and stop moving (spatial aggregation and temporal stagnation), forming a persistent "small group." This dynamic graph pattern differs significantly from the normal flow pattern learned by the network, so the system immediately marks this event as "abnormal personnel aggregation" and issues a high-level alarm to the monitoring center, alerting security personnel. Another example is a node (single person) repeatedly performing a "enter and exit" loop near the ticket gate; this repetitive trajectory in time and space is also accurately captured by the model and marked as "abnormal loitering" behavior.
[0050] It should be noted that, in specific implementation scenarios, based on the above solutions, the anomaly detection module based on spatiotemporal graph convolutional networks serves as the top-level application for intelligent analysis in the entire system. It highly relies on the accurate data input provided by all the aforementioned technical modules. The continuous and consistent trajectories generated by cross-camera target detection and re-identification, spanning the entire monitoring area, are the foundation for constructing an accurate dynamic graph. Without seamless cross-camera tracking, node trajectories will be interrupted, and the graph model will be incomplete. The unified world coordinate system obtained through 3D scene reconstruction provides a physical basis for calculating accurate spatial distances and movement speeds between nodes, making the edge attributes (spatial relationships) in the graph more precise. Furthermore, the output of this anomaly detection module can be linked with the "intelligent perspective switching and active tracking" module: once the system detects abnormal behavior (such as clustering) in a certain area at a macroscopic scale, it can immediately instruct nearby PTZ cameras to automatically turn to that area for close-range, high-definition active observation to obtain more detailed on-site information, achieving a closed loop of "macroscopic early warning" and "microscopic confirmation." Ultimately, all these discrete anomalies can be organized on a "behavioral semantic map" to form a coherent narrative about the security situation of the entire monitored area, greatly improving the insight and response efficiency of the monitoring system.
[0051] According to one embodiment of this application, system calibration of multiple cameras includes: An automatic calibration algorithm based on feature points is used to establish the mapping relationship between the image coordinates of each camera and the unified world coordinate system; Design an online calibration compensation mechanism to automatically detect changes in camera position and recalibrate.
[0052] As described above, an automatic calibration algorithm based on feature points is used to establish a mapping relationship between the image coordinates of each camera and a unified world coordinate system. This process extracts feature points from the images acquired by each camera, matches feature point pairs corresponding to the same physical space point in different images, and combines them with preset calibration references or inherent static features in the scene to calculate the internal and external parameters of each camera, thereby constructing an accurate mathematical transformation model from the two-dimensional image plane of each camera to the unified three-dimensional world coordinate system.
[0053] An online calibration compensation mechanism is designed to automatically detect camera position changes and recalibrate. This mechanism continuously monitors the positional stability of feature points in the images captured by each camera during system operation, or analyzes apparent changes in fixed reference objects in the scene, automatically identifying pose shifts caused by factors such as external impacts, loose supports, or thermal expansion and contraction. When a displacement or angular change exceeds a preset threshold, the system automatically triggers a recalibration process, dynamically updating camera parameters using current scene features to ensure the continuous accuracy of the mapping relationship.
[0054] According to one embodiment of this application, time synchronization and preprocessing of data collected by the plurality of cameras include: The PTP precision clock protocol is used to achieve microsecond-level time synchronization, ensuring the time consistency of data acquisition from multiple cameras; Establish an image quality assessment module to automatically identify quality issues such as occlusion, overexposure, and blurring, and trigger corresponding camera parameter adjustments. Design an illumination equalization algorithm to eliminate color inconsistencies caused by differences in exposure parameters between different cameras.
[0055] As mentioned above, the PTP precision clock protocol is used to provide a unified time reference for all cameras in the network, achieving microsecond-level time synchronization. This ensures that each frame of image data captured by each camera has a precise and consistent timestamp, providing accurate time alignment for subsequent cross-camera data association.
[0056] An image quality assessment module is established. This module analyzes image content and automatically identifies common quality issues such as occlusion, overexposure, underexposure, and motion blur. When a quality defect is detected, the module generates control commands, triggering the corresponding camera to automatically adjust its exposure time, gain, or focal length to optimize the imaging effect.
[0057] The design incorporates an illumination equalization algorithm that uses color correction and brightness mapping techniques to normalize the color and brightness differences caused by independent automatic exposure from different cameras. This eliminates the appearance differences caused by inconsistent exposure parameters, thereby ensuring that subsequent processing steps obtain image data with consistent colors.
[0058] According to one embodiment of this application, continuous tracking of a target based on a cross-camera target detection and re-identification algorithm includes: Construct a cascaded target association network, where: The first level performs coarse correlation based on appearance and motion features; The second level is based on graph neural networks to model the spatiotemporal relationships between targets; The third level utilizes three-dimensional position information for geometric verification; An incremental learning mechanism is designed to update the re-identification model online, adapting to changes in the target's appearance.
[0059] As described above, a cascaded target association network is constructed, which adopts a three-level progressive association strategy: The first level of association utilizes both appearance features extracted by the improved ResNet network and motion features extracted by optical flow to perform preliminary screening of cross-camera target candidates, completing coarse-grained association; the second level of association uses graph neural networks to model the complex spatiotemporal constraints between targets, and solves the problem of distinguishing similar-looking targets by analyzing the movement patterns of targets in the camera network; the third level of association introduces the world coordinate positions of targets obtained through 3D scene reconstruction to perform geometric consistency verification and eliminate erroneous associations that are impossible to exist in physical space.
[0060] Meanwhile, an incremental learning mechanism is designed. This mechanism continuously collects new target appearance data during system operation, dynamically adjusts and updates the parameters of the re-identification model, and enables the model to adapt to changes in the appearance features of the target caused by factors such as changes in lighting and perspective during tracking.
[0061] According to one embodiment of this application, real-time reconstruction of the three-dimensional scene structure of a monitoring area based on multi-view geometry principles includes: Design a lightweight SLAM algorithm to achieve real-time scene reconstruction and target localization on edge computing nodes; Establish a fast mapping relationship between two-dimensional image coordinates and three-dimensional world coordinates, and support the ability to locate the target's position in actual space by clicking on the image.
[0062] As described above, a lightweight SLAM algorithm is designed. This algorithm extracts and matches feature points in the scene by fusing synchronous video streams acquired by multiple cameras, and performs camera pose estimation and scene 3D point cloud reconstruction in parallel on edge computing nodes, thereby realizing real-time 3D scene structure reconstruction and accurate positioning of moving targets in the monitored area.
[0063] Establish a fast mapping relationship between two-dimensional image coordinates and three-dimensional world coordinates. Based on the intrinsic and extrinsic parameters obtained from system calibration and the real-time reconstructed three-dimensional scene, construct a bidirectional projection model from any camera image plane to a unified world coordinate system. This allows users to obtain the three-dimensional position coordinates of the corresponding target in the actual physical space by clicking on any pixel position in the monitoring screen and calculating the coordinate transformation.
[0064] According to one embodiment of this application, automatically selecting the optimal viewing angle based on a viewing angle quality assessment model includes: Construct a viewpoint quality assessment model to evaluate the target integrity, sharpness, and occlusion level in real time for each camera's image; Design a camera scheduling strategy based on reinforcement learning to automatically select the best viewing angle and control the gimbal camera for active tracking; When the target enters the blind spot or leaves the current camera's field of view, it automatically switches to the camera that can continue tracking.
[0065] As mentioned above, a viewpoint quality assessment model is constructed. This model analyzes the real-time images of each camera and performs quantitative assessment from three dimensions: target integrity, image sharpness, and occlusion degree. Target integrity assesses the size ratio and complete visibility of the target in the image, image sharpness assesses the focus and motion blur of the image, and occlusion degree assesses the proportion of the target that is obscured by obstacles.
[0066] We design a camera scheduling strategy based on reinforcement learning. This strategy models viewpoint selection as a sequential decision problem and obtains the optimal scheduling strategy by continuously interacting with the environment. It automatically selects the observation viewpoint with the highest comprehensive score from the available cameras and issues PTZ control commands to the gimbal camera to keep it in the best tracking state of the target.
[0067] An automatic viewpoint switching mechanism is established. When the system detects that a target has entered the current camera's blind spot or is about to leave the effective monitoring area, it immediately selects a camera from the candidate cameras that can continue to track the target based on target motion prediction and camera coverage analysis, thus achieving seamless switching of the tracking viewpoint.
[0068] According to one embodiment of this application, anomaly detection is achieved by modeling cross-camera behavior based on a spatiotemporal graph convolutional network, including: Design a panoramic anomaly detection algorithm to identify cross-regional anomaly patterns such as crowd gathering, abnormal loitering, and rapid movement; Establish a behavioral semantic map to organize discrete monitoring events into a coherent behavioral narrative.
[0069] As described above, a panoramic anomaly detection algorithm is designed. This algorithm extracts the spatiotemporal features of target trajectories across cameras based on a spatiotemporal graph convolutional network. By analyzing the distribution density, movement trajectory patterns, and movement speed characteristics of people in the monitored area, it identifies abnormal behavior patterns across camera areas, such as people gathering, abnormal loitering, and rapid movement.
[0070] A behavioral semantic map is established, which organizes monitoring events scattered across different cameras and at different times according to their spatiotemporal correlation in a unified world coordinate system. Through event clustering and causal relationship analysis, a semantically coherent behavioral sequence description is formed, completing the transformation from low-level visual features to high-level behavioral narratives.
[0071] According to one embodiment of this application, it also includes: It adopts an edge-cloud collaborative computing architecture, which completes real-time data processing at the edge nodes and performs in-depth analysis in the cloud; Design an intelligent video storage strategy that adaptively adjusts storage quality and duration based on event importance; By employing video summarization technology, timelines of key events are automatically generated, reducing storage space and retrieval time.
[0072] As mentioned above, by adopting an edge-cloud collaborative computing architecture, high real-time tasks such as target detection, data synchronization, and 3D reconstruction are deployed on edge computing nodes for processing, while computationally intensive tasks such as behavior analysis and model training are arranged to be completed in the cloud, thereby achieving a reasonable allocation of computing resources.
[0073] The design employs an intelligent video storage strategy that automatically adjusts storage parameters based on the importance of monitored events. High-value content, such as abnormal events, is stored at a high bitrate for a long duration, while ordinary monitoring footage is stored at a low bitrate for a short duration, thus achieving adaptive optimization of storage resources.
[0074] By employing video summarization technology, keyframes and event segments are extracted to automatically generate highly condensed timeline videos centered on events, significantly reducing the storage space occupied by redundant data and improving the efficiency of subsequent video retrieval and playback.
[0075] A second aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method described in any of the embodiments of the first aspect above.
[0076] Figure 2 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 2 As shown, the electronic device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840, wherein the processor 810, the communication interface 820, and the memory 830 communicate with each other via the communication bus 840. The processor 810 may call logical instructions in the memory 830 to execute the method in any of the embodiments of the first aspect described above, the method including: Perform system calibration on multiple cameras and establish the mapping relationship between the image coordinates of each camera and a unified world coordinate system; The data collected by the multiple cameras are synchronized in time and preprocessed. Based on cross-camera target detection and re-identification algorithms, the target is continuously tracked; Based on the principle of multi-view geometry, the three-dimensional scene structure of the monitored area is reconstructed in real time. Based on the viewpoint quality assessment model, the best viewing angle is automatically selected and the gimbal camera is controlled to actively track the view. Anomaly detection is achieved by modeling cross-camera behavior using spatiotemporal graph convolutional networks.
[0077] Furthermore, the logical instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.
[0078] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program, the computer program being able to be stored on a non-transitory computer-readable storage medium, and when the computer program is executed by a processor, the computer being able to perform the methods provided by the above methods, the method comprising: Perform system calibration on multiple cameras and establish the mapping relationship between the image coordinates of each camera and a unified world coordinate system; The data collected by the multiple cameras are synchronized in time and preprocessed. Based on cross-camera target detection and re-identification algorithms, the target is continuously tracked; Based on the principle of multi-view geometry, the three-dimensional scene structure of the monitored area is reconstructed in real time. Based on the viewpoint quality assessment model, the best viewing angle is automatically selected and the gimbal camera is controlled to actively track the view. Anomaly detection is achieved by modeling cross-camera behavior using spatiotemporal graph convolutional networks.
[0079] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the methods provided by the above methods, the method comprising: Perform system calibration on multiple cameras and establish the mapping relationship between the image coordinates of each camera and a unified world coordinate system; The data collected by the multiple cameras are synchronized in time and preprocessed. Based on cross-camera target detection and re-identification algorithms, the target is continuously tracked; Based on the principle of multi-view geometry, the three-dimensional scene structure of the monitored area is reconstructed in real time. Based on the viewpoint quality assessment model, the best viewing angle is automatically selected and the gimbal camera is controlled to actively track the view. Anomaly detection is achieved by modeling cross-camera behavior using spatiotemporal graph convolutional networks.
[0080] Example 2 This intelligent panoramic monitoring system was deployed in the departure hall of an airport. The hall, approximately 8,000 square meters in area, has an elliptical structure and a glass curtain wall ceiling, resulting in significant variations in natural lighting conditions over time. The system comprises 48 cameras: 32 fixed-view 4K cameras, 12 PTZ cameras, and 4 180-degree fisheye cameras, providing comprehensive coverage of the entire hall.
[0081] System initialization and calibration phase: Upon system startup, multi-camera system calibration is performed first. Technicians set up 12 calibration boards of known dimensions on the floor of the departure hall, distributed across various key areas. The system runs an automatic calibration algorithm based on ORB feature points, establishing the correspondence between feature points by extracting corner features from the images of each camera.
[0082] The calibration process uses the following mathematical model: Suppose a point in the world coordinate system The image coordinates (u,v) are mapped to the camera projection model:
[0083] Where K is the camera intrinsic parameter matrix, [R|t] is the extrinsic parameter matrix, s is the scale factor, R is the rotation matrix, and t is the translation vector.
[0084] The system optimizes parameters by minimizing reprojection error:
[0085] in Here are the observed image coordinates, and proj is the projection function. Let K be a point in the world coordinate system, and let K, R, t be the camera's intrinsic and extrinsic parameters.
[0086] The online calibration compensation mechanism continuously monitors the positional changes of fixed feature points, and automatically triggers recalibration when a pixel-level offset exceeding 2 pixels is detected.
[0087] Data synchronization and preprocessing The system uses the IEEE 1588 PTP protocol for time synchronization, with clock deviations between cameras controlled within 100 microseconds. The image quality assessment module is implemented based on a deep convolutional network, and its evaluation function is: Q = w1·C + w2·S - w3·O Where C is the target integrity score, S is the sharpness score, O is the degree of occlusion, and w1, w2, and w3 are weighting coefficients.
[0088] When the evaluation score Q falls below the threshold of 0.7, the system automatically adjusts the camera parameters. The illumination equalization algorithm employs an improved method based on Retinex theory, eliminating illumination unevenness through multi-scale Gaussian convolution: R(x,y) = log(I(x,y)) - log(F(x,y)*I(x,y)) Where I is the original image, F is the Gaussian kernel function, and R is the enhanced image.
[0089] Target tracking and 3D reconstruction: The system employs an improved YOLOv5 model for object detection, combined with the DeepSORT algorithm for single-camera tracking. Cross-camera re-identification uses a Siamese network architecture based on ResNet-50, with the feature extraction function as follows: f = Φ(I;θ) Where Φ represents the feature extraction network, and θ represents the network parameters.
[0090] Three-level verification mechanism for cascaded target association networks: Level 1: Calculate the cosine similarity of appearance features S_a and the Euclidean distance of motion features D_m; Level 2: The graph neural network updates node features through a message passing mechanism.
[0091] Let i be the feature vector of node i in the l-th layer. For activation function, Let i be the set of neighboring nodes. For attention weights, Let be the learnable weight matrix of the ll-th layer.
[0092] Level 3: Geometric verification judges the rationality of the association through reprojection error; The 3D reconstruction employs a lightweight, improved algorithm based on ORB-SLAM3, generating dense point cloud maps in real time on an edge server. The system processes 15 frames of 1280×720 resolution images per second, achieving a reconstruction accuracy of 5cm.
[0093] Intelligent perspective scheduling: The viewpoint quality assessment model uses a multi-task learning framework and outputs simultaneously: Target integrity score: based on the relative position of the target bounding box to the image boundary; Sharpness rating: based on image gradient magnitude statistics; Occlusion score: based on the occlusion ratio of instance segmentation results; The camera scheduling strategy based on reinforcement learning adopts the DQN algorithm. The state space s includes 28-dimensional features such as target position and quality scores of each camera, the action space a represents the camera selection instructions, and the reward function is designed as follows: r = w1·Q + w2·T - w3·S Where Q represents the view quality, T represents the tracking continuity, and S represents the switching frequency.
[0094] Abnormal behavior detection: The spatiotemporal graph convolutional network adopts the following architecture: Spatial graph convolution: using Chebyshev polynomial approximation graph filters; Temporal convolution: using a one-dimensional temporal convolution kernel; The anomaly detection model learns normal behavior patterns through a variational autoencoder, and the anomaly score is calculated as follows:
[0095] Where A is the anomaly score, x is the input data, E is the encoder, and D is the decoder.
[0096] The behavioral semantic map is constructed using knowledge graph technology, which represents monitored events as RDF triples and generates behavioral narratives through graph reasoning algorithms.
[0097] System performance: In actual deployment, the system successfully detected multiple abnormal events: An unusual gathering of people (more than 5 people staying for more than 3 minutes) was detected in Area B of Terminal 1. The system identifies abnormal behavior of traveling in the wrong direction at the security checkpoint. Unattended luggage was found in the waiting area for more than 10 minutes.
[0098] Example 3 This system was deployed in the intelligent logistics and warehousing center of an e-commerce company. The warehousing center has a building area of approximately 20,000 square meters, a floor height of 12 meters, and adopts a three-dimensional racking design with rack heights reaching 10 meters. The warehouse simultaneously operates 120 AGVs (Automated Guided Vehicles), 50 pickers, and 20 stacker cranes. The system deploys 86 cameras, including 40 fixed panoramic cameras, 30 PTZ cameras, and 16 thermal imaging cameras, achieving comprehensive coverage of the warehouse area.
[0099] System initialization and stereo calibration: For high-rack environments, the system employs a stereo calibration method. Reflective markers are placed on each layer of the rack, and the three-dimensional coordinates of each marker are accurately measured using a total station. The system then runs a stereo calibration algorithm based on SIFT feature points to establish the mapping relationship between each camera in three-dimensional space.
[0100] The calibration process employs bundle adjustment optimization.
[0101] in Let i be the image coordinates of feature point i in camera j. Let i be the world coordinates of feature point i. Let J be the intrinsic parameter matrix of camera j. Let i be the rotation matrix and translation vector of camera j, and let i represent the feature point index and j represent the camera index. This is the projection function.
[0102] To address the vibrations caused by AGV operation, the system is designed with an online compensation mechanism based on IMU data. A miniature inertial measurement unit is installed in each camera to monitor camera attitude changes in real time. When an angular deviation exceeding 0.1 degrees is detected, a calibration update is automatically triggered.
[0103] Multi-source data synchronization and quality control: The system employs an improved PTP protocol, achieving synchronization accuracy of 50 microseconds in an industrial Ethernet environment. Considering the characteristics of the logistics environment, the image quality assessment module pays particular attention to the following indicators: Motion blur detection: based on Laplacian variance calculation; Shelf occlusion assessment: semantic segmentation based on deep learning; Illumination adaptability: based on adaptive histogram equalization; The data fusion between thermal imaging cameras and visible light cameras uses a registration algorithm based on depth information: I_fused = α·I_visible + β·I_thermal Where α and β are adaptive weighting coefficients based on depth information, I_fused is the fused image, I_visible is the visible light image, and I_thermal is the thermal image.
[0104] Multi-target tracking and stereo positioning: The system employs a multimodal target detection network, simultaneously processing visible light and thermal imaging data. Specific feature extractors are designed for different target types. AGV: Based on its external shape and QR code identification; Personnel: Based on skeletal key points and thermal imaging features; Goods: Based on the texture and size characteristics of the packaging box; Introducing a spatiotemporal constraint matrix for cross-target association: A = [a_ij] = w1·S_appearance + w2·S_motion + w3·S_spatial Where S_spatial is calculated based on the geometric consistency of the 3D reconstruction results, A is the correlation matrix, S_appearance is the appearance similarity, S_motion is the motion similarity (based on the consistency of velocity direction), S_spatial is the spatial similarity, and w1, w2, w3 are the weight coefficients.
[0105] The stereo scene reconstruction uses a stereo vision-based SLAM algorithm, which calculates depth information through epipolar geometric constraints. d = f·B / (x_l - x_r) Where f is the focal length, B is the baseline distance, and x_l and x_r are the parallaxes of the left and right images.
[0106] Intelligent perspective scheduling and collaborative monitoring: The perspective quality assessment model is optimized for logistics scenarios: AGV tracking: Prioritize viewing angles where QR codes can be clearly identified; Personnel monitoring: Focus on assessing the observability of behavior; Inventory counting: Pay attention to the visibility of shelf shelves; The scheduling strategy based on multi-agent reinforcement learning treats each PTZ camera as an independent agent, coordinating their behavior through a centralized training and distributed execution framework. The reward function is designed as follows: R = Σ(λ_i·Q_i) - μ·C_switch Where Q_i is the quality score of each task, and C_switch is the switching cost.
[0107] Logistics anomaly detection and process monitoring: Spatiotemporal graph convolutional networks are used for modeling logistics processes: Node types: AGV, personnel, shelves, workstations; Edge relationships: distance relationships, task associations, and process order; Anomaly detection models primarily identify the following patterns: AGV path deviation: based on the Hausdorff distance between the planned path and the actual trajectory; Picking errors: based on the spatiotemporal consistency between action sequences and the location of goods; Personnel Detention: Analysis based on time spent in hazardous areas; The behavioral semantic map is constructed using a business process modeling language, which maps monitoring events into standardized logistics operation sequences.
[0108] For any parts not mentioned in this application, existing technologies may be used or referenced.
[0109] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.
[0110] The above description is merely an embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of this application should be included within the scope of the claims of this application.
Claims
1. A multi-camera collaborative blind-spot-free intelligent monitoring method, characterized in that, include: Perform system calibration on multiple cameras and establish the mapping relationship between the image coordinates of each camera and a unified world coordinate system; The data collected by the multiple cameras are synchronized in time and preprocessed. Based on cross-camera target detection and re-identification algorithms, the target is continuously tracked; Based on the principle of multi-view geometry, the three-dimensional scene structure of the monitored area is reconstructed in real time. Based on the viewpoint quality assessment model, the best viewing angle is automatically selected and the gimbal camera is controlled to actively track the view. Anomaly detection is achieved by modeling cross-camera behavior using spatiotemporal graph convolutional networks.
2. The method according to claim 1, characterized in that, System calibration of multiple cameras, including: An automatic calibration algorithm based on feature points is used to establish the mapping relationship between the image coordinates of each camera and the unified world coordinate system; Design an online calibration compensation mechanism to automatically detect changes in camera position and recalibrate.
3. The method according to claim 1, characterized in that, The data collected by the multiple cameras is synchronized in time and preprocessed, including: The PTP precision clock protocol is used to achieve microsecond-level time synchronization, ensuring the time consistency of data acquisition from multiple cameras; Establish an image quality assessment module to automatically identify quality issues such as occlusion, overexposure, and blurring, and trigger corresponding camera parameter adjustments. Design an illumination equalization algorithm to eliminate color inconsistencies caused by differences in exposure parameters between different cameras.
4. The method according to claim 1, characterized in that, Continuous target tracking based on cross-camera target detection and re-identification algorithms, including: Construct a cascaded target association network, where: The first level performs coarse correlation based on appearance and motion features; The second level is based on graph neural networks to model the spatiotemporal relationships between targets; The third level utilizes three-dimensional position information for geometric verification; An incremental learning mechanism is designed to update the re-identification model online, adapting to changes in the target's appearance.
5. The method according to claim 1, characterized in that, Real-time reconstruction of the 3D scene structure of the monitored area based on multi-view geometric principles, including: Design a lightweight SLAM algorithm to achieve real-time scene reconstruction and target localization on edge computing nodes; Establish a fast mapping relationship between two-dimensional image coordinates and three-dimensional world coordinates, and support the ability to locate the target's position in actual space by clicking on the image.
6. The method according to claim 1, characterized in that, The optimal viewing angle is automatically selected based on the viewing angle quality assessment model, including: Construct a viewpoint quality assessment model to evaluate the target integrity, sharpness, and occlusion level in real time for each camera's image; Design a camera scheduling strategy based on reinforcement learning to automatically select the best viewing angle and control the gimbal camera for active tracking; When the target enters the blind spot or leaves the current camera's field of view, it automatically switches to the camera that can continue tracking.
7. The method according to claim 1, characterized in that, Anomaly detection is achieved by modeling cross-camera behavior using spatiotemporal graph convolutional networks, including: Design a panoramic anomaly detection algorithm to identify cross-regional anomaly patterns such as crowd gathering, abnormal loitering, and rapid movement; Establish a behavioral semantic map to organize discrete monitoring events into a coherent behavioral narrative.
8. The method according to claim 1, characterized in that, Also includes: It adopts an edge-cloud collaborative computing architecture, which completes real-time data processing at the edge nodes and performs in-depth analysis in the cloud; Design an intelligent video storage strategy that adaptively adjusts storage quality and duration based on event importance; By employing video summarization technology, timelines of key events are automatically generated, reducing storage space and retrieval time.
9. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method as described in any one of claims 1-8.
10. An electronic device comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method as described in any one of claims 1-8.
Citation Information
Cited By
Intelligent monitoring method and system based on panorama
CN122069430A
Multi-camera cooperative positioning method based on graph model
CN122115577A
A multi-camera cooperative positioning method based on a graph model
CN122115577B
A Large-Model Traffic Accident Analysis Method Integrating Blind Spot Completion and Physics-Driven Reasoning
CN122416740A