Multi-camera cooperative monitoring method and system, electronic equipment and medium
By using a multi-camera collaborative 3D monitoring method, a 3D coordinate system for the substation is established, the grid cells of the overlapping view area are calculated and the target outline is constructed, and the change rate of the target outline is monitored in real time. This solves the problems of blind spots and misjudgments in the monitoring of multiple cameras working independently, and realizes accurate identification and rapid response of substation video monitoring.
Patent Information
- Application Number
- CN202511305971.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-12
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-09-12
AI Technical Summary
In existing substation video surveillance, the independent operation of multiple cameras makes it difficult to effectively integrate image information, resulting in blind spots, difficulty in accurately obtaining spatial location information, and easy to make monitoring misjudgments or omissions, especially in complex scenarios.
By establishing a collaborative 3D coordinate system with multiple cameras, the grid cells of the overlapping viewpoint area are calculated, depth values are extracted, and the target contour is constructed. The rate of change of the number of target contours is monitored in real time, and abnormal events are identified.
It improves the accuracy and security of substation video surveillance, can quickly identify abnormal behavior in complex scenarios, reduce monitoring misjudgments, and enhance security monitoring capabilities during peak hours.
Smart Images

Figure CN120897037A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of camera control, specifically to a multi-camera collaborative monitoring method, system, electronic device, and medium. Background Technology
[0002] With the continuous advancement of smart grid construction, substations, as crucial hubs in the power system, are directly related to the reliability of power supply through their safe operation. Currently, substations commonly employ video surveillance systems for safety management, deploying cameras at various locations within the substation to monitor equipment operation status and personnel activities in real time.
[0003] In existing technologies, video surveillance in substations typically employs multiple cameras operating independently. Each camera is responsible for monitoring its designated area and transmitting the captured video footage to a monitoring center. Monitoring personnel then observe the two-dimensional images transmitted from different cameras to determine if any abnormalities exist within the substation.
[0004] However, this independent monitoring method has certain limitations. Since each camera independently acquires and processes images, when the monitored object is located in the overlapping shooting area of multiple cameras, the image information from different cameras is difficult to effectively fuse, easily creating blind spots. Furthermore, relying solely on two-dimensional images for monitoring makes it difficult to accurately obtain the spatial location information of the monitored object, especially in complex scenarios, easily leading to misjudgments or missed detections. Summary of the Invention
[0005] This application provides a multi-camera collaborative monitoring method, system, electronic device, and medium, which can improve the accuracy of video monitoring in substations.
[0006] The first aspect of this application provides a multi-camera collaborative monitoring method, specifically including: Acquire two-dimensional images in real time from multiple cameras set up at different angles in the substation, and record the spatial position parameters and shooting angle parameters of each camera; Based on the spatial position parameters and shooting angle parameters of each camera, a three-dimensional coordinate system of the substation is established, and the two-dimensional images captured by each camera are mapped to the three-dimensional coordinate system. In the three-dimensional coordinate system, the overlapping area of each viewpoint and the two-dimensional image of the adjacent viewpoint is calculated respectively, and the overlapping area is divided into multiple grid cells of equal size; For each of the aforementioned grid cells, grid images captured by different cameras within the grid cell are extracted, and the 3D reconstruction uncertainty of each grid cell is calculated based on the aforementioned grid images. When there is a target overlapping region where the average 3D reconstruction uncertainty of all grid cells exceeds a preset threshold, the two adjacent viewpoints corresponding to the target grid cell with the largest 3D reconstruction uncertainty in the target overlapping region are selected as the cooperative observation viewpoints. The two-dimensional image corresponding to the collaborative observation perspective is subjected to perspective transformation so that the two-dimensional images of the two perspectives of the collaborative observation perspective are aligned in the three-dimensional coordinate system to obtain the depth value of the target grid cell. The target contour within the target grid cell is constructed based on the depth value, where the depth value represents the actual distance from the object to the camera. The rate of change of the number of target contours within each target grid cell is statistically analyzed in real time. Based on the rate of change and the target contours, the abnormal region where the abnormal event occurred is determined.
[0007] By employing the aforementioned technical solution, multiple cameras are deployed in the substation, and a three-dimensional coordinate system is established based on the spatial position parameters and shooting angle parameters of the cameras. By calculating the overlapping area of images from adjacent viewpoints and dividing it into grid cells, and by calculating the 3D reconstruction uncertainty of each grid cell, the system can accurately and quantitatively identify areas with the worst observation quality due to occlusion or other reasons. Once such problematic areas are identified, the system targets and selects the grid cell with the greatest uncertainty, and calls upon its corresponding optimal viewing angle combination for depth calculation. Simultaneously, by calculating the depth value of the target grid cell to construct the target contour, and by monitoring the rate of change of the number of target contours in real time, the system can quickly and accurately identify abnormal behavior in crowds. This multi-view collaborative 3D monitoring solution significantly improves the safety monitoring capabilities of the substation during peak hours, providing strong technical support for the timely detection and handling of safety hazards. This solution not only overcomes the limitations of existing technologies in terms of relatively independent image information processing, but also achieves accurate identification of abnormal events in the substation through the introduction of depth values, thereby effectively improving the accuracy of substation video monitoring.
[0008] Optionally, calculating the 3D reconstruction uncertainty of each mesh cell based on each mesh image includes: Extract the depth value of each pixel in the grid image, map the location region corresponding to each grid unit to the three-dimensional coordinate system, and obtain the depth value distribution of the location region corresponding to each grid unit; Based on the depth value distribution of the location region corresponding to each grid cell, the depth correlation degree between each grid cell and the corresponding location region is determined; The mesh cells with a depth correlation degree less than a preset correlation threshold are marked as overlapping mesh cells, and the distribution standard deviation of the overlapping mesh cells in the overlapping region is calculated to obtain the three-dimensional reconstruction uncertainty of each mesh cell.
[0009] By employing the aforementioned technical solution, the depth values of grid image pixels are extracted and mapped to a 3D coordinate system, allowing the system to obtain the actual depth distribution of each grid cell in space. Calculating depth correlation based on the depth value distribution effectively identifies viewpoint overlap caused by operators and equipment. By marking grid cells with depth correlation below a preset threshold as overlapping grid cells and calculating their distribution standard deviation, the system can accurately quantify the uncertainty of 3D reconstruction in different regions. This depth-based measurement method provides reliable data support for subsequently selecting the optimal collaborative observation viewpoint, thereby improving the system's monitoring accuracy in scenarios involving operator and equipment interaction. The introduction of depth values enables the system to better understand and analyze the distribution characteristics of operators and equipment in 3D space, laying the foundation for the timely detection of abnormal events.
[0010] Optionally, constructing the target contour within the target mesh cell based on the depth value includes: Extract the depth extreme points within the target grid cell from the depth value, determine the depth change rate based on the depth value, and select the region with the depth change rate higher than a preset depth change rate threshold as the candidate target region; Starting from the depth extreme point of the candidate target region, the gradient trajectory with decreasing depth change rate is recorded as it expands outward. Multiple inflection points are determined on the gradient trajectory, and the coordinate distance from each inflection point to the depth extremum point is calculated. The depth change rate difference and Euclidean distance between each inflection point and its adjacent inflection points are determined. When there is an abnormal inflection point where the ratio of the depth change rate difference to the Euclidean distance is greater than a preset ratio threshold and the coordinate distance is greater than a preset coordinate distance threshold, the abnormal inflection point is removed, and the pixels corresponding to the remaining inflection points are connected to form the target contour.
[0011] By employing the aforementioned technical solution, extracting depth extrema from depth values and analyzing the depth change rate, the system can effectively identify areas where human beings or equipment may exist. Using a diffusion strategy starting from depth extrema and combining it with the gradient trajectory of the depth change rate, the system can adaptively track the boundary features of the target contour. By determining inflection points on the gradient trajectory and analyzing their spatial distribution characteristics, combined with the ratio of the depth change rate difference to the Euclidean distance, the system can effectively eliminate abnormal inflection points caused by occlusion or environmental noise. This target contour construction method based on multi-dimensional features not only improves the accuracy of contour extraction but also effectively addresses the partial occlusion problem in scenarios involving operator and equipment interaction. By retaining inflection points with reasonable spatial distribution and depth change characteristics to construct the target contour, the system achieves precise monitoring of operators and equipment in substations, providing reliable basic data for subsequent anomaly event identification.
[0012] Optionally, after constructing the target contour within the target mesh cell based on the depth value, the method further includes: Based on the depth extreme points of each target contour, determine the contour boundary candidate points within the target grid cell for each target contour; Calculate the gradient direction angle between each of the contour boundary candidate points and each depth extreme point, and connect the contour boundary candidate points whose gradient direction angle is less than a preset angle threshold to form a boundary line; The target contour is divided into multiple target contours according to the dividing line, and each target contour is assigned a unique identifier.
[0013] By employing the aforementioned technical solution, the system analyzes the depth extreme points of the target contour to determine candidate contour boundary points. Combined with the constraint of the gradient direction angle, the system can accurately identify the true target contour boundary location. Using a boundary line connection strategy based on the gradient direction angle effectively distinguishes overlapping target contours caused by personnel and equipment. By assigning a unique identifier to each segmented target contour, the system achieves precise tracking and positioning of individual targets among operators and equipment. This contour segmentation method based on depth values and gradient features significantly improves the system's ability to identify individual targets in scenarios involving operator and equipment interaction, providing a more granular monitoring method for accurately detecting abnormal behavior in crowds, thereby further enhancing the safety monitoring effect of substations.
[0014] Optionally, determining the candidate boundary points of each target contour within the target mesh cell based on the depth extreme points of each target contour includes: For multiple target contours within a target grid cell, based on the Euclidean distance between each depth extreme point in each target contour, depth extreme points with an Euclidean distance less than a preset spatial distance threshold are grouped into the same group, and the target contour corresponding to each group of depth extreme points is marked as the contour to be separated. Construct a depth transition zone between depth extrema in each group of contours to be separated, the width of which is determined by the depth difference between adjacent depth extrema. Calculate the first and second derivatives of each depth value within the depth transition zone, and determine the set of points whose absolute value of the second derivative is greater than a preset boundary threshold as candidate contour boundary points.
[0015] By employing the aforementioned technical solution and analyzing the Euclidean distance between depth extrema points in the target contour, the system can adaptively identify overlapping contours that need to be separated. A transition zone construction strategy based on depth difference enables the system to accurately capture the depth variation characteristics between overlapping target contours. By calculating the first and second derivatives of the depth values within the depth transition zone and combining this with preset boundary thresholds for filtering, the system can precisely locate the key positions of contour boundaries. This boundary point extraction method, which combines spatial distance and depth variation characteristics, not only improves the accuracy of overlapping target contour separation but also effectively addresses varying degrees of overlap between personnel and equipment. Through refined analysis of depth values, this solution provides a reliable boundary basis for subsequent contour segmentation, thereby enhancing the system's ability to identify individual targets in scenarios involving interaction between operators and equipment.
[0016] Optionally, determining the abnormal region where the abnormal event occurred based on the rate of change and the target contour includes: An abnormal behavior feature database is established, and the rate of change and the target contour are compared with the abnormal behavior feature database to obtain the comprehensive abnormal probability of the target grid cell. When the overall anomaly probability exceeds a preset anomaly probability threshold, the target grid cell is identified as the anomaly region.
[0017] By adopting the above technical solution, an abnormal behavior feature database is established, and the real-time monitored change rate and target contour features are compared with it, enabling the system to comprehensively assess the abnormal situation of the target area. The comprehensive anomaly probability assessment mechanism considers not only the dynamic changes in the number of target contours but also their spatial morphological features, allowing the system to more accurately identify abnormal events. By setting reasonable anomaly probability thresholds, the system can effectively reduce the false alarm rate while ensuring the timely detection of genuine abnormal events. This feature database-based anomaly event identification method enables intelligent monitoring of personnel and equipment in substations, providing reliable early warning information for safety management personnel, thereby effectively improving the safety management level of substations.
[0018] Optionally, the abnormal behavior feature database includes target contour quantity mutation features, target contour relative position features, and target contour motion features. The step of comparing the rate of change and the target contour with the abnormal behavior feature database to obtain the comprehensive abnormal probability of the target mesh cell includes: Based on the rate of change of the number of target contours within each target grid cell and the abrupt change characteristics of the number of target contours, calculate the first abnormal probability of the rate of change of the number of target contours within the target grid cell; Based on the relative positional relationship and coordinate change of each target contour within the target mesh cell in the three-dimensional coordinate system, as well as the relative positional features and motion features of the target contour, the second anomaly probability of the target contour behavior within the target mesh cell is calculated. The weighted sum of the first anomaly probability and the second anomaly probability is calculated to obtain the comprehensive anomaly probability of the target grid cell.
[0019] By employing the aforementioned technical solution, the abnormal behavior feature database is subdivided into three dimensions: target contour quantity abrupt change features, relative position features, and motion features. This allows the system to more comprehensively characterize the feature patterns of abnormal events. The first abnormal probability calculation based on the target contour quantity change rate can effectively identify abnormal scenarios such as violent crowd gathering or dispersal. By analyzing the relative positional relationship and coordinate changes of the target contours in three-dimensional space, and combining the relative position features and motion features to calculate the second abnormal probability, the system can accurately capture abnormal behaviors with specific spatial motion characteristics. The use of a weighted fusion strategy to calculate the comprehensive abnormal probability not only balances the contributions of different feature dimensions but also improves the robustness of abnormal behavior identification. This multi-dimensional feature-based abnormal behavior assessment method significantly enhances the system's ability to identify different types of abnormal events, providing more reliable and accurate technical support for substation safety monitoring.
[0020] A second aspect of this application provides a multi-camera collaborative monitoring system, specifically including: The data acquisition module is used to acquire two-dimensional images in real time from multiple cameras set at different angles in the substation, and record the spatial position parameters and shooting angle parameters of each camera. The target contour construction module is used to establish a three-dimensional coordinate system of the substation based on the spatial position parameters and shooting angle parameters of each camera, and to map the two-dimensional images captured by each camera into the three-dimensional coordinate system; In the three-dimensional coordinate system, the overlapping area of each viewpoint and the two-dimensional image of the adjacent viewpoint is calculated respectively, and the overlapping area is divided into multiple grid cells of equal size; For each of the aforementioned grid cells, grid images captured by different cameras within the grid cell are extracted, and the 3D reconstruction uncertainty of each grid cell is calculated based on the aforementioned grid images. When there is a target overlapping region where the average 3D reconstruction uncertainty of all grid cells exceeds a preset threshold, the two adjacent viewpoints corresponding to the target grid cell with the largest 3D reconstruction uncertainty in the target overlapping region are selected as the cooperative observation viewpoints. The two-dimensional image corresponding to the collaborative observation perspective is subjected to perspective transformation so that the two-dimensional images of the two perspectives of the collaborative observation perspective are aligned in the three-dimensional coordinate system to obtain the depth value of the target grid cell. The target contour within the target grid cell is constructed based on the depth value, where the depth value represents the actual distance from the object to the camera. The abnormal region determination module is used to statistically analyze the rate of change of the number of target contours within each target grid cell in real time, and determine the abnormal region where an abnormal event has occurred based on the rate of change and the target contours.
[0021] By employing the aforementioned technical solution, multiple cameras are deployed in the substation, and a three-dimensional coordinate system is established based on the spatial position parameters and shooting angle parameters of the cameras. By calculating the overlapping area of images from adjacent viewpoints and dividing it into grid cells, and by calculating the 3D reconstruction uncertainty of each grid cell, the system can accurately and quantitatively identify areas with the worst observation quality due to occlusion or other reasons. Once such problematic areas are identified, the system targets and selects the grid cell with the greatest uncertainty, and calls upon its corresponding optimal viewing angle combination for depth calculation. Simultaneously, by calculating the depth value of the target grid cell to construct the target contour, and by monitoring the rate of change of the number of target contours in real time, the system can quickly and accurately identify abnormal behavior in crowds. This multi-view collaborative 3D monitoring solution significantly improves the safety monitoring capabilities of the substation during peak hours, providing strong technical support for the timely detection and handling of safety hazards. This solution not only overcomes the limitations of existing technologies in terms of relatively independent image information processing, but also achieves accurate identification of abnormal events in the substation through the introduction of depth values, thereby effectively improving the accuracy of substation video monitoring.
[0022] A third aspect of this application provides an electronic device including a processor, a memory, a user interface, and a network interface, wherein the memory is used to store instructions, the user interface and the network interface are both used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described in any of the foregoing.
[0023] A fourth aspect of this application provides a computer-readable storage medium storing instructions that, when executed, perform the method described in any of the preceding descriptions. Attached Figure Description
[0024] Figure 1 This is an exemplary system architecture diagram of a multi-camera collaborative monitoring system provided in an embodiment of this application; Figure 2 This is a flowchart illustrating a multi-camera collaborative monitoring method provided in an embodiment of this application; Figure 3 This is a schematic diagram of gradient trajectory generation provided in an embodiment of this application; Figure 4 This is a schematic diagram of the structure of a multi-camera collaborative monitoring system provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of this application.
[0025] Explanation of reference numerals in the attached figures: 901, processor; 902, communication bus; 903, user interface; 904, network interface; 905, memory. Detailed Implementation
[0026] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0027] In the description of the embodiments of this application, the words "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design that is described as "for example" or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design options. Rather, the use of the words "for example" or "for instance" is intended to present the relevant concepts in a specific manner.
[0028] In the description of the embodiments of this application, the term "multiple" means two or more. For example, multiple systems means two or more systems, and multiple screen terminals means two or more screen terminals. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0029] Figure 1 This paper presents a multi-camera collaborative monitoring system architecture.
[0030] like Figure 1 As shown, the system architecture may include multiple cameras 011, a network 012, and electronic devices 013. The network 012 provides a data transmission link between the multiple cameras 011 and the electronic devices 013. The network 012 may include various connection types, such as wired or wireless communication links or fiber optic cables.
[0031] The multi-camera 011 can transmit the acquired video streams and image data to the electronic device 013 via the network 012. The multi-camera 011 is mainly responsible for acquiring real-time video streams within the monitored area, as well as capturing key image data according to preset rules.
[0032] The multi-camera 011 is hardware, which can be an intelligent camera device with automatic tracking and scene switching functions, including but not limited to basic components such as high-definition network cameras, infrared thermal imaging cameras, and PTZ cameras.
[0033] Electronic device 013 is responsible for comprehensively analyzing and processing the received data, including multi-view scene stitching, target detection and tracking, behavior analysis and recognition, and multi-camera collaborative scheduling. Electronic device 013 can intelligently switch between multiple cameras based on scene characteristics, calculate the target's motion trajectory, and combine this with preset behavior patterns to ultimately achieve real-time early warning of abnormal events. These data processing and analysis results can be used for subsequent monitoring and security decision-making.
[0034] It should be noted that electronic devices can be either hardware or software. When an electronic device is hardware, it can be implemented as a distributed cluster of multiple electronic devices or as a single electronic device. When an electronic device is software, it can be implemented as multiple software programs or software modules (e.g., multiple software programs or software modules used to provide distributed processing) or as a single software program or software module. No specific limitations are set here.
[0035] It should be understood that Figure 1 The number of multiple cameras 011, network 012, and electronic devices 013 shown is merely illustrative. Depending on implementation needs, there can be any number of multiple cameras 011, network 012, and electronic devices 013. In particular, if the target data does not need to be acquired remotely, the above system architecture may exclude network 012 and include only multiple cameras 011 or electronic devices 013.
[0036] The following description uses an electronic device as an example to illustrate a multi-camera collaborative monitoring method provided in this application.
[0037] This application provides a multi-camera collaborative monitoring method, referencing... Figure 2 , Figure 2 This is a flowchart illustrating a multi-camera collaborative monitoring method provided in an embodiment of this application, including steps S101 to S107, as follows: S101: Acquire two-dimensional images in real time from multiple cameras set up at different angles in the substation, and record the spatial position parameters and shooting angle parameters of each camera.
[0038] In this embodiment, spatial position parameters and shooting angle parameters refer to a set of parameters used to describe the installation position and shooting direction of the camera in three-dimensional space. Spatial position parameters include the x, y, and z coordinates of the camera in the substation coordinate system, representing the specific installation position of the camera; shooting angle parameters include the horizontal rotation angle, vertical pitch angle, and field of view of the camera, used to represent the shooting direction and coverage area of the camera.
[0039] Specifically, for multiple cameras pre-installed at different locations within the substation, electronic equipment acquires real-time two-dimensional images from each camera. These two-dimensional images contain complete scene information within the field of view of each camera. During the acquisition of two-dimensional images, the current spatial position parameters of each camera are simultaneously read and recorded, i.e., the installation position coordinates of the camera in the three-dimensional coordinate system of the substation. At the same time, the shooting angle parameters of each camera are acquired and recorded, including horizontal turning angle, vertical pitch angle, and field of view.
[0040] S102: Based on the spatial position parameters and shooting angle parameters of each camera, establish a three-dimensional coordinate system for the substation, and map the two-dimensional images collected by each camera into the three-dimensional coordinate system.
[0041] Specifically, firstly, the reference point of the coordinate system is determined based on the spatial position parameters of each camera within the substation. Typically, a fixed feature point of the substation (such as an end point) is chosen as the origin. Then, based on the camera installation positions and the spatial structural features of the substation, the three coordinate axes of the three-dimensional coordinate system are determined. The x-axis is usually parallel to the length of the substation, the y-axis is parallel to the width, and the z-axis is perpendicular to the substation plane and pointing upwards. After establishing the three-dimensional coordinate system, the spatial position parameters of each camera are represented as a vector P(x,y,z), and the shooting angle parameters are represented as a rotation matrix R. For any point a(u,v) in the two-dimensional image, its corresponding three-dimensional spatial point P'(X,Y,Z) can be obtained through the following transformation relationship: P' = R⁻¹×K⁻¹×a+P; Where K is the intrinsic parameter matrix of the camera, containing parameters such as focal length and principal point. Through this transformation relationship, the two-dimensional image is mapped to a three-dimensional coordinate system.
[0042] S103: In the three-dimensional coordinate system, calculate the overlapping area of each viewpoint and the two-dimensional image of the adjacent viewpoint, and divide the overlapping area into multiple grid cells of equal size.
[0043] In this embodiment, the overlapping area refers to the common coverage area between the shooting angles of adjacent cameras.
[0044] Specifically, in the established 3D coordinate system, the field of view of each camera is first represented as a square pyramid with the camera position as its vertex, determined by the horizontal and vertical field of view angles. For any two adjacent cameras, the overlapping area of their viewpoints can be obtained by the intersection of their corresponding square pyramids in 3D space. This overlapping area calculation needs to be performed for each pair of adjacent cameras. After determining the overlapping area, according to the preset mesh size parameters, the overlapping area is uniformly divided along the three directions of the 3D coordinate system to form multiple mesh cells of equal size.
[0045] S104: For each grid cell, extract the grid images captured by different cameras within the grid cell, and calculate the 3D reconstruction uncertainty of each grid cell based on each grid image.
[0046] In this embodiment, 3D reconstruction uncertainty refers to an index used to quantify the observation quality of a single grid cell. It is quantified by calculating the spatial dispersion of grid cells with inconsistent observation data. The higher the value, the more blurred, occluded, or mismatched the depth information of that grid cell is, and the greater the possibility of mismatch, thus requiring more collaborative observation optimization.
[0047] Specifically, for each grid cell, the corresponding grid image is first extracted from the 2D images captured by each camera. Depth values of each pixel in these grid images are obtained using a depth calculation algorithm, and these depth values are mapped to their corresponding location regions in a 3D coordinate system, thus obtaining the depth value distribution characteristics of that location region. Subsequently, the depth correlation between the grid cell and its corresponding location region is calculated. When the depth correlation of a grid cell is lower than a pre-set correlation threshold, this grid cell is marked as an overlapping grid cell. Finally, the 3D reconstruction uncertainty is obtained by calculating the standard deviation of the distribution of these overlapping grid cells throughout the overlapping region.
[0048] Based on the above embodiments, as an optional embodiment, S104: the step of calculating the 3D reconstruction uncertainty of each grid cell based on each grid image may specifically include the following steps: S201: Extract the depth value of each grid image pixel, map the location region corresponding to each grid unit to the three-dimensional coordinate system, and obtain the depth value distribution of the location region corresponding to each grid unit.
[0049] In this embodiment, the depth value refers to the spatial distance measurement from each pixel in the image to the camera. It is used to represent the relative depth position of the pixel in three-dimensional space and is usually expressed in grayscale value. The larger the grayscale value, the farther away from the camera.
[0050] Specifically, firstly, all pixels contained in each grid cell are extracted from the image region corresponding to that grid cell. For each target pixel, a local window region is defined around it, and the grayscale value sequence, gradient direction sequence, and texture feature sequence within the local window region are extracted as feature vectors. Within the corresponding search region of adjacent viewpoint images, the same feature vectors are extracted for each candidate pixel using a sliding window approach. By calculating the Euclidean distance between the feature vectors of the target pixel and each candidate pixel, the candidate point with the smallest distance is selected as the matching point. The coordinate difference between two matching points on the image plane is the disparity value. The depth value of the target pixel can be obtained based on the ratio of the disparity value to the distance between the two cameras. After obtaining the depth value, the two-dimensional coordinates and corresponding depth values of the pixel are transformed into an established three-dimensional coordinate system by combining the spatial position parameters and shooting angle parameters of the cameras. After completing this transformation for all pixels within a grid cell, the depth value distribution of the corresponding location region of that grid cell in three-dimensional space can be obtained.
[0051] S202: Determine the depth correlation between each grid cell and its corresponding location region based on the depth value distribution of each grid cell's location region.
[0052] In the embodiments of this application, the depth correlation degree refers to the degree of consistency between the depth value distribution of pixels in the grid cell and its corresponding position region in three-dimensional space. It is used to characterize the reliability of the depth estimation result. The larger the value, the more accurate the depth estimation.
[0053] Specifically, firstly, a statistical analysis is performed on the depth value distribution of the corresponding region for each grid cell. The mean depth value of all pixels within the grid cell is calculated as the reference depth value for that region. For each pixel within the grid cell, the deviation between its depth value and the reference depth value is calculated. The sum of the squares of all deviations is divided by the total number of pixels to obtain the depth value dispersion of that grid cell. Finally, by normalizing the depth value dispersion and mapping it to a numerical range between 0 and 1, the depth correlation between the grid cell and the corresponding region is obtained.
[0054] S203: Mark the grid cells with a depth correlation degree less than the preset correlation threshold as overlapping grid cells, and calculate the distribution standard deviation of the overlapping grid cells in the overlapping region to obtain the 3D reconstruction uncertainty of each grid cell.
[0055] In the embodiments of this application, the standard deviation of distribution refers to the measure of the spatial dispersion of overlapping grid cells in the overlapping region, which is used to characterize the concentration or dispersion of overlapping grid cells. The larger the value, the more dispersed the distribution of overlapping grid cells.
[0056] Specifically, a preset association threshold is first set as the criterion for determining the depth association degree. For each grid cell, the depth association degree of the grid cell is compared with the preset association threshold. When the depth association degree is less than the preset association threshold, the grid cell is marked as an overlapping grid cell. Then, the spatial distribution of all overlapping grid cells is statistically analyzed in the 3D coordinate system of the entire overlapping region. The mean coordinates of these overlapping grid cells in the X, Y, and Z axes are calculated. The mean coordinates of each overlapping grid cell are subtracted from the mean coordinates in the corresponding directions to obtain the deviation values of the grid cell in each direction. The sum of squares of these deviation values is then normalized to obtain the standard deviation of the distribution. This standard deviation is the 3D reconstruction uncertainty of the grid cell.
[0057] S105: When there is a target overlapping region where the average 3D reconstruction uncertainty of all grid cells exceeds a preset threshold, select the two adjacent viewpoints corresponding to the target grid cell with the largest 3D reconstruction uncertainty in the target overlapping region as the cooperative observation viewpoints.
[0058] In this embodiment, the cooperative observation perspective refers to the need to reselect a more suitable pair of observation perspectives when the overall depth estimation quality of the overlapping area is poor, in order to improve the quality of depth value acquisition.
[0059] Specifically, the system first calculates the average 3D reconstruction uncertainty of all grid cells within each overlapping region and compares it with a pre-set threshold. When the average 3D reconstruction uncertainty of grid cells within a certain overlapping region exceeds the threshold, it indicates that the overall observation quality of that overlapping region is poor, with significant occlusion or blurring issues, requiring optimization. In this case, the system iterates through the 3D reconstruction uncertainties of all grid cells within the target overlapping region, identifies the grid cell with the highest uncertainty value (i.e., the grid cell with the worst observation quality and the greatest need for optimization), and marks it as the target grid cell. Finally, the two adjacent viewpoints on which the maximum uncertainty value was generated are selected as the optimal cooperative observation viewpoints, as the geometric relationship between these two viewpoints is likely best suited for optimizing the reconstruction of this specific problem point.
[0060] S106: Perform perspective transformation on the two-dimensional image corresponding to the collaborative observation viewpoint, align the two-dimensional images of the two viewpoints of the collaborative observation viewpoint in the three-dimensional coordinate system, obtain the depth value of the target grid cell, and construct the target contour within the target grid cell based on the depth value. The depth value represents the actual distance from the object to the camera.
[0061] In the embodiments of this application, perspective transformation refers to the process of converting two-dimensional images taken from different viewpoints into a unified three-dimensional coordinate system, which is used to eliminate image distortion caused by differences in viewpoints, so that images can be compared and analyzed in the same spatial coordinate system.
[0062] Specifically, firstly, perspective transformation is performed on the two images corresponding to the collaborative observation perspective, projecting them into the same 3D coordinate system. This ensures the images from the two collaborative observation perspectives are correctly aligned in 3D space. In the aligned image, depth values within the target grid cells are extracted. By analyzing the depth values, local extrema are identified; these extrema typically correspond to salient features of the target contour. Simultaneously, the depth change rate between adjacent pixels within the grid cell is calculated; when the change rate exceeds a preset depth change rate threshold, the region is marked as a candidate target region. Starting from the depth extrema, a diffusion search is performed along the depth value gradient direction, recording the trajectory of decreasing depth change rate. On these gradient trajectories, multiple inflection points are identified by analyzing the trend of depth value changes; these inflection points typically correspond to turning points in the target contour. The coordinate distance from each inflection point to the depth extrema is calculated, and the difference in depth change rate and Euclidean distance between adjacent inflection points are analyzed. When the ratio of the difference in depth change rate between a certain inflection point and its adjacent inflection points to the Euclidean distance exceeds a preset ratio threshold, and the coordinate distance from that inflection point to the depth extremum point is greater than a preset coordinate distance threshold, it is identified as an abnormal inflection point and removed. Finally, the pixels corresponding to the remaining inflection points are connected in spatial order to form a complete target contour.
[0063] Based on the above embodiments, as an optional embodiment, S106: the step of constructing the target contour within the target mesh cell based on the depth value may specifically include the following steps: S301: Extract depth extreme points within the target grid cell from the depth value, determine the depth change rate based on the depth value, and select regions with a depth change rate higher than a preset depth change rate threshold as candidate target regions.
[0064] In the embodiments of this application, the depth extreme point refers to the location where the depth value in the target grid cell is locally maximum or minimum, and is used to characterize the key feature location of the target contour.
[0065] Specifically, the depth values within the target grid cells are first scanned and analyzed. By comparing the depth values of each pixel with those of its neighboring pixels, locations where the depth value exhibits a local maximum or minimum are identified and marked as depth extrema. These depth extrema often correspond to significant features of the target contour. Then, the depth change rate is obtained by dividing the depth value difference between adjacent pixels by the distance between them. The calculated depth change rate is compared with a pre-set depth change rate threshold. When the depth change rate of a region exceeds the threshold, the region with a depth change rate higher than the threshold is marked as a candidate target region.
[0066] S302: Starting from the depth extreme point of the candidate target region, diffuse outwards and record the gradient trajectory with decreasing depth change rate.
[0067] In this embodiment, the gradient trajectory refers to the spatial path formed by extending from the depth extreme point along the direction of depth value change. It is used to describe the trend of depth value change in different directions and records the depth transition process from the target contour feature point to the background area.
[0068] Please refer to Figure 3 , Figure 3 This is a schematic diagram of gradient trajectory generation provided in an embodiment of this application.
[0069] Specifically, the location coordinates of each depth extremum point are first determined within the marked candidate target region. Centered on each depth extremum point, scanning directions are divided within a range of 0 to 360 degrees according to preset angular intervals (e.g., 30 degrees). The depth change rate is calculated in each scanning direction. The direction with the largest depth change rate is selected as the initial diffusion direction, and diffusion gradually proceeds outward from the depth extremum point. During diffusion, the depth change rate between the current pixel and the next pixel is continuously calculated, and the values of these depth change rates and their corresponding location coordinates are recorded to form gradient trajectories. When the depth change rate begins to increase or reaches the preset search range boundary, diffusion in the current direction is stopped, and the next direction with the second largest depth change rate is selected to continue the diffusion search. In this way, through diffusion searches in different directions, a series of gradient trajectories starting from depth extremum points and decreasing in depth change rate are finally obtained.
[0070] S303: Determine multiple inflection points on the gradient trajectory and calculate the coordinate distance from each inflection point to the depth extremum point.
[0071] In this embodiment, an inflection point refers to a location on the gradient trajectory where the rate of change of depth changes significantly. It is used to mark key turning points of the target contour. These locations typically correspond to boundary features where the shape of the human body or device changes significantly.
[0072] Specifically, the depth change rate on each gradient trajectory is first analyzed, and the difference in depth change rate between adjacent points is calculated. When the difference exceeds a preset change threshold, it indicates that the trend of depth value change at that location has changed significantly, and that location is marked as a candidate inflection point. Next, the depth extremum point is set as the origin of the coordinate system, and the distances from the origin in the horizontal and vertical directions of each determined candidate inflection point are calculated. Based on the horizontal and vertical distances, the straight-line distance from the inflection point to the depth extremum point, i.e., the coordinate distance, is calculated.
[0073] S304: Determine the difference in depth change rate and Euclidean distance between each inflection point and its adjacent inflection points. When there is an abnormal inflection point where the ratio of the difference in depth change rate to the Euclidean distance is greater than a preset ratio threshold and the coordinate distance is greater than a preset coordinate distance threshold, the abnormal inflection point is removed, and the pixels corresponding to the remaining inflection points are connected to form the target contour.
[0074] In the embodiments of this application, an abnormal inflection point refers to an inflection point in the inflection point sequence that has significant differences in spatial distribution and depth variation characteristics from adjacent inflection points. It is used to represent erroneous feature points that may be caused by noise or interference. These points often deviate from the true target contour.
[0075] Specifically, firstly, the difference in depth change rate between each pair of adjacent inflection points on the gradient trajectory is calculated sequentially to obtain the depth change rate difference value. Simultaneously, the Euclidean distance between the two inflection points on the image plane is calculated; this Euclidean distance represents the straight-line distance between the two points. The obtained depth change rate difference value is divided by the Euclidean distance to obtain the ratio. When the ratio between an inflection point and its adjacent inflection points exceeds a preset ratio threshold, and the coordinate distance from the inflection point to its corresponding depth extremum point is greater than a preset coordinate distance threshold, this inflection point is marked as an abnormal inflection point. Then, all marked abnormal inflection points are deleted, retaining only the reliable set of inflection points. Finally, the retained inflection points are connected sequentially according to their spatial order on the gradient trajectory to form a complete target contour line.
[0076] Based on the above embodiments, as an optional embodiment, S106: after the step of constructing the target contour within the target mesh cell based on the depth value, the step of separating the target contour is further included, which may specifically include the following steps: S401: Determine the candidate boundary points of each target contour within the target grid cell based on the depth extreme points of each target contour.
[0077] In the embodiments of this application, contour boundary candidate points refer to feature points in the overlapping or contact areas of multiple adjacent target contours, which serve as the separation positions of different target contours and are used to identify the boundary boundary positions of different target contours.
[0078] Specifically, firstly, the Euclidean distance between depth extrema of different target contours is calculated. When the Euclidean distance between some depth extrema is less than a preset spatial distance threshold, the closest depth extrema are grouped together, and the target contours corresponding to the same group of depth extrema are marked as contours to be separated. Then, in each group of contours to be separated, a depth transition zone is constructed with adjacent depth extrema as endpoints. The width of the depth transition zone is proportional to the depth difference between the two depth extrema. Within the defined depth transition zone region, the rate of change of depth value with spatial location (first derivative) and the rate of change of the rate of change (second derivative) are calculated. When the absolute value of the second derivative at a certain location exceeds a preset boundary threshold, these points are determined as candidate contour boundaries.
[0079] Based on the above embodiments, as an optional embodiment, S401: the step of determining the candidate boundary points of each target contour within the target mesh cell based on the depth extreme points of each target contour may specifically include the following steps: S501: For multiple target contours within a target grid cell, based on the Euclidean distance between depth extreme points in each target contour, depth extreme points with an Euclidean distance less than a preset spatial distance threshold are grouped into the same group, and the target contours corresponding to each group of depth extreme points are marked as contours to be separated.
[0080] In this embodiment of the application, the contour to be separated refers to multiple target contours that are close to each other in depth extrema within the target grid cell. It is used to represent a set of target contours that may overlap or contact each other and need to be separated by boundary. These contours usually appear in scenarios with many people or device interaction.
[0081] Specifically, firstly, all detected target contours and their corresponding depth extrema sets within the target mesh cell are acquired. For each pair of different target contour depth extrema points, their Euclidean distance in space is calculated. This Euclidean distance is obtained by calculating the difference between the horizontal and vertical distances of the two depth extrema points. When the Euclidean distance between some depth extrema points is less than a preset spatial distance threshold, they are marked as meeting the distance condition. All depth extrema points that meet the distance condition are grouped into the same group, and the target contours corresponding to these depth extrema points are marked as contours to be separated.
[0082] S502: Construct a depth transition zone between depth extrema in each set of contours to be separated. The width of the depth transition zone is determined by the depth difference between adjacent depth extrema.
[0083] In the embodiments of this application, the depth transition zone refers to a spatial band-shaped region in the area between adjacent depth extreme points, where the depth value gradually transitions from one extreme point to another.
[0084] Specifically, firstly, in each set of contours to be separated, the direction of the line connecting adjacent depth extrema pairs is determined. Along this line direction, the difference in depth values between the two depth extrema points is calculated to obtain the depth difference. The width of the depth transition zone is dynamically determined based on the magnitude of the depth difference; the width of the depth transition zone is proportional to the depth difference between the two depth extrema points. After determining the width, a strip-shaped region is formed by extending outwards from the connecting line as the center line; this region is the depth transition zone.
[0085] S503: Calculate the first and second derivatives of each depth value within the depth transition zone, and determine the set of points whose absolute value of the second derivative is greater than the preset boundary threshold as candidate points for contour boundary.
[0086] Specifically, firstly, within the constructed depth transition zone, along the line connecting the depth extrema, the first derivative of the depth value is obtained by calculating the depth difference between adjacent sampling points and dividing it by the distance between them. The first derivative represents the rate of depth change at that location. Based on the first derivative, the rate of change of the first derivative, i.e., the second derivative, is further calculated. This value reflects the change in the rate of depth change. When the absolute value of the second derivative at a certain location exceeds a pre-set boundary threshold, it indicates that the depth change trend at that location has changed drastically. This change usually corresponds to the boundary position of different target contours. All locations that meet the conditions are collected to form candidate contour boundary points.
[0087] S402: Calculate the gradient direction angle between each contour boundary candidate point and each depth extreme point, and connect the contour boundary candidate points whose gradient direction angle is less than the preset angle threshold to form a boundary line.
[0088] Specifically, firstly, for each candidate contour boundary point, the gradient direction of the depth value at that point is calculated. Simultaneously, the direction of the line connecting that candidate point to each relevant depth extremum point is calculated. The gradient direction angle is obtained by calculating the angle between the gradient direction and the connecting line direction. When the gradient direction angle of a candidate contour boundary point is less than a preset angle threshold, it indicates that the candidate point meets the angle condition. All candidate points that meet the angle condition are then connected sequentially according to their spatial location to form a complete boundary line.
[0089] S403: Divide the target contour into multiple target contours according to the boundary line, and assign a unique identifier to each target contour.
[0090] Specifically, firstly, based on the established boundary lines, the original overlapping or contacting target contours are segmented into multiple independent target contours. The boundary lines serve as cutting boundaries, clearly separating the contour regions corresponding to different human bodies or devices. Then, each segmented target contour is assigned a unique identifier. This identifier can be an incrementing numerical sequence number, or it can include feature values such as timestamps or location information, ensuring that each target contour has a unique identity throughout the entire processing.
[0091] S107: Real-time statistics of the rate of change of the number of target contours within each target grid cell, and determination of the abnormal area where an abnormal event occurs based on the rate of change and the target contours.
[0092] Specifically, the target contours within each target grid cell are first counted in real time, with the number of contours recorded at fixed time intervals (e.g., per second or per frame). The difference in the number of contours between two adjacent time points is calculated and divided by the time interval to obtain the rate of change in the number of contours. This rate of change can be positive (indicating an increase in quantity) or negative (indicating a decrease in quantity). The calculated rate of change data is then compared with the rate of change patterns stored in the abnormal behavior feature database, and a comprehensive analysis is performed in conjunction with the spatial distribution and morphological features of the currently detected target contours. Through multi-dimensional feature matching, a comprehensive anomaly probability reflecting the degree of anomaly in the current state is calculated. When the comprehensive anomaly probability of a target grid cell exceeds a preset anomaly probability threshold, the area is marked as an anomaly area.
[0093] Based on the above embodiments, as an optional embodiment, S107: the step of determining the abnormal region where the abnormal event occurred based on the rate of change and the target contour may specifically include the following steps: S601: Establish an abnormal behavior feature database, compare the rate of change and the target contour with the abnormal behavior feature database to obtain the comprehensive abnormal probability of the target grid cell.
[0094] In this application embodiment, the abnormal behavior feature database refers to a structured data set containing various typical abnormal behavior feature parameters, which is used to represent standard feature patterns of various abnormal behaviors.
[0095] The abnormal behavior feature database includes target contour quantity mutation features, target contour relative position features, and target contour motion features. In this application embodiment, the target contour quantity mutation features refer to the feature parameters of the target contour quantity within the target grid cell that change significantly in a short period of time. These features are used to represent the degree and pattern of drastic increase or decrease in the number of people. These features can reflect abnormal behaviors such as sudden gathering or rapid evacuation of people.
[0096] The relative positional features of target contours refer to the spatial distribution relationship parameters between multiple target contours within a target grid cell. These parameters are used to represent spatial features such as the degree of aggregation, spacing distribution, and arrangement of personnel and equipment in three-dimensional space. These features can reflect behavioral patterns such as crowding and evacuation of personnel and equipment.
[0097] Target contour motion characteristics refer to the dynamic parameters of the target contour, such as displacement, velocity, and acceleration, over a continuous period of time, used to characterize the target's motion state and trajectory.
[0098] Specifically, firstly, an abnormal behavior feature database is established by collecting a large amount of historical and simulated data. The database stores feature parameters for different types of abnormal behavior, including the threshold range of the rate of change of target contour quantity, typical abrupt change patterns, and abnormal distribution and motion patterns of target contours in three-dimensional space. Based on the established database, the current degree of anomaly within the target mesh cell is calculated. By analyzing the rate of change and abrupt change characteristics of the target contour quantity per unit time, these characteristics are matched and compared with the abnormal patterns stored in the database to obtain a first anomaly probability reflecting the degree of abnormal quantity change. Simultaneously, the spatial distribution characteristics of each target contour within the target mesh cell are analyzed, including relative positional relationships and changes in three-dimensional coordinates, and the motion characteristics of the target contours, such as motion direction and velocity, are considered. These features are compared with the abnormal behavior patterns in the database to calculate a second anomaly probability reflecting the degree of behavioral anomaly. Finally, based on the sensitivity of different types of anomalies in the actual application scenario, corresponding weighting coefficients are set, and the first and second anomaly probabilities are weighted and summed to obtain the comprehensive anomaly probability of the target mesh cell.
[0099] Based on the above embodiments, as an optional embodiment, S601: the step of comparing the rate of change and the target contour with the abnormal behavior feature database to obtain the comprehensive anomaly probability of the target mesh cell may specifically include the following steps: S701: Calculate the first anomalous probability of the target contour quantity change rate within the target grid cell based on the change rate of the target contour quantity within each target grid cell and the abrupt change characteristics of the target contour quantity.
[0100] Specifically, the number of target contours within the target grid cell is first monitored in real time, and the difference in number between adjacent time points is calculated to obtain the rate of change in the number of target contours. This rate of change is compared to a baseline value for normal pedestrian flow; when the rate of change exceeds a specific multiple of the baseline value, it is considered an abnormal change. The rate of change is examined across multiple consecutive time points; if the rate of change increases or decreases and the magnitude exceeds a threshold, a trend anomaly is considered to exist. When the magnitude of the change exceeds a preset multiple of the average value, it is recorded as an amplitude anomaly; when the duration of the abnormal change exceeds a preset duration, it is recorded as a persistent anomaly; when the number of abnormal changes per unit time exceeds a preset frequency, it is recorded as a frequency anomaly. The amplitude anomaly, persistent anomaly, and frequency anomaly are weighted, combined, and normalized to obtain the first anomaly probability.
[0101] S702: Based on the relative positional relationship and coordinate change of each target contour within the target grid cell in the three-dimensional coordinate system, as well as the relative positional characteristics and motion characteristics of the target contour, calculate the second anomaly probability of the target contour behavior within the target grid cell.
[0102] Specifically, firstly, the three-dimensional coordinates of all target contours within the target grid cell are obtained, and the Euclidean distance between any two target contours is calculated to obtain the relative positional relationship of the target contours. The minimum spacing is compared with a safe distance threshold; when it is less than the threshold, it is recorded as a spacing anomaly. The mean of all spacings is calculated; when the mean is less than a preset standard, it is recorded as a density anomaly; when the standard deviation exceeds the normal range, it is recorded as a distribution anomaly. For each target contour, the displacement vector of the target contour at continuous time points is calculated to obtain motion characteristics. The magnitude of the displacement vector is analyzed; when it exceeds the velocity threshold, it is recorded as a velocity anomaly. The angle between the displacement vectors of adjacent target contours is calculated; when the angle deviates from the normal range, it is recorded as a direction anomaly. The rate of change of the displacement vector is monitored; when the rate of change exceeds the acceleration threshold, it is recorded as an acceleration anomaly. The positional characteristic indicators such as spacing anomaly, density anomaly, and distribution anomaly are weighted and combined to obtain the positional anomaly value. The motion characteristic indicators such as velocity anomaly, direction anomaly, and acceleration anomaly are weighted and combined to obtain the motion anomaly value. Finally, the positional anomaly value and the motion anomaly value are weighted and combined and normalized to obtain the second anomaly probability.
[0103] S703: Calculate the weighted sum of the first and second anomaly probabilities to obtain the comprehensive anomaly probability of the target grid cell.
[0104] Specifically, firstly, based on the statistical results of abnormal events in the abnormal behavior feature database, the correlation coefficients between the first and second abnormal probabilities and the actual abnormal situation are obtained. After normalizing the correlation coefficients, their respective weight coefficients are obtained, ensuring that the sum of the two weight coefficients is 1. The first abnormal probability is multiplied by its weight coefficient, and the second abnormal probability is multiplied by its weight coefficient. The two weighted abnormal probabilities are then added together to obtain the comprehensive abnormal probability of the target grid cell.
[0105] S602: When the overall anomaly probability exceeds the preset anomaly probability threshold, the target grid cell is identified as an anomaly region.
[0106] Specifically, the overall anomaly probability of the target mesh cell is first obtained, and then compared with a preset anomaly probability threshold. When the overall anomaly probability is greater than the preset anomaly probability threshold, the target mesh cell is marked as an anomaly region.
[0107] refer to Figure 4 This application also provides a schematic diagram of a multi-camera collaborative monitoring system, which specifically includes: The data acquisition module is used to acquire two-dimensional images in real time from multiple cameras set at different angles in the substation, and record the spatial position parameters and shooting angle parameters of each camera. The target contour construction module is used to establish a three-dimensional coordinate system of the substation based on the spatial position parameters and shooting angle parameters of each camera, and to map the two-dimensional images captured by each camera into the three-dimensional coordinate system; In the three-dimensional coordinate system, the overlapping area of each viewpoint and the two-dimensional image of the adjacent viewpoint is calculated respectively, and the overlapping area is divided into multiple grid cells of equal size; For each of the aforementioned grid cells, grid images captured by different cameras within the grid cell are extracted, and the 3D reconstruction uncertainty of each grid cell is calculated based on the aforementioned grid images. When there is a target overlapping region where the average 3D reconstruction uncertainty of all grid cells exceeds a preset threshold, the two adjacent viewpoints corresponding to the target grid cell with the largest 3D reconstruction uncertainty in the target overlapping region are selected as the cooperative observation viewpoints. The two-dimensional image corresponding to the collaborative observation perspective is subjected to perspective transformation so that the two-dimensional images of the two perspectives of the collaborative observation perspective are aligned in the three-dimensional coordinate system to obtain the depth value of the target grid cell. The target contour within the target grid cell is constructed based on the depth value, where the depth value represents the actual distance from the object to the camera. The abnormal region determination module is used to statistically analyze the rate of change of the number of target contours within each target grid cell in real time, and determine the abnormal region where an abnormal event has occurred based on the rate of change and the target contours.
[0108] Optionally, the target contour construction module is specifically used for: Extract the depth value of each pixel in the grid image, map the location region corresponding to each grid unit to the three-dimensional coordinate system, and obtain the depth value distribution of the location region corresponding to each grid unit; Based on the depth value distribution of the location region corresponding to each grid cell, the depth correlation degree between each grid cell and the corresponding location region is determined; The mesh cells with a depth correlation degree less than a preset correlation threshold are marked as overlapping mesh cells, and the distribution standard deviation of the overlapping mesh cells in the overlapping region is calculated to obtain the three-dimensional reconstruction uncertainty of each mesh cell.
[0109] Optionally, the target contour construction module is further specifically used for: Extract the depth extreme points within the target grid cell from the depth value, determine the depth change rate based on the depth value, and select the region with the depth change rate higher than a preset depth change rate threshold as the candidate target region; Starting from the depth extreme point of the candidate target region, the gradient trajectory with decreasing depth change rate is recorded as it expands outward. Multiple inflection points are determined on the gradient trajectory, and the coordinate distance from each inflection point to the depth extremum point is calculated. The depth change rate difference and Euclidean distance between each inflection point and its adjacent inflection points are determined. When there is an abnormal inflection point where the ratio of the depth change rate difference to the Euclidean distance is greater than a preset ratio threshold and the coordinate distance is greater than a preset coordinate distance threshold, the abnormal inflection point is removed, and the pixels corresponding to the remaining inflection points are connected to form the target contour.
[0110] Optionally, the target contour construction module is further specifically used for: Based on the depth extreme points of each target contour, determine the contour boundary candidate points within the target grid cell for each target contour; Calculate the gradient direction angle between each of the contour boundary candidate points and each depth extreme point, and connect the contour boundary candidate points whose gradient direction angle is less than a preset angle threshold to form a boundary line. The target contour is divided into multiple target contours according to the dividing line, and each target contour is assigned a unique identifier.
[0111] Optionally, the target contour construction module is further specifically used for: For multiple target contours within a target grid cell, based on the Euclidean distance between each depth extreme point in each target contour, depth extreme points with an Euclidean distance less than a preset spatial distance threshold are grouped into the same group, and the target contour corresponding to each group of depth extreme points is marked as the contour to be separated. Construct a depth transition zone between depth extrema in each group of contours to be separated, the width of which is determined by the depth difference between adjacent depth extrema. Calculate the first and second derivatives of each depth value within the depth transition zone, and determine the set of points whose absolute value of the second derivative is greater than a preset boundary threshold as candidate contour boundary points.
[0112] Optionally, the abnormal region determination module is specifically used for: An abnormal behavior feature database is established, and the rate of change and the target contour are compared with the abnormal behavior feature database to obtain the comprehensive abnormal probability of the target grid cell. When the overall anomaly probability exceeds a preset anomaly probability threshold, the target grid cell is identified as the anomaly region.
[0113] Optionally, the abnormal region determination module is further specifically used for: Based on the rate of change of the number of target contours within each target grid cell and the abrupt change characteristics of the number of target contours, calculate the first abnormal probability of the rate of change of the number of target contours within the target grid cell; Based on the relative positional relationship and coordinate change of each target contour within the target mesh cell in the three-dimensional coordinate system, as well as the relative positional features and motion features of the target contour, the second anomaly probability of the target contour behavior within the target mesh cell is calculated. The weighted sum of the first anomaly probability and the second anomaly probability is calculated to obtain the comprehensive anomaly probability of the target grid cell.
[0114] It should be noted that the above embodiments of the apparatus are only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided above belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0115] This embodiment also discloses an electronic device, as shown in the reference. Figure 5 , Figure 5This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of this application. The electronic device 013 may include: at least one processor 901, at least one communication bus 902, a user interface 903, a network interface 904, and at least one memory 905.
[0116] The communication bus 902 is used to enable communication between these components.
[0117] The user interface 903 may include a display screen and a camera. Optionally, the user interface 903 may also include a standard wired interface and a wireless interface.
[0118] The network interface 904 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).
[0119] The processor 901 may include one or more processing cores. The processor 901 connects to various parts of the server using various interfaces and lines, and performs various server functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 905, and by calling data stored in the memory 905. Optionally, the processor 901 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array. The processor 901 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content to be displayed on the screen; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 901 and may be implemented as a separate chip.
[0120] The memory 905 may include random access memory (RAM) or read-only memory. Optionally, the memory 905 may include a non-transitory computer-readable storage medium. The memory 905 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 905 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 905 may also be at least one storage device located remotely from the aforementioned processor 901. (See reference...) Figure 5 The memory 905, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application for multi-camera collaborative monitoring.
[0121] exist Figure 5 In the electronic device shown, the user interface 903 is mainly used to provide an input interface for the user and to obtain the user input data; while the processor 901 can be used to call an application for multi-camera collaborative monitoring stored in the memory 905. When executed by one or more processors 901, the electronic device 013 performs one or more methods as described in the above embodiments.
[0122] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0123] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0124] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the shown or discussed mutual couplings or direct couplings or communication connections may be through some service interfaces; indirect couplings or communication connections between apparatuses or units may be electrical or other forms.
[0125] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0126] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0127] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, portable hard drives, magnetic disks, or optical disks.
[0128] The foregoing description is merely an exemplary embodiment of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Other embodiments of this disclosure will be readily apparent to those skilled in the art upon consideration of the disclosure in this specification. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described in this disclosure. The specification and embodiments are to be considered exemplary only, and the scope and spirit of this disclosure are defined by the claims.
Claims
1. A multi-camera collaborative monitoring method, characterized in that, Applied to electronic devices, the method includes: Acquire two-dimensional images in real time from multiple cameras set up at different angles in the substation, and record the spatial position parameters and shooting angle parameters of each camera; Based on the spatial position parameters and shooting angle parameters of each camera, a three-dimensional coordinate system of the substation is established, and the two-dimensional images captured by each camera are mapped to the three-dimensional coordinate system. In the three-dimensional coordinate system, the overlapping area of each viewpoint and the two-dimensional image of the adjacent viewpoint is calculated respectively, and the overlapping area is divided into multiple grid cells of equal size; For each of the aforementioned grid cells, grid images captured by different cameras within the grid cell are extracted, and the 3D reconstruction uncertainty of each grid cell is calculated based on the aforementioned grid images. When there is a target overlapping region where the average 3D reconstruction uncertainty of all grid cells exceeds a preset threshold, the two adjacent viewpoints corresponding to the target grid cell with the largest 3D reconstruction uncertainty in the target overlapping region are selected as the cooperative observation viewpoints. The two-dimensional image corresponding to the collaborative observation perspective is subjected to perspective transformation so that the two-dimensional images of the two perspectives of the collaborative observation perspective are aligned in the three-dimensional coordinate system to obtain the depth value of the target grid cell. The target contour within the target grid cell is constructed based on the depth value, where the depth value represents the actual distance from the object to the camera. The rate of change of the number of target contours within each target grid cell is statistically analyzed in real time. Based on the rate of change and the target contours, the abnormal region where the abnormal event occurred is determined.
2. The multi-camera collaborative monitoring method according to claim 1, characterized in that, The step of calculating the 3D reconstruction uncertainty of each mesh cell based on each mesh image includes: Extract the depth value of each pixel in the grid image, map the location region corresponding to each grid unit to the three-dimensional coordinate system, and obtain the depth value distribution of the location region corresponding to each grid unit; Based on the depth value distribution of the location region corresponding to each grid cell, the depth correlation degree between each grid cell and the corresponding location region is determined; The mesh cells with a depth correlation degree less than a preset correlation threshold are marked as overlapping mesh cells, and the distribution standard deviation of the overlapping mesh cells in the overlapping region is calculated to obtain the three-dimensional reconstruction uncertainty of each mesh cell.
3. The multi-camera collaborative monitoring method according to claim 1, characterized in that, The step of constructing the target contour within the target mesh cell based on the depth value includes: Extract the depth extreme points within the target grid cell from the depth value, determine the depth change rate based on the depth value, and select the region with the depth change rate higher than a preset depth change rate threshold as the candidate target region; Starting from the depth extreme point of the candidate target region, the gradient trajectory with decreasing depth change rate is recorded as it expands outward. Multiple inflection points are determined on the gradient trajectory, and the coordinate distance from each inflection point to the depth extremum point is calculated. The depth change rate difference and Euclidean distance between each inflection point and its adjacent inflection points are determined. When there is an abnormal inflection point where the ratio of the depth change rate difference to the Euclidean distance is greater than a preset ratio threshold and the coordinate distance is greater than a preset coordinate distance threshold, the abnormal inflection point is removed, and the pixels corresponding to the remaining inflection points are connected to form the target contour.
4. The multi-camera collaborative monitoring method according to claim 3, characterized in that, After constructing the target contour within the target mesh cell based on the depth value, the method further includes: Based on the depth extreme points of each target contour, determine the contour boundary candidate points within the target grid cell for each target contour; Calculate the gradient direction angle between each of the contour boundary candidate points and each depth extreme point, and connect the contour boundary candidate points whose gradient direction angle is less than a preset angle threshold to form a boundary line. The target contour is divided into multiple target contours according to the dividing line, and each target contour is assigned a unique identifier.
5. The multi-camera collaborative monitoring method according to claim 4, characterized in that, The step of determining the candidate boundary points of each target contour within a target mesh cell based on the depth extreme points of each target contour includes: For multiple target contours within a target grid cell, based on the Euclidean distance between each depth extreme point in each target contour, depth extreme points with an Euclidean distance less than a preset spatial distance threshold are grouped into the same group, and the target contour corresponding to each group of depth extreme points is marked as the contour to be separated. Construct a depth transition zone between depth extrema in each group of contours to be separated, the width of which is determined by the depth difference between adjacent depth extrema. Calculate the first and second derivatives of each depth value within the depth transition zone, and determine the set of points whose absolute value of the second derivative is greater than a preset boundary threshold as candidate contour boundary points.
6. The multi-camera collaborative monitoring method according to claim 1, characterized in that, The step of determining the abnormal region where the abnormal event occurred based on the rate of change and the target contour includes: An abnormal behavior feature database is established, and the rate of change and the target contour are compared with the abnormal behavior feature database to obtain the comprehensive abnormal probability of the target grid cell. When the overall anomaly probability exceeds a preset anomaly probability threshold, the target grid cell is identified as the anomaly region.
7. The multi-camera collaborative monitoring method according to claim 6, characterized in that, The abnormal behavior feature database includes target contour quantity mutation features, target contour relative position features, and target contour motion features. The step of comparing the rate of change and the target contour with the abnormal behavior feature database to obtain the comprehensive abnormal probability of the target mesh cell includes: Based on the rate of change of the number of target contours within each target grid cell and the abrupt change characteristics of the number of target contours, calculate the first abnormal probability of the rate of change of the number of target contours within the target grid cell; Based on the relative positional relationship and coordinate change of each target contour within the target mesh cell in the three-dimensional coordinate system, as well as the relative positional features and motion features of the target contour, the second anomaly probability of the target contour behavior within the target mesh cell is calculated. The weighted sum of the first anomaly probability and the second anomaly probability is calculated to obtain the comprehensive anomaly probability of the target grid cell.
8. A multi-camera collaborative monitoring system, characterized in that, The system is applied to electronic devices and includes: The data acquisition module is used to acquire two-dimensional images in real time from multiple cameras set at different angles in the substation, and record the spatial position parameters and shooting angle parameters of each camera. The target contour construction module is used to establish a three-dimensional coordinate system of the substation based on the spatial position parameters and shooting angle parameters of each camera, and to map the two-dimensional images captured by each camera into the three-dimensional coordinate system; In the three-dimensional coordinate system, the overlapping area of each viewpoint and the two-dimensional image of the adjacent viewpoint is calculated respectively, and the overlapping area is divided into multiple grid cells of equal size; For each of the aforementioned grid cells, grid images captured by different cameras within the grid cell are extracted, and the 3D reconstruction uncertainty of each grid cell is calculated based on the aforementioned grid images. When there is a target overlapping region where the average 3D reconstruction uncertainty of all grid cells exceeds a preset threshold, the two adjacent viewpoints corresponding to the target grid cell with the largest 3D reconstruction uncertainty in the target overlapping region are selected as the cooperative observation viewpoints. The two-dimensional image corresponding to the collaborative observation perspective is subjected to perspective transformation so that the two-dimensional images of the two perspectives of the collaborative observation perspective are aligned in the three-dimensional coordinate system to obtain the depth value of the target grid cell. The target contour within the target grid cell is constructed based on the depth value, where the depth value represents the actual distance from the object to the camera. The abnormal region determination module is used to statistically analyze the rate of change of the number of target contours within each target grid cell in real time, and determine the abnormal region where an abnormal event has occurred based on the rate of change and the target contours.
9. An electronic device, characterized in that, The device includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions. The user interface and the network interface are both used to communicate with other devices. The processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed, perform the method as described in any one of claims 1-7.
Citation Information
Patent Citations
Multi-camera cooperative monitoring method and device
CN105979203A
Substation monitoring system design method based on multi-view image modeling
CN115914579A
Method and device for determining camera layout blind area based on three-dimensional scene
CN118283439A
Intelligent video traffic monitoring system based on multiple viewpoints
CN118968781A
Video monitoring method, video monitoring system and computer program product
US20170032194A1