Method and System for Personnel Positioning and Tracking in Surveillance Videos Based on Visual SLAM
Through the visual SLAM-based method, external parameter calibration and personnel three-dimensional positioning are realized in a fixed camera environment, solving the problem of cumbersome operation and positioning tracking in traditional monitoring systems, and improving the efficiency and accuracy of the monitoring system.
Patent Information
- Application Number
- CN202210669897.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-14
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-06-14
AI Technical Summary
In the existing monitoring systems, two-dimensional monitoring operations are cumbersome, unintuitive, and inefficient. In addition, personnel positioning and tracking rely on handheld devices, making it difficult to achieve accurate external parameter calibration in a fixed camera environment.
Using a visual SLAM-based method, three-dimensional reconstruction is carried out through a depth camera, combined with external parameter calibration, the position and attitude calibration of the surveillance camera is realized, and the characters are identified through a deep neural network, and their three-dimensional positions are calculated to realize trajectory drawing in the point cloud map.
It realizes external calibration of the mobile camera without requiring it, reducing the complexity and labor cost of monitoring operations, and improving the accuracy and efficiency of positioning tracking.
Smart Images

Figure CN115035162B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine vision, and particularly to a method and system for positioning and tracking personnel in surveillance videos based on visual SLAM. Background Art
[0002] In recent years, the domestic and international video surveillance market has experienced explosive growth, and the development of surveillance has been continuously trending towards high definition and intelligence. However, traditional two-dimensional surveillance systems require frequent switching of surveillance images and have various drawbacks. With the development of artificial intelligence, more and more tasks can be automatically completed by machines, enabling people to be liberated from boring work and improving work efficiency.
[0003] In the prior art, the position of a person in a scene is generally determined by Wi-Fi or the like for positioning. Wi-Fi positioning generally uses the "nearest neighbor method" for determination, that is, the position is considered to be where it is closest to which hotspot or base station. If there are multiple signal sources nearby, cross positioning (triangulation) can be used to improve the positioning accuracy. When a user turns on Wi-Fi and mobile cellular networks when using a smart phone, it may become a data source. It is necessary to pre-record the signal strengths of a huge amount of determined position points in advance, and compare the signal strength of the newly added device with a database having a huge amount of data to determine the position.
[0004] Compared with the Wi-Fi positioning method, the monitored person needs to actively carry a handheld device to receive wireless signals. All existing calibration methods require moving the camera. However, in the actual environment, most surveillance cameras have been fixed to a certain position, and it is impossible to use the calibration method of moving the camera commonly used in visual SLAM to calibrate it, that is, it is impossible to accurately obtain its specific representation in the computer of its position in the environment. Based on this problem, the present invention adopts a visual positioning method relying on the surveillance system, combines the external parameter calibration in the scene of the surveillance camera, and can perform external parameter calibration on it without moving the camera. At the same time, the person to be positioned does not need to carry a device, and only performs automated processing through the video image information captured by the camera to accurately locate the position of the person. Summary of the Invention
[0005] In view of the problems of the prior art such as cumbersome, unintuitive, low-efficiency two-dimensional surveillance operations and personnel positioning and tracking relying on handheld devices, the present invention proposes a method and system for positioning and tracking personnel in surveillance videos based on visual SLAM.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] On the one hand, a method for positioning and tracking personnel in surveillance videos based on visual SLAM provided by the present invention uses a depth camera to perform three-dimensional reconstruction of the environment to obtain a point cloud map of the environment; combines external parameter calibration in the scenario of a surveillance camera, calibrates the external parameters of the surveillance camera through a calibration board to obtain the position and attitude information of the surveillance camera; tracks personnel through the surveillance camera, uses a deep neural network to identify the people in the surveillance image, and based on the previously calibrated position and attitude, calculates the three-dimensional position of the person according to the prior on the ground of the pedestrians appearing in the surveillance by using the principle of inverse perspective transformation, realizes drawing the personnel trajectory in the constructed point cloud map, and presents it in the point cloud map.
[0008] Further, the above-mentioned method for positioning and tracking personnel in surveillance videos based on visual SLAM is characterized by including the following steps:
[0009] S1. Three-dimensional reconstruction based on visual SLAM: Construct a three-dimensional point cloud map of the scene and record the data required for external parameter calibration. Use the RGBD image stream provided by the depth camera and the inertial sensor data to construct a three-dimensional point cloud map of the scene according to the visual odometry and save it to a file; during the three-dimensional reconstruction process, photograph a checkerboard calibration board and record the position and attitude of the camera at this time.
[0010] S2. Calibration: The surveillance camera takes a surveillance picture of a checkerboard calibration board, and combines the calibration data provided during the three-dimensional reconstruction in step S1 to calculate the position and attitude of the surveillance camera in the three-dimensional point cloud map.
[0011] S3. Position tracking and calculation of pedestrians in the surveillance camera: Track the personnel, and based on the position and attitude data obtained in step S2 and the surveillance video stream provided by the surveillance camera, identify the position of the person in the image and give the position of the person in the three-dimensional point cloud map.
[0012] The process flow of position tracking and calculation of pedestrians in the surveillance camera is as follows:
[0013] S31. According to the type of the own surveillance camera and whether it has moved, select whether to enter the attitude correction process. If yes, enter step S32; otherwise, enter step S33.
[0014] S32. If entering the attitude correction process, extract the vanishing point in the image and compare the positions of the vanishing points. Determine whether the surveillance picture has rotated according to whether the vanishing point has moved, and update the rotation.
[0015] S33. Enter the positioning process. First, take a frame of surveillance video image.
[0016] S34. Perform target detection on the pedestrians in the image to obtain the target box coordinates.
[0017] S35. Perform object tracking on the detected target bounding box and give the corresponding personnel position coordinates.
[0018] S36. Calculate the spatial position coordinates of the personnel based on the calibration parameters of the surveillance camera.
[0019] S37. If the positioning is not terminated, obtain the next frame of the image and return to step S33; otherwise, end.
[0020] Furthermore, in the process of constructing the three-dimensional point cloud map of step S1, add a point cloud map processing thread, which is used to receive the pose information and RGBD image frames of each incoming camera frame, and output an accurate point cloud map. The specific process is as follows:
[0021] S141. Filter the pose information and RGBD image frames of each incoming depth camera frame. When the camera angle change between the current frame and the previous selected frame is greater than 10° and the displacement change is greater than 2 meters, select the current frame and perform subsequent point cloud map generation operations.
[0022] S142. Calculate the point cloud block of the current frame and rotate it to the unified world coordinate system.
[0023] S143. Stitch and merge the point cloud blocks generated by all frames to obtain the overall point cloud map. Perform filtering and outlier removal on the point cloud map to compress the data volume of the point cloud map and optimize the visual perception of the map at the same time.
[0024] S144. When a loop occurs during the mapping process, ORB-SLAM3 re-optimizes the poses of the selected frames, re-stitch the point clouds, and perform point cloud processing operations again according to step S143.
[0025] Furthermore, the external parameter calibration method in step S2 is as follows:
[0026] S21. The surveillance camera takes pictures of the checkerboard calibration board: Select the origin position of the world coordinate system. The surveillance camera starts to move slowly towards the checkerboard calibration board from the origin position. During this process, use ORB_SLAM3 to estimate the pose of the surveillance camera in real time. When the surveillance camera moves in front of the checkerboard calibration board, close the program and save the photos taken by the current surveillance camera and the pose of the camera.
[0027] S22. Calibrate the internal parameters of the surveillance camera: Place the checkerboard calibration board within the range of the surveillance camera, move the checkerboard calibration board at multiple angles, record a video, extract the frames in the video, identify the checkerboard, and use the Zhang Zhengyou calibration method to calibrate the internal parameters and distortion of the surveillance camera.
[0028] S23. Extrinsic Calibration of Surveillance Camera: According to the actual three-dimensional position information of the target feature points and the two-dimensional positions of the target feature points in the image, the direct linear transformation method is used to solve the camera coordinate system and the target coordinate system, so as to calculate the relative position relationship between the surveillance camera and the three-dimensional point cloud map.
[0029] Further, the method for calculating the spatial position coordinates of the person in step S36 is as follows:
[0030] Obtain the optical center P of the camera model of the surveillance camera according to the pose transformation matrix of the surveillance camera ow =(X ow , Y ow , Z ow ):
[0031] P ow =-R T t#(21)
[0032] Take the midpoint M(u, v) of the lower bottom edge of the target box of the person to be located. Calculate the spatial position of point M on the normalized plane according to the projection equation. Assume the depth d is 1 meter, and solve the spatial coordinates P of point M on the normalized plane according to formula (22) m =(X m , Y m , Z m );
[0033]
[0034] According to the height h of the ground in the world coordinate system, give the equation of the plane where the ground is located as shown in formula (24), and convert the plane where the ground is located into the form of a point on the plane and a normal vector:
[0035] z = h#(23)
[0036] p0=(0, 0, h),
[0037] Write the ray as a parametric equation as shown in (25), where is the direction vector of the ray, t is the parameter t∈[0, ∞),
[0038]
[0039] Assume that the intersection point of the ray and the plane where the ground is located is P g , then there is:
[0040]
[0041] After sorting out formula (26), it is obtained:
[0042]
[0043] According to the distributive law of vector dot product, we have:
[0044]
[0045] From this, the intersection point is solved. The intersection point P g The coordinates are the three-dimensional coordinates of the pedestrian in the world coordinate system.
[0046] Furthermore, the above-mentioned method for positioning and tracking personnel in surveillance videos based on visual SLAM further includes step S4 visualization display: displaying the three-dimensional point cloud map of the scene in step S1 and the positions of the people in the three-dimensional point cloud map in step S3, and providing a GUI interface for user interaction and supervision. The specific content to be displayed includes multiple three-dimensional point clouds and their corresponding cameras and trajectories. When selecting a certain point cloud for display, other point clouds and their corresponding cameras and trajectories are hidden.
[0047] On the other hand, the present invention also provides a system for positioning and tracking personnel in surveillance videos based on visual SLAM, including the following modules to implement the above method steps:
[0048] A three-dimensional reconstruction module, which is used to construct a three-dimensional point cloud map of the scene and record the data required for external parameter calibration. It runs only once during the entire usage period. Based on the RGBD image stream provided by the depth camera and the inertial sensor data, it constructs a three-dimensional point cloud map of the scene according to visual odometry and saves it to a file for display by the visualization display module; during the three-dimensional reconstruction process, it takes pictures of the checkerboard calibration board and records the position and attitude of the camera at this time for use by the calibration module;
[0049] A calibration module, which is used to measure the position and attitude of each surveillance camera in the three-dimensional point cloud map. It runs only once during the entire usage period. The surveillance camera takes a picture of a checkerboard calibration board, and combines the calibration data provided during three-dimensional reconstruction to calculate the position and attitude of the surveillance camera in the three-dimensional point cloud map for use by the position calculation module;
[0050] A position calculation module, which is used to identify the position of the person in the image based on the position and attitude data obtained by the calibration module and the surveillance video stream provided by the surveillance camera, and give the position of the person in the three-dimensional point cloud map for display by the visualization display module;
[0051] A visualization display module, which is used to display the three-dimensional point cloud map of the scene and the positions of the people in the three-dimensional point cloud map, and provide a GUI interface for user interaction and supervision.
[0052] Compared with the prior art, the beneficial effects of the present invention are:
[0053] The personnel positioning and tracking method for surveillance videos based on visual SLAM in the present invention has three major modules: mapping, calibration, and monitoring, realizing the ability expansion from 2D monitoring to 3D monitoring. Meanwhile, a display system suitable for this monitoring scenario is designed, and optimizations are made in terms of system point cloud display and interface video and trajectory output.
[0054] 1. The present invention proposes a method for constructing a three-dimensional map based on visual SLAM and inertial odometer. The three-dimensional point cloud map generated by reconstruction is optimized through three main processes: tracking, local mapping, and loop closure optimization, so as to achieve better terrain reconstruction of the scenes inside the building, enabling the monitoring personnel to quickly understand the map and accurately locate.
[0055] 2. To solve the problem that most surveillance cameras in the actual environment have been fixed at a certain position and it is impossible to use the calibration method of moving cameras commonly used in visual SLAM to calibrate them, that is, it is impossible to accurately obtain their specific representation in the computer of their positions in the environment. The present invention provides an external parameter calibration method for general surveillance cameras and environmental three-dimensional point clouds. By using a checkerboard calibration board, it can meet the requirements of camera external parameter calibration in various environments.
[0056] 3. Through the depth camera parameter calibration method of the present invention, the position information related to the depth camera is obtained, and by only processing the video taken by an ordinary general camera, the person's trajectory can be accurately calculated. At the same time, each person is tracked and drawn in the already generated point cloud map, and the corresponding information of the trajectory is saved and displayed on the interface, and it can automatically eliminate the target tracking trajectory when it times out.
[0057] 4. Through the calculation and trajectory optimization of the personnel position, the present invention proposes a person position calculation and tracking algorithm based on a three-dimensional vision algorithm. The algorithm first uses a deep neural network to identify the people in the surveillance image, and based on the previously calibrated position relationship, calculates the three-dimensional position of the people based on the principle of inverse perspective transformation according to the prior that all pedestrians appearing in the surveillance are on the ground.
[0058] 5. Based on the above three-dimensional point cloud map and graphical interface, the present invention also provides expandable functions, which can identify the flames or smoke appearing in the video image, adapt to the in-building surveillance requirements of real scenarios, and ensure that the system can automatically detect dangers and give alarms in a timely manner when abnormal situations occur.
[0059] 6. The present invention designs a data transmission and display method to improve the trajectory drawing method in the three-dimensional point cloud map, format and transmit the point cloud file in the background to the system, and establish a 3D display interface based on the Three.JS graphics library that can display three-dimensional point clouds, camera models, and trajectories, enabling monitoring personnel to understand the output results simply, clearly, and comprehensibly, facilitating the adoption of corresponding measures. The designed interface has good versatility and can support display on multiple devices.
[0060] 7. The present invention combines the advantages of visual SLAM in map reconstruction, designs and develops a 3D intelligent monitoring system, which can comprehensively solve the pain points of traditional two-dimensional monitoring, has the advantages of intuitive and efficient, low cost, and easy deployment. In this system, a 3D point cloud map can be displayed, the three-dimensional layout of the entire building can be monitored, the function of target tracking can be realized with one key, the trajectories of people entering the building can be automatically calculated and drawn in the point cloud map, and the clear and accurate three-dimensional display of real-time trajectories and the query of relevant information in the trajectory list can be achieved. It can also switch to view the monitoring images of each floor with one key. Functions such as fire warning, smoke recognition, and behavior recognition can also be added.
[0061] 8. The system designed by the present invention simultaneously supports the calculation and display of the relevant poses of fixed monitoring cameras, which is convenient to be applied to existing monitoring scenarios and facilitates the deep integration of new generation information technologies such as big data and cloud computing under the background of Internet + with the monitoring system. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.
[0063] Figure 1 The monitoring video personnel positioning and tracking system based on visual slam provided by the embodiment of the present invention;
[0064] Figure 2 The flowchart of the three-dimensional reconstruction module provided by the embodiment of the present invention;
[0065] Figure 3 The flowchart of the point cloud map calibration provided by the embodiment of the present invention;
[0066] Figure 4 The corresponding relationship between the calibration board images captured by the RGBD camera and the monitoring camera and the extracted feature points provided by the embodiment of the present invention;
[0067] Figure 5 The target detection, tracking, and position calculation process provided by the embodiment of the present invention;
[0068] Figure 6 The vanishing point extraction diagram provided by the embodiment of the present invention;
[0069] Figure 7 The schematic diagram of the typed array provided by the embodiment of the present invention;
[0070] Figure 8 The sample diagram of 3D point cloud loading provided by the embodiment of the present invention;
[0071] Figure 9 The sample diagram of the camera model provided by the embodiment of the present invention;
[0072] Figure 10 The flow chart of trajectory update drawing provided by the embodiment of the present invention;
[0073] Figure 11 The system data transmission diagram provided by the embodiment of the present invention;
[0074] Figure 12 The display of trajectory-related functions of the video surveillance system based on 3D vision visualization provided by the embodiment of the present invention;
[0075] Figure 13 The overall interface diagram of the video surveillance system based on 3D vision visualization provided by the embodiment of the present invention, as well as the display of switch function and floor switching function;
[0076] Figure 14 The display of flame and smoke alarm functions of the video surveillance system based on 3D vision visualization provided by the embodiment of the present invention. Detailed implementation manners
[0077] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.
[0078] The embodiment of the present invention provides a monitoring video personnel positioning and tracking system based on visual slam. The hardware devices adopt Kinect for Azure depth camera, checkerboard calibration board, monitoring camera, personal computer or laptop. The system includes the following modules (as Figure 1 shown):
[0079] The 3D reconstruction module is used to construct a 3D point cloud map of the scene and record the data required for external parameter calibration. It runs only once during the entire usage period. It constructs a 3D point cloud map of the scene from the RGBD image stream provided by the depth camera and the inertial sensor data according to visual odometry and saves it to a file for display by the visualization module; during the 3D reconstruction process, it captures a checkerboard calibration board and records the position and pose of the camera at this time for use by the calibration module.
[0080] The calibration module is used to measure the position and pose of each monitoring camera in the 3D point cloud map. It runs only once during the entire usage period. The monitoring camera takes a photo of a checkerboard calibration board in the monitoring screen, and combines the calibration data provided during 3D reconstruction to calculate the position and pose of the monitoring camera in the 3D point cloud map for use by the position calculation module.
[0081] The position calculation module is used to identify the position of the person in the image based on the position and pose data obtained by the calibration module and the monitoring video stream provided by the monitoring camera, and give the position of the person in the 3D point cloud map for display by the visualization module.
[0082] The visualization module is used to display the 3D point cloud map of the scene and the position of the person in the 3D point cloud map, and provides a GUI interface for user interaction and supervision.
[0083] The system consists of the above four modules. Among them, the 3D reconstruction module, the calibration module, and the position calculation module have innovations at the technical level, and the visualization module has innovations at the product level.
[0084] Based on the above system, an embodiment of the present invention proposes a method for positioning and tracking personnel in a monitoring video based on visual slam, and the specific steps are as follows:
[0085] S1. 3D reconstruction based on visual slam: Construct a 3D point cloud map of the scene and record the data required for external parameter calibration. Construct a 3D point cloud map of the scene from the RGBD image stream provided by the depth camera and the inertial sensor data according to visual odometry and save it to a file; during the 3D reconstruction process, capture a checkerboard calibration board and record the position and pose of the camera at this time.
[0086] During the 3D reconstruction process, the device parameter calibration uses the IMU_utils tool, the Kalibr toolbox, and the QR code calibration board to measure the parameters of the depth camera, including: the internal parameters of the depth camera, the IMU parameters, the time offset between the IMU and the depth camera, and the pose transformation relationship between the IMU and the depth camera coordinate systems.
[0087] The 3D reconstruction module is adapted to the Kinect for Azure depth camera based on the open-source ORB_SLAM3 framework, enabling it to be used normally under actual working conditions. The construction process of the 3D point cloud map of step S1 scene is as follows Figure 2 shown, and the specific modified content is:
[0088] S11. Downsample the RGBD image stream provided by the depth camera, and the resolution is reduced from 1280×720 to 960×540;
[0089] S12. Align the time stamps of the incoming data stream according to the calibration parameters saved in the file;
[0090] S13. Change the Z-axis direction distance estimation of feature points from triangulation to depth map reading;
[0091] S14. Add a point cloud map processing thread to reconstruct the dense point cloud of the scene terrain.
[0092] The visual odometer only provides the position and pose of the depth camera. Therefore, it is necessary to reconstruct the dense point cloud of the scene terrain. It is divided into two steps. First, the point cloud is stitched out according to the pose of the depth camera and the depth camera model, and then the point cloud is processed to remove duplicate points and redundant points. In this system, the above functions are implemented through the point cloud map processing thread, and the flow chart is as follows Figure 2 shown. The function of the point cloud map processing thread is to receive the pose information and RGBD image frames (including color map and depth map) of each frame of the incoming depth camera, and output an accurate point cloud map. Step S14 of the point cloud map processing thread includes:
[0093] S141. Screen the pose information and RGBD image frames of each frame of the incoming depth camera. When the camera angle change between the current frame and the previous selected frame is greater than 10° and the displacement change is greater than 2 meters, select the current frame and perform subsequent point cloud map generation operations;
[0094] S142. Calculate the point cloud block of the current frame according to the depth camera model formula and rotate it to the unified world coordinate system;
[0095] S143. Stitch and merge the point cloud blocks generated by all frames to obtain the overall point cloud map, and perform filtering and outlier removal on the point cloud map to compress the data volume of the point cloud map and optimize the visual perception of the map;
[0096] S144. When a loop occurs during the mapping process, ORB-SLAM3 re-optimizes the poses of the selected frames, stitches the point cloud again, and performs point cloud processing operations again according to step S143.
[0097] The detailed process of converting the RGBD image frames obtained by the depth camera into point cloud blocks and stitching them into the world coordinate system is as follows:
[0098] For a certain pixel point p in the image, let its coordinates be (u, v), and the coordinates of the three-dimensional point P projected from p in the point cloud be (X c , Y c , Z c ). Then the generation formula of the points in the point cloud is as follows:
[0099]
[0100] Among them, f x , f y , c x , c y represent the internal parameters of the camera, which are obtained through the device parameter calibration method given above, and d represents the depth of the point p read from the depth map.
[0101] Through ORB-SLAM3, the rotation and translation matrices R cw , t cw of the camera in the world coordinate system are obtained and merged together:
[0102]
[0103] Then the coordinates X w , Y w , Z w of the point P in the world coordinate system can be obtained as:
[0104]
[0105] Based on the above method, all pixel points in the image can be converted into three-dimensional points in the world coordinate system. The advantage of this approach is that after the loop detection thread of ORB-SLAM corrects the key frames, the correction results can be applied to the dense map. Compared with directly stitching the point cloud, this approach can eliminate the cumulative error.
[0106] In the ORB_SLAM3 framework, since the world coordinate system is based on the position of the camera at initialization, the generated point cloud map will inevitably be skewed, which will have a negative impact on subsequent map visualization and subsequent extrinsic parameter calibration. Therefore, it is necessary to calibrate the coordinate system of the point cloud map. As Figure 3 shown, the calibration process of the three-dimensional point cloud map is as follows:
[0107] 1) Calculate the plane equation of the point cloud ground:
[0108] Using the plane detection method based on RANSAC in the PCL library, the plane equation of the ground ax + by + cz + d = 0 is obtained, where a, b, c, and d are the four parameters of the plane equation. The normal vector of the ground can be obtained as (a, b, c) through the plane equation.
[0109] 2) Calculate the rotation matrix between the point cloud ground and the horizontal plane of the coordinate system:
[0110] Calculate the rotation matrix between the point cloud ground and the horizontal plane of the coordinate system. Through this matrix, the rotation transformation of the point cloud map can be realized, thereby calibrating the point cloud map. The calculation method is as follows:
[0111] (1) Let the normal vector of the point cloud ground be v1 = (a, b, c), and the normal vector of the horizontal plane of the coordinate system be v2 = (0, 0, 1). Calculate the rotation axis n and the rotation angle θ of the rotation transformation between the two vectors. The calculation formula is as follows:
[0112]
[0113]
[0114] (2) Calculate the rotation matrix R from the rotation axis n and the rotation angle θ. The calculation formula is as follows:
[0115] R = cosθI + (1 - cosθ)nn T + sinθn^
[0116] Where the ^ symbol is the conversion symbol from a vector to an anti-symmetric matrix. Let the vector a = (a1, a2, a3), then the specific conversion formula is as follows:
[0117]
[0118] 3) Calibrate the point cloud map using the rotation matrix:
[0119] Assume that a point in the point cloud map to be calibrated is p0 = (x0, y0, z0), and the coordinates of this point after calibration are p1 = (x1, y1, z1). The conversion formula is as follows:
[0120] p1 = Rp 0
[0121] Applying this formula to all points in the point cloud map to be calibrated can achieve the calibration of the entire point cloud map, and finally achieve the alignment of the ground and the horizontal plane of the coordinate system.
[0122] S2. Calibration: Measure the position and attitude of each monitoring camera in the three-dimensional point cloud map. The monitoring camera takes a monitoring picture of a checkerboard calibration board, and calculates the position and attitude of the monitoring camera in the three-dimensional point cloud map in combination with the calibration data provided during the three-dimensional reconstruction in step S1.
[0123] This step mainly solves the following problem: for the surveillance cameras already installed in a building, how to obtain the coordinates of these cameras in the world coordinate system. The calibration process of step S2 is as follows:
[0124] S21. The surveillance camera captures a checkerboard calibration board:
[0125] After performing 3D reconstruction on the environment using an RGBD camera, a point cloud map of the environment is obtained. To use the surveillance camera to complete the personnel positioning task, it is also necessary to measure the relative position relationship between the surveillance camera and the point cloud map, that is, the extrinsic parameters. A calibration board is used as a marker to calculate the extrinsic parameters.
[0126] The role of the calibration board is to provide the camera with multiple 3D points with a fixed relative position relationship, so as to calculate the required parameters. Commonly used calibration boards for calibration include checkerboard calibration boards and QR code calibration boards. The QR code calibration board has the advantage that it can be distinguished after being upside down in the up-down, left-right directions. At the same time, due to the complex graphics, it poses higher requirements for lighting and camera resolution. Under the usage conditions of this article, considering that all surveillance cameras are facing the surveillance scene directly and there is no situation of being upside down, and in order for the calibration method to be adapted to lower-cost surveillance cameras and to arrange the system of this article under darker indoor environmental light conditions, the checkerboard calibration board is more suitable for the usage environment of this invention.
[0127] Modify the code of the ORB_SLAM3 framework to implement saving the photo captured by the current camera and the camera pose at a certain moment. During actual operation, select the origin position of the world coordinate system, and the surveillance camera starts to move slowly towards the checkerboard calibration board from the origin position. During this process, use ORB_SLAM3 to estimate the pose of the surveillance camera in real time. When the surveillance camera moves in front of the checkerboard calibration board, close the program and save the photo captured by the current surveillance camera and the camera pose;
[0128] S22. Intrinsic calibration of the surveillance camera:
[0129] Place the checkerboard calibration board within the range of the surveillance camera, move the checkerboard calibration board at multiple angles, record a video, extract the frames from the video, identify the checkerboard, and use Zhang Zhengyou's calibration method to calibrate the intrinsic parameters and distortion of the surveillance camera.
[0130] S23. Extrinsic calibration of the surveillance camera:
[0131] The calculation of the pose is a method to solve the transformation between the camera coordinate system and the target coordinate system based on the actual three-dimensional position information of the target feature points and the two-dimensional positions of the target feature points in the image. If the three-dimensional coordinates of the feature points on the calibration board are known, this paper will use the direct linear transformation method to construct an augmented matrix [R|T] with 12 unknowns to represent the transformation between the camera coordinate system and the target coordinate system. Select at least 6 pairs of corresponding point pairs of known three-dimensional space point coordinates and two-dimensional pixel point coordinates to solve the unknowns in the augmented matrix, so as to realize the calculation of the camera pose.
[0132] Suppose the homogeneous coordinate corresponding to a feature point P1 on the calibration board in space is P1 = (X, Y, Z, 1) T , and the corresponding two-dimensional point of this point in the image of the monitoring camera is denoted as x1 = (u1, v1, 1) T , according to the calculation method of direct linear transformation, assume that the 3×4 augmented matrix [R|T] is expanded into the form:
[0133]
[0134] In the above formula, u1, v1 are the pixel coordinates of a certain two-dimensional point in the image of the monitoring camera, X, Y, Z are the three-dimensional coordinates of the corresponding point, and s is the scale.
[0135] According to the linear transformation of the matrix, use the last row to eliminate the scale coefficient s, and get:
[0136]
[0137]
[0138] To simplify the representation, each horizontal row of the augmented matrix in formula (1) is represented by vectors t1, t2, t3:
[0139]
[0140] Then the equation represented by the vector can be obtained as shown in the following formula:
[0141]
[0142]
[0143] In formulas (5)(6), t is the vector to be solved. According to (5)(6), it can be seen that each feature point contains a constraint equation with two unknowns. If there are N pairs of corresponding point pairs of three-dimensional coordinates and two-dimensional coordinates, the characteristic equation can be written in the form shown in the following formula:
[0144]
[0145] According to Equation (7), there are a total of 12 unknowns. Therefore, at least six pairs of corresponding point pairs are required to solve the above equation. The checkerboard calibration board used in this system has 42 pairs of corner point coordinates, making the above equation an overdetermined equation, and the SVD method is used to perform least squares solution of the equation.
[0146] For the 42 three-dimensional points P on the checkerboard calibration board and their projections p on the normalized plane, the pose R and t of the camera were previously calculated using the direct linear transformation method, and its Lie algebra representation is ξ. Assume that the spatial coordinates of a certain calibration board corner point in space are P i =[X i , Y i , Z i T , and its projected coordinates u i =[u i , v i T . The relationship between the pixel coordinates and the spatial point positions is as shown in the formula:
[0147]
[0148] Written in matrix form, it is:
[0149] s i u i =Kexp(ξ^)P i
[0150] Due to the unknown camera pose and the noise of the observation points, there is an error in this equation. Sum the errors to construct a least squares problem, and then find the best camera pose to minimize it. The summation of the error terms is as shown in the formula:
[0151]
[0152] Solve it through the Gauss-Newton algorithm to obtain the camera pose transformation matrix when the error term is minimized.
[0153] As described above, the camera pose can be solved on the premise of knowing the coordinates of the feature points of the checkerboard calibration board. Due to the scale uncertainty of the monocular camera, it is impossible to determine the relative position relationship between each feature point only based on the size and style of the calibration board observed in the image. According to the actual situation of the system of the present invention, as Figure 4 shown, the scale information can be obtained from two aspects:
[0154] (1) The absolute Z-axis distance of the feature point from the optical center of the RGBD camera can be obtained from the depth map, so as to obtain the absolute position information of each feature point.
[0155] (2) Measure the size of the checkerboard on the calibration board to obtain the relative position relationship between each feature point on the checkerboard.
[0156] Taking the first feature corner point in the upper left corner of the checkerboard calibration board as the coordinate origin, with the horizontal direction as the x-axis, the vertical direction as the y-axis, and the vertically upward direction as the z-axis, a calibration board coordinate system is formed. A method for solving the camera pose by knowing the three-dimensional coordinates of the feature points on the checkerboard calibration board can solve the pose transformation matrix of the camera in the calibration board coordinate system. When using a surveillance camera to photograph the checkerboard calibration board, the transformation matrix in the calibration board coordinate system is denoted as T mb , when using an RGBD camera to photograph the checkerboard calibration board during the operation of visual odometry, the transformation matrix in the calibration board coordinate system is denoted as T cb , and at this time, the transformation matrix of the RGBD camera relative to the world coordinate system given by the visual odometry is T cw .
[0157] Let the coordinates of a certain feature point on the checkerboard calibration board in the calibration board coordinate system be P b =(X, Y, 0, 1) T , the coordinates of this point in the surveillance camera coordinate system are denoted as P m , and the coordinates in the RGBD camera coordinate system are denoted as P c , and the coordinates in the world coordinate system of the visual odometry are denoted as P w Then there is:
[0158] P c =T cb P b (10)
[0159] P m =T mb P b (11)
[0160] Multiply both sides of equation (10) on the left by T cb -1 Get:
[0161] P b =T cb -1 P c (12)
[0162] Substitute equation (12) into (11) to get:
[0163] P m =T mb T cb -1 P c (13)
[0164] According to the definition of Euclidean transformation, T in (13) mb T cb -1Represents the transformation method from the RGBD camera coordinate system to the monitoring camera coordinate system, so it is denoted as T mb T cb -1 as T mc .
[0165] The transformation matrix T from the world coordinate system to the monitoring camera coordinate system can be solved from Equation (14) mw .
[0166] T mw = T mc T cw (14)
[0167] S3. Position calculation: Based on the position and attitude data obtained in step S2 and the monitoring video stream provided by the monitoring camera, identify the position of the person in the image and give the position of the person in the three-dimensional point cloud map.
[0168] Personnel positioning and trajectory tracking based on the monitoring video have the advantages of high accuracy and good stability. The pedestrian detection and positioning algorithm based on the monitoring video will be introduced in detail below.
[0169] As Figure 5 shown, the position tracking and calculation process of pedestrians in the monitoring camera is as follows:
[0170] S31. According to the type of the monitoring camera itself and whether it has moved, select whether to enter the attitude correction process. If so, enter step S32; otherwise, enter step S33;
[0171] S32. If entering the attitude correction process, extract the vanishing point in the image and compare the positions of the vanishing points. Determine whether the monitoring screen has rotated based on whether the vanishing point has moved, and update the rotation;
[0172] Specifically, regarding the attitude correction of the monitoring camera, some monitoring cameras have the function of manual or automatic rotation, which can cause changes in the yaw angle and pitch angle of the camera attitude. Considering the positioning system described in this article, the workload of the pre-calibration work of the monitoring camera in the early stage is relatively large. It is unrealistic to re-calibrate the external parameters every time the attitude of the monitoring camera changes. Therefore, an automatic update and correction attitude workflow needs to be provided.
[0173] 1) Vanishing point extraction
[0174] According to the principle of projective geometry, in the case of perspective distortion, a group of parallel lines in the real world will intersect at an infinite point, and the projection of the intersection point on the imaging plane is called the vanishing point. When the parallel lines in the real world are parallel to the imaging plane, the vanishing point is located at infinity on the imaging plane. However, when there is a non-parallel relationship between the group of parallel lines and the imaging plane, the vanishing point will be located within a finite distance on the imaging plane, or even within the imaging area.
[0175] The vanishing point has some important properties:
[0176] a. In the real world, lines that are parallel to each other and lines that are parallel to each other all point to the same vanishing point;
[0177] b. The vanishing point corresponding to a straight line must be located in the direction of the projection ray of the straight line on the image plane;
[0178] c. The position of the vanishing point is independent of the roll angle and is only related to the pitch angle and yaw angle.
[0179] The vanishing point is an important feature formed on the image plane after perspective projection, which can provide a large amount of structural information and direction information for scene analysis or be used to measure the parameters of the camera itself. Therefore, the vanishing point has a wide range of applications in rectangular structure estimation and matching, 3D reconstruction, camera calibration, and azimuth estimation.
[0180] In view of the use of this system inside buildings, most building walls and floors are straight, so fixed vanishing points can be extracted. Determine whether the monitoring screen rotates based on whether the vanishing point moves. First, extract the vanishing point, and the process is as follows:
[0181] (1) Take the internal parameters and distortion of the monitoring camera, undistort the original image, and obtain the undistorted image;
[0182] (2) Use the LSD line segment extractor to extract line segments on the undistorted image;
[0183] (3) Screen the line segments according to the length, and take those longer than 60 pixels as valid line segments;
[0184] (4) Use the Hough transform to calculate the angle of each line segment;
[0185] (5) Cluster the line segments according to the angle and divide them into three categories;
[0186] (6) Use the least squares method to find the nearest point for each category of straight lines;
[0187] (7) Select the one with the smallest sum of coordinates among the three vanishing points as the reference point;
[0188] As Figure 6 shown, red, green, and blue are the three categories of line segments, and the specific position of the reference point is at the center of the red circle.
[0189] 2) Pose correction
[0190] As Figure 6As shown in the figure, assume that the coordinates of the vanishing point before rotation are (x0, y0), and the coordinates of the vanishing point after rotation are (x1, y1). Then, the change in yaw angle δyaw and the change in pitch angle δpitch can be calculated as shown in equations (15) and (16).
[0191] δyaw = arctan(x1 - x0) #(15)
[0192] δpitch = arctan(y1 - y0) #(16)
[0193] Since the pose of the surveillance camera is represented by the transformation matrix T mw Therefore, the change in angle is converted into the same rotation represented in matrix form. The rotation in the yaw angle direction is denoted as R x , and the rotation in the pitch angle direction is denoted as R y As shown in equations (17) and (18):
[0194]
[0195]
[0196] Given that the camera can only rotate and cannot move, it can be considered that the translation amount of the optical center coordinates is approximately zero. Multiply the rotation matrices in the two directions and then combine them with the translation vector t = (0, 0, 0) T to form the transformation matrix T 01 , which represents the conversion from the camera coordinates before rotation to the camera coordinates after rotation. The transformation matrix of the new surveillance camera required for positioning is denoted as T mw ′, then there is:
[0197] T mw ′ = T 01 T mw #(19)
[0198] S33. Enter the positioning process. First, take a frame of surveillance video image;
[0199] S34. Perform target detection on the pedestrians in the image to obtain the target box coordinates;
[0200] S35. Perform target tracking on the detected target box and give the corresponding personnel position coordinates;
[0201] S36. Calculate the spatial position coordinates of the personnel based on the calibration parameters of the surveillance camera;
[0202] Specifically, the position calculation process is as follows:
[0203] 1) Target recognition and tracking
[0204] Detect and track multiple targets in the video stream based on the SORT (Simple Online And Realtime Tracking) algorithm, and display the id of each target. This algorithm uses a powerful CNN detector - yolov3 to detect targets, and then uses the Kalman filter and the Hungarian algorithm to track the detected targets. This algorithm can achieve accurate multi-person tracking while meeting the real-time requirements.
[0205] 2) Pedestrian position calculation
[0206] The world coordinate system is the coordinate system of the point cloud map obtained by three-dimensional reconstruction, that is, with the position of the optical center of the first frame as the origin, and the Z-axis direction is opposite to the gravity direction. The transformation matrix T from the world coordinate system to the monitoring camera coordinate system mw Consists of a 3×3 rotation matrix R and a 1×3 translation vector t. Through actual measurement, the height of the ground in the world coordinate system is h, and the internal parameter matrix K of the camera is calibrated through a calibration board.
[0207] According to the projection equation (20), the coordinates X of the pedestrian on the ground w 、Y w Can be solved:
[0208]
[0209] The specific solution process is as follows:
[0210] Based on the pose transformation matrix of the monitoring camera, the optical center P of the camera model of the monitoring camera is obtained ow =(X ow ,Y ow ,Z ow ):
[0211] P ow =-R T t#(21)
[0212] Take the midpoint M(u, v) of the lower bottom edge of the target box of the person to be located. According to the projection equation, calculate the spatial position of point M on the normalized plane. Assume the depth d is 1 meter, and solve the spatial coordinates P of point M on the normalized plane according to formula (22) m =(X m ,Y m ,Z m );
[0213]
[0214] Write the equation of the plane where the ground is located according to the height h of the ground in the world coordinate system, as shown in Equation (24), and convert the plane where the ground is located into the representation of a point on the plane and a normal vector:
[0215] z = h#(23)
[0216] p0 = (0, 0, h),
[0217] Write the ray in parametric equation as shown in (25), where is the direction vector of the ray, t is the parameter t ∈ [0, ∞).
[0218]
[0219] Let the intersection point of the ray and the plane where the ground is located be P g , then there is:
[0220]
[0221] After sorting Equation (26), we get:
[0222]
[0223] According to the distributive law of vector dot product, we have:
[0224]
[0225] From this, the intersection point The coordinates of the intersection point P g are the three-dimensional coordinates of the pedestrian in the world coordinate system.
[0226] S37. If the positioning is not terminated, obtain the next frame of image and return to step S33, otherwise end.
[0227] S4 Visualization display: Display the three-dimensional point cloud map of the scene in step S1 and the position of the person in the three-dimensional point cloud map in step S3, and provide a GUI interface for user interaction and supervision.
[0228] The system designed and implemented by the present invention includes a 3D display interface based on the Three.JS graphics library that can display three-dimensional point clouds, camera models, and trajectories. Its functional logic is as follows.
[0229] 1) Three-dimensional point cloud loading
[0230] The present invention can load multiple three-dimensional point cloud PCD files located on the server side and correctly display them in the 3D display interface. Its functional logic is as follows:
[0231] (1) Find all PCL files in the specified directory to obtain the PCL file names;
[0232] (2) According to the file names, use the FileLoader built into JavaScript to read the PCD files to obtain the 3D point cloud data, such as the number of points, the point coordinate set, the color set, etc.;
[0233] (3) Remap the color gamut of the point cloud color set to the gamut range supported by the Three.Js graphics library. The formula is as follows;
[0234]
[0235] (4) Combine the typed arrays in JavaScript with the data classes in Three.Js, and fill the 3D point cloud data into the graphics class objects provided by Three.Js correspondingly; The typed arrays are as Figure 7 shown.
[0236] (5) Name each created graphics class object with the file name and add it to the display interface;
[0237] (6) According to the system settings, only display the point clouds for default viewing, and leave other point clouds for backup.
[0238] So far, the functional logic for loading 3D point clouds in the system has been sorted out, Figure 8 which is a sample for 3D point cloud loading.
[0239] 2) Camera model loading
[0240] The present invention can add a camera model with the correct orientation at the correct position in the 3D display interface according to the camera calibration parameters located on the server side, making the overall layout clear at a glance. Its functional logic is as follows:
[0241] (1) Read the specified parameter xml file to obtain the number of cameras and the corresponding names and transformation matrices;
[0242] (2) Use the FileLoader built into JavaScript to read the camera model obj file, and create the specified number of graphics class objects in Three.Js accordingly;
[0243] (3) Adjust the world coordinates and orientations of each graphics class object according to the transformation matrix and name them;
[0244] (4) Add each adjusted graphics class object to the display interface.
[0245] So far, the functional logic for loading the camera model in the system has been sorted out, Figure 9 which is a sample diagram of the loaded camera model, and its background isFigure 8 Point cloud
[0246] The specific content shown includes multiple 3D point clouds and their corresponding cameras and trajectories. When selecting a point cloud to display, other point clouds and their corresponding cameras and trajectories are hidden.
[0247] 3) Trajectory information transmission and drawing optimization
[0248] The server of the present invention can receive Socket information, and based on this, realizes the transmission and drawing functions of the obtained trajectory information, and solves the problem that the line thickness of the current Web graphics library cannot be adjusted. The function flow is as Figure 10 shown, and the function logic is as follows:
[0249] (1) Receive Socket information to obtain trajectory-related information in real time, including trajectory ID, operation type, and trajectory data;
[0250] (2) If the operation type is trajectory update: first search for the corresponding ID in the existing trajectory library. If not found, create a new Three.Js line object, fill in the corresponding trajectory data, and add it to the display interface; if found, update the existing Three.Js line object with the newly obtained trajectory data and redraw it on the display interface;
[0251] (3) If the operation type is trajectory deletion, search for the corresponding ID in the existing trajectory library and delete the Three.Js line object corresponding to this ID on the display interface;
[0252] Flickering problem of trajectory update: When drawing a trajectory, the update function provided by Three.Js will cause the trajectory to flicker continuously. The reason is that the trajectory is continuously deleted and redrawn during the update. To solve this problem, this patent adopts the "double-trajectory coverage" method, specifically: retain the old trajectory when processing the newly obtained trajectory information, and delete the old trajectory from the display interface after the new trajectory is covered and drawn. In this way, both the display effect of the trajectory can be guaranteed and the memory occupancy of the system can be reduced.
[0253] Thickness problem of the trajectory: When drawing a trajectory, the line object provided by Three.Js cannot adjust the thickness, resulting in the trajectory being difficult to identify on the display interface. To solve this problem, this patent proposes the method of "copying and partially overlapping a single trajectory", specifically: after updating the coordinate data of the Three.Js line object, copy the object multiple times when drawing on the display interface and gradually translate it around, successfully achieving the effect of thickening the trajectory and passing the performance test.
[0254] (4) Object binding and perspective operation
[0255] As described above, the 3D display interface of the present invention can display multiple point clouds and their corresponding cameras and trajectories, and can ensure that when a certain point cloud is selected for display, other point clouds and their corresponding cameras and trajectories are hidden. This patent uses the object binding operation provided by the Three.Js graphics library to correspond the cameras and trajectories to the specified point cloud according to the ID, thereby realizing the function of switching the point cloud scene.
[0256] In addition, the 3D display interface of the present invention realizes the operations of panning, rotating, and zooming of the viewing angle based on the event listener mechanism of Web and JavaScript, and has the ability to adapt to multiple terminals.
[0257] Regarding the GUI system interface, the overall system architecture is as Figure 11 shown.
[0258] This system adopts a front-end and back-end separation architecture. The front end uses the Vue development framework to build a visualization display platform, providing visualization functions such as 3D point clouds, pedestrian trajectories, and real-time monitoring.
[0259] The back end is implemented based on the SpringBoot development framework, providing calculation and storage functions such as monitoring management, pedestrian counting, pedestrian trajectories, and log management for the front-end visualization. The front end and the back end communicate with each other through Http and WebSocket. The descriptions of each function are as follows:
[0260] Monitoring management: Manage the monitoring access to the system, mainly including monitoring configuration management and status management.
[0261] Pedestrian trajectory: Use the pedestrian positioning and tracking algorithm to obtain the pedestrian trajectory coordinates in real time, and realize the visualization of the pedestrian trajectory and the persistent storage of the historical trajectory.
[0262] Pedestrian counting: Calculate the number of pedestrians according to the number of trajectories in the current scene and display it in real time on the visualization interface.
[0263] Log management: Record all behaviors generated by the system and provide a query function to help improve the system security.
[0264] The back end of this system uses the Mysql database to realize data persistence, which can provide efficient data reading and writing; use the kafka message queue to integrate the pedestrian positioning and tracking algorithm into the system. The specific process is as follows:
[0265] (1) The pedestrian positioning and tracking algorithm analyzes the video frame, and sends the coordinates of the pedestrian in the three-dimensional coordinate system into the topic named slam in Kafka. Setting the number of partitions to 1 can ensure the order of messages;
[0266] (2) The system backend listens to Kafka in real time, reads the pedestrian trajectory coordinates calculated by the pedestrian positioning and tracking algorithm from it. On the one hand, it persists the pedestrian trajectory coordinates into the MySql database. On the other hand, it sends the pedestrian trajectory coordinates to the front end through WebSocket;
[0267] (3) Whenever the front end receives the pedestrian trajectory coordinates pushed by the backend once, it draws trajectory points in the 3D point cloud map according to the coordinates. When a certain number of trajectory points accumulate, an obvious pedestrian trajectory is formed.
[0268] Using Kafka to integrate the algorithm is beneficial to improving the scalability of the system, can flexibly expand the algorithm functions, and provides pluggable function implantation.
[0269] The front-end interface of the visualization video monitoring system of the present invention is as Figure 13 shown, and includes parts such as a 3D display interface, video monitoring, an abnormal item monitoring switch, a trajectory list, and an abnormal event list. The floor can be switched in the upper right corner, and after switching, the point cloud, camera, and trajectory on the left side and the video monitoring situation on the right side will be updated synchronously; in the lower left corner, functions of this patent such as target tracking and fire alarm can be selectively enabled.
[0270] When pedestrians are detected in the right video monitoring area and their trajectories are calculated, these trajectories will be updated and drawn synchronously on the left side, and corresponding table items will be updated in the lower trajectory list area below, as Figure 12 shown, the trajectories are clearly visible, and the position and orientation of the camera are shown using a camera model.
[0271] When a fire or smoke is detected in the video monitoring, an alarm prompt will pop up to facilitate the user of this patent to discover abnormal situations in time. At the same time, the corresponding table items will be updated in the abnormal event list in the lower right corner, as Figure 14 shown. The overall system runs smoothly and meets the performance requirements.
[0272] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.
[0273] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the apparatus embodiments, electronic device embodiments, computer-readable storage medium embodiments, and computer program product embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.
[0274] The above-described embodiments are only specific implementation manners of the present application, used to illustrate the technical solutions of the present application, rather than limiting it. The protection scope of the present application is not limited thereto. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present application can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements for some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application. All should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for positioning and tracking personnel in surveillance videos based on visual SLAM, characterized in that, It includes the following steps: S1. 3D reconstruction based on visual SLAM: Construct a 3D point cloud map of the scene and record the data required for extrinsic calibration. Based on the visual odometer, construct a 3D point cloud map of the scene from the RGBD image stream provided by the depth camera and the inertial sensor data, and save it to a file; During the 3D reconstruction process, take a checkerboard calibration board and record the position and pose of the camera at this time; S2. Calibration: The monitoring camera takes a photo of a checkerboard calibration board. Combining the calibration data provided during the 3D reconstruction in step S1, calculate the position and pose of the monitoring camera in the 3D point cloud map; S3. Position tracking and calculation of pedestrians in the monitoring camera: Track the personnel. Based on the position and pose data obtained in step S2 and the monitoring video stream provided by the monitoring camera, identify the position of the person in the image and give the position of the person in the 3D point cloud map; The process of position tracking and calculation of pedestrians in the monitoring camera is as follows: S31. According to the type of the monitoring camera itself and whether it has moved, select whether to enter the pose correction process. If yes, enter step S32; otherwise, enter step S33; S32. If entering the pose correction process, extract the vanishing point in the image and compare the positions of the vanishing points. Determine whether the monitoring screen has rotated based on whether the vanishing point has moved, and update the rotation; S33. Enter the positioning process. First, take a frame of monitoring video image; S34. Perform object detection on the pedestrians in the image to obtain the target box coordinates; S35. Perform object tracking on the detected target box and give the corresponding personnel position coordinates; S36. Calculate the spatial position coordinates of the personnel according to the calibration parameters of the monitoring camera; S37. If the positioning is not terminated, obtain the next frame of image and return to step S33; otherwise, end.
2. The method for positioning and tracking personnel in surveillance videos based on visual SLAM according to claim 1, characterized in that, In the process of constructing the 3D point cloud map in step S1, add a point cloud map processing thread to receive the pose information and RGBD image frame of each incoming camera, and output an accurate point cloud map. The specific process is as follows: S141. Screen the pose information and RGBD image frame of each incoming depth camera. When the camera angle change between the current frame and the previous selected frame is greater than 10° and the displacement change is greater than 2 meters, select the current frame and perform subsequent point cloud map generation operations; S142. Calculate the point cloud block of the current frame and rotate it to the unified world coordinate system; S143. Stitch and merge the point cloud blocks generated by all frames to obtain the overall point cloud map. Perform filtering and outlier removal on the point cloud map to compress the data volume of the point cloud map and optimize the visual perception of the map at the same time; S144. When a loop occurs during the mapping process, ORB-SLAM3 re-optimizes the poses of the selected frames, re-stitch the point clouds, and perform point cloud processing operations again according to step S143.
3. The method for positioning and tracking personnel in surveillance videos based on visual SLAM according to claim 1, characterized in that, The method for extrinsic calibration in step S2 is: S21. The surveillance camera captures the checkerboard calibration board: Select the origin position of the world coordinate system. The surveillance camera starts moving slowly towards the checkerboard calibration board from the origin position. During this process, ORB_SLAM3 is used to estimate the pose of the surveillance camera in real time. When the surveillance camera moves in front of the checkerboard calibration board, the program is closed, and the photo captured by the current surveillance camera and the pose of the camera are saved. S22. Intrinsic calibration of the surveillance camera: Place the checkerboard calibration board within the range of the surveillance camera. Move the checkerboard calibration board at multiple angles and record a video. Extract the frames from the video, identify the checkerboard, and use Zhang Zhengyou's calibration method to calibrate the intrinsic parameters and distortion of the surveillance camera. S23. Extrinsic calibration of the surveillance camera: According to the actual three-dimensional position information of the target feature points and the two-dimensional positions of the target feature points in the image, use the direct linear transformation method to solve the camera coordinate system and the target coordinate system, and calculate the relative position relationship between the surveillance camera and the three-dimensional point cloud map.
4. The method for positioning and tracking personnel in surveillance videos based on visual SLAM according to claim 1, characterized in that, The method for calculating the spatial position coordinates of the person in step S36 is as follows: Obtain the optical center P of the camera model of the surveillance camera based on the pose transformation matrix of the surveillance camera ow =(X ow , Y ow , Z ow ): P ow = -R T t (21) Take the midpoint M(u, v) of the lower bottom edge of the target box of the person to be located, calculate the spatial position of point M on the normalized plane according to the projection equation, assume the depth d is 1 meter, and solve the spatial coordinates P of point M on the normalized plane according to formula (22) m =(X m ,Y m ,Z m ); Based on the height h of the ground in the world coordinate system, the equation of the plane where the ground is located is given as shown in Equation (24), and the plane where the ground is located is converted into the representation of a point on the plane and the normal vector: z = h (23) Write the ray as a parametric equation as shown in (25), where is the direction vector of the ray, t is the parameter, t ∈ [0, ∞), Let the ray intersect the plane where the ground is located at point P g , then we have: After sorting out Equation (26), we get: According to the distributive law of vector dot product: Solve for the intersection point from this Intersection point P g The coordinates are the three-dimensional coordinates of the pedestrian in the world coordinate system.
5. The method for positioning and tracking personnel in a surveillance video based on visual SLAM according to claim 1, wherein, It also includes step S4 for visual display: Display the three-dimensional point cloud map of the scene in step S1 and the position of the person in the three-dimensional point cloud map in step S3, and provide a GUI interface for user interaction and supervision. The specific content to be displayed includes multiple three-dimensional point clouds and their corresponding cameras and trajectories. When selecting a certain point cloud for display, other point clouds and their corresponding cameras and trajectories are hidden.
6. A system for positioning and tracking personnel in a surveillance video based on visual SLAM, wherein, It includes the following modules to implement the method of any one of claims 1-5: The three-dimensional reconstruction module is used to construct the three-dimensional point cloud map of the scene and record the data required for extrinsic calibration. It runs only once during the entire usage period. Based on the RGBD image stream provided by the depth camera and the inertial sensor data, it constructs the three-dimensional point cloud map of the scene according to the visual odometer and saves it to a file for display by the visualization display module. During the three-dimensional reconstruction process, it captures the checkerboard calibration board and records the position and pose of the camera at this time for use by the calibration module. The calibration module is used to measure the position and pose of each surveillance camera in the three-dimensional point cloud map. It runs only once during the entire usage period. The surveillance camera captures a photo of the checkerboard calibration board, and combines the calibration data provided during three-dimensional reconstruction to calculate the position and pose of the surveillance camera in the three-dimensional point cloud map for use by the position calculation module. The position calculation module is used to identify the position of the person in the image based on the position and pose data obtained by the calibration module and the surveillance video stream provided by the surveillance camera, and give the position of the person in the three-dimensional point cloud map for display by the visualization display module. The visualization display module is used to display the three-dimensional point cloud map of the scene and the position of the person in the three-dimensional point cloud map, and provide a GUI interface for user interaction and supervision.
Citation Information
Patent Citations
Monocular camera target detection and spatial positioning method based on three-dimensional virtual geographic scene
CN114332385A