Large-scale outdoor monitoring camera intelligent calibration method and calibration device based on single camera
By moving the calibration device with a single camera in the space where the monitoring camera exists, combining the timestamp and motion structure algorithms, identifying points of the same name and applying the principle of multi-view consistency, the problem of lack of position reference information in the monitoring camera system is solved, and the rapid and efficient calibration of the monitoring camera is achieved, and the efficiency of utilizing monitoring resources is improved.
Patent Information
- Application Number
- CN202510316314.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-13
AI Technical Summary
The existing surveillance camera systems lack location reference information and cannot achieve accurate position mapping between devices and monitoring networks, resulting in waste of monitoring resources and inefficiency.
The calibration device with a single camera is used to move in the outdoor monitoring camera space. By recording the position information of the single camera and the captured image data, matching it with the timestamp, two-dimensional image coordinates of the same name point are identified, and the spatial position and attitude angle of the monitoring camera are estimated using the motion structure algorithm and the principle of multi-view consistency.
It realizes fast, efficient and low-cost spatial position and attitude angle calibration of surveillance cameras, improves the efficiency of surveillance resources utilization, and supports generalized applications such as crowdsourcing.
Smart Images

Figure CN120147437A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of camera system calibration, and particularly to an intelligent calibration method and a calibration device for a large-scale outdoor surveillance camera based on a single camera. Background Art
[0002] With the development of society, with the wide deployment of surveillance cameras in public areas and important places, the number of surveillance cameras has increased sharply, generating a huge amount of surveillance video data. However, for the entire surveillance system, the images on the monitor are often messy and disorderly. Usually, monitoring personnel need to know the position of the camera corresponding to each video segment to conduct accurate and efficient monitoring. However, most current surveillance cameras lack position reference information, are unable to establish an accurate position mapping between the device and the surveillance network, and are also unable to optimize the spatial layout of the surveillance cameras, which results in a large amount of waste of surveillance resources. Summary of the Invention
[0003] Object of the Invention: The present invention aims to provide a calibration method for calculating the spatial position and attitude angle of a camera for a large-scale outdoor surveillance camera based on a calibration device with a single camera. Another object of the present invention is to provide an intelligent calibration device for a large-scale outdoor surveillance camera based on a single camera.
[0004] Technical Solution: The intelligent calibration method for a large-scale outdoor surveillance camera based on a single camera according to the present invention includes the following steps:
[0005] (1) The calibration device with a single camera moves in the space where each outdoor surveillance camera is located. The single camera and the outdoor surveillance camera photograph each other. During the movement, the calibration device records the pose information of the single camera, and matches the video data captured by the single camera with the pose information of the single camera through timestamps;
[0006] (2) According to the time when the single camera captures images, filter the video data captured by the outdoor surveillance camera after the calibration device enters the shooting range of the outdoor surveillance camera;
[0007] (3) Identify the two-dimensional image coordinates of the homologous points in the images captured by the single camera and the outdoor surveillance camera;
[0008] (4) According to the pose information of the single camera, obtain a rough solution of the pose of the outdoor surveillance camera through the structure from motion algorithm;
[0009] (5) Eliminate outliers from the rough solution of the pose of the outdoor surveillance camera in step (4), and then calculate the initial value of the pose of the outdoor surveillance camera after weighted averaging;
[0010] (6) Based on the principle of multi-view consistency, optimize the initial value of the pose of the outdoor surveillance camera in step (5) to estimate the spatial position and azimuth angle of each outdoor surveillance camera.
[0011] Furthermore, step (1) is specifically as follows:
[0012] (11) The calibration device with a single camera moves in the space where each outdoor surveillance camera is located. During the movement, the single camera takes a picture of the outdoor surveillance camera, and the outdoor surveillance camera also takes a picture of the calibration device with a single camera.
[0013] (12) Based on Kalman filtering, the joint calculation of the inertial navigation and GNSS in the posture module of the calibration device is completed to obtain the posture data of the calibration device during movement;
[0014] (13) converting the position and posture data of the calibration device during the movement obtained in step (12) in the Earth-centered Earth-fixed coordinate system into the position and posture data of the calibration device during the movement in the local coordinate system;
[0015] (14) According to the time corresponding to the timestamp, combined with the relative posture between the main antenna of the posture module of the calibration device and the single camera, the posture information of the single camera when taking the image at the corresponding time is calculated, and the image data taken by the single camera is matched with the posture information of the single camera through the timestamp.
[0016] Furthermore, step (3) is as follows:
[0017] (31) The images taken by the calibration device with a single camera are grouped into several groups according to geographical locations, and a fast feature extraction and matching algorithm is used to calculate the rough similarity between each group of calibration device images and the outdoor surveillance camera images;
[0018] (32) Select the calibration device image and the outdoor surveillance camera image corresponding to the highest rough similarity, input them into the SuperPoint deep learning network, and obtain image feature points and description factors;
[0019] (33) The image feature points and description factors are input into the LightGlue deep learning network to obtain the two-dimensional image coordinates and matching confidence of the same-name points, and the same-name points whose matching confidence is less than the set threshold are removed.
[0020] Furthermore, step (5) is specifically as follows:
[0021] (51) Calculate the position centroid of all position solutions of each outdoor surveillance camera and remove outliers, calculate the weighted average of the position centroids of the remaining position solutions, and obtain the initial value of the spatial position of the outdoor surveillance camera;
[0022] (52) Calculate the mean of all attitude angle solutions of each outdoor surveillance camera and remove outliers, calculate the weighted average of the remaining attitude angle solutions, and obtain the initial value of the surveillance camera attitude angle.
[0023] Further, step (6) is specifically as follows:
[0024] (61) Uniformly extract a number of corresponding points from the set of corresponding points obtained in step (3);
[0025] (62) Calculate the reprojection error of the extracted corresponding points based on the initial pose value of the outdoor surveillance camera;
[0026] (63) Calculate the view point weight of each calibration device image for optimizing the pose of the outdoor surveillance camera;
[0027] (64) Optimize the pose of the outdoor surveillance camera based on the multi-viewpoint consistency principle, and estimate the spatial position and attitude angle of the outdoor surveillance camera.
[0028] The intelligent calibration method for large-scale outdoor surveillance cameras based on a single camera according to the present invention includes:
[0029] A pose module for obtaining the pose data of the calibration device when moving along the space where each outdoor surveillance device is located;
[0030] A camera module for obtaining the camera image data of the calibration device when moving along the space where each outdoor surveillance device is located;
[0031] The device main body is fixedly connected to the pose module and the camera module, and serves as a carrier, enabling the device to be installed on vehicles such as cars, so as to freely move in outdoor spaces such as streets, squares, and parks;
[0032] A data acquisition module for obtaining the surveillance image data captured by each surveillance device when the marker moves along the space where each surveillance device is located;
[0033] A corresponding point recognition module for identifying the two-dimensional image coordinates of the corresponding points between the calibration device image and the outdoor surveillance camera image;
[0034] A pose estimation module for combining the two-dimensional image coordinates of the corresponding points between the surveillance image and the calibration device image, and estimating the rough spatial position and azimuth angle of each surveillance camera based on the structure from motion algorithm;
[0035] A pose optimization module for improving the accuracy of the results of the spatial position and azimuth angle of the outdoor surveillance camera through initial value screening and multi-viewpoint consistency optimization.
[0036] Beneficial effects: Compared with the prior art, the remarkable advantages of the present invention are as follows: The present invention obtains images with pose information through the spatial movement in the space where the surveillance camera exists, and combining it with the surveillance image can realize the estimation of the spatial position and attitude orientation of the surveillance camera. The present invention not only makes the calibration of outdoor surveillance cameras faster, more efficient and low-cost, but also the flexibility of the device provides support for the development of generalization applications such as crowdsourcing. Description of the Drawings
[0037] Figure 1 is the flow chart of the calibration method of the present invention;
[0038] Figure 2 is the schematic diagram of multi-viewpoint consistency;
[0039] Figure 3 is the structural schematic diagram of the homonymous point recognition network;
[0040] Figure 4 is the distribution map of the monitoring cameras in the experimental area;
[0041] Figure 5 is the schematic diagram of the comparison between the calibration result of the present invention and the true position. Specific implementation manners
[0042] The present invention will be further described below with reference to the accompanying drawings.
[0043] The intelligent calibration method for large-scale outdoor monitoring cameras based on a single camera according to the invention includes the following steps:
[0044] (1) The calibration device with a single camera moves in the space where each outdoor monitoring camera is located. The single camera and the outdoor monitoring cameras take pictures of each other. During the movement, the calibration device records the pose information of the single camera, and matches the image data captured by the single camera with the pose information of the single camera through timestamps.
[0045] (11) Based on Kalman filtering, perform joint solution of the inertial navigation and GNSS in the pose module of the calibration device, and output the positioning data during the movement of the device.
[0046] Among them, a positioning sensor and an inertial navigation sensor are provided on the calibration device. During the movement, real-time positioning and pose determination will be performed according to the acquisition frequency, that is, a collection is performed every once in a while, and a pose data is obtained. Through recursive optimization of Kalman filtering, fuse the absolute position information provided by the positioning sensor and the high-frequency relative motion information provided by the inertial navigation sensor, and output the optimal pose information of the system in the dynamic system.
[0047] (12) Analyze the pose data and convert it into the local coordinate system to obtain the pose of the calibration device at each acquisition moment during the movement in the space where each monitoring device is located.
[0048] Among them, analyze the original positioning message data of the positioning sensor, obtain the longitude, latitude and elevation data of the phase center of the sensor antenna at any moment, and then calculate the position coordinates in the local coordinate system. The calculation process is as follows:
[0049] First, convert the WGS84 coordinate system of the positioning data into the ECEF (Earth-Centered, Earth-Fixed coordinate system), and the expression is as follows:
[0050] x = (N + H) × cos(B) × cos(L)
[0051] y = (N + H) × cos(B) × sin(L)
[0052] z = (N × (1 - e×e) + H) × sin(L)
[0053] Where N is the radius of curvature of the reference ellipsoid, B, L, and H represent the latitude, longitude, and elevation of the antenna phase center respectively, and x, y, and z represent the coordinates in the transformed Earth-centered Earth-fixed coordinate system.
[0054] The transformation from the Earth-centered Earth-fixed coordinate system to the local coordinate system is expressed as follows:
[0055]
[0056] Where B0 and l0 are the longitude and latitude coordinates of the origin of the local coordinate system, (Δ x , Δ y , Δ z ) is the difference between the coordinates of the main antenna of the current calibration device pose module in the Earth-centered Earth-fixed coordinate system and the origin of the local coordinate system, and (x, y, z) is the final position coordinate.
[0057] (13) Obtain the camera images collected during the movement of the calibration device camera module along the space where each monitoring device is located and the corresponding image generation times.
[0058] When the image acquisition command is executed, read the timestamp information in the calibration device pose module.
[0059] (14) According to the image generation time of the device camera, obtain the pose information of the calibration device pose module at the corresponding moment. Based on this pose information, combined with the relative pose between the main antenna of the calibration device pose module and the device camera, solve the pose information of each camera image.
[0060] Let O I -X I Y I Z I represent the combined navigation system coordinate system, O-XYZ is the calibration device camera coordinate system, then the equation for converting a point (X I , Y I , Z I ) in the O I -X I Y I Z I coordinate system to (X, Y, Z) in the calibration device camera coordinate system is as follows:
[0061]
[0062] Among them, R I represents the rotation matrix from the integrated navigation coordinate system to the device camera coordinate system, and t I represents the translation vector from the integrated navigation coordinate system to the calibrated device camera coordinate system.
[0063] (2) According to the time when the single camera captures images, filter the image data captured by the outdoor surveillance camera after the calibration device enters the shooting range of the outdoor surveillance camera. Starting from the time when the calibration device begins to collect data, intercept the surveillance images of each surveillance camera at adjacent moments through the surveillance system.
[0064] (3) Identify the two-dimensional image coordinates of the homologous points in the images captured by the single camera and the outdoor surveillance camera.
[0065] (31) Group the images captured by the calibration device by geographical location into several groups, and use the fast feature extraction and matching algorithm to calculate the rough similarity score between each group of calibration device images and the surveillance camera images.
[0066] Read the geographical location information of the calibration device images, judge whether they are divided into the same group by calculating the distance between adjacent images, and set the maximum distance threshold to 200m. Use ORB to perform fast matching between each group of calibration device images and the surveillance camera images, and use the group with the maximum number of matching points calculated to be greater than 150 as the group matching the current surveillance image to participate in the subsequent precise matching.
[0067] (32) Input the calibration device image group with the highest rough similarity score and the surveillance camera image into the SuperPoint deep learning network to obtain the image feature points and descriptors.
[0068] The output of each group of images includes: the two-dimensional coordinates (x, y) of the feature points, and the descriptors (d) of the feature points.
[0069] (33) Input the image feature points and descriptors into the LightGlue deep learning network to obtain the two-dimensional image coordinates of the homologous points and the matching confidence.
[0070] Input the feature point coordinates (x, y) and descriptors (d) obtained in (32) into the LightGlue network, and the network structure diagram is as Figure 2 shown. Each group of inputs will be a pair of feature point sets, which are the feature point sets between the calibration device camera image and the surveillance camera image respectively.
[0071] (34) Retain the matching results of the two-dimensional coordinates of the homologous points with a confidence higher than 0.5.
[0072] (4) According to the pose information of the single camera, obtain a rough solution of the pose of the outdoor surveillance camera through the structure from motion algorithm.
[0073] The fundamental matrix F between the monitoring image and the dashcam image is estimated from multiple corresponding point pairs using the least squares method by the following formula. This formula can be abbreviated as Af = 0, and the non - zero solution F is obtained by performing singular value decomposition on matrix A.
[0074]
[0075] The internal parameter matrices of the two images are simplified to the identity matrix, and the approximate essential matrix E between the monitoring image and the dashcam image is obtained through E = K' T FK.
[0076] Perform singular value decomposition on the essential matrix E to get E = U∑V T , and solve for the two solutions of the relative rotation matrix R r R r1 = UWV T , R r2 = UW T V T , and the two solutions of the relative translation matrix t r are T r1 = U[:,2], t r2 = -U[:,2]. Determine the unique solutions of R r , t r through reprojection error minimization.
[0077] Finally, combine the pose information R, t of the calibration device camera image and the relative pose R r , t r , and calculate the rough solution of the pose of the monitoring camera R mc , t mc through the following formula.
[0078] [R mc t mc = [R t]·[R r t r -1 .
[0079] (5) Reject the outliers from the rough solution of the pose of the outdoor monitoring camera in step (4), and then calculate the initial value of the pose of the outdoor monitoring camera after weighted average.
[0080] (51) Calculate the centroid of all position solutions of each monitoring camera and reject the outliers, and calculate the weighted average of the remaining solutions to obtain the initial value of the spatial position of the monitoring camera.
[0081] Calculate the centroid of multiple sets of position solutions of the monitoring camera through , where N is the number of solutions, and p i is the three-dimensional coordinates of the i-th solution. Then calculate the distance d from each solution to the centroid i = ||p i - C||. Use the standard deviation to define the threshold for rejecting outliers. If the distance from a certain solution to the centroid exceeds μ + kσ, it can be regarded as an outlier and rejected. For the remaining solutions, use the reciprocal of the distance d i from the solution to the centroid as the weight, and thus calculate the weighted average as the initial value of the final position w i is the weight of the i-th solution, and p i is the three-dimensional position of the i-th solution.
[0082] (52) Calculate the mean of all attitude angle solutions of each monitoring camera and reject outliers. Calculate the weighted average of the remaining solutions to obtain the initial value of the attitude angle of the monitoring camera.
[0083] Through θ ij = 2×arccos(|q i ·q j |), calculate the mean of the rotation angle θ between each attitude solution in quaternion form and all other solutions. When this value exceeds μ + kσ, it can be regarded as an outlier and rejected. Use the recognition quality score output by the homonymous point recognition network as the weight, and through ij calculate the weighted covariance matrix of each quaternion. Then perform eigenvalue decomposition on the weighted covariance matrix A to find the eigenvector q corresponding to the largest eigenvalue. avg That is the result of the weighted average, and normalize it to obtain
[0084] (6) Based on the principle of multi-view consistency, optimize the initial pose values of the outdoor monitoring cameras in step (5) to estimate the spatial positions and azimuth angles of each outdoor monitoring camera.
[0085] (61) Uniformly extract several homonymous points from the homonymous point recognition results between each group of calibration device images and monitoring camera images.
[0086] Divide the calibration device camera images into a 10×10 grid, and randomly select a homonymous point corresponding to the monitoring image from each grid cell to participate in the subsequent reprojection calculation.
[0087] (62) Calculate the reprojection error of the above homonymous points obtained based on the initial pose values of the monitoring cameras.
[0088] According to the initial pose value [R s |t s of the monitoring camera, calculate the monitoring camera and the calibration device image sequences {D 1 , D 2 , D 3 , …, DN The relative external parameter [R' between i |t' i , and its corresponding essential matrix E i = [t' i × R' i . After restoring to the fundamental matrix F, the coordinates {P' 1 , P' 2 , P 3 , …, P' M} of the M feature points in the monitoring image reprojected onto the i-th calibration device image are obtained. The sum of the coordinate differences between these and the true coordinates {P 1 , P 2 , P 3 , …, P M} is denoted as δ i . Then, the reprojection error under all viewpoints
[0089] (63) Calculate the viewpoint weights for optimizing the pose of the monitoring camera for each calibration device image.
[0090] The weights mainly consider two factors: distance weight and matching quality weight of homologous points.
[0091] Distance weight is used to reduce the influence of viewpoints at a long distance on the reprojection optimization result. Closer viewpoints can usually provide clearer details and higher matching accuracy. Therefore, images at a closer distance are given higher weights. Assume that the initial distance between the position of the monitoring camera and the position of the camera of the i-th street view image is d i . Then
[0092] Matching quality weight reflects the overall quality of the matching of all feature points in the i-th street view image, to increase the influence of images with a high matching degree on the final optimization result. Assume that the average matching quality score of the M feature points in the i-th image is . Then
[0093] The final viewpoint weight w i is the product of the distance weight and the matching quality weight . As shown in the following formula:
[0094]
[0095] (64) Optimize the pose of the monitoring camera based on the multi-viewpoint consistency principle. The schematic diagram is as Figure 3 shown, and estimate the spatial position and attitude angle of the monitoring camera.
[0096] Define the optimization objective function as the weighted sum of reprojection errors, in the form as follows:
[0097]
[0098] where δ i is the reprojection error of the i-th street view image, and w i is the corresponding viewpoint weight. Take f(R s | t s ) as the objective optimization function of the Levenberg-Marquardt algorithm, and δ is the error vector of the initial parameters [R s | t s . By continuously iterating and updating [R|t], the minimization of δ is achieved, and finally the optimized pose parameters of the monitoring camera are obtained.
[0099] Finally, the simulation verification of the present invention is carried out. Select the Xianlin Campus of Nanjing Normal University as the experimental area. As Figure 4 shown, the area covers an area of about 251,021 m 2 , and there are basic geographical scene elements such as roads, shopping malls, playgrounds, teaching buildings, power distribution rooms, squares, bus stops, etc. in the area, which conform to the common characteristics of the outdoor monitoring camera layout scene. There are 400 video monitoring cameras in the test area, which are scattered at the intersections of internal roads and the external facades of buildings in the area. The implementation is carried out using the Python language in the Windows 10 64-bit system environment. Without knowing the camera position, the positions of the monitoring cameras in the north area are measured using both the manual measurement method and the method proposed in this paper. The results are as Figure 5 shown. It can be seen that the present invention accurately calculates the spatial position and attitude angle information of the monitoring camera.
Claims
1. A large-scale outdoor surveillance camera intelligent calibration method based on a single camera, characterized in that: The following steps are involved: (1) A calibration device with a single camera moves in the space where each outdoor surveillance camera is located. The single camera and the outdoor surveillance camera shoot each other. During the movement, the calibration device records the position information of the single camera and matches the image data shot by the single camera with the position information of the single camera through the timestamp; (2) According to the time when a single camera captured an image, the image data captured by the outdoor surveillance camera after the calibration device entered the shooting range of the outdoor surveillance camera was filtered; (3) Identify the two-dimensional image coordinates of the same-name points in the images taken by a single camera and the images taken by an outdoor surveillance camera; (4) Based on the pose information of a single camera, a rough solution of the pose of the outdoor surveillance camera is obtained through the motion structure algorithm; (5) removing outliers from the rough solution of the outdoor surveillance camera's posture in step (4), and then calculating the initial value of the outdoor surveillance camera's posture after weighted averaging; (6) Based on the principle of multi-view consistency, the initial value of the outdoor surveillance camera posture in step (5) is optimized to estimate the spatial position and azimuth of each outdoor surveillance camera.
2. According to claim 1, the large-scale outdoor surveillance camera intelligent calibration method based on a single camera is characterized in that: Step (1) is as follows: (11) The calibration device with a single camera moves in the space where each outdoor surveillance camera is located. During the movement, the single camera takes a picture of the outdoor surveillance camera, and the outdoor surveillance camera also takes a picture of the calibration device with a single camera. (12) Based on Kalman filtering, the joint calculation of the inertial navigation and GNSS in the posture module of the calibration device is completed to obtain the posture data of the calibration device during movement; (13) converting the position and posture data of the calibration device during the movement obtained in step (12) in the Earth-centered Earth-fixed coordinate system into the position and posture data of the calibration device during the movement in the local coordinate system; (14) According to the time corresponding to the timestamp, combined with the relative posture between the main antenna of the posture module of the calibration device and the single camera, the posture information of the single camera when taking the image at the corresponding time is calculated, and the image data taken by the single camera is matched with the posture information of the single camera through the timestamp.
3. The large-scale outdoor surveillance camera intelligent calibration method based on a single camera according to claim 2 is characterized in that: Step (3) is as follows: (31) The images taken by the calibration device with a single camera are grouped into several groups according to geographical locations, and a fast feature extraction and matching algorithm is used to calculate the rough similarity between each group of calibration device images and the outdoor surveillance camera images; (32) Select the calibration device image and the outdoor surveillance camera image corresponding to the highest rough similarity, input them into the SuperPoint deep learning network, and obtain image feature points and description factors; (33) The image feature points and description factors are input into the LightGlue deep learning network to obtain the two-dimensional image coordinates and matching confidence of the same-name points, and the same-name points whose matching confidence is less than the set threshold are removed.
4. The large-scale outdoor surveillance camera intelligent calibration method based on a single camera according to claim 3 is characterized in that: Step (5) is as follows: (51) Calculate the position centroid of all position solutions of each outdoor surveillance camera and remove outliers, calculate the weighted average of the position centroids of the remaining position solutions, and obtain the initial value of the spatial position of the outdoor surveillance camera; (52) Calculate the mean of all attitude angle solutions of each outdoor surveillance camera and remove outliers. Calculate the weighted average of the remaining attitude angle solutions to obtain the initial value of the surveillance camera attitude angle.
5. The large-scale outdoor surveillance camera intelligent calibration method based on a single camera according to claim 4 is characterized in that: Step (6) is as follows: (61) uniformly extracting a number of points with the same name from the set of points with the same name obtained in step (3); (62) Reprojection error of the same-name points extracted based on the initial value calculation of the outdoor surveillance camera’s pose; (63) Calculating the viewpoint weight of each calibration device image for optimizing the outdoor surveillance camera posture; (64) Based on the multi-viewpoint consistency principle, the posture of the outdoor surveillance camera is optimized and the spatial position and attitude angle of the outdoor surveillance camera are estimated.
6. A large-scale outdoor surveillance camera intelligent calibration method based on a single camera, characterized in that: include: A posture module is used to obtain the posture data of the calibration device when it moves along the space where each outdoor monitoring device is located; A camera module, used to obtain camera image data of the calibration device when it moves along the space where each outdoor monitoring device is located; The device body is firmly connected to the posture module and the camera module, and plays a bearing role, so that the device can be installed on vehicles such as cars to facilitate free movement in outdoor spaces such as streets, squares, and parks; A data acquisition module, used to acquire monitoring image data captured by each monitoring device when the marker moves along the space where each monitoring device is located; A homonymous point recognition module is used to recognize the two-dimensional image coordinates of homonymous points between the calibration device image and the outdoor monitoring camera image; The pose estimation module is used to combine the two-dimensional image coordinates of the same-name points between the monitoring image and the calibration device image, and estimate the rough spatial position and azimuth of each monitoring camera based on the motion structure recovery algorithm; The pose optimization module is used to improve the accuracy of the spatial position and azimuth results of outdoor surveillance cameras through initial value screening and multi-viewpoint consistency optimization.