Pedestrian Cross-Camera Localization Method, System, and Medium Based on Top-View Camera

By detecting the head and returning to the foot position in the top-view camera system, mapping and splicing of corrected maps and local maps, the problems of low accuracy and slow recognition speed of existing cross-mirror positioning technology are solved, and fast and accurate pedestrian positioning and tracking are achieved.

CN115050004BActive Publication Date: 2025-06-27JIANGSU FANTE TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210667434.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-13
Publication Date
2025-06-27
Estimated Expiration
2042-06-13

AI Technical Summary

Technical Problem

The accuracy of existing cross-mirror positioning technology is low and the recognition speed is slow. Especially when faced with problems such as light changes, pedestrian occlusion and similar clothing, the accuracy of the Top1 has dropped significantly and cannot reach the commercial level.

Method used

The pedestrian cross-mirror positioning method based on top-view camera is adopted. By acquiring the original image of each camera, detecting the head position and regressing the foot position, correcting map mapping and local map splicing are performed to achieve fast and accurate positioning of pedestrians.

Benefits of technology

Fast and accurate cross-mirror positioning and pedestrian tracking are achieved, avoiding misjudgment caused by factors such as light changes, pedestrian occlusion and similar clothing, improving the recognition accuracy and speed, and reaching the commercial level.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115050004B_ABST
    Figure CN115050004B_ABST
Patent Text Reader

Abstract

The present invention relates to a pedestrian cross-camera positioning method, system, and medium based on a top-view camera. Among them, the method includes: obtaining the original images of each camera within a target area; detecting the position of a person's head in the original images, and regressing the position of the feet based on the position of the head; mapping the position of the feet in each of the original images to the corresponding position on a correction map preset for the corresponding camera; mapping the corresponding position of the feet on the correction map to the corresponding position on a local map pre-constructed based on the target area, thereby realizing the positioning of pedestrians on the local map. Among them, the local map is obtained by stitching the correction maps of each camera. The present invention realizes a fast and accurate cross-camera positioning and subsequent pedestrian tracking through the preset correction map and local map, and its accuracy and recognition rate are better than existing solutions, having certain practical popularization significance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a pedestrian cross-camera positioning method, system and medium based on a top-view camera. Background Art

[0002] With the development of pattern recognition technology and video analysis and processing technology, people's demand for the safety of daily activity places is increasing. Intelligent video surveillance systems have been widely used in the security field to provide protection for people's property and life safety. Pedestrian trajectory tracking based on video sequences is an important part of intelligent video surveillance systems and is applied to important indoor places such as shopping malls, parking lots, banks, exhibitions, railway stations, etc.

[0003] Since 2017, pedestrian tracking technology under a single camera has achieved good development in the academic community. The MOT (Multiple Object Tracking Benchmark) benchmark is refreshed by new algorithms every year, but the accuracy is far from reaching the implementation standard (such as 90% - 95%). Taking the MOT20 data as an example, the accuracy of the current best algorithm is only 77.1%.

[0004] For the research on cross-camera tracking, the commonly used method in the industry is to integrate ReID (pedestrian re-identification) technology on the basis of single-camera tracking to complete the matching of the same target between different cameras. Although the ReID technology has achieved a Top1 accuracy rate of over 90% on public data (such as Market1501), in the actual security camera images, in the face of problems such as "light change", "pedestrian occlusion", "similar clothing", etc., the Top1 accuracy rate of ReID drops significantly. Coupled with the error accumulation caused by frequent pedestrian cross-camera, the accuracy rate of the entire cross-camera tracking system is lower than 60% and cannot reach the commercial level. Summary of the Invention

[0005] (1) Technical Problems to be Solved

[0006] In view of the above-mentioned shortcomings and deficiencies of the prior art, the present invention provides a pedestrian cross-camera positioning method, system and medium based on a top-view camera, which solves the technical problems of low accuracy and slow recognition speed of the existing cross-camera positioning technology.

[0007] (2) Technical Solutions

[0008] To achieve the above object, the main technical solutions adopted by the present invention include:

[0009] In a first aspect, an embodiment of the present invention provides a pedestrian cross-camera positioning method based on a top-view camera, including:

[0010] Obtain the original image of each camera within the target area;

[0011] Detect the position of the human head in the original image, and regress the position of the feet based on the position of the head;

[0012] Map the position of the feet in each original image to the corresponding position on the correction map preset for the corresponding camera;

[0013] Map the corresponding position of the feet on the correction map to the corresponding position on the local map pre-constructed based on the target area, so as to realize the positioning of pedestrians on the local map;

[0014] Wherein, the local map is obtained by stitching the correction maps of each camera.

[0015] Optionally, before obtaining the original images of each camera in the target area, it further includes:

[0016] After calibrating a number of cameras set in a preset target area, obtain the internal parameter matrix and external parameter matrix of the cameras;

[0017] Calculate the remapping matrix for correction based on the internal parameter matrix and external parameter matrix of the cameras;

[0018] Perform distortion correction on the images captured by each camera based on the remapping matrix to obtain the correction map of each camera;

[0019] Obtain the stitching parameters by performing feature point matching on the correction maps of every two adjacent cameras;

[0020] Stitch the images captured by each camera according to the stitching parameters to obtain the local map.

[0021] Optionally, after calibrating a number of cameras set in a preset target area, obtaining the internal parameter matrix and external parameter matrix of the cameras includes:

[0022] Obtain the images captured for the checkerboard grid pre-set in the target area;

[0023] Detect the corner points of the checkerboard grid fixed points in the captured images to obtain a number of corner point coordinates in two-dimensional image coordinates;

[0024] Calculate the internal parameter matrix and external parameter matrix of each camera based on the corner point coordinates of the captured images;

[0025] And, stitching the images captured by each camera according to the stitching parameters to obtain the local map includes:

[0026] According to the perspective transformation matrix in the stitching parameters between every two adjacent cameras and by combining matrix multiplication, the perspective transformation matrix from each camera to the specified camera is obtained;

[0027] Based on the perspective transformation matrix from each camera to the specified camera, perform perspective transformation on the rectified image of each camera, and then perform stitching to obtain a local map;

[0028] Among them, the internal parameter matrix K and the external parameter matrix R of each camera satisfy the following formula:

[0029]

[0030] In the formula, C0......C 53 are the default coordinates of each corner point in the three-dimensional camera coordinate system respectively, and P0......P 53 are the corner point coordinates of each corner point in the two-dimensional image coordinates;

[0031] The remapping matrix is:

[0032]

[0033] In the formula, u and v are the transformed pixel coordinates, x and y are the pixel coordinates corresponding to u and v respectively, K is the internal parameter matrix, and R is the external parameter matrix;

[0034] The stitching parameters include the following perspective transformation matrix:

[0035]

[0036] Among them, a i and b i are N groups of matching points of adjacent cameras A and B respectively, and H A→B is the perspective transformation matrix, i = 0, 1, 2,... N - 1.

[0037] Optionally, detecting the position of a person's head in the original image and regressing the position of the feet based on the position of the head includes:

[0038] Detecting the center point coordinates of a person's head in the original image through a pre-trained Yolov5 model;

[0039] Calculating the corresponding center point coordinates of the feet according to the center point coordinates of the head in the original image in combination with the regression equation;

[0040] Among them,

[0041] The pre-trained Yolov5 model is quantized and compressed and deployed on a specified hardware platform;

[0042] The regression equation is as follows:

[0043]

[0044] where foot i is the coordinate of the center point of the i-th foot, and head i is the coordinate of the center point of the head; i = 0, 1, 2,... N - 1, M is the fitting degree, defaulting to 5; W and H are the width and height of the fisheye image respectively; a k (k = 0, 1, 2,... M) are the optimal parameters obtained by solving N regression equations simultaneously through the least squares method.

[0045] Optionally, mapping the position of the foot in each of the original images to the corresponding position on the correction map preset by the corresponding camera includes:

[0046] Based on the remapping matrix, by obtaining the coordinates (u, v) in the correction map corresponding to each set of coordinates (x, y) of each of the original images, w*h two-dimensional arrays are obtained;

[0047] For any set of coordinates (x0, y0) of each of the original images, the point closest to (x0, y0) among the w*h two-dimensional arrays is found through the nearest neighbor algorithm as the point in the correction map corresponding to the set of coordinates (x0, y0);

[0048] where w and h are the width and height of the correction map respectively.

[0049] Optionally, mapping the corresponding position of the foot on the correction map to the corresponding position on the local map pre-constructed based on the target area, and then realizing the positioning of the pedestrian on the local map includes:

[0050] Based on the perspective transformation matrix from each camera to the specified camera, the corresponding position of the foot on the correction map is mapped to the corresponding position on the local map, and the pedestrian is positioned based on the local map.

[0051] Optionally, after positioning the pedestrian based on the local map, it further includes: through a preset multi-target final algorithm, tracking the trajectory of the pedestrian across cameras according to the corresponding position of the foot on the local map.

[0052] Optionally, the layout of each camera in the preset target area satisfies the following conditions:

[0053] It is preferably set directly above the seating area within the preset target area;

[0054] The distance between any two adjacent cameras does not exceed a predetermined value, and any three adjacent cameras form an equilateral triangle;

[0055] The detection ranges of all cameras together cover the entire target area;

[0056] And each camera is a fish-eye camera, and the image captured by the fish-eye camera is a fish-eye image.

[0057] In a second aspect, an embodiment of the present invention provides a pedestrian cross-camera positioning system based on a top-view camera, including:

[0058] An image acquisition module, configured to acquire the original images of each camera within the target area;

[0059] A head-to-foot mapping module, configured to detect the position of a person's head in the original image, and regress the position of the feet based on the position of the head;

[0060] A corrected map mapping module, configured to map the position of the feet in each original image to the corresponding position on the corrected map preset for the corresponding camera;

[0061] A local map mapping module, configured to map the corresponding position of the feet on the corrected map to the corresponding position on a local map pre-constructed based on the target area, so as to realize the positioning of pedestrians on the local map;

[0062] Wherein, the local map is obtained by stitching the corrected maps of each camera.

[0063] In a third aspect, an embodiment of the present invention provides a computer-readable medium, on which computer-executable instructions are stored, and when the executable instructions are executed by a processor, a pedestrian cross-camera tracking method based on a top-view camera as described above is implemented.

[0064] (III) Advantageous Effects

[0065] The beneficial effects of the present invention are as follows: In the initialization stage, the corrected map of each camera and the entire local map are pre-constructed in the present invention, providing a relatively convenient and fast way for subsequent real-time mapping from the original image to the corrected map, mapping from the corrected map to the local map, and even trajectory tracking. At the same time, it avoids misjudgment in positioning and tracking due to reasons such as light changes, pedestrian occlusion, and similar clothing. Thus, the present invention realizes fast and accurate cross-camera positioning and subsequent precise tracking of pedestrians, thereby obtaining the beneficial effect of the complete real-time trajectory of the target in the entire area. The present invention is superior to the existing solutions both in terms of recognition accuracy and recognition rate, and has certain practical promotion significance. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Figure 1 It is a schematic flowchart of a pedestrian cross-camera positioning method provided by the present invention;

[0067] Figure 2Schematic diagram of the specific process before step S1 of a pedestrian cross-camera positioning method based on an overhead camera provided by the present invention;

[0068] Figure 3 Schematic diagram of the specific process of step F11 of a pedestrian cross-camera positioning method based on an overhead camera provided by the present invention;

[0069] Figure 4 Checkerboard schematic diagram for calibration of a pedestrian cross-camera positioning method based on an overhead camera provided by the present invention;

[0070] Figure 5 Schematic diagram of the corner coordinates of the checkerboard of a pedestrian cross-camera positioning method based on an overhead camera provided by the present invention;

[0071] Figure 6 Schematic diagram of distortion correction of a pedestrian cross-camera positioning method based on an overhead camera provided by the present invention;

[0072] Figure 7 Schematic diagram of the local map of a pedestrian cross-camera positioning method based on an overhead camera provided by the present invention;

[0073] Figure 8 Schematic diagram of the specific process of step S2 of a pedestrian cross-camera positioning method based on an overhead camera provided by the present invention;

[0074] Figure 9 Schematic diagram of head detection of a pedestrian cross-camera positioning method based on an overhead camera provided by the present invention;

[0075] Figure 10 Schematic diagram of head-to-foot mapping of a pedestrian cross-camera positioning method based on an overhead camera provided by the present invention;

[0076] Figure 11 Schematic diagram from the original image to the corrected image of a pedestrian cross-camera positioning method based on an overhead camera provided by the present invention;

[0077] Figure 12 Mapping from the corrected image to the local map of a pedestrian cross-camera positioning method based on an overhead camera provided by the present invention;

[0078] Figure 13 Schematic diagram of target tracking based on the local map of a pedestrian cross-camera positioning method based on an overhead camera provided by the present invention;

[0079] Figure 14-1 、 Figure 14-2 and Figure 14-3 respectively show the first, second, and third layout modes of the camera of a pedestrian cross-camera positioning method based on an overhead camera provided by the present invention. Detailed implementation manners

[0080] To better explain the present invention for easy understanding, the present invention will be described in detail below in conjunction with the accompanying drawings through specific embodiments.

[0081] As Figure 1 shown, a pedestrian cross-camera positioning method based on a top-view camera proposed by an embodiment of the present invention includes: First, obtain the original image of each camera in the target area; Second, detect the position of the human head in the original image, and regress the position of the feet based on the position of the head; Then, map the position of the feet in each original image to the corresponding position of the correction map preset for the corresponding camera; Finally, map the corresponding position of the feet on the correction map to the corresponding position of the local map pre-constructed based on the target area, so as to realize the positioning of the pedestrian on the local map; wherein, the local map is obtained by stitching the correction maps of each camera.

[0082] In the initialization stage of the present invention, the correction map of each camera and the entire local map are pre-constructed in advance, providing a relatively convenient and fast way for subsequent real-time mapping from the original image to the correction map, mapping from the correction map to the local map, and even trajectory tracking. At the same time, it avoids misjudgment of positioning and tracking due to reasons such as light changes, pedestrian occlusion, and similar clothing. Thus, the present invention realizes the beneficial effects of fast and accurate cross-camera positioning and subsequent precise tracking of pedestrians, so as to obtain the complete real-time trajectory of the target in the entire area. The present invention is superior to the existing solutions both in terms of recognition accuracy and recognition rate, and has certain practical popularization significance.

[0083] To better understand the above technical solution, the exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more clear and thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.

[0084] Specifically, a pedestrian cross-camera positioning method based on a top-view camera provided by the present invention includes:

[0085] S1. Obtain the original image of each camera in the target area. Real-time collect the images of each fish-eye camera. To meet the requirements of pedestrian tracking, it is recommended that the sampling frequency is not less than 5 frames per second.

[0086] As Figure 2 shown, before step S1, it further includes:

[0087] F11. After calibrating a plurality of cameras set in the preset target area, obtain the internal parameter matrix and external parameter matrix of the cameras.

[0088] Further, as Figure 3 shown, step F11 includes:

[0089] F111. Obtain an image taken of a chessboard grid pre-set in a target area.

[0090] F112. Detect corner points of the fixed points of the chessboard grid in the taken image to obtain several corner point coordinates in two-dimensional image coordinates.

[0091] F113. Calculate the internal parameter matrix and external parameter matrix of each camera based on the corner point coordinates of the taken image.

[0092] The above camera calibration adopts Zhang Zhengyou calibration method, and its specific implementation steps are as follows:

[0093] (1) Fix a chessboard on the ground.

[0094] (2) Hold a camera (the same model as the installed camera) and take pictures of the chessboard from different angles and heights around the chessboard, ensuring that the chessboard occupies more than 1 / 8 of the entire picture, and accumulate about 50 pictures.

[0095] (3) Detect corner points of the vertices of the chessboard grid in the Figure 4 shown picture.

[0096] Taking the Figure 5 shown chessboard as an example (the corner points are marked with solid circles), this chessboard uses a 6*9 chessboard, so there are a total of 54 corner points; their coordinates in the three-dimensional camera coordinate system are by default:

[0097] C0 = (0, 0, 0), C1 = (1, 0, 0), C2 = (2, 0, 0), C3 = (3, 0, 0), C4 = (4, 0, 0), C5 = (5, 0, 0);

[0098] C6 = (0, 1, 0), C7 = (1, 1, 0), C8 = (2, 1, 0), C9 = (3, 1, 0), C10 = (4, 1, 0), C11 = (5, 1, 0);

[0099] C12 = (0, 2, 0), C13 = (1, 2, 0), C14 = (2, 2, 0), C15 = (3, 2, 0), C16 = (4, 2, 0), C17 = (5, 2, 0);

[0100] C18 = (0, 3, 0), C19 = (1, 3, 0), C20 = (2, 3, 0), C21 = (3, 3, 0), C22 = (4, 3, 0), C23 = (5, 3, 0);

[0101] C24 = (0, 4, 0), C25 = (1, 4, 0), C26 = (2, 4, 0), C27 = (3, 4, 0), C28 = (4, 4, 0), C29 = (5, 4, 0);

[0102] C30 = (0, 5, 0), C31 = (1, 5, 0), C32 = (2, 5, 0), C33 = (3, 5, 0), C34 = (4, 5, 0), C35 = (5, 5, 0);

[0103] C36 = (0, 6, 0), C37 = (1, 6, 0), C38 = (2, 6, 0), C39 = (3, 6, 0), C40 = (4, 6, 0), C41 = (5, 6, 0);

[0104] C42 = (0, 7, 0), C43 = (1, 7, 0), C44 = (2, 7, 0), C45 = (3, 7, 0), C46 = (4, 7, 0), C47 = (5, 7, 0);

[0105] C48 = (0, 8, 0), C49 = (1, 8, 0), C50 = (2, 8, 0), C51 = (3, 8, 0), C52 = (4, 8, 0), C53 = (5, 8, 0).

[0106] The two - dimensional image coordinates of them are obtained through corner detection, and are denoted as P0, P1, … P53 respectively.

[0107] (4) According to the corner coordinates of all the figures, calculate the internal camera matrix K and the external camera matrix R, and the internal camera matrix K and the external camera matrix R satisfy the formula:

[0108]

[0109] In the formula, C0......C 53 are the default coordinates of each corner point in the three - dimensional camera coordinate system respectively, and P0......P 53 are the corner coordinates of each corner point presenting two - dimensional image coordinates.

[0110] Based on the above formula, use the RANSAC (Random Sample Consensus) algorithm to solve the internal camera matrix K and the external camera matrix R. The process is as follows:

[0111] (a) Randomly sample N points.

[0112] (b) According to the N sampled points, perform the least - squares method to solve the local optimal solution of the unknown parameters, so as to obtain a temporary model.

[0113] (c) Calculate the Mean Square Error (MSE) of the remaining sampling points (points other than N points) according to the temporary model.

[0114] (d) Mark the points with errors greater than the threshold as outliers and the points with errors less than the threshold as inliers.

[0115] (e) Repeat the above four steps 3 to 5 times.

[0116] (f) Use all the inliers as the final sampling points, perform the least squares method to solve the optimal solution of the unknown parameters, and denote it as the final model.

[0117] F12. Calculate the remapping matrix for correction based on the internal parameter matrix and external parameter matrix of the camera.

[0118] Calculate the remapping matrix (correction parameter) according to the camera internal parameter matrix K and external parameter matrix R, and save it. The remapping matrix is:

[0119]

[0120] In the formula, u and v are the transformed pixel coordinates, x and y are the pixel coordinates corresponding to u and v respectively, K is the internal parameter matrix, and R is the external parameter matrix. p, q, and r are all intermediate calculation results and only participate in the calculation without actual meaning.

[0121] F13. Perform distortion correction on the images captured by each camera based on the remapping matrix to obtain the corrected image of each camera. With the above remapping matrix, refer to Figure 6 It can be seen that only one image remapping (placing the pixels in one image to the specified positions in another image) is required to perform distortion correction on the image.

[0122] F14. Obtain the stitching parameters by performing feature point matching on the corrected images of every two adjacent cameras, and the stitching parameters include the perspective transformation matrix.

[0123] Suppose there are N groups of matching points a0, a1,... a N-1 and b0, b1,... b N-1 in the corrected images of camera A and camera B. They are all two-dimensional coordinate vectors. Then, use the following formula to calculate the perspective transformation matrix between camera A and camera B:

[0124]

[0125] Among them, a i and b i are the N groups of matching points of adjacent cameras A and B respectively, H A→B is the perspective transformation matrix, i = 0, 1, 2,... N - 1, and z is also an intermediate calculation result and only participates in the calculation without actual meaning.

[0126] F15. Stitch the images captured by each camera according to the stitching parameters to obtain a local map.

[0127] A perspective transformation matrix H is obtained between all adjacent cameras A→B (from camera A to camera B). Then, a perspective transformation matrix can also be obtained for non-adjacent cameras through matrix multiplication. For example: the perspective transformation matrix from camera A to camera B is H A→B , and the perspective transformation matrix from camera B to camera C is H B→C . Then, the perspective transformation matrix from camera A to camera C is H A→C = H B→C * H A→B . Then, specify a main camera. Then, through the above method of matrix multiplication, the transformation matrices from all other cameras to the main camera can be obtained. All camera images are perspectively transformed according to their transformation matrices to the main camera, and then stitched to obtain the local map as shown in Figure 7 .

[0128] S2. Detect the position of the human head in the original image, and regress to obtain the position of the feet based on the position of the head.

[0129] As shown in Figure 8 , step S2 includes:

[0130] S21. As shown in Figure 9 , detect the central point coordinates of the human head in the original image through the pre-trained Yolov5 model.

[0131] Based on the deep learning human head detector, detect the position of the human head in the image (i.e., the position within the box in Figure 9 ). Specifically, it can be divided into a data preparation stage, a model training stage, and a model inference stage.

[0132] In the data preparation stage, the present invention will collect real fisheye videos / images of a specified scene and annotate rectangular boxes for the human head regions in the images. Generally, 1000 - 5000 effective images will be collected for each camera.

[0133] In the model training stage, the present invention uses the Yolov5 model and feeds the calibrated data to the model for iterative learning.

[0134] In the model inference stage, the present invention will first quantize and compress the trained model, then deploy the model on the specified hardware platform, and finally predict the real-time video stream in real time on the model deployment end, returning the location information and category information of all valid targets in each frame of the image. INT8 quantization technology is used in the model quantization stage, with extremely low precision loss in exchange for 2 to 3 times of efficiency improvement. In the model deployment stage, the present invention supports the server-side x86 architecture and the edge box-side arm64 architecture, and the computing power side is adapted to the NPU of Huawei Ascend 310, any GPU graphics card of NVIDIA, and pure CPU acceleration.

[0135] S22, according to the coordinates of the center point of the head in the original image combined with the regression equation, the corresponding foot coordinates are calculated, and the specific steps are as follows:

[0136] (1) Select N people in several pictures and mark the coordinates of the center of each person’s head and the center of their feet, denoted as head0, head1, …headN-1 and foot0, foot1, …footN-1, which are all two-dimensional vectors.

[0137] (2) Establish the regression equation:

[0138]

[0139] Among them, foot i is the coordinate of the center point of the i-th foot, head i is the coordinate of the center point of the head; i = 0, 1, 2, ... N-1, M is the fitting degree, the default is 5; W and H are the width and height of the fisheye image respectively; a k (k=0, 1, 2, ...M) is the optimal parameter obtained by combining N regression equations and solving them by the least squares method.

[0140] (3) Combine N equations and solve the optimal parameter a by the least squares method k (k=0,1,2,...M).

[0141] Therefore, for the head center point coordinate head detected by the target, the corresponding foot coordinate foot is calculated as follows:

[0142]

[0143] S3, such as Figure 11 As shown, the foot position in each original image is mapped to the corresponding position of the correction image preset by the corresponding camera.

[0144] Step S3 includes:

[0145] S31. Based on the remapping matrix, by obtaining the coordinates (u, v) in the corrected image corresponding to each set of coordinates (x, y) of each original image, w*h two-dimensional arrays are obtained.

[0146] S32. For any set of coordinates (x0, y0) of each original image, find the point closest to (x0, y0) among the w*h two-dimensional arrays through the nearest neighbor algorithm as the point in the corrected image corresponding to the set of coordinates (x0, y0). Here, w and h are the width and height of the corrected image respectively.

[0147] Based on the remapping matrix, the mapping relationship from the coordinates (u, v) in the corrected image to the coordinates (x, y) in the original fisheye image is obtained. At this moment, what needs to be done is reverse reasoning, that is, according to the coordinates (x, y) in the original fisheye image, solve the corresponding coordinates (u, v) in the corrected image. The specific approach is as follows:

[0148] (1) For each set of (u, v) (u = 0, 1,..., w - 1, v = 0, 1,..., h - 1, where w and h are the width and height of the corrected image respectively), solve the corresponding (x, y), thus obtaining a two-dimensional array of w*h.

[0149] (2) For the given (x0, y0), use the nearest neighbor algorithm KNN (K-Nearest Neighbors) to find the point closest to (x0, y0) among the w*h two-dimensional arrays in (1).

[0150] (3) The (u, v) corresponding to this closest point is the point in the corrected image corresponding to (x0, y0).

[0151] S4. Map the corresponding position of the foot on the corrected image to the corresponding position of the local map pre-constructed based on the target area, and then realize the positioning of the pedestrian on the local map.

[0152] Step S4 includes:

[0153] Based on the perspective transformation matrix from each camera to the specified camera, map the corresponding position of the foot on the corrected image to the corresponding position of the local map, and perform pedestrian positioning based on the local map.

[0154] The perspective transformation matrix H from any camera to the main camera has been obtained in the above steps. Refer to Figure 12 , the coordinate point (x, y) in the local map can be obtained by the following calculation for the coordinate point (u, v) in the corrected image:

[0155]

[0156] Moreover, after step s4, it also includes: Through the preset multi-target final algorithm, perform cross-shot trajectory tracking on the pedestrian according to the corresponding position of the foot on the local map. Such asFigure 13 As shown, after mapping the detection targets under all cameras to the local map through the above method, trajectory tracking is performed on the local map. The method of trajectory tracking is based on a multi-object tracking algorithm of deep learning, including but not limited to: SORT; and Figure 13 each dot in it represents a pedestrian, and the curve "dragged" by the origin is the historical trajectory of the pedestrian in the previous 5 seconds. The number "38:72" beside the dot indicates that the unique number of this pedestrian is 38 (the number used to distinguish different pedestrians) and this pedestrian has stayed in the area for a total of 72 seconds.

[0157] In a specific embodiment, the present invention uses a fisheye camera because it has a large field of view and can cover the entire indoor area with fewer cameras. Each fisheye camera is installed on the ceiling, and its camera layout / installation process is as follows:

[0158] (1) According to the CAD drawing, determine the activity area concerned by the business side. By default, cameras are installed within this activity area.

[0159] (2) Count the different heights within the area.

[0160] (3) Add cameras above each large seating area, that is, meet the following Principle 1.

[0161] (4) Starting from the cameras that have been added, recursively spread outwards. The spreading method needs to meet the following Principle 2 and Principle 4.

[0162] (5) Spread until the entire area is covered, that is, meet the following Principle 3.

[0163] (6) During on-site installation, the following Principle 5 needs to be met.

[0164] Principle 1: Cameras are preferentially installed directly above the seating area.

[0165] Principle 2: The distance between adjacent cameras cannot exceed a predetermined value (the predetermined value can be found in the "Camera Height and Coverage Range" table).

[0166] Principle 3: The coverage ranges of all cameras together can cover the entire activity area.

[0167] Principle 4: As Figure 14-1 、 Figure 14-2 and Figure 14-3 , it can be seen that all cameras are arranged in an equilateral triangle layout with each other.

[0168] The height and coverage range of all cameras are shown in the following table:

[0169]

[0170]

[0171]

[0172]

[0173] Based on the above table, it can be inferred that:

[0174] (1) If the height of Camera A is 2.54 meters, its coverage range can be replaced by a square with a side length of 7.06 meters, and the positions of the cameras can be arranged on the CAD drawing.

[0175] (2) If Camera A and Camera B are adjacent, where the height of Camera A is 2.54 meters and the height of Camera B is 2.58 meters, the distance between Camera A and Camera B cannot exceed 2.82 + 2.87 meters.

[0176] Referring to the following table, it can be seen that in the multi-object trajectory tracking task (MOT) of pedestrians in the industry, the accuracy rate of the existing technologies (black boxes) is generally lower than 80%. Through the combination of hardware and software and algorithm optimization, the present invention improves the accuracy rate to 90%, reaching the commercial level.

[0177] Benchmark Statistics

[0178]

[0179] Meanwhile, the present invention provides a pedestrian cross-camera positioning system based on a top-view camera, including:

[0180] An image acquisition module, configured to acquire the original image of each camera within the target area;

[0181] A head-to-foot mapping module, configured to detect the position of a person's head in the original image and regress to obtain the position of the foot based on the position of the head;

[0182] A corrected map mapping module, configured to map the position of the foot in each original image to the corresponding position on the corrected map preset for the corresponding camera;

[0183] A local map mapping module, configured to map the corresponding position of the foot on the corrected map to the corresponding position on the local map pre-constructed based on the target area, thereby realizing the positioning of the pedestrian on the local map;

[0184] Wherein, the local map is obtained by stitching the corrected maps of each camera.

[0185] In addition, the present invention also provides a computer-readable medium, on which computer-executable instructions are stored, and when the executable instructions are executed by a processor, a pedestrian cross-camera tracking method based on a top-view camera as described above is implemented.

[0186] Since the system / apparatus described in the above embodiments of the present invention is the system / apparatus adopted for implementing the method of the above embodiments of the present invention, those skilled in the art can understand the specific structure and variations of the system / apparatus based on the method described in the above embodiments of the present invention, and thus will not be elaborated herein. Any system / apparatus adopted for the method of the above embodiments of the present invention falls within the scope of protection of the present invention.

[0187] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can be implemented in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can be implemented in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0188] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions.

[0189] It should be noted that in the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present invention can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In the claims listing several devices, several of these devices can be embodied by the same hardware. The use of the words first, second, third, etc. is only for convenience of expression and does not denote any order. These words can be construed as part of the element name.

[0190] In addition, it should be noted that in the description of this specification, the description of terms such as "one embodiment", "some embodiments", "embodiment", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without conflict, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples.

[0191] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications after learning the basic creative concept. Therefore, the claims should be construed to cover the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention.

[0192] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and its equivalent technologies, the present invention should also include these modifications and variations.

Claims

1. A pedestrian cross-camera positioning method based on a top-view camera, characterized in that, Including: Obtain the original images of each camera within the target area; Detect the location of a person's head in the original image, and regress the location of the feet based on the location of the head; Map the location of the feet in each original image to the corresponding location on the correction map preset for the corresponding camera; Map the corresponding location of the feet on the correction map to the corresponding location on the local map pre-constructed based on the target area, thereby realizing the positioning of pedestrians on the local map; Wherein, the local map is obtained by stitching the correction maps of each camera; Detecting the location of a person's head in the original image and regressing the location of the feet based on the location of the head includes: Detect the coordinates of the center point of the person's head in the original image through a pre-trained Yolov5 model; Calculate the corresponding coordinates of the center point of the feet based on the coordinates of the center point of the head in the original image in combination with the regression equation; The pre-trained Yolov5 model is quantized, compressed and deployed on a specified hardware platform; The regression equation is: where, foot i is the coordinate of the center point of the i-th foot, head i is the coordinate of the center point of the head; i = 0, 1, 2, ... N - 1, M is the degree of fitting; W and H are the width and height of the fisheye image respectively; a k (k = 0, 1, 2, ... M) are the optimal parameters obtained by solving N regression equations simultaneously and using the least squares method.

2. The pedestrian cross-camera positioning method based on a top-view camera according to claim 1, wherein, Before obtaining the original images of each camera within the target area, it further includes: After calibrating several cameras arranged in a preset target area, obtain the internal parameter matrix and external parameter matrix of the cameras; Calculate the remapping matrix for correction based on the internal parameter matrix and external parameter matrix of the cameras; Perform distortion correction on the images captured by each camera based on the remapping matrix to obtain the correction map of each camera; Obtain the stitching parameters by performing feature point matching on the correction maps of every two adjacent cameras; Stitch the images captured by each camera according to the stitching parameters to obtain the local map.

3. The pedestrian cross-camera positioning method based on a top-view camera according to claim 2, wherein After calibrating several cameras arranged in a preset target area, obtaining the internal parameter matrix and external parameter matrix of the cameras includes: Obtain the images captured for the chessboard grid preset in the target area; Detect the corner points of the fixed points of the chessboard grid in the captured images to obtain several corner point coordinates in two-dimensional image coordinates; Calculate the internal parameter matrix and external parameter matrix of each camera based on the corner point coordinates of the captured images; And, stitching the images captured by each camera according to the stitching parameters to obtain the local map includes: Based on the perspective transformation matrix in the stitching parameters between every two adjacent cameras and combined with matrix multiplication, obtain the perspective transformation matrix from each camera to the specified camera; Perform perspective transformation on the correction map of each camera based on the perspective transformation matrix from each camera to the specified camera, and then perform stitching to obtain the local map; Wherein, the internal parameter matrix K and external parameter matrix R of each camera satisfy the following formula: wherein, C0......C 53 are respectively the default coordinates of each corner point in the three-dimensional camera coordinate system, and P0......P 53 are the corner point coordinates of each corner point in the two-dimensional image coordinates; The remapping matrix is: In the formula, u and v are the transformed pixel coordinates, x and y are the pixel coordinates corresponding to u and v respectively, K is the internal parameter matrix, R is the external parameter matrix, and p, q, and r are all intermediate calculation results; The splicing parameters include: a perspective transformation matrix H A→B , the perspective transformation matrix H A→B satisfies: where a i and b i are respectively N groups of matching points of adjacent cameras A and B, H A→B is the perspective transformation matrix, i = 0, 1, 2,... N-1, and z is the intermediate calculation result.

4. The pedestrian cross-camera positioning method based on a top-view camera according to claim 2, wherein Mapping the location of the feet in each original image to the corresponding location on the correction map preset for the corresponding camera includes: Based on the remapping matrix, by obtaining the coordinates (u, v) in the corrected image corresponding to each set of coordinates (x, y) of each original image, w*h two-dimensional arrays are obtained; For any set of coordinates (x0, y0) of each original image, the point closest to (x0, y0) among the w*h two-dimensional arrays is found through the nearest neighbor algorithm as the point in the corrected image corresponding to the set of coordinates (x0, y0); where w and h are the width and height of the corrected image respectively.

5. The pedestrian cross-camera positioning method based on a top-view camera according to claim 1, characterized in that Mapping the corresponding position of the foot on the corrected image to the corresponding position of the local map pre-constructed based on the target area, and then realizing the positioning of the pedestrian on the local map includes: Based on the perspective transformation matrix from each camera to the specified camera, mapping the corresponding position of the foot on the corrected image to the corresponding position of the local map, and positioning the pedestrian based on the local map.

6. The pedestrian cross-camera positioning method based on a top-view camera according to claim 5, wherein, After positioning the pedestrian based on the local map, it further includes: through a preset multi-target final algorithm, tracking the trajectory of the pedestrian across cameras according to the corresponding position of the foot on the local map.

7. A pedestrian cross-camera positioning method based on a top-view camera according to any one of claims 1-6, characterized in that The layout of each camera in the preset target area satisfies the following conditions: Preferably set directly above the seat area within the preset target area; The distance between any two adjacent cameras does not exceed a predetermined value, and any three adjacent cameras form an equilateral triangle; The detection ranges of all cameras together cover the entire target area; And each camera is a fish-eye camera, and the image captured by the fish-eye camera is a fish-eye image.

8. A pedestrian cross-camera positioning system based on a top-view camera, characterized in that, It includes: An image acquisition module for acquiring the original images of each camera in the target area; A head-to-foot mapping module for detecting the position of the human head in the original image and regressing to obtain the position of the foot according to the position of the head; A corrected image mapping module for mapping the position of the foot in each original image to the corresponding position of the corrected image preset by the corresponding camera; A local map mapping module for mapping the corresponding position of the foot on the corrected image to the corresponding position of the local map pre-constructed based on the target area, and then realizing the positioning of the pedestrian on the local map; where the local map is obtained by stitching the corrected images of each camera; Detecting the position of the human head in the original image and regressing to obtain the position of the foot according to the position of the head includes: Detecting the center point coordinates of the human head in the original image through a pre-trained Yolov5 model; Calculating the corresponding center point coordinates of the foot according to the center point coordinates of the head in the original image combined with the regression equation; The pre-trained Yolov5 model is quantized and compressed and deployed on a specified hardware platform; The regression equation is: Among them, foot i is the coordinate of the center point of the i-th foot, head i is the coordinate of the center point of the head; i = 0, 1, 2, ... N-1, M is the fitting degree; W and H are the width and height of the fisheye image respectively; a k (k=0, 1, 2, ...M) is the optimal parameter obtained by combining N regression equations and solving them by the least squares method.

9. A computer-readable medium having computer-executable instructions stored thereon, characterized in that, When the executable instructions are executed by the processor, it implements a pedestrian cross-camera tracking method based on a top-view camera as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Multi-image intelligent identification method and device

    CN103795978A

  • Panoramic video rapid splicing method and system

    CN110782394A

  • Multi-image splicing method, system and equipment for target tracking and identification, and medium

    CN114581307A