A method for object detection and tracking and positioning based on traffic monitoring cameras
By calibrating and deep learning detection of monocular surveillance cameras, combined with multi-task deep neural network model, the problem that roadside surveillance cameras cannot obtain the direction angle and three-dimensional dimensions of traffic participants is solved, real-time detection and tracking of traffic participants is achieved, and key information is provided for autonomous driving decisions and planning.
Patent Information
- Application Number
- CN202111402505.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-24
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-11-24
AI Technical Summary
The prior art cannot accurately obtain the direction angle and three-dimensional dimension information of traffic participants in roadside surveillance cameras, affecting autonomous driving decisions and planning.
By calibrating the monocular monitoring camera, the simultaneous matrix of the ground coordinate system to the image coordinate system is obtained, combined with deep learning detection methods and multi-task deep neural network model, the target 2D enclosing box and category recognition is achieved, the DeepSORT algorithm is used for correlation matching and tracking, and the multi-task deep neural network is used for regression calculation of angles and three-dimensional dimensions.
Real-time detection and tracking of traffic participants is realized, and the target category, position, speed, orientation angle and three-dimensional dimensions can be obtained, supporting decision-making and planning of autonomous driving.
Smart Images

Figure CN114332751B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target detection, and in particular to a method for target detection, tracking and positioning based on traffic monitoring cameras. Background Art
[0002] In recent years, with the increase in the number of automobiles, traffic congestion problems and traffic safety problems have become increasingly prominent. In order to improve traffic efficiency, increase safety, and liberate drivers from fatiguing driving work, autonomous driving has come into the public eye. Vehicle-road cooperation is an important branch of autonomous driving, and roadside perception requires sensors installed on the roadside to accurately perceive the surrounding environment and traffic participants in the environment, so as to share the pressure of vehicle-end perception and provide information redundancy, which is the core of vehicle-road cooperation.
[0003] Common sensors used in autonomous driving environment perception include RGB cameras, lidar, and millimeter-wave radars. RGB cameras have the advantages of low cost, easy deployment on the roadside, high data resolution, and rich features, but are greatly affected by lighting conditions and can hardly be used normally at night; millimeter-wave radars have low cost and are not affected by lighting and weather conditions, but have very low resolution; lidar is not easily affected by lighting conditions, can work day and night, has high precision and relatively high resolution, but is easily affected by harsh environments such as rain, fog, and snow.
[0004] Existing technical solutions deploy monitoring cameras on the roadside to complete functions such as the recognition and classification of traffic participants. For example, in the "Method and System for Measuring the Position and Speed of Targets in a Monocular Camera Monitoring Scenario", a monocular monitoring camera installed on the roadside is used to detect real-time images using deep learning methods, and the coordinates of the target in the ground coordinate system are obtained through the homography matrix obtained by calibration. Finally, the speed of the target is obtained through tracking. However, this solution can only obtain the category information, position information, and speed information of the target, and cannot obtain the orientation angle and three-dimensional size of the target, and the latter two play important roles in the decision-making control and planning of autonomous driving. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for target detection, tracking and positioning based on traffic monitoring cameras to overcome the defects of the above-mentioned existing technologies.
[0006] The purpose of the present invention can be achieved by the following technical solutions:
[0007] A method for target detection, tracking and positioning based on traffic monitoring cameras, the method comprising the following steps:
[0008] Step 1: Calibrate the monocular monitoring camera to obtain the homography matrix from the ground coordinate system to the image coordinate system;
[0009] Step 2: Obtain the 2D bounding box of the target and the target category based on the deep learning detection method;
[0010] Step 3: Obtain the ground coordinates of the target in the ground coordinate system according to the coordinates of the 2D bounding box obtained by the detection method in Step 2 in the image coordinate system and the homography matrix;
[0011] Step 6: Extract the images in the 2D bounding box of each target, and perform association matching and tracking of the target based on the association matching and tracking algorithm;
[0012] Step 9: Obtain the speed of the target according to the ground coordinates of the same target in the images of different frames;
[0013] Step 12: Perform regression of the observation orientation angle and the three-dimensional size on the images in the 2D bounding box of each target through a multi-task deep neural network model;
[0014] Step 15: Calculate the orientation angle of the target according to the observation orientation angle and the observation angle.
[0015] In the aforementioned Step 2, the process of obtaining the 2D bounding box of the target and the target category based on the deep learning detection method is specifically as follows:
[0016] Collect images through a calibrated monocular camera to obtain real-time images, and perform frame-by-frame detection on the real-time images based on the deep learning detection method to obtain the 2D bounding box of the target in the real-time images and the target category.
[0017] The deep learning detection method includes yolov3.
[0018] In the aforementioned Step 4, the association matching and tracking algorithm includes the DeepSORT algorithm.
[0019] In the aforementioned Step 6, the backbone part of the multi-task deep neural network model includes vgg19, and the output end of the multi-task deep neural network model includes regression of the observation orientation angle and regression of the three-dimensional size.
[0020] In the aforementioned Step 6, the process of performing regression of the observation orientation angle and the three-dimensional size specifically includes the following steps:
[0021] Step 601: For each target, the expression output by the model is:
[0022]
[0023]
[0024]
[0025] Among them, is the output of the model, and are the parameters output by the model, and are the confidence that the observation orientation angle of the target belongs to the first angle range and the confidence that the observation orientation angle of the target belongs to the second angle range respectively, and are the angle compensation parameters output by the model for determining the first angle range and the angle compensation parameters for determining the second angle range respectively, and are the width deviation of the target width from the average width, the length deviation of the target length from the average length, and the height deviation of the target height from the average height respectively;
[0026] Step 602: Calculate the angle compensation through the parameters and output by the model. The calculation formula for the angle compensation is:
[0027]
[0028]
[0029] where Δθ1 is the angle compensation belonging to the first angle range, and Δθ2 is the angle compensation belonging to the second angle range;
[0030] Step 603: Determine the range B of the observation orientation angle of the target through the confidence:
[0031]
[0032] where B is the range of the observation orientation angle, B1 is the first angle range, and B2 is the second angle range;
[0033] Step 604: Calculate the observation orientation angle θ of the target through the range B of the observation orientation angle l :
[0034] θ l = m + Δθ
[0035]
[0036] where θ l is the observation orientation angle, m is the central angle of the target, m1 is the central angle of the first angle range, m2 is the central angle of the second angle range, and Δθ is the angle compensation of the target;
[0037] Step 605: Calculate the three-dimensional size of the target through and The calculation formulas for the width, length, and height of the target are respectively;
[0038]
[0039]
[0040]
[0041] Among them, w is the width of the target, l is the length of the target, and h is the height of the target. and are the width deviation between the target width and the average width, the length deviation between the target length and the average length, and the height deviation between the target height and the average height, respectively. w a 、l a and h a are the average width of all vehicles, the average length of all vehicles, and the average height of all vehicles obtained by statistics, respectively. w a 、l a and h a are all preset fixed values.
[0042] In the said step 7, the calculation formula for calculating the orientation angle of the target according to the observation orientation angle and the observation angle is:
[0043] θ = θ ray + θ l - 2π
[0044] Among them, θ is the target orientation angle, θ ray is the observation angle, θ l is the observation orientation angle, θ ray is the observation angle. The observation angle is the angle between the line connecting the center point of the target and the origin of the camera coordinate system, projected onto the x-z plane of the camera coordinate system and the axis.
[0045] In the said step 6, the orientation angle cost function for observing the orientation angle using the multi-task deep neural network model is:
[0046]
[0047] Among them, L ori is the orientation angle cost function, softmax is the normalized exponential function, i is the number of angle ranges, N is the number of targets, is the indicator function. The indicator function means that if the condition in the parentheses is satisfied, the function return value is 1, otherwise the return value is 0. θ is the target orientation angle, θ l is the observation orientation angle, is the angle compensation parameter output by the model to determine the i-th angle range, m i is the central angle of the i-th angle range B i ,Δθ i is the i-th angle compensation.
[0048] In step 6 described above, the cost function of the three-dimensional dimension regression task using a multi-task deep neural network model for three-dimensional dimension regression is the L1 cost function, and its formula is:
[0049]
[0050] where L dim is the cost function of the three-dimensional dimension regression task, N is the number of targets, γ k is the true three-dimensional dimension value of the k-th target, is the average three-dimensional dimension of the target category, is the predicted three-dimensional dimension value of the k-th target.
[0051] In step 5 described above, the speed of the target is obtained according to the ground coordinates of the same target in images of different frames, and its formula is:
[0052]
[0053] where (x, y) are the coordinates of this target in the previous frame, (x', y') are the coordinates of this target in this frame, and Δt is the time interval between each frame.
[0054] Compared with the prior art, the present invention has the following advantages:
[0055] 1. By using a traffic monitoring camera to perform real-time detection and tracking of traffic participants, the target categories that can be detected include: pedestrians, two-wheeled vehicles, cars, and trucks;
[0056] 2. By using a traffic monitoring camera to perform real-time detection and tracking of traffic participants, not only can the category, position, and speed of the target be obtained, but also the three-dimensional dimension and orientation angle of the target can be obtained through deep learning methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 is a schematic flow chart of the present invention.
[0058] Figure 2 is a structural diagram of a multi-task deep neural network model.
[0059] Figure 3 is a schematic diagram of the observation orientation angle, observation angle, and orientation angle.
[0060] Figure 4 is a schematic diagram of the effect of the present invention, where the upper part is the monitoring screen and 2D detection results, and the lower part is a bird's-eye view of all targets within the intersection. DETAILED DESCRIPTION OF THE INVENTION
[0061] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0062] Embodiment
[0063] The present invention provides a method for target detection, tracking and positioning based on traffic monitoring cameras. The process is as follows Figure 1 , and the method includes the following steps:
[0064] Step 1: Calibrate the monocular monitoring camera to obtain the homography matrix from the ground coordinate system to the image coordinate system;
[0065] Step 2: Obtain the 2D bounding box and target category of the target based on the deep learning detection method;
[0066] Step 3: Obtain the ground coordinates of the target in the ground coordinate system according to the coordinates of the 2D bounding box obtained by the detection method in step 2 in the image coordinate system and the homography matrix;
[0067] Step 4: Extract the images in the 2D bounding box of each target, and use the DeepSORT algorithm for target association matching and tracking;
[0068] Step 5: Obtain the speed of the target according to the ground coordinates of the same target in the images of different frames;
[0069] Step 6: Perform regression on the observation orientation angle and three-dimensional size of the images in the 2D bounding box of each target through a multi-task deep neural network model;
[0070] Step 7: Calculate the orientation angle of the target according to the observation orientation angle and the observation angle.
[0071] In step 5, the speed of the target is obtained according to the ground coordinates of the same target in the images of different frames, and the formula is:
[0072]
[0073] where (x,y) is the coordinate of this target in the previous frame, (x',y') is the coordinate of this target in this frame, and Δt is the interval time between each frame.
[0074] In step 6, perform regression on the observation orientation angle and three-dimensional size of the images in the 2D bounding box of each target through a multi-task deep neural network model:
[0075] Step 601: For each target, the expression output by the model is:
[0076]
[0077]
[0078]
[0079] where is the model output, and are the parameters of the model output respectively, and are the confidence levels that the observation orientation angle of the target belongs to the first angle range and the confidence level that the observation orientation angle of the target belongs to the second angle range respectively, and are the angle compensation parameters for determining the first angle range and the angle compensation parameters for determining the second angle range output by the model respectively, and are the width deviation of the target width from the average width, the length deviation of the target length from the average length, and the height deviation of the target height from the average height respectively;
[0080] Step 602: Calculate the angle compensation through the parameters and output by the model. The calculation formula for the angle compensation is:
[0081]
[0082]
[0083] where, Δθ1 is the angle compensation belonging to the first angle range, and Δθ2 is the angle compensation belonging to the second angle range;
[0084] Step 603: Determine the range B of the observation orientation angle of the target through the confidence level:
[0085]
[0086] where, B is the range of the observation orientation angle, B1 is the first angle range, and B2 is the second angle range;
[0087] Step 604: Calculate the observation orientation angle θ of the target through the range B of the observation orientation angle l :
[0088] θ l = m + Δθ
[0089]
[0090] where, θ l is the observation orientation angle, m is the central angle of the target, m1 is the central angle of the first angle range, m2 is the central angle of the second angle range, and Δθ is the angle compensation of the target;
[0091] Step 605: Through and Calculate the target three-dimensional size. The calculation formulas for the width, length, and height of the target are as follows:
[0092]
[0093]
[0094]
[0095] where w is the width of the target, l is the length of the target, and h is the height of the target. and are the width deviation of the target width from the average width, the length deviation of the target length from the average length, and the height deviation of the target height from the average height, respectively. w a 、l a and h a are the average width of all vehicles, the average length of all vehicles, and the average height of all vehicles obtained through statistics, respectively. w a 、l a and h a are all preset fixed values.
[0096] The orientation angle cost function for observing the orientation angle using a multi-task deep neural network model is:
[0097]
[0098] where L ori is the orientation angle cost function, softmax is the normalization exponential function, i is the number of angle ranges, N is the number of targets, is the indicator function. The indicator function means that if the condition in the parentheses is satisfied, the function return value is 1, otherwise the return value is 0. θ is the target orientation angle, θ l is the observed orientation angle, is the angle compensation parameter output by the model to determine the i-th angle range, m i is the central angle of the i-th angle range B i and Δθ i is the i-th angle compensation;
[0099] The cost function for the three-dimensional size regression task using a multi-task deep neural network model for three-dimensional size regression is the L1 cost function, and its formula is:
[0100]
[0101] where L dim is the cost function for the three-dimensional size regression task, N is the number of targets, γ k is the true value of the three-dimensional size of the k-th target, is the average three-dimensional size of the target category, is the predicted three-dimensional size of the k-th target.
[0102] In step 7, the calculation formula for the orientation angle of the target based on the observation orientation angle and the observation angle is:
[0103] θ = θ ray + θ l - 2π
[0104] where θ is the target orientation angle, θ ray is the observation angle, θ l is the observation orientation angle, that is, the true value of the observation angle, θ ray is the observation angle. The observation angle is the angle between the line connecting the center point of the target and the origin of the camera coordinate system, projected onto the x-z plane of the camera coordinate system and the axis.
[0105] The deep learning detection method includes yolov3.
[0106] Using the deep learning method to perform regression on the three-dimensional size and orientation angle of the detected pedestrians, cars, trucks and two-wheelers, the designed multi-task deep neural network model only needs to input the RGB image of the target to directly calculate the length, width, height of the target and the orientation angle of the target at the same time. The orientation angle of the target is the angle between the direction the target faces on the x-y plane in the ground coordinate system and the x-axis.
[0107] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or substitutions within the technical scope disclosed by the present invention, and these modifications or substitutions should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for object detection, tracking and positioning based on traffic monitoring cameras, characterized in that, The method includes the following steps: Step 1: Calibrate the monocular monitoring camera to obtain the homography matrix from the ground coordinate system to the image coordinate system; Step 2: Obtain the 2D bounding box and target category of the target based on the deep learning detection method; Step 3: Obtain the ground coordinates of the target in the ground coordinate system according to the coordinates of the 2D bounding box in the image coordinate system obtained by the detection method in Step 2 and the homography matrix; Step 4: Extract the images in the 2D bounding box of each target, and perform association matching and tracking of the target based on the association matching and tracking algorithm; Step 5: Obtain the speed of the target according to the ground coordinates of the same target in the images of different frames; Step 6: Regress the observation orientation angle and three-dimensional size of the images in the 2D bounding box of each target through a multi-task deep neural network model; Step 7: Calculate the orientation angle of the target according to the observation orientation angle and the observation angle; In the said Step 6, the process of regressing the observation orientation angle and three-dimensional size specifically includes the following steps: Step 601: For each target, the expression output by the model is: Among them, is the model output, and are the parameters of the model output respectively, and are the confidence that the observation orientation angle of the target belongs to the first angle range and the confidence that the observation orientation angle of the target belongs to the second angle range respectively, and are the angle compensation parameters used to determine the first angle range and the angle compensation parameters used to determine the second angle range of the model output respectively, and are the width deviation of the target width from the average width, the length deviation of the target length from the average length, and the height deviation of the target height from the average height respectively; Step 602: Using the parameters output by the model and calculate the angle compensation. The calculation formula for the angle compensation is as follows: where, Δθ1 is the angle compensation belonging to the first angle range, and Δθ2 is the angle compensation belonging to the second angle range; Step 603: Determine the range B of the observation orientation angle of the target through the confidence level: where, B is the range of the observation orientation angle, B1 is the first angle range, and B2 is the second angle range; Step 604: Calculate the observed orientation angle θ of the target by observing the range B of the orientation angle l : θ l = m + Δθ where, θ l is the viewing orientation angle, m is the central angle of the target, m1 is the central angle of the first angular range, m2 is the central angle of the second angular range, and Δθ is the angular compensation of the target; Step 605: By and calculate the target three-dimensional dimensions, and the calculation formulas for the width, length, and height of the target are respectively; where w is the width of the target, l is the length of the target, and h is the height of the target, and are the width deviation between the target width and the average width, the length deviation between the target length and the average length, and the height deviation between the target height and the average height, respectively. w a 、l a and h a are the average width of all vehicles, the average length of all vehicles, and the average height of all vehicles obtained by statistics, respectively. w a 、l a and h a are all preset fixed values.
2. The object detection and tracking and positioning method based on traffic monitoring cameras according to claim 1, characterized in that, In the said Step 2, the process of obtaining the 2D bounding box and target category of the target based on the deep learning detection method is specifically: Collect images through the calibrated monocular camera to obtain real-time images, and perform frame-by-frame detection on the real-time images based on the deep learning detection method to obtain the 2D bounding box and target category of the target in the real-time images.
3. The object detection and tracking and positioning method based on traffic monitoring cameras according to claim 2, characterized in that, The said deep learning detection method includes yolov3.
4. A method for target detection, tracking and positioning based on traffic monitoring cameras according to claim 1, characterized in that, In the said Step 4, the association matching and tracking algorithm includes the DeepSORT algorithm.
5. A target detection and tracking and positioning method based on traffic monitoring cameras according to claim 1, characterized in that In the said Step 6, the backbone part of the multi-task deep neural network model includes vgg19, and the output end of the multi-task deep neural network model includes regressing the observation orientation angle and regressing the three-dimensional size.
6. The object detection and tracking and positioning method based on traffic monitoring cameras according to claim 1, characterized in that, In the said Step 7, the calculation formula for calculating the orientation angle of the target according to the observation orientation angle and the observation angle is: θ = θ ray + θ l - 2π where θ is the target orientation angle, θ ray is the viewing angle, θ l is the viewing orientation angle. The viewing angle is the angle between the line connecting the target center point and the origin of the camera coordinate system, after being projected onto the x-z plane of the camera coordinate system, and the axis.
7. A method for target detection, tracking and positioning based on traffic monitoring cameras according to claim 1, characterized in that, In the said Step 6, the orientation angle cost function for observing the orientation angle by using the multi-task deep neural network model is: Among them, L ori is the orientation angle cost function, softmax is the normalization exponential function, i is the number of angle ranges, N is the number of targets, is the indicator function. The indicator function means that if the condition in the parentheses is satisfied, the function returns a value of 1, otherwise the return value is 0. θ is the target orientation angle, θ l is the observed orientation angle, is the angle compensation parameter output by the model to determine the i-th angle range, m i is the central angle of the i-th angle range B i and Δθ i is the i-th angle compensation.
8. A method for target detection, tracking and positioning based on traffic monitoring cameras according to claim 6, characterized in that, In the said Step 6, the three-dimensional size regression task cost function for performing three-dimensional size regression by using the multi-task deep neural network model is the L1 cost function, and its formula is: Among them, L dim is the cost function for the three-dimensional size regression task, N is the number of targets, and γ k is the true three-dimensional size value of the k-th target, is the average three-dimensional size of the target category, is the predicted three-dimensional size value of the k-th target.
9. A method for target detection, tracking and positioning based on traffic monitoring cameras according to claim 1, characterized in that, In the said Step 5, the formula for obtaining the speed of the target according to the ground coordinates of the same target in the images of different frames is: where, (x,y) are the coordinates of this target in the previous frame, (x',y') are the coordinates of this target in this frame, and Δt is the time interval between each frame.
Citation Information
Patent Citations
Target position and speed measurement method and system based on monocular camera monitoring scene
CN110929567A
Target detection method and device, model training method and device, electronic equipment and medium
CN112487979A