A method and device for estimating obstacle position and speed
Through multi-camera, the circumferential image is collected and key point detection and pose solving, combined with pose fusion and multi-object tracking, the problem of dynamic obstacle position and velocity estimation in the prior art is solved, and high-accurate obstacle identification and tracking is achieved.
Patent Information
- Application Number
- CN202111326788.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-10
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2041-11-10
AI Technical Summary
In the existing automatic parking technology, it is difficult to effectively identify and estimate the position and speed of dynamic obstacles, especially when the distance and size are different, the calculation amount is large and the error is large.
By collecting circumferential images by multiple cameras installed at different positions of the vehicle body, performing key point detection, solving the position information of obstacles, performing posture fusion, performing multi-object tracking, and calculating the speed estimate of obstacles.
Accurate estimation of dynamic obstacle position and speed is achieved, computational complexity and error are reduced, and the accuracy of obstacle identification and tracking is improved.
Smart Images

Figure CN114022866B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of automatic parking, and in particular to a method and device for estimating the position and speed of an obstacle. Background Art
[0002] In recent years, with the rapid development of autonomous driving technology, the possibility of using autonomous driving vehicles in daily life has become increasingly greater. Automatic parking, as an indispensable part of the autonomous driving process, has gradually attracted widespread attention.
[0003] In the prior art, the method of simply using the surround view image of the parking vehicle to identify obstructions for parking control lacks the spatial position information of obstacles, and cannot distinguish and identify obstacles of different distances and sizes; the method of combining target detection with a depth estimation network has a high amount of calculation and a large depth estimation error, and the spatial information of objects around the vehicle is not accurately recognized. The speed estimation is after the depth network information, and there is a risk of error superposition; the method of using a binocular camera to use parallax and match pixel points for depth estimation has a complex algorithm and a large amount of calculation; the solution using lidar is expensive, and lidar is prone to blind spots, has limited installation locations, and occupies more space than cameras. The position of the vehicle is marked in the 3D point cloud and the marked 3D bounding box is mapped to the 2D image. This process increases the cost of data annotation and reduces the annotation efficiency.
[0004] Therefore, how to estimate the position and velocity of dynamic obstacles more effectively is an urgent problem to be solved. Summary of the invention
[0005] In view of this, the present invention provides a method and device for estimating the position and velocity of an obstacle, which can effectively estimate the position and velocity of a dynamic obstacle. The technical solution is as follows:
[0006] The present invention provides a method for estimating the position and speed of an obstacle, comprising:
[0007] Multiple cameras installed at different positions on the vehicle body are used to collect surround images around the vehicle body;
[0008] Perform key point detection on the surround view image to obtain the target prediction box and key point information of the obstacle;
[0009] At least based on the key point information, calculate the position and posture information of the obstacle;
[0010] The position and posture information of the obstacle is fused to obtain the position and posture fusion information of the obstacle;
[0011] According to the obstacle's position fusion information, multi-target obstacle tracking is performed to obtain the obstacle's motion trajectory information;
[0012] The obstacle's velocity estimate is calculated based on the obstacle's motion trajectory information and the obstacle's position and posture fusion information.
[0013] Preferably, performing key point detection on the surround view image includes:
[0014] The surround view image is input into the key point detection model to perform key point detection on the surround view image through the key point detection model, wherein the key point detection model is trained using images containing obstacles as training samples and using the key point information of the marked obstacles as sample labels.
[0015] Preferably, before calculating the position and posture information of the obstacle at least according to the key point information, the method further includes:
[0016] The vehicle type recognition model is used to recognize the vehicle type of the surround view image, so as to output the target vehicle type information of the vehicle when the obstacle contains the vehicle, wherein the vehicle type recognition model is trained by using the image containing the vehicle as the training sample and the vehicle type information of the annotated vehicle as the sample label;
[0017] If the obstacle includes a vehicle, the position information of the obstacle is calculated based on at least the key point information, including:
[0018] According to the key point information and target vehicle model information, the position and posture information of the obstacle is solved.
[0019] Preferably, if the obstacle includes a vehicle, the key points in the key point information include non-wheel key points and wheel key points, the key point information includes non-wheel key point information and wheel key point information, the non-wheel key point information is the coordinate information of the non-wheel key point in the pixel coordinate system, the wheel key point information is the coordinate information of the wheel key point in the pixel coordinate system, and the wheel key point is the tangent point between the wheel of the vehicle and the ground;
[0020] According to the key point information and target vehicle information, the position and posture information of the obstacle is calculated, including:
[0021] In the set of non-wheel key points, a subset of non-wheel key points is obtained by random sampling with probability p;
[0022] Query the vehicle model data corresponding to the target vehicle model information in the vehicle model database, and calculate the initial value of the obstacle's pose solution in the vehicle rear axle coordinate system based on the vehicle model data and the non-wheel key point subset information, wherein the non-wheel key point subset information is the non-wheel key point information corresponding to the non-wheel key point subset;
[0023] The Levin-Markberg algorithm is used to perform nonlinear optimization on the initial values of the obstacle pose solution, where the goal of the nonlinear optimization is to minimize the sum of the errors between the key point projection information and the key point information. The key point projection information is the coordinate information in the pixel coordinate system obtained by projecting the wheel key points and non-wheel key points onto the surround view image based on the key point information, the camera model corresponding to the camera, and the initial values of the pose solution.
[0024] Under the condition that the Z-axis coordinate value of the wheel key point in the vehicle rear axle coordinate system is 0 and the internal and external parameters contained in the camera model are known, the coordinate information of the wheel key point in the vehicle rear axle coordinate system is solved by a closed-form solution of the projection model according to the wheel key point information;
[0025] Combined with the coordinate information of the key points of the wheel in the vehicle's rear axle coordinate system, the initial value of the pose solution after nonlinear optimization is terminated to obtain the pose information of the obstacle in the vehicle's rear axle coordinate system.
[0026] Preferably, the initial value of the position and posture solution of the obstacle in the vehicle rear axle coordinate system is calculated based on the vehicle model data and the non-wheel key point subset information, including:
[0027] According to the vehicle model data and the non-wheel key point subset information, the coordinate information of each non-wheel key point in the non-wheel key point subset in the vehicle rear axle coordinate system is calculated;
[0028] According to the coordinate information of each non-wheel key point in the non-wheel key point subset in the vehicle rear axle coordinate system, an efficient perspective point algorithm is used to calculate the initial value of the obstacle's pose solution in the vehicle rear axle coordinate system.
[0029] Preferably, the coordinate information of the key points of the wheels in the coordinate system of the vehicle rear axle is combined to determine the termination of the initial value of the pose solution after nonlinear optimization, so as to obtain the pose information of the obstacle in the coordinate system of the vehicle rear axle, including:
[0030] Based on the initial value of the pose solution after nonlinear optimization and the wheel key point information, the coordinate information of the wheel key point projected to the vehicle rear axle coordinate system is determined;
[0031] Calculate the arithmetic mean of the Euclidean distance between the coordinate information of the wheel key point projected to the vehicle rear axle coordinate system and the coordinate information of the wheel key point in the vehicle rear axle coordinate system;
[0032] When the arithmetic mean of the Euclidean distance is less than a first preset threshold, the initial value of the posture solution after nonlinear optimization is used as the posture information of the obstacle in the rear axle coordinate system of the vehicle;
[0033] When the arithmetic mean of the Euclidean distance is greater than or equal to the first preset threshold, the solution count value is increased by 1, and when the solution count value is not equal to the preset number of solutions, the execution is returned to the set composed of non-wheel key points, and a subset of non-wheel key points is obtained by random sampling with probability p, wherein the initial value of the solution count value is 0.
[0034] Preferably, the position information of the obstacle includes the three-dimensional bounding box information and position information of the obstacle;
[0035] The obstacle's position and posture information is fused to obtain the obstacle's position and posture fusion information, including:
[0036] According to the location information of the obstacle, the three-dimensional bounding box information of the obstacle is separated into four quadrants;
[0037] For each quadrant:
[0038] Grouping the three-dimensional bounding box information located in the consensus area of the camera in the quadrant into groups of two each to obtain a plurality of three-dimensional bounding box information groups;
[0039] For each 3D bounding box information group, the Euclidean distance of the 3D bounding box information group is calculated according to two 3D bounding box information in the 3D bounding box information group, and it is determined whether the Euclidean distance is less than a second preset threshold value. If so, the obstacle corresponding to the 3D bounding box information group is regarded as the same obstacle, and the average value of the 3D bounding box information group is used as the 3D bounding box fusion information of the same obstacle, or the 3D bounding box information closest to the vehicle body in the 3D bounding box information group is used as the 3D bounding box fusion information of the same obstacle; the position fusion information of the same obstacle is determined according to the 3D bounding box fusion information of the same obstacle; the posture fusion information of the same obstacle is determined according to the 3D bounding box fusion information and the position fusion information of the same obstacle.
[0040] Preferably, the obstacle's position and posture fusion information includes the obstacle's three-dimensional bounding box fusion information;
[0041] According to the obstacle's position fusion information, multi-target obstacle tracking is performed to obtain the obstacle's motion trajectory information, including:
[0042] Calculate the state information of the obstacle according to the fusion information of the three-dimensional bounding box of the obstacle, wherein the state information includes the position coordinates of the center of the fused three-dimensional bounding box, the aspect ratio and height of the fused three-dimensional bounding box, and the speed information of the position coordinates, aspect ratio and height in the pixel coordinate system respectively;
[0043] According to the state information of the obstacle, the Kalman filter is used to predict the movement trajectory of the obstacle to obtain the predicted position information of the obstacle at the set time;
[0044] Based on the Hungarian algorithm, the overlap between the predicted position information of the obstacle at a set time and the three-dimensional bounding box information of the obstacle is used as the cost matrix to update the movement trajectory of the obstacle.
[0045] Preferably, the obstacle's posture fusion information includes the obstacle's position fusion information;
[0046] According to the obstacle's motion trajectory information and the obstacle's position fusion information, the obstacle's velocity estimation value is calculated, including:
[0047] The obstacle position fusion information is converted into position fusion information in the global coordinate system according to the movement trajectory information of the obstacle, so as to obtain the obstacle position fusion information in the global coordinate system at the sampling time of two adjacent frames of surround view images;
[0048] Calculate the difference of the position fusion information of the obstacle in the global coordinate system at the sampling time of two adjacent frames of surround view images, and the difference of the sampling time of two adjacent frames of surround view images;
[0049] According to the calculated difference of the position fusion information and the difference of the acquisition time, the velocity estimation value of the obstacle in different directions in the global coordinate system is calculated.
[0050] The present invention provides a device for estimating the position and speed of an obstacle, comprising:
[0051] A collection device, used to collect surround images around the vehicle body through multiple cameras installed at different positions of the vehicle body;
[0052] The key point detection module is used to detect key points of the surround view image and obtain the target prediction box and key point information of the obstacle;
[0053] A posture solving module, used for solving the posture information of the obstacle based on at least the key point information;
[0054] A posture fusion module is used to fuse the posture information of obstacles to obtain the fusion information of the posture of obstacles;
[0055] The target tracking module is used to track multiple obstacles based on the position and posture fusion information of the obstacles to obtain the movement trajectory information of the obstacles;
[0056] The speed estimation module is used to calculate the speed estimation value of the obstacle based on the motion trajectory information of the obstacle and the fusion information of the obstacle's position and posture.
[0057] In summary, the present invention discloses a method for estimating the position and speed of an obstacle. First, a plurality of cameras installed at different positions of the vehicle body are used to collect surround images around the vehicle body; then key point detection is performed on the surround images to obtain the target prediction frame and key point information of the obstacle; the position and posture information of the obstacle is solved at least according to the key point information of the obstacle; the position and posture information of the obstacle is fused to obtain the position and posture fusion information of the obstacle; multi-target tracking of the obstacle is performed according to the position and posture fusion information of the obstacle to obtain the motion trajectory information of the obstacle; and the speed estimation value of the obstacle is calculated according to the motion trajectory information of the obstacle and the position and posture fusion information of the obstacle. The present invention can estimate the position and posture information of the obstacle using two-dimensional surround images without using three-dimensional annotation data. After calculating the position and posture fusion information based on the position and posture information of the obstacle, the dynamic obstacle can be tracked according to the position and posture fusion information, and then the speed of the dynamic obstacle can be estimated using the global coordinate projection. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0059] Figure 1 A method flow chart of an embodiment of a method for estimating the position and velocity of an obstacle disclosed in the present invention;
[0060] Figure 2 It is a schematic diagram of multiple cameras around a vehicle body disclosed in the present invention;
[0061] Figure 3a It is a schematic diagram of key points of the automobile disclosed in the present invention;
[0062] Figure 3b A schematic diagram of key points of a pedestrian disclosed in the present invention;
[0063] Figure 4 A schematic diagram of a posture information solution process disclosed in the present invention;
[0064] Figure 5a It is a schematic diagram of the three-dimensional bounding box information around the vehicle body disclosed in the present invention;
[0065] Figure 5b It is a schematic diagram of the four-quadrant separation disclosed in the present invention;
[0066] Figure 5c It is a schematic diagram of the fusion information of the three-dimensional bounding box around the vehicle body disclosed in the present invention;
[0067] Figure 6 A schematic diagram of tracking and speed estimation disclosed in the present invention;
[0068] Figure 7 The present invention is a schematic structural diagram of an embodiment of an obstacle position and speed estimation device disclosed in the present invention. DETAILED DESCRIPTION
[0069] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0070] like Figure 1 FIG. 1 is a flow chart of an embodiment of a method for estimating the position and velocity of a dynamic obstacle disclosed in the present invention. The method may include the following steps:
[0071] S101, collecting surround images around the vehicle body through multiple cameras installed at different positions of the vehicle body.
[0072] For example, in an optional embodiment, when the position and speed of an obstacle need to be estimated, see Figure 2 As shown, fisheye cameras (cameras such as Figure 2 As shown in the circle, the fisheye camera has a larger field of view and is more friendly to the surround view system. It is worth noting that this application is not limited to fisheye cameras, and other cameras can also be used to collect surround view images around the vehicle body.
[0073] It should be noted that in this step, the obstacles can be dynamic obstacles. For example, obstacles such as pedestrians, vehicles (in the embodiment of the present application, the vehicle specifically refers to a car), bicycles, motorcycles, etc. can all move. This step can collect surround images for the obstacles during the movement of the obstacles.
[0074] Optionally, the above-mentioned multiple cameras can continuously collect surround view images. For example, multiple cameras collect surround view images once every set time period. For example, multiple cameras collect a frame of surround view image (including the image collected by each camera) at a first sampling moment, and then collect a frame of surround view image at a second sampling moment, and so on.
[0075] S102: Perform key point detection on the surround view image to obtain target prediction boxes and key point information of obstacles.
[0076] It can be understood that each obstacle contains key points. For example, for vehicles, the front, rear, left and right headlights of the vehicle body, the tangent points of the front, rear, left and right wheels with the ground, the center points of the front, rear, left and right four wheel hubs, the intersection points of the left and right rearview mirrors with the vehicle body, the front and rear license plates, and the center points of the front and rear roofs can be selected as key points; for pedestrians, the left and right feet and left and right shoulders of pedestrians can be selected as key points; for bicycles and motorcycles, the intersection points of the front and rear wheels with the ground can be selected as key points. For example, the key points of cars and pedestrians detected in this step can be respectively as follows: Figure 3a and 3b shown.
[0077] In this step, the target prediction frame and key point information of the obstacle can be obtained by performing key point detection on the surround view image. It can be understood that the surround view image may include multiple obstacles, and the target prediction frame can be used to determine which obstacle the key point information belongs to. Optionally, the target prediction frame and key point information can be coordinate information, such as coordinate information in a pixel coordinate system. For example, if the obstacle is a vehicle, the target prediction frame and key point information obtained can be [x, y, w, h, p1_x, p1_y, ..., p12_x, p12_y], where x, y represent the center point coordinates of the vehicle (the center point coordinates of the target prediction frame), w, h represent the width and height of the target prediction frame, pl_x, pl_y, etc. represent the coordinates of the key points of the vehicle body, and this information represents the coordinate values of a single vehicle in a pixel coordinate system.
[0078] In an optional embodiment, this step can be implemented by a neural network model, that is, a key point detection model is pre-established, and then the surround view image is input into the key point detection model, so as to perform key point detection on the surround view image through the key point detection model. The key point detection model is trained by using images containing obstacles as training samples and the predicted boxes and key point information of the marked obstacles as sample labels.
[0079] Here, the training process of the key point detection model includes: obtaining an image containing obstacles, annotating the obstacles in the image (such as cars, people, bicycles, motorcycles, etc.) with 2D BoundingBox and key points, and then using Convolutional Neural Networks (CNN) with feature extraction capabilities for training to obtain a deep learning model with generalization capabilities (i.e., the key point detection model), which is used in the target and key point detection module to output the target prediction box and key point information of the obstacles in the surround view image.
[0080] S103: Calculate the position and posture information of the obstacle at least based on the key point information.
[0081] Here, the position information of the obstacle includes but is not limited to: the orientation angle information, three-dimensional bounding box information and position information of the obstacle.
[0082] In this step, if the prior information of the obstacle is known, the position information of the obstacle at the sampling time of the frame of surround view image captured by multiple cameras (including the image captured by each camera) can be solved according to the key point information of the obstacle in the frame of surround view image. If S101 continuously captures surround view images, this step will continue to solve the position information of the obstacle at the sampling time of each frame of surround view image. Here, the prior information of the obstacle refers to information related to the obstacle, such as structural information and other information. For example, when the obstacle is a pedestrian, the prior information can be the average step length information of the pedestrian; when the obstacle is a bicycle or a motorcycle, the prior information can be information such as the distance between the two wheels.
[0083] Considering that in some scenarios (such as automatic parking scenarios), obstacles are mainly vehicles. In order to make the posture information of the vehicle calculated in this step more accurate, the posture information of the obstacle can be solved based on the key point information and the target vehicle model information. Here, the target vehicle model information refers to the vehicle model information, that is, in this step, if the obstacle includes a vehicle, the posture information of the vehicle can be solved based on the key point information and the target vehicle model information of the vehicle, which can effectively improve the accuracy of the vehicle's posture information.
[0084] Based on this, optionally, this step can perform vehicle type recognition on the surround view image before solving the position information of the obstacle, so as to obtain the target vehicle type information of the vehicle when the obstacle in the surround view image includes a vehicle. For example, this step performs vehicle type recognition on the surround view image, and can identify different vehicle type information such as sedans and SUVs.
[0085] In an optional embodiment, in order to improve the accuracy in solving the obstacle posture, this step can add a vehicle model recognition and classification network based on the network of S102, and use the feature extraction and classification capabilities of the neural network to obtain the vehicle model information. Based on this, this step can pre-establish a vehicle model recognition model, and then input the surround view image into the vehicle model recognition model. When the obstacle in the surround view image includes a vehicle, the vehicle model recognition model can output the target vehicle model information of the vehicle. Here, the vehicle model recognition model is trained using images containing vehicles as training samples and the labeled vehicle model information as sample labels.
[0086] In an optional embodiment, in order to enable those skilled in the art to better understand the present solution, the process of "calculating the position information of the obstacle based on the key point information and the target vehicle model information" is explained below by taking the obstacle including a vehicle as an example.
[0087] Optionally, this step can divide the key points of the vehicle into two categories: non-wheel key points and wheel key points. Correspondingly, the key point information obtained in S102 includes non-wheel key point information and wheel key point information, wherein the non-wheel key point information is the coordinate information of the non-wheel key point in the pixel coordinate system, and the wheel key point information is the coordinate information of the wheel key point in the pixel coordinate system. The wheel key point refers to the tangent point between the wheel of the vehicle and the ground, and the non-wheel key point refers to other key points other than the above tangent point. In order to facilitate the description of the following technology, the vehicle rear axle coordinate system is defined: the projection of the center of the vehicle rear axle on the ground is the origin, the coordinate axis perpendicular to the vehicle rear axle and pointing to the front of the vehicle is the X-axis, the coordinate axis perpendicular to the X-axis and pointing to the left side of the vehicle body is defined as the Y-axis, and the coordinate axis perpendicular to the X-axis and vertically upward is defined as the Z-axis. A right-handed coordinate system is used, and the unit is m (this application can determine the conversion relationship between the vehicle rear axle coordinate system and the global coordinate system).
[0088] Based on this, see Figure 4 As shown in the figure, the process of "solving the pose information of the obstacle based on the key point information and the target vehicle model information" may include five parts: non-wheel key point sampling, pose initial value estimation, pose nonlinear optimization, wheel key point pose estimation and pose estimation termination condition judgment:
[0089] a Non-wheel key point sampling: In the set of non-wheel key points, a subset of non-wheel key points is obtained by random sampling with probability p.
[0090] Here, the probability P is an empirical value, which can be set according to actual conditions, and this application does not limit this.
[0091] b. Initial pose estimation: The vehicle model data corresponding to the target vehicle model information (i.e., the vehicle prior information) is queried in the vehicle model database, and the initial pose solution of the obstacle in the vehicle rear axle coordinate system is calculated based on the vehicle model data and the non-wheel key point subset information, where the non-wheel key point subset information is the non-wheel key point information corresponding to the non-wheel key point subset.
[0092] Specifically, this step can query the vehicle model data corresponding to the target vehicle model information in the vehicle model database according to the vehicle model recognition result. For each non-wheel key point in the non-wheel key point subset, the three-dimensional coordinates of the non-wheel key point in the vehicle rear axle coordinate system can be calculated according to the vehicle model data and the coordinate information of the non-wheel key point in the pixel coordinate system in the non-wheel key point subset information, so that the coordinate information of each non-wheel key point in the non-wheel key point subset in the vehicle rear axle coordinate system can be obtained, and then according to the coordinate information of each non-wheel key point in the non-wheel key point subset in the vehicle rear axle coordinate system, the initial value of the obstacle's posture solution in the vehicle rear axle coordinate system can be calculated using the efficient perspective point algorithm (EPnP).
[0093] c Nonlinear optimization of posture: The Levenberg-Marquardt (LM) algorithm is used to perform nonlinear optimization on the initial value of the obstacle's posture solution.
[0094] In this step, during the nonlinear optimization process, the optimization variable is the position and posture information of the obstacle, that is, the Lie algebra corresponding to the origin coordinates and the rotation angle of the vehicle rear axle coordinate system. This step can calculate the error sum of the key point projection information and the key point information obtained in S102, that is, calculate the error between the projection information of each key point (including wheel key points and non-wheel key points) and the corresponding key point information and sum them, and minimize the error sum as the goal of nonlinear optimization. Here, the key point projection information refers to the coordinate information (the coordinate information is two-dimensional coordinate information) in the pixel coordinate system obtained by projecting the wheel key points and non-wheel key points respectively to the surround image according to the key point information, the camera model corresponding to the camera, and the initial value of the position and posture solution.
[0095] d Wheel key point pose estimation: When the Z-axis coordinate value of the wheel key point in the vehicle rear axle coordinate system is 0 and the internal and external parameters contained in the camera model are known, the coordinate information of the wheel key point in the vehicle rear axle coordinate system is solved through a closed-form solution of the projection model based on the wheel key point information.
[0096] Based on the establishment method of the vehicle rear axle coordinate system, it can be known that the Z-axis coordinate value of the wheel key point in the vehicle rear axle coordinate system is 0. Based on this condition, when the camera model and internal and external parameters are known, the coordinates of the wheel point in the vehicle rear axle coordinate system can be solved by the closed-form solution of the projection model based on the wheel key point information. Here, the camera model is the camera model corresponding to the camera in S101, the internal parameter refers to the conversion relationship between the camera model and the pixel coordinate system, and the external parameter refers to the conversion relationship between the camera model and the global coordinate system (world coordinate system).
[0097] e. Termination judgment of pose estimation: Combined with the coordinate information of the key points of the wheel in the coordinate system of the vehicle's rear axle, the initial value of the pose solution after nonlinear optimization is terminated to obtain the pose information of the obstacle in the coordinate system of the vehicle's rear axle.
[0098] In this step, once the nonlinear optimization process converges or reaches the total number of iterations, it is then determined whether the pose estimation process is terminated. Optionally, the determination process may include:
[0099] e1. Based on the initial value of the pose solution after nonlinear optimization and the wheel key point information, determine the coordinate information of the wheel key point projected to the vehicle rear axle coordinate system.
[0100] e2. Calculate the arithmetic mean of the Euclidean distance between the coordinate information of the wheel key point projected to the vehicle rear axle coordinate system and the coordinate information of the wheel key point in the vehicle rear axle coordinate system (i.e., the coordinate information obtained by "d wheel key point pose estimation").
[0101] e3. Determine whether the calculated arithmetic mean of the Euclidean distance is less than a first preset threshold.
[0102] Among them, if the calculated arithmetic mean of the Euclidean distance is less than the first preset threshold, it is considered that the current initial value of the pose solution is valid. At this time, the judgment process of this step is terminated, and the initial value of the pose solution after the current nonlinear optimization is used as the pose information of the obstacle in the vehicle rear axle coordinate system and output; if the calculated arithmetic mean of the Euclidean distance is greater than or equal to the first preset threshold, the judgment process of this step is continued, that is, e4 is executed.
[0103] e4. Add 1 to the solution count value (whose initial value is 0), and determine whether the solution count value is equal to the preset number of solutions. If the solution count value is not equal to the preset number of solutions, return to the execution in the set composed of non-wheel key points, and obtain a subset of non-wheel key points by random sampling with probability p.
[0104] In this step, if the solution count value is not equal to the preset solution count, it is necessary to re-sample the non-wheel key points randomly, that is, it is necessary to return to perform a non-wheel key point sampling until the preset solution count is completed. After the preset solution count is completed (that is, when the solution count value is equal to the preset solution count), if there is still no solution that meets the conditions, the vehicle posture solution fails and an error message is output.
[0105] Through the above steps a to e, it is possible to "calculate the position information of the obstacle based on the key point information and the target vehicle model information". For obstacles such as pedestrians, bicycles, motorcycles, etc., the position information can be calculated by following the above steps a to e, and no further details will be given here.
[0106] S104: Fusing the position and posture information of the obstacle to obtain fused position and posture information of the obstacle.
[0107] After posture solution is performed in S103, the posture information of the obstacle in the vehicle rear axle coordinate system can be obtained, and the posture information includes the three-dimensional bounding box (3D Bounding Box) information and position information of the obstacle. For example, the three-dimensional bounding box information is: [(x1, y1, z1), (x2, y2, z2), (x3, y3, z3), (x4, y4, z4), (x5, y5, z5), (x6, y6, z6), (x7, y7, z7), (x8, y8, z8)], and the coordinates are the values in the vehicle rear axle coordinate system.
[0108] As mentioned above, this application will collect surround images around the vehicle body through multiple cameras installed in different directions of the vehicle body, so there are common viewing areas between the left and front, front and right, right and rear, and rear and left of the camera. The consensus area can be found in Figure 2 The shadow part, for example, for a car located in the consensus area corresponding to the left and front cameras, can be captured by the left and front cameras at the same time. In order to reduce redundant information in tracking, the pose information of obstacles detected in four directions of the vehicle body needs to be fused into the consensus area in the coordinate system of the rear axle of the vehicle.
[0109] Optionally, the process of this step may specifically include: Figure 5a As shown in the figure, the dotted rectangular frame and the solid rectangular frame respectively represent the three-dimensional bounding box information of the obstacles corresponding to different cameras in each direction, among which there are obstacles located in the consensus area between the left and front camera, the right and front camera, and the right and rear camera, see Figure 5b As shown, first, according to the position information of the obstacle in the surround image, the three-dimensional bounding box information of the obstacle is separated into four quadrants. For each quadrant: the three-dimensional bounding box information located in the consensus area of the camera in the quadrant is grouped in pairs to obtain multiple three-dimensional bounding box information groups. For each three-dimensional bounding box information group, the Euclidean distance d of the three-dimensional bounding box information group is calculated according to the two three-dimensional bounding box information in the three-dimensional bounding box information group, and it is determined whether the calculated Euclidean distance d is less than a second preset threshold value k. If d is less than the second preset threshold value k, it means that the two three-dimensional bounding box information are the same target object (i.e., the same obstacle). At this time, the two three-dimensional bounding box information corresponding to the same obstacle can be fused in position and posture. For details, please refer to Figure 5c , Figure 5c Indicates the three-dimensional bounding box fusion information of the obstacle around the vehicle body in the vehicle rear axle coordinate system after the fusion is completed. Optionally, there are multiple implementation methods for the process of performing pose fusion on two three-dimensional bounding box information corresponding to the same obstacle. The embodiment of the present application provides but is not limited to the following two implementation methods:
[0110] The first method is: taking the average value of the 3D bounding box information group as the 3D bounding box fusion information of the same obstacle. For example, the two 3D bounding box information in the 3D bounding box information group both correspond to obstacle A, and the 3D bounding box information is coordinate information, then the corresponding coordinates of the two 3D bounding box information in the 3D bounding box information group are averaged to obtain the 3D bounding box fusion information of obstacle A.
[0111] The second method is to use the 3D bounding box information closest to the vehicle body in the 3D bounding box information group as the 3D bounding box fusion information of the same obstacle. For example, 3D bounding box information 1 and 2 in the 3D bounding box information group both correspond to obstacle B, and 3D bounding box information 1 is closer to the vehicle body in this case (i.e., the vehicle body in S101), then 3D bounding box information 1 is used as the 3D bounding box fusion information of obstacle B.
[0112] In this step, the position fusion information of the same obstacle can be determined based on the 3D bounding box fusion information of the same obstacle, and then the posture fusion information of the same obstacle can be determined based on the 3D bounding box fusion information and the position fusion information of the same obstacle.
[0113] By performing posture fusion in each quadrant according to the method provided in this step, the posture fusion information of each obstacle in the surround image can be obtained.
[0114] S105: Perform multi-target tracking on the obstacles according to the position and posture fusion information of the obstacles to obtain the motion trajectory information of the obstacles.
[0115] In this step, the Tracking-by-Detection strategy is used for target tracking, that is, multiple targets are tracked through the detection results. Specifically, the tracking part can be divided into Kalman filtering and Hungarian matching algorithm to model tracking as a motion model prediction and data association problem.
[0116] Optionally, the process of this step may include:
[0117] S1. Calculate the state information of the obstacle according to the fusion information of the three-dimensional bounding box of the obstacle, wherein the state information includes the position coordinates of the center of the fused three-dimensional bounding box, the aspect ratio and height of the fused three-dimensional bounding box, and the speed information of the position coordinates, aspect ratio and height in the pixel coordinate system respectively.
[0118] The above S104 can obtain the position fusion information of the obstacle at the sampling time of a frame of surround view image, and the position fusion information includes the three-dimensional bounding box fusion information, and the three-dimensional bounding box fusion information includes the corner point information of the fused three-dimensional bounding box. For example, the corner point information is [(x1, y1, z1), (x2, y2, z2), (x3, y3, z3), (x4, y4, z4), (x5, y5, z5), (x6, y6, z6), (x7, y7, z7), (x8, y8, z8)]. In this step, the state information of the obstacle can be calculated according to the three-dimensional bounding box fusion information of the obstacle, that is, in this step, the information of the vehicle target frame can be constructed into an 8-dimensional space to describe the state of the obstacle's motion trajectory at the sampling time of a frame of surround view image. Here, u and v represent the position coordinates of the center of the fused three-dimensional bounding box, γ represents the aspect ratio of the fused three-dimensional bounding box, and h represents the height of the fused three-dimensional bounding box. They represent the velocity information of the position coordinates u, v, aspect ratio γ and height h in the pixel coordinate system respectively.
[0119] S2. According to the state information of the obstacle, the Kalman filter is used to predict the movement trajectory of the obstacle to obtain the predicted position information of the obstacle at a set time.
[0120] In this step, a Kalman filter can be used to predict the movement trajectory of the obstacle, so as to obtain the predicted position information of the obstacle at a set time.
[0121] S3. Based on the Hungarian algorithm, the overlap between the predicted position information of the obstacle at the set time and the three-dimensional bounding box information of the obstacle is used as the cost matrix to update the movement trajectory of the obstacle.
[0122] This step can use the Hungarian algorithm for data association. For each obstacle detected by S102, the predicted position information of the obstacle at the set time and the IoU (Intersection over Union) of the three-dimensional bounding box information of all obstacles (i.e., obstacles detected by S102) at the set time are used as the cost matrix to update the movement trajectory of the obstacle.
[0123] This step uses a multi-target tracking algorithm to obtain information about each obstacle (each obstacle has a different ID) and its motion trajectory.
[0124] S106: Calculate the estimated speed of the obstacle based on the motion trajectory information of the obstacle and the fusion information of the position and posture of the obstacle.
[0125] In this step, the purpose of obstacle speed estimation is to obtain the specific speed of the obstacle movement. Therefore, when the obstacle posture fusion information includes the obstacle position fusion information, it is necessary to transform the obstacle position fusion information into the global coordinate system according to the obstacle's motion trajectory information, so as to obtain the obstacle position fusion information in the global coordinate system at the sampling time of two adjacent frames of surround view images captured by the camera. After that, the difference in the obstacle position fusion information in the global coordinate system at the sampling time of two adjacent frames of surround view images, as well as the difference in the sampling time of two adjacent frames of surround view images can be calculated. According to the calculated difference in position fusion information and the difference in sampling time, the speed estimation values vx and vy of the obstacle in the x and y directions in the global coordinate system can be calculated. The schematic diagram of the tracking and speed estimation is shown in FIG. Figure 6 As shown, where [1,(v x ,v y )] indicates that the obstacle ID being tracked is 1, the speed estimation in the x direction is vx, and the speed estimation in the y direction is vy. The same is true for other information, which will not be described here.
[0126] In summary, the system of the present invention can estimate the pose information of obstacles using two-dimensional surround view images without using three-dimensional annotation data. After calculating the pose fusion information based on the pose information of the obstacle, the dynamic obstacle can be tracked according to the pose fusion information, and then the speed of the dynamic obstacle can be estimated using global coordinate projection. Since only the target prediction frame and key point information of the obstacle in the two-dimensional surround view image need to be annotated, the data annotation cost of the method of this case is low. Moreover, since obstacles and key points can be directly detected in the collected surround view image, it can be applied even when the surround view image is distorted. Therefore, it is very suitable for vehicle-mounted surround view systems.
[0127] like Figure 7 FIG. 1 is a schematic diagram of a structure of an embodiment of an obstacle posture and speed estimation device disclosed in the present invention. The obstacle posture and speed estimation device may include:
[0128] A collection module 701 is used to collect surround images around the vehicle body through multiple cameras installed at different positions of the vehicle body;
[0129] A key point detection module 702 is used to perform key point detection on the surround view image to obtain a target prediction frame and key point information of obstacles;
[0130] A posture solving module 703, used to solve the posture information of the obstacle at least according to the key point information;
[0131] A posture fusion module 704 is used to fuse the posture information of the obstacle to obtain the fused posture information of the obstacle;
[0132] The target tracking module 705 is used to perform multi-target tracking on the obstacles according to the position and posture fusion information of the obstacles to obtain the motion trajectory information of the obstacles;
[0133] The speed estimation module 706 is used to calculate the speed estimation value of the obstacle according to the motion trajectory information of the obstacle and the position and posture fusion information of the obstacle.
[0134] In summary, the working principle of the obstacle posture and speed estimation device disclosed in this embodiment is the same as the working principle of the obstacle posture and speed estimation method disclosed in the above embodiments, and will not be repeated here.
[0135] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0136] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0137] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0138] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for estimating the position and velocity of an obstacle, characterized in that: include: Collecting surround images around the vehicle body by using multiple cameras installed at different positions of the vehicle body; Performing key point detection on the surround view image to obtain target prediction frames and key point information of obstacles; Calculating the position and posture information of the obstacle at least according to the key point information; Fusing the position and posture information of the obstacle to obtain position and posture fusion information of the obstacle; Performing multi-target tracking on the obstacle according to the position and posture fusion information of the obstacle to obtain the motion trajectory information of the obstacle; Calculating a velocity estimation value of the obstacle according to the motion trajectory information of the obstacle and the position and posture fusion information of the obstacle; The position information of the obstacle includes the three-dimensional bounding box information and position information of the obstacle; The position and posture information of the obstacle is fused to obtain the position and posture fusion information of the obstacle, including: According to the position information of the obstacle, four-quadrant separation is performed on the three-dimensional bounding box information of the obstacle; For each quadrant: Grouping the three-dimensional bounding box information located in the consensus area of the camera in the quadrant into groups of two each to obtain a plurality of three-dimensional bounding box information groups; For each 3D bounding box information group, the Euclidean distance of the 3D bounding box information group is calculated according to two 3D bounding box information in the 3D bounding box information group, and it is determined whether the Euclidean distance is less than a second preset threshold value. If so, the obstacles corresponding to the 3D bounding box information group are regarded as the same obstacle, and the average value of the 3D bounding box information group is used as the 3D bounding box fusion information of the same obstacle, or the 3D bounding box information closest to the vehicle body in the 3D bounding box information group is used as the 3D bounding box fusion information of the same obstacle; the position fusion information of the same obstacle is determined according to the 3D bounding box fusion information of the same obstacle; and the posture fusion information of the same obstacle is determined according to the 3D bounding box fusion information and the position fusion information.
2. The method for estimating the position and velocity of an obstacle according to claim 1, characterized in that: The performing key point detection on the surround view image comprises: The surround view image is input into a key point detection model to perform key point detection on the surround view image through the key point detection model, wherein the key point detection model is trained using images containing obstacles as training samples and using key point information of marked obstacles as sample labels.
3. The method for estimating the position and velocity of an obstacle according to claim 1, characterized in that: Before calculating the position and posture information of the obstacle at least according to the key point information, the method further includes: The vehicle type recognition model is used to recognize the vehicle type of the surround view image, so as to output the target vehicle type information of the vehicle when the obstacle contains a vehicle, wherein the vehicle type recognition model is trained by using images containing vehicles as training samples and the vehicle type information of the annotated vehicles as sample labels; If the obstacle includes a vehicle, calculating the position information of the obstacle at least based on the key point information includes: The position and posture information of the obstacle is calculated according to the key point information and the target vehicle model information.
4. The method for estimating the obstacle position and velocity according to claim 3, characterized in that: If the obstacle includes the vehicle, the key points in the key point information include non-wheel key points and wheel key points, the key point information includes non-wheel key point information and wheel key point information, the non-wheel key point information is the coordinate information of the non-wheel key point in the pixel coordinate system, the wheel key point information is the coordinate information of the wheel key point in the pixel coordinate system, and the wheel key point is the tangent point between the wheel of the vehicle and the ground; The step of calculating the position and posture information of the obstacle according to the key point information and the target vehicle model information includes: In the set consisting of the non-wheel key points, a subset of non-wheel key points is obtained by random sampling with probability p; Querying the vehicle model data corresponding to the target vehicle model information in the vehicle model database, and calculating the initial value of the position and posture solution of the obstacle in the vehicle rear axle coordinate system according to the vehicle model data and the non-wheel key point subset information, wherein the non-wheel key point subset information is the non-wheel key point information corresponding to the non-wheel key point subset; A Levin-Markberg algorithm is used to perform nonlinear optimization on the initial value of the pose solution of the obstacle, wherein the goal of the nonlinear optimization is to minimize the sum of errors between key point projection information and the key point information, and the key point projection information is coordinate information in the pixel coordinate system obtained by projecting the wheel key point and the non-wheel key point onto the surround view image according to the key point information, the camera model corresponding to the camera and the initial value of the pose solution; Under the condition that the Z-axis coordinate value of the wheel key point in the vehicle rear axle coordinate system is 0 and the internal and external parameters contained in the camera model are known, the coordinate information of the wheel key point in the vehicle rear axle coordinate system is solved by a closed-form solution of a projection model according to the wheel key point information; Combined with the coordinate information of the wheel key points in the vehicle rear axle coordinate system, a termination judgment is made on the initial value of the posture solution after nonlinear optimization to obtain the posture information of the obstacle in the vehicle rear axle coordinate system.
5. The method for estimating the position and speed of an obstacle according to claim 4, characterized in that: The step of calculating the initial value of the obstacle's position and posture in the vehicle rear axle coordinate system according to the vehicle model data and the non-wheel key point subset information includes: Calculate the coordinate information of each non-wheel key point in the non-wheel key point subset in the vehicle rear axle coordinate system according to the vehicle model data and the non-wheel key point subset information; According to the coordinate information of each non-wheel key point in the non-wheel key point subset in the vehicle rear axle coordinate system, an efficient perspective point algorithm is used to calculate the initial value of the position and posture solution of the obstacle in the vehicle rear axle coordinate system.
6. The method for estimating the position and speed of an obstacle according to claim 4, characterized in that: The step of combining the coordinate information of the wheel key point in the vehicle rear axle coordinate system, performing termination judgment on the initial value of the pose solution after nonlinear optimization, and obtaining the pose information of the obstacle in the vehicle rear axle coordinate system, includes: Based on the initial value of the pose solution after the nonlinear optimization and the wheel key point information, determining the coordinate information of the wheel key point projected to the vehicle rear axle coordinate system; Calculating the arithmetic mean of the Euclidean distance between the coordinate information of the wheel key point projected to the vehicle rear axle coordinate system and the coordinate information of the wheel key point in the vehicle rear axle coordinate system; When the arithmetic mean of the Euclidean distance is less than a first preset threshold, using the initial value of the posture solution after the nonlinear optimization as the posture information of the obstacle in the rear axle coordinate system of the vehicle; When the arithmetic mean of the Euclidean distance is greater than or equal to the first preset threshold value, the solution count value is increased by 1, and when the solution count value is not equal to the preset number of solutions, the execution is returned to obtain a subset of non-wheel key points by random sampling with probability p in the set composed of the non-wheel key points, wherein the initial value of the solution count value is 0.
7. The method for estimating obstacle position and velocity according to claim 4, characterized in that: The position and pose fusion information of the obstacle includes the three-dimensional bounding box fusion information of the obstacle; The performing multi-target tracking on the obstacle according to the position fusion information of the obstacle to obtain the motion trajectory information of the obstacle includes: Calculating state information of the obstacle according to the fusion information of the three-dimensional bounding box of the obstacle, wherein the state information includes the position coordinates of the center of the fused three-dimensional bounding box, the aspect ratio and height of the fused three-dimensional bounding box, and speed information of the position coordinates, the aspect ratio and the height in the pixel coordinate system respectively; According to the state information of the obstacle, a Kalman filter is used to predict the movement trajectory of the obstacle to obtain the predicted position information of the obstacle at a set time; Based on the Hungarian algorithm, the overlap between the predicted position information of the obstacle at the set time and the three-dimensional bounding box information of the obstacle is used as a cost matrix to update the movement trajectory of the obstacle.
8. The method for estimating the position and velocity of an obstacle according to claim 7, characterized in that: The pose fusion information of the obstacle includes the position fusion information of the obstacle; The calculating the estimated speed value of the obstacle according to the motion trajectory information of the obstacle and the position and posture fusion information of the obstacle includes: The position fusion information of the obstacle is converted into the position fusion information in the global coordinate system according to the motion trajectory information of the obstacle, so as to obtain the position fusion information of the obstacle in the global coordinate system at the sampling time of two adjacent frames of surround view images; Calculating the difference of the position fusion information of the obstacle in the global coordinate system at the sampling time of the two adjacent frames of surround view images, and the difference of the sampling time of the two adjacent frames of surround view images; According to the calculated difference in the position fusion information and the difference in the acquisition time, the estimated speed values of the obstacle in different directions in the global coordinate system are calculated.
9. An obstacle position and speed estimation device, characterized in that: include: A collection device, used for collecting surround images around the vehicle body through a plurality of cameras installed at different positions of the vehicle body; A key point detection module is used to perform key point detection on the surround view image to obtain a target prediction frame and key point information of the obstacle; A posture solving module, used for solving the posture information of the obstacle at least according to the key point information; A posture fusion module, used to fuse the posture information of the obstacle to obtain the fusion posture information of the obstacle; A target tracking module, used to perform multi-target tracking on the obstacles according to the position and posture fusion information of the obstacles to obtain the motion trajectory information of the obstacles; A speed estimation module, used to calculate a speed estimation value of the obstacle according to the motion trajectory information of the obstacle and the position and posture fusion information of the obstacle; The position information of the obstacle includes the three-dimensional bounding box information and position information of the obstacle; The position and posture information of the obstacle is fused to obtain the position and posture fusion information of the obstacle, including: According to the position information of the obstacle, four-quadrant separation is performed on the three-dimensional bounding box information of the obstacle; For each quadrant: Grouping the three-dimensional bounding box information located in the consensus area of the camera in the quadrant into groups of two each to obtain a plurality of three-dimensional bounding box information groups; For each 3D bounding box information group, the Euclidean distance of the 3D bounding box information group is calculated according to two 3D bounding box information in the 3D bounding box information group, and it is determined whether the Euclidean distance is less than a second preset threshold value. If so, the obstacles corresponding to the 3D bounding box information group are regarded as the same obstacle, and the average value of the 3D bounding box information group is used as the 3D bounding box fusion information of the same obstacle, or the 3D bounding box information closest to the vehicle body in the 3D bounding box information group is used as the 3D bounding box fusion information of the same obstacle; the position fusion information of the same obstacle is determined according to the 3D bounding box fusion information of the same obstacle; and the posture fusion information of the same obstacle is determined according to the 3D bounding box fusion information and the position fusion information.
Citation Information
Patent Citations
Pedestrian pose resolving method for vehicle
CN110717457A