A method for predicting road user trajectories and behaviors in dense and heterogeneous traffic environments
By using deep convolutional neural networks and long-term memory codecs combined with graph regularization methods in dense heterogeneous traffic environments, the problem of difficult to predict the medium and long-term trajectory and behavior of road users in the existing technology is solved, and accurate prediction and behavior classification in complex traffic environments is achieved, and effective traffic management and autonomous driving support is provided.
Patent Information
- Application Number
- CN202411098343.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-09
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-08-09
AI Technical Summary
The prior art is difficult to accurately predict the medium- and long-term trajectory and behavior of road users in dense heterogeneous traffic environments, and it is not applicable to complex and high-density urban and rural road traffic environments.
The object detection and recognition based on deep convolutional neural network is adopted, combined with long and short-term memory codecs and regularized graphs, and a dynamic interactive graph is constructed to realize multi-pipe multi-constraint prediction of road user trajectory and behavior.
In a dense heterogeneous traffic environment, accurate prediction of road users' trajectory and behavior can be achieved, accurate results can be provided over a longer span (15~20s), and behaviors can be classified into four categories, providing guarantees for traffic safety, alleviating traffic congestion and improving traffic efficiency.
Smart Images

Figure CN119152449B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of autonomous driving (Autonomous Driving) and intelligent transportation system (ITS), and specifically relates to a method for predicting the trajectory and behavior of road users in a dense and heterogeneous traffic environment. Background Art
[0002] With the increasing popularity of cameras and computer vision technology and the substantial improvement of embedded computing power, it has become possible to track a large number of heterogeneous road users in real time from end to end. On this basis, AI technology can be used to further predict the trajectory of each road user in the future, and serve various smart transportation-related applications, such as safe and reliable advanced assisted driving or autonomous driving, transportation volume (traffic volume) forecasting (traffic forecasting), vehicle path planning, traffic congestion management (congestion management), etc. In crowded and complex traffic environments, such as urban and rural road scenes, the risk of collisions between high-density traffic heterogeneous road users (cars, two-wheeled vehicles (bicycles, electric vehicles, motorcycles), tricycles, pedestrians)) increases exponentially. In order to ensure traffic safety, alleviate traffic congestion, and improve traffic efficiency, we need to not only accurately track the trajectory positions of road users, but also accurately analyze and predict their behaviors, so as to perform path planning and collision-free navigation based on behavior perception. Related studies have shown that road user behavior significantly affects traffic flow and road user trajectory. For example, when an aggressive road user tries to overtake, it may cause the object being overtaken to slow down, or a rather conservative road user may drive slowly and cause traffic congestion.
[0003] Regarding the problem of predicting road user trajectories or behaviors, the current related technologies have the following shortcomings or deficiencies:
[0004] (1) It can only predict trajectories relatively accurately in the short term (<5s). For medium- and long-term analysis, the prediction error accumulates over time and cannot converge;
[0005] (2) It is limited to simple sparse and homogeneous road user traffic environments, such as highways in Western countries, and is not applicable to the complex and high-density urban and rural road traffic environment in my country;
[0006] (3) In terms of road user behavior classification and intention prediction, some existing research or technologies do not address this issue. Some are too simple, and simply add manually designed rules and constraints based on trajectory prediction to obtain intuitive mobile behavior judgments, such as speeding and braking. Others are too abstract and complex, with too many categories and unclear boundaries, such as impatient, timid, threatening, etc.
[0007] (4) It requires the use of multiple sensors to perceive various physical status information of road users, and even the personal attribute information of drivers. The cost is uncontrollable, the system is highly complex, and it is sensitive to noise from various sensors, making it difficult to implement. Summary of the invention
[0008] Technical problem to be solved: The present invention discloses a method for predicting the trajectory and behavior of road users in dense heterogeneous traffic environments, which can accurately classify and robustly track multiple categories of road users. Combined with the dynamic graph modeling and analysis of the interactive behaviors between road users in the neighborhood, it is particularly suitable for trajectory and behavior prediction in dense heterogeneous traffic environments, such as the trajectory and behavior prediction of multiple types of road users (cars, two-wheeled vehicles (bicycles, electric vehicles, motorcycles), tricycles, pedestrians) in crowded and complex road traffic environments in towns and villages.
[0009] Technical solution:
[0010] A method for predicting a road user trajectory and behavior in a dense heterogeneous traffic environment, the method comprising the following steps:
[0011] S1, collects video image sequences of roads in dense and heterogeneous traffic environments;
[0012] S2, performing target detection and recognition based on a deep convolutional neural network on each road user in the video image sequence frame, and using the recognition result as a road user category constraint;
[0013] S3, using the detected and identified trajectory coordinates and timestamp information of the road user as the road user's historical trajectory information;
[0014] S4, sampling and intercepting target road user m i ,i∈[1,N] is the number of timestamps before t in the historical trajectory information h The data in the time period is input into the semi-supervised pipeline in the multi-pipeline network prediction model, and the prediction timestamp t is outputted by the first long short-term memory encoder and decoder t later. f Target road users per second in the time period m i The future trajectory coordinates; at the same time, based on the input [tt h ,t] target road user m in the time period i and the historical trajectory information of all other road users within the radius r neighborhood, combined with the road user category constraint w c , construct a dynamic interaction graph of road users at each time t
[0015] S5, using singular value decomposition to obtain the road user dynamic interaction graph The corresponding feature vectors and feature values are taken to form the traffic map sequence of road users at each time t. It is input into the decision pipeline in the multi-pipeline network prediction model, and outputs the prediction timestamp t after t through the second long short-term memory encoder and decoder. f Traffic map sequence of target road users per second within the time period
[0016] Among them, the target road user traffic map sequence On the one hand, after spectral clustering, the regularized loss function is used as a graph regularization constraint to correct the road target user m output by the semi-supervised pipeline prediction. i The future trajectory coordinates, on the other hand, pass through a fully connected network and then enter a logistic regression function network to output the predicted behavior type of the road user;
[0017] S6, ends the current process, returns to step S1, and restarts the next cycle of sampling prediction analysis.
[0018] Furthermore, in step S1, a vehicle-mounted, wearable or road-based camera is used to capture the image at a frequency f. sampl Capture video image sequences.
[0019] Furthermore, the frequency f sampl =10Hz.
[0020] Step S2 further comprises:
[0021] The video image sequence frames are imported into the YOLOv5 target detection network, and features are extracted through the Darknet-53 backbone network. After upsampling and feature fusion, regression analysis is performed to obtain the target category and detection frame; the obtained detection frame is input into the SORT algorithm for target feature modeling, matching and tracking, and the road target category, detection and tracking results are output.
[0022] Step S3 further comprises:
[0023] Perform multi-target tracking of the detected and identified road users based on the appearance features learned by deep convolutional neural network, and obtain the trajectory coordinates of the road users in the two-dimensional image coordinate system in real time;
[0024] Through camera calibration and homography matrix calculation, the trajectory coordinates of the road user in the two-dimensional image coordinate system are mapped into the trajectory coordinates in the two-dimensional road plane coordinate system in the real world using perspective transformation. The trajectory coordinates and timestamp information are used as the historical trajectory information of the road user.
[0025] Furthermore, the process of mapping the trajectory coordinates of the road user in the two-dimensional image coordinate system into the trajectory coordinates in the two-dimensional road plane coordinate system in the real world by using perspective transformation includes:
[0026] Let P = (X, Y) be any point in the world coordinate system, and the camera coordinate system (x c ,y c ) is located at the center of the front end of the camera lens, p = (u, v) is the projection of the midpoint P in the world coordinate system to the image coordinate system; the origin of the image coordinate system is the upper left corner vertex of the CCD imaging plane, then:
[0027]
[0028] In the formula, It is the homography matrix, which is used to realize the projection transformation of point coordinates between two planes; is the camera’s intrinsic parameter matrix, is the camera's extrinsic matrix; f x and f y is the focal length, (c x ,c y ) is the coordinate of the principal point; r ij is the rotation matrix element, t x ,t y and t z is the translation matrix element.
[0029] Further, in step S5, the predicted behavior types of the road user include radical, reckless, safe, and conservative;
[0030] The behavioral trajectory characteristics corresponding to aggressiveness include speeding, sudden stops, cutting in, changing lanes back and forth, and following too close; the behavioral trajectory characteristics corresponding to recklessness include violations of traffic regulations such as driving in the opposite direction, illegal reversing, and jaywalking; the behavioral trajectory characteristics corresponding to safety include complying with traffic regulations, maintaining a safe distance, and driving at a constant speed; the behavioral trajectory characteristics corresponding to conservativeness include slow speed, following too far, and waiting for other users for too long.
[0031] Beneficial effects:
[0032] First, the method for predicting road user trajectories and behaviors in dense heterogeneous traffic environments of the present invention introduces a graph-based regularization constraint, which is back-propagated to the long short-term memory network used to predict road user trajectories as a correction to the loss function. As the prediction time increases, it can be continuously optimized, so that our method can obtain more accurate prediction results over a longer time span (15 to 20 seconds).
[0033] Second, the method for predicting road user trajectories and behaviors in dense heterogeneous traffic environments of the present invention uses only the predicted road user graph data as input, and innovatively divides road user behaviors into four categories, namely "aggressive", "reckless", "safe" and "conservative" through convolutional neural network learning and reasoning. The category labels are clearly defined, the coverage is sufficient, the inter-class boundaries are clear, the intra-class aggregation is strong, and the data support is sufficient.
[0034] Third, the method for predicting road user trajectories and behaviors in a dense and heterogeneous traffic environment of the present invention involves a front-end sensor based only on a statically fixed or dynamically moving RGB camera, which is low-cost under the premise of sufficient and rich perceived information; the back-end is based on an embedded edge GPU processing unit, which can process and perform reasoning and analysis on multiple video image data in real time, has low power consumption, is green and low-carbon, has controllable costs, and is suitable for on-site deployment.
[0035] Fourth, the method for predicting road user trajectories and behaviors in dense heterogeneous traffic environments of the present invention can accurately predict road user trajectories and behaviors in dense heterogeneous traffic environments, provide effective guarantees for intelligent traffic decision makers and users to ensure traffic safety, alleviate traffic congestion, and improve traffic efficiency, and also give the autonomous driving subject the ability to interact and perceive with other types of road users, thereby greatly improving the inefficient and low-comfort driving experience caused by the existing autonomous driving (advanced assisted driving) due to excessive focus on road safety. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 The method for predicting the trajectory and behavior of road users in a dense and heterogeneous traffic environment of the present invention;
[0037] Figure 2 Flowchart of multi-target detection, recognition and tracking for dense and heterogeneous road users;
[0038] Figure 3 Schematic diagram of the coordinate system relationship in the pinhole camera model;
[0039] Figure 4 A schematic diagram of mapping from two-dimensional image plane coordinates to two-dimensional road surface coordinates based on the homography matrix;
[0040] Figure 5 This is a network structure diagram of the multi-channel multi-constraint road user trajectory and behavior prediction of the present invention. DETAILED DESCRIPTION
[0041] The following examples will enable those skilled in the art to more fully understand the present invention, but are not intended to limit the present invention in any way.
[0042] The present invention discloses a method for predicting a road user trajectory and behavior in a dense heterogeneous traffic environment, the method comprising the following steps:
[0043] S1, collects video image sequences of roads in dense and heterogeneous traffic environments;
[0044] S2, performing target detection and recognition based on a deep convolutional neural network on each road user in the video image sequence frame, and using the recognition result as a road user category constraint;
[0045] S3, using the detected and identified trajectory coordinates and timestamp information of the road user as the road user's historical trajectory information;
[0046] S4, sampling and intercepting target road user m i ,i∈[1,N] is the number of timestamps before t in the historical trajectory information h The data in the time period is input into the semi-supervised pipeline in the multi-pipeline network prediction model, and the prediction timestamp t is outputted by the first long short-term memory encoder and decoder t later. f Target road users per second in the time period m i The future trajectory coordinates; at the same time, based on the input [tt h ,t] target road user m in the time period i and the historical trajectory information of all other road users within the radius r neighborhood, combined with the road user category constraint w c , construct a dynamic interaction graph of road users at each time t
[0047] S5, using singular value decomposition to obtain the road user dynamic interaction graph The corresponding feature vectors and feature values are taken to form the traffic map sequence of road users at each time t. It is input into the decision pipeline in the multi-pipeline network prediction model, and outputs the prediction timestamp t after t through the second long short-term memory encoder and decoder. f Traffic map sequence of target road users per second within the time period
[0048] Among them, the target road user traffic map sequence On the one hand, after spectral clustering, the regularized loss function is used as a graph regularization constraint to correct the road target user m output by the semi-supervised pipeline prediction. i The future trajectory coordinates, on the other hand, pass through a fully connected network and then enter a logistic regression function network to output the predicted behavior type of the road user;
[0049] S6, ends the current process, returns to step S1, and restarts the next cycle of sampling prediction analysis.
[0050] See also Figure 1 , the invention mainly comprises the following steps:
[0051] 1) Using vehicle-mounted, wearable or road-based cameras at a specific frequency (f sampl =10Hz, that is) to collect video image sequences;
[0052] 2) Performing target detection and recognition based on a deep convolutional neural network on each road user in the video image sequence frame, and passing the recognition result to step 5) as the road user category constraint input;
[0053] 3) Perform multi-target tracking of detected and identified road users based on the surface features learned by deep convolutional neural networks in real time (in Δt trk ≈100ms interval) to obtain the trajectory coordinates of the road user in the two-dimensional image coordinate system;
[0054] 4) Through camera calibration and homography matrix calculation, the trajectory coordinates of the road user in the two-dimensional image coordinate system are mapped into the trajectory coordinates in the two-dimensional road plane coordinate system in the real world using perspective transformation, expressed in longitude and latitude. The coordinate and timestamp information is passed to step 5) as input as the historical trajectory information of the road user;
[0055] 5) Construct a multi-channel multi-constraint road user trajectory and behavior prediction network. The structure diagram of the multi-channel multi-constraint road user trajectory and behavior prediction network is as follows: Figure 5 As shown, it includes a multi-pipeline network prediction model and a multi-constraint network prediction model. The specific working principle is to sample and intercept the target road user m input by step 4). i ,i∈[1,N] is the number of timestamps before t in the historical trajectory information h =Data within a 10s period (sampling accuracy, second level), through the semi-supervised pipeline in the multi-pipeline network prediction model, through the long short-term memory encoder-decoder (LSTM Encode-Decoder), output prediction timestamp t after t f = m target road users per second in a 20s period i The future trajectory coordinates; at the same time, based on the input of step 4) [tt h ,t] target road user m in the time period i All other road users (maximum number of users N) within the neighborhood of radius r = 18m (covering three lanes in both directions, each lane is 3m wide) max =50), combined with the historical trajectory information of step 2) as the road user category constraint in the multi-constraint network prediction model Construct a dynamic interactive graph (DIG) of road users at each time t, The graph is an undirected weighted graph, and its structure includes the corresponding adjacency matrix, degree matrix and Laplacian matrix, which are updated in real time according to the actual road conditions at each time t. Singular value decomposition is further used to obtain The corresponding eigenvectors and eigenvalues are taken, and the eigenvectors corresponding to the first k = 10 maximum value features are taken to form the interactive graph spectrum (IGS) sequence of road users at each time t. Through the decision pipeline in the multi-pipeline network prediction model, through the long short-term memory encoder and decoder, the prediction timestamp t is output t later f = Traffic map sequence of target road users per second within a 20s period, On the one hand, the graph sequence is used as the graph regularization constraint in the multi-constraint network prediction model in the form of a regularized loss function after spectral clustering to correct the road target user m output by the semi-supervised pipeline prediction. i The future trajectory coordinates. On the other hand, after passing through a fully connected network, it enters a logistic regression function (softmax) network, and finally outputs four types of behavior predictions of the road user, namely "aggressive", "reckless", "safe" and "conservative". It should be noted that the multi-channel network prediction model and the multi-constraint network prediction model are only for the convenience of expression. In fact, the multi-channel multi-constraint road user trajectory and behavior prediction network works as an overall network structure. Table 1 is a road user behavior classification table.
[0056] Table 1 Classification of road user behavior
[0057] Behavior Category feature radical Speeding, sudden stop, cutting in, changing lanes, following too closely, etc. reckless Violations of traffic regulations such as driving against traffic, illegal reversing, and jaywalking Safety Obey traffic rules, maintain a safe distance, keep a constant speed, etc. keep Slow speed, too far distance from other vehicles, too long waiting time for other users, etc.
[0058] 6) End the current process, return to step 1), and restart the sampling prediction analysis of the next cycle.
[0059] See also Figure 2In this embodiment, the target detection and recognition module and the multi-target tracking module are respectively implemented to correspond to the specific implementation of the above steps 2 and 3, that is, the convolutional neural network is used to detect, identify and track road users in the video image sequence. After the video image sequence is input, it first enters the YOLOv5 target detection network, and extracts features through the Darknet-53 backbone network; secondly, upsampling and feature fusion are performed, and then regression analysis is performed to obtain the target category and detection frame; thirdly, the obtained prediction frame information is input into the SORT algorithm for target feature modeling, matching and tracking; finally, the road target category, detection and tracking results are output.
[0060] Figure 3 The diagram shows the relationship between the pinhole camera coordinate system, the world coordinate system, the camera coordinate system, and the image coordinate system. In the world coordinate system (x, y, z), the z axis coincides with the camera optical axis. If the x axis and y axis are arbitrarily specified, the plane is perpendicular to the z axis. Figure 3 Where P = (X, Y, Z) is any point in the world coordinate system, and the camera coordinate system (x c ,y c ,z c ) is located at the pinhole (the center point of the front end of the camera lens), p = (u, v) is the projection of the midpoint P in the world coordinate system to the point in the image coordinate system, and the origin of the image coordinate system is the upper left corner vertex of the CCD imaging plane. Then:
[0061]
[0062] In formula (1), is the camera’s intrinsic parameter matrix, is the camera's extrinsic parameter matrix. When we need to know how the road user coordinates on the image are mapped to their corresponding real-world road surface coordinates, we need to simplify equation (1) so that Z = 0, and merge the camera's extrinsic parameter matrices into a transformation matrix as shown in equations (2) and (3):
[0063]
[0064] Right now is the homography matrix, which is used to realize the projection transformation of point coordinates between two planes, h ij is the element of the homography matrix, which is the product of the previous two matrices; f x and f y is the focal length, (c x ,c y ) is the coordinate of the principal point; r ij is the rotation matrix element, t x ,t y and t z is the translation matrix element. Figure 4 Schematic diagram of mapping from two-dimensional image plane coordinates to two-dimensional road surface coordinates based on the homography matrix.
[0065] The above are only preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should be regarded as the protection scope of the present invention.
Claims
1. A method for predicting road user trajectories and behaviors in a dense heterogeneous traffic environment, characterized in that: The road user trajectory and behavior prediction method comprises the following steps: S1, collects video image sequences of roads in dense and heterogeneous traffic environments; S2, performing target detection and recognition based on a deep convolutional neural network on each road user in the video image sequence frames, and using the recognition results as road user category constraints; S3, using the detected and identified trajectory coordinates and timestamp information of the road user as the road user's historical trajectory information; S4, sampling and intercepting target road user m i ,i∈[1,N] is the number of timestamps before t in the historical trajectory information h The data in the time period is input into the semi-supervised pipeline in the multi-pipeline network prediction model, and the prediction timestamp t is outputted by the first long short-term memory encoder and decoder t later. f Target road users per second in the time period m i The future trajectory coordinates; at the same time, based on the input [tt h ,t] target road user m in the time period i and the historical trajectory information of all other road users within the radius r neighborhood, combined with the road user category constraint w c , construct a dynamic interaction graph of road users at each time t S5, using singular value decomposition to obtain the road user dynamic interaction graph The corresponding feature vectors and feature values are taken to form the traffic map sequence of road users at each time t. It is input into the decision pipeline in the multi-pipeline network prediction model, and outputs the prediction timestamp t after t through the second long short-term memory encoder and decoder. f Traffic map sequence of target road users per second within the time period Among them, the target road user traffic map sequence On the one hand, after spectral clustering, the regularized loss function is used as a graph regularization constraint to correct the road target user m output by the semi-supervised pipeline prediction. i The future trajectory coordinates, on the other hand, pass through a fully connected network and then enter a logistic regression function network to output the predicted behavior type of the road user; S6, ending the current process, returning to step S1, and restarting the next cycle of sampling prediction analysis; Step S3 further comprises: Perform multi-target tracking of the detected and identified road users based on the appearance features learned by deep convolutional neural network, and obtain the trajectory coordinates of the road users in the two-dimensional image coordinate system in real time; Through camera calibration and homography matrix calculation, the trajectory coordinates of the road user in the two-dimensional image coordinate system are mapped into the trajectory coordinates in the two-dimensional road plane coordinate system in the real world using perspective transformation. The trajectory coordinates and timestamp information are used as the historical trajectory information of the road user.
2. The method for predicting road user trajectories and behaviors in a dense heterogeneous traffic environment according to claim 1 is characterized in that: In step S1, a vehicle-mounted, wearable or road-based camera is used to capture the image at a frequency f sampl Capture video image sequences.
3. The method for predicting road user trajectories and behaviors in a dense heterogeneous traffic environment according to claim 2 is characterized in that: Frequency f sampl =10Hz.
4. The method for predicting road user trajectories and behaviors in a dense heterogeneous traffic environment according to claim 1, characterized in that: Step S2 further comprises: The video image sequence frames are imported into the YOLOv5 target detection network, and features are extracted through the Darknet-53 backbone network. After upsampling and feature fusion, regression analysis is performed to obtain the target category and detection frame; the obtained detection frame is input into the SORT algorithm for target feature modeling, matching and tracking, and the road target category, detection and tracking results are output.
5. The method for predicting road user trajectories and behaviors in a dense heterogeneous traffic environment according to claim 1, characterized in that: The process of mapping the trajectory coordinates of the road user in the two-dimensional image coordinate system into the trajectory coordinates in the two-dimensional road plane coordinate system in the real world by using perspective transformation includes: Let P = (X, Y) be any point in the world coordinate system, and the camera coordinate system (x c ,y c ) is located at the center of the front end of the camera lens, p = (u, v) is the projection of the midpoint P in the world coordinate system to the image coordinate system; the origin of the image coordinate system is the upper left corner vertex of the CCD imaging plane, then: In the formula, It is the homography matrix, which is used to realize the projection transformation of point coordinates between two planes; is the camera’s intrinsic parameter matrix, is the camera's extrinsic matrix; f x and f y is the focal length, (c x ,c y ) is the coordinate of the principal point; r ij is the rotation matrix element, t x ,t y and t z is the translation matrix element.
6. The method for predicting road user trajectories and behaviors in a dense heterogeneous traffic environment according to claim 1, characterized in that: In step S5, the predicted behavior types of the road user include radical, reckless, safe, and conservative; The behavioral trajectory characteristics corresponding to aggressiveness include speeding, sudden stops, cutting in, changing lanes back and forth, and following too close; the behavioral trajectory characteristics corresponding to recklessness include violations of traffic regulations such as driving in the opposite direction, illegal reversing, and jaywalking; the behavioral trajectory characteristics corresponding to safety include complying with traffic regulations, maintaining a safe distance, and driving at a constant speed; the behavioral trajectory characteristics corresponding to conservativeness include slow speed, following too far, and waiting for other users for too long.
Citation Information
Patent Citations
System and method for point-to-point traffic prediction
CN112470199A
Vehicle future trajectory prediction method based on graph neural network
CN115147790A