Vehicle driving track analysis method based on deep learning
By using the improved YOLOv3 model and SORT algorithm to detect and track target objects in vehicle driving trajectory analysis, combining LSTM neural network and multi-dimensional calibration method to generate driver models suitable for different regions, solving the problems of low detection accuracy and poor tracking accuracy in complex traffic scenarios, achieving higher detection accuracy and model universality, and improving the safety and reliability of the autonomous driving system.
Patent Information
- Application Number
- CN202510333766.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-20
AI Technical Summary
The existing vehicle driving trajectory analysis methods have problems such as low detection accuracy, poor tracking accuracy and difficulty in identifying complex driving behaviors in complex traffic scenarios, and have not fully considered socio-cultural differences, resulting in limited universality and accuracy of the model.
The vehicle driving trajectory analysis method based on deep learning is adopted, and the target object detection is carried out through the improved YOLOv3 model combined with Mask R-CNN, and the target object tracking is performed using SORT algorithm and Kalman filtering. The lane-changing behavior characteristics are extracted from the trajectory data, and a driver model adapted to different regions is generated based on the LSTM neural network and multi-dimensional calibration method.
It improves the detection accuracy of small target objects, enhances the accuracy and stability of trajectory tracking, can have a deeper understanding of driving behavior in complex traffic scenarios, generates a more versatile and accurate driver model, and improves the safety and reliability of the autonomous driving system.
Smart Images

Figure CN120182330A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of autonomous driving, and particularly relates to a method for analyzing vehicle driving trajectories based on deep learning. Background Art
[0002] With the acceleration of the urbanization process and the continuous increase in the number of automobiles, problems such as traffic congestion and traffic accidents have become increasingly serious. As an effective means to solve these problems, autonomous driving technology has received extensive attention and research. In the research and development process of autonomous driving technology, the analysis of vehicle driving trajectories is crucial, which can provide key decision-making basis for the autonomous driving system, help the system better understand the traffic scene, predict vehicle behavior, and thus improve the safety and reliability of autonomous driving.
[0003] Currently, although there are some methods for analyzing vehicle driving trajectories, in complex traffic scenarios, these methods still have many deficiencies. For example, in terms of object detection, the detection accuracy of small objects is relatively low, and there are easy cases of missed detection or false detection; in terms of trajectory tracking, for fast-moving or occluded objects, the accuracy and stability of tracking are relatively poor; in terms of extracting and analyzing driving behavior characteristics, it is difficult to accurately identify complex driving behaviors, such as lane-changing behaviors at multi-lane intersections. In addition, existing driving behavior models often do not fully consider the influence of social and cultural differences on driving behaviors, resulting in limitations in the generality and accuracy of the models. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for analyzing vehicle driving trajectories based on deep learning. The present invention can accurately extract vehicle driving trajectories, analyze driving behavior characteristics in complex traffic scenarios, and generate driver models adapted to different regions, providing reliable data support and model basis for the virtual verification of the autonomous driving system, thereby improving the safety and reliability of the autonomous driving system.
[0005] The technical solution of the present invention: A method for analyzing vehicle driving trajectories based on deep learning includes the following steps:
[0006] S1. Collect traffic video data of the target area through multiple road cameras and synchronously record GNSS timestamps;
[0007] S2. Use a pre-trained deep neural network to detect objects in the video data, identify the categories of vehicles, pedestrians, and non-motor vehicles, and generate bounding boxes of the objects;
[0008] S3. Adopt an online real-time tracking algorithm to continuously track the objects and obtain the movement trajectories of the objects in the video sequence;
[0009] S4. Convert the image coordinates to global three-dimensional coordinates through the camera calibration model and store them associated with the GNSS timestamp;
[0010] S5. Extract the lane-changing behavior characteristics in complex traffic scenarios from the trajectory data, including the starting point, ending point of lane change, and the change of lateral acceleration;
[0011] S6. Conduct clustering analysis on the lane-changing behavior based on the LSTM neural network to generate the driver model parameters of different driving styles;
[0012] S7. Integrate the driver model parameters into the microscopic traffic flow simulation tool to generate a dynamic traffic scenario for virtual verification of the autonomous driving system.
[0013] In the above vehicle driving trajectory analysis method based on deep learning, in step S2, in step S2, the deep neural network adopts an improved YOLOv3 model, combines Mask R-CNN for object pixel-level segmentation to improve the detection accuracy of small objects;
[0014] The improved YOLOv3 model includes:
[0015] Adopt a bidirectional feature pyramid structure, achieve multi-scale feature fusion through 3 cascaded feature fusion layers. For each layer, first upsample the high-level feature map by 2 times, then splice it with the low-level feature map in the channel dimension, and finally process it through a 3×3 convolutional layer. The expression is:
[0016]
[0017] In the formula: P i is the feature map of the i-th layer, represents splicing in the channel dimension;
[0018] The Mask R-CNN adopts a parallel three-branch structure, including a classification branch with a Softmax classifier, a regression branch with an L1 loss function, and a mask branch that generates a binary mask using a fully convolutional network.
[0019] In the aforementioned vehicle driving trajectory analysis method based on deep learning, in step S3, the tracking algorithm is the SORT algorithm, and the Kalman filter is fused to predict and correct the motion state of the object.
[0020] In the aforementioned vehicle driving trajectory analysis method based on deep learning, in step S4, the specific process of converting the image coordinates to global three-dimensional coordinates through the camera calibration model and storing them associated with the GNSS timestamp is:
[0021] Convert from the image coordinate system to the camera coordinate system:
[0022]
[0023] Where: (u, v) are the image pixel coordinates, and Z is the z-axis coordinate of the target point in the camera coordinate system; the camera internal parameter matrix K is obtained by the Zhang Zhengyou calibration method:
[0024]
[0025] Where: f x , f y respectively represent the focal length coordinates (in pixel units), and (u0, v0) are the coordinates of the image principal point.
[0026] Converting from the camera coordinate system to the global three-dimensional coordinates:
[0027] P w = R T ·(P c - T);
[0028] Where: P w is the three-dimensional point in the global three-dimensional coordinates, and P c is the three-dimensional point in the camera coordinate system; R is the rotation matrix; T is the translation vector; R T is the inverse of the rotation matrix;
[0029] The GNSS timestamp associated storage includes hardware synchronization and time synchronization. The hardware synchronization is ensured by the NTP protocol or the hardware trigger signal, and the clock deviation between all cameras and the GNSS receiver is <1 ms. The video frame capture moment is accurately recorded as tcam; the time synchronization is to store the three-dimensional coordinate P w and the corresponding timestamp tcam as structured data:
[0030]
[0031] For the aforementioned vehicle driving trajectory analysis method based on deep learning, the method for extracting the lane-changing behavior features described in step S5 includes:
[0032] S5.1 Calculate the rate of change of the distance between the vehicle and the adjacent lane markings, and set a threshold to determine the starting point of the lane change;
[0033] S5.2 Analyze the relative speed and distance of the surrounding vehicles based on the Wiedemann psychophysical model to judge the type of lane-changing intention.
[0034] For the aforementioned vehicle driving trajectory analysis method based on deep learning, in step S6, a multi-dimensional calibration method is adopted, and combined with the driver's age, gender, and sociocultural difference parameters, a driver model adapted to different regions is generated.
[0035] In the above-mentioned vehicle driving trajectory analysis method based on deep learning, in step S7, the microscopic traffic flow simulation tool is a joint simulation framework of PTV Vissim and GaiA software, which supports the automated generation and testing of dynamic traffic scenarios.
[0036] In the above-mentioned vehicle driving trajectory analysis method based on deep learning, in step 1, the method further includes optimizing the error detection of static and dynamic targets, specifically: by fusing data from multiple camera perspectives and using the optical flow method to compensate for the detection error caused by illumination changes.
[0037] Compared with the prior art, the present invention uses an improved YOLOv3 model combined with Mask R-CNN for target detection, effectively improving the detection accuracy of small targets, reducing the cases of missed detection and false detection, and providing a more accurate data basis for subsequent trajectory analysis. The SORT algorithm of the present invention fuses Kalman filtering to track targets, improving the accuracy and stability of tracking, and being able to better handle the situations of rapid movement and occlusion of targets, ensuring the continuity and reliability of the trajectory. The present invention extracts rich lane-changing behavior features from trajectory data and combines with the Wiedemann psychophysical model to judge the type of lane-changing intention, being able to more deeply understand the driver's behavior pattern and providing a more accurate basis for the generation of the driver model. The present invention generates driver models adapted to different regions based on the LSTM neural network and multi-dimensional calibration method, fully considering the individual differences and sociocultural factors of drivers, making the model more general and accurate, and being able to better simulate real driving behaviors. The present invention integrates driver model parameters into the joint simulation framework of PTV Vissim and GaiA software for virtual verification, providing an efficient and low-cost platform for the development and testing of autonomous driving systems, and helping to accelerate the development and application of autonomous driving technologies. Description of the Drawings
[0038] Figure 1 It is an actual case diagram of the global three-dimensional coordinates of the present invention;
[0039] Figure 2 It shows a schematic diagram of processing the starting and ending points of lane changes for the trajectory data of the present invention;
[0040] Figure 3 It shows a schematic diagram of the distribution of trajectory data according to the Wiedemann model.
[0041] Figure 4 It shows a schematic diagram of constructing a microscopic traffic flow model of the present invention. Detailed Embodiments
[0043] The present invention will be further described below in conjunction with embodiments, but it shall not be used as a basis for limiting the present invention.
[0044] Embodiment: A driver behavior recognition method for autonomous driving virtual testing and verification, comprising the following steps:
[0045] S1. Collect traffic video data of the target area through multiple road cameras and synchronously record GNSS timestamps; in this step, multiple road cameras are set at the highway entrance to ensure that the traffic conditions of the entire entrance area can be covered.
[0046] S2. Use a pre-trained deep neural network to detect target objects in the video data, identify the categories of vehicles, pedestrians and non-motor vehicles, and generate bounding boxes of the target objects; in this step, the pre-trained improved YOLOv3 model combined with MaskR-CNN is used to detect target objects in the collected video data. The network structure of the improved YOLOv3 model is optimized by adding a feature fusion layer, enabling the model to more effectively extract the features of small target objects. For example, for smaller target objects such as motorcycles and bicycles on the highway, the improved model can accurately identify their categories and generate precise bounding boxes. Mask R-CNN further performs pixel-level segmentation on the target objects to improve the detection accuracy and avoid misclassifying small target objects as other objects.
[0047] In this step, the improved YOLOv3 model includes
[0048] Adopt a bidirectional feature pyramid structure, and achieve multi-scale feature fusion through 3 cascaded feature fusion layers. For each layer, first upsample the high-level feature map by 2 times, then splice it with the low-level feature map in the channel dimension, and finally process it through a 3×3 convolutional layer. The expression is:
[0049]
[0050] In the formula: P i is the i-th layer feature map, represents splicing in the channel dimension.
[0051] Use the Swish activation function Swish(x) = x·σ(x) to replace LeakyReLU, where
[0052] Adopt the CIoU loss function to optimize the position prediction:
[0053]
[0054] In the formula: ρ is the Euclidean distance of the center point, c is the diagonal length of the smallest bounding box, v is the width-height ratio consistency parameter, For small target objects (size S < 100 pixels), through automatically boost the loss weight;
[0055] The Mask R-CNN adopts a parallel three-branch structure, including a classification branch of a Softmax classifier, a regression branch of an L1 loss function, and a mask branch that generates a binary mask using a fully convolutional network.
[0056] The RoIAlign operation is used to replace RoIPooling:
[0057] where n is the number of sampling points, solving the quantization error problem;
[0058] The mask loss adopts pixel-level binary cross-entropy:
[0059]
[0060] where m is the total number of mask pixels, P i is the predicted probability, and G i is the true label.
[0061] In this embodiment, the fusion method is as follows:
[0062] Construct a detection-segmentation pipeline: input image → YOLOv3 detection → candidate box screening → Mask R-CNN segmentation → merge results
[0063] The multi-task joint training loss function is:
[0064] L total = L YOLO + λ mask L mask where λ mask is the mask loss weight, and its value range is 0.5 - 2.0.
[0065] S3. An online real-time tracking algorithm is used to continuously track the target object to obtain the motion trajectory of the target object in the video sequence; in this step, the SORT algorithm is used to fuse the Kalman filter to continuously track the detected target object. In the highway scenario, the speed of vehicles is relatively fast, and some vehicles may suddenly accelerate, decelerate, or change lanes. The SORT algorithm performs data association based on the appearance features and motion information of the target object, and can quickly and accurately track the target object. The Kalman filter predicts the next moment state of the vehicle according to the current motion state and motion model of the vehicle, and corrects the prediction result by combining the measurement data. For example, when the vehicle changes lanes, the Kalman filter can timely adjust the tracking parameters to ensure the accurate tracking of the vehicle trajectory, improving the accuracy and stability of the tracking.
[0066] In this step, the Kalman filter is used to perform a prior prediction on the target motion state (position, velocity, etc.) to generate a predicted bounding box. This process is based on a uniform motion model and estimates the position of the target in the next frame through the state transition matrix and covariance matrix. The SORT algorithm matches the predicted bounding box with the detection bounding box in the current frame through the Hungarian algorithm. The matching criteria include the intersection over union (IoU) and appearance features (such as the cosine similarity of feature vectors). For the successfully matched targets, the state is updated. The unmatched predicted bounding boxes are regarded as lost tracks, and the unmatched detection bounding boxes are regarded as new targets. For the successfully matched targets, the detection bounding box is used to correct the posterior state of the Kalman filter, update the position and velocity estimates, and at the same time adjust the covariance matrix to reflect the uncertainty of the current state.
[0067] S4. Convert the image coordinates to global three-dimensional coordinates through the camera calibration model and store them associated with the GNSS timestamp; in this step, in the highway scenario, accurate coordinate conversion is crucial for subsequent trajectory analysis. The camera calibration model has been accurately calibrated and can accurately convert the coordinates of the target object in the image to global three-dimensional coordinates, enabling the data collected by different cameras to be analyzed and processed in a unified coordinate system.
[0068] In this step, convert from the image coordinate system to the camera coordinate system:
[0069]
[0070] In the formula: (u, v) are the image pixel coordinates, and Z is the z-axis coordinate of the target point in the camera coordinate system; the camera internal parameter matrix K is obtained through the Zhang Zhengyou calibration method:
[0071]
[0072] In the formula: f x , f y respectively represent the focal length coordinates (in pixel units), and (u0, v0) are the image principal point coordinates,
[0073] Convert from the camera coordinate system to the global three-dimensional coordinates:
[0074] P w = R T ·(P c - T);
[0075] In the formula: P w is the three-dimensional point in the global three-dimensional coordinates, and P c is the three-dimensional point in the camera coordinate system; R is the rotation matrix R (a 3×3 orthogonal matrix); T is the translation vector (a 3×1 vector); R T is the inverse of the rotation matrix; Figure 1 This is the actual case diagram of the global three-dimensional coordinates of the present invention;
[0076] In this embodiment, GNSS timestamp synchronization includes hardware synchronization and time synchronization. The hardware synchronization is ensured through the NTP protocol or a hardware trigger signal, such that the clock deviation between all cameras and the GNSS receiver is < 1 ms, and the video frame capture time is accurately recorded as tcam. The time synchronization is to store the three-dimensional coordinates P w and the corresponding timestamp tcam as structured data:
[0077]
[0078] S5. Extract the lane-changing behavior features in complex traffic scenarios from the trajectory data, including the lane-changing start point, end point, and lateral acceleration change; in this step, the extraction method of the lane-changing behavior features includes:
[0079] S5.1 Calculate the rate of change of the distance between the vehicle and the adjacent lane markings, and set a threshold to determine the lane-changing start point; in this step, first, use a deep learning model (such as DeepLabv3+) to segment the lane line pixels, convert them into three-dimensional coordinates in the world coordinate system, and perform a quadratic polynomial fitting on the three-dimensional coordinate points of the lane line:
[0080] y = ax 2 + bx + c;
[0081] where: (x, y) is the lateral position in the world coordinate system.
[0082] Then obtain the vehicle centroid coordinates (x v , y v ), and calculate the heading angle based on the difference of the vehicle trajectory:
[0083]
[0084] where: △x = x t - x t-1 , △ y = y t - y t-1 ;
[0085] Project the vehicle centroid (x v , y v ) onto the lane line model, and calculate the vertical distance d:
[0086]
[0087] Calculate the rate of change:
[0088]
[0089] where: Δt is the video frame interval;
[0090] Use a 3-frame moving average to eliminate noise:
[0091]
[0092] In this embodiment, the threshold is set according to the empirical threshold method. Through experimental statistics of the range, the threshold δ th is set to 0.3 - 0.5 m / s. When three consecutive frames satisfy and the lateral acceleration is greater than a th (0.5 m / s 2 ), it is determined that the lane change starting point begins.
[0093] Illustrative example:
[0094] Suppose the lateral distance change of a certain vehicle within 3 seconds is as follows:
[0095]
[0096] In the above table, if δ th = 0.4 m / s, when t = 0.099 s, the absolute value exceeds the threshold, and the lateral acceleration a = -68.18 m / s, then it is determined that the lane change starting point begins. Figure 2 Shows a schematic diagram of processing the lane change starting point and ending point of the trajectory data of the present invention.
[0097] S5.2 Analyze the relative speed and distance of surrounding vehicles based on the Wiedemann psychophysical model to determine the type of lane change intention.
[0098] In this step, the Wiedemann model is mainly used to describe the decision-making process of the driver during lane change, considering the relative speed and distance of surrounding vehicles, as well as the driver's psychological factors. The core of the Wiedemann model is to calculate the safe distance and relative speed of the driver. The safe distance usually consists of two parts: the reaction distance at the current vehicle speed and the braking distance. The reaction distance is the distance traveled by the driver during the reaction time, and the braking distance is the distance from the start of braking to the stop of the vehicle. The sum of these two parts constitutes the safe distance. Then, the relative speed refers to the speed difference between the vehicle in the target lane and the own vehicle. If the relative speed is positive, it means the vehicle in the target lane is faster than the own vehicle; otherwise, it is slower. Combining the safe distance and relative speed, the risk and type of lane change intention can be judged. For example, if the relative speed of surrounding vehicles is small and the distance is far, the lane change intention of the vehicle may be for overtaking; if the relative speed of surrounding vehicles is large and the distance is close, the lane change intention of the vehicle may be to avoid collision or find a more suitable driving lane. Figure 3 Shows a schematic diagram of the distribution of trajectory data according to the Wiedemann model.
[0099] S6. Cluster analysis of lane-changing behavior is performed based on the LSTM neural network to generate driver model parameters for different driving styles. In this step, cluster analysis of lane-changing behavior is performed based on the LSTM neural network, and combined with driver age, gender, and sociocultural difference parameters, driver model parameters suitable for highway scenarios are generated. In highway scenarios, there are significant differences in driving styles among different drivers. Some young drivers may be more inclined to an aggressive driving style, frequently changing lanes to pursue faster driving speeds; while some elderly drivers may be more inclined to a conservative driving style, being more cautious when changing lanes. The LSTM neural network can effectively perform cluster analysis on different driving styles by learning a large amount of lane-changing behavior data. Combined with the driver's age, gender, and sociocultural difference parameters, the generated driver model can more accurately reflect the driving behavior characteristics of different drivers in highway scenarios.
[0100] S7. Integrate the driver model parameters into the microscopic traffic flow simulation tool to generate a dynamic traffic scenario for virtual verification of the autonomous driving system. In this step, the generated driver model parameters are integrated into the joint simulation framework of PTV Vissim and GaiA software. In the joint simulation framework, PTV Vissim constructs a microscopic traffic flow model according to the driver model parameters to simulate behaviors such as vehicle following and lane-changing at highway entrances. The GaiA software provides a visual simulation environment to display the dynamic changes of the traffic scenario. Through virtual verification, the response ability of the autonomous driving system to different driving behaviors in highway scenarios can be evaluated, such as the decision-making and control ability of autonomous driving vehicles when encountering other vehicles changing lanes, providing a basis for the optimization of the autonomous driving system. Figure 4 Shows a schematic diagram of the construction of the microscopic traffic flow model of the present invention.
[0101] In practical applications, the error detection of static and dynamic targets can also be optimized. By fusing data from multiple camera perspectives, the optical flow method is used to compensate for the detection errors caused by changes in illumination. The optical flow method detects the movement of the target by calculating the motion vectors of pixel points in the image. When the illumination changes, the optical flow method can adjust the detection results according to the motion information of the pixel points, reducing the impact of illumination changes on target detection and improving the detection accuracy.
[0102] Furthermore, corresponding tests are carried out on the solution of the present invention:
[0103] 1. Target detection and segmentation test data;
[0104] Test dataset: KITTI Vision Benchmark Dataset (including 7481 training images and 7518 test images);
[0105] Hardware configuration: NVIDIA RTX 3090 GPU, Intel i9-12900K CPU;
[0106]
[0107] Table 1
[0108] As can be seen from Table 1, the improved YOLOv3 uses a bidirectional feature pyramid and the Swish activation function, and the small object detection rate is increased by 9.6%. After integrating Mask R-CNN, the instance segmentation mAP reaches 86.4%.
[0109] 2. Target tracking test data;
[0110] Test scenarios:
[0111] Scenario 1: Highway (vehicle speed 80 km / h, occlusion rate 10%)
[0112] Scenario 2: Urban intersection (vehicle speed 40 km / h, occlusion rate 30%)
[0113]
[0114] Table 2
[0115] As can be seen from Table 2, SORT + Kalman filtering improves the tracking accuracy of urban intersections in occluded scenarios by 20.3%, and the average tracking error is reduced from 12.4 pixels to 5.2 pixels.
[0116] 3. Widemann model intention classification effect;
[0117] Test data set:
[0118] Chinese drivers (n = 5000 lane changes)
[0119] German drivers (n = 3000 lane changes)
[0120] Indian drivers (n = 4000 lane changes)
[0121] Intention type Accuracy of Chinese classification Accuracy of German classification Accuracy of Indian classification Safe overtaking 91.2% 94.5% 88.7% Normal lane change 85.6% 89.3% 82.1% Emergency avoidance 95.8% 97.2% 93.4%
[0122] Table 3
[0123] As can be seen from Table 3, after the cultural parameters are corrected, the cross-border scenario classification accuracy is improved by 7 - 9%. In the data of Indian drivers, due to the high proportion (18%) of the intention of "dangerous operation", the overall accuracy is slightly lower than that of other regions.
[0124] 4. Simulation verification performance;
[0125] Test metrics:
[0126] Decision latency (ms) of the autonomous driving system in the virtual scenario;
[0127] Number of emergency brakes (times / hour);
[0128]
[0129] Table 4
[0130] As can be seen from Table 4, after integrating the driver model, the simulation scenario generation speed increased by 67%. The number of emergency brakes of the autonomous driving system decreased by 33%, verifying the effectiveness of the model.
[0131] Description of the experimental process:
[0132] Data collection:
[0133] Deploy road cameras in Hangzhou, China; Munich, Germany; and Delhi, India to collect 1,200 hours of video data.
[0134] Synchronously record vehicle CAN bus data (speed, acceleration), GNSS (accuracy 0.05 m), and IMU (100 Hz).
[0135] Model training:
[0136] The object detection model is pre-trained on the KITTI dataset and then fine-tuned on the local dataset for 100 epochs.
[0137] The LSTM clustering model inputs 1 million lane-changing trajectories (including 120-dimensional features) and is trained until the loss < 0.05.
[0138] Field verification:
[0139] Conduct 2,000 km of on-road vehicle tests on the Hangzhou Ring Expressway and urban arterial roads to compare the consistency between the model output and manual annotation.
[0140] Simulation test:
[0141] Generate 1,000 virtual scenarios, including extreme conditions such as rainy, foggy weather and construction sections, to test the robustness of the autonomous driving system.
[0142] Data source:
[0143] Road video data: Measured by the Zhejiang Asia-Pacific Smart Connected Vehicle Innovation Center;
[0144] Simulation tool: PTV Vissim 2023 + GaiA 4.2;
[0145] Hardware platform: NVIDIA DRIVE Sim + ROS2 Humble;
[0146] The above data verifies the technical feasibility of each step and provides a quantitative basis for the development of the autonomous driving system.
[0147] The above embodiments are only partial implementation manners of the present invention, and can be adjusted and extended according to specific requirements in actual applications. The protection scope of the present invention is not limited to the above embodiments, but also includes various deformations and improvements based on the technical solutions of the present invention.
Claims
1. A vehicle driving trajectory analysis method based on deep learning, characterized in that: The following steps are involved: S1. Collect traffic video data of the target area through multiple road cameras and synchronously record GNSS timestamps; S2. Use the pre-trained deep neural network to detect objects in the video data, identify the categories of vehicles, pedestrians and non-motor vehicles, and generate the bounding box of the object; S3. Use an online real-time tracking algorithm to continuously track the target and obtain the motion trajectory of the target in the video sequence; S4. Convert the image coordinates into global three-dimensional coordinates through the camera calibration model and store them in association with the GNSS timestamp; S5. Extract lane-changing behavior characteristics in complex traffic scenarios from trajectory data; S6. Perform cluster analysis on lane changing behaviors based on LSTM neural network and generate driver model parameters with different driving styles; S7. Integrate the driver model parameters into the microscopic traffic flow simulation tool to generate dynamic traffic scenarios for virtual verification of the autonomous driving system.
2. The vehicle driving trajectory analysis method based on deep learning according to claim 1 is characterized in that: In step S2, the deep neural network in step S2 uses an improved YOLOv3 model and combines Mask R-CNN to perform pixel-level segmentation of the target object to improve the detection accuracy of small targets: The improved YOLOv3 model includes: A bidirectional feature pyramid structure is adopted to realize multi-scale feature fusion through three cascaded feature fusion layers. Each layer first upsamples the high-level feature map by 2 times, then concatenates it with the low-level feature map in the channel dimension, and finally processes it through a 3×3 convolution layer. The expression is: Where: P i is the feature map of the i-th layer, ⊕ represents channel dimension concatenation; The MaskR-CNN adopts a parallel three-branch structure, including a classification branch of a Softmax classifier, a regression branch of an L1 loss function, and a mask branch that generates a binary mask using a fully convolutional network.
3. The vehicle driving trajectory analysis method based on deep learning according to claim 1 is characterized in that: In step S3, the tracking algorithm is a SORT algorithm, and Kalman filtering is integrated to predict and correct the motion state of the target object.
4. The vehicle driving trajectory analysis method based on deep learning according to claim 1 is characterized in that: In step S4, the image coordinates are converted into global three-dimensional coordinates through the camera calibration model, and the specific process of storing them in association with the GNSS timestamp is: Convert from image coordinate system to camera coordinate system: Where: (u, v) is the image pixel coordinate, Z is the z-axis coordinate of the target point in the camera coordinate system; the camera intrinsic parameter matrix K is obtained by Zhang Zhengyou calibration method: Where: f x , f y They represent the focal length coordinates (pixel units), (u0, v0) are the coordinates of the principal point of the image, Convert from camera coordinate system to global 3D coordinate system: P w =R T ·(P c -T); Where: P w is a 3D point in global 3D coordinates, P c is a three-dimensional point in the camera coordinate system; R is the rotation matrix; T is the translation vector; R T is the inverse of the rotation matrix; The GNSS timestamp associated storage includes hardware synchronization and time synchronization. Hardware synchronization is ensured by NTP protocol or hardware trigger signal. The clock deviation between all cameras and GNSS receivers is less than 1ms. The video frame capture time is accurately recorded as tcam. Time synchronization is to convert the three-dimensional coordinate P w The corresponding timestamp tcam is stored as structured data:
5. The vehicle driving trajectory analysis method based on deep learning according to claim 4 is characterized in that: In step S5, the method for extracting lane change behavior features includes: S5.1 calculates the rate of change of the distance between the vehicle and the adjacent lane markings and sets a threshold to determine the starting point of lane change; S5.2 analyzes the relative speed and distance of surrounding vehicles based on the Wiedemann psychophysical model to determine the type of lane change intention.
6. The vehicle driving trajectory analysis method based on deep learning according to claim 1, characterized in that: In step S6, a multi-dimensional calibration method is used to combine the driver's age, gender and social and cultural difference parameters to generate a driver model suitable for different regions.
7. The vehicle driving trajectory analysis method based on deep learning according to claim 1 is characterized in that: In step S7, the microscopic traffic flow simulation tool is a joint simulation framework of PTV Vissim and GaiA software, which supports the automatic generation and testing of dynamic traffic scenarios.
8. The vehicle driving trajectory analysis method based on deep learning according to claim 7 is characterized in that: In step 1, the method also includes optimizing the false detection of static and dynamic targets, specifically by fusing data from multiple camera perspectives and using the optical flow method to compensate for detection errors caused by illumination changes.
Citation Information
Cited By
Highway auxiliary driving method
CN122368104A