Method, device and equipment for identifying vehicle driving style at intersection without signal lights
By performing dimensionality reduction and cluster analysis on vehicle driving data at intersections without signal lights, combined with the AdaBoost model, the problems of insufficient accuracy and reliability in existing technologies for identifying vehicle driving styles at intersections without signal lights are solved, and efficient and safe driving style recognition is achieved.
Patent Information
- Application Number
- CN202510232724.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-02-28
AI Technical Summary
Existing technologies lack lane-vehicle characteristics analysis in vehicle driving style recognition at intersections without signal lights, resulting in insufficient accuracy and reliability of recognition results.
The preset kernel PCA algorithm is used to reduce the dimensionality of driving data features, and the K-means algorithm and the preset GMM model are combined to perform cluster analysis on the reduced dimensionality principal component vectors. Finally, the AdaBoost model is used for driving style recognition training and testing.
It achieves accurate and reliable identification of vehicle driving styles at intersections without signal lights, improving the effectiveness, safety and efficiency of driving decisions.
Smart Images

Figure CN120003500B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of autonomous driving technology, and in particular to a method, device, and apparatus for identifying the driving style of a vehicle at an intersection without signal lights. Background Art
[0002] With the rapid development of autonomous driving technology, how to enable autonomous driving systems to interact more efficiently and safely with human drivers and other traffic participants in complex road environments has become a critical research topic. In this context, driving style recognition, as a key technology, plays a vital role in the decision-making, planning, and control of downstream technical modules in autonomous vehicles, and has therefore attracted widespread attention from academia and industry. It analyzes and processes existing data, such as acceleration and steering, to identify the driver's individual driving style and leverage this information to enhance the decision-making capabilities and safety of the autonomous driving system.
[0003] The methods for realizing driving style recognition mainly include three aspects: 1) Using traditional machine learning methods, using support vector machines, random forests and other algorithms, to identify through extracted feature parameters; 2) Using deep learning methods, using neural networks to automatically extract features from raw data for style recognition; 3) Using clustering and unsupervised learning methods to classify the processed data, and perform unsupervised recognition based on this classification model.
[0004] Current driving style recognition is mostly used in lane-changing scenarios. However, intersections without traffic lights present a complex and challenging scenario, and their detailed characteristics have not been fully considered and analyzed. Furthermore, existing driving style recognition technologies focus on relatively simple lane and vehicle data features, resulting in a lack of accuracy and reliability in the recognition results. Consequently, driving decisions based on this lack effectiveness, safety, and efficiency. Summary of the Invention
[0005] The present application provides a method, device and equipment for identifying vehicle driving style at intersections without signal lights, which is used to solve the technical problems that the existing technology lacks analysis of lane vehicle characteristics at intersections without signal lights, and considers the data features to be relatively simple, resulting in a lack of accuracy and reliability in the recognition results.
[0006] In view of this, the first aspect of the present application provides a method for identifying a vehicle driving style at an intersection without a signal light, comprising:
[0007] Acquiring driving data features of vehicles passing through an intersection without a signal light, wherein the driving data features include straight-ahead data features and turn data features;
[0008] Using a preset kernel PCA algorithm to perform dimensionality reduction processing on the driving data features to obtain a reduced-dimensionality principal component vector, wherein the kernel function of the preset kernel PCA algorithm adopts a multinomial kernel function;
[0009] Based on the K-means algorithm, a preset GMM model is used to perform cluster analysis on the reduced-dimensionality principal component vector to obtain a cluster analysis result, wherein the initial value of the K-means algorithm is obtained by screening based on the annealing algorithm and the repulsive potential field algorithm;
[0010] Performing driving style recognition training and testing on an initial AdaBoost model based on the driving data characteristics and the cluster analysis results to obtain a target AdaBoost model;
[0011] The target AdaBoost model is used to identify the driving style of vehicles at an actual intersection without signal lights, and a style recognition result is obtained.
[0012] Preferably, the acquisition of driving data features of vehicles passing through an intersection without a signal light, wherein the driving data features include straight-ahead data features and turn data features, includes:
[0013] Extracting initial driving data from the trajectory dataset of vehicles passing through an unsignaled intersection;
[0014] Using an SG filter to filter the initial driving data to obtain filtered driving data;
[0015] The driving characteristics of the passing vehicles are calculated according to the filtered driving data, and the driving characteristics are sorted into binary features according to straight-moving and turning behaviors to obtain driving data features.
[0016] Preferably, the K-means algorithm is based on a preset GMM model to perform cluster analysis on the reduced-dimensional principal component vector to obtain a cluster analysis result, including:
[0017] The K-means algorithm is used to perform iterative clustering analysis with the preset initial cluster center as the initial value to obtain the initial clustering results;
[0018] Based on the expectation-maximization optimization algorithm, a preset GMM model is used to perform cluster analysis on the reduced-dimensional principal component vector according to the initial clustering result to obtain a cluster analysis result.
[0019] Preferably, the method further comprises: performing cluster analysis on the dimension-reduced principal component vectors based on a K-means algorithm and using a preset GMM model to obtain cluster analysis results;
[0020] An annealing algorithm is used to perform a sphere center search analysis on the data of the dimension-reduced principal component vector to obtain an annealing cluster center;
[0021] A distance screening operation is performed on the annealing cluster centers using a repulsive potential field algorithm to obtain preset initial cluster centers.
[0022] Preferably, the driving style recognition training and testing of the initial AdaBoost model based on the driving data characteristics and the cluster analysis results to obtain the target AdaBoost model includes:
[0023] Performing cluster verification on the cluster analysis results based on a gap statistics method, and if the verification passes, using the cluster analysis results as data labels for the driving data features, and generating a driving data set in combination with the driving data features;
[0024] Generate the initial AdaBoost model by constructing weak classifiers and strong classifiers;
[0025] The driving data set is used to perform driving style recognition training and testing on the initial AdaBoost model to obtain a target AdaBoost model.
[0026] A second aspect of the present application provides a device for identifying a vehicle driving style at an intersection without a signal light, comprising:
[0027] A data acquisition unit, configured to acquire driving data features of vehicles passing through an intersection without a signal light, wherein the driving data features include straight-ahead data features and turn data features;
[0028] a dimensionality reduction processing unit, configured to perform dimensionality reduction processing on the driving data features using a preset kernel PCA algorithm to obtain a reduced-dimensionality principal component vector, wherein the kernel function of the preset kernel PCA algorithm adopts a multinomial kernel function;
[0029] A cluster analysis unit is used to perform cluster analysis on the reduced-dimensional principal component vector based on a K-means algorithm and a preset GMM model to obtain a cluster analysis result, wherein the initial value of the K-means algorithm is obtained by screening based on an annealing algorithm and a repulsive potential field algorithm;
[0030] a model training unit, configured to perform driving style recognition training and testing on an initial AdaBoost model based on the driving data characteristics and the cluster analysis results, to obtain a target AdaBoost model;
[0031] The style recognition unit is used to use the target AdaBoost model to recognize the driving style of vehicles at an actual intersection without signal lights, and obtain a style recognition result.
[0032] Preferably, the data acquisition unit is specifically used to:
[0033] Extracting initial driving data from the trajectory dataset of vehicles passing through an unsignaled intersection;
[0034] Using an SG filter to filter the initial driving data to obtain filtered driving data;
[0035] The driving characteristics of the passing vehicles are calculated according to the filtered driving data, and the driving characteristics are sorted into binary features according to straight-moving and turning behaviors to obtain driving data features.
[0036] Preferably, the cluster analysis unit is specifically used to:
[0037] The K-means algorithm is used to perform iterative clustering analysis with the preset initial cluster center as the initial value to obtain the initial clustering results;
[0038] Based on the expectation-maximization optimization algorithm, a preset GMM model is used to perform cluster analysis on the reduced-dimensional principal component vector according to the initial clustering result to obtain a cluster analysis result.
[0039] Preferably, it also includes:
[0040] An annealing search unit is used to perform a sphere center search analysis on the data of the dimension-reduced principal component vector using an annealing algorithm to obtain an annealing cluster center;
[0041] The distance screening unit is used to perform a distance screening operation on the annealing cluster centers by using a repulsive potential field algorithm to obtain a preset initial cluster center.
[0042] A third aspect of the present application provides a device for identifying a vehicle driving style at an intersection without signal lights, the device comprising a processor and a memory;
[0043] The memory is used to store program code and transmit the program code to the processor;
[0044] The processor is configured to execute the method for identifying a vehicle driving style at a non-signal intersection according to the instructions in the program code.
[0045] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:
[0046] In the present application, a method for identifying the driving style of vehicles at an intersection without signal lights is provided, comprising: obtaining driving data features of vehicles passing through the intersection without signal lights, the driving data features including straight data features and turn data features; using a preset kernel PCA algorithm to perform dimensionality reduction processing on the driving data features to obtain a reduced-dimensionality principal component vector, wherein the kernel function of the preset kernel PCA algorithm uses a multinomial kernel function; based on the K-means algorithm, a preset GMM model is used to perform cluster analysis on the reduced-dimensionality principal component vector to obtain a cluster analysis result, wherein the initial value of the K-means algorithm is obtained by screening based on an annealing algorithm and a repulsive potential field algorithm; driving style recognition training and testing of an initial AdaBoost model based on the driving data features and the cluster analysis results to obtain a target AdaBoost model; and using the target AdaBoost model to identify the driving style of vehicles at an actual intersection without signal lights to obtain a style recognition result.
[0047] The method for identifying vehicle driving styles at intersections without signal lights provided in this application analyzes and identifies the driving styles of vehicles at intersections without signal lights, fully studies the driving data characteristics of intersections without signal lights, and uses a preset kernel PCA algorithm to reduce the dimensionality of the data characteristics, which can reduce the computational complexity; in addition, cluster analysis is performed on the unlabeled driving data characteristics through multiple algorithms, thereby achieving unsupervised training of subsequent classifiers; this process can ensure the reliability of the data characteristic analysis and the accuracy of the cluster analysis, so the target AdaBoost model obtained can accurately and reliably identify the driving styles of vehicles at actual intersections without signal lights, meeting the needs of practical applications. Therefore, this application can solve the technical problems that the existing technology lacks analysis of the characteristics of vehicles in lanes at intersections without signal lights, and the recognition results lack accuracy and reliability due to the relatively simple data characteristics. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 A flowchart of a method for identifying a vehicle driving style at an intersection without signal lights provided in an embodiment of the present application;
[0049] Figure 2 A schematic diagram of the structure of a device for identifying a vehicle driving style at an intersection without signal lights provided in an embodiment of the present application;
[0050] Figure 3 A schematic diagram of the internal structure of high-dimensional data visualized using t-SNE provided in an embodiment of the present application;
[0051] Figure 4 A schematic diagram of the columnar results of principal component analysis using the existing PCA algorithm provided in an embodiment of the present application;
[0052] Figure 5The explained variance histogram and curve graph of the dimensionality reduction process based on the preset kernel PCA algorithm provided in the embodiment of the present application;
[0053] Figure 6 A schematic diagram showing the relationship between data features before and after dimensionality reduction based on a preset kernel PCA algorithm provided in an embodiment of the present application;
[0054] Figure 7 A top view of the kernel principal component analysis effect provided in an embodiment of the present application;
[0055] Figure 8 A schematic diagram of the results of processing different three-dimensional distribution data using the cluster analysis algorithm provided in an embodiment of the present application;
[0056] Figure 9 An example diagram of clustering results of cluster analysis of dimensionality-reduced principal component vectors provided in an embodiment of the present application;
[0057] Figure 10 A silhouette coefficient histogram for evaluating cluster analysis results provided in an embodiment of the present application;
[0058] Figure 11 A graph showing the verification of cluster analysis results based on the gap statistics method provided in an embodiment of the present application;
[0059] Figure 12 A confusion matrix diagram of the improved decision tree classification process provided in an embodiment of the present application;
[0060] Figure 13 AUG-ROC curve diagram of the improved decision tree classification process provided in the embodiment of the present application;
[0061] Figure 14 A learning curve diagram of the improved decision tree classification process provided in an embodiment of the present application. DETAILED DESCRIPTION
[0062] In order to help those skilled in the art better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.
[0063] For easier understanding, see Figure 1 The present application provides an embodiment of a method for identifying a vehicle driving style at a non-signaled intersection, including:
[0064] Step 101: Acquire driving data features of vehicles passing through an intersection without a signal light, where the driving data features include straight-ahead data features and turning data features.
[0065] Furthermore, step 101 includes:
[0066] Extracting initial driving data from the trajectory dataset of vehicles passing through an unsignaled intersection;
[0067] The initial driving data is filtered using an SG filter to obtain filtered driving data;
[0068] The driving characteristics of the passing vehicles are calculated based on the filtered driving data, and the driving characteristics are sorted into binary features according to the straight-moving and turning behaviors to obtain the driving data features.
[0069] Since the acquired driving data features are used for subsequent cluster analysis and model training, the trajectory dataset of vehicles passing through the unsignaled intersection selected in this embodiment can be an existing vehicle driving dataset, such as the NGSIM dataset. This dataset is an open source dataset that includes vehicle trajectory data for two highways and two main road sections, among which Peachtree Street is a main road with multiple intersections.
[0070] The data selected from the Peachtree Street main road in this example is partial data from a single intersection on that road section. Specifically, this example only selects and studies vehicle-related data at the intersection. Furthermore, because existing data is collected from main roads and arterials, lacking data on vehicles traveling at intersections without traffic lights, this example removes data from intersections with no vehicles or during periods of low traffic volume, retaining scenes with multiple vehicles simultaneously traveling at the intersection to approximate the state of an intersection without traffic lights. The initial driving data extracted from the dataset includes, but is not limited to, vehicle identification number, longitudinal coordinates of the vehicle's front center, vehicle length, global coordinates, vehicle type, instantaneous speed, instantaneous acceleration, current lane position, and vehicle location in the starting and destination zones.
[0071] Since the existing data set is affected by various external factors, the data has varying degrees of noise. In this embodiment, an SG filter is used to filter and smooth the selected initial driving data. The SG filter can maintain the characteristics of the data and is particularly suitable for noise suppression and smoothing of time series data. The specific filtering process is not described here.
[0072] The driving characteristics of passing vehicles can be calculated based on the obtained filtered driving data. The driving characteristics are also expressed in the form of characteristic parameters. Please refer to Table 1 for details. In order to facilitate subsequent data processing, this embodiment uses binary features as indicators to distinguish the driving characteristics of vehicles going straight and turning, that is, 1 or 0 is used to represent whether the vehicle is turning or going straight. The difference between the two in the driving data characteristics is that the angle-related characteristic parameters of straight-moving vehicles are 0, while those of turning vehicles are different. Using binary identifiers to divide straight-moving and turning vehicles can avoid the bias in the results after dimensionality reduction due to the lack of turning-related characteristic parameters for straight-moving vehicles, and can also help the model extract effective information, so that the driving styles of straight-moving and turning vehicles can be identified in a set of characteristic parameters.
[0073] Table 1 Example of driving data characteristics
[0074]
[0075] It's understandable that both straight-travel and turn-travel data features share the same data dimension. For example, Table 1 shows 11-dimensional data, with some parameters for straight-travel set to 0. This means that the straight-travel and turn-travel data features share the same data format, allowing for undifferentiated subsequent data processing.
[0076] Step 102: Use a preset kernel PCA algorithm to perform dimensionality reduction processing on the driving data features to obtain a reduced-dimensionality principal component vector. The kernel function of the preset kernel PCA algorithm uses a multi-kernel function.
[0077] It should be noted that the PCA algorithm projects high-dimensional data into a low-dimensional space by finding a new coordinate system that maps the data to the direction of maximum variance, preserving the data's key features to the greatest extent possible. The main steps include data normalization, calculating the covariance matrix, calculating the eigenvalues and eigenvectors of the covariance matrix, selecting principal components, and projecting the data.
[0078] However, in this embodiment, it is necessary to consider the relevant characteristic parameters of both straight-moving vehicles and turning vehicles at the same time, so the data features show a strong nonlinear relationship, and it is still difficult to achieve ideal results using the existing PCA algorithm. Figure 3 and Figure 4 ,t-SNE is used to visualize the distribution of the 11 extracted driving data features in two-dimensional space. Figure 3 There are multiple scattered clusters in the data, and there is no obvious linear distribution feature between these clusters. The clusters have complex shapes and the distances between them are uneven, which indicates that the data has complex nonlinear relationships in high-dimensional space. Figure 4The PCA algorithm is used to obtain the components corresponding to the number of data features, so as to analyze the relationship between the various data in the original data. By analyzing the explained variance of each component, it can be concluded that the contribution of adjacent components to the total variance of the data is very small, that is, most components contain a certain amount of information, which further illustrates the strong nonlinearity of the original data.
[0079] Therefore, this embodiment proposes a preset kernel PCA algorithm, which uses a multinomial kernel function as its kernel function; this algorithm is a PCA algorithm with a kernel function. The preset kernel PCA algorithm uses the kernel technique to map data into a high-dimensional feature space, then performs standard PCA operations in this high-dimensional space, thereby capturing nonlinear structural characteristics.
[0080] Assume that the original driving data features are , then the process of mapping data to high-dimensional feature space can be expressed as:
[0081]
[0082] in, is a nonlinear mapping function, Represents a high-dimensional feature space.
[0083] Since D in the high-dimensional feature space is very large, the mapped data features cannot be directly calculated. In this embodiment, the kernel function is used to calculate the inner product between point pairs and construct the kernel matrix , each element in the matrix is expressed as:
[0084]
[0085] in, Represents the kernel function, which is specifically expressed as:
[0086]
[0087] in, is a constant term, is the polynomial order.
[0088] To ensure that the core matrix It is centralized in the feature space, so the kernel matrix needs to be centralized. The kernel matrix after centralized processing is expressed as:
[0089]
[0090]
[0091] The kernel matrix after centralization Perform eigenvalue decomposition:
[0092]
[0093] In the preset kernel PCA algorithm, the principal component is the data projection of the feature space. For each sample The component projected on the jth principal component is:
[0094]
[0095] in, is the jth eigenvector, is the corresponding eigenvalue. The value mapped to the jth principal component is:
[0096]
[0097] The subsequent operations are the same as those of the existing PCA algorithm, except that the eigenvalues of the preset kernel PCA algorithm of this embodiment do not directly represent the variance in the original space, but reflect the data characteristics of the high-dimensional kernel space.
[0098] See also Figure 5 The preset kernel PCA algorithm of this embodiment is used to reduce the dimensionality of driving data features. The difference between the various features in the figure is large, and it can be seen from the curve that the first three features already contain nearly 85% of the information of the original data. Therefore, these three principal components can be extracted to represent the characteristic parameters of the original data, that is, the dimensionality reduction principal component vector includes three principal component representative vectors. In other words, each data is reduced from 11 dimensions to 3 dimensions, which can reduce the complexity of data calculation and improve data processing efficiency.
[0099] It should also be noted that the preset kernel PCA algorithm of this embodiment can also use Kendall's Tau to help understand the relationship between features before and after dimensionality reduction. Kendall's Tau correlation coefficient is a non-parametric correlation measure that is suitable for measuring the order consistency of two variables in nonlinear and ordered data. Its value range is [-1, 1]. The closer it is to 1, the stronger the positive correlation between the two variables; the closer it is to -1, the stronger the negative correlation between the two variables. Please refer to Figure 6 ,Analysis of the principal components with the absolute value of Kendall's Tau coefficient exceeding 0.4 in each set of characteristic parameters shows a strong negative correlation except for the standard deviation of angular velocity and angular acceleration. For the standard deviation of angular velocity and angular acceleration, the second principal component is analyzed, such as Figure 7 Some data of category 0 correspond to a higher second principal component.
[0100] Step 103: Based on the K-means algorithm, a preset GMM model is used to perform cluster analysis on the reduced-dimensional principal component vector to obtain cluster analysis results. The initial value of the K-means algorithm is obtained by screening based on the annealing algorithm and the repulsive potential field algorithm.
[0101] Furthermore, step 103 includes:
[0102] The K-means algorithm is used to perform iterative clustering analysis with the preset initial cluster center as the initial value to obtain the initial clustering results;
[0103] Based on the expectation maximization optimization algorithm, the preset GMM model is used to perform cluster analysis on the reduced dimensionality principal component vector according to the initial clustering results to obtain the cluster analysis results.
[0104] Furthermore, before step 103, the following steps are also included:
[0105] The annealing algorithm is used to perform sphere center search analysis on the data of the reduced-dimensional principal component vector to obtain the annealing cluster center;
[0106] The distance screening operation of the annealing cluster center is performed through the repulsive potential field algorithm to obtain the preset initial cluster center.
[0107] It should be noted that the cluster center obtained by simple iteration of the K-means algorithm can be used as the initial center of the preset GMM model, that is, the mean; and accurate cluster analysis results can be obtained by using the preset GMM model for cluster analysis.
[0108] In order to deal with problems such as complex data distribution models, different cluster shapes, and results that are sensitive to the selected values of the initial cluster centers, this embodiment optimizes the initial values of the K-means algorithm, that is, a specific algorithm is used to select and optimize the initial values of the clustering algorithm, thereby ensuring the accuracy of the initial cluster centers; this method can solve the problem of the K-means algorithm being sensitive to initial values, and can further combine the strong adaptability of the K-means algorithm with the flexibility of GMM to ensure the reliability of the clustering analysis results.
[0109] This embodiment uses an annealing algorithm and a repulsive potential field algorithm to optimize the initial values of the K-means algorithm, i.e., the initial cluster centers. From the initial data, n random seeds are selected, exceeding the target number of clusters m. With each seed as the center, a hypersphere with a radius of r is assigned, and the number of data points contained within each hypersphere is calculated. For example, in this embodiment, the target number of clusters m is 3, indicating the driving styles of moderate, average, and aggressive. Specifically, for m=3, a hypersphere with a radius of r is assigned. A spherical search is performed, with a point d units of length from the center in a random direction as the next center. The number of data points within the hypersphere is calculated, with the goal of maximizing the number of data points within each sphere.
[0110] During the iteration of the annealing algorithm, a result with a small number of preset small probabilities can be used to update the number of results to avoid local optimal solutions; moreover, the Metropolis criterion can also be used to determine the target solution. In addition, this embodiment adds two conditions for accepting the solution during the annealing algorithm solution process. First, the average distance between each point in the sphere and the center of the sphere is calculated. If it is greater than the median of the distance, it can be used as an accepted candidate solution. The purpose of this condition is to make the updated points in the sphere as close to the center of the sphere as possible to obtain a better annealing clustering center. Second, when the radius of the sphere is half of the original, whether the amount of data in the small sphere is 1 / 5 of the original sphere data. When the radius is reduced by half, the volume ratio of the large sphere to the small sphere is 8:1. In order to obtain better results and balance the calculation time, 1 / 8 is changed to 1 / 5.
[0111] The repulsive potential field algorithm can be used to more accurately screen the annealing cluster centers obtained by the annealing algorithm. The screening criterion is the distance between the centers. The farther the distance between the centers, the better the clustering effect; otherwise, the worse. In order to obtain cluster centers with a larger distance, this embodiment uses the repulsive potential field algorithm to perform distance screening on the annealing cluster centers to obtain the target number of clusters m. The minimum repulsive force indicates that the distance between the seeds is the largest. Therefore, the repulsive potential field function is defined as:
[0112]
[0113] in, is the repulsive potential field function, is the repulsive force strength coefficient, is the distance between the two seeds, n is the potential energy decay rate exponent, And it is an integer.
[0114] The repulsive force is:
[0115]
[0116] In order to avoid the distance between the two centers being much larger than the distance of other points relative to these two points, the above-mentioned nth power inverse potential energy function is used, which makes the repulsive force increase exponentially as the distance shortens.
[0117] The seed of the final target number of clusters m is used as the initial value of the K-means algorithm, that is, the initial cluster center. After a simple iterative clustering analysis, the preset GMM model initial value can be obtained. That is, the initial clustering result is used as the initial value, and the preset GMM model can be used to perform cluster analysis on the reduced-dimensional principal component vector. During the clustering analysis process, the log-likelihood estimation algorithm is used based on the expectation-maximization algorithm to solve and obtain the clustering analysis result. The theoretical process will not be described in detail here and can be implemented based on existing technology.
[0118] See also Figure 8The algorithm provided in this embodiment has good results in processing three-dimensional data with different distributions. For specific performance index data comparison analysis, please refer to Figure 2. Specifically, after clustering analysis of the three principal component vectors selected above, that is, the three dimensionality reduction principal component vectors, the following can be obtained: Figure 9 As shown in the results, we can find that the data is clearly divided into three categories. Since GMM does not need to pre-specify the number of clusters, the processing results are consistent with the three categories we expect, which reflects the high performance and accuracy of the model. Figure 7 and Figure 9 Initial categories 1, 2, and 3 can be marked as aggressive, normal, and mild driving styles, respectively.
[0119] Table 2 Performance indicators of cluster analysis for different data in this embodiment
[0120]
[0121] Step 104 : Perform driving style recognition training and testing on the initial AdaBoost model based on the driving data characteristics and cluster analysis results to obtain a target AdaBoost model.
[0122] Furthermore, step 104 includes:
[0123] The cluster analysis results are clustered and verified based on the gap statistics method. If the verification is passed, the cluster analysis results are used as data labels for the driving data features, and the driving data set is generated by combining the driving data features.
[0124] Generate the initial AdaBoost model by constructing weak classifiers and strong classifiers;
[0125] The driving data set is used to train and test the initial AdaBoost model for driving style recognition, and the target AdaBoost model is obtained.
[0126] It should be noted that since the original driving data does not have corresponding labels, it is necessary to use intrinsic evaluation indicators to evaluate the cluster analysis results. Figure 10 , except for a small amount of data in the range of [-0.1, 0], the other data are generally concentrated in the positive range, and some reach 0.5. At this time, it can be considered that the clustering effect is good. In addition, the gap statistical method can be used to verify the number of clusters, see Figure 11It can be seen that the gap statistics for clusters with a number of 3 are the largest, indicating that a clear cluster structure exists in the data. This is consistent with the actual cluster analysis results of this embodiment and also reflects that the algorithm of this embodiment has a strong ability to capture the cluster structure in the data, thus passing the verification. Other evaluation or verification methods can also be designed based on actual circumstances to evaluate the accuracy and reliability of cluster analysis results. This embodiment is provided only as an example and is not intended to be limiting.
[0127] The cluster analysis results are used as data labels for the original driving data features, and combined with the driving data features to generate a driving dataset; the driving dataset can be divided into a training dataset and a test dataset, and then used to train the initial AdaBoost model to improve the model's style recognition ability. The obtained target AdaBoost model can be directly used in actual scenarios to realize vehicle driving style recognition.
[0128] The AdaBoost model is an ensemble learning model that constructs a strong classifier by weightedly combining multiple weak classifiers. The specific model construction and execution process involves setting a multi-classification problem, constructing weak classifiers, updating weights, constructing a strong classifier, and defining an exponential loss function. To handle multi-classification tasks, this example uses the SAMME.R algorithm, a multi-classification variant of AdaBoost. Compared to the more common SAMME algorithm, the SAMME.R algorithm offers faster convergence and better performance.
[0129] In addition, in order to avoid overfitting of the model and obtain better recognition effect when processing complex data, this embodiment changes the common decision tree stump in the weak classifier to three decision trees with a depth of 2 and a depth of 1 and 2 respectively. After training, it is found that this improvement can obtain better recognition effect. For specific performance analysis, please refer to Figure 12 and Figure 13 , the evaluation indicators are shown in Table 3. The training set score of this classification is 0.9687, and the average score of cross validation is 0.9212. Figure 14 From the learning curve, we can see that as the number of samples increases, the scores of the training set and the validation set gradually converge, eventually reaching 0.95 and 0.925 respectively. The gap is small, indicating that the model has a good fitting state and good generalization ability.
[0130] Table 3 Evaluation indicators
[0131]
[0132] Step 105: Use the target AdaBoost model to identify the driving style of vehicles at an actual intersection without a signal light, and obtain a style recognition result.
[0133] The method for identifying the driving style of vehicles at intersections without signal lights provided in the embodiment of the present application performs characteristic analysis and identification on the driving style of vehicles at intersections without signal lights, fully studies the driving data characteristics of intersections without signal lights, and uses a preset kernel PCA algorithm to perform dimensionality reduction processing on the data characteristics, which can reduce the computational complexity; in addition, cluster analysis is performed on the unlabeled driving data characteristics through multiple algorithms, thereby achieving unsupervised training of subsequent classifiers; this process can ensure the reliability of the data characteristic analysis and the accuracy of the cluster analysis, so the target AdaBoost model obtained can accurately and reliably identify the driving style of vehicles at actual intersections without signal lights, meeting the needs of practical applications. Therefore, the embodiment of the present application can solve the technical problems that the existing technology lacks analysis of the characteristics of vehicles in lanes at intersections without signal lights, and the recognition results lack accuracy and reliability due to the relatively simple data characteristics.
[0134] For easier understanding, see Figure 2 The present application provides an embodiment of a device for identifying a vehicle driving style at an intersection without a signal light, comprising:
[0135] The data acquisition unit 201 is used to acquire driving data features of vehicles passing through the intersection without signal lights, wherein the driving data features include straight-line data features and turn data features;
[0136] A dimensionality reduction processing unit 202 is configured to perform dimensionality reduction processing on the driving data features using a preset kernel PCA algorithm to obtain a reduced-dimensionality principal component vector, wherein the kernel function of the preset kernel PCA algorithm adopts a multinomial kernel function;
[0137] A cluster analysis unit 203 is configured to perform cluster analysis on the reduced-dimensional principal component vectors using a preset GMM model based on a K-means algorithm to obtain cluster analysis results. The initial value of the K-means algorithm is obtained by screening based on an annealing algorithm and a repulsive potential field algorithm.
[0138] A model training unit 204 is used to train and test the initial AdaBoost model for driving style recognition based on the driving data characteristics and cluster analysis results to obtain a target AdaBoost model;
[0139] The style recognition unit 205 is configured to recognize the driving style of vehicles at an actual intersection without a signal light using a target AdaBoost model to obtain a style recognition result.
[0140] Furthermore, the data acquisition unit 201 is specifically configured to:
[0141] Extracting initial driving data from the trajectory dataset of vehicles passing through an unsignaled intersection;
[0142] The initial driving data is filtered using an SG filter to obtain filtered driving data;
[0143] The driving characteristics of the passing vehicles are calculated based on the filtered driving data, and the driving characteristics are sorted into binary features according to the straight-moving and turning behaviors to obtain the driving data features.
[0144] Furthermore, the cluster analysis unit 203 is specifically configured to:
[0145] The K-means algorithm is used to perform iterative clustering analysis with the preset initial cluster center as the initial value to obtain the initial clustering results;
[0146] Based on the expectation maximization optimization algorithm, the preset GMM model is used to perform cluster analysis on the reduced dimensionality principal component vector according to the initial clustering results to obtain the cluster analysis results.
[0147] Furthermore, it also includes:
[0148] Annealing search unit 206, used to perform sphere center search analysis on the data of the reduced-dimensional principal component vector using an annealing algorithm to obtain an annealing cluster center;
[0149] The distance screening unit 207 is configured to perform a distance screening operation on the annealing cluster centers by using a repulsive potential field algorithm to obtain preset initial cluster centers.
[0150] The present application also provides a device for identifying a vehicle driving style at an intersection without a signal light, the device including a processor and a memory;
[0151] The memory is used to store program codes and transmit the program codes to the processor;
[0152] The processor is configured to execute the method for identifying the driving style of a vehicle at an intersection without signal lights in the above embodiment according to the instructions in the program code.
[0153] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0154] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0155] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0156] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for executing all or part of the steps of the method described in each embodiment of the present application through a computer device (which can be a personal computer, server, or network device, etc.). The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (full name: Read-Only Memory, English abbreviation: ROM), random access memory (full name: Random Access Memory, English abbreviation: RAM), disk or optical disk, and other media that can store program code.
[0157] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for identifying vehicle driving styles at an intersection without signal lights, characterized in that: include: Acquiring driving data features of vehicles passing through an intersection without a signal light, wherein the driving data features include straight-ahead data features and turn data features; Using a preset kernel PCA algorithm to perform dimensionality reduction processing on the driving data features to obtain a reduced-dimensionality principal component vector, wherein the kernel function of the preset kernel PCA algorithm adopts a multinomial kernel function; An annealing algorithm is used to perform a sphere center search analysis on the data of the dimension-reduced principal component vector to obtain an annealing cluster center; Performing a distance screening operation on the annealing cluster centers using a repulsive potential field algorithm to obtain a preset initial cluster center; Based on the K-means algorithm, a preset GMM model is used to perform cluster analysis on the reduced-dimensionality principal component vector to obtain a cluster analysis result, wherein the initial value of the K-means algorithm is obtained by screening based on the annealing algorithm and the repulsive potential field algorithm; Performing driving style recognition training and testing on an initial AdaBoost model based on the driving data characteristics and the cluster analysis results to obtain a target AdaBoost model; The target AdaBoost model is used to identify the driving style of vehicles at an actual intersection without signal lights, and a style recognition result is obtained.
2. The method for identifying vehicle driving style at a non-signal intersection according to claim 1, characterized in that: The obtaining of driving data features of vehicles passing through an intersection without a signal light, wherein the driving data features include straight-line data features and turn data features, includes: Extracting initial driving data from the trajectory dataset of vehicles passing through an unsignaled intersection; Using an SG filter to filter the initial driving data to obtain filtered driving data; The driving characteristics of the passing vehicles are calculated according to the filtered driving data, and the driving characteristics are sorted into binary features according to straight-moving and turning behaviors to obtain driving data features.
3. The method for identifying vehicle driving style at a non-signal intersection according to claim 1, characterized in that: The K-means algorithm is used to perform cluster analysis on the dimension-reduced principal component vector using a preset GMM model to obtain cluster analysis results, including: The K-means algorithm is used to perform iterative clustering analysis with the preset initial cluster center as the initial value to obtain the initial clustering results; Based on the expectation-maximization optimization algorithm, a preset GMM model is used to perform cluster analysis on the reduced-dimensional principal component vector according to the initial clustering result to obtain a cluster analysis result.
4. The method for identifying vehicle driving style at a non-signal intersection according to claim 1, characterized in that: The driving style recognition training and testing of the initial AdaBoost model based on the driving data characteristics and the cluster analysis results to obtain a target AdaBoost model includes: Performing cluster verification on the cluster analysis results based on a gap statistics method, and if the verification passes, using the cluster analysis results as data labels for the driving data features, and generating a driving data set in combination with the driving data features; Generate the initial AdaBoost model by constructing weak classifiers and strong classifiers; The driving data set is used to perform driving style recognition training and testing on the initial AdaBoost model to obtain a target AdaBoost model.
5. A vehicle driving style recognition device for a non-signal intersection, characterized in that: include: A data acquisition unit, configured to acquire driving data features of vehicles passing through an intersection without a signal light, wherein the driving data features include straight-ahead data features and turn data features; a dimensionality reduction processing unit, configured to perform dimensionality reduction processing on the driving data features using a preset kernel PCA algorithm to obtain a reduced-dimensionality principal component vector, wherein the kernel function of the preset kernel PCA algorithm adopts a multinomial kernel function; An annealing search unit is used to perform a sphere center search analysis on the data of the dimension-reduced principal component vector using an annealing algorithm to obtain an annealing cluster center; A distance screening unit, configured to perform a distance screening operation on the annealing cluster centers using a repulsive potential field algorithm to obtain a preset initial cluster center; A cluster analysis unit is used to perform cluster analysis on the reduced-dimensional principal component vector based on a K-means algorithm and a preset GMM model to obtain a cluster analysis result, wherein the initial value of the K-means algorithm is obtained by screening based on an annealing algorithm and a repulsive potential field algorithm; a model training unit, configured to perform driving style recognition training and testing on an initial AdaBoost model based on the driving data characteristics and the cluster analysis results, to obtain a target AdaBoost model; The style recognition unit is used to use the target AdaBoost model to recognize the driving style of vehicles at an actual intersection without signal lights, and obtain a style recognition result.
6. The device for identifying vehicle driving style at a non-signal intersection according to claim 5, characterized in that: The data acquisition unit is specifically used to: Extracting initial driving data from the trajectory dataset of vehicles passing through an unsignaled intersection; Using an SG filter to filter the initial driving data to obtain filtered driving data; The driving characteristics of the passing vehicles are calculated according to the filtered driving data, and the driving characteristics are sorted into binary features according to straight-moving and turning behaviors to obtain driving data features.
7. The device for identifying vehicle driving style at a non-signal intersection according to claim 5, characterized in that: The cluster analysis unit is specifically used for: The K-means algorithm is used to perform iterative clustering analysis with the preset initial cluster center as the initial value to obtain the initial clustering results; Based on the expectation-maximization optimization algorithm, a preset GMM model is used to perform cluster analysis on the reduced-dimensional principal component vector according to the initial clustering result to obtain a cluster analysis result.
8. A device for identifying vehicle driving styles at intersections without signal lights, characterized in that: The device includes a processor and a memory; The memory is used to store program code and transmit the program code to the processor; The processor is configured to execute the method for identifying a vehicle driving style at a non-signal intersection according to any one of claims 1 to 4 according to the instructions in the program code.
Citation Information
Patent Citations
Traffic sign recognition method based on extreme learning machine and self-adaptive lifting
CN105631477A
Heterogeneous vehicle problem CVRP determination method based on improved K-means clustering ant colony algorithm and cloud platform distribution system thereof
CN116341783A