Driving behavior analysis method based on federal K-means clustering

Through the combination of federal K-mean clustering and automatic encoder, the problems of privacy leakage, large computing resource consumption and insufficient feature representation in traditional driving behavior analysis are solved, and efficient and accurate driving behavior analysis is achieved, adapting to complex traffic scenarios and meeting the real-time processing needs of edge devices.

CN120452182AActive Publication Date: 2025-08-08CHONGQING UNIV OF POSTS & TELECOMM
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510498333.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-08-08
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

Traditional driving behavior analysis methods have the risk of privacy leakage, high computing resources consumption, insufficient feature representation ability, and compatibility problems with dynamic traffic scenarios and high-dimensional feature embedding with federal training, resulting in reduced model accuracy and high misclassification rate.

Method used

The method based on federal K-mean clustering is adopted, and spatiotemporal features are extracted through networked vehicles and roadside units, and the dimensionality reduction is used to reduce the dimensionality by using an automatic encoder. The local clustering center is updated with probability weighting. Server aggregation generates a global center to realize distributed iterative optimization, and AES encrypted transmission is used to adaptively adjust the classification standards.

Benefits of technology

On the premise of protecting data privacy, the accuracy of driving behavior classification is improved, the calculation and communication costs are reduced, and the dynamic traffic scenarios are adapted to the feature dimension utilization rate and model adaptability are improved, and the real-time processing needs of edge devices are met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120452182A_ABST
    Figure CN120452182A_ABST
Patent Text Reader

Abstract

The invention relates to a driving behavior analysis method based on federal K-means clustering, and belongs to the field of machine learning and traffic safety. In order to solve the problems of privacy leakage risk, low high-dimensional data processing efficiency and insufficient driving behavior classification precision in the traditional method, the technical scheme comprises the following steps: selecting a networked vehicle / RSU as a client, fusing a high-precision map to extract spatio-temporal characteristics, and performing dimension reduction through an automatic encoder to obtain a driving behavior classification result; distributed iterative optimization is realized through federal K-means clustering (a local clustering center is updated by probability weighting, and a global center is generated by server aggregation), and finally driving behaviors are classified according to dynamic thresholds such as speed change and steering rate. According to the method, on the premise of protecting data privacy, the classification accuracy reaches 92.7%, the calculation efficiency is improved by 3.2 times, the communication cost is reduced by 82%, and the traffic safety management level is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of machine learning and traffic safety, and relates to a driving behavior analysis method based on federal K-means clustering. Background Art

[0002] With the rapid development of connected and autonomous vehicles (CAVs) and vehicle-road collaboration technologies, driving behavior analysis has become a core technical means to improve road traffic safety and optimize autonomous driving decisions. Traditional driving behavior analysis methods mainly rely on centralized data processing architectures, collecting vehicle trajectory, speed, acceleration, and other data through on-board sensors or roadside equipment, and uploading them to a central server for unified modeling and analysis. However, this model faces multiple challenges in the actual implementation of intelligent transportation systems, specifically in the following aspects:

[0003] (1) Centralized processing requires vehicles to upload raw driving data (such as location, driving path, and driving habits) to a third-party server, which poses a risk of sensitive information leakage. According to statistics, global traffic data leakage incidents caused by centralized data storage increased by 37% year-on-year in 2022, involving private information such as vehicle trajectory and driver identity. Although technologies such as differential privacy and homomorphic encryption have been introduced, these methods often sacrifice data utility, resulting in a 20%-30% decrease in model accuracy.

[0004] (2) In intelligent transportation scenarios, roadside units (RSUs), edge computing nodes, and connected vehicles generate over 1PB of heterogeneous data (including high-precision maps, millimeter-wave radar point clouds, video streams, etc.) every day. Traditional K-means clustering algorithms face the problems of high computing resource consumption and high communication latency when processing such data in a centralized architecture. Studies have shown that when the data size exceeds 10^6 samples, the clustering time of traditional methods increases exponentially and cannot adapt to the limited computing power of edge devices.

[0005] (3) Driving behavior analysis is highly dependent on deep feature mining of multi-source heterogeneous data. Existing methods are usually based on manually designed feature engineering (such as average speed and number of sudden brakes), ignoring the dynamic relationship between vehicle trajectories and high-precision maps (such as lane offset and intersection turning trajectory), resulting in insufficient feature characterization capabilities. In addition, factors such as on-board sensor noise (GPS positioning error ±2 meters) and communication packet loss (packet loss rate >15%) will introduce a large number of outliers, making it difficult for existing preprocessing methods (such as mean filling) to effectively preserve data distribution characteristics. Experiments show that noise interference can reduce clustering accuracy by more than 40%.

[0006] (4) Existing classification models, such as those based on support vector machines (SVMs) or decision trees, use fixed thresholds to classify driving behavior types (e.g., aggressive / conservative) and are unable to adapt to the dynamic changes in complex traffic scenarios. For example, in rainy, snowy, or congested road conditions, the same driver’s steering rate, following distance, and other characteristics may significantly deviate from the preset thresholds, resulting in an increased misclassification rate. In addition, the differences in interactive behavior between autonomous vehicles and human-driven vehicles in mixed traffic environments have not been fully considered.

[0007] (5) Although Federated Learning (FL) provides a new approach to distributed data privacy protection, its combination with traditional clustering algorithms still faces technical bottlenecks. Existing federated clustering methods (such as FedAvg-Kmeans) directly average and aggregate local cluster centers, ignoring the heterogeneity of data distribution across different clients (Non-IID). For example, the vehicle trajectory characteristics at urban intersections and highways are significantly different. Direct global aggregation will cause cluster center offsets and reduce the model convergence speed by more than 50%. In addition, existing methods do not solve the compatibility problem between high-dimensional feature embedding and federated training, resulting in serious information loss of low-dimensional features during transmission. Summary of the Invention

[0008] In view of this, an object of the present invention is to provide a driving behavior analysis method based on federated K-means clustering.

[0009] In order to achieve the above object, the present invention provides the following technical solutions:

[0010] A driving behavior analysis method based on federated K-means clustering includes the following steps:

[0011] S1: Select connected vehicles or roadside communication units as clients participating in federated learning, fuse vehicle trajectory data with high-precision map information to extract spatiotemporal feature set X, pre-process the driving data and perform feature dimensionality reduction through an autoencoder to obtain a low-dimensional feature representation H;

[0012] S2: Perform distributed clustering analysis on the low-dimensional feature representation H using the federated K-means clustering algorithm. This involves multiple rounds of iterative processes, including initializing the global cluster center, calculating local cluster parameters on the client, and generating new cluster centers through server aggregation, until the global cluster center converges.

[0013] S3: Classify the driving behaviors according to the driving characteristic indicators in the clustering results and generate classification results for different driving behavior categories.

[0014] Furthermore, in S1, the selected client includes a trusted execution environment, the driving data is stored locally in the vehicle or on an edge device, and the preprocessing includes using linear interpolation to process missing values, eliminating outliers based on the 3σ criterion, performing noise processing through sliding average filtering, and using Z-score standardization and Min-Max normalization to eliminate dimensional differences.

[0015] Furthermore, the extraction of the spatiotemporal feature set X includes: performing spatiotemporal matching on the position, speed, and direction change data of the vehicle trajectory with the lane lines and traffic signs of the high-precision map, generating the position offset relative to the road elements, the slope of the speed change curve before the stop line, the standard deviation of the steering wheel angle, and the Pearson correlation coefficient between the vehicle distance and speed.

[0016] Furthermore, the autoencoder includes a three-layer encoder and a three-layer decoder. The encoder uses the ReLU activation function and the decoder uses the Sigmoid activation function. The reconstruction loss function is minimized by the Adam optimizer. The minimized reconstruction loss function is:

[0017]

[0018] Among them L res represents the reconstruction loss, x i Represents the original data, represents the reconstructed data; g(·) represents the decoding function; f(·) represents the encoding function.

[0019] Furthermore, the federated K-means clustering algorithm in S2 specifically includes:

[0020] The server initializes and broadcasts k cluster centers to the client;

[0021] Each client calculates the probability of data point belonging and updates the local cluster center. The probability calculation formula is:

[0022]

[0023] where p ij represents the probability that data point i belongs to cluster j, represents the data point i of client m, represents the cluster center j of client m;

[0024] The client performs local K-means clustering and uploads the cluster center and sample size;

[0025] The server generates a new global cluster center through weighted averaging. The calculation formula is:

[0026]

[0027] in represents the j global cluster centers after t+1 rounds of iteration, represents the number of clients m belonging to cluster center j, represents the local cluster center j of client m.

[0028] Furthermore, the client uses probability weighted calculation when updating the local cluster center:

[0029]

[0030] in represents the j updated cluster centers of client m; n m is the total number of samples of client m.

[0031] Furthermore, in S3, the driving behavior classification criteria include:

[0032] Speed change mode classification: The slope of the speed curve 5 seconds before the stop line is greater than -0.5m / s 2 Determined as aggressive behavior, less than -1.0m / s 2 Determined to be conservative behavior;

[0033] Steering rate classification: A steering wheel angle standard deviation of less than 15° is considered aggressive driving, and a standard deviation of more than 25° is considered conservative driving;

[0034] Following vehicle characteristic classification: If the correlation coefficient between vehicle distance and speed is greater than 0.7, it is judged as conservative driving.

[0035] Furthermore, in the federated learning process, AES encryption is used to transmit clustering parameters between the client and the server, and the server only stores the aggregated global cluster centers.

[0036] Furthermore, the local K-means clustering is achieved by minimizing the clustering loss function:

[0037]

[0038] Among them L clu is the clustering loss; h i is a low-dimensional feature representation, c j is the cluster center.

[0039] Furthermore, the high-precision map information includes lane line topology, traffic sign locations and intersection three-dimensional layout data, and the spatiotemporal matching error is controlled within the range of ±0.5 meters.

[0040] The beneficial effects of the present invention are:

[0041] (1) Based on the federated learning architecture, raw driving data (such as vehicle trajectory and driver ID) is always retained on the local client (connected vehicle / RSU), and only cluster center parameters (such as probability-weighted feature vectors) are transmitted, thus avoiding data leakage risks from the system design level. Combining the Trusted Execution Environment (TEE) with AES encryption transmission technology, it meets the requirements of data privacy regulations such as GDPR and CCPA, reducing the risk of data leakage.

[0042] The probability-weighted federated clustering center aggregation strategy effectively alleviates the model drift problem caused by non-IID data distribution. Experiments show that the clustering accuracy (ARI index) reaches 0.892 with 100 heterogeneous clients (mixed urban and highway scenarios).

[0043] (2) By matching vehicle trajectories with high-precision maps in time and space, we extract fine-grained features such as lane offset (error ±0.3 meters), slope of the speed curve before the stop line, and standard deviation of the steering angle, expanding the feature dimension from the traditional 6 dimensions to 15 dimensions and significantly improving the classification recall rate. A three-layer autoencoder is used for feature compression, reducing the feature dimension from 15 dimensions to 2 dimensions while retaining 95% of the original information, and improving computational efficiency by 3.2 times (the actual single inference time measured on a Tesla T4 GPU is <2ms). Combined with sliding average filtering and 3σ outlier removal, the clustering purity (Purity index) can still maintain 88.6% under 20% Gaussian noise interference.

[0044] (3) Adaptively adjust the classification criteria based on the clustering results, for example:

[0045] Speed change mode: Classification accuracy is improved based on the distribution range of the speed curve slope 5 seconds before the stop line.

[0046] Steering behavior recognition: The classification threshold is dynamically set by the cluster center value of the steering wheel angle standard deviation, reducing the misjudgment rate.

[0047] In a test set containing both human-driven and autonomous vehicles, federated clustering is used to capture heterogeneous driving patterns, and the classification F1-score reaches 0.916.

[0048] (4) Through automatic encoder dimensionality reduction (2-dimensional feature transmission) and parameter encryption compression (AES-256), the communication data volume per round for a single client is only 2.7KB, which reduces bandwidth usage compared to traditional centralized methods (transmitting 15-dimensional raw features). The average power consumption of the local K-means clustering task on an RSU device (such as NVIDIA Jetson Xavier) is 8.3W, supporting 50 clients concurrently to meet the real-time requirements of the edge. It supports a dynamic client join / exit mechanism. At the scale of 1000 clients, the number of communication rounds required for global cluster center convergence only increases by 18% (from 150 rounds to 177 rounds).

[0049] (5) The Spearman correlation coefficient between the clustering results of aggressive driving categories and historical accident data reached 0.78, which can provide early warning for high-risk road sections. By analyzing the following characteristics of conservative drivers (vehicle distance-speed correlation coefficient > 0.7), a safe following strategy library was generated to reduce the frequency of sudden braking of autonomous vehicles in mixed traffic scenarios. Based on the clustering results, a regional driving behavior heat map was generated to assist traffic management departments in optimizing signal timing plans, and the actual intersection traffic efficiency was improved.

[0050] Other advantages, objects, and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art upon examination of the following description or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below with reference to the accompanying drawings, in which:

[0052] Figure 1 This is a flowchart of a driving behavior analysis method based on federated K-means clustering;

[0053] Figure 2 This is a structural diagram of a driving behavior analysis method based on federated K-means clustering;

[0054] Figure 3 This figure shows the clustering results and driving behavior classification visualization results of a driving behavior analysis method based on federated K-means clustering. DETAILED DESCRIPTION

[0055] The following describes the embodiments of the present invention by means of specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and the following embodiments and features in the embodiments can be combined with each other without conflict.

[0056] Among them, the accompanying drawings are only for illustrative purposes and represent only schematic diagrams rather than actual pictures, and should not be understood as limiting the present invention. In order to better illustrate the embodiments of the present invention, some parts of the accompanying drawings may be omitted, enlarged or reduced, and do not represent the dimensions of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions may be omitted in the accompanying drawings.

[0057] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "back", etc. indicating directions or positional relationships, they are based on the directions or positional relationships shown in the drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operate in a specific direction. Therefore, the terms describing the positional relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.

[0058] like Figure 1 As shown, the present invention is a driving behavior analysis method based on federated K-means clustering, comprising:

[0059] S1, data acquisition. Select a connected vehicle or roadside unit (RSU) as the client and collect driving data of local vehicles through the Internet of Vehicles.

[0060] S2, data preprocessing, involves data type conversion, data cleaning, standardization, and normalization of the collected data to ensure data integrity and accuracy. This extracts features X, which include time-varying information such as the vehicle's position, speed, and direction, as well as detailed road information derived from high-precision maps, such as lane markings, traffic signs, and intersection layouts.

[0061] S3: Autoencoder obtains low-dimensional representation. The autoencoder performs preliminary feature extraction, dimensionality reduction, and denoising on the extracted features X to obtain low-dimensional representation features H in the embedding space.

[0062] Specifically, the encoding function f(·) is used to map the original data into a low-dimensional embedding space, and the decoding function g(·) is used to reconstruct the original data from the embedding. The optimization is performed by minimizing the reconstruction loss. The reconstruction loss is calculated as follows:

[0063]

[0064] Among them L res represents the reconstruction loss, x i Represents the original data, Indicates reconstruction data.

[0065] S4, Federated K-means clustering: In the federated learning framework, the federated K-means clustering algorithm is used to perform cluster analysis on driving behavior data.

[0066] Specifically, the central server first initializes the cluster center and broadcasts it to each client (i.e., connected vehicle or roadside unit). Each client calculates the probability that the data point belongs to the cluster center based on the received cluster center and uses this probability to update the local cluster center.

[0067] The formula for calculating the probability that a data point belongs to the cluster center is:

[0068]

[0069] where p ij represents the probability that data point i belongs to cluster j, represents the data point i of client m, represents the cluster center j of client m.

[0070] The calculation formula for updating the local cluster center using probability is:

[0071]

[0072] in represents the j updated cluster centers of client m.

[0073] Subsequently, each client performs the K-means clustering algorithm on the local data points to obtain the local cluster center and the corresponding number of samples.

[0074] The client calculates the distance of each data point to the k cluster centers and assigns it to the nearest cluster. It optimizes by minimizing the clustering loss function. The calculation formula of the K-means clustering algorithm is:

[0075]

[0076] Among them L clu is the clustering loss.

[0077] These local results are uploaded to the central server, which receives and aggregates the local results from each client for weighted averaging and calculates a new global cluster center.

[0078] The new global cluster center calculation formula is:

[0079]

[0080] in represents the j global cluster centers after t+1 rounds of iteration, represents the number of clients m belonging to cluster center j, represents the local cluster center j of client m.

[0081] The new global cluster center is sent to each client again for further local iterative optimization until the global cluster center converges or reaches the predetermined number of iterations.

[0082] S5: Based on the clustering results, driving behaviors are classified into different categories. Classification is based on driving characteristics such as speed change before reaching the stop line, turning rate, and correlation between the number of vehicles, distance between vehicles, and speed. For example, not reducing speed to zero before reaching the stop line is considered aggressive behavior; driving at a lower speed until coming to a complete stop is considered conservative behavior. Aggressive drivers have a lower turning rate and tend to change lanes earlier to save time; conservative drivers, on the other hand, have a higher turning rate and tend to drive straight.

[0083] Based on the classification results of driving behavior, suggestions and measures are provided to improve traffic safety.

[0084] The specific implementation steps are as follows:

[0085] S1: Data collection and preprocessing

[0086] Step 1: Deploy three RSU clients at an intersection in Chongqing and collect trajectory data (sampling rate 10 Hz) from 500 vehicles while they are at red lights. Six-dimensional features are extracted, including the speed drop rate before the stop line, steering angle fluctuation, and time distance to the preceding vehicle.

[0087] Step 2: Clean the collected driving data to handle missing values, outliers, and noise. Then, standardize and normalize the data to eliminate the dimensional differences between different features.

[0088] Step 3: Use the autoencoder to perform preliminary feature extraction, dimensionality reduction and denoising on the cleaned features.

[0089] Specifically, the encoding function f(·) is used to map the original data into a low-dimensional embedding space, and then the decoding function g(·) is used to reconstruct the original data from the embedding. The optimization is performed by minimizing the reconstruction loss. The reconstruction loss is calculated as follows:

[0090]

[0091] Among them L res represents the reconstruction loss, x i Represents the original data, Indicates reconstruction data.

[0092] The autoencoder was trained for 500 epochs and the learning rate was set to 0.001.

[0093] S2: Federated K-means clustering

[0094] Step 1: The central server initializes k=3 cluster centers and broadcasts them to each client.

[0095] Step 2: Each client calculates the probability that the data point belongs to the cluster center based on the received cluster center, and uses this probability to update the local cluster center.

[0096] The formula for calculating the probability that a data point belongs to a cluster center is as follows:

[0097]

[0098] where p ij represents the probability that data point i belongs to cluster j, represents the data point i of client m, represents the cluster center j of client m.

[0099] Use probability to update the local cluster center. The calculation formula is as follows:

[0100]

[0101] in represents the j updated cluster centers of client m.

[0102] Step 3: Each client executes the K-means clustering algorithm on the local data points to obtain the local cluster center and the corresponding number of samples.

[0103] The client calculates the distance of each data point to the k cluster centers and assigns it to the nearest cluster, optimizing by minimizing the clustering loss function. The calculation formula is:

[0104]

[0105] Among them L clu is the clustering loss.

[0106] Step 4: These local results are uploaded to the central server, which receives and aggregates the local results from each client for weighted averaging and calculates a new global cluster center.

[0107] The new global cluster center calculation formula is:

[0108]

[0109] in represents the j global cluster centers after t+1 rounds of iteration, represents the number of clients m belonging to cluster center j, represents the local cluster center j of client m.

[0110] The new global cluster center is sent to each client for further local iterative optimization. Steps 2 to 4 are repeated until the global cluster center converges or the predetermined number of iterations is reached.

[0111] The federated K-means clustering algorithm sets the number of local iterations to 20 and the number of communication rounds to 150.

[0112] S3: Driving behavior classification, as shown in Table 1.

[0113] Table 1

[0114] category Proportion Typical characteristics Radical 35% <![CDATA[Hard acceleration (>2.5 m / s 2 ) and frequent lane changes]]> Common type 20% Constant speed driving and smooth steering Conservative 45% Slow down in advance (starting from 50m from the stop line)

[0115] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions, which should all be included in the scope of the claims of the present invention.

Claims

1. A driving behavior analysis method based on federated K-means clustering, characterized by: The following steps are involved: S1: Select connected vehicles or roadside communication units as clients participating in federated learning, fuse vehicle trajectory data with high-precision map information to extract spatiotemporal feature set X, pre-process the driving data and perform feature dimensionality reduction through an autoencoder to obtain a low-dimensional feature representation H; S2: Perform distributed clustering analysis on the low-dimensional feature representation H using the federated K-means clustering algorithm. This involves multiple rounds of iterative processes, including initializing the global cluster center, calculating local cluster parameters on the client, and generating new cluster centers through server aggregation, until the global cluster center converges. S3: Classify the driving behaviors according to the driving characteristic indicators in the clustering results and generate classification results for different driving behavior categories.

2. The driving behavior analysis method based on federated K-means clustering according to claim 1, characterized in that: In S1, the selected client includes a trusted execution environment, and the driving data is stored locally in the vehicle or on an edge device. The preprocessing includes using linear interpolation to process missing values, eliminating outliers based on the 3σ criterion, performing noise processing through sliding average filtering, and using Z-score standardization and Min-Max normalization to eliminate dimensional differences.

3. The driving behavior analysis method based on federated K-means clustering according to claim 1, characterized in that: Extraction of the spatiotemporal feature set X includes performing spatiotemporal matching of the position, speed, and direction change data of the vehicle trajectory with the lane lines and traffic signs of the high-precision map, generating a position offset relative to the road element, a slope of the speed change curve before the stop line, a standard deviation of the steering wheel angle, and a Pearson correlation coefficient between vehicle distance and speed.

4. The driving behavior analysis method based on federated K-means clustering according to claim 1, characterized in that: The autoencoder includes a three-layer encoder and a three-layer decoder. The encoder uses the ReLU activation function and the decoder uses the Sigmoid activation function. The reconstruction loss function is minimized by the Adam optimizer. The minimized reconstruction loss function is: Among them L res represents the reconstruction loss, x i Represents the original data, represents the reconstructed data; g(·) represents the decoding function; f(·) represents the encoding function.

5. The driving behavior analysis method based on federated K-means clustering according to claim 1, characterized in that: The federated K-means clustering algorithm in S2 specifically includes: The server initializes and broadcasts k cluster centers to the client; Each client calculates the probability of data point belonging and updates the local cluster center. The probability calculation formula is: where p ij represents the probability that data point i belongs to cluster j, represents the data point i of client m, represents the cluster center j of client m; The client performs local K-means clustering and uploads the cluster center and sample size; The server generates a new global cluster center through weighted averaging. The calculation formula is: in represents the j global cluster centers after t+1 rounds of iteration, represents the number of clients m belonging to cluster center j, represents the local cluster center j of client m.

6. The driving behavior analysis method based on federated K-means clustering according to claim 5, characterized in that: The client uses probability weighted calculation when updating the local cluster center: in represents the j updated cluster centers of client m; n m is the total number of samples of client m.

7. The driving behavior analysis method based on federated K-means clustering according to claim 1, characterized in that: In S3, the driving behavior classification criteria include: Speed change mode classification: The slope of the speed curve 5 seconds before the stop line is greater than -0.5m / s 2 Determined as aggressive behavior, less than -1.0m / s 2 Determined to be conservative behavior; Steering rate classification: A steering wheel angle standard deviation of less than 15° is considered aggressive driving, and a standard deviation of more than 25° is considered conservative driving; Following vehicle characteristic classification: If the correlation coefficient between vehicle distance and speed is greater than 0.7, it is judged as conservative driving.

8. The driving behavior analysis method based on federated K-means clustering according to claim 1, characterized in that: In the federated learning process, clustering parameters are transmitted between the client and the server using AES encryption, and the server only stores the aggregated global cluster centers.

9. The driving behavior analysis method based on federated K-means clustering according to claim 5, characterized in that: The local K-means clustering is achieved by minimizing the clustering loss function: Among them L clu is the clustering loss; h i is a low-dimensional feature representation, c j is the cluster center.

10. The driving behavior analysis method based on federated K-means clustering according to claim 1, characterized in that: The high-precision map information includes lane line topology, traffic sign locations and intersection three-dimensional layout data, and the spatiotemporal matching error is controlled within the range of ±0.5 meters.

Citation Information

Patent Citations

  • Driving behavior analysis method based on improved K-means

    CN111461185A

  • Driver behavior cloud-side collaborative learning system based on federated transfer learning

    CN111476139A

  • Big data privacy protection method and system based on federated learning

    CN117972783A

  • Adaptive analysis of driver behavior

    US20180174485A1