Personalized driving habit analysis method based on big data

Through big data analysis methods, collect and process driving data, identify driving behavior patterns and build personalized portraits, solve the problem of difficult to analyze and constrain dangerous driving habits in the existing technology, realize safe and economical driving suggestions, and provide data support for related industries.

CN120217067APending Publication Date: 2025-06-27苏晓磊
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411724006.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-28
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively analyze and constrain drivers’ dangerous driving habits, and cannot provide economical driving habit suggestions in different scenarios.

Method used

Using personalized driving habit analysis methods based on big data, data is collected through OBD, mobile applications and urban traffic monitoring systems, pre-processing, behavioral feature extraction, modeling and neural network identification, personalized driver portraits are constructed, and suggestions for safe driving and energy conservation and emission reduction are provided.

Benefits of technology

Effectively analyze drivers’ personalized driving habits, provide drivers with safe and economical driving advice, and provide data support for the traffic management and insurance industries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005158759160000021
    Figure BDA0005158759160000021
  • Figure BDA0005158759160000051
    Figure BDA0005158759160000051
  • Figure FDA0005158759150000021
    Figure FDA0005158759150000021
Patent Text Reader

Abstract

The invention provides a personalized driving habit analysis method based on big data, and the method comprises the following steps: S1, collecting vehicle and real-time operation data through an OBD, meanwhile, collecting mobile phone driving behavior data through a mobile phone application of a driver, and collecting road use condition data through an urban traffic monitoring system; s2, preprocessing the data; s3, performing behavior feature extraction processing on the preprocessed data; s4, modeling the driving behavior according to the data subjected to feature extraction processing, and identifying a complex driving mode by using a neural network; s5, a personalized driving habit portrait is constructed for the driver according to the driving behavior data, and the driving risk of the driver is evaluated; according to the method, the personalized driving habits of the driver can be effectively analyzed, safer and more economical driving suggestions are provided for the driver, and data support is provided for industries such as traffic management and insurance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly relates to a personalized driving habit analysis method based on big data. Background Art

[0002] Driving habits generally refer to the stable driving behavior patterns formed by drivers during the driving process. These habits cover various behaviors and decision-making methods of drivers when operating vehicles, including both technical-level operations and psychological attitudes of observing traffic rules and driving etiquette;

[0003] The driving habits of each relevant person are not concerned. When restraining their dangerous driving habits, there will be greater driving risks, and in different scenarios, more economical driving habits cannot be adopted. Therefore, a personalized driving habit analysis method based on big data is proposed. Summary of the Invention

[0004] In view of this, embodiments of the present invention hope to provide a personalized driving habit analysis method based on big data to solve or alleviate the technical problems existing in the prior art, and at least provide a beneficial option.

[0005] The technical solution of the embodiments of the present invention is implemented as follows: The personalized driving habit analysis method based on big data includes the following steps:

[0006] S1. Collect vehicle and real-time operation data through OBD, and at the same time collect mobile phone driving behavior data of the driver through the driver's mobile phone, and collect road usage data by using the urban traffic monitoring system;

[0007] S2. Preprocess the data;

[0008] S3. Perform behavior feature extraction processing on the preprocessed data;

[0009] S4. Model the driving behavior according to the data after feature extraction processing, and use a neural network to identify complex driving patterns;

[0010] S5. Construct a personalized driving habit portrait for the driver according to the driving behavior data, and evaluate the driving risk of the driver;

[0011] S6. Provide personalized safe driving suggestions according to the driving habits, and give driving strategies for energy conservation and emission reduction.

[0012] In some embodiments, in the step S2, the data preprocessing includes the following steps

[0013] S21. Remove invalid, incorrect, and duplicate data;

[0014] S22. Unify the formats of data from different sources and formats for easy analysis;

[0015] S23. Classify and label the data.

[0016] In some embodiments, in the said S3, the driving behavior feature extraction process includes the following steps,

[0017] S31. Calculate the average values of the driver's speed and acceleration;

[0018] S32. Determine the median of the driving data, which helps to understand the central tendency of the driving behavior;

[0019] S33. Find out the driving behavior pattern with the highest occurrence frequency;

[0020] S34. Measure the dispersion degree of the driving behavior data;

[0021] S35. Divide the data into four parts to better understand the data distribution;

[0022] S36. Identify the change trend of the data over time by calculating the moving average method;

[0023] S37. Decompose the time series into trend, seasonal and random components, which is often used to identify and quantify seasonal patterns;

[0024] S38. Identify the long-term fluctuations in the time series, which are usually related to the economic cycle.

[0025] In some embodiments, in the said S36, the moving average method includes the following steps,

[0026] S361. Determine the size of the time window for calculating the moving average;

[0027] S362. For the first moving average value of the time series, only calculate the average value of the data points within the time window;

[0028] S363. Starting from the second time window, each time the average value is calculated, move forward one time unit and exclude the earliest data point, including the latest data point;

[0029] S364. Repeat the calculation of the average value for each new time window until the entire time series is covered;

[0030] S365. Plot the calculated moving average values on the time series chart to observe the trend of the data.

[0031] In some embodiments, in the said S4, a classification algorithm is used to model the driving behavior, and the formula of the classification algorithm is as follows,

[0032]

[0033] Among them, P(Y = 1|X) is the probability that the driving behavior belongs to category 1 given the feature X, β0 is the intercept term, and β1,..., β n are the coefficients of the features X1,..., X n .

[0034] In some embodiments, in the S1, the data includes basic data, operation data, and status data. The basic data includes vehicle information, driver information, and driving route. The operation data includes acceleration, deceleration, turning, braking frequency and intensity. The status data includes vehicle speed, engine status, and fuel consumption.

[0035] Due to the above technical solutions adopted in the embodiments of the present invention, it has the following advantages:

[0036] Through the present analysis method, the present invention can effectively analyze the personalized driving habits of drivers, provide safer and more economical driving suggestions for drivers, and at the same time provide data support for industries such as traffic management and insurance.

[0037] The above summary is only for the purpose of the specification and is not intended to be limiting in any way. Detailed Embodiments

[0038] In the following, only some exemplary embodiments are simply described. As those skilled in the art can recognize, the described embodiments can be modified in various different ways without departing from the spirit or scope of the present invention.

[0039] It should be noted that terms such as "first", "second", "symmetric", "array", etc. are only used for the purpose of distinguishing descriptions and position descriptions, and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, features defined with "first", "symmetric", etc. can explicitly or implicitly include one or more of such features; similarly, when certain features are not limited in quantity by words such as "two", "three", etc., it should be noted that such features also belong to explicitly or implicitly including one or more feature quantities;

[0040] In the present invention, unless otherwise clearly defined and limited, terms such as "installation", "connection", "fixation", etc. should be understood in a broad sense; for example, it can be a fixed connection, a detachable connection, or an integral molding; it can be a mechanical connection, a direct connection, a welding connection, or an indirect connection through an intermediate medium, and it can be the communication inside two components or the interaction relationship between two components.

[0041] The embodiments of the present invention provide a personalized driving habit analysis method based on big data, including the following steps:

[0042] S1. Collect vehicle and real-time operation data through OBD. At the same time, collect driving behavior data through the driver's mobile phone application, and collect road usage data using the urban traffic monitoring system;

[0043] S2. Preprocess the data;

[0044] S3. Extract behavioral features from the preprocessed data;

[0045] S4. Model the driving behavior based on the data after feature extraction, and use neural networks to identify complex driving patterns;

[0046] S5. Construct a personalized driving habit profile for the driver based on the driving behavior data and evaluate the driving risk of the driver;

[0047] S6. Provide personalized safe driving suggestions based on driving habits and give driving strategies for energy conservation and emission reduction.

[0048] In this embodiment, specifically, in S2, the data preprocessing includes the following steps.

[0049] S21. Remove invalid, incorrect, and duplicate data;

[0050] S22. Unify the formats of data from different sources and formats for easy analysis;

[0051] S23. Classify and label the data, such as normal driving, hard braking, speeding, etc.

[0052] In this embodiment, specifically, in S3, the behavioral feature extraction processing includes the following steps.

[0053] S31. Calculate the average values of the driver's speed and acceleration;

[0054] S32. Determine the median of the driving data, which helps to understand the central tendency of the driving behavior;

[0055] S33. Find the driving behavior pattern with the highest frequency of occurrence;

[0056] S34. Measure the dispersion degree of the driving behavior data;

[0057] S35. Divide the data into four parts to better understand the data distribution;

[0058] S36. Identify the change trend of the data over time by calculating the moving average method;

[0059] S37. Decompose the time series into trend, seasonal, and random components, which is often used to identify and quantify seasonal patterns;

[0060] S38. Identify long-term fluctuations in the time series, which are usually related to the economic cycle.

[0061] In this embodiment, specifically, in S36, the moving average method includes the following steps.

[0062] S361. Determine the size of the time window for calculating the moving average, which is usually a fixed time period. For example, data for the last 5 days, 10 days, or 30 days.

[0063] S362. For the first moving average of the time series, only calculate the average of the data points within the time window.

[0064] S363. Starting from the second time window, each time calculating the average, move forward one time unit (e.g., one day), and exclude the earliest data point, including the latest data point.

[0065] S364. Repeat calculating the average for each new time window until the entire time series is covered.

[0066] S365. Plot the calculated moving average on the time series chart to observe the trend of the data.

[0067] The code example is as follows.

[0068] import pandas as pd

[0069] import matplotlib.pyplot as plt

[0070] # Assume df is a Pandas DataFrame containing time series data, 'time' is the time column, and 'value' is the column of the value for which the trend needs to be calculated.

[0071] df['moving_average'] = df['value'].rolling(window = 5).mean() # Use 5 days as the time window

[0072] plt.plot(df['time'], df['value'], label = 'Original')

[0073] plt.plot(df['time'], df['moving_average'], label = 'Moving Average', color ='red')

[0074] plt.legend()

[0075] plt.show().

[0076] In this embodiment, specifically, in S4, a classification algorithm is used to model the driving behavior. The formula of the classification algorithm is as follows:

[0077]

[0078] where P(Y = 1|X) is the probability that the driving behavior belongs to category 1 given the feature X, β0 is the intercept term, and β1,..., β n are the coefficients of the features X1,..., X n .

[0079] In this embodiment, specifically, in S1, the data includes basic data, operation data, and status data. The basic data includes vehicle information, driver information, and driving route. The operation data includes acceleration, deceleration, turning, braking frequency, and intensity. The status data includes vehicle speed, engine status, and fuel consumption.

[0080] In this embodiment, specifically, an example of using a neural network to identify complex driving patterns is as follows:

[0081] Using CNN for image data recognition

[0082] Input: Images of the driving scene captured by on-vehicle cameras.

[0083] Network structure:

[0084] Convolutional layer: Extract image features.

[0085] Pooling layer: Reduce the feature dimension and retain important information.

[0086] Fully connected layer: Perform classification.

[0087] Output: Driving mode category.

[0088] Using LSTM for time series data recognition

[0089] Input: Sequences of driving behavior data within a continuous time period.

[0090] Network structure:

[0091] LSTM layer: Capture long-term dependencies in time series data.

[0092] Fully connected layer: Perform classification.

[0093] Output: Driving mode category.

[0094] In this embodiment, specifically, the steps of decomposing the time series into trend, seasonal, and random components are as follows:

[0095] 1. Observation and identification

[0096] Preliminary observation: By plotting the time series graph, observe whether there are obvious trends, seasonal fluctuations, and random components in the data.

[0097] Identify periodicity: Determine the cycle length of the seasonal component. For example, monthly data may have a seasonal cycle of 12 months.

[0098] 2. Select a decomposition model

[0099] Determine the model type: Select a suitable decomposition model, such as an additive model, a multiplicative model, or a log-additive model.

[0100] Additive model: Observed value = Trend + Seasonal + Random component

[0101] Multiplicative model: Observed value = Trend × Seasonal × Random component

[0102] Log-additive model: Observed value = exp(Trend + Seasonal + Random component)

[0103] 3. Calculate the trend component

[0104] Remove seasonality: If the seasonality is obvious, the seasonal component can be removed from the original data first, and then the trend is calculated.

[0105] Trend estimation: Methods such as moving average and exponential smoothing can be used to estimate the trend component.

[0106] 4. Calculate the seasonal component

[0107] Seasonal adjustment: Divide the original data by the trend component to obtain the seasonally adjusted data.

[0108] Calculate the seasonal index: Calculate the average value of each season and compare it with the overall average value to obtain the seasonal index.

[0109] 5. Calculate the random component

[0110] Residual calculation: After removing the trend and seasonal components from the original data, the remaining part is the random component.

[0111] Analyze the residuals: Check whether the residuals are white noise, that is, whether they have a constant mean and variance, and the residuals are independent of each other.

[0112] 6. Adjustment and optimization

[0113] Parameter optimization: According to the results of the residual analysis, adjust the model parameters to obtain a better fitting effect.

[0114] Model validation: Use a part of the data as the test set to verify the prediction ability of the model.

[0115] 7. Reconstruct the time series

[0116] Reconstruct the data: Recombine the trend, seasonal, and random components according to the selected model type.

[0117] Visual inspection: Plot the reconstructed time series graph and compare it with the original data to check the decomposition effect.

[0118] 8. Prediction

[0119] Predict future values: Use the decomposed components to predict future values.

[0120] In this embodiment, specifically, the method for measuring the dispersion degree of driving behavior data is as follows.

[0121] Range:

[0122] Difference between the maximum and minimum values: This is the simplest method for measuring dispersion, which calculates the difference between the maximum and minimum values in the dataset.

[0123] Interquartile range:

[0124] Difference between the 75th percentile and the 25th percentile: The IQR can exclude the influence of extreme values and better reflect the dispersion degree of the central part of the data.

[0125] Variance:

[0126] Average of the squared differences between each data point and the mean: Variance is an important indicator for measuring the dispersion of data point distributions, but it is affected by the data scale.

[0127] Standard deviation:

[0128] Square root of the variance: The standard deviation is the square root of the variance, which represents the dispersion degree in the units of the original data and is the most widely used method for measuring dispersion.

[0129] Mean absolute deviation:

[0130] Average of the absolute values of the differences between each data point and the mean: The MAD is a measure of the average degree of deviation of data points from the mean and is not affected by extreme values.

[0131] Coefficient of variation:

[0132] Standard deviation divided by the mean: The coefficient of variation is a dimensionless indicator used to compare the dispersion degrees of different datasets, especially suitable for datasets with large differences in means.

[0133] Ratio of the range to the mean:

[0134] Range divided by the mean: This ratio can be used to measure the relative relationship between the dispersion of data and the average level.

[0135] Coefficient of dispersion:

[0136] Variance divided by the mean: Similar to the coefficient of variation, but usually used for count data.

[0137] Sum of squares:

[0138] Sum of the squares of the differences between each data point and the overall mean: Commonly used in statistical tests such as ANOVA and can be further decomposed into the sum of squares within groups and the sum of squares between groups.

[0139] Kurtosis and skewness:

[0140] Kurtosis: Describes the "sharpness" of the shape of the data distribution.

[0141] Skewness: Describes the asymmetry of the data distribution.

[0142] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various changes or substitutions, and these should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A personalized driving habit analysis method based on big data, characterized in that: The following steps are involved: S1. Collect vehicle and real-time operation data through OBD, as well as driving behavior data through the driver's mobile phone application, and collect road usage data using the city traffic monitoring system; S2, preprocessing the data; S3, extracting behavioral features from the preprocessed data; S4, model driving behavior based on feature extraction and processed data, and use neural networks to identify complex driving patterns; S5. Build a personalized driving habit profile for the driver based on driving behavior data and assess the driver's driving risk; S6. Provide personalized safe driving suggestions based on driving habits and give driving strategies for energy conservation and emission reduction.

2. The method for analyzing personalized driving habits based on big data according to claim 1, characterized in that: In said S2, data preprocessing includes the following steps: S21. Remove invalid, erroneous and duplicate data; S22. Unify data from different sources and formats to facilitate analysis; S23. Classify and label the data.

3. The method for analyzing personalized driving habits based on big data according to claim 1, characterized in that: In said S3, the behavior feature extraction process comprises the following steps: S31, calculating the average value of the driver's speed and acceleration; S32, determining the median value of driving data, which helps to understand the central trend of driving behavior; S33, finding the most frequent driving behavior pattern; S34, measure the discreteness of driving behavior data; S35. Divide the data into four parts to better understand the distribution of the data; S36, identifying the trend of data changes over time by calculating the moving average method; S37. Decomposing time series into trend, seasonal and random components is often used to identify and quantify seasonal patterns; S38. Identify long-term fluctuations in time series, usually associated with economic cycles.

4. The method for analyzing personalized driving habits based on big data according to claim 3 is characterized by: In said S36, the moving average method comprises the following steps: S361, determining the time window size for calculating the moving average; S362, for the first moving average of the time series, only the average of the data points within the time window is calculated; S363, starting from the second time window, each time the average value is calculated, move forward one time unit, exclude the earliest data point, and include the latest data point; S364, repeatedly calculating the average value for each new time window until the entire time series is covered; S365. Plot the calculated moving average on a time series chart to observe the trend of the data.

5. The personalized driving habit analysis method based on big data according to claim 1 is characterized by: In S4, a classification algorithm is used to model the driving behavior. The formula of the classification algorithm is as follows: Where P(Y=1|X) is the probability that the driving behavior belongs to category 1 given feature X, β0 is the intercept term, β1, ..., β n are the features X1,...,X n The coefficient of .

6. The method for analyzing personalized driving habits based on big data according to claim 1, characterized in that: In S1, the data includes basic data, operation data and status data. The basic data includes vehicle information, driver information and driving route. The operation data includes acceleration, deceleration, turning and braking frequency and strength. The status data includes vehicle speed, engine status and fuel consumption.