Online recognition method of abnormal driving behavior based on Encoder-Decoder attention network and LSTM

By using the Encoder-Decoder attention network and LSTM model on smartphones, combined with the smartphone sensor data, online recognition of abnormal driving behavior is achieved, solving the invasiveness and high cost of data acquisition devices in the prior art, and improving the recognition accuracy.

CN114548216BActive Publication Date: 2025-05-16NANJING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111675120.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-31
Publication Date
2025-05-16
Estimated Expiration
2041-12-31

AI Technical Summary

Technical Problem

The prior art has problems with invasiveness and high cost of data acquisition equipment in driving behavior analysis, and image data is easily affected by light, which consumes a lot of computing resources.

Method used

The online recognition method of abnormal driving behavior based on Encoder-Decoder attention network and LSTM is adopted. Data is obtained through the built-in sensor of the smartphone, and preprocessed and normal distribution transformation is performed. A fusion model is built for training to realize online recognition of abnormal driving behavior.

Benefits of technology

It reduces the difficulty and cost of data acquisition, avoids interference from data acquisition equipment to drivers, improves the recognition accuracy of abnormal driving behavior, and has the advantages of non-invasive, easy data acquisition and low cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114548216B_ABST
    Figure CN114548216B_ABST
Patent Text Reader

Abstract

The present invention discloses an online recognition method for abnormal driving behavior based on Encoder-Decoder attention network and LSTM. The present invention is composed of three main modules, namely, an encoder-decoder based on LSTM, an attention mechanism and a classifier based on SVM, including input encoding, attention learning, feature decoding, sequence reconstruction, residual calculation and driving behavior classification. The present invention is based on mobile phone multi-sensor fusion data, and on the basis of driving behavior data characteristics and behavior pattern analysis, integrates Encoder-Decoder deep learning model, Attention mechanism and SVM classification model to identify abnormal driving behavior. The present invention has the advantages of easy data acquisition, non-intrusiveness and low cost. It not only considers the time correlation of driving behavior, but also considers the differences at different times. It can identify abnormal driving behavior online in an end-to-end manner, and can provide a methodological basis for driving behavior evaluation and safety warning, which is of great significance to the design of intelligent driving system and traffic safety decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to driving behavior recognition technology, specifically to an online recognition method for abnormal driving behavior based on Encoder-Decoder attention network and LSTM. Background Art

[0002] With the development of social economy, the number of motor vehicles has increased rapidly. While vehicles bring convenience to people, they also bring serious traffic safety hazards. More than 90% of traffic accidents are related to the driver's driving behavior, and the main cause is improper driving behavior such as sudden acceleration and deceleration. Therefore, online analysis and identification of abnormal driving behavior is an important way to prevent traffic accidents, reduce casualties, and improve traffic safety.

[0003] Current driving behavior research can be roughly divided into two categories: analysis-based methods and data-driven methods. Analysis-based methods are usually based on the driver's psychological and physiological signals. However, the collection of psychological and physiological signals requires the driver to wear a large number of sensor devices during driving, which is an unfriendly contact data collection method with strong invasiveness and interference, and has a great impact on driving behavior analysis. Data-driven driving behavior analysis methods mainly include two categories: image-based and sensor-based. Due to the rapid development of machine vision in recent years, there are many mature solutions based on image data. However, image data is easily affected by factors such as lighting, the application scenarios are limited, and the amount of video data is extremely large, which requires a lot of computing resources. Driving behavior analysis based on sensor data usually relies on special sensors and requires the installation of a large number of special on-board equipment for the vehicle, such as Lidar, inertial measurement unit IMU, etc. The installation and maintenance costs are extremely high, and it is difficult to promote and apply. Summary of the invention

[0004] The purpose of the present invention is to provide an online recognition method for abnormal driving behavior based on Encoder-Decoder attention network and LSTM.

[0005] The technical solution to achieve the purpose of the present invention is: an online recognition method for abnormal driving behavior based on Encoder-Decoder attention network and LSTM, the specific steps are:

[0006] Step 1: Obtain mobile phone sensor data, perform preprocessing and normal distribution transformation, and obtain driving behavior time series data;

[0007] Step 2: Build the Encoder-Decoder attention network and LSTM fusion model and train it;

[0008] Step 3: Use the trained model to identify abnormal driving behavior.

[0009] Preferably, the specific method for preprocessing the mobile phone sensor data and transforming the normal distribution is:

[0010] Step 1.1: Missing data processing: Find missing data records, determine the nature of missing data, and complete or eliminate missing data;

[0011] Step 1.2: Data normalization: Use the maximum and minimum value normalization method to normalize the data on each feature dimension. The calculation method is as follows:

[0012]

[0013] Among them, value i is the i-th value, value max The maximum value of the current column, value min is the minimum value of the current column, and S is the standardized value;

[0014] Step 1.3: Unbalanced data processing: Use the resampling method to downsample the normal driving data and upsample the abnormal driving data;

[0015] Step 1.4: Normalization transformation of characteristic distribution: Determine whether the data sample has the characteristics of normal distribution. If not, use the mapping relationship to transform it to have the characteristics of normal distribution.

[0016] Preferably, the Encoder-Decoder attention network and LSTM fusion model includes an encoder module, an attention learning module, a decoder module, a sequence reconstruction module, a reconstruction error module and a SVM classifier module.

[0017] Preferably, the specific steps of using the Encoder-Decoder attention network and LSTM fusion model to identify abnormal driving behavior are:

[0018] Step 4.1: Input the time series data into the encoder, and the encoder calculates the hidden layer state at each moment;

[0019] Step 4.2: The attention learning module performs weighted summation of all hidden layer states in the encoder according to the attention weights to obtain the semantic vector at time t;

[0020] Step 4.3: Concatenate the semantic vector at time t with the original hidden layer state of the decoder to obtain the new hidden layer state of the decoder;

[0021] Step 4.4: Use the decoder hidden layer state with fused attention to obtain the reconstructed sequence;

[0022] Step 4.5: Calculate the residual between the reconstructed sequence and the original sequence at each moment to obtain the reconstruction error at each moment, and concatenate them to obtain the mixed classification feature vector;

[0023] Step 4.6: Use the mixed classification feature vector as the input of the SVM classifier, and use the SVM classifier to classify the driving behavior to obtain the classification result to realize abnormal driving behavior recognition.

[0024] Preferably, the specific calculation formula for the encoder to calculate the hidden layer state at each moment is:

[0025]

[0026] In the formula, f is the encoder, is the hidden layer state at time t, x t For time series data.

[0027] Preferably, the attention learning module uses attention weights to allocate all hidden layer states in the encoder i=1,2,…,T, and perform weighted summation to obtain the semantic vector c at time t t , the specific calculation formula is:

[0028]

[0029] In the formula, a t (i) represents the state of the i-th hidden layer of the encoder The hidden layer state of the decoder at time t The weight of .

[0030] Preferably, the encoder i-th hidden layer state The hidden layer state of the decoder at time t The weight of is determined by the correlation score between the hidden layer state of the decoder at time t and the hidden layer state of the encoder at time i. The specific calculation formula is:

[0031]

[0032] Where W a is the weight matrix, exp(·) is the exponential function, and T is the sequence length.

[0033] Preferably, the decoder hidden layer state s with fused attention is used t Get the reconstructed sequence The specific calculation formula is:

[0034]

[0035] Where W hois the decoder hidden layer output coefficient matrix, b h0 is the bias term.

[0036] Preferably, the attention-fused decoder hidden state s t , the specific calculation formula is:

[0037]

[0038] Where W c is the transformation matrix, is the hidden layer state of the decoder at time t, c t is the semantic vector at time t.

[0039] Compared with the prior art, the present invention has the following significant advantages:

[0040] (1) The present invention uses the built-in sensor of the smart phone as the data acquisition means, and does not need to install a dedicated data acquisition device, which reduces the difficulty of data acquisition and has the advantages of easy data acquisition and low cost;

[0041] (2) The present invention uses a smartphone carried in the vehicle to acquire data, without the need to wear invasive devices such as electroencephalograms and eye trackers, thereby avoiding the interference of data acquisition equipment on the driver's driving behavior and having the advantage of being non-invasive.

[0042] (2) From a data-driven perspective, the present invention uses data fusion technology to fuse data from multiple sensors, making full use of multi-source data information and promoting accurate identification of abnormal driving behavior;

[0043] (3) The present invention integrates the Encoder-Decoder deep learning model, the Attention mechanism and the SVM classification model, which enhances the learning ability of the model and improves the accuracy of abnormal driving behavior recognition.

[0044] The present invention will be further described in detail below in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 This is a schematic diagram of the normal transformation of data.

[0046] Figure 2 This is the structural diagram of the abnormal driving behavior recognition model based on the Encoder-Decoder attention network.

[0047] Figure 3 This is the distribution curve of normal driving behavior data changing over time.

[0048] Figure 4 It is the distribution curve of abnormal driving behavior data changing over time.

[0049] Figure 5 The ROC curve and PRC curve of abnormal driving behavior recognition results. Note: LOG-logistic regression binary classification; RF-random forest; feature extraction models all use the above Encoder-Decoder attention model.

[0050] Table 1 shows the abnormal driving behavior recognition results. DETAILED DESCRIPTION

[0051] The online recognition method of abnormal driving behavior based on Encoder-Decoder attention network and LSTM has the following specific steps:

[0052] Step 1: Obtain mobile phone sensor data, perform preprocessing and normal distribution transformation, and obtain driving behavior time series data;

[0053] The present invention takes into account a variety of abnormal driving behaviors. It is difficult for a single sensor to capture a variety of driving behaviors. Therefore, a multi-sensor fusion method is used to describe abnormal driving behaviors. Among them, the accelerometer data reflects the sudden acceleration / deceleration to a certain extent, and the gyroscope data can be used to analyze the sharp turn. The movement caused by rapid lane changes can be recorded by multiple sensors. At the same time, there is a significant coupling between the abnormal driving behaviors. In view of this, the present invention combines the data of different sensors at the same time to form a driving behavior feature vector, and the specific processing method of the sensor data is summarized as follows:

[0054] Step 1.1: Missing data processing. Find missing data records, determine the nature of missing data, use interpolation to fill in accidental missing data, and remove missing records from the data for structural missing data.

[0055] Step 1.2: Data normalization. The original data is composed of multiple sensors, and the measurement units and value ranges of different sensors are different. In order to eliminate the dimensional influence between indicators, the maximum and minimum value normalization method (Min-MaxScaling) is used to normalize the data on each feature dimension. The calculation method is as follows:

[0056]

[0057] Among them, value i is the i-th value, value max The maximum value of the current column, value min is the minimum value of the current column, and S is the standardized value.

[0058] Step 1.3: Unbalanced data processing: Use the resampling method to downsample normal driving and upsample abnormal driving, so that the ratio of positive and negative samples is as close to 1:1 as possible.

[0059] Step 1.4: Normalization transformation of feature distribution. It is necessary to make sure that the data sample has the characteristics of normal distribution. If not, it is necessary to use the mapping relationship to make it have this characteristic.

[0060] Step 2: Construct an Encoder-Decoder model and train it using training samples. The Encoder-Decoder model is a framework for processing time series, so it is usually implemented based on a recurrent neural network (RNN). To process multi-sensor driving data arranged in time series, LSTM, which has the advantage of processing long-term memory, is used as a feature extraction network; after inputting the original data, the model is allowed to perform attention learning to increase the weight of effective information.

[0061] Specifically, in order to process multi-sensor driving data arranged in time series, LSTM, which has the advantage of processing long-term memory, is used as the feature extraction network;

[0062] Furthermore, the Encoder-Decoder attention network and LSTM fusion model includes an encoder module, an attention learning module, a decoder module, a sequence reconstruction module, a reconstruction error module and a SVM classifier module;

[0063] The encoder calculates the hidden layer state at each moment in the following way:

[0064]

[0065] In the formula, f is the encoder, using LSTM network. Using f, the core of the hidden layer is three gated units: forget gate, input gate, and output gate. At each time t, LSTM has two transmission states: unit state c t , hidden layer state h t . Cell state c t The long-term dependency information is saved and is controlled by the forget gate and the input gate. The forget gate determines the unit state c at the previous moment. t-1 How much information is retained to the current cell state c t , the calculation method is as follows,

[0066] f t =σ(W f ·[h t-1 ,x t ]+b f )

[0067] In the formula, f t is the forget gate state at the current moment, x t is the input at the current moment, h t-1is the hidden layer state at the previous moment, [·,·] represents vector concatenation, the symbol “·” represents the dot product operator, W f is the forget gate weight matrix, b f is the forget gate bias term, σ(·) is a nonlinear mapping function (such as the sigmoid function). The input gate determines the input x at the current moment. t How much information is passed to the cell state c t , the calculation formula is as follows,

[0068] i t =σ(W i ·[h t-1 ,x t ]+b i )

[0069]

[0070] In the formula, i t is the input gate state at the current moment, W i is the input gate weight matrix, b i is the input gate bias term, c t With c t-1 are the unit states at the current moment and the previous moment, respectively, and the symbols represents the Hadamard product operator, c t ′ describes the cell state facing the current input, and the calculation method is as follows:

[0071] c′ t =tanh(W c ·[h t-1 ,x t ]+b c )

[0072] Where W c With b c are the weight matrix and bias term respectively.

[0073] Hidden layer state h t is the output of LSTM, which is connected by the output gate and the cell state c t The calculation method is as follows:

[0074]

[0075] Where tanh(·) represents the hyperbolic tangent function, t is the output of the output gate at the current moment, i.e. the hidden layer at the next moment The calculation method is as follows,

[0076] o t =σ(W o·[h t-1 ,x t ]+b o )

[0077] Where W o With b o are the weight matrix and bias term of the output gate respectively.

[0078] The attention learning module uses attention weights to allocate all hidden layer states in the encoder. i=1,2,…,T, and perform weighted summation to obtain the semantic vector c at time t t :

[0079]

[0080] In the formula, a t (i) represents the state of the i-th hidden layer of the encoder The hidden layer state of the decoder at time t The specific calculation method is:

[0081] Calculate the hidden layer state of the decoder at time t and the hidden layer state of the encoder at time i The correlation score between them is normalized. The specific formula is:

[0082]

[0083] Where W a is the weight matrix, exp(·) is the exponential function, and T is the sequence length.

[0084] The semantic vector c at time t t and the original hidden layer state of the decoder Splice together to get the new hidden layer state s of the decoder t :

[0085]

[0086] Since the shape of the output vector changes after splicing, the transformation matrix W is used c After the above processing, the hidden layer state s of the decoder with fusion attention is used t Get the reconstructed sequence

[0087]

[0088] Where W ho is the decoder hidden layer output coefficient matrix, b f is the bias term.

[0089] After obtaining the reconstructed sequence, the reconstructed sequence at each moment With the original sequence X origin ={x1,x2,…,x T} to calculate the residual and obtain the reconstruction error E = {e1, e2, …, e T}, concatenate to get the mixed classification feature vector X = {(x1,e1),(x2,e2),…,(x T ,e T )}, the mixed classification feature vector is used as the input of the SVM classifier, and the driving behavior is classified using the SVM classifier to obtain the classification result o t , t=1,2,…,T, to realize abnormal driving behavior recognition.

[0090] In summary, the model training process includes input encoding, attention learning, feature decoding, sequence reconstruction, residual calculation and driving behavior classification, which can be described in the following steps:

[0091] Step 2.1: Input code. origin ={x1,x2,…,x T} is input to the model, where t=1,2,…,T is the driving behavior feature vector at time t. The LSTM-based Encoder is used to learn the features of the input data and obtain the encoding vector of each input moment.

[0092] Step 2.2: Attention learning. The encoder features H enc As input, combined with the decoder features at each moment The Attention model is used to calculate the encoder feature weights that take into account context information, thereby obtaining the context feature vector c at the current moment. t , t=1,2,…,T;

[0093] Step 2.3: Feature decoding. For each time t, t=1, 2, ..., T, decoder feature at the previous time t-1 Decoder Output and the current context feature c t Input the encoder at the current moment, use the LSTM-based Decoder to learn the input, and get the decoder state at the current moment

[0094] Step 2.4: Sequence reconstruction. For each time t, t = 1, 2, ..., T, the decoder at the current time uses the decoder hidden layer state s of the fused attention t As input, the reconstruction model is used to calculate the reconstruction features of the current moment Get the reconstruction sequence at each moment

[0095] Step 2.5: Residual calculation. Using the original sequence X origin ={x1,x2,…,x T} and refactoring sequence Calculate the reconstruction error E at each moment T};

[0096] Step 2.6: SVM classification. Reconstruction error E = {e1, e2, …, e T} and the original sequence X origin ={x1,x2,…,x T} concatenate to obtain the mixed classification feature vector X = {(x1,e1),(x2,e2),…,(x T ,e T )}, taking the classification feature as input, using SVM to classify the driving behavior to obtain o t , t=1,2,…,T, to realize abnormal driving behavior recognition.

[0097] Step 3: Use the trained Encoder-Decoder model to identify abnormal driving behavior. Predict the specific state prediction values ​​at different times in the future time period through the preprocessed driving behavior time series data, and calculate the residual with the actual value and merge it into X = {(x1, e1), (x2, e2), ..., (x T ,e T )}, and then input it into the SVM model to obtain the abnormal driving behavior recognition result.

[0098] The abnormal driving behavior recognition results are shown in Table 1. The abnormal driving behavior recognition accuracy F1-score of the model proposed in the present invention is 0.717, which is 2.3% higher than the traditional Encoder-Decoder model, indicating that the Attention mechanism helps to improve the model's attention and improve the model's learning effect; the ROC (Receiver Operating Characteristic Curve) curve and PRC (Precision-Recall Curve) curve of the abnormal driving behavior recognition results are shown in Table 1. Figure 5 As shown in the figure, based on the features extracted by the deep network, SVM has a better abnormal driving behavior recognition effect than the Logistic and random forest classification models.

[0099] Table 1

[0100]

[0101]

[0102] The Attention mechanism used in the present invention effectively improves the model's attention and model learning effect, and the use of non-invasive data sets also makes it more applicable and flexible. At the same time, the present invention has the advantages of easy data acquisition, non-invasiveness, and low cost. It not only considers the time correlation of driving behavior, but also considers the differences at different times. It can identify abnormal driving behavior online in an end-to-end manner, and can provide a methodological basis for driving behavior evaluation and safety warning, which is of great significance to the design of intelligent driving systems and traffic safety decision-making.

Claims

1. An online recognition method for abnormal driving behavior based on Encoder-Decoder attention network and LSTM, characterized in that: The specific steps are: Step 1: Obtain mobile phone sensor data, perform preprocessing and normal distribution transformation, and obtain driving behavior time series data; Step 2: construct an Encoder-Decoder attention network and LSTM fusion model and train it. The Encoder-Decoder attention network and LSTM fusion model includes an encoder module, an attention learning module, a decoder module, a sequence reconstruction module, a reconstruction error module and a SVM classifier module; Step 3: Use the trained model to identify abnormal driving behavior. The specific steps are as follows: Step 4.1: Input the time series data into the encoder, and the encoder calculates the hidden layer state at each moment. The specific calculation formula is: In the formula, f is the encoder, is the hidden layer state of the encoder at time t, x t is time series data; Step 4.2: The attention learning module uses the attention weights to adjust all hidden layer states in the encoder. i=1,2,…,T, and perform weighted summation to obtain the semantic vector c at time t t , the specific calculation formula is: In the formula, a t (i) represents the state of the i-th hidden layer of the encoder The hidden layer state of the decoder at time t The weight of The state of the i-th hidden layer of the encoder The hidden layer state of the decoder at time t The weight of is determined by the correlation score between the hidden layer state of the decoder at time t and the hidden layer state of the encoder at time i. The specific calculation formula is: Where W a is the weight matrix, exp(·) is the exponential function, and T is the sequence length; Step 4.3: Concatenate the semantic vector at time t with the original hidden layer state of the decoder to obtain the new hidden layer state of the decoder; Step 4.4: Decoder hidden state s using fused attention t Get the reconstructed sequence The specific calculation formula is: Where W ho is the decoder hidden layer output coefficient matrix, b h0 is the bias term; The hidden state s of the decoder with fused attention t , the specific calculation formula is: Where W c is the transformation matrix, is the hidden layer state of the decoder at time t, c t is the semantic vector at time t; Step 4.5: Calculate the residual between the reconstructed sequence and the original sequence at each moment to obtain the reconstruction error at each moment, and concatenate them to obtain the mixed classification feature vector; Step 4.6: Use the mixed classification feature vector as the input of the SVM classifier, and use the SVM classifier to classify the driving behavior to obtain the classification result to realize abnormal driving behavior recognition.

2. The method for online recognition of abnormal driving behavior based on Encoder-Decoder attention network and LSTM according to claim 1 is characterized in that: The specific steps for preprocessing and normal distribution transformation of mobile phone sensor data are as follows: Step 1.1: Missing data processing: Find missing data records, determine the nature of missing data, and complete or eliminate missing data; Step 1.2: Data normalization: Use the maximum and minimum value normalization method to normalize the data on each feature dimension. The calculation method is as follows: Among them, value i is the i-th value, value max The maximum value of the current column, value min is the minimum value of the current column, and S is the standardized value; Step 1.3: Unbalanced data processing: Use the resampling method to downsample the normal driving data and upsample the abnormal driving data; Step 1.4: Normalization transformation of characteristic distribution: Determine whether the data sample has the characteristics of normal distribution. If not, use the mapping relationship to transform it so that it has the characteristics of normal distribution.

Citation Information

Patent Citations

  • Aberrant driving behavior monitoring and recognizing method and system based on smart mobile terminal

    CN104463244A

  • End-to-end automatic driving behavior decision-making method and system and terminal equipment

    CN113139446A