A Context-Aware Recognition Method for Open-Pit Truck Drivers' Driving Styles
Through feature decoupling module, context embedding module and fusion feature extraction module, the problem of insufficient feature decoupling of GPS data in open-pit mines is solved, and fine-grained recognition and accurate representation of truck driver driving styles is achieved.
Patent Information
- Application Number
- CN202310260288.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-17
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2043-03-17
AI Technical Summary
The prior art is difficult to effectively identify the driving style of truck drivers in open-pit mines, mainly due to insufficient decoupling of GPS data features, resulting in the lack of context information, the inability to fully utilize different categories of information in the trajectory, and the fine-grained driving style is not considered.
The feature decoupling module, context embedding module, fusion feature extraction module and driving style recognition module are used to decompose the GPS trajectory data into motion features, section sequences and global features through road network matching technology and mathematical statistical methods. The context feature extraction is performed by combining bidirectional Transformer and masking mechanism, and the feature representation is enhanced by CrossModal Attention and convolutional layer.
It improves the accuracy and fine-grainedness of driving style recognition for open-pit mine truck drivers, makes full use of trajectory information, enhances the ability to extract context features, and improves the quality of driving style representation.
Smart Images

Figure CN116189097B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of trajectory data mining, and particularly to a context-aware method for identifying the driving style of open-pit mine truck drivers. Background Art
[0002] Transportation activities are an important part of the daily production process in open-pit mines. Affected by the complex production environment in the open-pit mine, drivers need to choose a relatively conservative driving style when driving. However, it is difficult to supervise the driving style manually. Thanks to the development of current in-vehicle sensor technology and wireless communication technology, it is possible to use vehicle data to supervise the driving style.
[0003] In the existing methods for identifying the driving style of open-pit mine truck drivers using vehicle data, it is mainly to collect high-quality OBD-II data from the vehicle CAN-bus (Controller Area Network) and process it using deep models such as recurrent convolutional networks. These methods are quite effective in daily vehicle applications. However, due to the complex production environment in the open-pit mine and the high usage cost of these high-precision sensors, there are various drawbacks in their application in open-pit mines. For example, the sensors in open-pit mine trucks usually have a low sampling frequency and poor data quality. And different from urban roads, the roads in open-pit mines are rough, and high-precision sensors are easily damaged. In order to control costs and effectively identify, existing research tends to use vehicle GPS data to identify the driving style. The driving behavior of the driver is analyzed using the trajectory data provided by the GPS sensors installed in the open-pit mine trucks to extract the unique driving style of each driver.
[0004] The GPS data in open-pit mine trucks is represented in the form of a time series. For data of this series type, a model of the recurrent convolutional network series is usually selected for processing. The conventional sampling frequency of the GPS sensors in open-pit mine trucks is 30 Hz, and a relatively long sequence is usually required to represent a section of the trajectory. However, it is difficult for the models of the recurrent convolutional network series to extract long-term dependencies when processing long sequences. Driving behavior is the reaction of different drivers to the external driving environment and is potentially affected by various complex background factors. Therefore, the context-aware network is an important factor in driving style recognition. Thus, how to decouple features from GPS data, decompose a section of the trajectory into the motion features of the vehicle, the surrounding context features, and the global features with rich semantic information, and use these features as the input of a deep model to improve the representation quality of the driver's driving style is an urgent problem to be solved. Different from the situation on urban roads, in open-pit mines, the road conditions are more complex. Therefore, how to represent the context features of each trajectory is also an unsolved problem. In addition, before driving style recognition, it is necessary to understand that although a driver with a relatively conservative driving style may be aggressive in some situations. Therefore, recognition needs to be carried out at a more fine-grained level, and then the overall driving style of the driver is evaluated. So, the conventional driving style methods are currently difficult to achieve accurate recognition results in the open-pit mine environment. Summary of the Invention
[0005] Technical Problem: The purpose of the present invention is to overcome the deficiencies in the prior art and provide a context-aware method for recognizing the driving style of open-pit mine truck drivers, so as to solve the problems of low recognition accuracy caused by the lack of context information due to the difficulty of decoupling features from GPS data during driving style recognition in open-pit mines, the insufficient consideration of driving environment factors due to the difficulty of accurately representing context information during driving, and the problem of not considering more fine-grained driving styles.
[0006] Technical solution: A context-aware driving style recognition method for open-pit mining truck drivers of the present invention includes a feature decoupling module, a context embedding module, a fused feature extraction module, and a driving style recognition module; First, obtain the GPS trajectory data of the open-pit mining truck and the driving style label of each trajectory, and preprocess all the trajectory data; Then, use the feature decoupling module composed of road network matching technology and mathematical statistics method to decompose the trajectory data into a motion feature sequence, a road segment sequence, and a global feature; Secondly, use the context embedding module, combine the bidirectional Transformer with the masking mechanism, to obtain a context feature sequence, which is used to represent the driving environment of this section of the trajectory; In order to improve the expression ability of this module for context features, before training the entire recognition model, pre-train the context embedding module; Thirdly, use the fused feature extraction module to fuse and extract features from the vehicle motion feature sequence representing driving behavior and the context feature sequence with driving environment information, and obtain the representation reflecting the driving style in the trajectory; Finally, splice the driving style representation generated by each trajectory with the global feature representation, and input it into the driving style recognition module to obtain a fine-grained driving style; The specific steps are as follows:
[0007] Step 1: Preprocess the obtained GPS trajectory data of the open-pit mining truck and the driving style label of each trajectory: including detecting and interpolating abnormal points in the GPS sequence, segmenting the trajectory, and vectorizing the label of each segment of the trajectory to obtain the preprocessed GPS trajectory;
[0008] Step 2: Input the open-pit mining road network information and the preprocessed GPS trajectory into the data feature decoupling module. For each segment of the trajectory, use the road network matching technology in the module to convert the trajectory into a road segment sequence; At the same time, use a mathematical calculation method to calculate the motion feature sequence using the GPS data; In addition, normalize whether it is a working day and the current weather when the trajectory departs into a vector of a fixed length, and pass it through two layers of feedforward neural networks to obtain a global feature representation;
[0009] Step 3: Based on the road segment sequence obtained in Step 2, use the context embedding module for embedding to obtain a context feature sequence; Before training the entire recognition model, use the masking mechanism to pre-train this module to improve the expression ability of this module for context information;
[0010] Step 4: Input the motion feature sequence in Step 2 and the context feature sequence obtained in Step 3 into the fused feature extraction module, use the CrossModalAttention layer in this module to fuse the two sequences, and then use the convolutional layer to extract the local dependence of the sequence and convert it into a vector representation of a fixed length to describe the driving style;
[0011] Step 5: Based on the driving style representation obtained in Step 4 and the global feature representation obtained in Step 2, perform vector concatenation to obtain the final driving style representation; use the driving style recognition module for recognition to obtain the recognition result category; count the driving styles recognized for each trajectory segment in the entire trajectory, and the final output is the proportion of different driving styles of the driver in this trajectory.
[0012] In Step 1, preprocess the obtained GPS trajectory data and the driving style labels of each trajectory, including the detection and interpolation of abnormal points in the GPS sequence, trajectory segmentation, and the vectorization of labels for each trajectory segment;
[0013] The abnormal points in the GPS sequence are mispositioned points caused by poor positioning signals;
[0014] The method for detecting abnormal points in the GPS sequence:
[0015] Based on the obtained GPS trajectory data, define as the trajectory set, where n is the number of trajectories; τ j ={p1, p2,..., p i ..., p |T|} represents a trajectory, and |T| represents the length of τ j trajectory, p i =<lat, lng, alt, ts> represents the trajectory data point in τ j trajectory; where lat, lng, alt, and ts represent longitude, latitude, altitude, and timestamp respectively; first, pre-specify the speed threshold, and then calculate the average speed between two trajectory points in the sequence. When the speed is greater than the threshold, it is marked as an abnormal point;
[0016] The interpolation processing method for abnormal points in the GPS sequence:
[0017] After the detection of abnormal points, use the midpoint of the two trajectory points before and after the abnormal point as the interpolation point for the abnormal point; set p i as the abnormal point, then the latitude calculation formula for the interpolation point p′ i .lat is:
[0018]
[0019] The longitude and altitude of abnormal points in the GPS sequence are calculated in the same way, and the timestamp is retained;
[0020] The trajectory segmentation method:
[0021] Divide the trajectory set into fixed-length driving trajectory segments τ1,..., τ j ,..., τn To avoid excessive loss of information between two adjacent trajectory segments, when intercepting trajectory segments, there is an overlapping length between the two trajectory segments. If the number of trajectory points in each trajectory segment is len, then the overlapping length is len / 2. Subsequently, the obtained trajectory segments will be processed and recognized, that is, the driving style will be recognized for each trajectory segment.
[0022] The described label vectorization method:
[0023] According to the total number of driving styles, digitize the features of the driving style labels; use one-hot encoding to convert the driving style labels into vectors s of length o j s j Only one component in s is 1, and the rest are 0, which is used to represent the category.
[0024] In step 2, input the open-pit mine road network information and GPS trajectories into the feature decoupling module, decouple the GPS data into three types of features, namely motion feature sequences, road segment sequences, and global features, represent the information of a trajectory from three different aspects, and make full use of all the information in the trajectory.
[0025] The calculation process of the described motion feature sequence: Move backward with a sliding window of length L and a step size of L / 2. L is the length of the sequence included in each sliding window. Calculate statistics with the trajectory points within L / 2 as the unit, including the mean, minimum, maximum, standard deviation, and quantiles (25%, 50%, 75%) of the velocity norm, acceleration norm, velocity difference norm, acceleration difference norm, and angular velocity norm, to obtain the motion feature sequence Among them, |m| is the sequence length, and b is the feature dimension. Using this data processing method can better express the driver's driving style;
[0026] The calculation process of the described road segment sequence: Use static information such as road type, speed limit value, and number of lanes to represent the information of the road segment. For each type of feature, perform a normalization operation; if the current number of lanes is num lane and the maximum number of lanes in all data is num all then convert the number of lanes in the original data to num lane / num all Perform such operations on both the road type and the number of lanes. The speed limit information is processed using the maximum-minimum normalization method. Given a trajectory, use the road network matching technology to map the trajectory to the road network, and then calculate the road information in the road network according to the above processing method to obtain the road segment sequence;
[0027] The calculation process of the global features: For each trajectory, the day of the week and the current weather are statistically analyzed. Since such features are invariant for the entire trajectory, these features are defined as global features. For the day-of-week information, two categories are defined, namely weekdays and non-weekdays. For the weather information, the rain, shower, drizzle, snow, fog, wind, and temperature data in the area where the trajectory is generated are obtained. The information of the day of the week and the information in the weather information other than the temperature are normalized in the same way as the lane number information in the context features. For the temperature information, min-max normalization is used. These features are concatenated into a vector of length u and embedded into a fixed-length vector using a feed-forward neural network. Where u is the vector length.
[0028] In step 3, the context embedding module is used for embedding to obtain the context feature sequence. The context embedding module consists of a spatio-temporal encoding layer and a bidirectional Transformer layer. Before the entire recognition model is trained, the mask mechanism is used to pre-train this module to improve the module's expression ability for context information.
[0029] The process of using the context embedding module for embedding is as follows:
[0030] Based on the road segment sequence R obtained in step 2 j ={(r1, ts1), (r2, ts2),..., (r i , ts i )..., (r |n|, ts |n| )}, where each point (r i , ts i ) contains the road segment information and the timestamp. The road segment sequence is input into the spatio-temporal encoding layer. The spatio-temporal encoding layer consists of a road segment encoding block and a time series encoding block. First, the road segment information r i is input into the road segment encoding block to calculate Ω(r i ), where Ω(·) represents the road segment encoding block. Then, the timestamps in the road segment sequence are input into the time series encoding block to obtain ψ(·) represents the time series encoding block. Secondly, the road segment encoding and the time series encoding are added together to obtain the output of the spatio-temporal encoding layer Finally, the encoded sequence is input into the bidirectional Transformer layer to obtain the context features that fuse the adjacent road segments of this road segment. The output of the bidirectional Transformer layer is {z(r1), z(r2),..., z(r |n| )}, where the embedding vector z(r i ) of each road segment contains the information of its adjacent road segments.
[0031] The spatio-temporal encoding layer includes a road segment encoding block and a time series encoding block. The road segment encoding block consists of a fully connected layer, which increases the feature dimension of the original input road segment information from q to d, enhancing its feature expression ability and facilitating the subsequent fusion of information from adjacent road segments. The travel time information in the road segment can effectively represent the road conditions of the road segment. Therefore, it is necessary to embed the time information into the context feature sequence. The time series encoding block uses absolute position encoding and encodes the time series information using real timestamps.
[0032] The pre-training process described above:
[0033] Use the combination of bidirectional Transformer and masking mechanism to pre-train this module. Design a self-supervised strategy for this module. Given a road segment sequence, first input it into the spatio-temporal encoding layer to obtain the encoded sequence. Secondly, randomly select δ% of the points for masking, that is, use the special value [m t to replace the original information. Then input it into the context embedding module. For the masked point r m , use h m in the output sequence for prediction. The calculation formula is:
[0034]
[0035] where FC(·) is a fully connected layer used to map h m to the feature dimension of road segment information. The pre-training objective is to maximize the prediction accuracy, and the calculation formula is as follows:
[0036]
[0037] where θ represents all learnable parameters, and Γ represents all masked points.
[0038] After the pre-training is completed, copy the parameters of this module to the recognition model as the context embedding module.
[0039] The calculation process of the spatio-temporal encoding layer:
[0040] Based on the road segment sequence R j obtained in step 2, calculate the spatio-temporal encoding for each point (r j , ts i , ts i ) in the road segment sequence R
[0041] Ω(r i ) = FC(r i )
[0042] where, The output of the road segment encoding block is denoted as, and FC(·) represents the fully connected layer; then the temporal encoding is calculated, and the calculation formula of the temporal encoding block is:
[0043] ψ(ts i ) = [cos(ω1ts i ), sin(ω2ts i ),..., cos(ω d ts i )]
[0044] where d is the feature dimension of the road segment encoding, denotes the output of the temporal encoding block, and the temporal encoding block uses the real timestamps in the sequence and the learnable parameters {ω1, ω2,..., ω d} for temporal encoding; the output of the road segment encoding block and the output of the temporal encoding block are added together to obtain the output of the spatio-temporal encoding layer, and the calculation formula of the spatio-temporal encoding layer is as follows:
[0045] z′(r i ) = Ω(r i ) + ψ(ts i )
[0046] where, denotes the output of the spatio-temporal embedding layer;
[0047] The above process is calculated for each point in the road segment sequence R j = {(r1, ts1), (r2, ts2)..., (r i , ts i )..., (r |n| , ts |n| )}, and the road segment sequence is transformed into {z′(r1), z′(r2),..., z′(r |n| )}.
[0048] The processing process of the bidirectional Transformer layer:
[0049] Based on the output {z′(r1), z′(r2),..., z′(r |n| )} of the spatio-temporal encoding layer, it is input into the bidirectional Transformer layer, which includes a multi-head self-attention layer and a fully connected neural network layer, and the calculation formula is as follows:
[0050] H c = {h1, h2,..., h |n|} = TransEnc(z′(r1), z′(r2),...Jz′(r |n| ))
[0051] where TransEnc(·) is the bidirectional Transformer operation process; the i-th item h in the output sequence i represents r i the embedded vector with context information in the trajectory, and H c is the context feature sequence of the trajectory.
[0052] In step 4, the fusion feature extraction module includes three half-step feed-forward neural network layers, a CrossModal Attention layer, and a convolutional layer; the motion feature sequence obtained in step 2 and the context feature sequence obtained in step 3 are input into the fusion feature extraction module; first, the two sequences are respectively input into the half-step feed-forward neural network layer to unify the feature dimensions; then the CrossModal Attention layer is used to fuse the motion feature sequence and the context feature sequence, and the global dependencies of the sequences are extracted to obtain the fused feature sequence; then, a convolutional layer is used for further local feature extraction; finally, through a half-step feed-forward neural network, a fixed-length feature vector representation is obtained, and the feature vector representation of each driver is used as the driving style representation of the driver;
[0053] The process of the fusion feature extraction module is as follows:
[0054] Based on the motion feature sequence in step 2 where |m| is the sequence length and q is the feature dimension, and the context feature sequence obtained in step 3 where |n| is the sequence length and d is the feature dimension; the two sequences are respectively input into the half-step feed-forward neural network layer, and the feature dimensions of the two sequences are both transformed into D, and its output is Secondly, the two sequences are input into the CrossModal Attention layer to fuse the two sequences, obtain the correlation between the sequences, and obtain the context-aware driving style sequence Thirdly, s c is input into the convolutional layer to extract the local features in the sequence, and obtain Finally, through another half-step feed-forward neural network layer, the output is obtained
[0055] The characteristics of the feed-forward neural network layer:
[0056] In the entire fusion feature extraction module, two half-step feed-forward neural network layers are used to wrap the CrossModal Attention layer and the convolutional layer. Such a structure is beneficial to the fitting ability of the model, making the extracted features more effective; the feed-forward neural network layer includes two fully connected layers and a Swish activation function;
[0057] The processing process of the feedforward neural network layer:
[0058] Taking the motion feature sequence as an example, it is first input into the first fully connected layer, where the feature dimension is increased from q to D, then passed through the Swish activation function to increase non-linearity, and finally passed through another fully connected layer to obtain the output The calculation formula of the feedforward neural network layer is:
[0059] m′ c = FC(Swish(FC(m c )))
[0060] where FC(·) represents the fully connected layer, Swish(·) represents the activation function, and the calculation formula of the Swish activation function is:
[0061] Swish(x) = x × sigmoid(x)
[0062] where the calculation formula of sigmoid(·) is:
[0063]
[0064] The characteristics of the CrossModal Attention layer: Use CrossModal Attention to obtain the association in the motion feature sequence and the context feature sequence, and dynamically adjust the information in the context feature sequence into the motion feature sequence to generate a fused feature sequence;
[0065] The processing process of the CrossModal Attention layer is:
[0066] First, to maintain the temporal information of the sequence, positional encodings are calculated for the two sequences respectively. Taking the motion feature sequence as an example, the calculation formula:
[0067] m″ c = m′ c + PE(|m|, D)
[0068]
[0069]
[0070] where |m| is the sequence length, D is the feature dimension, PE(·) is the positional encoding process, b is the row index of the sample in the sequence, and i is the column index; after positional encoding, the output The calculation process of the context feature sequence is the same; after the positional encoding is completed, a layer normalization is performed, and then the attention scores are calculated to fuse the information in the context feature sequence into the motion feature sequence; the calculation process is as follows:
[0071]
[0072]
[0073]
[0074]
[0075] Among them, s c is the fused feature sequence, and CM(·) represents the CrossModal Attention calculation process, is a weight matrix of size D×D; (·) T represents the matrix transpose operation, and softmax(·) represents the activation function used, and the calculation formula is:
[0076]
[0077] The convolutional layer mentioned above includes layer normalization, depthwise separable convolution, batch normalization, Swish activation function, pointwise convolution, and Dropout;
[0078] The depthwise separable convolution mentioned above includes one-dimensional depth convolution (1D Depthwise Conv), gated linear unit (GLU), and pointwise convolution (Pointwise Conv); among them, the one-dimensional depth convolution only focuses on the dependencies within each channel, while the pointwise convolution only focuses on the dependencies between channels; combining these two convolutions can achieve the effect of traditional convolution with fewer parameters; the gated linear unit is used to filter information to accelerate the convergence of the model; in addition, batch normalization (BatchNorm) and Swish activation function are also used to facilitate the training of deep models; in order to make up for the lack of the Transformer structure in extracting local features, a convolutional layer is used after the CrossModal Attention layer to enhance the model's ability to learn local features;
[0079] The processing process of the convolutional layer is as follows:
[0080] Based on the fused feature sequence s c output by the CrossModal Attention layer, first a layer normalization is used, and then pointwise convolution is used, and the calculation formula is:
[0081] PointwiseConv(s c ,ef)
[0082] where s c represents the fused feature sequence, ef is the dilation coefficient, and after pointwise convolution, the feature dimension of s c is increased to ef×D, and the output of this unit is
[0083] After that, a gated linear unit is used to filter information, and the features are split in terms of the number of channels and respectively undergo convolutional transformation. One of them passes through a non-linear function sigmoid after transformation and serves as a gating unit to control the output; it is then subjected to a Hadamard product with the other one as the output of this unit where sigmoid is the activation function, and its role is to map the input value to the range from 0 to 1; after that, one-dimensional depth convolution is used to extract the dependencies within a single channel; secondly, batch normalization and the Swish activation function are used to facilitate the training of deep models; where Swish represents the non-linear activation function; thirdly, pointwise convolution is used to further extract local features; after the above operations, a driving style sequence s′ with local features is obtained c ; finally, it passes through a half-step feed-forward neural network to obtain the final output of the fused feature extraction module
[0084] In step 5, the driving style recognition module conducts recognition:
[0085] The driving style recognition module consists of two fully-connected layers. The number of neurons in the first fully-connected layer is u+|m|×d, which are the global feature representation and the flattened length of the driving style s″ c respectively. The number of neurons in the second fully-connected layer is o, that is, the number of driving styles; first, is flattened into a one-dimensional vector, then concatenated with the global feature representation, and after passing through two fully-connected layers, the driving style representation is converted into a one-dimensional vector s of length o p ; finally, the Softmax activation function is used to map the values in this one-dimensional vector to the range from 0 to 1, which is the probability that the driving style of this segment of trajectory belongs to different styles;
[0086] The cross-entropy loss is used as the loss function to train the model. The calculation formula of the cross-entropy loss is as follows:
[0087]
[0088] where s p is the final output of the model, s jis the driving style label for this trajectory segment, exp(·) represents the exponential function with base e, and o is the type of driving style.
[0089] Beneficial effects: Due to the adoption of the above technical solution, the present invention solves the problems of difficult feature decoupling of GPS data when performing driving style recognition in open-pit mines, inability to fully utilize different types of information contained in the trajectory, insufficient extraction of context information contained in the trajectory, and failure to consider more fine-grained driving styles when performing driving style recognition. A feature decoupling module, a context feature embedding module, a fused feature extraction module, and a driving style recognition module are adopted. In the feature decoupling module, trajectory data, road network data, and weather data are used to decouple the information contained in each trajectory segment into three different types of features, representing the information contained in the trajectory and the driving environment from different perspectives respectively; in the context embedding module, advanced encoding means are used to encode the road segment information, and then a bidirectional Transformer and a pre-training method are used to enhance the learning ability of this module for context features; in the fused feature extraction module, CrossModal Attention is used to combine the context feature sequence with the motion feature sequence, and a convolutional module is used to enhance the extraction ability for local features, analyze the driving style in a specific driving scenario, and improve the quality of driving style representation; the driving style recognition module uses two fully connected layers and an activation function for classification operations, outputs the most likely driving style of this trajectory segment, and then conducts statistics to output the driving style reflected by this trajectory. Compared with the prior art, the present invention has obvious advantages in the context-aware open-pit truck driving style recognition method. The main advantages are as follows:
[0090] 1) The feature decoupling module decouples GPS data, road network data, and weather data into three types of features, and more fully excavates the information contained in the trajectory.
[0091] 2) The context embedding module uses a bidirectional Transformer structure to fuse the road segment information of adjacent road segments for each road segment contained in the trajectory, and uses a masking mechanism to pre-train this module, making the extraction of context features by this module more accurate.
[0092] 3) CrossModalAttention is used in the fused feature extraction module, which can process two sequences of different lengths, obtain the association between the two sequences, fuse the information in the sequences, and use a convolutional layer to enhance the extraction ability of this module for local features, improving the quality of the representation of the driver's driving style. Brief Description of the Drawings
[0093] Figure 1 is the flowchart of the method of the present invention.
[0094] Figure 2 This is the structural diagram for identifying driving styles of the present invention. Detailed implementation manners
[0095] The embodiments of the present invention will be further described below in conjunction with the accompanying drawings:
[0096] The context-aware open-pit truck driver driving style recognition method of the present invention includes a feature decoupling module, a context embedding module, a fused feature extraction module, and a driving style recognition module; first, obtain the GPS trajectory data of the open-pit truck and the driving style label of each trajectory, and preprocess all the trajectory data; then, use the feature decoupling module composed of road network matching technology and mathematical statistics methods to decompose the trajectory data into a motion feature sequence, a road segment sequence, and a global feature; secondly, use the context embedding module, in combination with a bidirectional Transformer and a masking mechanism, to obtain a context feature sequence for representing the driving environment of this section of the trajectory; in order to improve the expression ability of this module for context features, pre-train the context embedding module before training the entire recognition model; again, use the fused feature extraction module to fuse and extract features from the vehicle motion feature sequence representing driving behavior and the context feature sequence with driving environment information, and obtain the representation reflecting the driving style in the trajectory; finally, splice the driving style representation generated by each section of the trajectory with the global feature representation, and input it into the driving style recognition module to obtain a fine-grained driving style; the specific steps are as follows:
[0097] Step 1: Preprocess the obtained GPS trajectory data of the open-pit truck and the driving style label of each trajectory: including the detection and interpolation of abnormal points in the GPS sequence, the segmentation of the trajectory, and the vectorization processing of the label of each section of the trajectory, to obtain the preprocessed GPS trajectory;
[0098] The abnormal points in the GPS sequence are mispositioned points caused by poor positioning signals;
[0099] The detection method for abnormal points in the GPS sequence:
[0100] Based on the obtained GPS trajectory data, define as the trajectory set, where n is the number of trajectories; τ j ={p1, p2,..., p i ... p |T|} represents a trajectory, and |T| represents the length of τ j trajectory, p i =<lat, lng, alt, ts> represents τ jTrack data points in the trajectory; where lat, lng, alt, and ts represent longitude, latitude, altitude, and timestamp respectively; first, a speed threshold is specified in advance, then the average speed between two track points in the sequence is calculated, and when the speed is greater than the threshold, it is marked as an abnormal point;
[0101] The interpolation processing method for abnormal points in the GPS sequence:
[0102] After the abnormal point detection is completed, the midpoint of the two track points before and after the abnormal point is used as the interpolation point for the abnormal point; set p i as the abnormal point, then the interpolation point p' i . The latitude calculation formula for lat is:
[0103]
[0104] The longitude and altitude of the abnormal points in the GPS sequence are calculated in the same way, and the timestamp is retained;
[0105] The described trajectory segmentation method:
[0106] The trajectory set is divided into driving trajectory segments τ1,..., τ j ,..., τ n ,..., τ of fixed length according to time. In order to avoid excessive loss of information between two adjacent trajectory segments, when intercepting the trajectory segments, there is an overlapping length between the two trajectory segments. If the number of track points in each trajectory segment is len, then the overlapping part length is len / 2; subsequently, the obtained trajectory segments will be processed and recognized, that is, the driving style is recognized for each trajectory segment;
[0107] The described label vectorization method:
[0108] According to the total number of driving styles, the driving style labels are digitized in terms of features; the one-hot encoding is used to transform the driving style labels into a vector s of length o j , s j in which only one component is 1 and the rest are 0, used to represent the category.
[0109] Step 2: Input the open-pit mine road network information and the preprocessed GPS trajectory into the data feature decoupling module. For each segment of the trajectory, using the road network matching technology in the module, the trajectory is transformed into a road segment sequence; at the same time, using mathematical calculation methods, the motion feature sequence is calculated using GPS data; in addition, whether it is a working day and the current weather when the trajectory starts are normalized into a vector of fixed length, and through two-layer feedforward neural networks, a global feature representation is obtained;
[0110] Input the open-pit mine road network information and GPS trajectories into the feature decoupling module, which decouples the GPS data into three types of features, namely the motion feature sequence, the road segment sequence, and the global feature, representing the information of a trajectory from three different aspects and making full use of all the information in the trajectory.
[0111] The calculation process of the motion feature sequence: Move backward with a sliding window of length L by a step size of L / 2. L is the length of the sequence included in each sliding window. Calculate the statistics with the trajectory points within L / 2 as the unit, including the mean, minimum, maximum, standard deviation, and quantiles (25%, 50%, 75%) of the velocity norm, acceleration norm, velocity difference norm, acceleration difference norm, and angular velocity norm, to obtain the motion feature sequence. Among them, |m| is the sequence length, and b is the feature dimension. Using this data processing method can better express the driver's driving style.
[0112] The calculation process of the road segment sequence: Use static information such as road type, speed limit value, and number of lanes to represent the information of the road segment. For each type of feature, perform a normalization operation. For example, if the current number of lanes is num lane , and the maximum number of lanes in all data is num all , then convert the number of lanes in the original data to num lane / num all . Perform such operations on both the road type and the number of lanes. The speed limit information is processed using the maximum-minimum normalization method. Given a trajectory, use the road network matching technology to map the trajectory to the road network, and then perform operations on the road information in the road network according to the above processing method to obtain the road segment sequence.
[0113] The calculation process of the global feature: Statistically analyze whether it is a working day and the current weather for each trajectory. Since such features are invariant for the entire trajectory, these features are defined as global features. Define two categories for the working day information, namely working day and non-working day. For the weather information, obtain the rain, shower, drizzle, snow, fog, wind, and temperature data in the area where the trajectory is generated. Perform the same normalization operation on the working day information and the information in the weather except for the temperature as the number of lanes information in the context features. Use the maximum-minimum normalization for the temperature information. Concatenate these features into a vector of length u and embed it into a vector of fixed length using a feed-forward neural network. Where u is the vector length.
[0114] Step 3: Based on the road segment sequence obtained in Step 2, use the context embedding module for embedding to obtain the context feature sequence; before training the entire recognition model, use the masking mechanism to pre-train this module to enhance the module's ability to express context information;
[0115] Use the context embedding module for embedding to obtain the context feature sequence. The context embedding module consists of a spatio-temporal encoding layer and a bidirectional Transformer layer; before training the entire recognition model, use the masking mechanism to pre-train this module to enhance the module's ability to express context information;
[0116] The process of using the context embedding module for embedding is as follows:
[0117] Based on the road segment sequence R j ={(r1, ts1), (r2, ts2)..., (r i , ts i )..., (r |n|, ts |n| )}, where each point (r i , ts i ) contains road segment information and time stamps, Input the road segment sequence into the spatio-temporal encoding layer. The spatio-temporal encoding layer consists of a road segment encoding block and a time series encoding block; first, input the road segment information r i into the road segment encoding block to calculate Ω(r i ), where Ω(·) represents the road segment encoding block; then, input the time stamps in the road segment sequence into the time series encoding block to obtain ψ(·) represents the time series encoding block; secondly, add the road segment encoding and the time series encoding to obtain the output of the spatio-temporal encoding layer Finally, input the encoded sequence into the bidirectional Transformer layer to obtain the context features that fuse the adjacent road segments of this road segment. The output of the bidirectional Transformer layer is {z(r1), z(r2),..., z(r |n| )}, where the embedding vector z(r i ) of each road segment contains the information of its adjacent road segments;
[0118] The spatio-temporal encoding layer includes a road segment encoding block and a temporal encoding block. The road segment encoding block consists of a fully-connected layer, which elevates the feature dimension of the original input road segment information from q to d, enhancing its feature expression ability and facilitating the subsequent fusion of information from adjacent road segments. The travel time information in the road segment can effectively represent the road conditions of the road segment. Therefore, it is necessary to embed the time information into the context feature sequence. The temporal encoding block uses absolute position encoding and encodes the temporal information using real timestamps.
[0119] The pre-training process described above:
[0120] Use the combination of bidirectional Transformer and masking mechanism to pre-train this module. Design a self-supervised strategy for this module. Given a road segment sequence, first input it into the spatio-temporal encoding layer to obtain the encoded sequence. Secondly, randomly select δ% of the points for masking, that is, use the special value [m t to replace the original information. Then input it into the context embedding module. For the masked point r m , use the h m in the output sequence for prediction. The calculation formula is:
[0121]
[0122] where FC(·) is the fully-connected layer used to map h m to the feature dimension of the road segment information. The pre-training objective is to maximize the prediction accuracy. The calculation formula is as follows:
[0123]
[0124] where θ represents all learnable parameters and Γ represents all masked points.
[0125] After the pre-training is completed, copy the parameters of this module to the recognition model as the context embedding module.
[0126] The calculation process of the spatio-temporal encoding layer:
[0127] Based on the road segment sequence R j obtained in step 2, calculate the spatio-temporal encoding for each point (r j , ts i , ts i ) in the road segment sequence R
[0128] Ω(r i ) = FC(r i )
[0129] where, Denote the output of the road segment encoding block, and FC(·) represents the fully connected layer; then calculate the temporal encoding, and the calculation formula of the temporal encoding block is:
[0130] ψ(ts i ) = [cos(ω1ts i ),sin(ω2ts i ),...,cos(ω d ts i )]
[0131] where d is the feature dimension of the road segment encoding, denote the output of the temporal encoding block, and the temporal encoding block uses the real timestamps in the sequence and the learnable parameters {ω1, ω2,..., ω d} to perform temporal encoding; add the output of the road segment encoding block and the output of the temporal encoding block to obtain the output of the spatio-temporal encoding layer, and the calculation formula of the spatio-temporal encoding layer is as follows:
[0132] z′(r i ) = Ω(r i ) + ψ(ts i )
[0133] where, denote the output of the spatio-temporal embedding layer;
[0134] The above process calculates each point in the road segment sequence R j = {(r1, ts1), (r2, ts2),..., (r i , ts i )..., (r |n| , ts |n| )}, and converts the road segment sequence into {z′(r1), z′(r z ),..., z′(r |n| )}.
[0135] The processing process of the bidirectional Transformer layer:
[0136] Based on the output {z′(r1), z′(r2),..., z′(r |n| )} of the spatio-temporal encoding layer, input it into the bidirectional Transformer layer, which includes a multi-head self-attention layer and a fully connected neural network layer, and the calculation formula is as follows:
[0137] H c = {h1, h2,..., h |n|} = TransEnc(z′(r1), z′(r2),..., z′(r |n| ))
[0138] where TransEnc(·) is the operation process of the bidirectional Transformer; the i-th item h in the output sequence i represents r i the embedded vector with context information in the trajectory, and H c is the context feature sequence of the trajectory.
[0139] Step 4: Input the motion feature sequence in Step 2 and the context feature sequence obtained in Step 3 into the fusion feature extraction module. Use the CrossModal Attention layer in this module to fuse the two sequences, and then use the convolutional layer to extract the local dependencies of the sequence and convert it into a fixed-length vector representation to describe the driving style;
[0140] The fusion feature extraction module includes three half-step feed-forward neural network layers, a CrossModalAttention layer, and a convolutional layer; input the motion feature sequence obtained in Step 2 and the context feature sequence obtained in Step 3 into the fusion feature extraction module; first, input the two sequences into the half-step feed-forward neural network layer respectively to unify the feature dimensions; then use the CrossModal Attention layer to fuse the motion feature sequence and the context feature sequence to obtain the fused feature sequence; then, use the convolutional layer for further local feature extraction; finally, through a half-step feed-forward neural network, obtain a fixed-length feature vector representation, and use the feature vector representation of each driver as the driving style representation of the driver;
[0141] The process of the fusion feature extraction module is as follows:
[0142] Based on the motion feature sequence in Step 2 where |m| is the sequence length and q is the feature dimension, and the context feature sequence obtained in Step 3 where |n| is the sequence length and d is the feature dimension; input the two sequences into the half-step feed-forward neural network layer respectively, and convert the feature dimensions of the two sequences into D, and its output is Secondly, input the two sequences into the CrossModal Attention layer to fuse the two sequences, obtain the correlation between the sequences, and get the context-aware driving style sequence Thirdly, input s c into the convolutional layer to extract the local features in the sequence, and get Finally, through another half-step feed-forward neural network layer, obtain the output
[0143] The characteristics of the feed-forward neural network layer:
[0144] In the entire fusion feature extraction module, two half-step feedforward neural network layers are used to wrap the CrossModalAttention layer and the convolutional layer. Such a structure is beneficial to the fitting ability of the model, making the extracted features more effective. The feedforward neural network layer contains two fully connected layers and a Swish activation function;
[0145] The processing process of the feedforward neural network layer:
[0146] Taking the motion feature sequence as an example, it is first input into the first fully connected layer, and the feature dimension is increased from q to D. Then, it passes through the Swish activation function to increase the non-linearity. Finally, it passes through another fully connected layer to obtain the output The calculation formula of the feedforward neural network layer is:
[0147] m′ c =FC(Swish(FC(m c )))
[0148] Among them, FC(·) represents the fully connected layer, Swish(·) represents the activation function, and the calculation formula of the Swish activation function is:
[0149] Swish(x)=x×sigmoid(x)
[0150] Among them, the calculation formula of sigmoid(·) is:
[0151]
[0152] The features of the CrossModal Attention layer: Use CrossModal Attention to obtain the association between the motion feature sequence and the context feature sequence, and dynamically adjust the information in the context feature sequence to the motion feature sequence to generate a fused feature sequence;
[0153] The processing process of the CrossModal Attention layer is:
[0154] First, in order to maintain the temporal information of the sequence, positional encodings are calculated for the two sequences respectively. Taking the motion feature sequence as an example, the calculation formula:
[0155] m″ c =m′ c +PE(|m|,D)
[0156]
[0157]
[0158] where |m| is the sequence length, D is the feature dimension, PE(·) is the position encoding process, b is the row index of the sample in the sequence, and i is the column index; after position encoding, the output is The calculation process of the context feature sequence is the same; after position encoding is completed, it goes through a layer normalization, and then the attention scores are calculated to fuse the information in the context feature sequence into the motion feature sequence; the calculation process is as follows:
[0159]
[0160]
[0161]
[0162]
[0163] where s c is the fused feature sequence, and CM(·) represents the CrossModal Attention calculation process, is a weight matrix of size D×D; (·) T represents the matrix transpose operation, and softmax(·) represents the activation function used, and the calculation formula is:
[0164]
[0165] The convolutional layer mentioned above includes layer normalization, depthwise separable convolution, batch normalization, Swish activation function, pointwise convolution, and Dropout;
[0166] The depthwise separable convolution includes one-dimensional depth convolution (1D Depthwise Conv), Gated Linear Unit (GLU), and pointwise convolution (Pointwise Conv); among them, one-dimensional depth convolution only focuses on the dependencies within each channel, while pointwise convolution only focuses on the dependencies between channels; combining these two convolutions can achieve the effect of traditional convolution with fewer parameters; using the Gated Linear Unit to filter information can accelerate the convergence of the model; in addition, batch normalization (BatchNorm) and Swish activation function are also used to facilitate the training of deep models; in order to make up for the lack of the Transformer structure in extracting local features, a convolutional layer is used after the CrossModal Attention layer to enhance the model's ability to learn local features;
[0167] The processing process of the convolutional layer is:
[0168] The fused feature sequence s output based on the CrossModal Attention layer c , first use a layer normalization, and then use pointwise convolution. The calculation formula is:
[0169] PointwiseConv(s c , ef)
[0170] where s c represents the fused feature sequence, ef is the dilation coefficient. After pointwise convolution, the feature dimension of s c is increased to ef×D. The output of this unit is
[0171] After that, use a gated linear unit to filter information, split the features in terms of the number of channels, and perform convolution transformations respectively. One of them passes through a transformation and then through a sigmoid non-linear function as a gating unit to control the output; and perform a Hadamard product with the other as the output of this unit where sigmoid is the activation function, and its role is to map the input value to the interval from 0 to 1; after that, use one-dimensional depth convolution to extract the dependencies within a single channel; secondly, use batch normalization and the Swish activation function to facilitate the training of deep models; where Swish represents the non-linear activation function; thirdly, use pointwise convolution to further extract local features; after the above operations, a driving style sequence s′ with local features is obtained c ; finally, pass through a half-step feedforward neural network to obtain the final output of the fused feature extraction module
[0172] Step 5: Based on the driving style representation obtained in Step 4 and the global feature representation obtained in Step 2, perform vector concatenation to obtain the final driving style representation; use the driving style recognition module for recognition to obtain the recognition result category; for each trajectory segment in the entire trajectory, count the recognized driving styles, and the final output is the proportion of different driving styles of the driver in this trajectory.
[0173] The described driving style recognition module performs recognition:
[0174] The driving style recognition module consists of two fully connected layers. The number of neurons in the first fully connected layer is u+|m|×d, which are the global feature representation and the flattened length of the driving style s″ c respectively. The number of neurons in the second fully connected layer is o, that is, the number of driving styles; first flatten into a one-dimensional vector, then concatenate it with the global feature representation, and after passing through two fully connected layers, the driving style representation is converted into a one-dimensional vector s with a length of o p; Finally, use the Softmax activation function to map the values in the one-dimensional vector to the range of 0 to 1, which is the probability that the driving style of this section of the trajectory belongs to different styles;
[0175] Use cross-entropy loss as the loss function to train the model. The calculation formula of cross-entropy loss is as follows:
[0176]
[0177] where s p is the final output of the model, and s j is the driving style label of this section of the trajectory. exp(·) represents the exponential function with base e, and o is the type of driving style.
[0178] Application of the context-aware open-pit truck driver driving style recognition method of the present invention in a real GPS dataset of open-pit truck transportation: The GPS sampling frequency is about 30 seconds, and the data collection period is 1 month; it includes the following steps:
[0179] Step 1: Obtain the GPS trajectory data of the open-pit truck and perform preprocessing: including the detection and interpolation of abnormal points in the GPS sequence, the segmentation of the trajectory, and the vectorization processing of the labels; during the abnormal point detection process, the set abnormal speed for abnormal point detection is 120 m / s; the trajectory is divided into trajectory segments with a length of 100; the labels are vectorized into vectors with a length of 3, representing three driving styles: aggressive, conservative, and moderate; the preprocessed GPS trajectory is obtained;
[0180] Step 2: Input the open-pit mine road network information and the preprocessed GPS trajectory into the feature decoupling module. For each section of the trajectory, use the road network matching technology in the module to convert the trajectory into a road section sequence; at the same time, use mathematical calculation methods to calculate the motion feature sequence using GPS data; in addition, normalize whether it is a working day and the current weather when the trajectory departs into a vector with a fixed length, and obtain the global feature representation through two layers of feedforward neural networks;
[0181] Step 3: Based on the road section sequence obtained in Step 2, use the context embedding module for embedding to obtain the context feature sequence. Before training the entire recognition model, use the masking mechanism to pre-train this module to improve the module's ability to express context information;
[0182] Step 4: Based on the motion feature sequence in Step 2 and the context feature sequence obtained in Step 3, input these two types of features into the fusion feature extraction module, and use the CrossModal Attention layer and the convolutional layer to fuse the two sequences and extract the global and local dependencies in the sequence to obtain a fixed-length vector representation for the driving style;
[0183] Step 5: Based on the driving style representation obtained in Step 4 and the global feature representation obtained in Step 2, perform vector concatenation to obtain the final driving style representation. Use the driving style recognition module for recognition to obtain the recognition result category, and count the driving styles recognized for each trajectory segment in the entire trajectory. The final output is the proportion of different driving styles of the driver in this trajectory;
[0184] Step 6: Experimental environment and hyperparameter setting:
[0185] The deep learning framework relied on is PyTorch 1.1.0, and the programming language is Python 3.8; all experiments are conducted on a computer equipped with NVIDIA GeForce RTX 4070, and its deep learning acceleration environment is CUDA 10.0 and cuDNN7.0; during training, use the Adam optimizer to optimize the model, with a learning rate of 0.0001, an exponential decay rate of the first moment estimate of 0.9, and an exponential decay rate of the second moment estimate of 0.999; the length len of the trajectory segment is 100, the length o of the label vector is 3, and the sliding window L f is 4; the length |m| of the motion feature sequence is 97, the feature dimension b is 35, the length |n| of the context feature sequence is 100, the feature dimension q is 3, the length u of the global feature vector is 16, d in the encoding block in the spatio-temporal encoding layer is 64, the proportion δ% of the mask is 20%, the feature dimension D of the fused feature sequence is 64, the dilation coefficient ef is 2, and the proportion of Dropout is 0.1; the total number of training iterations is 160K times.
Claims
1. A context-aware method for identifying the driving style of open-pit mining truck drivers, characterized by: It includes a feature decoupling module, a context embedding module, a fused feature extraction module, and a driving style recognition module; First, obtain the GPS trajectory data of open-pit mine trucks and the driving style labels of each trajectory, and preprocess all the trajectory data; Then, use the feature decoupling module composed of road network matching technology and mathematical statistics methods to decompose the trajectory data into motion feature sequences, road segment sequences, and global features; Secondly, use the context embedding module, combine bidirectional Transformer with a masking mechanism to obtain a context feature sequence for representing the driving environment of this section of the trajectory; To improve the expression ability of this module for context features, pre-train the context embedding module before training the entire recognition model; Thirdly, use the fused feature extraction module to fuse and extract features from the vehicle motion feature sequence representing driving behavior and the context feature sequence with driving environment information, and obtain the representation reflecting the driving style in the trajectory; Finally, concatenate the driving style representations generated by each trajectory with the global feature representation, input them into the driving style recognition module, and obtain the fine-grained driving style; The specific steps are as follows: Step 1: Preprocess the obtained GPS trajectory data of open-pit mine trucks and the driving style labels of each trajectory: including the detection and interpolation of abnormal points in the GPS sequence, the segmentation of the trajectory, and the vectorization processing of the labels of each segment of the trajectory, to obtain the preprocessed GPS trajectory; Step 2: Input the open-pit mine road network information and the preprocessed GPS trajectory into the data feature decoupling module. For each segment of the trajectory, use the road network matching technology in the module to convert the trajectory into a road segment sequence; At the same time, use mathematical calculation methods to calculate the motion feature sequence using GPS data; In addition, normalize whether it is a working day and the current weather when the trajectory starts into a vector of a fixed length, and obtain the global feature representation through two layers of feed-forward neural networks; Step 3: Based on the road segment sequence obtained in Step 2, use the context embedding module for embedding to obtain a context feature sequence; Before training the entire recognition model, use the masking mechanism to pre-train this module to improve the expression ability of this module for context information; Step 4: Input the motion feature sequence in Step 2 and the context feature sequence obtained in Step 3 into the fused feature extraction module, use the CrossModal Attention layer in this module to fuse the two sequences, and then use the convolutional layer to extract the local dependencies of the sequences and convert them into a vector representation of a fixed length to describe the driving style; Step 5: Based on the driving style representation obtained in Step 4 and the global feature representation obtained in Step 2, perform vector concatenation to obtain the final driving style representation; Use the driving style recognition module for recognition to obtain the recognition result category; Count the driving styles recognized for each trajectory segment in the entire trajectory, and the final output is the proportion of different driving styles of the driver in this trajectory.
2. The context-aware open-pit truck driver driving style recognition method according to claim 1, characterized in that: In step 1, the obtained GPS trajectory data and the driving style labels of each trajectory are preprocessed, including the detection and interpolation of abnormal points in the GPS sequence, trajectory segmentation, and the vectorization of labels for each segment of the trajectory. The abnormal points in the GPS sequence are mispositioning points caused by poor positioning signals. The method for detecting abnormal points in the GPS sequence: Based on the obtained GPS trajectory data, define as the trajectory set, n is the number of trajectories; τ j ={p1, p2, ..., p i ..., p |T|} represents a trajectory, |T| represents the length of τ j trajectory, p i =<lat, lng, alt, ts> represents the trajectory data point in τ j trajectory; where lat, lng, alt, ts represent longitude, latitude, altitude, and timestamp respectively; First, a speed threshold is specified in advance, and then the average speed between two trajectory points in the sequence is calculated. When the speed is greater than the threshold, it is marked as an abnormal point; The interpolation processing method for abnormal points in the GPS sequence: After the detection of abnormal points, the midpoint of the two trajectory points before and after the abnormal point is used as the interpolation point of the abnormal point; set p i as the abnormal point, then the interpolation point p' i .lat's latitude calculation formula is: The longitude and altitude of abnormal points in the GPS sequence are calculated in the same way, and the timestamp is retained. The method for trajectory segmentation: The set of trajectories is divided into driving trajectory segments τ1,..., τ j ,..., τ n of a fixed length according to time. To avoid excessive loss of information between two adjacent trajectory segments, there is a length of overlap between the two trajectory segments when intercepting the trajectory segments. If the number of trajectory points in each trajectory segment is len, then the length of the overlapping part is len / 2. Subsequently, the obtained trajectory segments will be processed and recognized, that is, the driving style will be recognized for each trajectory segment; The method for label vectorization: Digitize the driving style labels according to the total number of driving styles; use one-hot encoding to convert the driving style labels into a vector s of length o j , s j in which only one component is 1 and the rest are 0, used to represent the category.
3. A method for identifying the driving style of an open-pit mine truck driver based on context awareness according to claim 1, characterized in that: In step 2, the open-pit mine road network information and the GPS trajectory are input into the feature decoupling module, and the GPS data is decoupled into three types of features, namely the motion feature sequence, the road segment sequence, and the global feature, representing the information of a segment of trajectory from three different aspects and making full use of all the information in the trajectory.
4. A method for identifying the driving style of an open-pit mine truck driver with context awareness according to claim 3, characterized in that: The calculation process of the motion feature sequence: A sliding window with a length of L moves backward with a step size of L / 2. L is the length of the sequence included in each sliding window. The trajectory points within L / 2 are used as units to calculate statistics, including the mean, minimum value, maximum value, standard deviation, and quantiles (25%, 50%, 75%) of the velocity norm, acceleration norm, velocity difference norm, acceleration difference norm, and angular velocity norm, to obtain the motion feature sequence. Among them, |m| is the sequence length, and b is the feature dimension. Using this data processing method can better express the driver's driving style. Calculation process of the described road segment sequence: Use static information such as road type, speed limit value, and number of lanes to represent the information of road segments. For each type of feature, perform a normalization operation; if the current number of lanes is num lane , and the maximum value of the number of lanes in all data is num all , then convert the number of lanes in the original data to numlane / num all . Perform such operations on both the road type and the number of lanes. The speed limit information is processed using the maximum-minimum normalization method. Given a trajectory, use the road network matching technology to map the trajectory to the road network, and then perform operations on the road information in the road network according to the above processing method to obtain the road segment sequence; Calculation process of the global features: For each trajectory, statistics are made on whether it is a working day and the current weather. Since such features are invariant for the entire trajectory, these features are defined as global features. For the working day information, two categories are defined, namely working day and non-working day. For the weather information, rain, shower, drizzle, snow, fog, wind, and temperature data in the area where the trajectory is generated are obtained. The information of the working day and the information in the weather information other than the temperature are all normalized in the same way as the lane number information in the context features. For the temperature information, min-max normalization is used. These features are concatenated into a vector of length u and embedded into a vector of a fixed length using a feedforward neural network where u is the vector length.
5. A context-aware open-pit mine truck driver driving style recognition method according to claim 1, characterized in that: In step 3, the context embedding module is used for embedding to obtain the context feature sequence. The context embedding module consists of a spatio-temporal encoding layer and a bidirectional Transformer layer; before the entire recognition model is trained, the module is pre-trained using the masking mechanism to enhance the module's expression ability for context information. The process of using the context embedding module for embedding is as follows: Based on the road segment sequence R obtained in step 2 j ={(r1, ts1), (r2, ts2),...,(r i , ts i )...,(r |n| , ts |n| )}, where each point (r i , ts i ) contains road segment information and a timestamp, Input the road segment sequence into the spatio-temporal encoding layer, which consists of a road segment encoding block and a time series encoding block; First, input the road segment information r i into the road segment encoding block to calculate Ω(r i ), where Ω(·) represents the road segment encoding block; Then, input the timestamps in the road segment sequence into the time series encoding block to obtain ψ(·) represents the time series encoding block; Secondly, add the road segment encoding and the time series encoding to obtain the output of the spatio-temporal encoding layer Finally, input the encoded sequence into the bidirectional Transformer layer to obtain the context features that fuse the adjacent road segments of this road segment. The output of the bidirectional Transformer layer is {z(r1), z(r2),..., z(r |n| )}, where the embedding vector z(r i ) of each road segment contains the information of its adjacent road segments; The spatio-temporal encoding layer includes a road segment encoding block and a temporal encoding block. The road segment encoding block consists of a fully connected layer, which increases the feature dimension of the original input road segment information from q to d, enhancing its feature expression ability and facilitating the subsequent fusion of information of adjacent road segments; the driving time information in the road segment can effectively express the road conditions of the road segment, so the time information needs to be embedded into the context feature sequence. The temporal encoding block uses absolute position encoding to encode the temporal information using the real timestamp. The pre-training process: Pre-train this module by combining a bidirectional Transformer and a masking mechanism, design a self-supervised strategy for this module. Given a sequence of road segments, first input it into the spatio-temporal encoding layer to obtain the encoded sequence; secondly, randomly select δ% of the points for masking, that is, use the special value [m t to replace the original information; then input it into the context embedding module. For the masked point r m , use h m in the output sequence for prediction. The calculation formula is: Among them, FC(·) is a fully connected layer used to map h m to the feature dimension of the road segment information; the pre-training objective is to maximize the prediction accuracy, and the calculation formula is as follows: Among them, θ represents all learnable parameters, and Γ represents all masked points. After the pre-training is completed, the parameters of this module are copied to the recognition model as the context embedding module.
6. A context-aware open-pit mine truck driver driving style recognition method according to claim 5, characterized in that: The calculation process of the spatio-temporal encoding layer: Based on the road segment sequence R obtained in step 2 j , for each point (r j , ts i ) in the road segment sequence R i , calculate the spatio-temporal encoding. First, calculate the road segment encoding. The road segment encoding block consists of a fully connected layer. First, use the road segment encoding block calculation formula: Ω(r i ) = FC(r i ) Among them, represents the output of the road segment coding block, and FC(·) represents the fully connected layer; then the temporal coding is calculated, and the calculation formula of the temporal coding block is: ψ(ts i ) = [cos(ω1ts i ), sin(ω2ts i ),..., cos(ω d ts i )] where d is the feature dimension of the road segment encoding, represents the output of the temporal encoding block, and the temporal encoding block uses the real timestamps in the sequence and learnable parameters {ω1, ω2,..., ω d} for temporal encoding; adding the output of the road segment encoding block and the output of the temporal encoding block to obtain the output of the spatio-temporal encoding layer, and the calculation formula of the spatio-temporal encoding layer is as follows: z′(r i ) = Ω(r i ) + ψ(ts i ) Among them, represents the output of the spatio-temporal embedding layer; The above process calculates each point in the road segment sequence R j ={(r1, ts1), (r2, ts2),...,(r i , ts i ),...,(r |n| , ts |n| ),} and converts the road segment sequence into {z'(r1), z'(r2),..., z'(r |n| )}.
7. A context-aware open-pit truck driver driving style recognition method according to claim 5, characterized in that: The processing process of the bidirectional Transformer layer: Based on the output of the spatio-temporal coding layer {z'(r1), z'(r2),..., z'(r |n| )}, it is input into the bidirectional Transformer layer, which includes a multi-head self-attention layer and a fully connected neural network layer. The calculation formula is as follows: H c = {h1, h2,..., h |n|} = TransEnc(z′(r1), z′(r2),..., z′(r |n| )) where TransEnc(·) is the bidirectional Transformer operation process; the i-th item h in the output sequence i represents r i the embedding vector with context information in the trajectory, and H c is the context feature sequence of the trajectory.
8. A context-aware open-pit mine truck driver driving style recognition method according to claim 1, characterized in that: In step 4, the fusion feature extraction module includes three half-step feed-forward neural network layers, a CrossModal Attention layer, and a convolutional layer; the motion feature sequence obtained in step 2 and the context feature sequence obtained in step 3 are input into the fusion feature extraction module; first, the two sequences are respectively input into the half-step feed-forward neural network layer to unify the feature dimensions; then the CrossModal Attention layer is used to fuse the motion feature sequence and the context feature sequence and extract the global dependencies of the sequence to obtain the fusion feature sequence; then, a convolutional layer is used for further local feature extraction; finally, after a half-step feed-forward neural network, a fixed-length feature vector representation is obtained, and the feature vector representation of each driver is used as the driving style representation of the driver. The process of the fusion feature extraction module is: Based on the motion feature sequence in step 2 where |m| is the sequence length, q is the feature dimension, and the context feature sequence obtained in step 3 where |n| is the sequence length and d is the feature dimension; the two sequences are respectively input into a half-step feedforward neural network layer, and the feature dimensions of the two sequences are both transformed into D, and its output is Secondly, the two sequences are input into a CrossModal Attention layer to fuse the two sequences, obtain the correlation between the sequences, and get a context-aware driving style sequence Thirdly, input s c into a convolutional layer to extract local features in the sequence, and get Finally, after passing through another half-step feedforward neural network layer, the output is obtained The features of the feed-forward neural network layer: In the entire fusion feature extraction module, two half-step feedforward neural network layers are used to wrap the CrossModalAttention layer and the convolutional layer. Such a structure is beneficial to the fitting ability of the model, making the extracted features more effective. The feedforward neural network layer contains two fully connected layers and a Swish activation function. The processing process of the feedforward neural network layer: Taking the motion feature sequence as an example, it is first input into the first fully connected layer to increase the feature dimension from q to D, then passed through the Swish activation function to increase non-linearity, and finally passed through another fully connected layer to obtain the output The calculation formula of the feedforward neural network layer is as follows: m′ c = FC(Swish(FC(m c ))) Among them, FC(·) represents the fully connected layer, Swish(·) represents the activation function, and the calculation formula of the Swish activation function is: Swish(x) = x × sigmoid(x) Among them, the calculation formula of sigmoid(·) is: The features of the CrossModal Attention layer: Use CrossModal Attention to obtain the association in the motion feature sequence and the context feature sequence, dynamically adjust the information in the context feature sequence to the motion feature sequence, and generate a fused feature sequence. The processing process of the CrossModal Attention layer is: First, in order to maintain the temporal information of the sequence, positional encoding is calculated for the two sequences respectively. Taking the motion feature sequence as an example, the calculation formula: m″ c = m′ c + PE(|m|, D) where |m| is the sequence length, D is the feature dimension, PE(·) is the position encoding process, b is the row index of the sample in the sequence, and i is the column index; the output after position encoding The calculation process of the context feature sequence is the same; after position encoding, it goes through a layer normalization, and then the attention scores are calculated to fuse the information in the context feature sequence into the motion feature sequence; the calculation process is as follows: Among them, s c is the fused feature sequence, CM(·) represents the CrossModal Attention calculation process, is the weight matrix of size D×D; (·) T represents the matrix transpose operation, sigmoid(·) represents the activation function used, and the calculation formula is:
9. A context-aware open-pit truck driver driving style recognition method according to claim 8, characterized in that: The convolutional layer includes layer normalization, depthwise separable convolution, batch normalization, Swish activation function, pointwise convolution, and Dropout. The depthwise separable convolution includes one-dimensional depth convolution (1D Depthwise Conv), gated linear unit (Gated Linear Unit, GLU), and pointwise convolution (Pointwise Conv). Among them, one-dimensional depth convolution only focuses on the dependencies within each channel, while pointwise convolution only focuses on the dependencies between channels. Combining these two convolutions can achieve the effect of traditional convolution with fewer parameters. Use the gated linear unit to filter information and accelerate the convergence of the model. In addition, batch normalization (BatchNorm) and Swish activation function are also used to facilitate the training of deep models. In order to make up for the lack of the Transformer structure in extracting local features, a convolutional layer is used after the CrossModal Attention layer to enhance the model's learning ability for local features. The processing process of the convolutional layer is: The fused feature sequence s output based on the CrossModal Attention layer c , first use a layer normalization and then use a pointwise convolution, and the calculation formula is as follows: PointwiseConv(s c ,ef) Among them, s c represents the fused feature sequence, ef is the dilation coefficient, and after pointwise convolution, the feature dimension of s c is increased to ef×D, and the output of this unit is After that, a gated linear unit is used to filter information, the features are split in terms of the number of channels, and are respectively subjected to convolutional transformation. One of them passes through a non-linear function sigmoid after transformation and serves as a gating unit to control the output; it is subjected to a Hadamard product with the other one as the output of this unit where sigmoid is an activation function whose role is to map the input value to the interval from 0 to 1; after that, one-dimensional depth convolution is used to extract the dependencies within a single channel; secondly, batch normalization and the Swish activation function are used to facilitate the training of deep models; where Swish represents a non-linear activation function; again, pointwise convolution is used to further extract local features; after the above operations, a driving style sequence s' with local features is obtained c ; finally, it passes through a half-step feedforward neural network to obtain the final output of the fusion feature extraction module 10. A method for identifying the driving style of an open-pit mine truck driver based on context awareness according to claim 1, characterized in that: In step 5, the driving style recognition module conducts recognition: The driving style recognition module consists of two fully connected layers. The number of neurons in the first fully connected layer is (u + m)×d, which are the global feature representation and the driving style s″ respectively. c The length after flattening, and the number of neurons in the second fully connected layer is o, that is, the number of driving styles. First, is flattened into a one-dimensional vector, and then concatenated with the global feature representation. After passing through two fully connected layers, the driving style representation is converted into a one-dimensional vector s of length o. p Finally, the Softmax activation function is used to map the values in this one-dimensional vector to the interval from 0 to 1, which is the probability that the driving style of this section of trajectory belongs to different styles. Use cross-entropy loss as the loss function to train the model. The calculation formula of cross-entropy loss is as follows: Among them, s p is the final output of the model, s j is the driving style label of this trajectory segment, exp(·) represents the exponential function with base e, and o is the type of driving style.
Citation Information
Patent Citations
Automatic driving track prediction method based on space-time pyramid
CN115049130A
KR20220050758A