Abnormal driving behavior identification method based on multi-modal sensor data fusion
Through the multimodal sensor data fusion method, short-time Fourier transform and feature extraction modules are used to identify abnormal driving behavior, which solves the problems of noise interference and information loss, and achieves higher recognition accuracy and robustness.
Patent Information
- Application Number
- CN202510463178.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-08
AI Technical Summary
The prior art faces features redundancy and information loss caused by noise interference, modal heterogeneity and unreasonable fusion strategies when identifying abnormal driving behaviors, which affects accuracy and robustness.
The multimodal sensor data fusion method is adopted to generate velocity spectrum through short-term Fourier transform, Markov transition field, Gram angle and field conversion, and an asymmetric spatial attention convolution block, asymmetric channel attention convolution block and dual-flow feature fusion module are constructed to extract and fuse features to improve recognition capabilities.
It improves the accuracy and robustness of abnormal driving behavior recognition, can effectively suppress noise interference, reasonably evaluate multimodal information, and improves recognition accuracy.
Smart Images

Figure CN120279532A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of abnormal driving behavior recognition, and particularly relates to a method for recognizing abnormal driving behavior based on multi-modal sensor data fusion. Background Art
[0002] With the development of intelligent transportation systems, the recognition of abnormal driving behavior has become an important research direction for improving road traffic safety. Abnormal driving behavior usually refers to the operations of drivers deviating from normal driving norms, such as fatigue driving, sudden acceleration, sudden braking, and frequent lane changes. The recognition of abnormal driving behavior can not only help drivers adjust their driving styles and reduce accident risks, but also provide auxiliary decision-making for intelligent driving systems and improve the intelligent level of traffic management.
[0003] However, accurately recognizing abnormal driving behavior still faces many challenges. First, due to individual differences, drivers' driving styles and abnormal behavior patterns vary greatly. Different drivers may exhibit different driving behaviors in similar road environments, which poses higher requirements for the generalization ability of the model. Second, the characteristics of abnormal driving behavior are often difficult to directly observe or quantify, and usually need to be inferred by combining multi-modal information fusion, which poses higher challenges to the accuracy of sensors and the effectiveness of algorithms. In addition, some abnormal driving behaviors may occur in a very short time, such as a brief loss of attention caused by fatigue or a sudden emergency braking, which also poses higher requirements for the real-time detection ability of the system.
[0004] Methods for recognizing abnormal driving behavior are mainly divided into two categories: traditional machine learning and deep learning. Traditional methods rely on manual feature extraction and machine learning models (such as support vector machines, random forests, etc.), and it is difficult to recognize complex behaviors. Deep learning methods such as convolutional neural networks and long short-term memory networks can automatically extract features and have gradually become the mainstream methods. However, existing methods still face challenges. On the one hand, vehicle speed data may be affected by various vibrations caused by factors such as vehicle engines and road conditions, and road images may contain objects unrelated to safe driving. How to identify key features and suppress noise interference has become a difficult problem. On the other hand, the fusion of multi-modal data (vehicle speed data and road images) brings information redundancy or loss, affecting accuracy. Summary of the Invention
[0005] In view of the above deficiencies in the prior art, the purpose of the present invention is to provide a method for recognizing abnormal driving behavior based on multi-modal sensor data fusion, which solves the problems of feature redundancy and information loss caused by noise interference, modal heterogeneity, and unreasonable fusion strategies in existing algorithms, and improves the accuracy and robustness of abnormal driving behavior recognition.
[0006] To achieve the above object, the present invention provides an abnormal driving behavior recognition method based on multi-modal sensor data fusion, comprising the following steps: S1. Obtain road pictures in front of the vehicle and vehicle speed data; S2. Apply short-time Fourier transform (STFT), Markov transition field (MTF), Gram angle sum field (GASF), and Gram angle difference field (GADF) to the vehicle speed data for conversion respectively, and splice the conversion results to obtain a speed spectrogram; S3. Construct a feature extraction sub-network based on an asymmetric spatial attention convolution block, an asymmetric channel attention convolution block, a stride convolution layer, a max pooling layer, and a fully connected layer; S4. Based on the feature extraction sub-network, extract features from the road pictures and the speed spectrogram respectively; S5. Perform two-stream feature fusion on the features extracted from the road pictures and the speed spectrogram, and based on the obtained fusion features, use a classifier for recognition to output a classification prediction result of the driving behavior.
[0007] As a preferred solution of the present invention, in S1, the road pictures are obtained based on an in-vehicle camera, and continuous vehicle speed data is obtained based on an in-vehicle motion sensor.
[0008] As a preferred solution of the present invention, in S3, the feature extraction sub-network includes a first part of an asymmetric spatial attention convolution block, a first part of a stride convolution layer, a second part of an asymmetric channel attention convolution block, a second part of a stride convolution layer, a third part of an asymmetric spatial attention convolution block, a third part of a stride convolution layer, a fourth part of an asymmetric channel attention convolution block, a max pooling layer, and a fully connected layer arranged in sequence; Among them, the first part of the asymmetric spatial attention convolution block includes a first part of convolution module one, a first part of joint spatial attention module one, a first part of convolution module two, and a first part of joint spatial attention module two connected in sequence; The second part of the asymmetric channel attention convolution block includes a second part of convolution module one, a second part of multi-scale channel attention module one, a second part of convolution module two, and a second part of multi-scale channel attention module two connected in sequence; The third part of the asymmetric spatial attention convolution block includes a third part of convolution module one, a third part of joint spatial attention module one, a third part of convolution module two, and a third part of joint spatial attention module two connected in sequence; The fourth part of the asymmetric channel attention convolution block includes a fourth part of convolution module one, a fourth part of multi-scale channel attention module one, a fourth part of convolution module two, and a fourth part of multi-scale channel attention module two connected in sequence.
[0009] As a preferred embodiment of the present invention, each of the convolutional modules includes a two-dimensional convolutional layer, a normalization layer, and a ReLu activation layer connected in sequence; each strided convolutional layer uses a two-dimensional convolutional layer; The architectures and feature extraction processes of each joint spatial attention module are as follows: For the input feature tensor, initial convolutional features are extracted through a convolutional layer. Subsequently, the initial convolutional features pass through a max pooling module and an average pooling module respectively, and then through a spatial channel compression module respectively to obtain maximum height features, maximum width features, average height features, and average width features. Subsequently, the maximum height feature and the average height feature are weighted and fused to obtain a fused height feature, the maximum width feature and the average width feature are weighted and fused to obtain a fused width feature. The fused height feature and the fused width feature pass through a weight conversion module to obtain feature weights, and then a multiplication operation is performed with the input feature tensor to obtain spatial attention features; among them, the weight conversion module includes a batch normalization layer and a sigmoid function; The architectures and feature extraction processes of each multi-scale channel attention module are as follows: For the input feature tensor, primary features are extracted through a 3×3 convolutional module, a 5×5 convolutional module, and a 7×7 convolutional module respectively. An addition operation is performed on the three primary features, and then a primary channel feature is obtained through a pooling layer. The primary channel feature passes through three convolutional layers respectively, and deep channel features are extracted through a concatenation operation and a softmax transformation. The three deep channel features and the corresponding primary features are multiplied respectively to obtain three deep features, and an addition operation is performed on the three deep features to obtain channel attention features.
[0010] As a preferred embodiment of the present invention, in S4, the process of feature extraction is as follows: for a road picture, the road picture is input into the first part of the asymmetric spatial attention convolutional block, and shallow features are extracted through the first part of convolutional module one, expressed as: ; In the formula, A represents the road picture; CNN represents the convolutional operation; is the shallow feature extracted by the first part of convolutional module one; the input and output channel sizes of the first part of convolutional module one are 3 and 64; Key features are extracted through the first part of joint spatial attention module one to obtain spatial attention features ; Subsequently, it is input into the first part of convolutional module two. The input and output channel sizes of the first part of convolutional module two are 64 and 64. Convolutional features are obtained after passing through the first part of convolutional module two , and Input the first part and combine it with the spatial attention module 2 to get the spatial attention feature ; Will Enter the first part of the stride convolution layer to reduce the spatial size of the feature to half, expressed as: ; In the formula, Indicates the reduced ; Will Input the second part of the asymmetric channel attention convolution block, first pass through the second part of the convolution module 1, the input and output channel sizes of the second part of the convolution module 1 are 64 and 128, After the second part of the convolution module, the convolution feature is obtained. ; Will Input the second part of the multi-scale channel attention module 1 to obtain the channel attention feature ; Will Input the second part of convolution module 2. The input and output channel sizes of the second part of convolution module 2 are 128 and 128. After the second part of the convolution module 2, the convolution feature is obtained ; Will Input the second part of the multi-scale channel attention module 2 to obtain the channel attention feature ; Will Input the second part of the stride convolution layer to get the convolution feature ; Will Input the third part of the asymmetric spatial attention convolution block, first pass through the third part of the convolution module 1, the input and output channel sizes of the third part of the convolution module 1 are 128 and 256, After the third part of the convolution module, we get the convolution feature ; Will Input the third part into the joint spatial attention module 1 to obtain the spatial attention feature ; Will Input the third part of the convolution module 2, the input and output channel sizes of the third part of the convolution module 2 are 256 and 256, After the third part of the convolution module 2, the convolution feature is obtained ; Will Input the third part and joint spatial attention module 2 to get the spatial attention feature ; Will Input the third part of the stride convolution layer to obtain convolution features ; Input the fourth part of the asymmetric channel attention convolution block. First, pass through the first convolution module of the fourth part. The input and output channel sizes of the first convolution module of the fourth part are 256 and 512, After passing through the first convolution module of the fourth part, obtain convolution features ; ; Input the obtained convolution features into the first multi-scale channel attention module of the fourth part to obtain channel attention features ; ; Input the obtained channel attention features into the second convolution module of the fourth part. The input and output channel sizes of the second convolution module of the fourth part are 512 and 512, After passing through the second convolution module of the fourth part, obtain convolution features ; ; Input the obtained convolution features into the second multi-scale channel attention module of the fourth part to obtain channel attention features ; ; Perform max pooling operation on it to reduce its spatial size to 1×1 to obtain a one-dimensional feature vector M. Use a fully connected layer to reduce the dimension of M to obtain road image features ; ; Similarly, obtain speed features based on the speed spectrogram .
[0011] As a preferred solution of the present invention, in S4, the process of obtaining the spatial attention features is specifically as follows. For : Input it into a convolution layer to obtain preliminary convolution features U. Input U into the max pooling module and the average pooling module respectively. The max pooling module performs max pooling operation on U along the width and height directions, and the average pooling module performs average pooling operation on U along the width and height directions, which is expressed as: ; ; ;In the formula, represents the max pooling operation; represents the average pooling operation; , are the output features of max pooling along the height and width directions respectively; , are the output features of average pooling along the height and width directions respectively; Input , , , are respectively input into the spatial channel compression module. For and , after adjusting the shape of , it is concatenated with along the channel direction, and then compressed through a convolution module to represent the channels, expressed as: ; In the formula, is the feature after adjusting the shape of ; CAT represents the concatenation operation; is the feature after compressing the channels; Divides into width features and height features along the width and height directions, and , are restored to features with the original channel length through a convolutional layer, expressed as: ; In the formula, and are respectively the maximum height feature and the maximum width feature output by the spatial channel compression module; For and , the average height feature and the average width feature are obtained in the same way; Perform weighted fusion to obtain the fused height feature and the fused width feature , expressed as: ; ; In the formula, t and s are two learnable parameters; Input and into the weight conversion module to obtain the corresponding feature weights and , expressed as: ; In the formula, BN is the batch normalization operation; is the sigmoid function; Multiply and by to obtain the spatial attention feature : ; Among them, is of a modified shape and is used to restore its original shape for multiplication operations; The remaining spatial attention features are obtained in the same way.
[0012] As a preferred solution of the present invention, the process of obtaining the channel attention features is specifically as follows. For : Input into the second part of the multi-scale channel attention module one. Respectively pass through the 3×3 convolution module, the 5×5 convolution module, and the 7×7 convolution module to extract the corresponding primary features , , , and then perform an addition operation to obtain the feature : ; Perform an average pooling operation on to obtain the channel feature , respectively pass through three one-dimensional convolutional layers to obtain the corresponding three channel weights , , , and splice them to form the spliced feature : ; Perform softmax transformation on the three dimensions in to generate the transformed channel weights , , , which are the deep channel features. Multiply , , with the corresponding , , to obtain three weighted features , , , which are the deep features. Add , , to obtain the channel attention feature ; The remaining channel attention features are obtained in the same way.
[0013] As a preferred solution of the present invention, in S5, the dual-stream feature fusion module includes a feature extraction module one and a feature extraction module two. Input and Input into the dual-stream feature fusion module to obtain the fused feature , and the process is as follows: First, and are preliminarily fused to obtain the preliminary fused feature D: ; Input D into two parallel branches. The first branch first passes through Feature Extraction Module 1 and then through Feature Extraction Module 2 for feature extraction to obtain the feature ; The second branch first passes through Feature Extraction Module 2 and then through Feature Extraction Module 1 for feature extraction to obtain the feature ; Among them, Feature Extraction Module 1 includes a one-dimensional convolutional layer and a normalization layer connected in sequence; Feature Extraction Module 2 includes a two-dimensional convolutional layer 1, a normalization layer 1, a ReLu activation layer 1, a two-dimensional convolutional layer 2, a normalization layer 2, and a ReLu activation layer 2 connected in sequence. The output channel size of the two-dimensional convolutional layer 1 is , t is the compression factor, and C is the output channel size of the two-dimensional convolutional layer 2; Then, and are fused into the feature : ; Apply a convolutional operation and a sigmoid mapping to obtain the weight : ; In the formula, is the sigmoid function; Next, use to fuse and to obtain the final fused feature : ; In the formula, represents element-wise multiplication.
[0014] As a preferred solution of the present invention, in S5, based on , classify and predict the driving behavior. Before prediction, train the entire network through a cross-entropy loss function. After training is completed, use a linear layer as the classifier. Input into the linear layer to obtain the classification prediction result Pred of the driving behavior.
[0015] The algorithm involved in the present invention can be executed by an electronic device. The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. The above algorithm calculation is implemented by the processor executing the software.
[0016] The beneficial effects of the present invention are as follows: Aiming at the problem that the temporal characteristics of vehicle speed data are difficult to be fully utilized by traditional methods, the present invention applies methods such as Short-Time Fourier Transform (STFT), Markov Transition Field (MTF), Gram Angular Summation Field (GASF), and Gram Angular Difference Field (GADF) to encode vehicle speed into four different types of two-dimensional representations to obtain speed spectrograms, so as to retain comprehensive information and enable the convolutional network to efficiently extract speed change patterns.
[0017] Aiming at the problem that driving behavior involves multiple modal information and it is difficult to effectively focus on key features, the present invention proposes a joint spatial attention module and a multi-scale channel attention module. By establishing the correlation in the spatial and channel dimensions, it selectively focuses on the information closely related to driving behavior and improves the recognition ability of the model.
[0018] Aiming at the problems of information loss and redundancy that may be caused by the fusion of image features and vehicle speed features, the present invention designs a two-stream feature fusion module. This module can capture global and local patterns, reasonably evaluate the importance of different modalities, promote the effective fusion of cross-modal features, and improve the recognition accuracy of abnormal driving behaviors. Brief Description of the Drawings
[0019] Figure 1 is the flow schematic diagram of the present invention; Figure 2 is the modular flow chart of the present invention; Figure 3 is the schematic diagram of the architecture of the asymmetric spatial attention convolutional block of the present invention; Figure 4 is the schematic diagram of the architecture of the asymmetric channel attention convolutional block of the present invention; Figure 5 is the schematic diagram of the architecture of the convolutional module of the present invention; Figure 6 is the schematic diagram of the architecture and feature extraction process of the joint spatial attention module of the present invention; Figure 7 is the schematic diagram of the architecture and feature extraction process of the multi-scale channel attention module of the present invention; Figure 8 is the schematic diagram of the architecture and process of the two-stream feature fusion module of the present invention. Detailed Embodiments
[0020] The following further describes the embodiments of the present invention with reference to the drawings: Embodiment 1: As Figure 1 and Figure 2 shown, a method for identifying abnormal driving behaviors based on multi-modal sensor data fusion includes the following steps: S1. Obtain the road image in front of the vehicle and the vehicle speed data; S2. Apply the short-time Fourier transform (STFT), Markov transition field (MTF), Gram angular sum field (GASF), and Gram angular difference field (GADF) to the vehicle speed data for conversion respectively, and splice the conversion results to obtain a speed spectrogram; S3. Construct a feature extraction sub-network based on an asymmetric spatial attention convolutional block, an asymmetric channel attention convolutional block, a strided convolutional layer, a max pooling layer, and a fully connected layer; S4. Based on the feature extraction sub-network, extract features from the road image and the speed spectrogram respectively; S5. Perform two-stream feature fusion on the features extracted from the road image and the speed spectrogram, and based on the obtained fusion features, use a classifier for recognition and output the classification prediction result of the driving behavior.
[0021] In S1, obtain the road image based on an in-vehicle camera and obtain continuous vehicle speed data based on an in-vehicle motion sensor.
[0022] STFT is a mathematical transform related to the Fourier transform that allows time-frequency analysis of signals; MTF is a stochastic process used to describe the change in pixel intensity in an image; GASF is a method for representing local structural features in an image; GADF is a variant of GASF that obtains a vector field describing the local structural change in an image by calculating the gradient direction difference between adjacent pixels in the image. These four conversion methods are all prior arts.
[0023] In S2, apply STFT to the vehicle speed data to obtain a time-frequency spectrogram; apply MTF to obtain the Markov model features of the speed change; apply GASF to obtain the local structural features of the speed data; apply GADF to obtain the local structural change features of the speed data.
[0024] Splice these four feature sets along the channel dimension to form a multi-channel speed spectrogram. The result of each conversion process is an H×W image, and after splicing, it is a tensor of H×W×4, where H represents the height and W represents the width.
[0025] As Figure 2 shown, in S3, the feature extraction sub-network includes an asymmetric spatial attention convolutional block in the first part, a strided convolutional layer in the first part, an asymmetric channel attention convolutional block in the second part, a strided convolutional layer in the second part, an asymmetric spatial attention convolutional block in the third part, a strided convolutional layer in the third part, an asymmetric channel attention convolutional block in the fourth part, a max pooling layer, and a fully connected layer arranged in sequence; Figures 2-5 In, the prefix "first part" etc. are omitted, and only the overall architecture is shown.
[0026] Among them, asFigure 3 As shown in the figure, the first part of the asymmetric spatial attention convolution block includes a first part convolution module one, a first part joint spatial attention module one, a first part convolution module two, and a first part joint spatial attention module two connected in sequence; As Figure 4 shown in the figure, the second part of the asymmetric channel attention convolution block includes a second part convolution module one, a second part multi-scale channel attention module one, a second part convolution module two, and a second part multi-scale channel attention module two connected in sequence; The third part of the asymmetric spatial attention convolution block includes a third part convolution module one, a third part joint spatial attention module one, a third part convolution module two, and a third part joint spatial attention module two connected in sequence; The fourth part of the asymmetric channel attention convolution block includes a fourth part convolution module one, a fourth part multi-scale channel attention module one, a fourth part convolution module two, and a fourth part multi-scale channel attention module two connected in sequence.
[0027] As Figure 5 shown in the figure, each convolution module includes a two-dimensional convolution layer, a normalization layer, and a ReLu activation layer connected in sequence; each stride convolution layer uses a two-dimensional convolution layer; As Figure 6 shown in the figure, the architecture and feature extraction process of each joint spatial attention module are as follows: For the input feature tensor, initial convolution features are extracted through a convolution layer. Subsequently, the initial convolution features pass through a max pooling module and an average pooling module respectively, and then through a spatial channel compression module respectively to obtain the maximum height feature, the maximum width feature, the average height feature, and the average width feature. Subsequently, the maximum height feature and the average height feature are weighted and fused to obtain the fused height feature, the maximum width feature and the average width feature are weighted and fused to obtain the fused width feature. The fused height feature and the fused width feature pass through a weight conversion module to obtain feature weights, and then a multiplication operation is performed with the input feature tensor to obtain the spatial attention feature; among them, the weight conversion module includes a batch normalization layer and a sigmoid function; As Figure 7 shown in the figure, the architecture and feature extraction process of each multi-scale channel attention module are as follows: For the input feature tensor, primary features are extracted through a 3×3 convolution module, a 5×5 convolution module, and a 7×7 convolution module respectively. An addition operation is performed on the three primary features, and then a primary channel feature is obtained through a pooling layer. The primary channel feature passes through three convolution layers respectively, and deep channel features are extracted through a concatenation operation and a softmax transformation. The three deep channel features and the corresponding primary features are respectively multiplied to obtain three deep features, and an addition operation is performed on the three deep features to obtain the channel attention feature.
[0028] In S4, the process of feature extraction is as follows: for a road image, the road image is input into the first part of the asymmetric spatial attention convolutional block, and shallow features are extracted through the first convolutional module one, which is expressed as: ; In the formula, A represents the road image; CNN represents the convolution operation; is the shallow feature extracted by the first convolutional module one; the input and output channel sizes of the first convolutional module one are 3 and 64, and the channel dimension is increased to 64 dimensions; Key features are extracted through the first joint spatial attention module one to obtain spatial attention features ; Subsequently, It is input into the first convolutional module two. The input and output channel sizes of the first convolutional module two are 64 and 64, After passing through the first convolutional module two, convolutional features are obtained. Then, is input into the first joint spatial attention module two to obtain spatial attention features ; Then, is input into the first stride convolutional layer, and the spatial size of the feature is reduced by half (the same applies to the other stride convolutional layers), which is expressed as: ; In the formula, represents the reduced ; Then, is input into the second part of the asymmetric channel attention convolutional block. First, it passes through the second convolutional module one. The input and output channel sizes of the second convolutional module one are 64 and 128, After passing through the second convolutional module one, convolutional features are obtained; Then, is input into the second multi-scale channel attention module one to obtain channel attention features ; Then, is input into the second convolutional module two. The input and output channel sizes of the second convolutional module two are 128 and 128, After passing through the second convolutional module two, convolutional features are obtained; Then, is input into the second multi-scale channel attention module two to obtain channel attention features ; Then, Input the second part of the stride convolution layer to obtain convolution features ; Input the third part of the asymmetric spatial attention convolution block. First, pass through the first convolution module of the third part. The input and output channel sizes of the first convolution module of the third part are 128 and 256, After passing through the first convolution module of the third part, obtain convolution features ; ; Input it into the first joint spatial attention module of the third part to obtain spatial attention features ; ; Input it into the second convolution module of the third part. The input and output channel sizes of the second convolution module of the third part are 256 and 256, After passing through the second convolution module of the third part, obtain convolution features ; ; Input it into the second joint spatial attention module of the third part to obtain spatial attention features ; ; Input it into the stride convolution layer of the third part to obtain convolution features ; ; Input it into the asymmetric channel attention convolution block of the fourth part. First, pass through the first convolution module of the fourth part. The input and output channel sizes of the first convolution module of the fourth part are 256 and 512, After passing through the first convolution module of the fourth part, obtain convolution features ; ; Input it into the first multi-scale channel attention module of the fourth part to obtain channel attention features ; ; Input it into the second convolution module of the fourth part. The input and output channel sizes of the second convolution module of the fourth part are 512 and 512, After passing through the second convolution module of the fourth part, obtain convolution features ; ; Input it into the second multi-scale channel attention module of the fourth part to obtain channel attention features ; ; Perform a max pooling operation on it to reduce its spatial size to 1×1, obtaining a one-dimensional feature vector M. Use a fully connected layer to reduce the dimension of M to obtain road image features ; ; Similarly, obtain speed features based on the speed spectrogram .
[0029] In S4, the process of acquiring spatial attention features is as follows: : Will The input convolution layer obtains the preliminary convolution feature U, and U is input into the maximum pooling module and the average pooling module respectively to obtain comprehensive feature information. The maximum pooling module performs the maximum pooling operation on U along the width and height directions, and the average pooling module performs the average pooling operation on U along the width and height directions, which can be expressed as: ; ; In the formula, Represents the maximum pooling operation; represents the average pooling operation; , They are the output features of the maximum pooling along the height and width directions respectively; , They are the output features of average pooling along the height and width directions respectively; Will , , , Input into the spatial channel compression module respectively, for , ,Adjustment After the shape The concatenation is performed along the channel direction, and then a convolution module is passed to compress the channel, which is expressed as: ; In the formula, Yes Adjustment Features after shape; CAT represents the splicing operation; It is the characteristic after compression channel; Adjustment The shape is for splicing, change the shape from C×1×W to C×W×1, so that The shape after splicing is C×(H+W)×1; Will Divided into width features along the width and height directions and height characteristics ,Will , The features restored to the original channel length through the convolution layer are expressed as: ; In the formula, , They are the maximum height feature and the maximum width feature output by the spatial channel compression module, respectively; For 、 , the average height feature and the average width feature are obtained in the same way; Weighted fusion is performed to obtain the fused height feature and the fused width feature , which is expressed as: ; ; In the formula, t and s are two learnable parameters, ranging from 0 to 1; Bring 、 into the weight conversion module to obtain the corresponding feature weights 、 , which is expressed as: ; In the formula, BN is the batch normalization operation; is the sigmoid function; Bring 、 and are multiplied to obtain the spatial attention feature : ; Among them, is after shape change, which is used to restore its original shape for multiplication operation; The remaining spatial attention features are obtained in the same way.
[0030] The process of obtaining the channel attention feature is specifically as follows. For : Bring into the second part of the multi-scale channel attention module one, and respectively pass through the 3×3 convolution module, the 5×5 convolution module and the 7×7 convolution module to extract the corresponding primary features 、 、 , and then an addition operation is performed to obtain the feature : ; Perform average pooling operation on to obtain the channel feature , and respectively pass through three one-dimensional convolution layers to obtain the corresponding three channel weights 、 , , splice to form a splicing feature : ; For , perform softmax transformation on the three dimensions respectively to generate the transformed channel weights , , , which are the deep channel features. Multiply , , with the corresponding , , to obtain three weighted features , , , which are the deep features. Add , , to get the channel attention feature ; The remaining channel attention features are obtained in the same way.
[0031] In S5, the two-stream feature fusion module includes a feature extraction module one and a feature extraction module two. Input and into the two-stream feature fusion module to get the fusion feature . The process is as follows: First, and are preliminarily fused to obtain the preliminary fusion feature D: ; Input D into two parallel branches. The first branch first passes through the feature extraction module one, and then through the feature extraction module two for feature extraction to obtain the feature ; The second branch first passes through the feature extraction module two, and then through the feature extraction module one for feature extraction to obtain the feature ; Among them, the feature extraction module one includes a one-dimensional convolutional layer and a normalization layer connected in sequence; the feature extraction module two includes a two-dimensional convolutional layer one, a normalization layer one, a ReLu activation layer one, a two-dimensional convolutional layer two, a normalization layer two, and a ReLu activation layer two connected in sequence. The output channel size of the two-dimensional convolutional layer one is , t is the compression factor, and C is the output channel size of the two-dimensional convolutional layer two; Then, and are fused into the feature : ; Apply convolution operation and sigmoid mapping to obtain weights : ; In the formula, is the sigmoid function; Utilize to fuse and to obtain the final fused feature : ; In the formula, represents element-wise multiplication.
[0032] The architecture and process of the two-stream feature fusion module are as Figure 8 shown.
[0033] In S5, based on classify and predict driving behaviors. Before prediction, train the entire network through the cross-entropy loss function to reduce the difference between the predicted class of the model (the prediction model composed of the entire network in this embodiment) and the true class, and increase the generalization ability of the model. After training is completed, use a linear layer as the classifier, input into the linear layer to obtain the classification prediction result Pred of driving behaviors.
[0034] The cross-entropy loss function Loss is expressed as: ; In the formula, CE represents cross-entropy, which is used to measure the difference between the true value and the predicted value; Label is the true class of driving behaviors.
[0035] The verification process is as follows, Adopt the proposed driving behavior recognition method to train and test on three groups of datasets, namely the highway dataset, the secondary road dataset, and all datasets containing the above two datasets. In the experiment, road images are cropped to a size of 224×224, and the STFT, MTF, GASF, and GADF methods are applied to convert the speed time series data into 2D representations. To facilitate data input, normalize the sizes of these four 2D representations to 257×200 and combine them into a 4-channel tensor. During the network training process, the epoch (number of training rounds) is set to 200, and the Adam optimizer is used to optimize the network. During model training, the learning rate changes. Specifically, the learning rate for the first 50 rounds is set to 2.5e -4 , the learning rate for 50 - 150 rounds is set to 2.5e -5 , and the learning rate for the last 50 rounds is set to 2.5e -6。All experiments were conducted on a server equipped with a GeForce RTX 3090 GPU.
[0036] Table 1 Performance comparison between different driving behavior recognition methods on the highway dataset
[0037] Table 2 Performance comparison between different driving behavior recognition methods on the secondary road dataset
[0038] Table 3 Performance comparison between different driving behavior recognition methods on all datasets
[0039] Tables 1 - 3 show the performance comparison between different driving behavior recognition methods on three datasets. Among the performance evaluation metrics, the F1 score reflects considering both precision and recall to measure the performance of the model in the classification task, ACC reflects the proportion of samples correctly predicted in the total number of samples, Pre reflects the proportion of samples of a certain category that actually belong to that category, Rec reflects the proportion of samples belonging to a certain category that are correctly predicted as that category, and Spe reflects the proportion of samples correctly predicted as negative in all true negative categories. Compared with other models, the model proposed in this embodiment achieves better recognition performance.
[0040] Example 2: The difference between this embodiment and Embodiment 1 is that the classifier uses an adaptive multi-scale spatio-temporal attention classifier instead of a non-linear layer.
[0041] The adaptive multi-scale spatio-temporal attention classifier includes a spatio-temporal feature decoupling module, a multi-branch gated attention module, and a dynamic weight fusion layer connected in sequence: Input the fused feature into the spatio-temporal feature decoupling module, and respectively extract the temporal dynamic feature and the spatial structure feature through the parallel temporal dilated convolution path and the spatial deformable convolution path, where the temporal dilated convolution path includes three cascaded dilated convolution layers with dilation factors of 2, 4, and 8 respectively, and the spatial deformable convolution path includes two cascaded deformable convolution layers; Input and into the multi-branch gated attention module, which includes: Cross-modal interaction unit: Use the cross-attention mechanism to establish the dynamic association between the temporal feature and the spatial feature, and generate the interaction weight matrix ; Multi-scale gated unit: For and Perform three-scale pyramid pooling separately, dynamically select the contribution degrees of features at each scale through a gating function, and generate gated weighted features. and ; Feature recombination layer: After concatenating with , through tensor concatenation, generate enhanced features through depthwise separable convolution. ; Input the into the dynamic weight fusion layer, which includes: Feature bifurcation module: Decompose the feature into K sub-feature vectors along the channel dimension. ; Dynamic kernel generator: Generate corresponding adaptive convolution kernels based on each sub-feature vector, whose size is dynamically adjusted according to the statistical characteristics of the input feature; Kernel fusion unit: Perform matrix decomposition operation on each and the global context feature to generate the final classification weight matrix ; Among them, the global context feature is a feature vector with global perception ability formed by aggregating the statistical information of all spatial positions and channel dimensions; Adaptive classifier: Calculate the classification score through , and output the prediction result using a softmax function with adjustable temperature coefficient.
Claims
1. An abnormal driving behavior recognition method based on multi-modal sensor data fusion, characterized in that It includes the following steps: S1. Obtain the road picture in front of the vehicle and the vehicle speed data; S2. Apply the short-time Fourier transform (STFT), Markov transition field (MTF), Gram angle sum field (GASF), and Gram angle difference field (GADF) to the vehicle speed data respectively for conversion, and splice the conversion results to obtain a speed spectrogram; S3. Construct a feature extraction sub-network based on an asymmetric spatial attention convolution block, an asymmetric channel attention convolution block, a stride convolution layer, a max pooling layer, and a fully connected layer; S4. Based on the feature extraction sub-network, extract features from the road picture and the speed spectrogram respectively; S5. Perform two-stream feature fusion on the features extracted from the road picture and the speed spectrogram, and based on the obtained fusion features, use a classifier for recognition to output the classification prediction result of the driving behavior.
2. The abnormal driving behavior recognition method based on multi-modal sensor data fusion according to claim 1, wherein In the above-mentioned S1, the road picture is obtained based on an in-vehicle camera, and the continuous vehicle speed data is obtained based on an in-vehicle motion sensor.
3. The abnormal driving behavior recognition method based on multi-modal sensor data fusion according to claim 1, characterized in that, In the above-mentioned S3, the feature extraction sub-network includes a first part of an asymmetric spatial attention convolution block, a first part of a stride convolution layer, a second part of an asymmetric channel attention convolution block, a second part of a stride convolution layer, a third part of an asymmetric spatial attention convolution block, a third part of a stride convolution layer, a fourth part of an asymmetric channel attention convolution block, a max pooling layer, and a fully connected layer arranged in sequence; Among them, the first part of the asymmetric spatial attention convolution block includes a first part of a convolution module one, a first part of a joint spatial attention module one, a first part of a convolution module two, and a first part of a joint spatial attention module two connected in sequence; The second part of the asymmetric channel attention convolution block includes a second part of a convolution module one, a second part of a multi-scale channel attention module one, a second part of a convolution module two, and a second part of a multi-scale channel attention module two connected in sequence; The third part of the asymmetric spatial attention convolution block includes a third part of a convolution module one, a third part of a joint spatial attention module one, a third part of a convolution module two, and a third part of a joint spatial attention module two connected in sequence; The fourth part of the asymmetric channel attention convolution block includes a fourth part of a convolution module one, a fourth part of a multi-scale channel attention module one, a fourth part of a convolution module two, and a fourth part of a multi-scale channel attention module two connected in sequence.
4. The abnormal driving behavior recognition method based on multi-modal sensor data fusion according to claim 3, characterized in that, Each of the above-mentioned convolution modules includes a two-dimensional convolution layer, a normalization layer, and a ReLu activation layer connected in sequence; each of the stride convolution layers uses a two-dimensional convolution layer; The architecture and feature extraction process of each joint spatial attention module are as follows: For the input feature tensor, preliminary convolutional features are extracted through the convolutional layer. Subsequently, the preliminary convolutional features pass through the max pooling module and the average pooling module respectively, and then through the spatial channel compression module respectively to obtain the maximum height feature, the maximum width feature, the average height feature, and the average width feature. Subsequently, the maximum height feature and the average height feature are weighted and fused to obtain the fused height feature, and the maximum width feature and the average width feature are weighted and fused to obtain the fused width feature. The fused height feature and the fused width feature pass through the weight conversion module to obtain the feature weights, and then perform a multiplication operation with the input feature tensor to obtain the spatial attention feature; among them, the weight conversion module includes a batch normalization layer and a sigmoid function; The architecture and feature extraction process of each multi-scale channel attention module are as follows: For the input feature tensor, primary features are extracted through the 3×3 convolutional module, the 5×5 convolutional module, and the 7×7 convolutional module respectively. An addition operation is performed on the three primary features, and then the primary channel feature is obtained through the pooling layer. The primary channel feature passes through three convolutional layers respectively, and the deep channel feature is extracted through the concatenation operation and the softmax transformation. The three deep channel features and the corresponding primary features perform a multiplication operation respectively to obtain three deep features, and an addition operation is performed on the three deep features to obtain the channel attention feature.
5. The abnormal driving behavior recognition method based on multi-modal sensor data fusion according to claim 4, characterized in that, In S4 described above, the feature extraction process is as follows. For the road image, the road image is input into the first part of the asymmetric spatial attention convolutional block, and the shallow feature is extracted through the first part of the convolutional module one, which is expressed as: ; Wherein, A represents a road picture; CNN represents a convolution operation; is the shallow feature extracted by the first convolution module 1; the input and output channel sizes of the first convolution module 1 are 3 and 64; Extract key features through the first part of the joint spatial attention module one to obtain spatial attention features ; Subsequently, input the first part of convolutional module 2, where the input and output channel sizes of the first part of convolutional module 2 are 64 and 64, and obtain convolutional features after passing through the first part of convolutional module 2 , and input into the first part of joint spatial attention module 2 to obtain spatial attention features ; Feed the first part of the stride convolutional layer, reducing the spatial dimension of the feature by half, which is expressed as: ; In the formula, represents the reduced ; Input to the second part of the asymmetric channel attention convolution block, first through the second part convolution module 1. The input and output channel sizes of the second part convolution module 1 are 64 and 128, and after passing through the second part convolution module 1, convolution features are obtained ; Input into the second part of the multi-scale channel attention module 1 to obtain channel attention features ; Input into the second part of the second convolutional module. The input and output channel sizes of the second part of the second convolutional module are 128 and 128, and obtain convolutional features through the second part of the second convolutional module ; Input into the second part of the multi-scale channel attention module two to obtain channel attention features ; Input into the second part of the stride convolutional layer to obtain convolutional features ; Enter the third part of the asymmetric spatial attention convolutional block, first pass through the third part convolutional module 1. The input and output channel sizes of the third part convolutional module 1 are 128 and 256, After passing through the third part convolutional module 1, the convolutional feature is obtained ; Input into the third part of the combined spatial attention module 1 to obtain spatial attention features ; Feed into the second convolution module of the third part. The input and output channel sizes of the second convolution module of the third part are 256 and 256, After passing through the second convolution module of the third part, convolution features are obtained ; Input into the third part, the combined spatial attention module two, to obtain the spatial attention feature ; Input into the third part of the stride convolution layer to obtain convolution features ; Input into the fourth part of the asymmetric channel attention convolutional block. First, it passes through the first convolutional module of the fourth part. The input and output channel sizes of the first convolutional module of the fourth part are 256 and 512, and after passing through the first convolutional module of the fourth part, convolutional features are obtained ; Input into the fourth part, the multi-scale channel attention module one, to obtain channel attention features ; Input it into the second convolutional module of the fourth part. The input and output channel sizes of the second convolutional module of the fourth part are 512 and 512, After passing through the second convolutional module of the fourth part, convolutional features are obtained ; Input into the second multi-scale channel attention module of the fourth part to obtain channel attention features ; For perform a max pooling operation to reduce its spatial dimension to 1×1, obtaining a one-dimensional feature vector M, and use a fully connected layer to reduce the dimension of M to obtain the road image feature ; Similarly, velocity features are obtained based on the velocity spectrogram .
6. The abnormal driving behavior recognition method based on multi-modal sensor data fusion according to claim 5, characterized in that In S4 described above, the process of obtaining the spatial attention feature is specifically as follows. For : Input into the convolutional layer to obtain the initial convolutional feature U. Then input U into the max-pooling module and the average-pooling module respectively. The max-pooling module performs max-pooling operations on U along the width and height directions, and the average-pooling module performs average-pooling operations on U along the width and height directions, which is expressed as: ; ; In the formula, represents the max pooling operation; represents the average pooling operation; , are the output features of max pooling along the height and width directions respectively; , are the output features of average pooling along the height and width directions respectively; Input , , , into the spatial channel compression module respectively. For , , adjust the shape of and splice it with along the channel direction, and then compress the channel through a convolution module, which is expressed as: ; In the formula, is the feature after adjusting the shape; CAT represents the splicing operation; is the feature after compressing the channel; Divide into width features along the width and height directions and height features , and divide 、 into features with the original channel length through a convolutional layer, expressed as: ; In the formula, and are the maximum height feature and the maximum width feature output by the spatial channel compression module, respectively; For and the average height feature is obtained in the same way and the average width feature ; Perform weighted fusion to obtain the fused height feature and the fused width feature , which is expressed as: ; ; In the formula, t and s are two learnable parameters; Fetch and the corresponding feature weights from the input weight conversion module and , expressed as: ; wherein, BN is a batch normalization operation; is the sigmoid function; Multiply and with to obtain the spatial attention feature : ; Among them, is of a changed shape and is used to restore its original shape for multiplication operation; The remaining spatial attention features are obtained in the same way.
7. The abnormal driving behavior recognition method based on multi-modal sensor data fusion according to claim 5, characterized in that The process of obtaining the channel attention feature is specifically as follows. For : Input into the second multi-scale channel attention module 1, respectively through the 3×3 convolution module, 5×5 convolution module and 7×7 convolution module, and extract the corresponding primary features , , . Then perform an addition operation to obtain the feature : ; Pair Perform average pooling operation to obtain channel features , Respectively pass through three one-dimensional convolutional layers to obtain the corresponding three channel weights , , , and splice them to form a spliced feature : ; For perform softmax transformation on the three dimensions respectively to generate the transformed channel weights , , , which are the deep channel features. Multiply , , with the corresponding , , to obtain three weighted features , , , which are the deep features. Add , , to get the channel attention feature ; The remaining channel attention features are obtained in the same way.
8. The abnormal driving behavior recognition method based on multi-modal sensor data fusion according to claim 5, wherein In the above-mentioned S5, the dual-stream feature fusion module includes a first feature extraction module and a second feature extraction module. and are input into the dual-stream feature fusion module to obtain the fused feature . The process is as follows: Combine and to perform preliminary fusion to obtain the preliminary fusion feature D: ; Input D into two parallel branches. The first branch first passes through Feature Extraction Module 1 and then through Feature Extraction Module 2 for feature extraction to obtain features ; The second branch first passes through Feature Extraction Module 2 and then through Feature Extraction Module 1 for feature extraction to obtain features ; Among them, the first feature extraction module includes a one-dimensional convolutional layer and a normalization layer connected in sequence; the second feature extraction module includes a first two-dimensional convolutional layer, a first normalization layer, a first ReLu activation layer, a second two-dimensional convolutional layer, a second normalization layer, and a second ReLu activation layer connected in sequence. The output channel size of the first two-dimensional convolutional layer is , where t is the compression factor and C is the output channel size of the second two-dimensional convolutional layer; Combine with to form a feature : ; Apply convolution operation and sigmoid mapping to obtain weights : ; In the formula, is the sigmoid function; Utilize Fuse and to obtain the final fused feature : ; In the formula, represents element-wise multiplication.
9. The abnormal driving behavior recognition method based on multi-modal sensor data fusion according to claim 1, characterized in that, In the above-mentioned S5, based on classify and predict driving behaviors. Before prediction, train the entire network through the cross-entropy loss function. After training is completed, use a linear layer as the classifier, and input it into the linear layer to obtain the classification prediction result Pred of driving behaviors.
Citation Information
Cited By
Vibration event classification method and system based on distributed optical fiber sensing
CN121302000A
Driver disability emergency takeover method based on multi-modal perception and intelligent decision
CN121375836A