A method for recognizing a driver's steering intention by fusing in-vehicle and out-of-vehicle images
By integrating interior and exterior vehicle imagery with a lightweight model trained on pre-existing weights, the method effectively predicts driver intent, addressing the limitations of current ADAS systems and ensuring timely vehicle control.
Patent Information
- Application Number
- CN202410801942.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-20
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-06-20
AI Technical Summary
The existing ADAS system has too many model parameters, insufficient real-time and accuracy in driver steering intention recognition, and it is difficult to take into account both the lightweight and real-time model while maintaining accuracy, resulting in difficulty in practical application.
The driver's steering intention recognition method that fuses inside and outside the vehicle images is adopted, and the in-vehicle feature extraction module is constructed using a 3D residual network. Combined with the feature fusion of optical flow image processing and attention mechanism, the exterior features are extracted through the ConvLSTM and DSMax modules, and the freezing training strategy is used to reduce the amount of model parameters.
Reduce the number of model parameters while maintaining accuracy, improve the real-time and accuracy of driver steering intention recognition, adapt to practical application needs, and improve the accuracy of driver motivation behavior recognition and the degree of lightweight model.
Smart Images

Figure CN118823439B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent driving assistance, and particularly relates to a method for recognizing a driver's steering intention by fusing in-vehicle and out-of-vehicle images. Background Art
[0002] In recent years, with the continuous increase in the number of automobiles, traffic jams on the road have gradually deteriorated, and traffic accidents occur frequently. Based on this, people have started to look for new solutions, and the Advanced Driver Assistance System (ADAS) is considered to be one of the important ways to reduce accident risks and improve road safety.
[0003] ADAS is an in-vehicle intelligent technology that usually collects vehicle information using sensors and cameras and provides warnings and suggestions to the driver. However, the current ADAS equipped in vehicles only provides warnings based on the running state of the vehicle, lacking environmental background information and driver behavior data, resulting in warning delays and unable to give the driver enough time to respond. If ADAS can predict the driver's intention a few seconds in advance, the system can be prepared before the driver manipulates the vehicle and help the driver control the vehicle in time to avoid collisions.
[0004] In recent years, many studies on the driver's steering intention have emerged, covering various algorithms such as generative models, discriminant models, and deep learning. These studies use a variety of data to model driving behavior, including vehicle dynamics data, driver state data, and external road environment data, etc. Many studies analyze the driver's steering intention using the vehicle's motion state, but the vehicle running state is essentially the result of driving behavior and reflects the effect of the driver's steering intention on the vehicle. Therefore, the vehicle motion state is not suitable for directly recognizing the driver's steering intention, especially in the case of need for early prediction. There are also some studies that use the method of multi-feature splicing for research. Although this multi-feature method has shown certain effectiveness in the research, simple feature splicing cannot well handle the correlation between multiple features. In addition, some methods attempt to use 3D convolutional neural networks to process time series data or data with a time dimension such as videos. However, this method results in too many parameters of the model, restricting its performance in practical applications. In addition, previous studies mostly carried out research based on manually encoded features, and the parameter scales of most models are huge, which does not meet the requirements of the lightweight standard in actual vehicle applications and is difficult to be put into actual use.
[0005] Therefore, how to reduce the number of model parameters while maintaining accuracy, taking into account the real-time performance, accuracy, and model lightweight of driver steering intention recognition, so as to ensure the efficiency and accuracy in actual application, has become an urgent problem to be solved at present. Summary of the Invention
[0006] In view of the deficiencies of the above-mentioned prior art, the present invention provides a method for identifying a driver's steering intention by fusing in-vehicle and out-of-vehicle images, which can reduce the number of model parameters while maintaining accuracy, taking into account the real-time performance, accuracy and model lightweight of driver steering intention recognition, and ensuring the efficiency and accuracy during actual application.
[0007] To solve the above technical problems, the present invention adopts the following technical solutions:
[0008] A method for identifying a driver's steering intention by fusing in-vehicle and out-of-vehicle images, comprising the following steps:
[0009] S1. Construct a steering intention recognition model for identifying a driver's steering intention based on in-vehicle and out-of-vehicle images; the steering intention recognition model includes an in-vehicle feature extraction module, an out-of-vehicle feature extraction module, an in-vehicle and out-of-vehicle feature fusion module, and a classifier recognition module;
[0010] The in-vehicle feature extraction module is constructed based on a 3D residual network and is used to obtain in-vehicle driver intention features according to the in-vehicle driver image; the image processing module is used to convert the out-of-vehicle road environment video into an optical flow map of the relative motion information between the vehicle and other traffic participants; the out-of-vehicle feature extraction module is used to obtain out-of-vehicle road environment features according to the converted optical flow map;
[0011] The in-vehicle and out-of-vehicle feature fusion module is used to fuse the driver intention features and the road environment features to obtain comprehensive judgment features; the classifier recognition module is used to identify the driver's motivation behavior according to the comprehensive judgment features, and the motivation behavior includes going straight, left lane change, left turn, right lane change, and right turn;
[0012] S2. Train the steering intention recognition model;
[0013] S3. Obtain in-vehicle and out-of-vehicle images during actual driving, and perform driver intention recognition through the trained steering intention recognition model for intelligent assisted driving.
[0014] Compared with the prior art, the present invention has the following beneficial effects:
[0015] 1. When using this method, after constructing and training the steering intention recognition model, during actual driving, after obtaining the images inside and outside the vehicle, the in-vehicle feature extraction module will obtain the driver's intention features inside the vehicle based on the driver's image inside the vehicle; the out-of-vehicle feature extraction module will obtain the road environment features outside the vehicle based on the converted optical flow map. Then, the in- and out-of-vehicle feature fusion module will fuse the driver's intention features and the road environment features to obtain comprehensive judgment features; and then the classifier recognition module will identify the driver's motive behavior (going straight, left lane change, left turn, right lane change, or right turn) based on the comprehensive judgment features. Compared with the simple feature splicing in the prior art, this method can better handle the correlation between multiple features.
[0016] Compared with the conventional technology, this method starts from the driver's cognitive perspective, combines the driver's behavior information inside the vehicle and the road environment information outside the vehicle, and uses an attention mechanism-based method to fuse the driver's intention features and the road environment features to accurately identify the driver's steering intention. Experimental verification shows that the complementary in- and out-of-vehicle information of this method helps the recognition of the driver's steering intention, and its comprehensive performance is better than other similar methods. It can ensure the accuracy of the driver's motive behavior.
[0017] 2. The image processing module will convert the road environment video outside the vehicle into an optical flow map of the relative motion information between the vehicle and other traffic participants. Since the optical flow represents the changes in the image and captures the motion information of the target in consecutive frames of the video. Therefore, the optical flow image is used instead of the traditional semantic segmentation image. Such a processing method can ensure the effectiveness of the obtained road environment features. Further ensure the accuracy of the recognition of the driver's motive behavior.
[0018] 3. The in-vehicle feature extraction module is constructed based on a 3D residual network. Compared with the prior art that uses a 3D convolutional neural network to process time series data, it can take into account the lightweight of the model.
[0019] In summary, this method can reduce the number of model parameters while maintaining accuracy, taking into account the real-time performance, accuracy, and model lightweight of the driver's steering intention recognition, and ensuring the efficiency and accuracy during its actual application.
[0020] Preferably, in S2, when training the steering intention recognition model, the model freezing training strategy is applied, freezing the parameters of the in-vehicle feature extraction module and the out-of-vehicle feature extraction module. While utilizing the knowledge learned from the previous weight file, the model can focus resources and attention on the subsequent part of the training.
[0021] Compared with building a model from scratch, the model built by this method is more robust. It can not only improve the generalization ability of the model, but also reduce the risk of overfitting, and better meet the requirements of actual application scenarios. Moreover, by adopting the freeze training strategy, the classification accuracy can be improved while the model is lightweighted.
[0022] Preferably, the in-vehicle feature extraction module is constructed by introducing dilated convolution in Stage 1 of 3DresNet; wherein, the Stage 1 includes a convolutional layer and a max pooling layer, which are used to extract and fuse the spatial features in three-dimensional data; 3DresNet is 3DResNet 34, 3DResNet 50 or 3D ResNet 101.
[0023] With such a setting of the in-vehicle feature extraction module, while lightweighting the model, it can ensure capturing a wider range of driver behaviors. By introducing holes between convolutional kernels, the receptive field can be expanded without losing the size of the feature map, better capturing temporal and spatial information, thus avoiding information loss. In this way, the utilization rate of feature extraction can be improved. Dilated convolution combines local and global information. Incorporating dilated convolution in the first stage can improve the utilization rate of key feature extraction and reduce information loss. It can also control the model complexity. Introducing dilated convolution in the early stage of the model can limit the number of network parameters and computational complexity, avoiding excessive additional computational burdens in subsequent stages.
[0024] Preferably, the out-of-vehicle feature extraction module includes ConvLSTM and four DSMax modules connected in sequence; ConvLSTM is used to predict optical flow features according to the optical flow map; the four DSMax modules are used to further extract and abstract the key features in the optical flow features predicted by ConvLSTM as road environment features.
[0025] With such a setting, ConvLSTM inherits the convolutional characteristics of CNN, can effectively extract the features of high-dimensional data information, and at the same time integrates the memory ability of LSTM, can handle relatively complex time-series data, and has a good processing effect on time-series data such as speech and video. Based on DSC, the DSMax structure is constructed, which can operate on the optical flow features predicted by ConvLSTM to prepare for subsequent feature fusion. In this way, it can not only process complex spatio-temporal data more efficiently, but also effectively capture the deep feature relationships hidden in the time series.
[0026] Preferably, the DSMax module is obtained by applying a max pooling layer to the output of the DSC; the DSC consists of a depthwise convolution (DWConv) and a pointwise convolution (PWConv); for each channel of the input feature map, the DWConv performs a convolution operation using a convolution kernel, and then the outputs of each channel are concatenated to obtain the final output; the PWConv is a 1×1 convolution used to perform channel fusion on the feature map output by the DWConv to obtain effectively integrated feature information.
[0027] Ordinary convolution is prone to gradient problems in deep networks, which can affect the training and inference speed of the model. In contrast, the depthwise separable convolution (DSC) can quickly implement the training and inference of the model, with a small amount of computation, and can achieve a more concise model structure.
[0028] Preferably, the number of channels of four sequentially connected DSMax modules are 64, 128, 256, and 512 respectively, and the sizes of the output feature maps in the spatial dimension are 37×59, 12×20, 4×7, and 1×2 respectively.
[0029] Preferably, the working process of the in-vehicle and out-of-vehicle feature fusion module includes: horizontally concatenating the driver intention feature and the road environment feature along the column dimension to obtain a combined feature; subsequently, calculating the attention weight using a linear layer and a sigmoid activation function; then performing an element-wise multiplication operation on the combined feature and the attention weight to obtain the final feature representation as the comprehensive judgment feature.
[0030] In this way, the dynamic fusion of the features inside and outside the cockpit is realized, enabling the model to adaptively focus on the feature information of different parts and perform fusion.
[0031] Preferably, the formula for calculating the attention weight using a linear layer and a sigmoid activation function is:
[0032]
[0033] where W is the weight of the linear layer, X is the input feature vector, and b is the bias.
[0034] In this way, the Sigmoid activation function is applied to map the value of WX + b to the range [0, 1] to obtain the attention weight. W and b are learned during the model training process, and by continuously adjusting the values of the weight matrix to better model the input data, the attention weight suitable for a specific task is learned.
[0035] Preferably, the image processing module is FlowNet 2.0; the classifier recognition module is a classifier composed of a fully connected layer (FC) and a Softmax layer.
[0036] Preferably, the loss function of the in-vehicle feature extraction module is the cross-entropy loss function;
[0037]
[0038] In the formula, C represents the number of categories, and p i is the value of the i-th category in the true label, and q i is the predicted probability;
[0039] The loss functions of the out-of-vehicle feature extraction module and the in-vehicle and out-of-vehicle feature fusion module are the mean square error loss functions;
[0040]
[0041] In the formula, N is the number of samples, is the predicted value, and y i is the target value.
[0042] In this way, the robustness and performance of model training can be ensured. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to make the objectives, technical solutions, and advantages of the invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings, where:
[0044] Figure 1 is the structural schematic diagram in the embodiment;
[0045] Figure 2 is the schematic diagram of the steering intention recognition model in the first embodiment;
[0046] Figure 3 is the network structure schematic diagram of different 3D ResNet series models in the first embodiment;
[0047] Figure 4 is the structural schematic diagram of ConvLSTM in the first embodiment;
[0048] Figure 5 is the detailed structural schematic diagram of DSMax in the first embodiment;
[0049] Figure 6 is the structural flowchart of D-CAF in the first embodiment;
[0050] Figure 7 is the partial sample schematic diagram of the Brain4Cars and Zenodo datasets in the second embodiment;
[0051] Figure 8 is the schematic diagram of the confusion matrix of the overall experiment in the second embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0052] The following is a further detailed description through specific embodiments:
[0053] Embodiment 1
[0054] As Figure 1 、 Figure 2 shown, in this embodiment, a method for identifying a driver's steering intention by fusing in-vehicle and out-of-vehicle images is disclosed, including the following steps:
[0055] S1. Construct a steering intention recognition model for identifying the driver's steering intention based on in-vehicle and out-of-vehicle images; the steering intention recognition model includes an in-vehicle feature extraction module, an out-of-vehicle feature extraction module, an in-vehicle and out-of-vehicle feature fusion module, and a classifier recognition module;
[0056] In-vehicle feature extraction module (A-3DResNet 50)
[0057] The in-vehicle feature extraction module is constructed based on a 3D residual network and is used to obtain in-vehicle driver intention features according to the in-vehicle driver image;
[0058] Specifically, when implemented, the in-vehicle feature extraction module is constructed by introducing dilated convolution in Stage 1 of 3DresNet; wherein, the Stage 1 includes a convolutional layer and a max-pooling layer for extracting and fusing spatial features in three-dimensional data; 3DresNet is 3DResNet 34, 3DResNet 50, or 3D ResNet 101.
[0059] 3D ResNet performs excellently when processing video data, especially in human action recognition tasks. 3D ResNet can handle video sequences of various lengths and sizes and is suitable for video data input in the cockpit. In the field of deep learning, common 3D ResNets include network architectures with different numbers of layers such as 3D ResNet 18, 3D ResNet 34, 3D ResNet 50, 3D ResNet 101, and 3D ResNet 152. The design of each network architecture has its specific advantages and limitations, which become more prominent when performing specific tasks. For the shallow network 3D ResNet 18, its representation and learning capabilities are limited, and the complexity of the features it can learn is relatively low, resulting in poor performance when dealing with complex and high-dimensional data. The deep network 3D ResNet 152 is significantly higher than other networks in terms of model complexity and computational resource requirements, leading to delays in the training and inference processes and more stringent requirements for hardware resources. Therefore, to balance the representation ability, model complexity, and computational resource requirements, this method selects 3D ResNet 34, 3D ResNet 50, and 3D ResNet 101 in the 3D ResNet series for research. These networks not only maintain a high representation ability but are also relatively easy to train and have moderate requirements for computational resources. This helps to meet the lightweight requirements of the driver's steering intention recognition task while ensuring the model performance, making it more suitable for actual autonomous driving systems.
[0060] Figure 3 The network structures of 3D ResNet 34, 3D ResNet 50, and 3D ResNet 101 are shown. These network structures are very similar overall and can be divided into five stages. Stage 1 includes convolutional layers and max pooling layers, which are mainly used to extract and fuse spatial features in three-dimensional data. Stages 2 to 5 are composed of different numbers of residual blocks. Figure 3 The gray modules in are BasicBlock residual blocks, which are commonly used to construct shallow networks in the ResNet series. The blue modules are Bottleneck residual blocks, which are used to construct deep networks in the ResNet series. The Bottleneck residual blocks reduce the number of parameters and can effectively reduce the model's requirements for video memory and computational resources.
[0061] Although 3D ResNet performs well in processing three-dimensional data, it still encounters performance bottlenecks when capturing long-term temporal information and processing large-sized video frames. Therefore, in this method, dilated convolutions are introduced in Stage 1 of 3DResNet 34, 3DResNet 50, and 3DResNet 101 to construct the driver intention feature extractor A-3DResNet series network for the interior of the cockpit.
[0062] By introducing holes between the convolutional kernels, the receptive field can be enlarged without losing the size of the feature map, better capturing temporal and spatial information, thus avoiding information loss. Such a design mainly considers two points:
[0063] 1) Improve the utilization rate of feature extraction: Dilated convolution combines local and global information. Incorporating dilated convolution in the first stage can improve the utilization rate of key feature extraction and reduce information loss.
[0064] 2) Control the model complexity: Introducing dilated convolution in the early stage of the model can limit the number of network parameters and computational complexity, avoiding excessive additional computational burdens in subsequent stages.
[0065] Applied to the task of driver steering intention recognition, in this embodiment, A-3DResNet 50 is used as the driver intention feature extractor for the interior of the cockpit. The model input is 16 frames uniformly sampled from each second of video, and the frame image size is 112×112. Figure 2 A-3DResNet 50 in [reference] is the in-vehicle feature extraction module in this method. Table 1 provides the parameter information of each module in the model. Among them, the "Output_size" column shows the shape of the feature map output by each module. The first number is the number of channels of the feature map, and the subsequent numbers describe the size of the feature map in the spatial dimension.
[0066] Table 1 Parameter information of each module of the A-3DResNet 50 model.
[0067]
[0068] Out-of-vehicle feature extraction module (ConvLSTM-4DSMax)
[0069] The image processing module is used to convert the out-of-vehicle road environment video into an optical flow map of the relative motion information between the vehicle and other traffic participants; specifically in implementation, the image processing module is FlowNet 2.0.
[0070] The out-of-vehicle feature extraction module is used to obtain the out-of-vehicle road environment features based on the converted optical flow map.
[0071] The external vehicle feature extraction module includes ConvLSTM and four sequentially connected DSMax modules; ConvLSTM is used to predict optical flow features based on the optical flow map; the four DSMax modules are used to further extract and abstract the key features in the optical flow features predicted by ConvLSTM as road environment features.
[0072] ConvLSTM (Convolutional Long Short Term Memory) is a deep learning model that combines CNN and RNN and is commonly used to process spatio-temporal data in the form of data frames. This method designs the DSMax module based on depthwise separable convolution, and combines the ConvLSTM and DSMax modules to obtain the road environment feature extractor ConvLSTM-4DSMax outside the cockpit.
[0073] Optical flow characterizes the changes in an image and captures the motion information of objects in consecutive frames of a video. Therefore, instead of using traditional semantic segmentation images, optical flow images are used to generate an optical flow map of the relative motion information between the vehicle and other traffic participants through FlowNet 2.0, and then the optical flow features are predicted through ConvLSTM.
[0074] ConvLSTM inherits the convolutional characteristics of CNN and can effectively extract the features of high-dimensional data information. At the same time, it combines the memory ability of LSTM and can handle relatively complex time-series data, and has a good processing effect on time-series data such as speech and video. The working principle of ConvLSTM is shown in the following 5 formulas:
[0075] i t =σ([H t-1 ,X t ,C t-1 *W i +b i );
[0076] f t =σ([H t-1 ,X t ,C t-1 *W f +b f );
[0077]
[0078] o t =σ([H t-1 ,X t ,C t *W o +b o );
[0079] H t= o t ⊙ tanh(C t );
[0080] where i t represents the input gate, f t represents the forget gate, C t represents the cell output, o t represents the output gate, H t represents the hidden state. σ represents the Sigmoid function, * represents the convolution operation, ⊙ represents the element-wise multiplication, and represents the element-wise addition.
[0081] In Figure 4 the ConvLSTM structure is shown. Where h i,j is the hidden state, C i,j is the cell state, i represents the time step, and j represents the layer number. The optical flow frame images are uniformly sampled between 1 second and 5 seconds and used as the input to the ConvLSTM, with an input size of 112×176.
[0082] Regarding DSMax
[0083] Ordinary convolution is prone to gradient problems in deep networks, which can affect the training and inference speed of the model. In contrast, depthwise separable convolution (DSC) can quickly implement the training and inference of the model, with a small amount of computation, and can achieve a more concise model structure. Therefore, this method constructs the DSMax structure based on DSC to operate on the optical flow features predicted by the ConvLSTM in order to prepare for subsequent feature fusion.
[0084] DSC consists of depthwise convolution (DWConv) and pointwise convolution (PWConv). In the DWConv stage, for each channel of the input feature map, a convolution kernel is used for convolution operation, and then the outputs of each channel are concatenated to obtain the final output. PWConv is actually a 1×1 convolution and plays two roles in DSC:
[0085] 1) Adjust the number of channels: Independent DWConv cannot change the number of output channels, so PWConv is used to adjust the number of output channels.
[0086] 2) Channel fusion: PWConv is used to perform channel fusion operations on the feature maps output by DWConv, thereby effectively integrating feature information.
[0087] The DSMax module applies a max pooling layer to the output of the DSC, reducing the feature map dimension while retaining important features. Its detailed structure is as shown in Figure 5 shown below.
[0088] Combining ConvLSTM with four DSMax modules results in the road environment feature extractor ConvLSTM-4DSMax outside the cockpit. Among them, ConvLSTM is used for predicting the optical flow features outside the cockpit, and the four DSMax modules further extract and abstract the key features in the data. Such a design can not only process complex spatio-temporal data more efficiently but also effectively capture the deep feature relationships hidden in the time series.
[0089] Figure 2 The model structure of ConvLSTM-4DSMax (i.e., the external vehicle feature extraction module) is shown in, and the parameter information of each module in ConvLSTM-4DSMax is shown in Table 2. Among them, the "Output_size" column shows the output size of each module, and the first number is the number of channels.
[0090] Table 2 Parameter information of each module in ConvLSTM-4DSMax.
[0091]
[0092] Driver-vehicle interior and exterior feature fusion module (D-CAF module)
[0093] The driver's intention features (Interior features) inside the cockpit are obtained through A-3DResNet 50 (the interior vehicle feature extraction module), and the road environment features (Exterior features) outside the cockpit are obtained through ConvLSTM-4DSMax (the external vehicle feature extraction module). For these two types of features, this method designs a dynamic combined-feature attention fusion module (Dynamic Combined-Feature Attention Fusion, D-CAF), Figure 6 showing the structural process of D-CAF.
[0094] The interior and exterior cockpit features are horizontally concatenated according to the column dimension to obtain a new feature vector, that is, the combined feature (Combined Feature). Subsequently, the attention weights (Attention_weights) are calculated using a linear layer and a sigmoid activation function, and this process is shown in the following formula:
[0095]
[0096] Where, W is the weight of the linear layer, X is the input feature vector, and b is the bias. The Sigmoid activation function is applied to map the value of WX + b to the range [0, 1] to obtain the attention weights. W and b are learned during the model training process. By continuously adjusting the values of the weight matrix, the model can better model the input data, thereby learning the attention weights suitable for specific tasks. The combined features are multiplied element-wise with the attention weights to obtain the final feature representation (Weighted Feature), that is, the comprehensive judgment feature. The entire process realizes the dynamic fusion of the features inside and outside the cockpit, enabling the model to adaptively focus on the feature information of different parts and perform fusion.
[0097] S2. Train the steering intention recognition model.
[0098] In a multi-network model, each network is designed for a specific task, and the final model needs to comprehensively consider all networks to achieve the overall goal. In actual operation, each network performs well on individual tasks, but when integrated into a unified framework, the overall effect is not satisfactory. The main reason is that the features to be considered in the joint task are too complex, resulting in the model being difficult to capture all key information during the overall learning process, thus affecting the performance.
[0099] To solve this problem, this method applies the model freezing training strategy in the overall model training. Freeze the parameters of A-3DResNet 50 and ConvLSTM. While utilizing the knowledge learned from the previous weight file, the model can focus resources and attention on training the subsequent parts. Compared with building the model from scratch, the model built by this method is more robust, which can not only improve the generalization ability of the model but also reduce the risk of overfitting and better adapt to the requirements of actual application scenarios.
[0100] Specifically, during implementation, the loss function of the in-vehicle feature extraction module is the cross-entropy loss function;
[0101]
[0102] In the formula, C represents the number of categories, p i is the value of the i-th category in the true label, and q i is the predicted probability;
[0103] The loss functions of the out-of-vehicle feature extraction module and the in-vehicle and out-of-vehicle feature fusion module are the mean squared error loss functions;
[0104]
[0105] In the formula, N is the number of samples, is the predicted value, and y i is the target value.
[0106] The loss function is used to measure the difference between the true intention of the driver and the intention recognized by the model. The smaller the difference, the better the mapping ability of the model from input to output. In this way, the robustness and performance of model training can be ensured.
[0107] S3. During actual driving, images inside and outside the vehicle are acquired, and the trained steering intention recognition model is used to recognize the driver's intention for intelligent assisted driving.
[0108] Using this method, after constructing and training the steering intention recognition model, during actual driving, after acquiring the images inside and outside the vehicle, the in-vehicle feature extraction module will obtain the in-vehicle driver intention features based on the driver image inside the vehicle; the out-of-vehicle feature extraction module will obtain the out-of-vehicle road environment features based on the converted optical flow map. Then, the in-vehicle and out-of-vehicle feature fusion module will fuse the driver intention features and the road environment features to obtain comprehensive judgment features; and then the classifier recognition module will recognize the driver's motivation behavior (going straight, left lane change, left turn, right lane change or right turn) according to the comprehensive judgment features. Compared with the simple feature splicing in the prior art, this method can better handle the correlation between multiple features. Compared with the conventional technology, this method starts from the driver's cognitive perspective, combines the in-vehicle driver behavior information and the out-of-vehicle road environment information, and uses an attention mechanism-based method to fuse the driver intention features and the road environment features to accurately identify the driver's steering intention. Experimental verification shows that the complementarity of in-vehicle and out-of-vehicle information in this method helps the recognition of the driver's steering intention, and the comprehensive performance is better than other similar methods. It can ensure the accuracy of the driver's motivation behavior.
[0109] In addition, the image processing module will convert the out-of-vehicle road environment video into an optical flow map of the relative motion information between the vehicle and other traffic participants. Since the optical flow represents the changes in the image and captures the motion information of the target in consecutive frames of the video. Therefore, the optical flow image is used instead of the traditional semantic segmentation image. Such a processing method can ensure the effectiveness of the obtained road environment features. Further ensure the accuracy of the recognition of the driver's motivation behavior. And, the in-vehicle feature extraction module is constructed based on a 3D residual network. Compared with using a 3D convolutional neural network to process time series data in the prior art, it can take into account the lightweight of the model.
[0110] This method can reduce the number of model parameters while maintaining accuracy, taking into account the real-time performance, accuracy and model lightweight of driver steering intention recognition, and ensuring the efficiency and accuracy during its actual application.
[0111] Embodiment 2
[0112] To better illustrate the effectiveness of the steering intention recognition model in this method, the following experimental description is carried out.
[0113] The proposed method was evaluated on the Brain4cars and Zenodo datasets, and some examples of the two datasets are as Figure 7 shown.
[0114] Dataset
[0115] (1) Brain4Cars dataset: It contains videos recorded simultaneously for the driver (1920px × 1088px, 30fps) and for the road (720px × 480px, 30fps). Five types of maneuvering behaviors are defined, namely going straight, left lane change, left turn, right lane change, and right turn, covering the driving behaviors before the driver's actual operation. In addition to the videos, the dataset also contains some information extracted from external cameras and GPS, such as the lane number of the car, the number of road lanes, etc.
[0116] (2) Zenodo dataset: According to the recording standard of the Brain4Cars dataset, it provides videos from two perspectives, for the driver (1048px × 810px, 30fps) and for the road (1620px × 1088px, 30fps). The videos were recorded in the laboratory, using a game simulator to simulate highway and urban driving conditions, and a three-screen display system and a driving device with pedals, gears, and a force-feedback steering wheel to simulate the vehicle. The collected data was annotated and processed, containing 113 videos, covering five types of maneuvering behaviors: going straight, left lane change, left turn, right lane change, and right turn.
[0117] During the process of converting the videos into frames, some data that did not meet the research requirements were found, such as blank data, videos with too short length, and inconsistent data inside and outside the cockpit. Therefore, the data was manually screened, and Table 3 shows the specific information of the Brain4Cars and Zenodo datasets. Among them, "Straight", "Lchange", "Lturn", "Rchange", "Rturn" represent going straight, left lane change, left turn, right lane change, and right turn respectively, "In_car" and "Out_car" represent the actually available data inside and outside the cockpit after excluding unqualified data, and "In_Out" represents the qualified data with one-to-one correspondence between the data inside and outside the cockpit.
[0118] Table 3 Specific information of the Brain4Cars and Zenodo datasets
[0119]
[0120] Implementation information
[0121] Based on the PyTorch framework, research on driver steering intention recognition inside the cockpit, outside the cockpit, and jointly inside and outside the cockpit was conducted on a server equipped with NVIDIA GeForce RTX 3060. 5-fold cross-validation was used for evaluation, and different training strategies were adopted for experiments from different perspectives:
[0122] · Experiment inside the cockpit: The training method used 60 epochs and a batch size of 12. To ensure the robustness and performance of model training, the cross-entropy loss function was used, SGD was selected as the optimizer, the momentum was set to 0.9, and the weight decay was set to 0.001. A multi-step learning rate scheduler was created with an initial learning rate of 0.01, and the learning rate was adjusted at a decay rate of 0.1 after the 30th and 50th epochs.
[0123] · Experiment outside the cockpit: The training method used 80 epochs and a batch size of 8. The initial learning rate was 0.01, and the learning rate was adjusted at a decay rate of 0.1 after the 30th and 60th epochs. The mean squared error loss function was adopted, and the other parameters were the same as those in the experiment inside the cockpit.
[0124] · Joint experiment inside and outside the cockpit: The training strategy was the same as that in the experiment outside the cockpit.
[0125] Evaluation metrics
[0126] The accuracy, F1 score, and the number of model parameters were used to evaluate the model performance. In addition, the number of model parameters was compared additionally. The number of model parameters refers to the total number of weights and biases that need to be learned in the network. The number of parameters reflects the complexity of the model. A smaller number of parameters means the model is lighter and more efficient and can be better applied in resource-constrained environments.
[0127] Results and discussion
[0128] In this experiment, the model was mainly evaluated around the experiments inside the cockpit, outside the cockpit, and jointly inside and outside the cockpit. Since two datasets were used in the model training process, for simplicity of expression, the Brain4Cars dataset was uniformly abbreviated as B, and the Zenodo dataset was abbreviated as Z.
[0129] Experiment inside the cockpit
[0130] In the experiment inside the cockpit, the performance of 3DResNet 50, A-3DResNet34, A-3DResNet 50, and A-3D ResNet 101 on the Brain4Cars dataset was compared, and the results are shown in Table 4. The bold numbers represent the best performance.
[0131] Table 1 Internal Cockpit Experiment Evaluation: Evaluation Results of 3DResNet and A-3DResNet Series Networks on the Brain4Cars Dataset
[0132]
[0133] According to the evaluation results of each model on the Brain4Cars dataset, the A-3DResNet 50 model has the best accuracy and F1-score, which are 79.5% and 81.6% respectively, and the number of parameters is 46.20M. A-3DResNet 34 is a shallow network with fewer layers and fewer parameters. The number of parameters is only 33.15M, which makes the model faster in training and inference, but also results in relatively poor performance. The accuracy and F1-score of A-3DResNet 101 are 0.4% and 2.0% smaller than those of A-3DResNet 50 respectively. Because it has a deeper network structure and a larger number of parameters, the model training and inference time are longer, resulting in poor performance.
[0134] According to the evaluation results of 3DResNet 50 and A-3DResNet 50 on the Brain4Cars dataset, the performance of A-3DResNet50 is better than that of 3DResNet 50. The accuracy and F1-score are increased by 2.1% and 6.1% respectively, and the number of parameters is decreased by 0.02M. The reason for the performance improvement is the dilated convolution introduced in the first stage of A-3DResNet 50, which expands the receptive field of the input data, extracts richer feature information, and the dilated convolution utilizes the correlation between adjacent pixels across boundaries to extract information, reducing the number of parameters to be learned.
[0135] External Cockpit Experiment
[0136] In the external cockpit experiment, ablation experiments of ConvLSTM and DSMax modules were carried out on the Brain4Cars dataset, and ConvLSTM-4DSMax was compared with other studies. The results are shown in Table 5, and the bold numbers indicate the best performance.
[0137] Table 2 External Cockpit Experiment Evaluation: Evaluation Results of ConvLSTM-4DSMax and Other Studies on the Brain4Cars Dataset, and Ablation Experiments of ConvLSTM and 4DSMax Modules
[0138]
[0139] From the ablation experiment results of the ConvLSTM and DSMax modules in Table 5, it can be seen that the effect of using ConvLSTM alone is not good. After adding 4 DSMax modules, the model accuracy and F1 score reached the highest, which are 63.9% and 66.5% respectively. However, after adding 5 DSMax modules, the model performance decreased instead. This is because as the model complexity increases, the risk of overfitting also increases. Too many modules not only consume computing resources but also reduce the parameter efficiency, thus affecting the learning efficiency and performance of the model.
[0140] According to the comparison of ConvLSTM-4DSMax with the methods of Rong and Gebert on the Brain4Cars dataset, ConvLSTM-4DSMax is superior to the other two methods in terms of accuracy, F1 score and the number of parameters. The accuracy is improved by 3.0% and 10.7% respectively. The F1 score is improved by 0.1% and 23.1% respectively. In terms of the number of parameters, the ConvLSTM-4DSMax model is only 5.35M. The reason for the excellent performance is that the DSMax module composed of depthwise separable convolution and max pooling is used in ConvLSTM-4DSMax. Depthwise separable convolution can reduce the number of parameters and computational complexity of the model. Max pooling is a strategy for spatial dimensionality reduction, which can reduce the dimension of the feature map while retaining important features and improve the efficiency of subsequent processing.
[0141] Experiments inside and outside the combined cockpit
[0142] In EDNet, A-3DResNet 50 extracts the driver intention features inside the cockpit, ConvLSTM-4DSMax extracts the road environment features outside the cockpit, and then uses the D-CAF method to fuse these two types of features, and adopts the freeze training strategy for A-3DResNet 50 and ConvLSTM among them. To evaluate the impact of the D-CAF method and the freeze training strategy on the model performance, the model is evaluated using four different combinations on the Brain4Cars dataset and the Zenodo dataset. The evaluation results are shown in Table 6. Among them, "Splicing" represents the feature-level splicing method, "Freeze" represents the freeze training strategy of the model, and "D-CAF" represents the attention-based feature fusion method. The bold numbers represent the best performance.
[0143] · Splicing: Feature-level splicing, that is, the features inside and outside the cockpit are spliced along the column dimension.
[0144] · Splicing+Freeze: Feature-level splicing, adopting the freeze training strategy.
[0145] · D-CAF: D-CAF feature fusion.
[0146] · D-CAF + Freeze: Feature fusion of D-CAF, adopting the freeze training strategy.
[0147] Table 3 Influence of the D-CAF module and the freeze training strategy on the model performance on the Brain4Cars dataset and the Zenodo dataset
[0148]
[0149] As can be seen from Table 6, for the two feature fusion methods of "Splicing" and "D-CAF", after adding "Freeze", the number of parameters is greatly reduced, and the accuracy and F1 score are improved to varying degrees. After replacing "Splicing" with "D-CAF", the accuracy is increased by 1.4%, the F1 score is increased by 2.2%, and the number of parameters is increased by 1.71M because more parameters are needed to capture and integrate the relationships between different features. The results show that the effect of adopting "D-CAF + Freeze" is the best, with the number of parameters being 11.88M, only one-fifth of that of "D-CAF". Due to adopting the previous optimal weights, the accuracy and F1 score are both improved, reaching 86.3% and 87.8% respectively. This result indicates that the D-CAF module and the freeze strategy are effective.
[0150] Figure 8 The confusion matrix is shown. Each row represents the true label, and each column represents the predicted class. The results show that the recognition effect of going straight is excellent, with an accuracy of 91.3%, and the recognition accuracy of left lane change is the lowest, at 79.17%. From the confusion matrix, 12.5% of the samples with the intention of going straight are recognized as the intention of left lane change. By observing the dataset, it can be found that there are some easily confused samples between going straight and left lane change. There are some irregular behaviors of drivers in the going straight data, such as the driver communicating with other people in the car or being attracted by things outside the car, which causes the driver's head to turn, resulting in the model misjudging the intention of going straight as the intention of left lane change.
[0151] Table 7 shows the comparison with other methods. The bold numbers represent the best performance in the mixed dataset, and the italic numbers represent the best performance on the Brain4Cars dataset. The "Feature" column represents the adopted feature category, the "Camera" column represents the video data facing inside and outside the cockpit, and "Other" represents other features such as GPS information and vehicle speed.
[0152] Table 4 Comparison with other methods
[0153]
[0154] As can be seen from Table 7, EDNet performs best on the Zenodo dataset, achieving an accuracy of 86.6%, an F1-score of 89.0%, and 11.88M parameters. This is because the data quality of the Zenodo dataset is better. The Zenodo dataset is sourced from a laboratory simulator with high video quality and clearer video information. In addition, the drivers in the Zenodo dataset show greater amplitude of movement and more standardized behavior in terms of driving behavior. On the Brain4Cars dataset, the models of this method outperform those of Rong, Gebert, and CEMFormer-CC. On the mixed dataset, the model of this method reaches an accuracy of 86.3% and an F1-score of 87.8%, further demonstrating the effectiveness of the method adopted in this study.
[0155] As can be seen from Table 7, on the Brain4Cars dataset, the F1-score of the method adopted in this study is smaller than that of TIFN. Therefore, to compare these two methods in detail, Table 8 shows the early prediction capabilities of different methods on the Brain4cars dataset. The early prediction ability of the model is evaluated by inputting different numbers of frames into the model. The standard video length in the dataset is 5 seconds, so -1s, -2s, -3s, -4s, -5s are used to represent the cases where the model predicts 1 second to 5 seconds in advance. Among them, TIFN does not mention the relevant information about the accuracy rate within different time ranges.
[0156] Table 5 Early prediction capabilities of different methods on the Brain4cars dataset
[0157]
[0158] The results show that in the comparison of accuracy rates, the method adopted in this study is better than Rong's method in all time ranges. In the comparison of F1-scores, except for the prediction 1 second in advance, the method adopted in this study is better than the other two methods. As the prediction time range increases, both the accuracy rate and the F1-score gradually increase. It can be seen that there is an important connection between the time range and the prediction accuracy. The larger the time range, the more information the model learns, which is more conducive to identifying the driver's steering intention.
[0159] In addition, the accuracy rates and F1-scores of the three methods are relatively low for the predictions 5 seconds and 4 seconds in advance, mainly because the drivers maintain a straight driving state in the early stage of the video, making it difficult to identify. It should be noted that within the time range from 4 seconds to 3 seconds in advance, the model performance of this method has a significant improvement, with the accuracy rate and F1-score increasing by 7.9% and 10.2% respectively, reaching an accuracy rate of 73.7% and an F1-score of 73.0%. This indicates that the method proposed in this study can identify the driver's steering intention 3 seconds in advance.
[0160] Conclusion
[0161] This method proposes an end-to-end temporal information fusion network (steering intention recognition model) based on driver cognition. It uses A-3DResNet as the driver intention feature extractor inside the cockpit, combines the ConvLSTM and DSMax modules as the road environment feature extractor outside the cockpit, then fuses the two features based on the attention feature fusion module D-CAF, and freezes the training of A-3DResNet and ConvLSTM during the network training process. Verified by experiments, the complementarity of information inside and outside the vehicle helps the recognition of the driver's steering intention. This method achieved an accuracy of 85.6% and an F1 score of 86.2% on the Brain4Cars dataset, an accuracy of 86.6% and an F1 score of 89.0% on the Zenodo dataset, and an accuracy of 86.3 and an F1 score of 87.8% on the basis of mixing the two datasets. The number of model parameters is 11.88M, and the comprehensive performance is better than other similar methods.
[0162] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Those of ordinary skill in the art should understand that any modifications or equivalent replacements to the technical solutions of the present invention, without departing from the purpose and scope of the present technical solution, should be covered by the scope of the claims of the present invention.
Claims
1. A method for identifying a driver's steering intention by fusing in-vehicle and out-of-vehicle images, characterized in that, It includes the following steps: S1. Construct a steering intention recognition model for recognizing the driver's steering intention based on the influences inside and outside the vehicle; The steering intention recognition model includes an in-vehicle feature extraction module, an out-vehicle feature extraction module, an in-out vehicle feature fusion module, and a classifier recognition module; The in-vehicle feature extraction module is constructed based on a 3D residual network and is used to obtain the driver's intention features inside the vehicle according to the driver's image inside the vehicle; The image processing module is used to convert the road environment video outside the vehicle into an optical flow map of the relative motion information between the vehicle and other traffic participants; The out-vehicle feature extraction module is used to obtain the road environment features outside the vehicle according to the converted optical flow map; The in-out vehicle feature fusion module is used to fuse the driver's intention features and the road environment features to obtain comprehensive judgment features; The classifier recognition module is used to recognize the driver's motivation behavior according to the comprehensive judgment features, and the motivation behavior includes going straight, left lane change, left turn, right lane change, and right turn; S2. Train the steering intention recognition model; S3. Obtain the images inside and outside the vehicle during actual driving and perform driver intention recognition through the trained steering intention recognition model for intelligent assisted driving; Among them, the in-vehicle feature extraction module is constructed by introducing dilated convolution in Stage 1 of 3DresNet; Among them, the said Stage1 includes a convolutional layer and a max pooling layer for extracting and fusing the spatial features in three-dimensional data; 3DresNet is 3DResNet34, 3DResNet 50 or 3D ResNet 101; The out-vehicle feature extraction module includes ConvLSTM and four sequentially connected DSMax modules; ConvLSTM is used to predict the optical flow features according to the optical flow map; The four DSMax modules are used to further extract and abstract the key features in the optical flow features predicted by ConvLSTM as the road environment features; The working process of the in-out vehicle feature fusion module includes: horizontally splicing the driver's intention features and the road environment features according to the column dimension to obtain a combined feature; Subsequently, calculate the attention weights using a linear layer and a sigmoid activation function; Then perform an element-wise multiplication operation on the combined feature and the attention weights to obtain the final feature representation as the comprehensive judgment feature.
2. The method for identifying a driver's steering intention by fusing in-vehicle and out-of-vehicle images according to claim 1, wherein: In S2, when training the steering intention recognition model, apply the model freezing training strategy to freeze the parameters of the in-vehicle feature extraction module and the out-vehicle feature extraction module, so that while using the knowledge learned from the previous weight file, the model focuses its resources and attention on training the subsequent parts.
3. The driver steering intention recognition method for fusing in-vehicle and off-vehicle images according to claim 1, wherein: The DSMax module is obtained by applying a max pooling layer to the output of DSC; DSC is composed of a depthwise convolution DWConv and a pointwise convolution PWConv; DWConv performs a convolution operation on each channel of the input feature map using a convolution kernel, and then splices the outputs of each channel to obtain the final output; PWConv is a 1×1 convolution used to perform channel fusion operations on the feature map output by DWConv to obtain effectively integrated feature information.
4. The method for identifying a driver's steering intention by fusing in-vehicle and out-of-vehicle images according to claim 1, characterized in that: The number of channels of four successively connected DSMax modules are 64, 128, 256, and 512 respectively, and the sizes of the output feature maps in the spatial dimension are 37×59, 12×20, 4×7, and 1×2 respectively.
5. The method for identifying a driver's steering intention by fusing in-vehicle and out-of-vehicle images according to claim 1, characterized in that: The formula for calculating the attention weights using a linear layer and a sigmoid activation function is: where W is the weight of the linear layer, X is the input feature vector, and b is the bias.
6. The method for identifying a driver's steering intention by fusing in-vehicle and off-vehicle images according to claim 1, characterized in that: The image processing module is FlowNet 2.0; the classifier recognition module is a classifier composed of a fully connected layer FC and a Softmax layer.
7. The driver steering intention recognition method for fusing in-vehicle and off-vehicle images according to claim 1, characterized in that: The loss function of the in-vehicle feature extraction module is the cross-entropy loss function; where C represents the number of categories, and p i is the value of the i-th category in the true label, and q i is the predicted probability; The loss functions of the out-of-vehicle feature extraction module and the in-vehicle and out-of-vehicle feature fusion module are the mean squared error loss functions; where N is the number of samples, is the predicted value, and y i is the target value.
Citation Information
Patent Citations
Defective sitting posture detection method and related equipment
CN117593763A
Lightweight model-based driver expression recognition method
CN118072292A