Driving style classification method and device based on multi-modal fusion
Through the multimodal fusion driving style classification method, CNN and LSTM are used to extract visual and motion modal features, and then fused through a cross-modal attention mechanism, which solves the problem of insufficient single modality data, realizes efficient and real-time driving style recognition, and improves the safety and reliability of the autonomous driving system.
Patent Information
- Application Number
- CN202510693427.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-09-09
AI Technical Summary
Existing driving style classification methods mainly rely on single-modal data and cannot fully utilize the rich information in the driving process, resulting in limited classification effect. They also lack in-depth mining and fusion of multimodal information, making it difficult to achieve efficient and real-time driving style recognition in complex traffic environments.
A driving style classification method based on multimodal fusion is adopted. Visual and motion modal features are extracted through convolutional neural networks (CNN) and long short-term memory networks (LSTM). The cross-modal attention mechanism is used for feature fusion, and driving style classification is performed in combination with a fully connected neural network (FCNN).
It significantly improves the accuracy and real-time performance of driving style classification, meets the needs of autonomous driving systems in complex traffic environments, and improves the generalization ability and classification accuracy of the model.
Smart Images

Figure CN120611238A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of autonomous driving technology, and in particular to a driving style classification method, device, electronic device, and computer-readable storage medium based on multimodal fusion. Background Art
[0002] With the rapid development of artificial intelligence (AI), autonomous driving has become an integral part of modern transportation systems. Autonomous driving technology aims to reduce human intervention, enabling vehicles to make autonomous decisions and navigate the roads, thereby improving traffic safety, optimizing travel efficiency, and enhancing the user experience. In recent years, the analysis and recognition of driving behavior, particularly the classification of driving styles, has garnered widespread attention from both academia and industry.
[0003] Currently, research on driving style classification primarily relies on vehicle motion data and environmental perception information. However, traditional classification methods are often based on single-modal data and often fail to fully utilize the rich information present during driving, resulting in limited classification effectiveness. Furthermore, current driving style classification frameworks typically rely on basic machine learning algorithms and lack the in-depth exploration and integration of multimodal information. Summary of the Invention
[0004] In view of this, the present application provides a method, apparatus, device and computer-readable storage medium for driving style classification based on multimodal fusion. By introducing advanced feature extraction and fusion mechanisms, the accuracy and real-time performance of driving style classification are improved to meet the needs of autonomous driving systems in complex traffic environments.
[0005] The present application is introduced below from multiple aspects, and the implementation methods and beneficial effects of the following multiple aspects can be referenced to each other.
[0006] In a first aspect, the present application provides a driving style classification method based on multimodal fusion, comprising: obtaining visual modal data and motion modal data of a vehicle during driving, the visual modal data being used to indicate the external environment of the vehicle in the form of images or videos, and the motion modal data being used to indicate the motion state of the vehicle; determining, based on the visual modal data and a first neural network, the spatial features of the vehicle during driving, the spatial features being used to indicate multiple elements in the external environment and the distance and / or relative motion relationship between the vehicle and the multiple elements, the multiple elements including one or more of lane lines, obstacles, pedestrians, and other vehicles; determining, based on the motion modal data and a second neural network, the time series features of the vehicle during driving, the time series features being used to indicate the motion laws of the vehicle's motion state in the time dimension; and determining the driving style classification result of the driver of the vehicle based on the fusion features of the spatial features and the time series features.
[0007] The implementation of this application significantly improves data processing efficiency and response speed, meeting the real-time requirements in complex driving environments. Furthermore, the introduced cross-modal attention mechanism effectively integrates feature information from different modalities, improving the model's generalization and classification accuracy.
[0008] In a possible implementation of the first aspect above, the first neural network is a convolutional neural network (CNN), and determining the spatial features of the vehicle during driving includes: adjusting the attention weight of each channel in the CNN through a lightweight channel attention mechanism to perform a feature extraction process; passing through multiple convolutional layers and pooling layers of the CNN, and implementing batch normalization and ReLU activation functions to obtain a feature map after spatial feature extraction and downsampling; and compressing the feature map into a spatial feature vector of a fixed length through a global average pooling layer to obtain the spatial features of the vehicle.
[0009] In a possible implementation of the first aspect above, the second neural network is a long short-term memory network LSTM, and determining the time series characteristics of the vehicle during driving includes: processing and optimizing the long sequence data in the motion modal data through a sparse self-attention mechanism; determining the time dependency of at least one motion parameter in the optimized long sequence data through a multi-layer gating mechanism, the at least one motion parameter including one or more of speed, acceleration, and steering angle; and determining the time series characteristics of the vehicle based on the time dependency of the at least one motion parameter.
[0010] In a possible implementation of the first aspect above, determining the driving style classification result of the driver of the vehicle includes: performing weighted fusion of the spatial features and the time series features through a cross-modal attention mechanism to obtain a fused feature representation; performing splicing processing and linear transformation on the fused features to obtain a comprehensive feature representation; and determining the driving style classification result of the driver based on the comprehensive feature representation and a third neural network.
[0011] In a possible implementation of the first aspect above, the weighted fusion of the spatial feature and the time series feature and the splicing processing and linear change of the fused feature include: determining the dot product of the first eigenvector corresponding to the spatial feature and the second eigenvector corresponding to the time series feature to obtain an attention score; determining the respective weights of the visual modality and the motion modality in the attention score through a normalization function; and determining the fused feature representation based on the sum of the product of the weight of the visual modality and the first eigenvector and the product of the weight of the motion modality and the second eigenvector.
[0012] In a possible implementation of the first aspect above, the third neural network is a fully connected neural network (FCNN), and determining the driving style classification result of the driver includes: performing a nonlinear transformation on the comprehensive feature representation through multiple hidden layers of the FCNN; determining a probability distribution of each driving style among multiple driving styles through a Softmax activation function and output results of the multiple hidden layers; and determining the driving style classification result of the driver from the multiple driving styles based on the probability distribution of each driving style among the multiple driving styles.
[0013] According to an embodiment of the present application, the classifier can classify driving styles into different types, such as smooth driving, aggressive driving, and cautious driving. During the training process, the model is optimized based on the labeled data to improve the classification accuracy.
[0014] In a possible implementation of the first aspect, each hidden layer in the multiple hidden layers uses a ReLU activation function to perform nonlinear transformation.
[0015] In a possible implementation of the first aspect above, the visual modality data includes one or more of an image of the road ahead, an image of the surrounding environment, and environmental perception information; the motion modality data includes data on the vehicle's speed, acceleration, and steering angle collected in the time dimension; and the driving style classification result includes one of smooth driving, aggressive driving, and cautious driving.
[0016] In a second aspect, the present application provides a driving style classification device based on multimodal fusion, comprising: an acquisition unit, configured to acquire visual modal data and motion modal data of a vehicle during driving, wherein the visual modal data is used to indicate the external environment of the vehicle in the form of an image or video, and the motion modal data is used to indicate the motion state of the vehicle; a processing unit, configured to determine, based on the visual modal data and a first neural network, the spatial features of the vehicle during driving, wherein the spatial features are used to indicate multiple elements in the external environment and the distance and / or relative motion relationship between the vehicle and the multiple elements, wherein the multiple elements include one or more of lane lines, obstacles, pedestrians, and other vehicles; and a processing unit, configured to determine, based on the motion modal data and a second neural network, the time series features of the vehicle during driving, wherein the time series features are used to indicate the motion law of the vehicle's motion state in the time dimension; and a driving style classification result of the driver of the vehicle is determined based on the fusion features of the spatial features and the time series features.
[0017] In a third aspect, the present application provides an autonomous driving system for driving style classification, the device comprising: a memory for storing instructions executed by one or more processors of the device, and a processor, which is one of the processors of the device, for executing the driving style classification method based on multimodal fusion disclosed in any aspect of the first aspect above.
[0018] In a fourth aspect, a computer-readable storage medium is provided, wherein the storage medium stores instructions. When the instructions are executed on a computer, the computer executes the steps of the driving style classification method based on multimodal fusion described in the first aspect.
[0019] In a fifth aspect, a computer program product comprising instructions is provided. When the instructions are executed on a computer, the computer is caused to perform the steps of driving style classification based on multimodal fusion described in the first aspect. Alternatively, a computer program is provided. When the computer program is executed on a computer, the computer is caused to perform the steps of driving style classification based on multimodal fusion described in the first aspect.
[0020] The possible implementation methods and technical effects obtained in the above-mentioned second to fifth aspects are similar to the corresponding technical means and technical effects obtained in the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 This is a schematic diagram of the process of a driving style classification method based on multimodal fusion provided in an embodiment of the present application;
[0022] Figure 2 This is a process diagram of another driving style classification method based on multimodal fusion provided in an embodiment of the present application;
[0023] Figure 3 is a schematic diagram of a driving style classification method based on multimodal fusion provided in an embodiment of the present application;
[0024] Figure 4 It is a structural schematic diagram of a device provided in an embodiment of the present application;
[0025] Figure 5 It is a structural diagram of a system on chip (SoC) provided in an embodiment of the present application. DETAILED DESCRIPTION
[0026] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0027] First, the prior art and existing technical problems involved in this application are introduced.
[0028] Currently, research on driving style classification mainly relies on the construction of classification models based on vehicle motion data (such as acceleration, steering wheel angle, brake pedal travel, etc.) and environmental perception information (such as following distance, road curvature, etc.). However, traditional classification methods are mostly based on single-modal data and often cannot fully utilize the rich information in the driving process, resulting in limited classification results. For example, relying solely on the collected speed and acceleration standard deviation as the judgment standard for aggressive driving can easily lead to misjudgment in intersection lane change scenarios. Motion modalities are difficult to represent the driver's cognitive state (such as psychological factors such as distraction and road rage), and environmental perception data is significantly affected by sensor noise (millimeter wave radar has large ranging errors in rainy and foggy weather), resulting in information loss.
[0029] In recent years, researchers have attempted to combine visual and motion modalities to improve classification accuracy. However, these methods still have shortcomings in feature fusion and dynamic change processing, especially in highly complex scenarios, making it difficult to effectively capture subtle changes in driving behavior. For example, there is a lack of adaptive adjustment mechanisms for modal weights in dynamic driving scenarios. In particular, in emergency braking conditions, the confidence level of the motion modality should be higher than that of the visual modality, but existing fixed-weight fusion strategies make it difficult to achieve this requirement.
[0030] Furthermore, existing driving style classification frameworks typically rely on basic machine learning algorithms and lack the in-depth exploration and integration of multimodal information. Therefore, in practical applications, how to effectively integrate visual and motion information to build an efficient, real-time driving style classification framework has become a key issue that needs to be addressed.
[0031] In order to solve the above technical problems, the present application proposes a driving style classification method 100 based on multimodal fusion. Figure 1 is a flow chart of the method 100 for driving style classification based on multimodal fusion, as shown in FIG. Figure 1 As shown, the method includes steps 110 to 140. The method 100 aims to improve the accuracy and real-time performance of driving style classification by introducing advanced feature extraction and fusion mechanisms to meet the needs of autonomous driving systems in complex traffic environments.
[0032] Step 110: Acquire visual modal data and motion modal data of the vehicle during driving.
[0033] Specifically, the visual modality data is used to indicate the external environment of the vehicle in the form of images or videos, and the motion modality data is used to indicate the motion state of the vehicle.
[0034] Exemplarily, the visual modality data may include road infrastructure perception in images or videos, dynamic object detection and tracking, and environmental state perception. For example, road infrastructure perception may include lane line detection (lane markings captured by the camera are solid, dashed, double yellow lines, etc.), traffic signal recognition (traffic light status (color, countdown)), road signs and signs (speed limit signs, no-entry signs, direction signs), road geometry (curve curvature, slope, shoulder width, etc.), etc. Dynamic object detection and tracking may include other vehicle targets (type classification such as cars, trucks, buses, etc., motion status such as speed, acceleration, turn signal status, relative position), pedestrians and non-motor vehicles (pedestrian posture is standing, walking or running, non-motor vehicle intention prediction such as the tendency to cross the road), detection and tracking of special obstacles, and environmental state perception may include weather conditions (rain and snow detection, haze concentration), lighting changes (day-to-night transitions, sudden changes in brightness at tunnel entrances and exits), and road surface conditions (potholes, ice), etc.
[0035] For another example, the motion modal data may include one or more of the vehicle's acceleration collected along a time sequence (such as three-axis acceleration, i.e., acceleration values along the x-axis (front and rear direction), y-axis (left and right lateral direction), and z-axis (vertical direction) of the vehicle coordinate system, which can be used to describe the vehicle's longitudinal acceleration / deceleration, lateral deviation, and vertical vibration.), roll angle (the vehicle's rotation angle around the x-axis, which can be used to reflect the vehicle's left and right tilt state (such as the body roll when turning)), pitch angle (the vehicle's rotation angle around the y-axis, which can be used to reflect the vehicle's front and rear pitch state (such as nodding action during acceleration or braking)), yaw angle (the vehicle's rotation angle around the z-axis, which can be used to reflect changes in the vehicle's driving direction (such as changes in heading angle during turning)) and vehicle speed (the vehicle's current driving speed, which can be calculated from a wheel speed sensor or GPS data). Motion state data can also include vehicle control parameters, such as the steering wheel angle (reflecting the driver's steering operation intention), brake pressure (the pressure applied by the brake pedal, used to judge the braking intensity), throttle opening (throttle pedal position, used to reflect acceleration requirements), gear status (the vehicle's current gear (such as forward gear, reverse gear, etc.)), etc.
[0036] Step 120: Determine the spatial characteristics of the vehicle during driving based on the visual modality data and the first neural network.
[0037] Specifically, the spatial feature is used to indicate multiple elements in the external environment and the distance and / or relative motion relationship between the vehicle and the multiple elements, and the multiple elements include one or more of lane lines, obstacles, pedestrians, and other vehicles.
[0038] Exemplarily, the relative motion relationship in the spatial feature may include the position, posture, motion state of the vehicle in the three-dimensional physical space and its relative relationship with the surrounding environment. For example, it may include lane association information, such as lateral offset (the distance between the center of the vehicle and the center of the lane line), lane curvature (the degree of curvature of the lane ahead). For another example, it may include obstacle interaction information, such as (longitudinal / lateral distance from the obstacle (such as the distance from the vehicle in front, pedestrians)), relative speed (speed difference with surrounding moving objects (for collision time calculation), collision risk area (dynamic safety boundary generated based on vehicle size and motion state). For another example, it may include traffic constraint information, such as the relative position to the traffic light (such as the distance to the stop line), geometric restrictions of the turning area of the intersection (such as the boundary of the left turn waiting area), etc.
[0039] In an embodiment of the present application, the first neural network can be a convolutional neural network (CNN), and the attention weight of each channel in the CNN can be adjusted by a lightweight channel attention mechanism to perform a feature extraction process. Afterwards, the CNN passes through multiple layers of convolutional layers and pooling layers, and batch normalization and ReLU activation functions are implemented to obtain a feature map after spatial feature extraction and downsampling, thereby improving the convergence speed and expression ability of the model. Finally, the feature map is compressed into a spatial feature vector of a fixed length through a global average pooling layer to obtain the spatial features of the vehicle, thereby enhancing the model's ability to focus on key visual information. The following embodiments will introduce this process in detail and will not be repeated here.
[0040] Step 130: Determine the time series characteristics of the vehicle during the driving process based on the motion modal data and the second neural network.
[0041] Specifically, the time series feature is used to indicate the motion regularity of the vehicle's motion state in the time dimension.
[0042] Exemplarily, the time series characteristics of the vehicle may include a speed sequence (instantaneous speed changes over time, reflecting driving behaviors such as acceleration, deceleration, and constant speed), an acceleration sequence (longitudinal (forward / braking) and lateral (turning) acceleration), a yaw rate sequence (the vehicle's rotation rate around the vertical axis, characterizing the turning trend), a steering wheel angle sequence (a continuous record of steering operations, used to identify the driver's intentions such as changing lanes and turning), etc.
[0043] In an embodiment of the present application, the second neural network can be a long short-term memory network (LSTM). In step 130, the long sequence data in the motion modal data is processed and optimized through a sparse self-attention mechanism, significantly improving computational efficiency. Afterwards, the time dependency of at least one motion parameter in the optimized long sequence data is determined through a multi-layer gating mechanism. The at least one motion parameter includes one or more of speed, acceleration, and steering angle, thereby forming a rich motion modal feature representation. Finally, the time series characteristics of the vehicle are determined based on the time dependency of the at least one motion parameter. The following embodiments will introduce this process in detail and will not be elaborated here.
[0044] Step 140: Determine a driving style classification result of the driver of the vehicle based on the fusion feature of the spatial feature and the time series feature.
[0045] For example, in the feature fusion process of the present application, the dot product of the first eigenvector corresponding to the spatial feature and the second eigenvector corresponding to the time series feature can be determined to obtain an attention score. Afterwards, the weights of the visual modality and the motion modality in the attention score are determined by a normalization function to determine the importance of each modality under the current input. Then, the visual features and motion features are fused by weighted averaging to form a fused feature representation. Afterwards, the fused features are spliced and linearly transformed to obtain a higher-dimensional comprehensive feature representation, thereby providing a rich information basis for the subsequent driving style classifier.
[0046] In the driving style classification process of the present application, the driver's driving style classification result can be determined based on the above-mentioned comprehensive feature representation and the third neural network, wherein the third neural network can be a fully connected neural network (FCNN). For example, the comprehensive feature representation can be nonlinearly transformed through multiple hidden layers of the FCNN, wherein each hidden layer can use a ReLU activation function for nonlinear transformation. Thereafter, the probability distribution of each driving style among the multiple driving styles is determined by the Softmax activation function and the output results of the multiple hidden layers. Furthermore, based on the probability distribution of each driving style, the driver's driving style classification result is determined from the multiple driving styles. The following embodiment will introduce step 140 in detail and will not be repeated here.
[0047] Method 100 can accurately identify driving style by collecting and fusing visual modal information and motion modal information in real time, thereby improving the accuracy and real-time performance of driving style classification to meet the needs of autonomous driving in complex traffic environments.
[0048] The following combination Figures 2 to 3 An example of an embodiment corresponding to method 100 is introduced. Figure 2 is a schematic diagram of the process of a method for driving style classification based on multimodal fusion provided in an embodiment of the present application, that is, a schematic diagram of an exemplary method of method 100; Figure 3 It is a schematic diagram of the example method provided in the embodiment of the present application.
[0049] Step 210: Initialize the visual modality feature extraction module and the motion modality feature extraction module, and load pre-trained model parameters for real-time analysis.
[0050] For example, by initializing the visual modality feature extraction module and the motion modality feature extraction module, these modules can collect dynamic and environmental information about the vehicle, such as initializing multiple sensors to enable real-time data acquisition. Pre-trained model parameters are loaded, along with a deep learning algorithm that can be used by the feature extraction module to extract valuable features for subsequent analysis. Furthermore, a machine learning algorithm that can be used by the classification module is loaded to process the extracted features.
[0051] Step 220: Collect visual modal data (such as video stream) and motion modal data (such as vehicle speed, acceleration, steering angle, etc.) during driving in real time and perform data preprocessing.
[0052] Specifically, the data acquisition module can be used to obtain multimodal data generated during driving, such as vehicle motion information (such as speed, acceleration, steering angle, etc.) The data acquisition module collects and stores visual information (such as images of the road ahead and the surrounding environment), and performs data preprocessing to ensure the quality and consistency of the input data. For example, the data acquisition module can use a variety of sensors to collect vehicle motion information, such as an inertial measurement unit (IMU), a global positioning system (GPS), and a camera, to collect real-time information about the vehicle's status and surrounding environment. For another example, visual information can be captured by a camera around the traffic environment, including pedestrians, other vehicles, and traffic signs.
[0053] Step 230: The preprocessed visual modality data is input into the visual modality feature extraction module, and the improved CNN is used to extract spatial features; at the same time, the motion modality data is input into the motion modality feature extraction module, and the improved LSTM is used to extract time series features.
[0054] Specifically, the preprocessed visual data is input into the visual modality feature extraction module, and the improved CNN is used to extract the spatial feature X (V) At the same time, the motion data And input to the motion modality feature extraction module, using the improved LSTM to extract the time series feature X (M) .
[0055] Exemplarily, the visual modality feature extraction module is used to extract spatial features from visual data during driving. Specifically, Figure 3 As shown in the figure, an improved CNN is used in combination with a lightweight channel attention mechanism (efficient channel attention network, ECA-Net), which dynamically adjusts the attention weight α of each channel. c , optimize the feature extraction process. After that, through multiple convolutional layers and pooling layers, and implementing batch normalization and ReLU activation function, the convergence speed and expression ability of the model are improved. Finally, the feature map F is pooled by the global average pooling layer. v Compressed to a fixed-length feature vector X (v) , thereby enhancing the model's attention to key visual information.
[0056] Optionally, in the visual modality feature extraction module, the lightweight channel attention mechanism (ECANet) adaptively learns the correlation between channels and uses one-dimensional convolution instead of the fully connected layer, reducing the number of parameters and computational complexity. Specifically, the attention weight α c Calculated by global average pooling and one-dimensional convolution:
[0057] α c =σ(Conv1D(GAP(X (v) ))) (1)
[0058] Where σ is the Sigmoid activation function and GAP(·) represents the global average pooling operation.
[0059] Optionally, in the motion modality feature extraction module, a sparse self-attention mechanism can calculate attention scores only between specific positions in the sequence, reducing computational complexity. For long sequence inputs, by setting a fixed-size attention window w, attention scores are calculated only at positions that satisfy |ij| ≤ w.
[0060] In addition, if Figure 3 As shown in Figure 1, the motion modality feature extraction module is used to extract time series features from the vehicle's motion data. Specifically, based on the improved LSTM, a sparse self-attention mechanism (including attention window, attention score calculation, and attention regularization) is introduced to optimize the processing of long sequence inputs. Through a multi-layer gating mechanism, the time dependence of velocity, acceleration and steering angle is fully captured.
[0061] Specifically, for the long sequence structure of motion modality data, the input features of each time step t are first represented as a vector This includes a joint embedding representation of multiple parameters such as vehicle speed, acceleration, steering angle, etc. To capture temporal dependencies, an improved long short-term memory (LSTM) network is introduced. This network integrates a local sparse self-attention mechanism based on the standard gating structure to enhance the perception of key time steps.
[0062] During the attention calculation process, a fixed window size is defined For each time step t, only the attention weight within the window is calculated, specifically:
[0063] s t,τ =(W q x t ) T (W k x τ ),where|τ-t|≤w (2)
[0064] in, is the mapping matrix between query and key, d a is the attention dimension. The softmax function is then used to normalize the attention score to obtain the attention weight:
[0065]
[0066] Then calculate the sparse attention weighted representation:
[0067]
[0068] in, is the value mapping matrix, h t Represents the aggregated features of the historical state within the window.
[0069] In order to avoid information bias caused by excessive concentration of attention, the attention regularization term is introduced:
[0070]
[0071] Among them, λ is the regularization coefficient, α t Represents the attention distribution of time step t to each time point in the window.
[0072] Finally, the hidden state h of each time step is output t , in order to realize the temporal modeling of vehicle dynamic behavior and form the characteristic representation of motion mode X (m) , and enhance the model's ability to recognize subtle behavioral patterns in complex driving situations.
[0073] Step 240: Through the cross-modal attention mechanism, the extracted visual features and motion features are weightedly fused to generate a comprehensive feature representation.
[0074] Specifically, through the cross-modal attention mechanism, the extracted visual features X (v) and motion features X (m) Weighted fusion is performed to generate a comprehensive feature representation F to capture important dynamic information in driving behavior.
[0075] For example, the visual feature X is calculated by computing (v) and motion features X (m) The dot product between them generates the attention score And normalize it through the Softmax function to get the weight of each mode:
[0076]
[0077] Among them, s v and s m Represent the scores of visual and motion modalities respectively. Then, the visual features and motion features are fused by weighted average to form a fusion feature representation
[0078]
[0079] Afterwards, the fusion feature representation module can also be used to represent the fused features For further processing. Specifically, the fusion features are spliced and transformed linearly. Where W and b are learnable parameters that generate a comprehensive feature representation with higher dimensions, providing rich information for subsequent classifiers.
[0080] Step 250: Represent the fused features Or z inputs the classifier to classify the driving style and outputs the corresponding classification results.
[0081] The classifier, which can be an FCNN, support vector machine (SVM), or random forest (RF), is used to classify driving styles based on the fused features. For example, an FCNN can include multiple hidden layers, each of which uses a ReLU activation function for nonlinear transformations and employs a dropout strategy to reduce the risk of model overfitting. The final layer uses a Softmax activation function to output a probability distribution P(y|z) = Softmax(z) for each driving style, where y represents the driving style category.
[0082] The classifier can classify driving styles into different types, such as smooth driving, aggressive driving, and cautious driving. During the training process, the model is optimized based on the labeled data to improve classification accuracy.
[0083] After step 250, the classification results can be fed back to the driver assistance system for subsequent decision support and real-time adjustments to optimize the driving experience and improve safety. For example, the system can provide feedback and suggestions to the driver based on the classification results. For example, in aggressive driving situations, the system can prompt the driver to adjust their driving style to improve driving safety. The system can also interact with the vehicle control system to automatically adjust the vehicle's dynamic control strategy based on driving style to enhance driving safety and comfort.
[0084] Optionally, the autonomous driving system for driving style classification in the present application can be designed modularly, such as designing the above-mentioned feature extraction module, feature fusion module and classification module, so that each functional module can be independently upgraded and optimized to ensure the adaptability and scalability of the system.
[0085] The aforementioned driving style classification method and system significantly improves data processing efficiency and response speed by decomposing the driving style classification problem into three steps: feature extraction from visual and motion modalities, cross-modal fusion, and classification. This meets the real-time requirements of complex driving environments. Furthermore, the introduction of a cross-modal attention mechanism effectively integrates feature information from different modalities, improving the model's generalization and classification accuracy.
[0086] Optionally, in an embodiment of the present application, more deep learning models (such as Transformer architecture) can be introduced to replace or optimize existing feature extraction modules and fusion mechanisms.
[0087] Through the above technical solution, the driving style classification framework based on vision-motion multimodal fusion provided by this application overcomes the limitations of traditional methods under a single modality, realizes accurate and efficient classification of driving styles, and is of great significance for improving the safety and reliability of autonomous driving systems.
[0088] Now refer to Figure 4, shown is a block diagram of a device 400 according to one embodiment of the present application. The device 400 may include one or more processors 401 coupled to a controller hub 403. For at least one embodiment, the controller hub 403 communicates with the processor 401 via a multi-drop bus such as a front side bus (FSB), a point-to-point interface such as a quickpath interconnect (QPI), or a similar connection 410. The processor 401 executes instructions that control general types of data processing operations. In one embodiment, the controller hub 403 includes, but is not limited to, a graphics memory controller hub (GMCH) (not shown) and an input / output hub (IOH) (which may be on separate chips) (not shown), wherein the GMCH includes memory and a graphics controller and is coupled to the IOH.
[0089] The device 400 may also include a coprocessor 402 and a memory 404 coupled to the controller hub 403. Alternatively, one or both of the memory and the GMCH may be integrated within the processor, with the memory 404 and the coprocessor 402 directly coupled to the processor 401 and the controller hub 403, with the controller hub 403 and the IOH being in a single chip. The memory 404 may be, for example, a dynamic random access memory (DRAM), a phase change memory (PCM), or a combination of the two. In one embodiment, the coprocessor 402 is a special-purpose processor, such as, for example, a high-throughput MIC processor (many integrated core, MIC), a network or communication processor, a compression engine, a graphics processor, a general purpose computing on GPU (GPGPU), or an embedded processor, etc. The optional nature of the coprocessor 402 is indicated by a dotted line in Figure 4 middle.
[0090] Memory 404, as a computer-readable storage medium, may include one or more tangible, non-transitory computer-readable media for storing data and / or instructions. For example, memory 404 may include any suitable non-volatile memory such as flash memory and / or any suitable non-volatile storage device, such as one or more hard disk drives (HDD(s)), one or more compact disc (CD) drives, and / or one or more digital versatile disc (DVD) drives.
[0091] In one embodiment, device 400 may further include a network interface controller (NIC) 406. NIC 406 may include a transceiver for providing a radio interface for device 400, thereby communicating with any other suitable devices (such as a front-end module, an antenna, etc.). In various embodiments, NIC 406 may be integrated with other components of device 400. NIC 406 may implement the functionality of the communication unit in the above-described embodiments.
[0092] Device 400 may further include input / output (I / O) devices 405. I / O 405 may include: a user interface designed to enable a user to interact with device 400; a peripheral component interface designed to enable peripheral components to interact with device 400; and / or sensors designed to determine environmental conditions and / or location information related to device 400.
[0093] It is worth noting that Figure 4 This is for illustrative purposes only. Figure 4 It is shown that the device 400 includes multiple components such as a processor 401, a controller hub 403, a memory 404, etc. However, in actual applications, the device using the methods of the present application may only include a part of the components of the device 400, for example, it may only include the processor 401 and the NIC406. Figure 4 The properties of the optional components are shown with dashed lines. According to some embodiments of the present application, the memory 404 as a computer-readable storage medium stores instructions that, when executed on a computer, cause the device 400 to perform the attention training method according to the above-described embodiment. For details, please refer to the method of the above-described embodiment and will not be repeated here.
[0094] Now refer to Figure 5 , which is a block diagram of a system on chip (SoC) 500 according to an embodiment of the present application. Figure 5In FIG, similar components have the same reference numerals. In addition, the dashed boxes are optional features of more advanced SoCs. Figure 5 In the embodiment, SoC 500 includes: an interconnect unit 550 coupled to an application processor 510; a system agent unit 580; a bus controller unit 590; an integrated memory controller unit 540; a set of one or more coprocessors 520, which may include integrated graphics logic, an image processor, an audio processor, and a video processor; a static random access memory (SRAM) unit 530; and a direct memory access (DMA) unit 560. In one embodiment, the coprocessors 520 include specialized processors, such as, for example, a network or communication processor, a compression engine, a GPGPU, a high-throughput MIC processor, or an embedded processor.
[0095] The static random access memory (SRAM) unit 530 may include one or more computer-readable media for storing data and / or instructions. The computer-readable storage medium may store instructions, specifically, temporary and permanent copies of the instructions. The instructions may include: when executed by at least one unit in the processor, causing the Soc 500 to perform the attention training method according to the above embodiment. For details, please refer to the method of the above embodiment, which will not be repeated here.
[0096] The various embodiments of the mechanisms disclosed in this application can be implemented in hardware, software, firmware, or a combination of these implementation methods. The embodiments of the present application can be implemented as a computer program or program code executed on a programmable system, which includes at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.
[0097] Program code can be applied to input instructions to perform the functions described herein and generate output information. The output information can be applied to one or more output devices in a known manner. For purposes of this application, a processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor.
[0098] Program code can be implemented with a high-level programming language or an object-oriented programming language to communicate with the processing system. Where necessary, program code can also be implemented in assembly language or machine language. In fact, the mechanism described in this application is not limited to the scope of any particular programming language. In either case, the language can be a compiled language or an interpreted language.
[0099] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried or stored on one or more temporary or non-temporary machine-readable (e.g., computer-readable) storage media, which may be read and executed by one or more processors. For example, instructions may be distributed over a network or through other computer-readable media. Therefore, a machine-readable medium may include any mechanism for storing or transmitting information in a machine (e.g., computer) readable form, including but not limited to floppy disks, optical disks, optical discs, compact disc read-only memories (CD-ROMs), magneto-optical disks, read-only memories (ROMs), random-access memories (RAMs), erasable programmable read-only memories (EPROMs), electrically erasable programmable read-only memories (EEPROMs), magnetic or optical cards, flash memory, or a tangible machine-readable memory for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) using the Internet in electrical, optical, acoustic, or other forms of propagation signals. Accordingly, machine-readable media includes any type of machine-readable media suitable for storing or transmitting electronic instructions or information in a form readable by a machine (eg, a computer).
[0100] In the accompanying drawings, some structural or method features may be shown in a particular arrangement and / or order. However, it should be understood that such a particular arrangement and / or order may not be required. Rather, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the accompanying drawings. In addition, the inclusion of a structural or method feature in a particular figure does not imply that such a feature is required in all embodiments, and in some embodiments, such features may not be included or may be combined with other features.
[0101] It should be noted that the units / modules mentioned in the various device embodiments of the present application are all logical units / modules. Physically, a logical unit / module can be a physical unit / module, or a part of a physical unit / module, or can be implemented as a combination of multiple physical units / modules. The physical implementation of these logical units / modules themselves is not the most important. The combination of functions implemented by these logical units / modules is the key to solving the technical problems raised by this application. In addition, in order to highlight the innovative part of this application, the above-mentioned device embodiments of this application do not introduce units / modules that are not closely related to solving the technical problems raised by this application. This does not mean that other units / modules do not exist in the above-mentioned device embodiments.
[0102] It should be noted that in the examples and description of this patent, relational terms such as first and second, etc. are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "including a" does not exclude the presence of other identical elements in the process, method, article or device that includes the element.
[0103] Although the present application has been shown and described with reference to certain preferred embodiments thereof, it will be understood by those skilled in the art that various changes in form and details may be made therein without departing from the spirit and scope of the application.
Claims
1. A driving style classification method based on multimodal fusion, characterized in that: include: Acquiring visual modal data and motion modal data of a vehicle during driving, wherein the visual modal data is used to indicate the external environment of the vehicle in the form of images or videos, and the motion modal data is used to indicate the motion state of the vehicle; determining, based on the visual modality data and the first neural network, spatial features of the vehicle during driving, the spatial features being used to indicate a plurality of elements in the external environment and distances and / or relative motion relationships between the vehicle and the plurality of elements, the plurality of elements comprising one or more of lane lines, obstacles, pedestrians, and other vehicles; Determining, based on the motion modal data and the second neural network, a time series feature of the vehicle during driving, the time series feature being used to indicate a motion regularity of the vehicle's motion state in a time dimension; A driving style classification result of the driver of the vehicle is determined based on the fusion features of the spatial features and the time series features.
2. The method according to claim 1, characterized in that The first neural network is a convolutional neural network (CNN), and determining the spatial features of the vehicle during driving includes: Through the lightweight channel attention mechanism, the attention weight of each channel in the CNN is adjusted to perform feature extraction; Passing through multiple convolutional layers and pooling layers of the CNN, and implementing batch normalization and ReLU activation functions to obtain feature maps after spatial feature extraction and downsampling; The feature map is compressed into a spatial feature vector of fixed length through a global average pooling layer to obtain the spatial feature of the vehicle.
3. The method according to claim 1 or 2, characterized in that The second neural network is a long short-term memory network (LSTM), and determining the time series characteristics of the vehicle during driving includes: Processing and optimizing long sequence data in the motion modality data through a sparse self-attention mechanism; Determining, through a multi-layer gating mechanism, a time dependency of at least one motion parameter in the optimized long sequence data, the at least one motion parameter comprising one or more of velocity, acceleration, and steering angle; The time series characteristic of the vehicle is determined based on the time dependency of the at least one motion parameter.
4. The method according to claim 1 or 2, characterized in that Determining the driving style classification result of the driver of the vehicle includes: The spatial features and the time series features are weightedly fused through a cross-modal attention mechanism to obtain a fused feature representation; Performing splicing processing and linear transformation on the fused features to obtain a comprehensive feature representation; The driving style classification result of the driver is determined based on the comprehensive feature representation and the third neural network.
5. The method according to claim 4, characterized in that The weighted fusion of the spatial features and the time series features and the splicing and linear transformation of the fused features include: Determine a dot product of a first eigenvector corresponding to the spatial feature and a second eigenvector corresponding to the time series feature to obtain an attention score; Determining the respective weights of the visual modality and the motion modality in the attention score through a normalization function; The fused feature representation is determined according to the sum of the product of the weight of the visual modality and the first feature vector and the product of the weight of the motion modality and the second feature vector.
6. The method according to claim 4, characterized in that The third neural network is a fully connected neural network FCNN, and determining the driving style classification result of the driver includes: Performing a nonlinear transformation on the comprehensive feature representation through multiple hidden layers of the FCNN; determining a probability distribution of each of the plurality of driving styles using a Softmax activation function and output results of the plurality of hidden layers; The driving style classification result of the driver is determined from the plurality of driving styles according to a probability distribution of each driving style in the plurality of driving styles.
7. The method according to claim 6, characterized in that Each hidden layer in the multiple hidden layers uses a ReLU activation function to perform nonlinear transformation.
8. The method according to claim 1 or 2, characterized in that The visual modality data includes one or more of an image of the road ahead, an image of the surrounding environment, and environmental perception information; the motion modality data includes data on the speed, acceleration, and steering angle of the vehicle collected in the time dimension; and the driving style classification result includes one of smooth driving, aggressive driving, and cautious driving.
9. A driving style classification device based on multimodal fusion, characterized in that: include: an acquisition unit, configured to acquire visual modal data and motion modal data of a vehicle during driving, wherein the visual modal data is used to indicate the external environment of the vehicle in the form of images or videos, and the motion modal data is used to indicate the motion state of the vehicle; A processing unit is configured to: determine, based on visual modal data and a first neural network, spatial features of the vehicle during driving, the spatial features being used to indicate multiple elements in the external environment and the distance and / or relative motion relationship between the vehicle and the multiple elements, the multiple elements comprising one or more of lane lines, obstacles, pedestrians, and other vehicles; determine, based on motion modal data and a second neural network, time series features of the vehicle during driving, the time series features being used to indicate the motion regularity of the vehicle's motion state in the time dimension; and determine a driving style classification result of the vehicle driver based on a fusion feature of the spatial features and the time series features.
10. An automatic driving system for driving style classification, characterized in that include: a memory for storing instructions executable by the processor; A processor, wherein the processor is configured to implement the method according to any one of claims 1 to 8 when executing the instructions.