Online car-hailing driver driving behavior intelligent sensing method and system and storage medium
Through multimodal data preprocessing and Transformer encoder feature extraction network, combined with multiple perceptrons for driving behavior recognition, the problem of unified modeling of multimodal data in online car-hailing operation scenarios is solved, and efficient recognition and risk assessment of complex driving behaviors are achieved.
Patent Information
- Application Number
- CN202510842727.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-09-23
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies have difficulty in fully mining the deep behavioral information in multimodal sensor data in online ride-hailing operation scenarios, lack a unified modeling mechanism, have difficulty identifying complex driving behaviors, and have difficulty adapting to diverse vehicles and environmental changes.
Raw driving behavior data is collected through multiple modal sensors, and data preprocessing and time alignment are performed. A feature extraction network based on the Transformer encoder is constructed, and multiple perceptrons are combined for feature fusion and classification judgment to achieve efficient fusion and collaborative modeling of multi-source data.
It has improved the integrity and expressiveness of driving behavior characteristics, and can effectively identify a variety of typical driving behaviors, meeting the needs of online ride-hailing platforms in driver portrait construction, safety risk identification and service quality monitoring.
Smart Images

Figure CN120688009A_ABST
Abstract
Description
Technical Field
[0001] The present invention is applicable to the field of vehicle networking technology, and in particular relates to a method, system and storage medium for intelligent perception of driving behavior of online car-hailing drivers. Background Art
[0002] With the continuous improvement of intelligent urban traffic management and the integration of the Internet of Vehicles (IoV) and artificial intelligence technologies, real-time perception and intelligent analysis of driving behavior have become a crucial component of intelligent transportation systems. Among numerous travel scenarios, online ride-hailing platforms, with their wide service coverage, flexible operating models, and diverse driver base, place a higher demand on driving behavior recognition technology. Automated modeling and identification of driver behavior not only assists the platform with driver behavior assessment and risk control, but also serves multiple application areas, including insurance actuarial science, service quality improvement, and road safety regulation. Therefore, developing intelligent driving behavior perception technology for online ride-hailing scenarios has broad research value and practical significance.
[0003] In recent years, the widespread deployment of multimodal sensor devices in ride-hailing vehicles has enabled them to collect diverse sensory data in real time, including acceleration, angular velocity, position trajectory, vehicle control status, and environmental images. These data modalities primarily originate from accelerometers, gyroscopes, GPS modules, CAN buses, OBD terminals, dashcams, and cameras. These data modalities feature high sampling frequencies, rich data dimensions, and strong spatiotemporal synchronization, providing multidimensional information support for behavioral modeling. In related research, some solutions are based on acceleration and gyroscope signals, construct driving behavior recognition models by extracting statistical features, and use supervised learning methods such as support vector machines (SVM), decision trees, and random forests to classify typical driving behaviors; other technical solutions focus on trajectory modeling, using factors such as speed, position, and heading in GPS data, combined with geographic information systems to analyze driving route patterns and operating modes, thereby achieving long-term analysis and trend judgment of driving behaviors; in addition, some methods further introduce a joint modeling strategy of image modalities and control signals, and construct a spatiotemporal modeling architecture based on convolutional neural networks (CNN) or recurrent neural networks (RNN), using data such as forward camera images and vehicle control inputs to achieve understanding and modeling of more complex behaviors.
[0004] However, when it comes to the actual operation of online ride-hailing services, existing solutions have certain flaws:
[0005] Most current technical solutions focus on analyzing sensor data from a single or a few modalities, making it difficult to fully explore the deep behavioral information contained in modalities such as control signals and visual images.
[0006] For example, driving behavior is inherently temporal and evolutionary, but some methods still rely on static windows or low-order statistical processing, making it difficult to effectively capture the dynamic changes in complex behavior.
[0007] Furthermore, due to differences in sampling frequency, timestamp accuracy, and data structure among sensor data, the lack of a unified modeling mechanism makes efficient fusion and collaborative modeling of asynchronous heterogeneous data difficult. Existing methods for object recognition focus on common behaviors such as sudden acceleration and braking, while their ability to represent complex behaviors such as fatigue driving, continuous lane changes, and illegal turns needs to be improved.
[0008] Finally, existing methods are mostly based on closed data sets or specific environment training, and are difficult to adapt to changing situations such as diverse vehicle types, urban routes, and driver habits in online ride-hailing platforms.
[0009] Therefore, it is necessary to propose a new intelligent perception method for driver driving behavior based on the actual operation scenarios of online ride-hailing. Summary of the Invention
[0010] The present invention aims to solve the technical problems of poor behavior perception performance and poor behavior expression ability in existing methods of driving behavior judgment in online car-hailing operation scenarios.
[0011] To solve the above technical problems, in a first aspect, the present invention provides a method for intelligently sensing the driving behavior of online ride-hailing drivers, comprising the following steps:
[0012] S101. Collecting raw driving behavior data from online ride-hailing vehicles and drivers through sensors of multiple modalities;
[0013] S102, performing data preprocessing and time alignment on the original driving behavior data collected by different sensors to obtain modal data;
[0014] S103, performing feature construction and fusion coding according to the modal data to obtain a fusion feature vector;
[0015] S104, performing feature extraction on the fused feature vector according to a feature extraction network based on a Transformer encoder to obtain a global semantic information vector corresponding to the fused feature vector;
[0016] S105. Classify the driving behavior of the online car-hailing driver according to the classification judgment network based on multiple perceptrons and the global semantic information vector to obtain the perception result of the driving behavior of the online car-hailing driver.
[0017] Furthermore, step S101 is specifically as follows:
[0018] The original driving behavior data including image information, acceleration information, angular velocity information, positioning information, and control signal information are collected through sensors of multiple modalities, and the original driving behavior data are processed into a time series according to the collection time.
[0019] Furthermore, step S102 includes the following sub-steps:
[0020] defining a main time axis for constructing the modal data;
[0021] For the image information, input it into a pre-trained convolutional neural network model for dimensionality compression and high-order semantic feature extraction to obtain low-dimensional image features, and align and match the low-dimensional image features with the main time axis in the form of image frames to obtain image modality sub-data;
[0022] For the acceleration information, smoothing the acceleration information using a sliding window of a first window size and an average filtering method to obtain acceleration processing data; resampling each acceleration vector in the acceleration processing data according to the main time axis using a piecewise interpolation method, and performing normalization processing to obtain acceleration modal sub-data;
[0023] For the angular velocity information, smoothing it using a sliding window of a second window size and a low-pass filtering method to obtain angular velocity processed data; resampling each angular velocity vector in the angular velocity processed data according to the main time axis using a linear interpolation method, and normalizing it to obtain angular velocity modal sub-data;
[0024] For the positioning information, resample the vector of each dimension in the positioning information according to the main time axis using a linear interpolation method, and perform normalization processing on the vectors of speed and direction to obtain positioning modal sub-data;
[0025] For the control signal information, clipping the portion exceeding the preset control signal threshold to obtain control signal processing data; interpolating and padding the control signal processing data according to the main time axis using a linear interpolation method, and performing normalization processing to obtain control signal modal sub-data;
[0026] The image modality sub-data, the acceleration modality sub-data, the angular velocity modality sub-data, the positioning modality sub-data and the control signal modality sub-data are aligned and arranged according to the main time axis, and integrated and output as the modality data.
[0027] Furthermore, step S103 includes the following sub-steps:
[0028] For the image modality sub-data, input it into the pre-trained convolutional neural network model to extract low-dimensional semantic embedding vector features to obtain an image feature sub-vector;
[0029] The acceleration modal sub-data is processed into an acceleration characteristic sub-vector of fixed dimension by using a linear transformation method;
[0030] The angular velocity modal sub-data is processed into an angular velocity characteristic sub-vector by adopting a linear transformation method based on an independent weight mapping function;
[0031] For the positioning modal sub-data, a plurality of data segments with continuous values are extracted therefrom, and the data segments are mapped into positioning feature sub-vectors;
[0032] For the control signal sub-data, a nonlinearly activated multi-layer perceptron is used to perform fitting processing to obtain a control signal sub-vector;
[0033] The image feature sub-vector, the acceleration feature sub-vector, the angular velocity feature sub-vector, the positioning feature sub-vector and the control signal sub-vector at the same time on the main time axis are feature-concatenated to obtain the fused feature vector.
[0034] Furthermore, step S104 includes the following sub-steps:
[0035] Performing additive fusion with the fused feature vector using a preset position encoding vector to obtain an input vector with a time sequence identifier;
[0036] Inputting the input vector into the feature extraction network for feature extraction to obtain context-aware features corresponding to each time point on the main timeline; wherein the feature extraction network includes multiple stacked Transformer encoders, each of which includes a multi-head self-attention mechanism and a feedforward neural network, and different Transformer encoders are connected through layer normalization and a residual structure;
[0037] Through the global attention weighting mechanism, the context-aware features corresponding to different time points on the main time axis are weighted and summed to obtain the global semantic information vector.
[0038] Furthermore, step S105 includes the following sub-steps:
[0039] Constructing data labels that define the driving behavior of online ride-hailing drivers, and constructing the classification judgment network based on the multi-sensor machine according to the data labels;
[0040] Inputting the global semantic information vector into the classification judgment network, and outputting a driving behavior judgment result with the data label as an entry;
[0041] Inputting the global semantic information vector into a preset recurrent neural network to perform a risk assessment on the driving behavior of the online car-hailing driver, and outputting a driving behavior risk assessment result;
[0042] The driving behavior judgment result and the driving behavior risk assessment result are integrated and output as the driving behavior perception result of the online car-hailing driver.
[0043] Furthermore, the second window size is smaller than the first window size.
[0044] In a second aspect, the present invention further provides an intelligent perception system for online car-hailing drivers' driving behavior, comprising:
[0045] A multimodal acquisition module, used to collect raw driving behavior data from ride-hailing vehicles and drivers through sensors of multiple modalities;
[0046] A data processing module, configured to perform data preprocessing and time alignment on the raw driving behavior data collected by different sensors to obtain modal data;
[0047] A feature construction module, configured to perform feature construction and fusion encoding based on the modal data to obtain a fused feature vector;
[0048] A deep temporal modeling module is used to perform feature extraction on the fused feature vector according to a feature extraction network based on a Transformer encoder to obtain a global semantic information vector corresponding to the fused feature vector;
[0049] The driving behavior characterization module is used to classify the driving behavior of the online car-hailing driver according to the classification judgment network based on the multiple perceptrons and the global semantic information vector, and obtain the driving behavior perception result of the online car-hailing driver.
[0050] In the third aspect, the present invention also provides a computer device, comprising: a memory, a processor, and an intelligent perception program for online car-hailing driver driving behavior stored on the memory and runnable on the processor. When the processor executes the intelligent perception program for online car-hailing driver driving behavior, it implements the steps in the intelligent perception method for online car-hailing driver driving behavior as described in any one of the above embodiments.
[0051] In a fourth aspect, the present invention also provides a storage medium, on which is stored a program for intelligent perception of driving behavior of online car-hailing drivers. When the program for intelligent perception of driving behavior of online car-hailing drivers is executed by a processor, the steps in the method for intelligent perception of driving behavior of online car-hailing drivers as described in any one of the above embodiments are implemented.
[0052] The beneficial effect achieved by the present invention lies in proposing an intelligent perception method for online car-hailing drivers' driving behavior. This method focuses on the multi-source data collected during the actual operation of online car-hailing, designs a unified data alignment mechanism and deep fusion strategy to improve the integrity and expression ability of behavioral characteristics, and by building an end-to-end deep learning model, combining time series modeling and semantic aggregation structure, it can effectively identify a variety of typical driving behaviors, meeting the actual needs of online car-hailing platforms in driver portrait construction, safety risk identification and service quality monitoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 This is a flowchart of the steps of the intelligent perception method of driving behavior of online car-hailing drivers provided by an embodiment of the present invention;
[0054] Figure 2 This is a schematic diagram of the structure of the intelligent perception system for online car-hailing drivers' driving behavior provided by an embodiment of the present invention;
[0055] Figure 3 It is a structural diagram of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0056] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0057] Please refer to Figure 1 , Figure 1 This is a flowchart of the steps of the intelligent perception method of driving behavior of online car-hailing drivers provided by an embodiment of the present invention. The intelligent perception method of driving behavior of online car-hailing drivers includes the following steps:
[0058] S101. Collect original driving behavior data from online ride-hailing vehicles and drivers through sensors of multiple modalities.
[0059] In the embodiment of the present invention, considering that driving behavior is essentially a complex process composed of the driver's operation, the vehicle's dynamic response, and the external environment, step S101 is specifically as follows:
[0060] The original driving behavior data including image information, acceleration information, angular velocity information, positioning information, and control signal information are collected through sensors of multiple modalities, and the original driving behavior data are processed into a time series according to the collection time.
[0061] Specifically, the image information is obtained by collecting image sequences from the vehicle-mounted forward camera, and the image information at time t is defined as I t, used to describe the road scene, front target, lane information, etc. during driving.
[0062] The linear acceleration and angular velocity are provided by the accelerometer and gyroscope respectively, reflecting the real-time performance of dynamic behaviors such as sudden acceleration, emergency braking, and rapid steering during driving. Acceleration information is defined as a t , represents the instantaneous acceleration in three directions; the angular velocity information is defined as ω t , represents the rotation speed around the three axes. This information is usually acquired at a higher sampling frequency and is suitable for capturing high-frequency detail changes.
[0063] The positioning information is provided by the GPS module, including the longitude, latitude, current speed and heading angle of the vehicle's current position, which is used to describe the relationship between driving trajectory, speed fluctuation and geographical environment. The positioning information at time t is defined as a four-dimensional vector g t , with better stability and coverage.
[0064] The control signal information is obtained through the CAN bus or OBD interface, including control variables such as throttle opening, brake status, steering wheel angle, gear position, etc. The control signal information at time t is defined as a multidimensional vector c t ,These data can directly reflect the driver’s control behavior of the vehicle.
[0065] According to the definition in the above embodiment, the original driving behavior data can be uniformly represented as a data quintuple (I t ,a t ,ω t ,g t ,c t ), where each data is recorded synchronously in the time domain and maintains the original resolution and structural information, providing multi-dimensional and comprehensive information support for subsequent data preprocessing, feature modeling and behavior recognition.
[0066] S102 : performing data preprocessing and time alignment on the original driving behavior data collected by different sensors to obtain modal data.
[0067] Because raw driving behavior data comes from multiple sensor modules and is subject to inconsistent sampling frequencies, timestamp misalignment, unit variations, and noise interference, it must undergo standardized preprocessing before it can be used in subsequent deep learning models for behavior recognition and risk analysis. In this embodiment, the raw driving behavior data is systematically processed through four aspects: data cleaning, interpolation reconstruction, numerical normalization, and format alignment. Targeted processing strategies are then applied to each data type to ensure consistency across all modalities in both temporal and numerical scales.
[0068] Specifically, step S102 includes the following sub-steps:
[0069] S1021. Define a main time axis for constructing the modal data. In an embodiment of the present invention, time t is used as the main time axis.
[0070] S1022. For the image information, input it into a pre-trained convolutional neural network model for dimensionality compression and high-order semantic feature extraction to obtain low-dimensional image features, and align and match the low-dimensional image features with the main time axis in the form of image frames to obtain image modal sub-data.
[0071] Since the on-board forward-facing camera captures continuous image frames at a certain frame rate, this data is used to assist in understanding the external environment, such as the road conditions ahead, lane lines, and the distribution of vehicles or obstacles ahead. Considering that image information is unstructured high-dimensional data, its preprocessing process is different from that of numerical sensor data. In an embodiment of the present invention, the image pixels are not directly normalized, but the original image of the image information is input into the pre-trained convolutional neural network model (such as ResNet) and extracted as low-dimensional image features; thereafter, the image frames of the low-dimensional image features are aligned with the main time axis according to their timestamps, so that each frame of the image maintains temporal consistency with other modal data at the same time, constituting the image modal sub-data.
[0072] S1023. Smoothing the acceleration information using a sliding window of a first window size and an average filtering method to obtain acceleration processing data; resampling each acceleration vector in the acceleration processing data according to the main time axis using a piecewise interpolation method, and normalizing the data to obtain acceleration modal sub-data.
[0073] Acceleration information is provided by a three-axis accelerometer installed on the vehicle, which records the linear motion trend of the vehicle in the X, Y, and Z directions. It is a core indicator for capturing behaviors such as sudden acceleration and sudden braking. This type of data has a high sampling frequency (usually 50Hz or higher), but due to the influence of road vibration, engine vibration, vehicle hardware jitter, etc., the raw data is often accompanied by strong high-frequency noise, which can easily lead to misjudgment of behavior recognition. To this end, in an embodiment of the present invention, the sequence of acceleration information is first smoothed by a sliding window average filter with a first window size to retain the macroscopic motion trend and weaken the perturbation, thereby obtaining acceleration processing data;
[0074] Considering that high-frequency modal alignment errors may accumulate significantly, in the embodiment of the present invention, a piecewise interpolation method is used to resample each vector of the acceleration processing data so that it is strictly aligned with the main time axis; then, in order to eliminate the dimensionality effect and improve the convergence speed of model training, Z-score normalization processing is performed on each dimensional vector to convert the data into a standard numerical domain, and the acceleration modal sub-data a′ is obtained. t, this process can be expressed as:
[0075]
[0076] S1024. Smoothe the angular velocity information using a sliding window of a second window size and a low-pass filtering method to obtain angular velocity processing data, where the second window size is smaller than the first window size; resample each angular velocity vector in the angular velocity processing data according to the main time axis using a linear interpolation method, and perform normalization to obtain angular velocity modal sub-data.
[0077] The angular velocity information comes from the gyroscope, which records the rotational angular velocity of the vehicle around the X, Y, and Z axes. It is often used to detect fine-grained operation modes such as steering, yaw, and lane changes. This type of data also has a high sampling frequency, similar to acceleration, but its numerical fluctuation amplitude is usually smaller, and it is more sensitive to local fluctuations. Therefore, compared with the processing method of acceleration information, during filtering, the embodiment of the present invention uses a sliding window method with a second window size shorter than the first window size and a low-pass filter to perform lightweight smoothing on the angular velocity information to obtain angular velocity processing data to retain detailed rotation change characteristics. The alignment method is consistent with the acceleration information, and linear interpolation is used to reconstruct it to a standard time series, and normalization is also performed to obtain angular velocity modal sub-data ω′. t , the process can be expressed as:
[0078]
[0079] in, is the filtered angular velocity, μ and σ are historical statistical parameters.
[0080] S1025. For the positioning information, use a linear interpolation method to resample the vector of each dimension in the positioning information according to the main time axis, and perform normalization processing on the speed and direction vectors to obtain positioning modal sub-data.
[0081] Positioning information includes the longitude, latitude, speed and heading information of the vehicle at every moment, and is an important source for characterizing the global movement path and traffic environment response. This data is generally collected at a low frequency (1Hz to 5Hz), which has a significant difference in time accuracy compared to high-frequency modes and cannot be directly fused at the frame level. To this end, in the embodiment of the present invention, a linear interpolation strategy is used to reconstruct each dimension of information to the main time axis based on the numerical continuity characteristics of the positioning information to avoid information loss, and obtain the positioning mode sub-data g t ’; At the same time, the two key quantities of speed and heading are normalized to ensure that they have the same expression weight as other modes when inputting the model.
[0082] S1026. For the control signal information, the portion exceeding the preset control signal threshold is cropped to obtain control signal processing data; the control signal processing data is interpolated and padded according to the main time axis using a linear interpolation method, and normalized to obtain control signal modal sub-data.
[0083] The control signal information comes from the vehicle's CAN bus or OBD interface, covering operating instructions such as throttle opening, braking intensity, steering wheel angle, gear position, etc., and is a key variable that expresses the driver's subjective behavioral intention. This type of data may be sparsely sampled, transient mutations, or short-term missing, which can easily interfere with subsequent feature modeling. To this end, the embodiment of the present invention first sets a preset control signal threshold of physical significance for each field, and cuts or removes the over-limit values; then interpolation and padding are performed on the main time axis to ensure that a complete control signal information vector can be obtained at every moment, and finally a normalization conversion is performed to obtain the control signal modal sub-data c t ′.
[0084] S1027 , aligning the image modality sub-data, the acceleration modality sub-data, the angular velocity modality sub-data, the positioning modality sub-data, and the control signal modality sub-data according to the main time axis, and integrating and outputting them as the modality data.
[0085] Through the above steps, this embodiment of the present invention performs cleansing, alignment, normalization, and structural transformation on various types of raw driving behavior data, enabling simultaneous modeling of cross-modal, high-resolution, and multi-dimensional features. All processed data is then used in a standardized structure, uniform frequency, and equal time-step format for subsequent model analysis, ensuring the integrity of the model and the quality of the behavioral semantic modeling.
[0086] S103: Perform feature construction and fusion coding according to the modal data to obtain a fusion feature vector.
[0087] The goal of step S103 is to further abstract various types of raw or preprocessed data into high-dimensional embedded feature representations that can be used for deep semantic modeling. This process not only requires differentiated feature construction strategies for the data structure, physical meaning, and dynamic characteristics of different modalities, but also needs to solve problems such as semantic misalignment and distribution inconsistency that may arise in the fusion process of multi-source information, so as to ensure that the fused features have good robustness and behavioral distinguishability.
[0088] Specifically, step S103 includes the following sub-steps:
[0089] S1031. For the image modality sub-data, input it into the pre-trained convolutional neural network model to perform low-dimensional semantic embedding vector feature extraction to obtain an image feature sub-vector.
[0090] Image modality sub-data has high information density but huge original dimensions. Considering the contextual support role of visual data in driving behavior modeling, such as detecting the preceding vehicle, obstacles, lane lines, and traffic lights, in an embodiment of the present invention, a pre-trained convolutional neural network (such as ResNet, CNN, etc.) is used to extract the semantic embedding vector of the image frame to complete dimensionality compression and high-order semantic extraction, and generate image feature sub-vectors with the same dimensions as other modal data. During implementation, since image data and other numerical modalities differ greatly in distribution form and modeling mechanism, attention should be paid to scale unification and distribution regularization during feature fusion. This process can be expressed as:
[0091]
[0092] in, It is a convolutional neural network.
[0093] S1032: For the acceleration modal sub-data, use a linear transformation method to process it into an acceleration characteristic sub-vector of fixed dimension.
[0094] The acceleration modal sub-data reflects the mutation trend of vehicle motion and is an important basis for distinguishing between sudden acceleration, emergency braking, uphill and downhill operations in behavior recognition. Considering that its physical units are consistent and the variation range is concentrated, in the embodiment of the present invention, a linear transformation is used to map it to a unified embedding space, and a low-dimensional embedding is used to avoid redundant interference. In the implementation process, a weight-shared one-dimensional convolution or a fully connected network can be used to map the acceleration modal sub-data into a fixed-dimensional acceleration feature sub-vector. This not only preserves the direction information but also has good scalability. The process can be expressed as:
[0095]
[0096] Among them, W a and b a are the weights and biases used for acceleration modal sub-data, respectively.
[0097] S1033: Process the angular velocity modal sub-data into an angular velocity characteristic sub-vector by using a linear transformation method based on an independent weight mapping function.
[0098] Angular velocity modal sub-data is a key signal for determining whether the driver is making a quick lane change, sharp turn, or snaking around. Compared to acceleration, the change in angular velocity is more subtle and direction-sensitive, so the directional characteristics must be preserved during the embedding conversion process. In this embodiment of the present invention, an independent weight mapping function is used to perform a linear transformation on the angular velocity modal sub-data while preserving its time series order to obtain the angular velocity feature sub-vector This allows for effective expression of behavioral patterns such as “continuous steering” or “back and forth adjustment” in subsequent time series modeling. The process can be expressed as:
[0099]
[0100] Among them, W ω and b ω are the weights and biases used for the angular velocity modal sub-data, respectively.
[0101] S1034. For the positioning modal sub-data, extract a plurality of data segments with continuous values therefrom, and map the data segments into positioning feature sub-vectors.
[0102] The speed and heading information in the positioning modal sub-data are important variables that reflect the macroscopic motion state of the vehicle. Speed can be used to identify the overall trend of acceleration and deceleration, while heading provides the vehicle's forward direction in the road network. In the embodiment of the present invention, the positioning feature sub-vector is obtained by extracting continuous numerical fields (such as instantaneous speed and heading angle) from the positioning modal sub-data and normalizing them, mapping them into a unified embedding vector. The process can be expressed as:
[0103]
[0104] Among them, W g and b g are the weights and biases used to locate the modal sub-data, respectively.
[0105] S1035 . Perform fitting processing on the control signal sub-data using a multi-layer perceptron with nonlinear activation to obtain a control signal sub-vector.
[0106] The control signal sub-data is a direct expression of the driver's subjective intention and is also the key basis for judging the "initiative" of behavior. Although this type of data has low dimensionality, it changes abruptly and has strong behavioral decision-making significance. For example, the combination of accelerator and brake can identify operations such as "hesitant driving" or "forced deceleration". In this embodiment of the present invention, a multi-layer perceptron (MLP) structure with nonlinear activation is used, combined with ReLU activation and dropout to suppress overfitting, to obtain the control signal sub-vector To achieve sufficiently strong semantic expression capabilities in low dimensions, the process can be expressed as:
[0107]
[0108] Among them, W c and b c are the weights and biases used to control the signal sub-data respectively.
[0109] S1036. Perform feature splicing on the image feature sub-vector, the acceleration feature sub-vector, the angular velocity feature sub-vector, the positioning feature sub-vector, and the control signal sub-vector at the same time on the main time axis to obtain the fused feature vector.
[0110] After completing the feature construction of each sub-modal data, the feature sub-vectors at the same time t are fused into a unified high-dimensional fusion feature vector z by using feature concatenation. t This method not only retains the independent expression ability of each modality, but also allows the subsequent model to automatically determine the contribution ratio of different modalities in the final behavior recognition through parameter learning. t It can be expressed as:
[0111]
[0112] S104 . Perform feature extraction on the fused feature vector according to a feature extraction network based on a Transformer encoder to obtain a global semantic information vector corresponding to the fused feature vector.
[0113] Because a fused feature vector at a single point in time cannot fully capture the entire process of a behavioral event, and the formation and recognition of behavior rely on the evolution and inherent relationships of features over time, effectively capturing the dynamic changes of multimodal fused features over time and modeling the contributions of behaviors across different time periods are key issues in building a high-performance driving behavior perception system.
[0114] Specifically, step S104 includes the following sub-steps:
[0115] S1041: Use a preset position encoding vector and the fused feature vector to perform additive fusion to obtain an input vector with a time sequence identifier.
[0116] In step S104, the embodiment of the present invention introduces a feature extraction network based on the Transformer encoder to construct a global semantic information vector with global perception and behavioral semantic aggregation capabilities. However, the Transformer architecture itself does not have the ability to model time sequence, so it is necessary to add position encoding to the input sequence to explicitly inject temporal information. In the embodiment of the present invention, a trainable preset position encoding vector p is used. t and the fused feature vector z t Perform additive fusion to form an input vector with a time series identifier And form a complete input vector sequence Positional encoding not only enables the model to distinguish positions at different behavioral stages, but also provides a structural basis for capturing the evolutionary order during behavior.
[0117] S1042. Input the input vector into the feature extraction network for feature extraction to obtain context-aware features corresponding to each time point on the main time axis; wherein the feature extraction network includes multiple stacked Transformer encoders, each of the Transformer encoders includes a multi-head self-attention mechanism and a feedforward neural network, and different Transformer encoders are connected through layer normalization and residual structure.
[0118] The multi-head self-attention mechanism is a key component for modeling temporal dependencies. It allows the feature extraction network to learn the dynamic interactions with all other time steps at each moment, without relying on fixed-length sliding windows or sequential recursion. The multi-head self-attention mechanism calculates weights based on the similarity between query, key, and value vectors, thereby achieving dynamic temporal information aggregation. The outputs of multiple heads are concatenated along the channel dimension and fused through a linear layer to form an updated feature representation for the current moment. The multi-head self-attention mechanism enables the features of each time point to reference information from the entire sequence, giving it a natural advantage in modeling complex behaviors such as "preparing to turn, controlling the steering wheel, and stabilizing the lane."
[0119] After feature extraction through the feature extraction network implemented by several layers of stacked Transformer encoders, each time step will output its context-aware feature h t ,This vector not only contains the current state information, but also integrates the global semantic dependencies between the previous and subsequent behavior states, and is suitable for capturing non-local behavior patterns such as behavior chains, delayed decisions, and continuous manipulation.
[0120] S1043. Using a global attention weighted mechanism, perform weighted summation on the context-aware features corresponding to different time points on the main time axis to obtain the global semantic information vector.
[0121] In the embodiment of the present invention, in order to further integrate the entire context-aware feature h t The sequence is compressed into a fixed-dimensional behavior representation vector for subsequent classification and risk assessment. In this embodiment of the present invention, a global attention weighting mechanism is introduced to perform weighted summation on the output features of all time steps to generate the final global semantic information vector h * , the process can be expressed as:
[0122]
[0123] where α tThis global attention weighting mechanism enables the feature extraction network to automatically focus on key periods that contribute most to the current behavior classification when constructing the global semantic information vector, such as the initial stage of emergency braking and the peak points of continuous lane changes, thereby improving interpretation ability and behavior differentiation.
[0124] It should be noted that the Transformer encoder on which the feature extraction network in the embodiment of the present invention is based can also use variant structures such as Performer, Informer or time-aware Transformer, depending on the usage scenario, to adapt to edge deployment environments of different scales, resource constraints or delay sensitivity. Different encoders can also replace layer normalization and residual structures with graph neural networks (GNNs), modal attention mechanisms (such as modal gating networks) or fused convolutional structures to achieve dynamic weighting and information selectivity enhancement between modalities.
[0125] S105. Classify the driving behavior of the online car-hailing driver according to the classification judgment network based on multiple perceptrons and the global semantic information vector to obtain the perception result of the driving behavior of the online car-hailing driver.
[0126] The main goal of step S105 is to obtain the global semantic information vector h based on the output of the previous step. * The behavior category and risk level corresponding to the current driving segment are predicted simultaneously, thereby realizing intelligent recognition of driving status, high-precision judgment of behavioral events, and quantitative assessment of driver risk status.
[0127] Specifically, step S105 includes the following sub-steps:
[0128] S1051. Construct data labels that define the driving behavior of online car-hailing drivers, and construct the classification judgment network based on multiple perceptrons based on the data labels.
[0129] The data labels in the embodiment of the present invention include categories of driving behaviors, such as "normal driving", "sudden acceleration", "sudden braking", "frequent lane changes", "non-standard steering", "fatigue driving", etc. These categories include both perceptible instantaneous actions and continuous or implicit behavior patterns, and have practical application value for the behavior evaluation of online car-hailing drivers.
[0130] The driving behavior classification task is a multi-class classification problem. In the embodiment of the present invention, a multi-layer perceptron is used as the classifier structure, with the global semantic information vector h * As input, output is the probability distribution of driving behavior The process can be expressed as:
[0131]
[0132] Among them, W1 is the classification weight matrix, b1 is the bias term. The final output label is the maximum probability The corresponding category is determined, namely:
[0133]
[0134] S1052: Input the global semantic information vector into the classification judgment network, and output a driving behavior judgment result with the data label as an entry.
[0135] S1053. Input the global semantic information vector into a preset regression neural network to perform a risk assessment on the driving behavior of the online car-hailing driver, and output the driving behavior risk assessment result.
[0136] To further quantify the safety risk level corresponding to the current behavior segment, the embodiment of the present invention designs a parallel risk regression channel to perform continuous value prediction of the potential safety impact of driving behavior. The risk assessment value r is output by an independent regression neural network. This process can be expressed as:
[0137] r=σ(W2h * +b2);
[0138] Here, σ is the activation function, ensuring that the output value falls within an interpretable risk range. W2 and b2 are the weight and bias, respectively. The risk assessment design allows ride-hailing platforms to quantitatively manage behavior, such as setting risk thresholds to automatically trigger alerts, inferring fatigue status based on continuous risk trends, or using it as an input for vehicle insurance floating factors.
[0139] S1054. Integrate the driving behavior judgment result and the driving behavior risk assessment result and output them as the driving behavior perception result of the online car-hailing driver.
[0140] Based on the methods described in the preceding embodiments, the present invention constructs a complete intelligent behavior perception technology system across three dimensions: data, modeling, and output. Implementation is not limited to the vehicle itself, but can also be extended to integrate user mobile terminal data, environmental sensing devices, external signal platforms, and more, enabling cross-platform, multi-dimensional driving behavior modeling. This design not only overcomes the bottlenecks of existing technologies in information fusion and behavior representation, but also achieves a closed-loop transformation from driving data to actionable risk control indicators through an innovative sequence modeling mechanism and joint prediction architecture.
[0141] The beneficial effect achieved by the present invention lies in proposing an intelligent perception method for online car-hailing drivers' driving behavior. This method focuses on the multi-source data collected during the actual operation of online car-hailing, designs a unified data alignment mechanism and deep fusion strategy to improve the integrity and expression ability of behavioral characteristics, and by building an end-to-end deep learning model, combining time series modeling and semantic aggregation structure, it can effectively identify a variety of typical driving behaviors, meeting the actual needs of online car-hailing platforms in driver portrait construction, safety risk identification and service quality monitoring.
[0142] The embodiment of the present invention also provides an intelligent perception system 200 for online car-hailing drivers’ driving behavior. Figure 2 , Figure 2 : This is a schematic diagram of the structure of the intelligent perception system for online car-hailing driver driving behavior provided by an embodiment of the present invention, which includes:
[0143] The multimodal acquisition module 201 is used to collect raw driving behavior data from online ride-hailing vehicles and drivers through sensors of multiple modalities;
[0144] A data processing module 202 is configured to perform data preprocessing and time alignment on the raw driving behavior data collected by different sensors to obtain modal data;
[0145] A feature construction module 203 is used to perform feature construction and fusion coding based on the modal data to obtain a fused feature vector;
[0146] A deep temporal modeling module 204 is configured to perform feature extraction on the fused feature vector according to a feature extraction network based on a Transformer encoder to obtain a global semantic information vector corresponding to the fused feature vector;
[0147] The driving behavior characterization module 205 is used to classify the driving behavior of the online car-hailing driver according to the classification judgment network based on the multiple sensor machines and the global semantic information vector to obtain the driving behavior perception result of the online car-hailing driver.
[0148] The online car-hailing driver driving behavior intelligent perception system 200 can implement the steps in the online car-hailing driver driving behavior intelligent perception method in the above embodiment, and can achieve the same technical effects. Please refer to the description in the above embodiment and will not repeat it here.
[0149] The embodiment of the present invention also provides a computer device, please refer to Figure 3 , Figure 3 It is a structural diagram of a computer device provided in an embodiment of the present invention. The computer device 300 includes: a memory 302, a processor 301, and an intelligent perception program for driving behavior of online car-hailing drivers stored in the memory 302 and capable of running on the processor 301.
[0150] The processor 301 calls the intelligent perception program of the online car-hailing driver's driving behavior stored in the memory 302 to execute the steps of the intelligent perception method of the online car-hailing driver's driving behavior provided by the embodiment of the present invention. Figure 1 , specifically including the following steps:
[0151] S101. Collect original driving behavior data from online ride-hailing vehicles and drivers through sensors of multiple modalities.
[0152] Step S101 is specifically as follows:
[0153] The original driving behavior data including image information, acceleration information, angular velocity information, positioning information, and control signal information are collected through sensors of multiple modalities, and the original driving behavior data are processed into a time series according to the collection time.
[0154] S102 : performing data preprocessing and time alignment on the original driving behavior data collected by different sensors to obtain modal data.
[0155] Step S102 includes the following sub-steps:
[0156] defining a main time axis for constructing the modal data;
[0157] For the image information, input it into a pre-trained convolutional neural network model for dimensionality compression and high-order semantic feature extraction to obtain low-dimensional image features, and align and match the low-dimensional image features with the main time axis in the form of image frames to obtain image modality sub-data;
[0158] For the acceleration information, smoothing the acceleration information using a sliding window of a first window size and an average filtering method to obtain acceleration processing data; resampling each acceleration vector in the acceleration processing data according to the main time axis using a piecewise interpolation method, and performing normalization processing to obtain acceleration modal sub-data;
[0159] Smoothing the angular velocity information using a sliding window of a second window size and a low-pass filtering method to obtain angular velocity processed data, wherein the second window size is smaller than the first window size; resampling each angular velocity vector in the angular velocity processed data according to the main time axis using a linear interpolation method, and performing normalization processing to obtain angular velocity modal sub-data;
[0160] For the positioning information, resample the vector of each dimension in the positioning information according to the main time axis using a linear interpolation method, and perform normalization processing on the vectors of speed and direction to obtain positioning modal sub-data;
[0161] For the control signal information, clipping the portion exceeding the preset control signal threshold to obtain control signal processing data; interpolating and padding the control signal processing data according to the main time axis using a linear interpolation method, and performing normalization processing to obtain control signal modal sub-data;
[0162] The image modality sub-data, the acceleration modality sub-data, the angular velocity modality sub-data, the positioning modality sub-data and the control signal modality sub-data are aligned and arranged according to the main time axis, and integrated and output as the modality data.
[0163] S103: Perform feature construction and fusion coding according to the modal data to obtain a fusion feature vector.
[0164] Step S103 includes the following sub-steps:
[0165] For the image modality sub-data, input it into the pre-trained convolutional neural network model to extract low-dimensional semantic embedding vector features to obtain an image feature sub-vector;
[0166] The acceleration modal sub-data is processed into an acceleration characteristic sub-vector of fixed dimension by using a linear transformation method;
[0167] The angular velocity modal sub-data is processed into an angular velocity characteristic sub-vector by adopting a linear transformation method based on an independent weight mapping function;
[0168] For the positioning modal sub-data, a plurality of data segments with continuous values are extracted therefrom, and the data segments are mapped into positioning feature sub-vectors;
[0169] For the control signal sub-data, a nonlinearly activated multi-layer perceptron is used to perform fitting processing to obtain a control signal sub-vector;
[0170] The image feature sub-vector, the acceleration feature sub-vector, the angular velocity feature sub-vector, the positioning feature sub-vector and the control signal sub-vector at the same time on the main time axis are feature-concatenated to obtain the fused feature vector.
[0171] S104 . Perform feature extraction on the fused feature vector according to a feature extraction network based on a Transformer encoder to obtain a global semantic information vector corresponding to the fused feature vector.
[0172] Step S104 includes the following sub-steps:
[0173] Performing additive fusion with the fused feature vector using a preset position encoding vector to obtain an input vector with a time sequence identifier;
[0174] Inputting the input vector into the feature extraction network for feature extraction to obtain context-aware features corresponding to each time point on the main timeline; wherein the feature extraction network includes multiple stacked Transformer encoders, each of which includes a multi-head self-attention mechanism and a feedforward neural network, and different Transformer encoders are connected through layer normalization and a residual structure;
[0175] The global semantic information vector is obtained by weighting and summing the context-aware features corresponding to different time points on the main time axis through a global attention weighting mechanism.
[0176] S105. Classify the driving behavior of the online car-hailing driver according to the classification judgment network based on multiple perceptrons and the global semantic information vector to obtain the perception result of the driving behavior of the online car-hailing driver.
[0177] Step S105 includes the following sub-steps:
[0178] Constructing data labels that define the driving behavior of online ride-hailing drivers, and constructing the classification judgment network based on the multi-sensor machine according to the data labels;
[0179] Inputting the global semantic information vector into the classification judgment network, and outputting a driving behavior judgment result with the data label as an entry;
[0180] Inputting the global semantic information vector into a preset recurrent neural network to perform a risk assessment on the driving behavior of the online car-hailing driver, and outputting a driving behavior risk assessment result;
[0181] The driving behavior judgment result and the driving behavior risk assessment result are integrated and output as the driving behavior perception result of the online car-hailing driver.
[0182] The computer device 300 provided in an embodiment of the present invention can implement the steps in the intelligent perception method of driving behavior of online car-hailing drivers in the above embodiment, and can achieve the same technical effects. Please refer to the description in the above embodiment and will not repeat it here.
[0183] An embodiment of the present invention also provides a storage medium, on which is stored a program for intelligent perception of driving behavior of online car-hailing drivers. When the program for intelligent perception of driving behavior of online car-hailing drivers is executed by a processor, the various processes and steps in the method for intelligent perception of driving behavior of online car-hailing drivers provided in an embodiment of the present invention are implemented, and the same technical effects can be achieved. To avoid repetition, they will not be repeated here.
[0184] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by using an intelligent perception program for online ride-hailing drivers' driving behavior to instruct related hardware (which can be a mobile phone, computer, server, air conditioner, or network equipment, etc.). The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0185] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0186] The embodiments of the present invention are described above in conjunction with the accompanying drawings. What is disclosed is only a preferred embodiment of the present invention. However, the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms and equivalent changes without departing from the scope of protection of the purpose of the present invention and the claims, which are all within the protection of the present invention.
Claims
1. An intelligent perception method for online car-hailing driver's driving behavior, characterized by: The following steps are involved: S101. Collecting raw driving behavior data from online ride-hailing vehicles and drivers through sensors of multiple modalities; S102, performing data preprocessing and time alignment on the original driving behavior data collected by different sensors to obtain modal data; S103, performing feature construction and fusion coding according to the modal data to obtain a fusion feature vector; S104, performing feature extraction on the fused feature vector according to a feature extraction network based on a Transformer encoder to obtain a global semantic information vector corresponding to the fused feature vector; S105. Classify the driving behavior of the online car-hailing driver according to the classification judgment network based on multiple perceptrons and the global semantic information vector to obtain the perception result of the driving behavior of the online car-hailing driver.
2. The intelligent perception method of driving behavior of online car-hailing drivers according to claim 1 is characterized in that: Step S101 is specifically as follows: The original driving behavior data including image information, acceleration information, angular velocity information, positioning information, and control signal information are collected through sensors of multiple modalities, and the original driving behavior data are processed into a time series according to the collection time.
3. The intelligent perception method of driving behavior of online car-hailing drivers according to claim 2 is characterized in that: Step S102 includes the following sub-steps: defining a main time axis for constructing the modal data; For the image information, input it into a pre-trained convolutional neural network model for dimensionality compression and high-order semantic feature extraction to obtain low-dimensional image features, and align and match the low-dimensional image features with the main time axis in the form of image frames to obtain image modality sub-data; For the acceleration information, smoothing the acceleration information using a sliding window of a first window size and an average filtering method to obtain acceleration processing data; resampling each acceleration vector in the acceleration processing data according to the main time axis using a piecewise interpolation method, and performing normalization processing to obtain acceleration modal sub-data; For the angular velocity information, smoothing it using a sliding window of a second window size and a low-pass filtering method to obtain angular velocity processed data; resampling each angular velocity vector in the angular velocity processed data according to the main time axis using a linear interpolation method, and normalizing it to obtain angular velocity modal sub-data; For the positioning information, resample the vector of each dimension in the positioning information according to the main time axis using a linear interpolation method, and perform normalization processing on the vectors of speed and direction to obtain positioning modal sub-data; For the control signal information, clipping a portion exceeding a preset control signal threshold to obtain control signal processing data; Using a linear interpolation method to interpolate and fill the control signal processing data according to the main time axis, and performing normalization processing to obtain control signal modal sub-data; The image modality sub-data, the acceleration modality sub-data, the angular velocity modality sub-data, the positioning modality sub-data and the control signal modality sub-data are aligned and arranged according to the main time axis, and integrated and output as the modality data.
4. The intelligent perception method of driving behavior of online car-hailing drivers according to claim 3 is characterized in that: Step S103 includes the following sub-steps: For the image modality sub-data, input it into the pre-trained convolutional neural network model to extract low-dimensional semantic embedding vector features to obtain an image feature sub-vector; The acceleration modal sub-data is processed into an acceleration characteristic sub-vector of fixed dimension by using a linear transformation method; The angular velocity modal sub-data is processed into an angular velocity characteristic sub-vector by adopting a linear transformation method based on an independent weight mapping function; For the positioning modal sub-data, a plurality of data segments with continuous values are extracted therefrom, and the data segments are mapped into positioning feature sub-vectors; For the control signal sub-data, a nonlinearly activated multi-layer perceptron is used to perform fitting processing to obtain a control signal sub-vector; The image feature sub-vector, the acceleration feature sub-vector, the angular velocity feature sub-vector, the positioning feature sub-vector and the control signal sub-vector at the same time on the main time axis are feature-concatenated to obtain the fused feature vector.
5. The intelligent perception method of driving behavior of online car-hailing drivers according to claim 4 is characterized in that: Step S104 includes the following sub-steps: Performing additive fusion with the fused feature vector using a preset position encoding vector to obtain an input vector with a time sequence identifier; Inputting the input vector into the feature extraction network for feature extraction to obtain context-aware features corresponding to each time point on the main timeline; wherein the feature extraction network includes multiple stacked Transformer encoders, each of which includes a multi-head self-attention mechanism and a feedforward neural network, and different Transformer encoders are connected through layer normalization and a residual structure; Through the global attention weighting mechanism, the context-aware features corresponding to different time points on the main time axis are weighted and summed to obtain the global semantic information vector.
6. The intelligent perception method of driving behavior of online car-hailing drivers according to claim 5 is characterized in that: Step S105 includes the following sub-steps: Constructing data labels that define the driving behavior of online ride-hailing drivers, and constructing the classification judgment network based on the multi-sensor machine according to the data labels; Inputting the global semantic information vector into the classification judgment network, and outputting a driving behavior judgment result with the data label as an entry; Inputting the global semantic information vector into a preset recurrent neural network to perform a risk assessment on the driving behavior of the online car-hailing driver, and outputting a driving behavior risk assessment result; The driving behavior judgment result and the driving behavior risk assessment result are integrated and output as the driving behavior perception result of the online car-hailing driver.
7. The intelligent perception method of driving behavior of online car-hailing drivers according to claim 3 is characterized in that: The second window size is smaller than the first window size.
8. An intelligent perception system for online car-hailing drivers' driving behavior, characterized by: include: A multimodal acquisition module, used to collect raw driving behavior data from ride-hailing vehicles and drivers through sensors of multiple modalities; A data processing module, configured to perform data preprocessing and time alignment on the raw driving behavior data collected by different sensors to obtain modal data; A feature construction module, configured to perform feature construction and fusion encoding based on the modal data to obtain a fused feature vector; A deep temporal modeling module is used to perform feature extraction on the fused feature vector according to a feature extraction network based on a Transformer encoder to obtain a global semantic information vector corresponding to the fused feature vector; The driving behavior characterization module is used to classify the driving behavior of the online car-hailing driver according to the classification judgment network based on the multiple perceptrons and the global semantic information vector, and obtain the driving behavior perception result of the online car-hailing driver.
9. A computer device, characterized in that: include: A memory, a processor, and an intelligent perception program for online car-hailing driver's driving behavior stored in the memory and runnable on the processor. When the processor executes the intelligent perception program for online car-hailing driver's driving behavior, the steps in the intelligent perception method for online car-hailing driver's driving behavior as described in any one of claims 1 to 7 are implemented.
10. A storage medium, characterized in that: The storage medium stores an intelligent perception program for online car-hailing driver's driving behavior. When the intelligent perception program for online car-hailing driver's driving behavior is executed by the processor, the steps in the intelligent perception method for online car-hailing driver's driving behavior as described in any one of claims 1-7 are implemented.