Driving Behavior Analysis Method and System Based on Multimodal Sensor Fusion

Through multimodal sensor fusion and behavioral analysis network, vehicle operation, environmental interaction and behavioral continuity characteristics are extracted, abnormal behavior is identified and driving strategies are optimized, which solves the problem of insufficient comprehensive analysis of driving behavior in the existing technology, and improves driving safety and efficiency.

CN120156538BActive Publication Date: 2025-07-22SICHUAN BEIDOU SATELLITE OF CHINA TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510631825.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-07-22
Estimated Expiration
2045-05-16

AI Technical Summary

Technical Problem

The existing driving behavior analysis technology ignores the interaction characteristics between the vehicle and the environment and the continuity of driving behavior, resulting in the incomplete and accurate identification of abnormal driving behaviors, and simple rule judgments cannot adapt to complex and changeable driving scenarios.

Method used

Multi-source sensor data is obtained through multi-modal sensor fusion, and after performing time-space alignment processing, the preset behavior analysis network is used to extract vehicle operation, environmental interaction and behavior continuity characteristics, combine the abnormal behavior recognition layer to identify abnormal behavior, and generate driving strategy optimization data to adjust driving strategy.

Benefits of technology

It has achieved in-depth exploration and comprehensive capture of driving behavior, and can detect abnormal behaviors in a timely and accurate manner, improving driving safety, comfort and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120156538B_ABST
    Figure CN120156538B_ABST
Patent Text Reader

Abstract

The present invention provides a driving behavior analysis method and system based on multi-modal sensor fusion. First, a multi-source sensor data set of a target vehicle containing driving environment perception data from at least two different sensor sources is obtained, and it is subjected to spatio-temporal alignment processing to generate a multi-source fusion data set. Then, through the feature extraction layer of a preset behavior analysis network, a driving behavior feature set containing vehicle operation, environmental interaction, and behavior continuity features is extracted from the multi-source fusion data set. Next, an anomaly recognition layer is used to perform anomaly recognition on the driving behavior feature set to generate an anomaly recognition result. Finally, driving strategy optimization data is generated according to the anomaly recognition result and fed back to the vehicle control system to adjust the driving strategy. This method can comprehensively analyze driving behaviors, timely detect anomalies, and optimize driving strategies, improving driving safety and efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent assisted driving, and in particular, to a driving behavior analysis method and system based on multi-modal sensor fusion. Background Art

[0002] With the rapid development of the automotive industry and the continuous progress of intelligent driving technology, the need for accurate analysis and optimization of driving behavior has become increasingly urgent. At present, in the field of driving behavior analysis, there are many limitations in the existing technologies.

[0003] Some traditional methods often only focus on the basic operation characteristics of the vehicle, such as acceleration, deceleration, steering, etc., while ignoring the interaction characteristics between the vehicle and the surrounding environment and the continuity characteristics of driving behavior. Driving behavior is a complex process, which is not only related to the vehicle's own operations, but also closely related to the surrounding environment. At the same time, driving behavior has a certain continuity and coherence. The lack of comprehensive consideration of these characteristics makes the analysis of driving behavior not comprehensive and in-depth enough, and it is impossible to accurately identify potential abnormal driving behaviors.

[0004] For the identification of abnormal driving behaviors, most of the existing technologies are based on simple rule judgments. For example, a fixed speed threshold is set to judge speeding behaviors. This simple rule judgment method cannot adapt to complex and changeable driving scenarios, and it is difficult to effectively identify some atypical abnormal behaviors, resulting in the inability to timely discover potential risks during driving. Summary of the Invention

[0005] In view of the problems mentioned above, in combination with the first aspect of the present invention, embodiments of the present invention provide a driving behavior analysis method based on multi-modal sensor fusion, and the method includes:

[0006] Obtain a multi-source sensor data set of a target vehicle, where the multi-source sensor data set includes driving environment perception data from at least two different sensor sources;

[0007] Perform spatio-temporal alignment processing on the multi-source sensor data set to generate a multi-source fusion data set;

[0008] Based on the feature extraction layer of a preset behavior analysis network, perform behavior feature extraction processing on the multi-source fusion data set to obtain a driving behavior feature set, where the driving behavior feature set includes vehicle operation features, environment interaction features, and behavior continuity features;

[0009] Based on the abnormal recognition layer of a preset behavior analysis network, perform abnormal behavior recognition processing on the driving behavior feature set to generate a driving behavior abnormal recognition result;

[0010] Generate driving strategy optimization data according to the abnormal driving behavior recognition result, and feedback the driving strategy optimization data to the vehicle control system to trigger a driving strategy adjustment operation.

[0011] In another aspect, an embodiment of the present invention further provides a driving behavior analysis system based on multi-modal sensor fusion, including a processor and a machine-readable storage medium. The machine-readable storage medium is connected to the processor. The machine-readable storage medium is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the machine-readable storage medium to implement the above method.

[0012] Based on the above aspects, the present invention obtains a multi-source sensor data set including driving environment perception data from at least two different sensor sources, and performs spatio-temporal alignment processing on the multi-source sensor data set to generate a multi-source fusion data set, effectively solving the problem of inconsistency of different sensor data in time and space, so that the fused data can accurately reflect the driving scene information at the same moment and space, improving the availability and reliability of the data. Use the feature extraction layer of the preset behavior analysis network to extract behavior features from the multi-source fusion data set, and obtain a driving behavior feature set including vehicle operation features, environmental interaction features and behavior continuity features, realizing in-depth mining and comprehensive capture of driving behavior features, and accurately depicting driving behavior from multiple aspects. Based on the abnormal behavior recognition layer of the preset behavior analysis network, perform abnormal behavior recognition processing on the driving behavior feature set and generate an abnormal driving behavior recognition result, which can timely and accurately detect abnormal behaviors during driving, providing a strong guarantee for driving safety. Finally, generate driving strategy optimization data according to the abnormal driving behavior recognition result and feedback it to the vehicle control system to trigger a driving strategy adjustment operation, realizing the dynamic optimization and real-time adjustment of the driving strategy, and effectively improving the safety, comfort and efficiency of driving. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 is a schematic execution flowchart of a driving behavior analysis method based on multi-modal sensor fusion provided by an embodiment of the present invention.

[0014] Figure 2 is a schematic diagram of an exemplary hardware and software component of a driving behavior analysis system based on multi-modal sensor fusion provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0015] The present invention will be specifically described below with reference to the accompanying drawings of the specification. Figure 1 is a flowchart of a driving behavior analysis method based on multi-modal sensor fusion provided by an embodiment of the present invention. The driving behavior analysis method based on multi-modal sensor fusion will be introduced in detail below.

[0016] Step S110: Obtain a multi-source sensor data set of the target vehicle, where the multi-source sensor data set includes driving environment perception data from at least two different sensor sources.

[0017] In an actual application scenario, the target vehicle is equipped with a variety of different types of sensors to collect relevant information about the driving environment. Suppose these sensors are mainly divided into types A, B, C, etc. Sensor A may be a vision sensor, such as a camera, which can collect image data in areas such as the front and side of the vehicle. These image data contain visual feature information of targets such as road signs, other vehicles, and pedestrians. Sensor B can be a radar sensor, which can measure information such as the distance, relative speed, and angle between the target object and the vehicle. These data are output in the form of electrical signals and are converted into specific numerical information after conversion. Sensor C may be an inertial measurement unit, which can sense motion state information such as the acceleration and angular velocity of the vehicle, reflecting the dynamic changes during the vehicle's driving. The data from these different sensor sources together constitute the multi-source sensor data set, and the data of each sensor has its unique properties and dimensions. For example, the image data collected by sensor A may be a multi-dimensional pixel matrix, the measurement data of sensor B may be a one-dimensional combination of distance, speed, and angle values, and the data of sensor C may be a multi-dimensional vector about acceleration and angular velocity.

[0018] Step S120: Perform spatio-temporal alignment processing on the multi-source sensor data set to generate a multi-source fusion data set.

[0019] Since the working characteristics of different sensors are different, there are often differences in the time and space references of the data they collect. In order to effectively fuse these data, spatio-temporal alignment processing is required, which specifically includes the following sub-steps:

[0020] Step S121: Perform timestamp alignment processing on each type of source data in the multi-source sensor data set to generate a time-synchronized data set, where the timestamp alignment processing includes interpolating and resampling the source data with different sampling frequencies to achieve a unified time reference.

[0021] The sampling frequencies of different sensors are different, which results in inconsistent time intervals for the data they collect. Suppose the sampling frequency of sensor A is fA, the sampling frequency of sensor B is fB, and fA is not equal to fB. In order to achieve timestamp alignment, it is necessary to first determine a unified time reference T. For the data collected by each sensor, it is necessary to process it according to this unified time reference.

[0022] For the source data of off-frequency sampling, an interpolation resampling method is adopted. Taking the data of sensor A and sensor B as an example, assume that the data collected by sensor A at time points t1, t2, t3... are DA1, DA2, DA3... respectively, and the data collected by sensor B at time points t1', t2', t3'... are DB1, DB2, DB3... respectively. Under the unified time reference T, for a specific time point Tk, if sensor A does not directly collect data at the moment of Tk, the data value at this moment needs to be estimated by interpolation. Common interpolation methods include linear interpolation. Assuming that Tk is between ti and t_(i + 1), then the data DAk at the moment of Tk can be estimated according to the data DAi and DA_(i + 1) at the moments of ti and t_(i + 1) through the linear interpolation formula. The same applies to the data of sensor B. After such processing, the data of all sensors are under the unified time reference T, forming a time-synchronized data set.

[0023] Step S122: Perform spatial coordinate transformation processing on the time-synchronized data set to generate a spatially aligned data set, where the spatial coordinate transformation processing includes mapping the coordinate systems of different sensors to a unified vehicle coordinate system.

[0024] Different sensors are installed at different positions on the vehicle, each with its own independent coordinate system. To unify these data spatially, it is necessary to map their coordinate systems to a unified vehicle coordinate system. Assume that the coordinate system of sensor A is SA, the coordinate system of sensor B is SB, and the unified vehicle coordinate system is SV.

[0025] First, it is necessary to determine the transformation relationship of each sensor coordinate system relative to the unified vehicle coordinate system. This can be obtained by measuring the installation position and attitude of the sensor. Assume that the rotation matrix of sensor A relative to the unified vehicle coordinate system is RA, and the translation vector is TA. The rotation matrix of sensor B relative to the unified vehicle coordinate system is RB, and the translation vector is TB. For the spatial data point PA (in the coordinate system SA) collected by sensor A, to transform it to the unified vehicle coordinate system SV, the following transformation is required: first, rotate PA through the rotation matrix RA to obtain the rotated vector PA', and then add the translation vector TA to obtain the coordinate PV_A in the unified vehicle coordinate system. For the data PB (in the coordinate system SB) of sensor B, similarly rotate it through the rotation matrix RB to obtain PB', and then add the translation vector TB to obtain the coordinate PV_B in the unified vehicle coordinate system. Through such coordinate transformation, the data of all sensors are unified to the vehicle coordinate system, forming a spatially aligned data set.

[0026] Step S123: Based on a preset source fusion rule, perform feature-level fusion processing on the data from different sensors in the spatially aligned data set to generate the multi-source fusion data set. Each data unit in the multi-source fusion data set contains a fusion feature vector of at least two types of source data.

[0027] The preset source fusion rule defines how to fuse the data from different sensors. In this step, a feature-level fusion method is adopted. Assume that after the previous spatio-temporal alignment processing, the data of sensor A is FA and the data of sensor B is FB, both of which are multi-dimensional feature vectors.

[0028] First, it is necessary to extract features from the data of different sensors. For the data FA of sensor A, the feature vector VA is extracted through the feature extraction algorithm EA; for the data FB of sensor B, the feature vector VB is extracted through the feature extraction algorithm EB. Then, according to the preset source fusion rule, the feature vectors VA and VB are fused. A common fusion method is weighted concatenation. Weights wA and wB (wA + wB = 1) are assigned to the feature vectors VA and VB respectively, and the weighted feature vectors are concatenated to obtain the fusion feature vector VF. For example, VF = [wA * VA; wB * VB], where ";" represents the concatenation operation. In this way, each data unit contains a fusion feature vector of at least two types of source data, forming a multi-source fusion data set.

[0029] Step S130: Based on the feature extraction layer of a preset behavior analysis network, perform behavior feature extraction processing on the multi-source fusion data set to obtain a driving behavior feature set, which includes vehicle operation features, environmental interaction features, and behavior continuity features.

[0030] The feature extraction layer of the preset behavior analysis network is used to extract features related to driving behavior from the multi-source fusion data set, specifically through the following branches:

[0031] Step S131: Through the vehicle operation feature extraction branch of the feature extraction layer, perform normalization processing and multi-scale convolution feature extraction on the vehicle control signals in the multi-source fusion data set to generate a normalized vehicle operation feature vector. Among them, the normalization processing includes mapping the steering wheel angle, throttle depth, and brake pressure to a unified numerical interval, and the multi-scale convolution feature extraction uses convolution kernels of different scales to capture the local patterns of direction control, acceleration, and braking behaviors respectively.

[0032] The multi-source fusion data set contains control signals of the vehicle, such as steering wheel angle, throttle depth, and brake pressure. First, normalize these control signals. For throttle depth and brake pressure, linear transformation can be used to map them to the interval [0, 1]. For example, for the throttle depth β, its normalized value β' = (β - β_min) / (β_max - β_min), where β_min and β_max are the minimum and maximum values of the throttle depth respectively; the same applies to the brake pressure γ.

[0033] For the steering wheel angle α, in order to retain its direction information, a symmetric normalization method is used to map it to the interval [-1, 1]. Let the maximum absolute value of the steering wheel angle be α_max_abs, and the normalized steering wheel angle α' = α / α_max_abs. In this way, the steering wheel angle for a left turn is negative and for a right turn is positive, retaining the direction information.

[0034] Then, perform multi-scale convolution feature extraction. Convolution kernels of different scales, such as convolution kernels K1, K2, and K3, are used to perform convolution operations on the normalized vehicle control signals respectively. Convolution kernel K1 can be used to capture local patterns of direction control behavior, convolution kernel K2 is used to capture local patterns of acceleration behavior, and convolution kernel K3 is used to capture local patterns of braking behavior. Through the convolution operation, feature maps of different scales are obtained, and pooling operations, such as max pooling, are performed on these feature maps to compress the dimensions of the feature maps, and finally a standardized vehicle operation feature vector V_op is generated.

[0035] Step S132: Through the environmental interaction feature extraction branch of the feature extraction layer, perform coordinate system alignment and dynamic target association analysis on the environmental perception signals in the multi-source fusion data set to generate an environmental interaction feature vector, where the coordinate system alignment includes converting the three-dimensional coordinates of the camera detection target and the radar speed measurement data to the same vehicle coordinate system, and the dynamic target association analysis includes calculating the lateral distance deviation between the host vehicle and the target vehicle, the predicted collision time, and the signal light state encoding.

[0036] The environmental perception signals in the multi-source fusion data set contain target information detected by the camera and radar speed measurement data. First, perform coordinate system alignment to convert the three-dimensional coordinates of the camera detection target to the same vehicle coordinate system as the radar data. Assume that the coordinates of the target detected by the camera in its own coordinate system are PC, and through the coordinate conversion method mentioned above, convert it to the coordinates PV_C in the unified vehicle coordinate system.

[0037] Then perform dynamic target association analysis. For the detected target vehicle, calculate the lateral distance deviation Δd between the host vehicle and the target vehicle, that is, the distance difference between the host vehicle and the target vehicle in the lateral direction. The original formula for the predicted time to collision TTC is TTC = d_rel / v_rel. However, when the relative speed v_rel approaches zero, TTC tends to infinity and cannot be calculated in practice. To solve this problem, a velocity smoothing term ε is introduced, and the modified formula is TTC = d_rel / (v_rel + ε), where ε is a very small positive number that can be set according to the actual situation, such as 0.01 m / s. This can avoid invalid values of TTC, ensure the validity of TTC in the environmental interaction feature vector, and prevent outliers from damaging the model stability. The signal light state encoding can determine the state of the signal light (such as red, green, yellow) through an image recognition algorithm based on the image information collected by the camera, and encode it as a numerical value. Combine the information calculated above to form the environmental interaction feature vector V_env.

[0038] Step S133: Through the behavior continuity feature extraction branch of the feature extraction layer, perform sliding window statistic calculation and long short-term dependence modeling on the time series signals in the multi-source fusion data set to generate a behavior continuity feature vector, where the sliding window statistic calculation includes the extraction of normalized indexes for the volatility of the steering wheel angle, the change frequency of the throttle, and the duration of braking, and the long short-term dependence modeling uses a recurrent neural network to extract the temporal behavior pattern.

[0039] The time series signals in the multi-source fusion data set contain information on vehicle operations and environmental perception that changes over time. First, perform sliding window statistic calculation. For the steering wheel angle, calculate its volatility within a sliding window, for example, represent the volatility by calculating the standard deviation of the steering wheel angle within the window. For the throttle depth, calculate its change frequency within the sliding window, such as counting the number of times the throttle depth changes. For the braking pressure, calculate its duration. Normalize these statistics and map them to a unified numerical interval.

[0040] In terms of long short-term dependence modeling, the original method used a recurrent neural network (such as LSTM) to process the time series signals. However, the recurrent neural network has a high processing delay for long sequences and is difficult to meet the requirements of real-time driving analysis. To solve this problem, a lightweight temporal model TCN (Temporal Convolutional Network) is used to replace the recurrent neural network. TCN has the efficient computing characteristics of a convolutional network and can reduce the computing delay when processing long sequences. Input the time series signals after sliding window statistic calculation and normalization into the TCN, and the network will learn the long short-term dependence relationship in the signals, extract the temporal behavior pattern, and finally generate the behavior continuity feature vector V_cont.

[0041] Step S134: Through the fusion module of the feature extraction layer, perform dimension alignment and weighted fusion on the standardized vehicle operation feature vector, environmental interaction feature vector, and behavior continuity feature vector to generate the driving behavior feature set. Among them, the dimension alignment includes using a fully connected layer to project each feature vector into a unified dimension space, and the weighted fusion includes dynamically allocating weights to vehicle operation features and environmental interaction features based on the attention mechanism, and applying a time decay factor to the behavior continuity feature to strengthen the influence of recent behaviors.

[0042] First, perform dimension alignment. Use a fully connected layer to process the standardized vehicle operation feature vector V_op, environmental interaction feature vector V_env, and behavior continuity feature vector V_cont. Let the weight matrices of the fully connected layer be W_op, W_env, and W_cont respectively, and the bias vectors be b_op, b_env, and b_cont respectively. Perform a linear transformation on each feature vector through the fully connected layer, that is, V_op' = W_op * V_op + b_op, V_env' = W_env * V_env + b_env, V_cont' = W_cont * V_cont + b_cont, so that they are projected into a unified dimension space.

[0043] Then, perform weighted fusion. Based on the attention mechanism, dynamically allocate weights to the vehicle operation feature vector V_op' and the environmental interaction feature vector V_env'. Let the weights calculated by the attention mechanism be w_op_att and w_env_att respectively, and w_op_att + w_env_att = 1. For the behavior continuity feature vector V_cont', apply a time decay factor λ. The more recent the behavior, the greater its impact on the result. Finally, splice the weighted feature vectors to generate the driving behavior feature set V_drive = [w_op_att * V_op'; w_env_att * V_env'; λ * V_cont'].

[0044] Step S140: Based on the anomaly recognition layer of the preset behavior analysis network, perform anomaly behavior recognition processing on the driving behavior feature set to generate a driving behavior anomaly recognition result.

[0045] The anomaly recognition layer of the preset behavior analysis network is used to analyze the driving behavior feature set and identify the abnormal behaviors therein, specifically including the following steps:

[0046] Step S141: Input the vehicle operation feature vector, environment interaction feature vector, and behavior continuity feature vector in the driving behavior feature set into the feature association module of the anomaly recognition layer, and generate a first association weight matrix between the vehicle operation feature vector and the environment interaction feature vector, and a second association weight matrix between the behavior continuity feature vector and the vehicle operation feature vector through the multi-head attention mechanism.

[0047] Input the vehicle operation feature vector V_op, environment interaction feature vector V_env, and behavior continuity feature vector V_cont in the driving behavior feature set into the feature association module of the anomaly recognition layer. Using the multi-head attention mechanism, for the vehicle operation feature vector V_op and the environment interaction feature vector V_env, calculate the association weights between them through multiple attention heads. Let the number of attention heads be n. For the i-th attention head, calculate the association weight matrix A_op_env_i between the vehicle operation feature vector V_op and the environment interaction feature vector V_env. Concatenate the association weight matrices calculated by all attention heads to obtain the first association weight matrix A_op_env. Similarly, calculate the second association weight matrix A_cont_op between the behavior continuity feature vector V_cont and the vehicle operation feature vector V_op.

[0048] Step S142: Based on the first association weight matrix, perform weighted aggregation on the environment interaction feature vector to generate a first fusion feature vector aligned with the dimension of the vehicle operation feature vector. At the same time, based on the second association weight matrix, perform weighted aggregation on the behavior continuity feature vector to generate a second fusion feature vector aligned with the dimension of the vehicle operation feature vector.

[0049] Perform weighted aggregation on the environment interaction feature vector V_env according to the first association weight matrix A_op_env. Let each row of A_op_env represent the association weight between an element of the vehicle operation feature vector and each element of the environment interaction feature vector. Multiply these weights by the corresponding elements of the environment interaction feature vector and sum them to obtain the first fusion feature vector V_fusion1 aligned with the dimension of the vehicle operation feature vector. Similarly, perform weighted aggregation on the behavior continuity feature vector V_cont according to the second association weight matrix A_cont_op to obtain the second fusion feature vector V_fusion2 aligned with the dimension of the vehicle operation feature vector.

[0050] Step S143: Concatenate the vehicle operation feature vector, the first fusion feature vector, and the second fusion feature vector along the channel dimension, and input them into the multi-scale temporal convolutional module of the anomaly recognition layer. Extract driving behavior patterns at different time scales through parallel dilated convolutional layers, and perform max pooling and channel fusion on the temporal features output by each dilated convolutional layer to generate a multi-scale fusion feature tensor.

[0051] Concatenate the vehicle operation feature vector V_op, the first fusion feature vector V_fusion1, and the second fusion feature vector V_fusion2 along the channel dimension to obtain the concatenated feature vector V_concat = [V_op; V_fusion1; V_fusion2]. Input V_concat into the multi-scale temporal convolutional module of the anomaly recognition layer. This module contains multiple parallel dilated convolutional layers, each with a different dilation rate, used to extract driving behavior patterns at different time scales. Let the number of dilated convolutional layers be m. For the j-th dilated convolutional layer, its dilation rate is r_j. Perform a convolution operation on V_concat to obtain the temporal feature T_j. Perform a max pooling operation on the temporal feature T_j output by each dilated convolutional layer to compress its dimension, and then perform channel fusion on all pooled temporal features to generate a multi-scale fusion feature tensor T_fusion.

[0052] Step S144: Input the multi-scale fusion feature tensor into the dynamic scoring module of the anomaly recognition layer, and map it to the anomaly probability value of the current time window through a fully connected layer. Among them, the weight parameters of the dynamic scoring module are adaptively adjusted according to the fluctuation amplitude of the vehicle operation feature vector and the target density of the environmental interaction feature vector, and the output layer uses a non-linear activation function to generate an anomaly score.

[0053] Input the multi-scale fusion feature tensor T_fusion into the dynamic scoring module of the anomaly recognition layer. This module contains a fully connected layer, and maps the multi-scale fusion feature tensor T_fusion to the anomaly probability value of the current time window through the fully connected layer. The weight parameters of the dynamic scoring module are adaptively adjusted according to the fluctuation amplitude of the vehicle operation feature vector and the target density of the environmental interaction feature vector. Let the fluctuation amplitude of the vehicle operation feature vector be σ_α, and the target density of the environmental interaction feature vector be ρ_env. Since the numerical ranges of σ_α and ρ_env may vary greatly, in order to avoid the weights being biased towards a certain feature, σ_α and ρ_env are normalized before calculating the weights.

[0054] Let the mean of σ_α be μ_σ_α and the standard deviation be σ_σ_α. After standardization, σ_α_std = (σ_α - μ_σ_α) / σ_σ_α. Let the mean of ρ_env be μ_ρ_env and the standard deviation be σ_ρ_env. After standardization, ρ_env_std = (ρ_env - μ_ρ_env) / σ_ρ_env.

[0055] Adjust the weight matrix W_score and bias vector b_score of the fully connected layer according to the standardized σ_α_std and ρ_env_std. The output layer uses a non-linear activation function (such as the Sigmoid function) to process the output of the fully connected layer to generate an anomaly score S.

[0056] Step S145: Input the number of dynamic targets, relative speed variance in the environmental interaction feature vector, and the standard deviation of the steering wheel angle in the vehicle operation feature vector into the threshold generation sub-network of the anomaly recognition layer, and calculate the dynamic anomaly determination threshold related to the scenario through the fully connected layer.

[0057] Input the number of dynamic targets N, relative speed variance σ_rel in the environmental interaction feature vector, and the standard deviation of the steering wheel angle σ_α in the vehicle operation feature vector into the threshold generation sub-network of the anomaly recognition layer. Since the dimensions of these three input parameters are not unified, the number of dynamic targets N is a dimensionless count, the dimension of the relative speed variance σ_rel is the square of the speed unit (such as (m / s)²), and the dimension of the standard deviation of the steering wheel angle σ_α is the angle unit (such as °). Directly inputting them into the fully connected layer without standardization will lead to an imbalance in parameter weight distribution.

[0058] Therefore, before inputting into the threshold generation sub-network, standardize N, σ_rel, and σ_α. Let the mean of N be μ_N and the standard deviation be σ_N. After standardization, N_std = (N - μ_N) / σ_N. Let the mean of σ_rel be μ_σ_rel and the standard deviation be σ_σ_rel. After standardization, σ_rel_std = (σ_rel - μ_σ_rel) / σ_σ_rel. Let the mean of σ_α be μ_σ_α and the standard deviation be σ_σ_α. After standardization, σ_α_std = (σ_α - μ_σ_α) / σ_σ_α.

[0059] This sub-network includes a fully connected layer. Through the fully connected layer, a linear transformation is performed on the standardized input parameters [N_std; σ_rel_std; σ_α_std]. Let the weight matrix of the fully connected layer be W_thresh and the bias vector be b_thresh. Then the dynamic anomaly determination threshold T_thresh = W_thresh * [N_std; σ_rel_std; σ_α_std] + b_thresh.

[0060] Step S146: Perform a sliding window cumulative calculation on the anomaly scores of consecutive time windows. When the cumulative score exceeds the dynamic anomaly determination threshold and the duration meets the preset conditions, trigger an anomaly behavior flag and generate an initial anomaly determination sequence.

[0061] Perform a sliding window cumulative calculation on the anomaly score S of consecutive time windows. Let the size of the sliding window be k. In each time window, accumulate the anomaly scores within the current window to obtain the cumulative score S_sum. When S_sum exceeds the dynamic anomaly determination threshold T_thresh and the duration of this situation meets the preset conditions (such as the duration being greater than t_0), trigger an anomaly behavior flag and generate an initial anomaly determination sequence S_init, where the elements in S_init are 0 or 1, indicating whether there is an anomaly behavior in this time window.

[0062] Step S147: Input the initial anomaly determination sequence into the temporal correction module of the anomaly recognition layer. Capture the time dependence of historical anomaly flags through a gated recurrent unit network and combine the time decay factor in the behavior continuity feature vector to correct the probability of the anomaly flag in the current time window, and output the smoothed driving behavior anomaly recognition result.

[0063] Input the initial anomaly determination sequence S_init into the temporal correction module of the anomaly recognition layer. This temporal correction module uses a gated recurrent unit network (such as GRU) to capture the time dependence of historical anomaly flags. The gated recurrent unit network can effectively handle the long-range dependence relationships in sequence data through its internal gating mechanism. In this scenario, the changing pattern of anomaly behaviors over time can be learned based on the historical anomaly flag information in the initial anomaly determination sequence S_init.

[0064] At each time step of the gated recurrent unit network, receive the initial anomaly flag of the current time window and the hidden state of the previous time step. Let the current time step be t, the value of the initial anomaly determination sequence S_init at time step t be S_init_t, and the hidden state of the previous time step be h_{t - 1}. The gated recurrent unit network will update the hidden state h_t of the current time step through a series of calculations based on these two inputs. These calculations include the calculation of the reset gate, update gate, and candidate hidden state. The reset gate is used to determine how much information of the hidden state h_{t - 1} of the previous time step needs to be reset, the update gate is used to determine how much information of the hidden state h_{t - 1} of the previous time step needs to be retained to the current time step, and the candidate hidden state is calculated based on the current input and the reset hidden state of the previous time step.

[0065] Meanwhile, the time decay factor λ in the behavioral continuity feature vector is incorporated. The time decay factor reflects the importance of recent behaviors for current anomaly determination. The more recent the behavior, the greater its impact on current anomaly determination. The time decay factor λ is combined with the hidden state h_t output by the gated recurrent unit network to correct the probability of the anomaly flag for the current time window.

[0066] Specifically, let the hidden state h_t output by the gated recurrent unit network be processed through a fully connected layer and a non-linear activation function to obtain a preliminary corrected probability P_pre_t. Then, the preliminary corrected probability P_pre_t and the time decay factor λ are combined with weights. Assuming the weighting coefficients are w_λ and w_pre (w_λ + w_pre = 1), the corrected anomaly probability P_corrected_t = w_λ * λ + w_pre * P_pre_t.

[0067] Finally, the anomaly flag for the current time window is updated according to the corrected anomaly probability P_corrected_t. A threshold T_corrected can be set. If P_corrected_t is greater than T_corrected, it is considered that there is an abnormal behavior in the current time window, and the corresponding anomaly flag is set to 1; otherwise, it is set to 0. Through such processing, the smoothed driving behavior anomaly recognition result S_smoothed is output. This result takes into account the time dependence of historical anomaly flags and the characteristics of behavioral continuity, and can more accurately reflect abnormal situations in driving behavior.

[0068] Step S150: Generate driving strategy optimization data based on the driving behavior anomaly recognition result, and feedback the driving strategy optimization data to the vehicle control system to trigger a driving strategy adjustment operation.

[0069] Based on the driving behavior anomaly recognition result, corresponding driving strategy optimization data needs to be generated to adjust the driving strategy of the vehicle, such as the following specific steps:

[0070] Step S151: Based on the anomaly behavior flag in the driving behavior anomaly recognition result, match the corresponding anomaly correction strategy template from the preset driving strategy optimization rule library. The anomaly correction strategy template includes steering wheel angle correction parameters, desired speed adjustment coefficients, and safe vehicle distance thresholds.

[0071] A preset driving strategy optimization rule library stores abnormal correction strategy templates corresponding to various abnormal behaviors. After obtaining the driving behavior abnormality recognition result S_smoothed, according to the abnormal behavior identifier therein, a match is made in the rule library. Assume that the abnormal behavior identifiers are different category numbers, such as ID_1, ID_2, etc., and each number corresponds to a specific abnormal behavior. For each abnormal behavior category, there is a corresponding abnormal correction strategy template in the rule library. For example, when the abnormal behavior identifier is ID_1, the corresponding abnormal correction strategy template T_template_1 includes a steering wheel angle correction parameter Δα_template_1, an expected speed adjustment coefficient k_v_template_1, and a safe vehicle distance threshold d_safe_template_1. Through this matching operation, the abnormal correction strategy template corresponding to the current abnormal behavior is found.

[0072] Step S152: Extract the vehicle operation feature vector and the environment interaction feature vector from the driving behavior feature set, normalize the standard deviation of the steering wheel angle, the mean value of the throttle depth, and the integral of the brake pressure in the vehicle operation feature vector to generate an operation state evaluation vector, and at the same time standardize the lateral distance deviation, the predicted collision time, and the variance of the relative speed of the target vehicle in the environment interaction feature vector to generate an environment risk evaluation vector.

[0073] Extract the vehicle operation feature vector V_op and the environment interaction feature vector V_env from the driving behavior feature set V_drive. For the vehicle operation feature vector V_op, extract the standard deviation of the steering wheel angle σ_α, the mean value of the throttle depth μ_β, and the integral of the brake pressure I_γ therein. Normalize these parameters. For example, for the standard deviation of the steering wheel angle σ_α, map it to the interval [0, 1] through a linear transformation to obtain the normalized standard deviation of the steering wheel angle σ_α_norm. Similarly, normalize the mean value of the throttle depth μ_β and the integral of the brake pressure I_γ to obtain μ_β_norm and I_γ_norm. Combine these normalized parameters into an operation state evaluation vector V_op_eval = [σ_α_norm; μ_β_norm; I_γ_norm].

[0074] For the environmental interaction feature vector V_env, extract the lateral distance deviation Δd, the predicted collision time TTC, and the relative speed variance σ_rel of the target vehicle from it. Standardize these parameters. For example, by calculating the mean and standard deviation of each parameter, convert them into a standard normal distribution with a mean of 0 and a standard deviation of 1. Let the mean of the lateral distance deviation Δd be μ_Δd and the standard deviation be σ_Δd, then the standardized lateral distance deviation Δd_std = (Δd - μ_Δd) / σ_Δd. Similarly, standardize the predicted collision time TTC and the relative speed variance σ_rel of the target vehicle to obtain TTC_std and σ_rel_std. Combine these standardized parameters into the environmental risk assessment vector V_env_eval = [Δd_std; TTC_std; σ_rel_std].

[0075] Step S153: Input the operation state evaluation vector and the environmental risk assessment vector into the linear weighting module of the anomaly correction policy template, and calculate the dynamic weight of the steering wheel angle correction parameter and the influence factor of the environmental risk assessment vector, where the dynamic weight is adaptively adjusted according to the ratio of the standard deviation of the steering wheel angle to the lateral distance deviation, and the influence factor is generated based on the product of the reciprocal of the predicted collision time and the relative speed variance of the target vehicle.

[0076] Input the operation state evaluation vector V_op_eval and the environmental risk assessment vector V_env_eval into the linear weighting module of the anomaly correction policy template. For the dynamic weight w_Δα of the steering wheel angle correction parameter, it is adaptively adjusted according to the ratio of the standard deviation of the steering wheel angle σ_α_norm in the operation state evaluation vector to the lateral distance deviation Δd_std in the environmental risk assessment vector. For example, w_Δα = f(σ_α_norm / Δd_std), where f is a predefined function that can be designed according to specific requirements to achieve reasonable adjustment of the dynamic weight.

[0077] For the influence factor f_env of the environmental risk assessment vector, it is generated based on the product of the reciprocal of the predicted collision time TTC_std in the environmental risk assessment vector and the relative speed variance σ_rel_std of the target vehicle. That is, f_env = 1 / TTC_std * σ_rel_std. Through such calculations, the dynamic weight w_Δα of the steering wheel angle correction parameter and the influence factor f_env of the environmental risk assessment vector are obtained.

[0078] Step S154: Based on the dynamic weight and influence factor, perform proportional-integral correction on the steering wheel angle correction parameter to generate an initial angle correction amount. At the same time, calculate the attenuation gradient of the desired speed adjustment coefficient according to the difference between the average throttle depth and the safe vehicle distance threshold, and generate a dynamic braking response curve in combination with the time derivative of the brake pressure integral.

[0079] Perform proportional-integral correction on the steering wheel angle correction parameter Δα_template in the abnormal correction strategy template based on the dynamic weight w_Δα and the influence factor f_env. The proportional-integral correction includes a proportional term and an integral term. The proportional term adjusts Δα_template according to the dynamic weight and the influence factor, and the integral term accumulates the historical deviation. Let the proportional coefficient be k_p and the integral coefficient be k_i. The initial angle correction amount Δα_init can be calculated as follows: first calculate the proportional term k_p*w_Δα*f_env*Δα_template, then calculate the integral term k_i*∫(w_Δα*f_env*Δα_template)dt, and add the proportional term and the integral term to obtain the initial angle correction amount Δα_init.

[0080] For the desired speed adjustment coefficient k_v_template, calculate the attenuation gradient according to the difference between the average throttle depth μ_β_norm in the operation state evaluation vector and the safe vehicle distance threshold d_safe_template in the abnormal correction strategy template. Let the attenuation gradient be g_kv, g_kv = h(μ_β_norm - d_safe_template), where h is a predefined function.

[0081] Define the brake pressure integral as the cumulative value of the brake pressure within a time window. Let the brake pressure be P_brake(t), and the brake pressure integral I_γ = ∫P_brake(t)dt. The time derivative of the brake pressure integral is the instantaneous brake pressure P_brake(t). Generate a dynamic braking response curve in combination with the instantaneous brake pressure P_brake(t). The dynamic braking response curve describes the change of the brake pressure over time and can be designed according to specific driving scenarios and requirements. For example, it can be represented by a linear or non-linear function.

[0082] Step S155: Check the vehicle dynamics constraints for the initial angle correction amount, limit the initial angle correction amount according to the current vehicle speed and the maximum allowable range of the tire side slip angle to generate a safe angle correction instruction. At the same time, perform longitudinal acceleration smoothing filtering on the desired speed adjustment coefficient to generate a smoothed speed adjustment instruction.

[0083] Perform vehicle dynamics constraint verification on the initial steering angle correction Δα_init. The steering operation of the vehicle is restricted by the current vehicle speed v and the maximum allowable range δ_max of the tire sideslip angle. According to the vehicle dynamics model, the maximum sideslip angle that the tire can withstand is different at different vehicle speeds. Let the maximum allowable steering angle correction calculated based on the current vehicle speed v be Δα_max. If the initial steering angle correction Δα_init is greater than Δα_max, it is limited to Δα_max; if Δα_init is less than -Δα_max, it is limited to -Δα_max. After the limiting process, a safe steering angle correction command Δα_safe is generated.

[0084] For the desired speed adjustment coefficient k_v_template, perform longitudinal acceleration smoothing filtering. A sudden change in longitudinal acceleration will affect the ride comfort and driving stability of the vehicle, so it is necessary to smooth the desired speed adjustment coefficient. A low-pass filter can be used to filter k_v_template, such as a first-order low-pass filter, and its output is the smoothed desired speed adjustment coefficient k_v_smooth. A smooth speed adjustment command is generated based on k_v_smooth, and this command can make the speed adjustment of the vehicle smoother.

[0085] Step S156: Calculate the feedforward control term and feedback compensation term of the brake pressure correction based on the dynamic brake response curve and the relative speed variance of the target vehicle in the environmental risk assessment vector, where the feedforward control term is generated according to the deviation ratio between the safe distance threshold and the actual distance, and the feedback compensation term is dynamically adjusted based on the exponential decay function of the predicted time-to-collision value.

[0086] Based on the dynamic brake response curve and the relative speed variance σ_rel_std of the target vehicle in the environmental risk assessment vector, calculate the feedforward control term and feedback compensation term of the brake pressure correction. The feedforward control term is generated according to the deviation ratio between the safe distance threshold d_safe_template in the anomaly correction strategy template and the actual distance d_actual. Let the feedforward control coefficient be k_ff, and the feedforward control term ΔP_ff = k_ff * (d_safe_template - d_actual) / d_safe_template.

[0087] The feedback compensation term is dynamically adjusted based on the exponential decay function of the predicted collision time TTC_std in the environmental risk assessment vector. Let the feedback compensation coefficient be k_fb, and the exponential decay function be exp(-a*TTC_std), where a is a preset parameter. The feedback compensation term ΔP_fb = k_fb*σ_rel_std*exp(-a*TTC_std). Add the feedforward control term and the feedback compensation term to obtain the brake pressure correction ΔP_brake.

[0088] Step S157: Perform control instruction fusion on the safety corner correction instruction, smooth speed adjustment instruction, and brake pressure correction amount to generate multi-dimensional driving strategy optimization data. During the fusion process, a priority arbitration mechanism is used to sort the execution order of the steering instruction and the braking instruction, and phase compensation is performed on the instruction timing based on the response delay parameter of the vehicle control system.

[0089] Fuse the safety corner correction instruction Δα_safe, the smooth speed adjustment instruction k_v_smooth, and the brake pressure correction amount ΔP_brake. During the fusion process, a priority arbitration mechanism is used to sort the execution order of the steering instruction and the braking instruction. Generally, the priority of the braking instruction may be higher than that of the steering instruction, but in an emergency, synchronous adjustment may be required (such as steering + braking when avoiding an obstacle). The original step did not define a conflict resolution rule, which may lead to control chaos.

[0090] To solve this problem, a dynamic priority allocation mechanism is introduced. The priorities of the steering instruction and the braking instruction are adjusted in real time according to the collision risk. Let the collision risk assessment value be R_collision, which can be calculated from factors such as the predicted collision time TTC in the environmental interaction feature vector, e.g., R_collision = 1 / TTC. When R_collision is lower than a certain threshold, according to the normal priority, the priority of the braking instruction is higher than that of the steering instruction; when R_collision is higher than this threshold, it indicates a high collision risk, and both the steering and braking instructions need to be executed simultaneously. At this time, a cooperative control strategy can be adopted to coordinately adjust the steering instruction and the braking instruction to avoid instruction conflicts.

[0091] Meanwhile, based on the response delay parameter τ of the vehicle control system, phase compensation is performed on the instruction timing. Since the vehicle control system takes a certain amount of time to respond after receiving an instruction, it is necessary to adjust the sending time of the instruction. For example, for the brake pressure correction instruction, its sending time is advanced by τ time to ensure that the vehicle can respond in a timely manner. After priority arbitration and phase compensation, the safe corner correction instruction, smooth speed adjustment instruction, and brake pressure correction amount are spliced together to generate the multi-dimensional driving strategy optimization data D_optimize = [Δα_safe; k_v_smooth; ΔP_brake].

[0092] Step S158: Package the multi-dimensional driving strategy optimization data into a standardized control message according to the preset vehicle control bus protocol, and send it to the steering motor controller, throttle servo unit, and electronic brake control module respectively through the actuator interface of the vehicle control system, trigger the real-time driving strategy adjustment operation, and at the same time send the adjusted vehicle operation feature vector back to the feature extraction layer of the behavior analysis network for closed-loop feedback verification.

[0093] Package the multi-dimensional driving strategy optimization data D_optimize according to the preset vehicle control bus protocol to generate a standardized control message M_control. The preset vehicle control bus protocol stipulates the data format, transmission rate, communication rules, etc. Through the actuator interface of the vehicle control system, send the standardized control message M_control to the steering motor controller, throttle servo unit, and electronic brake control module respectively. The steering motor controller adjusts the steering wheel angle according to the received safe corner correction instruction Δα_safe; the throttle servo unit adjusts the throttle opening according to the smooth speed adjustment instruction k_v_smooth; the electronic brake control module adjusts the brake pressure according to the brake pressure correction amount ΔP_brake, thereby triggering the real-time driving strategy adjustment operation.

[0094] Meanwhile, send the adjusted vehicle operation feature vector V_op_updated back to the feature extraction layer of the behavior analysis network. The behavior analysis network will perform operations such as behavior feature extraction and abnormal behavior recognition again based on the received adjusted vehicle operation feature vector, evaluate the effect of the driving strategy adjustment, and achieve closed-loop feedback verification. If it is found that there are still abnormal behaviors or the effect of the driving strategy adjustment is not ideal, the driving strategy will be optimized and adjusted again until a satisfactory driving effect is achieved.

[0095] Step S210: Obtain a sample multi-source sensor data set and a basic behavior analysis network. The sample multi-source sensor data set includes a multi-source fusion data set that has completed spatio-temporal alignment processing, and the sample multi-source sensor data set is labeled with corresponding driving behavior category labels.

[0096] To train the basic behavior analysis network, it is necessary to obtain a sample multi-source sensor data set and the basic behavior analysis network. The process of obtaining the sample multi-source sensor data set is similar to the previous steps S110 - S120. First, the raw data of different sensors is collected, and then spatio-temporal alignment processing is performed to generate a multi-source fusion data set. The difference is that the sample multi-source sensor data set is also labeled with corresponding driving behavior category labels. For example, "abnormal following strategy under complex road conditions". In urban congested road conditions, the vehicle needs to start and stop frequently and adjust the following distance. Normal following behavior needs to dynamically adjust its own speed and spacing according to the speed, acceleration of the vehicle in front and the distance from the vehicle in front. When there are situations such as inappropriate following distance maintenance, frequent and unnecessary acceleration and deceleration, it can be determined as abnormal following strategy. Another example is "abnormal driving trajectory in a curve". When driving in a curve, the vehicle should maintain a reasonable driving trajectory according to the curvature, slope of the curve and its own speed. If the vehicle shows excessive deviation from the center of the lane or an overly tortuous driving trajectory in the curve, it belongs to abnormal driving trajectory in a curve. Another example is "improper timing and method of lane change on the highway". When changing lanes on the highway, factors such as the speed and distance of the vehicles in front and behind and the acceleration ability of the own vehicle need to be considered. If a forced lane change is made under unsafe conditions, or the speed changes abnormally during the lane change process, it can be determined as improper timing and method of lane change. These labels can be obtained through manual annotation or other reliable annotation methods and are used to represent the driving behavior category corresponding to each data sample. The basic behavior analysis network is an initial neural network model, and its structure and parameters still need to be optimized through training.

[0097] Step S220: Load the sample multi-source sensor data set into the feature extraction layer of the basic behavior analysis network to generate a sample driving behavior feature set.

[0098] Input the sample multi-source sensor data set into the feature extraction layer of the basic behavior analysis network. The structure and function of the feature extraction layer are the same as those described in step S130, including a vehicle operation feature extraction branch, an environmental interaction feature extraction branch, a behavior continuity feature extraction branch, and a fusion module. Through the processing of these branches and modules, a sample driving behavior feature set is extracted from the sample multi-source sensor data set, which also includes vehicle operation features, environmental interaction features, and behavior continuity features.

[0099] Step S230: Load the sample driving behavior feature set into the anomaly recognition layer of the basic behavior analysis network to generate a sample anomaly recognition prediction result.

[0100] Input the set of sample driving behavior characteristics into the anomaly recognition layer of the basic behavior analysis network. The structure and function of the anomaly recognition layer are the same as those described in step S140, including a feature association module, a multi-scale temporal convolutional module, a dynamic scoring module, a threshold generation sub-network, and a temporal correction module. Through the processing of these modules, anomaly behavior recognition is performed on the set of sample driving behavior characteristics to generate sample anomaly recognition prediction results. This result represents the anomaly behavior prediction of the basic behavior analysis network for each sample data.

[0101] Step S240: Calculate the first network learning cost based on the sample anomaly recognition prediction result and the driving behavior category label.

[0102] Calculate the first network learning cost according to the sample anomaly recognition prediction result and the driving behavior category label annotated by the sample multi-source sensor data set. Methods such as the cross-entropy loss function can be used to calculate the learning cost. Let the sample anomaly recognition prediction result be P_pred, the driving behavior category label be P_true, and the cross-entropy loss function be L_ce. For each sample data, calculate its cross-entropy loss L_ce_i, and then sum or average the cross-entropy losses of all sample data to obtain the first network learning cost L_first. The first network learning cost reflects the degree of difference between the prediction result of the basic behavior analysis network and the true label. The smaller the learning cost, the more accurate the network's prediction.

[0103] Step S250: Jointly train the feature extraction layer and the anomaly recognition layer of the basic behavior analysis network based on the first network learning cost to generate an optimized behavior analysis network.

[0104] Based on the first network learning cost L_first, jointly train the feature extraction layer and the anomaly recognition layer of the basic behavior analysis network. Use the backpropagation algorithm to calculate the gradient according to the first network learning cost, and then update the parameters of the feature extraction layer and the anomaly recognition layer according to the gradient. During the training process, continuously adjust the parameters of the network to gradually reduce the first network learning cost. After multiple iterative trainings, an optimized behavior analysis network is obtained. The feature extraction layer and the anomaly recognition layer of this network achieve the generalization extraction of driving behavior characteristics and the anomaly pattern recognition ability through parameter optimization.

[0105] Step S260: Use the optimized behavior analysis network as the preset behavior analysis network, where the feature extraction layer and the anomaly recognition layer of the preset behavior analysis network achieve the generalization extraction of driving behavior characteristics and the anomaly pattern recognition ability through parameter optimization.

[0106] The optimized behavior analysis network obtained through training is used as the preset behavior analysis network. This network can more accurately extract features related to driving behavior from multi-source sensor data in the feature extraction layer, including vehicle operation features, environmental interaction features, and behavior continuity features. In the anomaly recognition layer, it can more accurately identify abnormal patterns in driving behavior and generate accurate anomaly recognition results. In this way, the accuracy and reliability of the entire driving behavior analysis method are improved.

[0107] Step S310: Obtain an augmented example multi-source sensor data set, which is generated by performing data augmentation processing on the example multi-source sensor data set. The data augmentation processing includes at least one of the following operations: random time jitter, feature channel noise injection, and local feature masking.

[0108] To further improve the performance and generalization ability of the preset behavior analysis network, it is necessary to perform data augmentation processing on the existing example multi-source sensor data set to generate an augmented example multi-source sensor data set. The purpose of data augmentation processing is to increase the diversity of data so that the network can learn feature patterns in more different scenarios.

[0109] First, perform the random time jitter operation. Select a target data segment from the example multi-source sensor data set. Assume that the target data segment contains data from different sensors, such as data DA of sensor A, data DB of sensor B, etc. For the data DA of sensor A, perform random time jitter processing. The time offset of the random time jitter needs to be controlled within the range not exceeding the preset time threshold, and to avoid spatio-temporal misalignment of multi-sensor data, the jitter range is limited to be less than the sensor sampling interval. Assume the preset time threshold is T_thresh and the sensor sampling interval is T_sample. Randomly generate a time offset Δt, and -min(T_thresh, T_sample) ≤ Δt ≤ min(T_thresh, T_sample).

[0110] Translate the data DA of sensor A according to the time offset Δt to obtain the time-augmented data DA_time. Since time jitter may affect the accuracy of spatial coordinate transformation, spatial coordinate transformation needs to be performed again after time jitter. Assume the coordinate system of sensor A is SA, the unified vehicle coordinate system is SV, the original rotation matrix of sensor A relative to the unified vehicle coordinate system is RA, and the translation vector is TA. After time jitter, according to the new timestamp and the motion state of the sensor, recalculate the rotation matrix RA' and the translation vector TA', and transform the time-augmented data DA_time from the coordinate system SA to the unified vehicle coordinate system SV.

[0111] Next, a feature channel noise injection operation is performed. For the data DB of sensor B in the target data segment, feature channel noise injection processing is carried out. The amplitude of the injected noise cannot exceed a preset ratio of the original feature value. Let the preset ratio be r_noise. For each feature channel in the data DB, a noise vector N is randomly generated, and the amplitude of each element in the noise vector N does not exceed r_noise times the original feature value of this feature channel. The noise vector N is added to the data DB to obtain the noise-augmented data DB_noise. This can simulate the possible noise interference in the actual environment and make the network more robust to noise.

[0112] In addition, a local feature masking operation can also be performed. In the target data segment, some local regions are randomly selected for feature masking. For example, for the data DC of sensor C, one or more consecutive time steps and feature channels are randomly selected, and the data values in these regions are set to zero. Through the local feature masking operation, the network can learn more important feature information in the data and avoid over-reliance on certain local features.

[0113] Through at least one of the above data augmentation operations, each data segment in the sample multi-source sensor data set is processed, and finally an augmented sample multi-source sensor data set is generated. The data in this set not only contains the information of the original sample multi-source sensor data set, but also adds more different variation data samples, providing richer training materials for subsequent network training.

[0114] Step S320: Load the augmented sample multi-source sensor data set into the preset behavior analysis network to generate an augmented driving behavior feature set.

[0115] The generated augmented sample multi-source sensor data set is input into the preset behavior analysis network. The feature extraction layer of the preset behavior analysis network will process the augmented sample multi-source sensor data set, and its processing process is similar to that of step S130.

[0116] In the vehicle operation feature extraction branch, the vehicle control signals in the augmented sample multi-source sensor data set are normalized and multi-scale convolutional feature extraction is performed. For the vehicle control signals after random time jitter, feature channel noise injection or local feature masking processing, they are still normalized according to the previous method, and the steering wheel angle, throttle depth and brake pressure are mapped to a unified numerical interval. Then, convolutional kernels of different scales are used to capture the local patterns of direction control, acceleration and braking behaviors respectively, and a standardized vehicle operation feature vector is generated.

[0117] In the environmental interaction feature extraction branch, coordinate system alignment and dynamic target association analysis are performed on the environmental perception signals in the augmented sample multi-source sensor data set. Even though the environmental perception signals have undergone data augmentation processing, the three-dimensional coordinates of the camera-detected targets and the radar speed measurement data still need to be converted to the same vehicle coordinate system, and the lateral distance deviation between the host vehicle and the target vehicle, the predicted collision time, and the signal light status encoding are calculated to generate an environmental interaction feature vector.

[0118] In the behavior continuity feature extraction branch, sliding window statistic calculation and long short-term dependence modeling are performed on the time series signals in the augmented sample multi-source sensor data set. For the time series signals processed by time jitter and other methods, the normalized indicators of the steering wheel angle volatility, throttle change frequency, and brake duration are calculated, and a recurrent neural network is used to extract the temporal behavior pattern to generate a behavior continuity feature vector.

[0119] Finally, through the fusion module of the feature extraction layer, dimension alignment and weighted fusion are performed on the generated standardized vehicle operation feature vector, environmental interaction feature vector, and behavior continuity feature vector to generate an augmented driving behavior feature set. This set reflects the driving behavior features corresponding to the augmented sample multi-source sensor data set and contains more different feature patterns generated due to data augmentation.

[0120] Step S330: Calculate the second network learning cost based on the augmented anomaly recognition prediction result corresponding to the augmented driving behavior feature set and the driving behavior class label of the original sample multi-source sensor data set.

[0121] Input the augmented driving behavior feature set into the anomaly recognition layer of the preset behavior analysis network. The processing process of this layer is the same as that in step S140. Through the processing of the feature association module, multi-scale temporal convolutional module, dynamic scoring module, threshold generation sub-network, and temporal correction module, anomaly behavior recognition is performed on the augmented driving behavior feature set to generate an augmented anomaly recognition prediction result.

[0122] Then, calculate the second network learning cost according to the augmented anomaly recognition prediction result and the driving behavior class label of the original sample multi-source sensor data set. The cross-entropy loss function and other methods can also be used for calculation. Let the augmented anomaly recognition prediction result be P_pred_aug, and the driving behavior class label of the original sample multi-source sensor data set be P_true_ori. For each augmented data sample, calculate its cross-entropy loss L_ce_aug_i, and then sum or average the cross-entropy losses of all augmented data samples to obtain the second network learning cost L_second. The second network learning cost reflects the difference degree between the prediction result of the preset behavior analysis network when processing augmented data and the original true label, and it is used to evaluate the performance of the network when facing more diverse data.

[0123] Step S340: Based on the second network learning cost, perform incremental training on the preset behavior analysis network to generate an updated behavior analysis network.

[0124] Based on the calculated second network learning cost L_second, perform incremental training on the preset behavior analysis network. The purpose of incremental training is to further optimize the parameters of the network on the basis of the existing preset behavior analysis network so that it can better process various data patterns in the augmented example multi-source sensor data set.

[0125] Adopt the backpropagation algorithm to calculate the gradient according to the second network learning cost. The gradient represents the change direction and change degree of the network parameters under the current training data. According to the calculated gradient, update the parameters of the feature extraction layer and the anomaly recognition layer of the preset behavior analysis network. During the update process, parameters such as the learning rate can be used to control the step size of parameter update to avoid too fast or too slow parameter update.

[0126] During the incremental training process, continuously adjust the network parameters to gradually reduce the second network learning cost. After multiple iterative trainings, the network can learn new feature patterns and anomaly patterns in the augmented data, thereby improving its ability to identify driving behaviors in different scenarios. Finally, an updated behavior analysis network is generated, which has higher accuracy and generalization ability when processing multi-source sensor data for driving behavior analysis and anomaly recognition.

[0127] Step S410: The specific process of obtaining the example multi-source sensor data set includes collecting the original multi-source sensor data stream, where the original multi-source sensor data stream includes a first source data stream and a second source data stream, and the original multi-source sensor data stream is labeled with driving behavior category labels.

[0128] When obtaining the example multi-source sensor data set, first collect the original multi-source sensor data stream. Assume that the first source data stream comes from sensor X and the second source data stream comes from sensor Y. Sensor X may be a sensor for detecting the vehicle's surrounding environment, such as a lidar, which can collect information such as the distance and angle of objects around the vehicle in real time to form the first source data stream. Sensor Y may be a sensor for monitoring the vehicle's own state, such as an in-vehicle CAN bus sensor, which can obtain information such as the vehicle's speed, acceleration, and steering wheel angle to form the second source data stream.

[0129] While collecting these original data streams, it is necessary to label them with driving behavior category tags. These driving behavior category tags are more complex and difficult to directly judge than the previous examples. For example, "abnormal driving behavior at complex intersections". At complex multi-lane intersections, vehicles need to perform reasonable passing operations according to traffic lights, traffic signs, and the dynamics of other vehicles and pedestrians. If there are situations such as driving in the wrong lane or hesitating at the intersection causing traffic jams, it can be labeled as abnormal driving behavior at complex intersections. Another example is "abnormal driving stability in bad weather". Under bad weather conditions such as rainy days and snowy days, the road surface friction decreases, and the vehicle's handling performance will be affected. When driving normally, it is necessary to control the vehicle speed, brakes, and steering more carefully. If the vehicle experiences excessive skidding, sudden braking resulting in tire locking, etc., it belongs to abnormal driving stability in bad weather. Another example is "violation of parking and driving regulations in the parking lot". In the parking lot, vehicles need to drive and park according to the specified routes. If there are situations such as driving in reverse or randomly occupying the fire lane for parking, it can be determined as a violation of parking and driving regulations in the parking lot. Through such labeling, clear target tags are provided for subsequent model training, enabling the model to learn the sensor data characteristics corresponding to different complex driving behaviors.

[0130] Step S420: Perform timestamp alignment processing on the first source data stream and the second source data stream to generate a time-synchronized data stream.

[0131] Since the sampling frequencies and time bases of sensor X and sensor Y may be different, it is necessary to perform timestamp alignment processing on the first source data stream and the second source data stream. Let the sampling frequency of sensor X be fX, the sampling frequency of sensor Y be fY, and fX is not equal to fY.

[0132] First, determine a unified time base T. For each data point in the first source data stream, record its original timestamp tX_i. Similarly, for each data point in the second source data stream, record its original timestamp tY_j. Then, according to the unified time base T, perform interpolation resampling on these data points.

[0133] For the first source data stream, if there is no directly corresponding sampled data point at a certain unified time point Tk, interpolation methods need to be used to estimate the data value at this time point. Linear interpolation can be adopted. Assume that Tk is between tX_m and tX_(m + 1), and according to the data values at tX_m and tX_(m + 1), estimate the data value at Tk using the linear interpolation formula. The same applies to the second source data stream.

[0134] After such interpolation resampling processing, the first source data stream and the second source data stream have the same time step under the unified time reference T, thereby generating a time-synchronized data stream. Each data point in this data stream corresponds to the same time point, providing temporal consistency for subsequent spatial coordinate transformation and data fusion.

[0135] Step S430: Perform spatial coordinate transformation processing on the time-synchronized data stream to generate a spatially aligned data stream.

[0136] The installation positions and coordinate systems of different sensors may be different. Therefore, it is necessary to perform spatial coordinate transformation processing on the time-synchronized data stream to unify it into the same coordinate system. Assume that the coordinate system of sensor X is SX, the coordinate system of sensor Y is SY, and the unified vehicle coordinate system is SV.

[0137] First, it is necessary to determine the transformation relationship of sensor X and sensor Y relative to the unified vehicle coordinate system. By measuring the installation position and attitude of the sensors, the rotation matrix RX and translation vector TX of sensor X relative to the unified vehicle coordinate system, and the rotation matrix RY and translation vector TY of sensor Y relative to the unified vehicle coordinate system can be obtained.

[0138] For the data point PX of sensor X in the time-synchronized data stream (in the coordinate system SX), to convert it to the unified vehicle coordinate system SV, first rotate PX by the rotation matrix RX to obtain the rotated vector PX', and then add the translation vector TX to obtain the coordinate PV_X in the unified vehicle coordinate system. For the data point PY of sensor Y (in the coordinate system SY), similarly rotate it by the rotation matrix RY to obtain PY', and then add the translation vector TY to obtain the coordinate PV_Y in the unified vehicle coordinate system.

[0139] Through such coordinate transformation, all data points in the time-synchronized data stream are unified into the vehicle coordinate system, generating a spatially aligned data stream. The data in this data stream is spatially consistent, facilitating subsequent feature-level fusion processing.

[0140] Step S440: Perform feature-level fusion processing on the spatially aligned data stream based on the source fusion strategy to generate a multi-source fusion data stream.

[0141] After obtaining the spatially aligned data stream, it is necessary to perform feature-level fusion processing on it based on the source fusion strategy. The source fusion strategy defines how to fuse the data of sensor X and sensor Y to make full use of the information of different sensors.

[0142] First, feature extraction is performed on the data of sensor X and sensor Y in the spatially aligned data stream. For the data of sensor X, the feature vector VX is extracted through the feature extraction algorithm EX; for the data of sensor Y, the feature vector VY is extracted through the feature extraction algorithm EY.

[0143] Then, according to the source fusion strategy, the feature vectors VX and VY are fused. A common fusion method is weighted concatenation. Weights wX and wY (wX + wY = 1) are assigned to the feature vectors VX and VY respectively, and the weighted feature vectors are concatenated to obtain the fused feature vector VF. For example, VF = [wX * VX; wY * VY], where ";" represents the concatenation operation.

[0144] Through such feature-level fusion processing, the data of sensor X and sensor Y are fused together to generate a multi-source fusion data stream, and each data point in it contains the fused feature information of sensor X and sensor Y.

[0145] Step S450: Perform time window segmentation processing on the multi-source fusion data stream to generate the sample multi-source sensor data set, where each data segment in the time window corresponds one-to-one with the driving behavior category label.

[0146] To convert the multi-source fusion data stream into a sample multi-source sensor data set suitable for model training, time window segmentation processing needs to be performed on it. The specific steps are as follows:

[0147] Step S451: Determine the behavior start time point and behavior end time point corresponding to the driving behavior category label.

[0148] For more complex driving behavior category labels, determining the behavior start and end time points requires more detailed analysis and avoiding problems caused by subjective annotation. Taking "abnormal driving trajectory in a curve" as an example, objective indicators are used to determine the start and end times. Set the lane departure distance threshold D_thresh and the duration threshold N seconds. When the vehicle enters the curve, the distance D between the vehicle's driving trajectory and the center of the lane is monitored in real time. When D is greater than D_thresh and the duration exceeds N seconds, it is considered that the abnormal driving trajectory in the curve behavior starts. When the vehicle exits the curve and D is less than D_thresh and lasts for a period of time (such as M seconds), it is considered that the abnormal behavior ends.

[0149] For "abnormal following vehicle strategy under complex road conditions", factors such as the relative distance and speed difference between the host vehicle and the preceding vehicle can be comprehensively considered. Set a relative distance threshold d_thresh and a speed difference threshold v_thresh. When the relative distance between the host vehicle and the preceding vehicle continuously remains less than d_thresh and the speed difference continuously remains greater than v_thresh for a certain period of time (such as P seconds), it is considered that the abnormal following vehicle strategy behavior starts. When the relative distance recovers above d_thresh and the speed difference recovers within v_thresh and lasts for a period of time (such as Q seconds), it is considered that the abnormal behavior ends.

[0150] For "improper timing and manner of lane change on highway", factors such as the turn signal of the vehicle, the speed change before and after the lane change, and the safe distance from the front and rear vehicles are combined. When the vehicle turns on the turn signal, if it starts to move under the condition of not meeting the safe lane change condition (such as the safe distance from the front and rear vehicles is less than the safety threshold s_thresh) and lasts for a period of time (such as R seconds), it is considered that the abnormality starts. When the vehicle completes the lane change and the relative position and speed relationship with the surrounding vehicles return to normal (such as the safe distance from the front and rear vehicles is greater than s_thresh) and lasts for a period of time (such as S seconds), it is considered that the abnormality ends.

[0151] Step S452: Using the behavior start time point and the behavior end time point as boundaries, intercept data segments from the multi-source fusion data stream that match the driving behavior category labels.

[0152] Perform standardization processing on the intercepted data segments to improve the quality and consistency of the data. Taking the data segment of "abnormal following vehicle strategy under complex road conditions" as an example, first perform denoising processing to remove possible noise interference in the sensor data. Filtering algorithms such as moving average filtering and median filtering can be used to smooth data such as the speed, acceleration of the host vehicle, and the relative distance from the preceding vehicle, and remove high-frequency noise. Then perform normalization processing to map each eigenvalue in the data to a unified numerical interval. For example, for the vehicle speed eigenvalue v, it can be normalized to the [0, 1] interval through linear transformation, that is, v_norm=(v - v_min) / (v_max - v_min), where v_min and v_max are the minimum and maximum values of the speed eigenvalue in the data segment respectively. The same normalization process is carried out for other features such as acceleration and relative distance. Finally, perform feature scaling. According to the importance and variation range of different features, the eigenvalue is appropriately scaled. For example, for the relative distance feature, since its variation range may be large, it can be scaled to have a similar scale to features such as speed and acceleration. Through these standardization processes, standardized behavior data segments are generated. The data segments of behaviors such as "abnormal driving trajectory on a curve" and "improper timing and manner of lane change on highway" are also subjected to denoising, normalization, and feature scaling processing.

[0153] Step S453: Perform normalization processing on the data segment to generate a normalized behavior data segment. The normalization processing includes denoising, normalization, and feature scaling.

[0154] Perform normalization processing on the intercepted data segment to improve the quality and consistency of the data. First, perform denoising processing to remove possible noise interference in the data. Filtering algorithms such as moving average filtering and median filtering can be used to smooth the data and remove high-frequency noise.

[0155] Then perform normalization processing to map each feature value in the data to a unified numerical interval. For example, for a certain feature value x in the data segment, it can be normalized to the [0, 1] interval through linear transformation, that is, x_norm = (x - x_min) / (x_max - x_min), where x_min and x_max are the minimum and maximum values of the feature value in the data segment, respectively.

[0156] Finally, perform feature scaling. According to the importance and variation range of different features, scale the feature values appropriately. For example, for some features with a large variation range, scale them to have a similar scale to other features. Through these normalization processes, a normalized behavior data segment is generated.

[0157] Step S454: Associate and store the normalized behavior data segment with the corresponding driving behavior category label to generate the example multi-source sensor data set.

[0158] Associate the generated normalized behavior data segment with the corresponding driving behavior category label. Data structures such as dictionaries or database records can be used to store the normalized behavior data segment as the value in the key-value pair and the corresponding driving behavior category label as the key. For example, associate and store the normalized behavior data segment corresponding to "abnormal following strategy under complex road conditions" with this label. Through such associated storage, an example multi-source sensor data set is formed. Each data segment in this example multi-source sensor data set corresponds to a clear driving behavior category label, providing accurate training data and labels for subsequent model training.

[0159] Figure 2 FIG. shows a schematic diagram of exemplary hardware and software components of a driving behavior analysis system 100 based on multi-modal sensor fusion that can implement the idea of the present application provided by some embodiments of the present application. For example, the processor 120 can be used on the driving behavior analysis system 100 based on multi-modal sensor fusion and is used to execute the functions in the present application.

[0160] The driving behavior analysis system 100 based on multi-modal sensor fusion may be a general-purpose server or a special-purpose server, both of which can be used to implement the driving behavior analysis method based on multi-modal sensor fusion of the present application. Although only one server is shown in the present application, for convenience, the functions described in the present application can be implemented in a distributed manner on multiple similar platforms to balance the processing load.

[0161] For example, the driving behavior analysis system 100 based on multi-modal sensor fusion may include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and different forms of storage media 140, such as disks, ROM, or RAM, or any combination thereof. Exemplarily, the driving behavior analysis system 100 based on multi-modal sensor fusion may also include program instructions stored in ROM, RAM, or other types of non-transitory storage media, or any combination thereof. The method of the present application can be implemented according to these program instructions. The driving behavior analysis system 100 based on multi-modal sensor fusion further includes an I / O interface 150 between the computer and other input / output devices.

[0162] For ease of explanation, only one processor is described in the driving behavior analysis system 100 based on multi-modal sensor fusion. However, it should be noted that the driving behavior analysis system 100 based on multi-modal sensor fusion in the present application may also include multiple processors. Therefore, the steps executed by one processor described in the present application can also be jointly executed or separately executed by multiple processors. For example, if the processor of the driving behavior analysis system 100 based on multi-modal sensor fusion executes steps A and B, it should be understood that steps A and B can also be jointly executed by two different processors or separately executed in one processor. For example, the first processor executes step A, the second processor executes step B, or the first processor and the second processor jointly execute steps A and B.

[0163] In addition, an embodiment of the present invention further provides a readable storage medium, in which computer-executable instructions are preset. When the processor executes the computer-executable instructions, the driving behavior analysis method based on multi-modal sensor fusion as described above is implemented.

[0164] It should be noted that, in order to simplify the description of the present invention disclosure and thus help the understanding of one or more embodiments of the invention, in the foregoing description of the embodiments of the present invention, sometimes multiple features are merged into one embodiment, drawing, or description thereof.

Claims

1. A driving behavior analysis method based on multi-modal sensor fusion, characterized in that, The method includes: Obtaining a multi-source sensor data set of a target vehicle, where the multi-source sensor data set includes driving environment perception data from at least two different sensor sources; Performing spatio-temporal alignment processing on the multi-source sensor data set to generate a multi-source fusion data set; Based on the feature extraction layer of a preset behavior analysis network, performing behavior feature extraction processing on the multi-source fusion data set to obtain a driving behavior feature set, where the driving behavior feature set includes vehicle operation features, environment interaction features, and behavior continuity features; Based on the anomaly recognition layer of a preset behavior analysis network, performing abnormal behavior recognition processing on the driving behavior feature set to generate a driving behavior anomaly recognition result; Generating driving strategy optimization data according to the driving behavior anomaly recognition result and feeding the driving strategy optimization data back to the vehicle control system to trigger a driving strategy adjustment operation; The performing spatio-temporal alignment processing on the multi-source sensor data set to generate a multi-source fusion data set includes: Performing timestamp alignment processing on each source data in the multi-source sensor data set to generate a time synchronization data set, where the timestamp alignment processing includes performing interpolation resampling on the source data with different sampling frequencies to achieve a unified time reference; Performing spatial coordinate transformation processing on the time synchronization data set to generate a spatial alignment data set, where the spatial coordinate transformation processing includes mapping the coordinate systems of different sensors to a unified vehicle coordinate system; Based on a preset source fusion rule, performing feature-level fusion processing on the data from different sensor sources in the spatial alignment data set to generate the multi-source fusion data set, and each data unit in the multi-source fusion data set includes a fusion feature vector of at least two source data; The method further includes: Obtaining a sample multi-source sensor data set and a basic behavior analysis network, where the sample multi-source sensor data set includes a multi-source fusion data set that has completed spatio-temporal alignment processing, and the sample multi-source sensor data set is labeled with corresponding driving behavior category labels; Loading the sample multi-source sensor data set into the feature extraction layer of the basic behavior analysis network to generate a sample driving behavior feature set; Loading the sample driving behavior feature set into the anomaly recognition layer of the basic behavior analysis network to generate a sample anomaly recognition prediction result; Calculating a first network learning cost based on the sample anomaly recognition prediction result and the driving behavior category label; Jointly training the feature extraction layer and the anomaly recognition layer of the basic behavior analysis network based on the first network learning cost to generate an optimized behavior analysis network; Using the optimized behavior analysis network as the preset behavior analysis network, where the feature extraction layer and the anomaly recognition layer of the preset behavior analysis network achieve the generalization extraction of driving behavior features and the abnormal pattern recognition ability through parameter optimization.

2. The driving behavior analysis method based on multi-modal sensor fusion according to claim 1, wherein The obtaining the sample multi-source sensor data set includes: Collecting an original multi-source sensor data stream, where the original multi-source sensor data stream includes a first source data stream and a second source data stream, and the original multi-source sensor data stream is labeled with driving behavior category labels; Perform timestamp alignment processing on the first source data stream and the second source data stream to generate a time-synchronized data stream; Perform spatial coordinate transformation processing on the time-synchronized data stream to generate a spatially aligned data stream; Perform feature-level fusion processing on the spatially aligned data stream based on a source fusion strategy to generate a multi-source fusion data stream; Perform time window segmentation processing on the multi-source fusion data stream to generate the example multi-source sensor data set, where each data segment in a time window corresponds one-to-one with the driving behavior category label.

3. The driving behavior analysis method based on multi-modal sensor fusion according to claim 2, characterized in that The performing time window segmentation processing on the multi-source fusion data stream to generate the example multi-source sensor data set includes: Determine the behavior start time point and the behavior end time point corresponding to the driving behavior category label; Using the behavior start time point and the behavior end time point as boundaries, intercept data segments that match the driving behavior category label from the multi-source fusion data stream; Perform normalization processing on the data segments to generate normalized behavior data segments, and the normalization processing includes denoising, normalization, and feature scaling; Associate and store the normalized behavior data segments with the corresponding driving behavior category labels to generate the example multi-source sensor data set.

4. The driving behavior analysis method based on multi-modal sensor fusion according to claim 1, characterized in that, The method further includes: Obtain an augmented example multi-source sensor data set, which is generated by performing data augmentation processing on the example multi-source sensor data set, and the data augmentation processing includes at least one of the following operations: random time jitter, feature channel noise injection, local feature masking; Load the augmented example multi-source sensor data set into the preset behavior analysis network to generate an augmented driving behavior feature set; Calculate a second network learning cost based on the augmented anomaly recognition prediction result corresponding to the augmented driving behavior feature set and the driving behavior category label of the original example multi-source sensor data set; Perform incremental training on the preset behavior analysis network based on the second network learning cost to generate an updated behavior analysis network.

5. The driving behavior analysis method based on multi-modal sensor fusion according to claim 4, wherein The performing data augmentation processing on the example multi-source sensor data set includes: Randomly select target data segments from the example multi-source sensor data set; Perform random time jitter processing on the first source data in the target data segment to generate time-augmented data, and the time offset of the random time jitter processing does not exceed a preset time threshold; Perform feature channel noise injection processing on the second source data in the target data segment to generate noise-augmented data, where the injected noise amplitude does not exceed a preset ratio of the original feature value; Perform feature fusion on the time-augmented data and the noise-augmented data to generate an augmented data segment; Associate the augmented data segment with the original driving behavior category label and add it to the example multi-source sensor data set to generate the augmented example multi-source sensor data set.

6. The driving behavior analysis method based on multi-modal sensor fusion according to claim 1, characterized in that, The performing behavior feature extraction processing on the multi-source fusion data set based on the feature extraction layer of the preset behavior analysis network to obtain a driving behavior feature set includes: Through the vehicle operation feature extraction branch of the feature extraction layer, the vehicle control signals in the multi-source fusion data set are normalized and multi-scale convolutional features are extracted to generate a standardized vehicle operation feature vector. Among them, the normalization process includes mapping the steering wheel angle, throttle depth, and brake pressure to a unified numerical interval, and the multi-scale convolutional feature extraction uses convolutional kernels of different scales to capture local patterns of direction control, acceleration, and braking behaviors respectively; Through the environmental interaction feature extraction branch of the feature extraction layer, the environmental perception signals in the multi-source fusion data set are subjected to coordinate system alignment and dynamic target association analysis to generate an environmental interaction feature vector. Among them, the coordinate system alignment includes converting the three-dimensional coordinates of the camera detection target and the radar speed measurement data to the same vehicle coordinate system, and the dynamic target association analysis includes calculating the lateral distance deviation between the host vehicle and the target vehicle, the predicted collision time value, and the signal light state encoding; Through the behavior continuity feature extraction branch of the feature extraction layer, the time series signals in the multi-source fusion data set are subjected to sliding window statistic calculation and long short-term dependence modeling to generate a behavior continuity feature vector. Among them, the sliding window statistic calculation includes extracting normalized indexes of the steering wheel angle volatility, throttle change frequency, and brake duration, and the long short-term dependence modeling uses a recurrent neural network to extract the temporal behavior pattern; Through the fusion module of the feature extraction layer, the standardized vehicle operation feature vector, environmental interaction feature vector, and behavior continuity feature vector are dimensionally aligned and weighted fused to generate the driving behavior feature set. Among them, the dimensional alignment includes projecting each feature vector to a unified dimensional space using a fully connected layer, and the weighted fusion includes dynamically allocating weights to the vehicle operation features and environmental interaction features based on the attention mechanism, and applying a time decay factor to the behavior continuity features to strengthen the influence of recent behaviors.

7. The driving behavior analysis method based on multi-modal sensor fusion according to claim 6, characterized in that, The abnormal recognition layer based on the preset behavior analysis network performs abnormal behavior recognition processing on the driving behavior feature set to generate a driving behavior abnormal recognition result, including: Input the vehicle operation feature vector, environmental interaction feature vector, and behavior continuity feature vector in the driving behavior feature set into the feature association module of the abnormal recognition layer, and generate a first association weight matrix between the vehicle operation feature vector and the environmental interaction feature vector, and a second association weight matrix between the behavior continuity feature vector and the vehicle operation feature vector through the multi-head attention mechanism; Based on the first association weight matrix, the environmental interaction feature vector is weighted aggregated to generate a first fusion feature vector aligned with the vehicle operation feature vector in dimension. At the same time, based on the second association weight matrix, the behavior continuity feature vector is weighted aggregated to generate a second fusion feature vector aligned with the vehicle operation feature vector in dimension; Concatenate the vehicle operation feature vector, the first fusion feature vector, and the second fusion feature vector along the channel dimension, and input them into the multi-scale temporal convolutional module of the anomaly recognition layer. Extract driving behavior patterns at different time scales through parallel dilated convolutional layers, and perform max pooling and channel fusion on the temporal features output by each dilated convolutional layer to generate a multi-scale fusion feature tensor; Input the multi-scale fusion feature tensor into the dynamic scoring module of the anomaly recognition layer, and map it to the anomaly probability value of the current time window through a fully connected layer. Among them, the weight parameters of the dynamic scoring module are adaptively adjusted according to the fluctuation amplitude of the vehicle operation feature vector and the target density of the environmental interaction feature vector, and the output layer uses a non-linear activation function to generate an anomaly score; Input the number of dynamic targets, relative speed variance in the environmental interaction feature vector, and the standard deviation of the steering wheel angle in the vehicle operation feature vector into the threshold generation sub-network of the anomaly recognition layer, and calculate the dynamic anomaly determination threshold related to the scene through a fully connected layer; Perform sliding window cumulative calculation on the anomaly scores of consecutive time windows. When the cumulative score exceeds the dynamic anomaly determination threshold and the duration meets the preset conditions, trigger the anomaly behavior flag and generate an initial anomaly determination sequence; Input the initial anomaly determination sequence into the temporal correction module of the anomaly recognition layer, capture the time dependence of historical anomaly flags through a gated recurrent unit network, and combine the time decay factor in the behavior continuity feature vector to correct the probability of the anomaly flag in the current time window, and output the smoothed driving behavior anomaly recognition result.

8. A driving behavior analysis system based on multi-modal sensor fusion, characterized in that, It includes a processor and a memory. The memory is connected to the processor. The memory is used to store programs, instructions, or codes, and the processor is used to execute the programs, instructions, or codes in the memory to implement the driving behavior analysis method based on multi-modal sensor fusion according to any one of claims 1-7 above.

Citation Information

Patent Citations

  • Automatic driving vehicle behavior prediction method and system based on artificial intelligence

    CN119578658A