Automatic driving decision-making method and system based on multi-modal heterogeneity feature fusion
The improved LSTM network with segmentation algorithm and multimodal feature fusion solves the problems of insufficient modeling of driver behavior mode switching and discontinuous decision-making in existing technologies, realizes a high-precision and highly adaptable autonomous driving decision-making method, and improves the prediction performance of the autonomous driving system.
Patent Information
- Application Number
- CN202510980488.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-07-16
AI Technical Summary
Existing deep learning algorithms cannot effectively capture the driver's behavioral mode switching process under different driving states such as acceleration, deceleration and steady-state following in autonomous driving decision-making, and lack explicit modeling of driving state transitions and discontinuous decision-making processes, resulting in insufficient prediction accuracy and adaptability of the model in complex driving environments.
A segmentation algorithm is used to divide vehicle trajectories. Newell stimulus-response theory and dynamic time warping algorithm are combined to identify the following flow state and free flow state. An improved LSTM network is constructed. Through the coordinated operation of GRU and LSTM, multimodal feature fusion is achieved. The model is trained to realize driving state prediction and following behavior modeling.
It improves the model's ability to capture and model complex driving behaviors, enhances prediction accuracy and adaptability, and enhances the scenario generalization capability of the autonomous driving system, especially the ability to reproduce the dynamic characteristics of traffic flow at the micro and macro levels.
Smart Images

Figure CN120735797A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of autonomous driving technology, and in particular to an autonomous driving decision-making method and system based on multimodal heterogeneous feature fusion. Background Art
[0002] The autonomous driving decision-making model is the core cognitive engine of the intelligent transportation system, and its core task is to analyze the dynamic evolution mechanism of vehicle interaction behavior. Within the traditional control theory framework, the optimal velocity model (OVM) generates motion commands by establishing a deterministic mapping relationship between vehicle distance and speed, while the intelligent driver model (IDM) constructs vehicle dynamics equations based on a preset safety distance threshold. These models provide a mathematical foundation for traffic flow stability analysis, but their linear modeling paradigm is difficult to adapt to the nonlinear decision-making process in real driving scenarios. In particular, in highly dynamic interactive scenarios such as congestion propagation and ramp merging, the trajectory prediction error of traditional methods exceeds 35%, exposing a lack of ability to express the multimodal characteristics of driving behavior.
[0003] With the advancement of multi-source perception fusion and vehicle networking technologies, deep learning-based autonomous driving decision-making methods are triggering a paradigm shift. Temporal neural networks, such as LSTM and GRU, significantly improve motion planning accuracy compared to traditional methods by constructing memory encodings of driving scenarios. While this data-driven paradigm breaks through the rigid constraints of rule-based models, it still faces two challenges in engineering deployment. First, the heterogeneity of driving strategies encompasses both inter-subject differences (such as aggressive versus defensive decision-making styles) and intra-subject dynamics (such as scenario-adaptive switching between following and lane-changing modes). Existing methods often use discrete classifiers for feature extraction. Second, current architectures often rely on static feature space construction and lack the ability to online identify the continuously evolving driving intent, resulting in decision lag rates as high as 28.6% under complex road conditions.
[0004] Specifically, existing technology 1 is a car-following behavior modeling method based on an LSTM network. Its technical solution includes: first, using the NGSIM vehicle trajectory dataset, extracting headway, relative speed, and own vehicle speed as input features; then, constructing an LSTM neural network architecture and establishing a vehicle acceleration prediction model through time series learning; finally, using the mean squared error (MSE) as the loss function for model training, and verifying the model's accuracy by comparing the predicted results with real data. This technology captures the temporal dependencies of driving behavior through the memory effect of recurrent neural networks, outperforming traditional theoretical models in reproducing traffic oscillations. However, technology 1 inadequately represents heterogeneity in car-following modeling. It only implicitly captures driving style differences through continuous features such as headway, without explicitly quantifying the driver's behavioral fluctuations during acceleration and deceleration. This results in significant deviations in the model's predictions of the same driver's behavior in different traffic scenarios. Furthermore, technology 1 exhibits a fixed feature space, using fixed-dimensional continuous feature inputs, making it unable to represent sudden changes in driving state (such as discontinuous behavioral changes caused by emergency braking). This results in significant prediction lag in safety-critical scenarios such as accident warning.
[0005] The second existing technology is a car-following modeling method based on driving style classification. Its technical solution includes: using the K-means clustering algorithm to divide drivers into discrete categories such as aggressive and conservative; on this basis, training a dedicated LSTM prediction model for each type of driver; and then achieving personalized behavior prediction through cluster label matching. This method improves the model's ability to characterize heterogeneity among drivers through explicit driving style classification. However, the driver heterogeneity classification involved in technology 2 is coarse-grained. The static classification based on global trajectory features ignores the behavioral fluctuations of individual drivers in a single trip. Measured data shows that the acceleration distribution of the same driver at different times can vary by up to 32%. During the model training phase, the model's generalization is limited, and a separate model must be trained for each type of driver. In addition, the real-time performance is insufficient, relying on offline clustering of the complete trip trajectory, which cannot achieve online recognition and dynamic adjustment of driving status.
[0006] In summary, while existing deep learning algorithms have significantly improved prediction accuracy by incorporating temporal network technology, these methods still suffer from significant limitations when processing static feature inputs. Specifically, they are unable to effectively capture the driver's behavioral mode switching during different driving states, such as acceleration, deceleration, and steady-state following. This limitation makes it difficult for the models to fully reflect the driver's actual driving behavior. Furthermore, when addressing the heterogeneity of driving behavior, existing technologies often remain at a relatively coarse-grained classification level, failing to delve into more detailed analysis. Specifically, they neither quantitatively analyze the potential fluctuations in a driver's strategy within a single trip nor explicitly model driving state transitions and discontinuous decision-making processes. This imperfect approach severely limits the model's application value and effectiveness in scenarios requiring high-precision predictions, such as autonomous driving decision-making and traffic bottleneck analysis. Therefore, there is an urgent need to further improve and optimize existing deep learning methods to better adapt to complex and changing driving environments.
[0007] Therefore, there is an urgent and practical need to develop intelligent decision-making technologies that integrate dynamic heterogeneous features to enhance the scenario generalization capabilities of autonomous driving systems. Summary of the Invention
[0008] The purpose of the present invention is to address the core problems of insufficient heterogeneity representation and dynamic adaptability defects in existing car-following behavior modeling technologies, aiming to break through the inherent limitations of traditional data-driven methods and provide an autonomous driving decision-making method and system based on multimodal heterogeneous feature fusion.
[0009] The first object of the present invention is to provide an autonomous driving decision-making method based on multimodal heterogeneous feature fusion, comprising:
[0010] S1. Classify driving status, including:
[0011] S101, using a segmentation algorithm to perform feature segmentation on the speed and trajectory profile of the vehicle's driving trajectory;
[0012] S102: Based on Newell stimulus-response theory and dynamic time warping algorithm, establish dynamic response relationships between vehicles and identify the following flow state and free flow state;
[0013] S103, identifying a specific driving state;
[0014] S2. Build a car-following model using driving status, including:
[0015] S201, building a driving state prediction module framework;
[0016] S202. Build an improved LSTM network to achieve multimodal feature fusion;
[0017] S3. Train the car-following model and achieve collaborative training of driving state prediction and car-following behavior modeling through curriculum learning and multi-stage optimization strategies.
[0018] Preferably, S101 includes:
[0019] S1011, data preprocessing: performing median filtering on the original velocity sequence;
[0020] S1012, initial segment generation: divide the velocity sequence into atomic segments, each segment contains 5-10 consecutive sampling points
[0021] S1013, iterative merging optimization;
[0022] S1014: State interval integration: merging continuous segments whose slope change rate is less than a threshold.
[0023] Preferably, S1013 includes:
[0024] Calculate the slope matching degree of adjacent segments and apply a penalty coefficient of 0.3-0.6 when the slope signs are opposite;
[0025] Dynamically select the optimal segment pair for merging based on the merging cost function;
[0026] The iteration is stopped when the number of trajectory segments is reduced to 1.2-1.8 times the total duration.
[0027] Preferably, the threshold is 0.01 m / s 2 .
[0028] Preferably, S102 includes:
[0029] S1021, data preprocessing: perform spatiotemporal synchronization on the segmented trajectory data sets of the leading and trailing vehicles, and construct the velocity time series v of the trailing vehicle. L (t), time series of the preceding vehicle's speed v F (t), time series of the following vehicle position x L (t), the time series of the preceding vehicle's position x F (t);
[0030] S1022: Multi-dimensional DTW alignment, calculating the optimal matching path of the spatial trajectory:
[0031] DTW x =arg min π ∑ (i,j)∈π |x L (t i )-x(t j )|;
[0032] Get the matching paths on the two trajectories and k represents the serial number of the matching path; i represents the trajectory of the preceding vehicle, j represents the trajectory of the following vehicle, and π represents the set of all trajectories.
[0033] S1023, feature parameter extraction: analyzing the time delay distribution of the optimal curved path, extracting parameters characterizing the dynamic response between vehicles, and generating a speed delay distribution histogram;
[0034] S1024, state judgment criteria: establish a state judgment criteria based on τ n The state discriminant function of dynamic characteristics: close following state is CF, and free flow state is FF;
[0035]
[0036] in, is the time delay set τ n The 80% quantile of The time interval for matching paths between the leading and trailing vehicles.
[0037] Preferably, S103 includes:
[0038] S1031, obtaining segmented speed curves of the following vehicles CF and FF;
[0039] S1032. Slope classification: The segments are divided into two groups according to the slope, one group has a positive slope, representing acceleration behavior; the other group has a negative slope, representing deceleration behavior;
[0040] S1033, steady-state segment identification: classify line segments with slopes between -0.5 and 0.5 as constant-speed segments;
[0041] S1034: State classification, obtaining the driving state at each moment in the vehicle trajectory:
[0042] In the CF profile, the positive slope segment represents the acceleration following state A, the segment close to zero slope represents the constant speed following state F, and the negative slope segment represents the deceleration following state D;
[0043] In the FF profile, the positive slope segment represents the free acceleration Fa state, the segment close to zero slope represents the cruising state at the desired speed C state, and the FF segment does not include stationary or deceleration conditions.
[0044] Preferably, S201 includes:
[0045] A time series prediction model is constructed using GRU. The input layer receives the vehicle kinematic state vector X_{t-1}∈R^9 from the historical time window tk to t-1. The vehicle kinematic state vector includes the relative position of the following vehicle and the leading vehicle, the speed difference, the three-dimensional continuous features of the following vehicle's speed, and the one-hot encoding of the six-dimensional driving state. Through the coordinated operation of the GRU's update gate z_t and reset gate r_t, a hidden state update mechanism is established:
[0046]
[0047] Where z t represents the update gate, h t-1 represents the hidden state of the previous time step, ⊙ represents the Hadamard product, is the candidate hidden state;
[0048] The output layer uses the Softmax function to generate the driving state DR t The probability distribution of ∈{0,1,...,5} is expressed as:
[0049] P(DR t |X {t-1,...,t-k} )=Softmax(W g h t +b g )
[0050] Where, X {t-1,...,t-k} Represents the historical time series characteristics of the model input, W g Represents the weight matrix of the output layer, h t is the hidden state at the current moment, b g is the bias vector of the output layer.
[0051] Preferably, S202 includes:
[0052] S2021, driving status DR t Through the learnable embedding matrix W e ∈R 1×6 Mapped to a continuous latent space, generating a low-dimensional embedding vector E d =W e OneHot(DR t )+b e ; OneHot is the one-hot encoding operator, b e is the learnable bias;
[0053] S2022, embed the vector E d With continuous motion feature X c ∈R 3 Perform cross-modal splicing to generate fusion feature X fused =Concat(X c ,Ed )∈R 4 ;Concat is the concatenation operator;
[0054] S2023. Use the hyperbolic tangent function to update the cell state in the LSTM gate unit:
[0055] c t =f t ⊙c t-1 +i t ⊙tanh(W c [h t-1 ,x t ]+b c )
[0056] Where ⊙ represents the Hadamard product, f t represents the output of the forget gate, c t-1 Indicates the cell state at the previous moment, i t represents the output of the input gate, W c represents the weight matrix, h t-1 Indicates the hidden state at the previous moment, x t Indicates the current input, b c is the bias vector. The gradient saturation property of the tanh function is used to suppress the gradient explosion phenomenon in the recurrent neural network. The parameterized rectified linear unit PReLU is introduced into the output layer, and the activation function is defined as:
[0057]
[0058] Among them, α is a learnable parameter, α∈[0,1], which automatically optimizes the slope of the non-saturated region through back propagation;
[0059] S2024, using the Adam adaptive optimization algorithm to set key parameters;
[0060] S2025. Construct a composite loss function to achieve joint optimization of driving state classification and motion parameter regression.
[0061] A second object of the present invention is to provide an autonomous driving decision-making system based on multimodal heterogeneous feature fusion, comprising:
[0062] The data division module divides the driving status, including:
[0063] The feature division unit uses a segmentation algorithm to perform feature division on the speed and trajectory profile of the vehicle's driving trajectory;
[0064] The dynamic response relationship unit, based on Newell stimulus-response theory and dynamic time warping algorithm, establishes dynamic response relationships between vehicles and identifies the following flow state and free flow state;
[0065] Identification unit, identifies specific driving status;
[0066] Building modules to construct a car-following model using driving status, including:
[0067] Framework unit, building the driving state prediction module framework;
[0068] Fusion unit, building an improved LSTM network to achieve multimodal feature fusion;
[0069] The training module trains the car-following model and realizes the coordinated training of driving state prediction and car-following behavior modeling through curriculum learning and multi-stage optimization strategy.
[0070] The third object of the present invention is to provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the above-mentioned autonomous driving decision-making method based on multimodal heterogeneous feature fusion.
[0071] The fourth object of the present invention is to provide a computer program product, including a computer program, which, when executed by a processor, implements the above-mentioned autonomous driving decision-making method based on multimodal heterogeneous feature fusion.
[0072] The advantages and positive effects of this application are:
[0073] The present invention overcomes the following key technical difficulties: First, in response to the problem of dynamic quantitative characterization of driver heterogeneity, the present invention focuses on exploring how to accurately capture and quantify individual differences among drivers under different driving states, and on this basis, establishes a real-time mapping relationship between driving state transitions and neural network feature space, ensuring that the dynamic changes in driving states can be reflected in the model in a timely and accurate manner; second, in terms of constructing a multi-dimensional feature input architecture, the present invention aims to solve the modeling failure problem of traditional models when dealing with discontinuous behavior mutations. By fusing discrete driving states with continuous motion parameters, a comprehensive, multi-dimensional feature input system is constructed, thereby improving the model's ability to capture and model complex driving behaviors; third, the present invention focuses on the design of a local-global joint optimization method, striving to simultaneously optimize the accuracy of single-vehicle trajectory prediction while improving the ability to reproduce the dynamic characteristics of short-vehicle traffic flow, ensuring that the model has efficient prediction performance at both micro and macro levels. By organically combining dynamic driving state characteristics with the deep learning car-following model architecture, the present invention has achieved significant technological improvements at both the micro and macro levels in the field of vehicle car-following behavior modeling. It not only greatly improves the model's prediction accuracy and adaptability, but also provides a set of car-following behavior modeling tools with both high precision and strong adaptability for intelligent connected vehicles and traffic simulation platforms, laying a solid foundation for the optimization and development of future intelligent transportation systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0075] Figure 1 is a flow chart of a preferred embodiment of the present invention;
[0076] Figure 2 is a speed delay distribution histogram of a preferred embodiment of the present invention;
[0077] Figure 3 This is a flow chart of S103 in a preferred embodiment of the present invention.
[0078] Figure 4 is an architectural diagram of a car-following model in a preferred embodiment of the present invention;
[0079] Figure 5 This is a simulation diagram of a single trajectory in a preferred embodiment of the present invention;
[0080] Figure 6 2. A speed difference-distance oscillation comparison diagram of a pair of following vehicles in a preferred embodiment of the present invention;
[0081] Figure 7 This is a fleet simulation diagram in a preferred embodiment of the present invention. DETAILED DESCRIPTION
[0082] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0083] See also Figure 1 , an autonomous driving decision-making method based on multimodal heterogeneous feature fusion, comprising the following steps:
[0084] S1. Classify driving status, including:
[0085] S101, using a segmentation algorithm to perform feature segmentation on the speed and trajectory profile of the vehicle's driving trajectory;
[0086] S102: Based on Newell stimulus-response theory and dynamic time warping algorithm, establish dynamic response relationships between vehicles and identify the following flow state and free flow state;
[0087] S103, identifying a specific driving state;
[0088] S2. Build a car-following model using driving status, including:
[0089] S201, building a driving state prediction module framework;
[0090] S202. Build an improved LSTM network to achieve multimodal feature fusion;
[0091] S3. Train the car-following model and achieve collaborative training of driving state prediction and car-following behavior modeling through curriculum learning and multi-stage optimization strategies.
[0092] In order to better understand the concept of the present invention, the following non-limiting description is given:
[0093] The purpose of S1 is to divide the driving state of the trajectory, which includes the following steps:
[0094] S101: Use a segmentation algorithm to characterize the speed and trajectory profile of the vehicle's driving trajectory, including the following core steps:
[0095] S1011: Data preprocessing: Perform median filtering (window length 5 sampling points) on the original velocity sequence to eliminate sensor noise interference;
[0096] S1012: Initial segment generation: Divide the filtered sequence into atomic segments, each segment contains 5 to 10 consecutive sampling points
[0097] S1013: Iterative Merge Optimization:
[0098] (a) Calculate the slope matching degree of adjacent segments and apply a penalty coefficient of 0.3-0.6 when the slope signs are opposite;
[0099] (b) Dynamically select the optimal segment pair for merging based on the merging cost function;
[0100] (c) The iteration stops when the number of trajectory segments is reduced to 1.2-1.8 times the total duration.
[0101] S1014: State interval integration: slope change rate is less than 0.01m / s 2 Merge consecutive segments.
[0102] S102: Based on the segmented vehicle speed profile, a hybrid analysis method based on Newell stimulus-response theory and dynamic time warping algorithm is proposed to accurately identify the following flow state and free flow state by establishing the dynamic response relationship between vehicles. The core steps include the following:
[0103] S1021: Data preprocessing: Perform spatiotemporal synchronization on the segmented trajectory data sets of the leading and trailing vehicles to construct the speed time series v L (t), v F (t) and position time series x L (t), x F (t).
[0104] S1022: Multi-dimensional DTW alignment
[0105] Position sequence alignment, calculating the optimal matching path of the spatial trajectory
[0106]
[0107] Get the matching paths on the two trajectories k represents the sequence number of the matching path
[0108] S1023: Feature Parameter Extraction
[0109] By analyzing the time delay distribution of the optimal curved path, the key parameters characterizing the dynamic response between vehicles are extracted, and a speed delay distribution histogram is generated, such as Figure 2 .
[0110]
[0111] S1024: Status judgment criteria
[0112] Established based on τ n State discriminant function of dynamic characteristics: is the time delay set τ n The 80% quantile of τ is 3s in this experiment. n ≤3 seconds, it is judged as the close following state CF, when τ n >3 seconds, it is judged as free flow state FF;
[0113]
[0114] S103: After distinguishing the CF and FF segments, the next step is to accurately identify the specific driving states in these sections. Specifically, in the CF section, the states of accelerating and following A, decelerating and following D, following at a constant speed F, and stationary S need to be identified; while in the FF section, the states of free acceleration Fa and cruising at the desired speed C need to be identified. The process is as follows Figure 3 , including the following core steps:
[0115] S1031: Input processing: First, provide the segmented speed curves of the following vehicle in the CF and FF segments as input.
[0116] S1032: Slope classification: The segments are divided into two groups according to the slope: one group has a positive slope, representing acceleration behavior; the other group has a negative slope, representing deceleration behavior.
[0117] S1033: Steady-state segment identification: When the acceleration or deceleration is within 0.05g (g is the acceleration due to gravity), the segment is considered a steady-state (i.e., constant speed) segment. Therefore, a line segment with a slope between -0.5 and 0.5 is classified as a constant speed segment.
[0118] S1034: Status classification:
[0119] In the CF profile, the positive slope segment represents the acceleration following state A, the segment close to zero slope represents the constant speed following state F, and the negative slope segment represents the deceleration following state D.
[0120] In the FF profile, the positive slope segment represents the free acceleration Fa state, while the segment with a slope close to zero represents the cruising state at the desired speed C state. Obviously, the FF segment usually does not include stationary or deceleration conditions.
[0121] Through the above steps, the driving state of the vehicle at each moment in the trajectory is obtained, providing training features and labels for the subsequent deep learning-based coupled car-following model.
[0122] S2: Modeling of a car-following model embedded in driving status;
[0123] The present invention constructs a car-following behavior modeling system based on a hybrid deep learning architecture, which realizes deep coupling modeling of discrete driving states and continuous motion features through the collaborative work of the driving state prediction module and the dynamic feature fusion prediction module. Figure 4 As shown, the specific implementation includes the following steps:
[0124] S201: Driving state prediction module framework construction:
[0125] A time series prediction model is constructed using a gated recurrent unit (GRU). The input layer receives the vehicle kinematic state vector X_{t-1}∈R^9 from the historical time window tk to t-1. This vector contains the three-dimensional continuous features of the relative position, speed difference, and speed of the following vehicle, as well as the one-hot encoding of the six-dimensional driving state. Through the coordinated operation of the GRU's update gate z_t and reset gate r_t, a hidden state update mechanism is established as shown in Equation (1):
[0126]
[0127] Where ⊙ represents the Hadamard product, is the candidate hidden state. The output layer uses the Softmax function to generate the driving state DR t The probability distribution of ∈{0,1,...,5} is expressed as:
[0128] P(DR t |X {t-1,...,t-k} )=Softmax(W g h t +b g )
[0129] S202: Construction of dynamic feature fusion prediction module;
[0130] Construct an improved LSTM network to achieve multimodal feature fusion. The specific processing flow includes:
[0131] ① Discrete feature embedding encoding: driving state DR t Through the learnable embedding matrix W e ∈R 1×6 Mapped to a continuous latent space, generating a low-dimensional embedding vector E d =W e OneHot(DR t )+b e
[0132] ② Multimodal feature concatenation: embedding vector E d With continuous motion feature X c ∈R 3 Perform cross-modal splicing to generate fusion feature X fused =Concat(X c ,E d )∈R 4
[0133] ③ Nonlinear activation mechanism
[0134] The hyperbolic tangent function is used to update the cell state in the LSTM gating unit:
[0135] c t =f t ⊙c t-1 +i t ⊙tanh(W c [h t-1 ,x t ]+b c )
[0136] Where ⊙ represents the Hadamard product. The gradient saturation property of the tanh function can suppress the gradient explosion phenomenon in recurrent neural networks. The output layer introduces the parameterized rectified linear unit (PReLU), and its activation function is defined as:
[0137]
[0138] The parameter α∈[0,1] can be learned and the slope of the non-saturated region can be automatically optimized through back propagation.
[0139] ④Optimizer configuration
[0140] The Adam adaptive optimization algorithm is used, and the key parameters are set as follows:
[0141] Initial learning rate: Optimal value η = 10^{-3} is determined by grid search in the range {10^{-1}, 10^{-2}, 10^{-3}, 10^{-4}, 10^{-5}};
[0142] Momentum parameters: β_1 = 0.95, β_2 = 0.9999;
[0143] Batch size: 128 is selected based on orthogonal experiments to balance video memory usage and gradient estimation variance.
[0144] ⑤Multi-objective loss function
[0145] Construct a composite loss function to achieve joint optimization of driving state classification and motion parameter regression:
[0146] L total =λL cls +(1-λ)L reg
[0147] in:
[0148] Classification loss L cls Use label smoothed cross entropy:
[0149]
[0150] Where, the smoothing label q c = 1-ε when c is the true category, otherwise ε=0.1 is the smoothing coefficient;
[0151] Regression loss L reg Comprehensive acceleration, velocity and spacing errors:
[0152]
[0153] a represents acceleration, v represents velocity, △x represents spacing, and N represents the number of time points in each trajectory;
[0154] Regularization strategy, implement the following measures to prevent overfitting:
[0155] Hidden layer Dropout: random dropout with probability p = 0.2 is inserted between LSTM layers;
[0156] Gradient clipping: Set the gradient norm threshold to 5.0 to suppress gradient outliers;
[0157] Weight decay: Add an L2 regularization term with a coefficient of λ = 10^{-4}.
[0158] S3: Training. The hierarchical progressive training method proposed in this invention achieves collaborative training of driving state prediction and car-following behavior modeling through curriculum learning and multi-stage optimization strategy, including the following steps:
[0159] (1) Training phase division
[0160] Adopting a three-stage course learning framework:
[0161] ① Initial stage (0-50 training rounds): Freeze the LSTM network parameters and optimize the GRU driving state prediction network separately. The objective function is:
[0162]
[0163] Among them L cls represents the label smoothed cross entropy loss function, f GRU (x i ; θ) represents the predicted DR, Indicates the label value of the actual DR;
[0164] ②Medium stage (50-100 rounds): Jointly optimize the GRU-LSTM network, the objective function is:
[0165] (θ * ,γ * )=arg min θ,γ λL CE +(1-λ)L global
[0166] Where λ = 0.3 is the trade-off coefficient.
[0167] ③ Later stage (100+ rounds): Freeze GRU parameters and focus on optimizing the LSTM dynamics prediction network:
[0168]
[0169] in represents the true trajectory label value, Represents the trajectory prediction value
[0170] (2) Optimization strategy implementation
[0171] Implement local-global joint optimization:
[0172] ① Local single-step optimization: based on real observation sequence Predict the state at time t+1:
[0173]
[0174] ②Global multi-step optimization: recursive use of predicted values Building closed-loop predictions:
[0175]
[0176] ③Optimization process: First perform K=50 rounds of local optimization, when the single-step prediction MAE is <0.15m / s 2 Start global optimization (3) Specific training implementation
[0177] The driving state prediction network training process includes:
[0178] ① Data sampling: randomly extract batch B from the car-following trajectory dataset D;
[0179] ② Input construction: constructing time series input vector
[0180] ③Loss calculation: Use label smoothed cross entropy loss:
[0181] The dynamic feature fusion network training process includes:
[0182] ① Phase 1 (local optimization):
[0183] Based on real state
[0184] Predicting driving status:
[0185] And generate the control volume:
[0186] Calculate the single-step loss:
[0187] ②Phase 2 (global optimization):
[0188] Initialization state:
[0189] Implement closed-loop state recursion:
[0190]
[0191] ③Multi-objective loss calculation:
[0192]
[0193] By integrating dynamic driving state characteristics with a deep learning car-following model architecture, this invention achieves significant technical improvements in the field of car-following behavior modeling at both the micro and macro levels:
[0194] 1. Breakthrough in the accuracy of micro-behavior modeling
[0195] (1) Improved multi-index prediction accuracy
[0196] As shown in the table, this solution achieves all-round performance breakthroughs compared to traditional models:
[0197] Acceleration prediction: The MSE-a of LSTM-DR (0.375) is 58.47%, 56.45%, 55.41%, and 55.09% lower than that of IDM (0.903), RNN (0.861), GRU (0.841), and LSTM (0.835), respectively.
[0198] Speed prediction: The MSE-v of LSTM-DR (0.684) is 48.38%, 23.40%, 20.09%, and 15.66% lower than that of IDM (1.325), RNN (0.893), GRU (0.856), and LSTM (0.811), respectively.
[0199] Position prediction: The MSE-x of LSTM-DR (19.25) is 56.29%, 30.79%, 26.40%, and 26.10% lower than that of IDM (44.05), RNN (27.812), GRU (26.154), and LSTM (26.05), respectively.
[0200] Table 1 is a comparison table of micro-evaluation indicators
[0201] Model MSE-a Optimize ratio MSE-v Optimize ratio MSE-x Optimize ratio IDM 0.903 58.47% 1.325 48.38% 44.05 56.29% RNN 0.861 56.45% 0.893 23.40% 27.812 30.79% GRU 0.841 55.41% 0.856 20.09% 26.154 26.40% LSTM 0.835 55.09% 0.811 15.66% 26.05 26.10% LSTM-DR 0.375 —— 0.684 —— 19.25 ——
[0202] (2) Ability to adapt to all working conditions
[0203] like Figure 5 As shown in Figure 2, the single trajectory prediction accuracy of the present invention performs excellently in two typical scenarios; Figure 5 Three simulation cases are presented. Each figure contains three sub-figures, namely speed (unit: m / s), spacing (unit: m), and acceleration (unit: m / s2). The horizontal axis represents time (unit: 0.1s). Each sub-figure compares the LSTM model, IDM model, LSTM-DR model and empirical data (Empirical_data). The title is the MSE value of each model for the corresponding variable.
[0204] 2. Optimization of the ability to reproduce macroscopic traffic flow characteristics
[0205] The velocity disturbance propagation analysis based on motion wave theory shows that the present invention has significant advantages in macroscopic traffic characteristic modeling:
[0206] In terms of oscillation pattern matching, Figure 6 As shown, Figure 6 This is a simulation example. The figure compares the relationship between the space and speed of the original data, the IDM model, the LSTM model, and the LSTM-DR model. This model approximates the oscillation behavior more accurately and can more accurately capture the relative distance and speed fluctuations between vehicles.
[0207] In terms of capturing spatiotemporal fluctuations, in the lead car 539 queue (Table 2), the average position error (MAE-x = 6.0969) of this scheme is 14.0% lower than that of the LSTM (7.0881), and the speed error (MAE-v = 0.5640) is reduced by 50.1%. For the high-intensity oscillation scenario (lead car 745 queue, Table 3), the acceleration prediction MAE-a of this scheme is 0.3339, which is 1.7% better than the LSTM (0.3398), and the speed prediction error is reduced by 56.2% (MAE-v = 0.9203 vs 2.1016).
[0208] Figure 7 The simulation cases of two convoys (the first car is numbered 539 in the first picture and the second car is numbered 745) are presented, and the position and speed data of the original data, the IDM model, the LSTM model, and the LSTM-DR model are compared.
[0209] Table 2 shows the simulation error of the convoy with the lead vehicle number 539.
[0210]
[0211]
[0212] Table 3 shows the simulation error of the convoy with the lead vehicle number 745.
[0213]
[0214] An autonomous driving decision-making system based on multimodal heterogeneous feature fusion is used to implement the above-mentioned autonomous driving decision-making method based on multimodal heterogeneous feature fusion. The system includes:
[0215] The data division module divides the driving status, including:
[0216] The feature division unit uses a segmentation algorithm to perform feature division on the speed and trajectory profile of the vehicle's driving trajectory;
[0217] The dynamic response relationship unit, based on Newell stimulus-response theory and dynamic time warping algorithm, establishes dynamic response relationships between vehicles and identifies the following flow state and free flow state;
[0218] Identification unit, identifies specific driving status;
[0219] Building modules to construct a car-following model using driving status, including:
[0220] Framework unit, building the driving state prediction module framework;
[0221] Fusion unit, building an improved LSTM network to achieve multimodal feature fusion;
[0222] The training module trains the car-following model and realizes the coordinated training of driving state prediction and car-following behavior modeling through curriculum learning and multi-stage optimization strategy.
[0223] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the above-mentioned autonomous driving decision-making method based on multimodal heterogeneous feature fusion.
[0224] A computer program product includes a computer program, which, when executed by a processor, implements the above-mentioned autonomous driving decision-making method based on multimodal heterogeneous feature fusion.
[0225] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented in whole or in part in the form of a computer program product, the computer program product includes one or more computer instructions. When the computer program instructions are loaded or executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL) or wireless (e.g., infrared, wireless, microwave, etc.)) method. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).
[0226] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. An autonomous driving decision-making method based on multimodal heterogeneous feature fusion, characterized in that: include: S1. Classify driving status, including: S101, using a segmentation algorithm to perform feature segmentation on the speed and trajectory profile of the vehicle's driving trajectory; S102: Based on Newell stimulus-response theory and dynamic time warping algorithm, establish dynamic response relationships between vehicles and identify the following flow state and free flow state; S103, identifying a specific driving state; S2. Build a car-following model using driving status, including: S201, building a driving state prediction module framework; S202. Build an improved LSTM network to achieve multimodal feature fusion; S3. Train the car-following model and achieve collaborative training of driving state prediction and car-following behavior modeling through curriculum learning and multi-stage optimization strategies.
2. The autonomous driving decision-making method based on multimodal heterogeneous feature fusion according to claim 1, characterized in that: S101 includes: S1011, data preprocessing: performing median filtering on the original velocity sequence; S1012, initial segment generation: divide the velocity sequence into atomic segments, each segment contains 5-10 consecutive sampling points S1013, iterative merging optimization; S1014: State interval integration: merging continuous segments whose slope change rate is less than a threshold.
3. The autonomous driving decision-making method based on multimodal heterogeneous feature fusion according to claim 2, characterized in that: S1013 includes: Calculate the slope matching degree of adjacent segments and apply a penalty coefficient of 0.3-0.6 when the slope signs are opposite; Dynamically select the optimal segment pair for merging based on the merging cost function; The iteration is stopped when the number of trajectory segments is reduced to 1.2-1.8 times the total duration.
4. The autonomous driving decision-making method based on multimodal heterogeneous feature fusion according to claim 2, characterized in that: The threshold is 0.01 m / s 2 .
5. The autonomous driving decision-making method based on multimodal heterogeneous feature fusion according to claim 1, characterized in that: S102 includes: S1021, data preprocessing: perform spatiotemporal synchronization on the segmented trajectory data sets of the leading and trailing vehicles, and construct the velocity time series v of the trailing vehicle. L (t), time series of the preceding vehicle's speed v F (t), time series of the following vehicle position x L (t), the time series of the preceding vehicle's position x F (t); S1022: Multi-dimensional DTW alignment, calculating the optimal matching path of the spatial trajectory: DTW x =arg min π ∑ (i,j)∈π | x L(t i )-x F (t j )|| Get the matching paths on the two trajectories and k represents the serial number of the matching path; i represents the trajectory of the preceding vehicle, j represents the trajectory of the following vehicle, and π represents the set of all trajectories; S1023, feature parameter extraction: analyzing the time delay distribution of the optimal curved path, extracting parameters characterizing the dynamic response between vehicles, and generating a speed delay distribution histogram; S1024, state judgment criteria: establish a state judgment criteria based on τ n The state discriminant function of dynamic characteristics: close following state is CF, and free flow state is FF; in, is the time delay set τ n The 80% quantile of The time interval for matching paths between the leading and trailing vehicles.
6. The autonomous driving decision-making method based on multimodal heterogeneous feature fusion according to claim 5, characterized in that: S103 includes: S1031, obtaining segmented speed curves of the following vehicles CF and FF; S1032. Slope classification: The segments are divided into two groups according to the slope, one group has a positive slope, representing acceleration behavior; the other group has a negative slope, representing deceleration behavior; S1033, steady-state segment identification: classify line segments with slopes between -0.5 and 0.5 as constant-speed segments; S1034: State classification, obtaining the driving state at each moment in the vehicle trajectory: In the CF profile, the positive slope segment represents the acceleration following state A, the segment close to zero slope represents the constant speed following state F, and the negative slope segment represents the deceleration following state D; In the FF profile, the positive slope segment represents the free acceleration Fa state, the segment close to zero slope represents the cruising state at the desired speed C state, and the FF segment does not include stationary or deceleration conditions.
7. The autonomous driving decision-making method based on multimodal heterogeneous feature fusion according to claim 6, characterized in that: S201 includes: A time series prediction model is constructed using GRU. The input layer receives the vehicle kinematic state vector X_{t-1}∈R^9 from the historical time window tk to t-1. The vehicle kinematic state vector includes the relative position of the following vehicle and the leading vehicle, the speed difference, the three-dimensional continuous features of the following vehicle's speed, and the one-hot encoding of the six-dimensional driving state. Through the coordinated operation of the GRU's update gate z_t and reset gate r_t, a hidden state update mechanism is established: Where z t represents the update gate, h t-1 represents the hidden state of the previous time step, ⊙ represents the Hadamard product, is the candidate hidden state; The output layer uses the Softmax function to generate the driving state DR t The probability distribution of ∈{0,1,...,5} is expressed as: P(DR t |X {t-1,...,t-k} )=Softmax(W g h t +b g ) Where, X {t-1,...,t-k} Represents the historical time series characteristics of the model input, W g Represents the weight matrix of the output layer, h t is the hidden state at the current moment, b g is the bias vector of the output layer.
8. The autonomous driving decision-making method based on multimodal heterogeneous feature fusion according to claim 7, characterized in that: S202 includes: S2021, driving status DR t Through the learnable embedding matrix W e ∈R 1×6 Mapped to a continuous latent space, generating a low-dimensional embedding vector E d =W e OneHot(DR t )+b e ; OneHot is the one-hot encoding operator, b e is the learnable bias; S2022, embed the vector E d With continuous motion feature X c ∈R 3 Perform cross-modal splicing to generate fusion feature X fused =Concat(X c ,E d )∈R 4 ;Concat is the concatenation operator; S2023. Use the hyperbolic tangent function to update the cell state in the LSTM gate unit: c t =f t ⊙c t-1 +i t ⊙tanh(W c [h t-1 ,x t ]+bc ) Where ⊙ represents the Hadamard product, f t represents the output of the forget gate, c t-1 Indicates the cell state at the previous moment, i t represents the output of the input gate, W c represents the weight matrix, h t-1 Indicates the hidden state at the previous moment, x t Indicates the current input, b c is the bias vector. The gradient saturation property of the tanh function is used to suppress the gradient explosion phenomenon in the recurrent neural network. The parameterized rectified linear unit PReLU is introduced into the output layer, and the activation function is defined as: Among them, α is a learnable parameter, α∈[0,1], which automatically optimizes the slope of the non-saturated region through back propagation; S2024, using the Adam adaptive optimization algorithm to set key parameters; S2025. Construct a composite loss function to achieve joint optimization of driving state classification and motion parameter regression.
9. An autonomous driving decision-making system based on multimodal heterogeneous feature fusion, characterized in that: include: The data division module divides the driving status, including: The feature division unit uses a segmentation algorithm to perform feature division on the speed and trajectory profile of the vehicle's driving trajectory; The dynamic response relationship unit, based on Newell stimulus-response theory and dynamic time warping algorithm, establishes dynamic response relationships between vehicles and identifies the following flow state and free flow state; Identification unit, identifies specific driving status; Building modules to construct a car-following model using driving status, including: Framework unit, building the driving state prediction module framework; Fusion unit, building an improved LSTM network to achieve multimodal feature fusion; The training module trains the car-following model and realizes the coordinated training of driving state prediction and car-following behavior modeling through curriculum learning and multi-stage optimization strategy.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the autonomous driving decision-making method based on multimodal heterogeneous feature fusion as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Deep learning car-following prediction method considering driver fuzzy perception
CN112193245A
Vehicle acceleration prediction method considering driving behavior characteristics in following scene
CN116279471A
Vehicle following safety analysis method based on Newell following trajectory theory
CN117095532A
Vehicle track prediction method based on data driving and theoretical driving
CN117150258A
Simulation-based digital twinborn visualization technology method and system
CN119003905A
Cited By
Method and system for training autonomous vehicle trajectory planning module
CN121157972A