End-to-end automatic driving behavior decision-making method and system considering uncertainty estimation
By introducing an integrated quantile network into the end-to-end autonomous driving behavior decision model, combining distributed reinforcement learning and multi-model integration to obtain uncertainty estimation of decision results, the problem of unreliable decision-making in the face of noise or edge scenarios is solved, and higher decision reliability and credibility are achieved.
Patent Information
- Application Number
- CN202510055256.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-13
AI Technical Summary
The existing end-to-end autonomous driving behavior decision model is unreliable when facing data noise or edge scenarios, resulting in low decision safety and credibility.
The integrated quantile network is adopted, combined with distributed reinforcement learning and multi-model integration methods, and the complete uncertainty estimate of the agent's decision results, including data uncertainty estimates and model uncertainty estimates, and dynamically select the optimal driving behavior or braking strategy based on confidence.
It improves the reliability and credibility of end-to-end autonomous driving decisions, and by quantifying the uncertainty of decision results, unfounded potentially dangerous decisions are avoided, and the safety and stability of the system are enhanced.
Smart Images

Figure CN119975406A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the technical field of end-to-end autonomous driving vehicle behavior decision-making, and in particular, to an end-to-end autonomous driving behavior decision-making method and system considering uncertainty estimation. Background Art
[0002] With the increase in the number of cars, traffic congestion and traffic accidents are frequent. The autonomous driving system can optimize the driving route and alleviate traffic congestion through high-precision sensors and algorithms; it can significantly reduce human errors and reduce traffic accidents. Since the traditional modular autonomous driving paradigm has clear functional modules and input and output interfaces, the information transmission loss is large and the maintenance cost is high, the end-to-end autonomous driving paradigm has emerged. The end-to-end paradigm integrates modular networks into an end-to-end autonomous driving model, realizes direct mapping from perception information to control signals, and realizes lossless transmission of information, which provides a new perspective for the development of autonomous driving technology and promotes the improvement of the performance of autonomous driving systems.
[0003] In recent years, end-to-end autonomous driving behavior decision models have been widely used because of their ability to learn complex high-dimensional driving strategies. However, the internal structure of the model network is not transparent and interpretable, and the model only provides a solution to the final decision, without providing information about the agent's confidence in the model's decision solution. When the input information is blocked (data noise) or the input samples exceed the training distribution (edge scenarios), the decision results made by the model are unreliable, resulting in low safety and credibility of end-to-end autonomous driving decisions. Summary of the invention
[0004] In view of this, an embodiment of the present invention provides an end-to-end autonomous driving behavior decision-making method and system considering uncertainty estimation, aiming to improve the reliability and credibility of end-to-end autonomous driving behavior decision-making.
[0005] According to a first aspect of an embodiment of the present invention, there is provided an end-to-end autonomous driving behavior decision method considering uncertainty estimation, comprising: inputting image information of surrounding traffic participants collected by a camera into a convolutional neural network of an end-to-end autonomous driving decision network for feature extraction to obtain state information of surrounding traffic participants; performing feature splicing and fusion of the vehicle state information and the surrounding traffic participants state information, and inputting the obtained fusion features into a fully connected layer of a downstream decision module of a constructed end-to-end autonomous driving decision model agent to obtain an agent decision result; constructing an integrated quantile network based on the fully connected layer to obtain a complete uncertainty estimate of the agent decision result, and a complete estimate of the agent decision result. Uncertainty estimation includes data uncertainty estimation based on implicit quantile networks and model uncertainty estimation based on multi-model integrated networks considering prior Bayesian estimation; based on the complete uncertainty estimation, the confidence of the agent's decision results is evaluated. If the confidence is strong, the optimal driving action of the vehicle predicted by the integrated quantile network is output; if the confidence is weak, the braking strategy output is switched to; the predicted optimal driving action and braking strategy of the vehicle are input as strategy information into the control module of the end-to-end autonomous driving decision model agent, and a control signal is output. According to the output control signal, the vehicle performs different control actions, including acceleration, braking, and steering.
[0006] In one implementation, the surrounding traffic participant status information includes the location x of the traffic participant. i ,y i , speed v i and heading angle ψ i ,The self-vehicle state information includes the position information and speed information of the self-vehicle.
[0007] In another implementation, the end-to-end autonomous driving decision network uses a deep reinforcement Q network. The core reinforcement learning framework of the deep reinforcement Q network is the optimal strategy π(s) for interactive learning between the agent and the environment. The optimal strategy describes the action a taken by the agent in state s. When the environment moves to a new state s′, the agent receives a reward r.
[0008] The value of the agent taking action a in state s following strategy π is defined by the state-action value Q-value function:
[0009]
[0010] Among them, the optimal strategy π * The Q value is defined as:
[0011]
[0012] The deep reinforcement Q network uses a deep neural network instead of the traditional Q value representation to achieve an approximate representation of the optimal Q value function:
[0013]
[0014] The deep reinforcement Q network learns the strategy by minimizing the time difference error between the predicted Q value and the target value. The time difference error is expressed as:
[0015]
[0016] Among them, r is the discount factor that weighs the current reward and future rewards, and its value range is [0,1]; Q(s′,a′; θ - ) is the maximum Q value of all possible actions in the next state s′, which is determined by the target network θ - Calculation: Q(s,a;θ) is the predicted Q value of the current state s and action a, and parameter θ is the current network weight.
[0017] In another implementation, the network parameter θ is optimized by minimizing the time difference error, and the parameters are updated by back propagation using the gradient descent method to adjust Q(s, a; θ) closer to the target value. The loss function is:
[0018]
[0019] in, Represents the expectation of the loss function, specifically the square of the temporal difference error, which is used in reinforcement learning to measure the difference between the predicted value and the actual value.
[0020] In another implementation, the data uncertainty estimation based on the implicit quantile network includes:
[0021] Based on the implicit quantile network, distributed reinforcement learning is introduced to learn the distribution of Q values. Each action a corresponds to a group of quantile prediction values to represent the distribution of the Q value of the action. The expression of Q value is:
[0022] Z(s,a)~Q(s,a)
[0023] Among them, Z(s,a) is the Q value distribution of state s and action a, and Q(s,a) is the expected value of Z(s,a), that is,
[0024] The implicit quantile network learns the quantile function of the distribution by sampling the quantile points. The two quantile samples τ and τ′ are used to estimate the distribution value Z of the state-action pair. τ (s,a) and target value Z τ′(s′, a′), represents the current estimated quantile value and the quantile value of the target distribution; the quantile sample τ, Implicit mapping f is realized through convolutional neural network θ (s, a, τ), output a set of Q value estimates of the loci, i.e. Z τ (s,a)=f θ (s,a,τ), to capture the Q-value distribution of state-action pairs;
[0025] The time difference error is used to measure the error between the current quantile sample and the target quantile sample to estimate the Q value distribution. The implicit quantile network learns the Q value distribution based on the time difference error, which is expressed as:
[0026]
[0027] Among them, r t is the reward of the current time step, γ is the discount factor that weighs the current reward against future rewards, and θ - represents the parameters of the target network, π * (s) is the maximum policy value of the state-action pair;
[0028] A risk aversion strategy is adopted to select the maximum strategy value and calculate the conditional risk value. Under a given quantile α, for the distributed value function Z(s,a) of state s and action a, the conditional risk value calculation formula is:
[0029]
[0030] The optimal strategy formed based on quantile returns and conditional value at risk is expressed as:
[0031]
[0032] Among them, λ1 represents the weight of quantile returns to reflect the priority of returns, and λ2 represents the weight of conditional risk value to reflect the degree of risk aversion;
[0033] The target distribution is optimized using the distribution fitting loss Wasserstein distance, and the quantiles τ and τ' are sampled N and N' times respectively:
[0034]
[0035] Among them, ρ k is the quantile Huber regression loss function, and the formula is as follows:
[0036]
[0037] Among them, k is a hyperparameter, called the smoothing threshold, k = α*max(|Z τ′(s′,a′)|), α is a scaling factor, ranging from [0.1,0.5];
[0038] The estimated distribution of reward from trained agents is used to quantify the stochastic uncertainty of the agent’s decision outcomes, including:
[0039] For K τ uniformly distributed samples τ σ ={i / K τ ∣i∈[1,K τ ]}, the metric uses the estimated return variance of uniformly distributed samples, and its threshold is defined as If the variance is below the threshold, the trained agent follows the optimal strategy formed from the quantile return and the conditional value at risk. If the variance is not below the threshold, the agent is braked by the predefined braking strategy π backup (s) makes decision output, the formula is as follows:
[0040]
[0041] in represents the benefit of generating a decision trajectory τ based on strategy σ after selecting action a in state s; CVaR α (s,a) is the conditional risk value of taking this action at a certain risk level; λ1 and λ2 are hyperparameters to balance the weight between expected return and conditional risk value; represents the variance of the return under trajectory τ, which is used to measure uncertainty; is the upper threshold of the variance associated with action a, used to constrain uncertainty; π backup (s) indicates that the variance exceeds the threshold A conservative braking strategy is used to ensure the stability of the model decision.
[0042] In another implementation, the model uncertainty estimation obtained based on the multi-model ensemble network considering the prior Bayesian estimation specifically includes:
[0043] Perform Bayesian probability modeling based on training data and sample data to establish the prior distribution of model parameters;
[0044] Calculate the likelihood function of the training data, combine the prior distribution of the model parameters with the likelihood function of the training data, and obtain the posterior distribution p(θ|D), as follows:
[0045]
[0046] Among them, p(θ) is the prior distribution of model parameters, which represents the assumption about model parameters θ before unknown data, p(D|θ) is the likelihood function of training data, which represents the probability of training data given model parameters θ, p(D) is a normalized constant representing the marginal probability of training data, and D represents training data;
[0047] The optimal approximate posterior distribution is found by minimizing the KL divergence, expressed as:
[0048]
[0049] The multi-model ensemble network samples multiple word models from the posterior distribution and performs weighted averaging, including: based on the approximate posterior distribution generated by the acquired Bayesian modeling, sampling a set of multiple parameters {θ1,θ2,…,θ K}, corresponding to the network parameters of K sub-models, integrate the prediction results of each sub-model, and perform weighted average, where the weight is based on the probability of the posterior distribution of the model, that is, w k ∝P(θ k |D), the formula is as follows:
[0050]
[0051] Among them, the Q value of the multi-model integrated network is calculated by combining Bayesian modeling and the output of each sub-model, expressed as:
[0052]
[0053] Among them, f(s,a;θ k ) is to train the network based on historical data to obtain the predicted expected return value, is the posterior reward value of Bayesian modeling, and β is the adjustment factor to balance the expected return of training data and the reward value generated by Bayesian modeling;
[0054] Minimizing the loss function based on the time difference error optimizes the network parameters to train the multi-model ensemble network, where the time difference error of the ensemble model changes, and its loss function is expressed as follows:
[0055]
[0056] Among them, r t is the return at the current time step, γ is the discount factor, is the output of the target network;
[0057] The variance of the Q value distribution in the integrated model is defined as the threshold for measuring the model uncertainty estimation standard. If the sample variance is above a threshold, a predefined braking strategy π is selected backup(s) is the output of the model decision result, and the formula is as follows:
[0058]
[0059] in, Indicates that after selecting action a in state s, the strategy The expected Q value under , that is, the action value function, represents the cumulative reward obtained by taking a certain action in a certain state; Var k [Q k (s,a)] is Q k The variance of (s,a) measures the uncertainty of the expected reward of the action; is the variance threshold of the action, indicating the maximum allowed uncertainty; π backup (s) is the braking strategy adopted when the variance exceeds the threshold, which is used to avoid excessive risk or high uncertainty.
[0060] In another implementation, the complete uncertainty estimation combines the quantile network and the integrated network considering Bayesian prior into a new algorithm through data uncertainty estimation and model uncertainty estimation, combining the quantile reward distribution output and uncertainty estimation of each sub-model to optimize the reward value Y of the state-action pair k,τ (S,a), the calculation formula is:
[0061]
[0062] Among them, Y k,τ (s,a) is the k-th model’s estimate of the reward value for state s and action a at the quantile τ, f τ (s,a;θ k ) is the kth model based on parameter θ k Quantile return forecasts for , return value estimates generated by the quantile network, It is the uncertainty estimate of the quantile τ based on the model generated by the Bayesian posterior distribution;
[0063] The final output is the fused reward value, that is, the reward value Z of k models k,τ The mean of (s,a) is expressed as:
[0064]
[0065] Select the optimal strategy based on the fused return value:
[0066]
[0067] Among them, the agent defines an end-to-end decision-making mechanism based on confidence evaluation and braking strategy according to the variance of the reward distribution: when the model strategy confidence is strong, that is, and The strategy selection is driven by the expected value of the return value prediction. When the confidence is weak, that is, when the variance of the uncertainty estimate is large, it switches to the braking strategy to reduce the potential risk of low-confidence decisions. The specific formula is as follows:
[0068]
[0069] Among them, Z k,τ (s,a) is the quantile return value prediction of k models for state and action, Take the expectation within the confidence interval for the quantile τ, The expected value of the return predictions of the k models, representing the overall prediction of the ensemble model.
[0070] According to a second aspect of an embodiment of the present invention, an end-to-end autonomous driving behavior decision system considering uncertainty estimation is provided, including: an extraction module, which is used to input the image information of surrounding traffic participants collected by a camera into the convolutional neural network of the end-to-end autonomous driving decision network for feature extraction, so as to obtain the state information of the surrounding traffic participants; a fusion module, which is used to perform feature splicing and fusion of the vehicle state information and the state information of the surrounding traffic parameters, and input the obtained fusion features into the fully connected layer of the downstream decision module of the constructed end-to-end autonomous driving decision model agent to obtain the agent decision result; an estimation module, which is used to construct an integrated quantile network based on the fully connected layer to obtain the complete uncertainty estimation of the agent decision result, and the agent decision result The complete uncertainty estimation includes data uncertainty estimation based on implicit quantile networks and model uncertainty estimation based on multi-model integrated networks considering prior Bayesian estimation; the confidence assessment module is used to evaluate the confidence of the agent's decision results based on the complete uncertainty estimation. If the confidence is strong, the optimal driving action of the vehicle predicted by the integrated quantile network is output. If the confidence is weak, it switches to the braking strategy output; the execution module is used to input the predicted optimal driving action and braking strategy of the vehicle as strategy information into the control module of the end-to-end autonomous driving decision model agent, output control signals, and make the vehicle perform different control actions according to the output control signals. The different control actions include acceleration, braking, and steering.
[0071] According to a third aspect of an embodiment of the present invention, there is provided an electronic device, comprising a processor and a memory storing a program, wherein the program comprises instructions, and when the instructions are executed by the processor, the processor executes the steps executed by the method of the first aspect.
[0072] According to a fourth aspect of an embodiment of the present invention, there is provided a computer storage medium on which a computer program is stored. When the program is executed by a processor, the method of the first aspect described above is implemented.
[0073] Compared with the prior art, the present invention has the following beneficial effects:
[0074] The present invention proposes an end-to-end autonomous driving behavior decision-making method and system considering uncertainty estimation, which improves the safety and reliability of end-to-end autonomous driving decisions. The present invention introduces an integrated quantile network, combines distributed reinforcement learning with a multi-model integration method, and obtains a complete uncertainty estimate of the intelligent agent's decision results, including obtaining data uncertainty estimates based on the implicit learning return probability distribution of the quantile function, and obtaining model uncertainty estimates based on a multi-model integration method considering a priori Bayesian estimates; when making decisions, the present invention comprehensively considers the potential returns at the quantile level and the conditional risk value that quantifies extreme risks, and dynamically balances benefits and risks based on the prediction results to select the optimal driving behavior. In addition, the uncertainty of the decision results is quantified, and the present invention also sets a braking strategy corresponding to the decision results with higher uncertainty, so as to prevent the intelligent agent from making unfounded and potentially dangerous decisions. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0076] Figure 1 This is a flowchart of the steps of the end-to-end autonomous driving behavior decision-making method considering uncertainty estimation of the present invention.
[0077] Figure 2 For Figure 1 Corresponding schematic diagram of the overall technical route of the end-to-end autonomous driving behavior decision method considering uncertainty estimation of the present invention.
[0078] Figure 3 A schematic diagram of the technical route for data uncertainty estimation of the present invention.
[0079] Figure 4 A schematic diagram of the technical route for model uncertainty estimation of the present invention. DETAILED DESCRIPTION
[0080] In order to have a clearer understanding of the technical features, purposes and effects of the embodiments of the present invention, the specific implementation of the embodiments of the present invention is now described with reference to the accompanying drawings.
[0081] In this document, “exemplary” means “serving as an example, instance or illustration”, and any illustration or implementation described in this document as “exemplary” should not be construed as a more preferred or more advantageous technical solution.
[0082] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in the field based on the embodiments in the embodiments of the present invention should fall within the scope of protection of the embodiments of the present invention.
[0083] The specific implementation of the embodiment of the present invention is further described below in conjunction with the accompanying drawings of the embodiment of the present invention.
[0084] See also Figure 1 , Figure 2 The present invention provides an end-to-end autonomous driving behavior decision method considering uncertainty estimation, comprising:
[0085] Step S1, inputting the image information of surrounding traffic participants collected by the camera into the convolutional neural network of the end-to-end autonomous driving decision network for feature extraction to obtain the state information of the surrounding traffic participants;
[0086] Step S2: perform feature splicing and fusion of the vehicle state information and the surrounding traffic participant state information, and input the obtained fusion features into the fully connected layer of the downstream decision module of the constructed end-to-end autonomous driving decision model agent to obtain the agent decision result;
[0087] Step S3, constructing an integrated quantile network based on the fully connected layer to obtain a complete uncertainty estimate of the agent's decision result, wherein the complete uncertainty estimate of the agent's decision result includes a data uncertainty estimate obtained based on the implicit quantile network and a model uncertainty estimate obtained based on a multi-model integrated network considering a priori Bayesian estimation;
[0088] Step S4: Evaluate the confidence of the decision result of the agent based on the complete uncertainty estimation. If the confidence is strong, output the optimal driving action of the vehicle predicted by the integrated quantile network. If the confidence is weak, switch to the braking strategy output.
[0089] Step S5: input the predicted optimal driving action and braking strategy of the vehicle as strategy information into the control module of the end-to-end autonomous driving decision model intelligent body, output a control signal, and make the vehicle perform different control actions according to the output control signal. The different control actions include acceleration, braking, and steering.
[0090] Optionally, the surrounding traffic participant status information includes the location x of the traffic participant. i ,y i , speed v i and heading angle ψi ,The self-vehicle state information includes the position information and speed information of the self-vehicle.
[0091] Specifically, the scheme of the present invention is further described according to the following examples:
[0092] First, six cameras are used, located in front of the vehicle, in front of the left, in front of the right, in the back, in the left, and in the right. The image information of the surrounding traffic participants is input, and the convolutional neural network is used to extract features, including the position x of the traffic participants. i ,y i , speed v i and heading angle ψ i ,In addition, the position information and speed information of the ego vehicle are extracted for feature fusion and ,input into the downstream decision module.
[0093] The information extracted by perception is input into the implicit quantile network. Each action corresponds to a set of quantile observations, which represent the Q value distribution of the action. The network predicts the quantile return distribution Z τ (s,a), based on the quantile return results, calculate the conditional risk value, estimate the expected risk exceeding the quantile, combine the quantile return and conditional risk value, and use the weighted objective function to select the optimal action.
[0094] The ensemble method uses multiple independently trained quantile networks, considers Bayesian prior probability modeling, samples multiple parameters from the posterior distribution, sets the network parameters corresponding to multiple sub-models, generates multiple distribution prediction values, and performs quantile return prediction, that is, The Q-value estimation distribution of multiple models is integrated to obtain statistical information such as the mean and variance of the distribution, and the variance of the distribution is used for uncertainty estimation.
[0095] Each model in the integrated network is trained independently and optimized using the quantile Huber loss function. When training the integrated quantile network, the number of integrated models k and the range of the sampling quantile τ are adjusted through cross-validation.
[0096] If the variance is within the threshold range, that is, the confidence level is strong, the optimal driving action output of the optimal strategy value predicted by the network is adopted; if the variance is large and the confidence level is weak, it is directly switched to the conservative strategy, that is, the braking strategy output.
[0097] The strategy information of the decision network is input into the control module to generate control signals, so that the vehicle performs control actions such as "acceleration", "braking" and "steering".
[0098] In summary, the specific implementation method of building an end-to-end autonomous driving decision-making model considering uncertainty estimation has been analyzed and explained. The data-driven end-to-end autonomous driving decision-making model has long-tail problems such as uneven data distribution and unknown model decision confidence. The present invention uses an integrated quantile network to quantify the uncertainty of the model decision results, and sets a braking strategy corresponding to the weak confidence of the decision results to avoid unfounded and potentially dangerous decisions output by the model, thereby enhancing the safety, stability and credibility of the end-to-end autonomous driving decision model.
[0099] In order to improve the reliability and credibility of end-to-end autonomous driving behavior decisions, the present invention proposes an end-to-end autonomous driving behavior decision-making method that considers uncertainty estimation. It is mainly aimed at high-risk behavior decision-making scenarios caused by field of view occlusion or uneven distribution of training data (edge scenarios). Data uncertainty and model uncertainty estimation are introduced into a modular end-to-end autonomous driving behavior decision model, and its results are quantitatively evaluated to select safe and reliable behavior decisions, thereby enhancing people's trust in the end-to-end autonomous driving system.
[0100] The end-to-end autonomous driving decision network, data uncertainty estimation, model uncertainty estimation and complete uncertainty estimation of the present invention are described in detail through the following content:
[0101] (1) End-to-end autonomous driving decision network
[0102] The end-to-end autonomous driving decision network uses a deep reinforcement Q network (DQN). The core of the deep reinforcement Q network is the reinforcement learning (RL) framework. The reinforcement learning framework is the optimal strategy π(s) for the interactive learning between the agent and the environment. The optimal strategy describes the action a taken by the agent in state s. When the environment moves to a new state s′, the agent receives a reward r.
[0103] The value of the agent taking action a in state s following policy π is defined by the state-action value (Q-value) function as:
[0104]
[0105] Among them, the optimal strategy π * The Q value is defined as:
[0106]
[0107] The deep reinforcement Q network uses a deep neural network instead of the traditional Q value representation to achieve an approximate representation of the optimal Q value function:
[0108]
[0109] The deep reinforcement Q network learns the strategy by minimizing the temporal difference (TD) error between the predicted Q value and the target value. The TD error is expressed as:
[0110]
[0111] Among them, r is the discount factor that weighs the current reward and future rewards, and its value range is [0,1]; Q(s′,a′; θ - ) is the maximum Q value of all possible actions in the next state s′, which is determined by the target network θ - Calculation: Q(s,a;θ) is the predicted Q value of the current state s and action a, and parameter θ is the current network weight.
[0112] The network parameter θ is optimized by minimizing the time difference error, and the parameters are updated by back propagation using the gradient descent method to adjust Q(s, a; θ) closer to the target value. The loss function is:
[0113]
[0114] in, Represents the expectation of the loss function, specifically the square of the temporal difference error, which is used in reinforcement learning to measure the difference between the predicted value and the actual value.
[0115] (2) Data uncertainty estimation
[0116] The end-to-end autonomous driving decision model relies on a large amount of training data. Due to the uneven distribution of data, the scenarios cannot be exhausted and there are edge scenarios, which reduces the accuracy of the end-to-end decision model. This paper considers the random uncertainty of the input scene data, introduces an implicit quantile network to learn the distribution function of the cumulative Q value, captures the distribution information of the Q value, and characterizes the potential risks and uncertainties of different actions. For a technical route diagram of data uncertainty estimation, see Figure 3 shown.
[0117] Based on the implicit quantile network, distributed reinforcement learning is introduced to learn the distribution of Q values. Each action a corresponds to a group of quantile prediction values to represent the distribution of the Q value of the action. The expression of Q value is:
[0118] Z(s,a)~Q(s,a)
[0119] Among them, Z(s,a) is the Q value distribution of state s and action a, and Q(s,a) is the expected value of Z(s,a), that is,
[0120] The implicit quantile network learns the quantile function of the distribution by sampling the quantile points. The two quantile samples τ and τ′ are used to estimate the distribution value Z of the state-action pair. τ (s,a) and target value Zτ′ (s′, a′), represents the current estimated quantile value and the quantile value of the target distribution; the quantile sample τ, Implicit mapping f is realized through convolutional neural network θ (s, a, τ), output a set of Q value estimates of the loci, i.e. Z τ (s,a)=f θ (s,a,τ), to capture the Q-value distribution of state-action pairs;
[0121] In order to estimate the distribution of Q values, the error between the current quantile sample and the target quantile sample is measured by the temporal difference (TD) error. The implicit quantile network learns the Q value distribution according to the temporal difference error, which is expressed as:
[0122]
[0123] Among them, r t is the reward of the current time step, γ is the discount factor that weighs the current reward against future rewards, and θ - represents the parameters of the target network, π * (s) is the maximum policy value of the state-action pair;
[0124] The selection of the above-mentioned maximum strategy value adopts a risk aversion strategy in the present invention, considering the strategy corresponding to the behavior of maximizing the quantile return and the conditional risk value, which measures the average loss under extreme conditions. The conditional risk value is a risk measure that represents the expected loss within the first α (such as 5% or 10%) where the loss is the most serious. By averaging the value function in the low quantile area, it reflects the overall risk level of the area.
[0125] Specifically, a risk aversion strategy is adopted to select the maximum strategy value and calculate the conditional risk value. Under a given quantile α, for the distributed value function Z(s,a) of state s and action a, the conditional risk value calculation formula is:
[0126]
[0127] The optimal strategy formed based on quantile returns and conditional value at risk is expressed as:
[0128]
[0129] Among them, λ1 represents the weight of quantile returns to reflect the priority of returns, and λ2 represents the weight of conditional risk value to reflect the degree of risk aversion;
[0130] In order to optimize the Q value distribution, the distribution fitting loss (Wasserstein distance) is used to optimize the target distribution, and the quantiles τ and τ' are sampled N and N' times respectively:
[0131]
[0132] Among them, ρ κ is the quantile Huber regression loss function, which is used to smooth outliers and enhance the robustness to outliers. The formula is as follows:
[0133]
[0134] Among them, κ is a hyperparameter, called the smoothing threshold, κ = α*max(|Z τ′ (s′,a′)|), α is a scaling factor, ranging from [0.1,0.5];
[0135] It will be appreciated that κ is typically taken as a fixed ratio relative to the target value.
[0136] Finally, the estimated distribution of reward of the trained agent is used to quantify the stochastic uncertainty of the agent’s decision outcomes.
[0137] Specifically, for K τ uniformly distributed samples τ σ ={i / K τ ∣i∈[1,K τ ]}, the metric uses the estimated return variance of uniformly distributed samples, and its threshold is defined as If the variance is below the threshold, the trained agent follows the optimal strategy formed based on the quantile return and conditional risk value. Otherwise, it indicates that the decision has a high degree of uncertainty and a high risk, and the predefined braking strategy π is used. backup (s) makes decision output, the formula is as follows:
[0138]
[0139] in, represents the benefit of generating a decision trajectory τ based on strategy σ after selecting action a in state s; CVaR α (s,a) is the conditional risk value of taking this action at a certain risk level; λ1 and λ2 are hyperparameters to balance the weight between expected return and conditional risk value; represents the variance of returns under trajectory τ, which is used to measure uncertainty; is the upper threshold of the variance associated with action a, used to constrain uncertainty; π backup (s) indicates that the variance exceeds the threshold A conservative braking strategy is used to ensure the stability of the model decision.
[0140] (3) Model uncertainty estimation
[0141] Model uncertainty is caused by the uncertainty of model parameters, such as insufficient training or limited model expression ability. When the model encounters noise or uncertain data, in the decision model, when dealing with uncertain information such as high-risk and edge scenarios, the uncertainty estimation of the model can provide the reliability of the decision model. See the technical route diagram of model uncertainty estimation for details. Figure 4 shown.
[0142] The posterior distribution is generated by Bayesian probability modeling considering the prior distribution and applied to the integrated network to estimate the uncertainty of the model. The dispersion (variance) of the multi-model output is used to estimate the discrete degree of the model's prediction of samples, that is, the model uncertainty.
[0143] Perform Bayesian probability modeling based on training data and sample data to establish the prior distribution of model parameters;
[0144] Calculate the likelihood function of the training data, combine the prior distribution of the model parameters with the likelihood function of the training data, and obtain the posterior distribution p(θ|D), as follows:
[0145]
[0146] Among them, p(θ) is the prior distribution of model parameters, which represents the assumption about model parameters θ before unknown data, p(D|θ) is the likelihood function of training data, which represents the probability of training data given model parameters θ, p(D) is a normalized constant representing the marginal probability of training data, and D represents training data;
[0147] The goal of Bayes' theorem is to update the distribution of model parameters through training data to obtain the posterior distribution of the parameters p(θ|D). However, the posterior distribution is usually difficult to solve and is solved by approximate inference. KL divergence measures the difference between the approximate distribution q(θ) and the true posterior distribution p(θ|D). The optimal approximate posterior distribution is found by minimizing the KL divergence, which is expressed as:
[0148]
[0149] The multi-model ensemble network samples multiple word models from the posterior distribution and performs weighted averaging, including: based on the approximate posterior distribution generated by the acquired Bayesian modeling, sampling a set of multiple parameters {θ1,θ2,…,θ K}, corresponding to the network parameters of K sub-models, integrate the prediction results of each sub-model, and perform weighted average, where the weight is based on the probability of the posterior distribution of the model, that is, w k ∝P(θ k |D), the formula is as follows:
[0150]
[0151] Among them, the Q value of the multi-model integrated network is calculated by combining Bayesian modeling and the output of each sub-model, expressed as:
[0152]
[0153] Among them, f(s,a;θ k ) is the expected return value obtained by training the network based on historical data, is the posterior reward value of Bayesian modeling, and β is the adjustment factor to balance the expected return of training data and the reward value generated by Bayesian modeling;
[0154] Minimizing the loss function based on the time difference error optimizes the network parameters to train the multi-model ensemble network, where the time difference error of the ensemble model changes, and its loss function is expressed as follows:
[0155]
[0156] Among them, r t is the return at the current time step, γ is the discount factor, is the output of the target network;
[0157] The variance of the Q value distribution in the integrated model is defined as the threshold for measuring the model uncertainty estimation standard. If the sample variance is above a threshold, a predefined braking strategy π is selected backup (s) is the output of the model decision result, and the formula is as follows:
[0158]
[0159] in, Indicates that after selecting action a in state s, the strategy The expected Q value under , that is, the action value function, represents the cumulative reward obtained by taking a certain action in a certain state; Var k [Q k (s,a)] is Q k The variance of (s,a) measures the uncertainty of the expected reward of the action; is the variance threshold of the action, indicating the maximum allowed uncertainty; π backup (s) is the braking strategy adopted when the variance exceeds the threshold, which is used to avoid excessive risk or high uncertainty.
[0160] (4) Complete uncertainty estimation
[0161] The complete uncertainty estimation combines the quantile network and the integrated network considering Bayesian prior into a new algorithm through data uncertainty estimation and model uncertainty estimation, combining the quantile reward distribution output and uncertainty estimation of each sub-model to optimize the reward value Y of the state-action pair k,τ (s,a), the calculation formula is:
[0162]
[0163] Among them, Y k,τ (s,a) is the k-th model’s estimate of the reward value for state s and action a at the quantile τ, f τ (s,a;θ k ) is the kth model based on parameter θ k Quantile return forecasts for , return value estimates generated by the quantile network, It is the uncertainty estimate of the quantile τ based on the model generated by the Bayesian posterior distribution;
[0164] The final output is the fused reward value, that is, the reward value Z of k models k,τ The mean of (s,a) is expressed as:
[0165]
[0166] Select the optimal strategy based on the fused return value:
[0167]
[0168] Among them, the agent defines an end-to-end decision-making mechanism based on confidence evaluation and braking strategy according to the variance of the reward distribution: when the model strategy confidence is strong, that is, and The strategy selection is driven by the expected value of the return value prediction. When the confidence is weak, that is, when the variance of the uncertainty estimate is large, it switches to the braking strategy to reduce the potential risk of low-confidence decisions. The specific formula is as follows:
[0169]
[0170] Among them, Z k,τ (s,a) is the quantile return value prediction of k models for state and action, Take the expectation within the confidence interval for the quantile τ, The expected value of the return predictions of the k models, representing the overall prediction of the ensemble model.
[0171] An embodiment of the present invention further provides an end-to-end autonomous driving behavior decision system considering uncertainty estimation, including:
[0172] An extraction module is used to input the image information of surrounding traffic participants collected by the camera into the convolutional neural network of the end-to-end autonomous driving decision network for feature extraction to obtain the state information of the surrounding traffic participants;
[0173] The fusion module is used to perform feature splicing and fusion of the vehicle state information and the surrounding traffic parameter state information, and input the obtained fusion features into the fully connected layer of the downstream decision module of the constructed end-to-end autonomous driving decision model agent to obtain the agent decision result;
[0174] An estimation module is used to construct an integrated quantile network based on a fully connected layer to obtain a complete uncertainty estimate of the agent's decision results. The complete uncertainty estimate of the agent's decision results includes a data uncertainty estimate obtained based on an implicit quantile network and a model uncertainty estimate obtained based on a multi-model integrated network that considers a priori Bayesian estimation.
[0175] The confidence evaluation module is used to evaluate the confidence of the agent's decision results based on the complete uncertainty estimate. If the confidence is strong, the optimal driving action of the vehicle predicted by the integrated quantile network is output. If the confidence is weak, it switches to the braking strategy output;
[0176] The execution module is used to input the predicted optimal driving action and braking strategy of the vehicle as strategy information into the control module of the end-to-end autonomous driving decision model intelligent body, output control signals, and make the vehicle perform different control actions according to the output control signals. The different control actions include acceleration, braking, and steering.
[0177] Compared with the prior art, the present invention has the following beneficial effects:
[0178] The present invention proposes an end-to-end autonomous driving behavior decision-making method and system considering uncertainty estimation, which improves the safety and reliability of end-to-end autonomous driving decisions. The present invention introduces an integrated quantile network, combines distributed reinforcement learning with a multi-model integration method, and obtains a complete uncertainty estimate of the intelligent agent's decision results, including obtaining data uncertainty estimates based on the implicit learning return probability distribution of the quantile function, and obtaining model uncertainty estimates based on a multi-model integration method considering a priori Bayesian estimates; when making decisions, the present invention comprehensively considers the potential returns at the quantile level and the conditional risk value that quantifies extreme risks, and dynamically balances benefits and risks based on the prediction results to select the optimal driving behavior. In addition, the uncertainty of the decision results is quantified, and the present invention also sets a braking strategy corresponding to the decision results with higher uncertainty, so as to prevent the intelligent agent from making unfounded and potentially dangerous decisions.
[0179] As another example, the present invention also provides an electronic device, and now will describe an electronic device that can be used as a server or client of the present invention, which is an example of a hardware device that can be applied to various aspects of the present invention. Electronic devices are intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic devices can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples, and are not intended to limit the implementation of the present invention described and / or required herein.
[0180] The electronic device may include: a processor (processor), a communication interface (CommunicationsInterface), a memory (memory) and a communication bus.
[0181] The processor, the communication interface and the memory communicate with each other through the communication bus. The communication interface is used to communicate with other electronic devices or servers.
[0182] The processor is used to execute the program, and specifically can execute the relevant steps in the above method embodiment.
[0183] Specifically, the program may include program codes including computer operation instructions.
[0184] The processor may be a CPU, or an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention. The one or more processors included in the smart device may be processors of the same type, such as one or more CPUs; or processors of different types, such as one or more CPUs and one or more ASICs.
[0185] The memory is used to store programs and may include a high-speed RAM memory and may also include a non-volatile memory, such as at least one disk memory.
[0186] When the program is executed by a processor, it is used to enable the electronic device to perform the end-to-end autonomous driving behavior decision-making method considering uncertainty estimation of the present invention.
[0187] In addition, the specific implementation of each step in the program can refer to the corresponding description of the corresponding steps and units in the above method embodiment, which will not be repeated here. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process of the above-described devices and modules can refer to the corresponding process description in the above method embodiment, which will not be repeated here.
[0188] The exemplary embodiments of the present invention further provide a computer storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the methods of the various embodiments of the present invention. The corresponding process descriptions in the aforementioned method embodiments may be referred to and will not be repeated here.
[0189] The above-described method according to an embodiment of the present invention may be implemented in hardware, firmware, or as software or computer code that may be stored in a recording medium (such as a CD ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or as computer code that is originally stored in a remote recording medium or a non-temporary machine-readable medium downloaded over a network and will be stored in a local recording medium, so that the method described herein may be stored in such software processing on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It is understood that a computer, processor, microprocessor controller, or programmable hardware includes a storage component (e.g., RAM, ROM, flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by a computer, processor, or hardware, the method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the method shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the method shown herein.
[0190] Thus far, specific embodiments of the present invention have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims may be performed in a different order and still achieve the desired results. Additionally, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing may be advantageous.
[0191] It should be understood that although this specification is described according to various embodiments, not every embodiment contains only one independent technical solution. This narrative method of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment may also be appropriately combined to form other implementation methods that can be understood by those skilled in the art.
[0192] Finally, it should be noted that the above implementation methods are only used to illustrate the embodiments of the present invention, and are not limitations of the embodiments of the present invention. Ordinary technicians in the relevant technical field can make various changes and modifications without departing from the spirit and scope of the embodiments of the present invention. Therefore, all equivalent technical solutions also belong to the scope of the embodiments of the present invention. The patent protection scope of the embodiments of the present invention should be defined by the claims.
Claims
1. An end-to-end autonomous driving behavior decision method considering uncertainty estimation, characterized in that: include: The image information of surrounding traffic participants collected by the camera is input into the convolutional neural network of the end-to-end autonomous driving decision network for feature extraction to obtain the status information of surrounding traffic participants; The vehicle status information is combined with the status information of surrounding traffic participants, and the obtained fusion features are input into the fully connected layer of the downstream decision module of the constructed end-to-end autonomous driving decision model agent to obtain the agent decision result; An integrated quantile network is constructed based on the fully connected layer to obtain a complete uncertainty estimate of the agent's decision results. The complete uncertainty estimate of the agent's decision results includes a data uncertainty estimate obtained based on an implicit quantile network and a model uncertainty estimate obtained based on a multi-model integrated network that considers a priori Bayesian estimation. Based on the complete uncertainty estimate, the confidence of the agent's decision result is evaluated. If the confidence is strong, the optimal driving action of the vehicle predicted by the integrated quantile network is output. If the confidence is weak, the braking strategy output is switched. The predicted optimal driving action and braking strategy of the vehicle are input as strategy information into the control module of the end-to-end autonomous driving decision model agent, and a control signal is output. According to the output control signal, the vehicle performs different control actions, including acceleration, braking, and steering.
2. The method according to claim 1, characterized in that: The status information of surrounding traffic participants includes the location x of the traffic participants i ,y i , speed v i and heading angle ψ i ,The self-vehicle state information includes the position information and speed information of the self-vehicle.
3. The method according to claim 2, characterized in that The end-to-end autonomous driving decision network uses a deep reinforcement Q network. The core reinforcement learning framework of the deep reinforcement Q network is the optimal strategy π(s) for interactive learning between the agent and the environment. The optimal strategy describes the action a taken by the agent in state s. When the environment moves to a new state s′, the agent receives a reward r. The value of the agent taking action a in state s following strategy π is defined by the state-action value Q-value function: Among them, the optimal strategy π * The Q value is defined as: The deep reinforcement Q network uses a deep neural network instead of the traditional Q value representation to achieve an approximate representation of the optimal Q value function: The deep reinforcement Q network learns the strategy by minimizing the time difference error between the predicted Q value and the target value. The time difference error is expressed as: Among them, r is the discount factor that weighs the current reward and future rewards, and its value range is [0,1], Q(s′,a′; θ - ) is the maximum Q value of all possible actions in the next state s′, which is determined by the target network θ - Calculate, Q(s,a;θ) is the predicted Q value of the current state s and action a, and the parameter θ is the current network weight.
4. The method according to claim 3, characterized in that The network parameter θ is optimized by minimizing the time difference error, and the parameters are updated by back propagation using the gradient descent method to adjust Q(s, a; θ) closer to the target value. The loss function is: in, Represents the expectation of the loss function, specifically the square of the temporal difference error, which is used in reinforcement learning to measure the difference between the predicted value and the actual value.
5. The method according to claim 4, characterized in that Data uncertainty estimation based on implicit quantile network, including: Based on the implicit quantile network, distributed reinforcement learning is introduced to learn the distribution of Q values. Each action a corresponds to a group of quantile prediction values to represent the distribution of the Q value of the action. The expression of Q value is: Z(s,a)~Q(s,a) Among them, Z(s,a) is the Q value distribution of state s and action a, and Q(s,a) is the expected value of Z(s,a), that is, The implicit quantile network learns the quantile function of the distribution by sampling the quantile points. The two quantile samples τ and τ′ are used to estimate the distribution value Z of the state-action pair. τ (s,a) and target value Z τ′ (s′, a′), represents the current estimated quantile value and the target distribution quantile value, quantile sample τ, Implicit mapping f is realized through convolutional neural network θ (s, a, τ), output a set of Q value estimates of the loci, i.e. Z τ (s,a)=f θ (s,a,τ), to capture the Q-value distribution of state-action pairs; The time difference error is used to measure the error between the current quantile sample and the target quantile sample to estimate the Q value distribution. The implicit quantile network learns the Q value distribution based on the time difference error, which is expressed as: Among them, r t is the reward of the current time step, γ is the discount factor that weighs the current reward against future rewards, and θ - represents the parameters of the target network, π * (s) is the maximum policy value of the state-action pair; A risk aversion strategy is adopted to select the maximum strategy value and calculate the conditional risk value. Under a given quantile α, for the distributed value function Z(s,a) of state s and action a, the conditional risk value calculation formula is: The optimal strategy formed based on quantile returns and conditional value at risk is expressed as: Among them, λ1 represents the weight of quantile returns to reflect the priority of returns, and λ2 represents the weight of conditional risk value to reflect the degree of risk aversion; The target distribution is optimized using the distribution fitting loss Wasserstein distance, and the quantiles τ and τ' are sampled N and N' times respectively: Among them, ρ κ is the quantile Huber regression loss function, and the formula is as follows: Among them, κ is a hyperparameter, called the smoothing threshold, κ = α*max(|Z τ′ (s′,a′)|), α is a scaling factor, ranging from [0.1,0.5]; The estimated distribution of reward from trained agents is used to quantify the stochastic uncertainty of the agent’s decision outcomes, including: For K τ uniformly distributed samples τ σ ={i / K τ ∣i∈[1,K τ ]}, the metric uses the estimated return variance of uniformly distributed samples, and its threshold is defined as If the variance is below the threshold, the trained agent follows the optimal strategy formed from the quantile return and the conditional value at risk. If the variance is not below the threshold, the agent is braked by the predefined braking strategy π backup (s) makes decision output, the formula is as follows: in, represents the benefit of generating a decision trajectory τ based on strategy σ after selecting action a in state s, CVaR α (s,a) is the conditional risk value of taking this action at a certain risk level, λ1 and λ2 are hyperparameters to balance the weight between expected return and conditional risk value, represents the variance of the return under trajectory τ, which is used to measure uncertainty. is the upper threshold of the variance associated with action a, used to constrain uncertainty, π backup (s) indicates that the variance exceeds the threshold A conservative braking strategy is used to ensure the stability of the model decision.
6. The method according to claim 5, characterized in that Model uncertainty estimation based on multi-model ensemble network considering a priori Bayesian estimation, including: Perform Bayesian probability modeling based on training data and sample data to establish the prior distribution of model parameters; Calculate the likelihood function of the training data, combine the prior distribution of the model parameters with the likelihood function of the training data, and obtain the posterior distribution p(θ|D), as follows: Among them, p(θ) is the prior distribution of model parameters, which represents the assumption about model parameters θ before unknown data, p(D|θ) is the likelihood function of training data, which represents the probability of training data given model parameters θ, p(D) is a normalized constant representing the marginal probability of training data, and D represents training data; The optimal approximate posterior distribution is found by minimizing the KL divergence, expressed as: The multi-model ensemble network samples multiple sub-models from the posterior distribution and performs weighted averaging, including: based on the approximate posterior distribution generated by the acquired Bayesian modeling, sampling a set of multiple parameters {θ1,θ2,…,θ K }, corresponding to the network parameters of the K sub-models, the prediction results of each sub-model are integrated and weighted averaged, where the weight is based on the probability of the posterior distribution of the model, that is, w k ∝P(θ k |D), the formula is as follows: Among them, the Q value of the multi-model integrated network is calculated by combining Bayesian modeling and the output of each sub-model, expressed as: Among them, f(s,a;θ k ) is the expected return value obtained by training the network based on historical data, is the posterior reward value of Bayesian modeling, and β is the adjustment factor to balance the expected return of training data and the reward value generated by Bayesian modeling; Minimizing the loss function based on the time difference error optimizes the network parameters to train the multi-model ensemble network, where the time difference error of the ensemble model changes, and its loss function is expressed as follows: Among them, r t is the return at the current time step, γ is the discount factor, is the output of the target network; The variance of the Q value distribution in the integrated model is defined as the threshold for measuring the model uncertainty estimation standard. If the sample variance is above a threshold, a predefined braking strategy π is selected backup (s) is the output of the model decision result, and the formula is as follows: in, Indicates that after selecting action a in state s, the strategy The expected Q value under the action value function, that is, the cumulative reward obtained by taking a certain action in a certain state, Var k [Q k (s,a)] is Q k The variance of (s,a) measures the uncertainty of the expected reward of the action, is the variance threshold of the action, indicating the maximum allowed uncertainty, π backup (s) is the braking strategy adopted when the variance exceeds the threshold.
7. The method according to claim 1, characterized in that The complete uncertainty estimation combines the quantile network and the integrated network considering Bayesian prior into a new algorithm through data uncertainty estimation and model uncertainty estimation, combining the quantile reward distribution output and uncertainty estimation of each sub-model to optimize the reward value Y of the state-action pair k,τ (s,a), the calculation formula is: Among them, Y k,τ (s,a) is the k-th model’s estimate of the reward value for state s and action a at the quantile τ, f τ (s,a;θ k ) is the kth model based on parameter θ k Quantile return forecasts for , return value estimates generated by the quantile network, It is the uncertainty estimate of the quantile τ based on the model generated by the Bayesian posterior distribution; The final output is the fused reward value, that is, the reward value Z of k models k,τ The mean of (s,a) is expressed as: Select the optimal strategy based on the fused return value: Among them, the agent defines an end-to-end decision-making mechanism based on confidence evaluation and braking strategy according to the variance of the reward distribution: when the model strategy confidence is strong, that is, and The strategy selection is driven by the expected value of the return value prediction. When the confidence is weak, that is, when the variance of the uncertainty estimate is large, it switches to the braking strategy to reduce the potential risk of low-confidence decisions. The specific formula is as follows: Among them, Z k,τ (s,a) is the quantile return value prediction of k models for state and action, Take the expectation within the confidence interval for the quantile τ, The expected value of the return predictions of the k models, representing the overall prediction of the ensemble model.
8. An end-to-end autonomous driving behavior decision system considering uncertainty estimation, characterized in that: include: An extraction module is used to input the image information of surrounding traffic participants collected by the camera into the convolutional neural network of the end-to-end autonomous driving decision network for feature extraction to obtain the state information of the surrounding traffic participants; The fusion module is used to perform feature splicing and fusion of the vehicle state information and the surrounding traffic parameter state information, and input the obtained fusion features into the fully connected layer of the downstream decision module of the constructed end-to-end autonomous driving decision model agent to obtain the agent decision result; An estimation module is used to construct an integrated quantile network based on a fully connected layer to obtain a complete uncertainty estimate of the agent's decision results. The complete uncertainty estimate of the agent's decision results includes a data uncertainty estimate obtained based on an implicit quantile network and a model uncertainty estimate obtained based on a multi-model integrated network that considers a priori Bayesian estimation. The confidence evaluation module is used to evaluate the confidence of the agent's decision results based on the complete uncertainty estimate. If the confidence is strong, the optimal driving action of the vehicle predicted by the integrated quantile network is output. If the confidence is weak, it switches to the braking strategy output; The execution module is used to input the predicted optimal driving action and braking strategy of the vehicle as strategy information into the control module of the end-to-end autonomous driving decision model intelligent body, output control signals, and make the vehicle perform different control actions according to the output control signals. The different control actions include acceleration, braking, and steering.
9. An electronic device, characterized in that: include: processor; A memory for storing programs; The program includes instructions, which, when executed by the processor, cause the processor to perform the steps of the method as claimed in any one of claims 1 to 7.
10. A computer storage medium, characterized in that: A computer program is stored thereon, and when the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Hazardous article detection method based on high-energy and low-energy images
CN120374608A
TBM tunneling parameter intelligent optimization decision-making system based on LSTM network
CN121024627A
Intelligent Optimization Decision System for TBM Tunneling Parameters Based on LSTM Network
CN121024627B