Autonomous Driving Decision-Making Method and System Based on Generative World Large Model and Multi-Step Reinforcement Learning
By combining generative world model and multi-step reinforcement learning, the problem of insufficient prediction accuracy and reliability of autonomous driving systems in complex traffic environments in the prior art is solved, high-precision behavior prediction and stable decision-making are achieved, and road passage efficiency and safety of autonomous driving are improved.
Patent Information
- Application Number
- CN202410826646.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-25
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-06-25
AI Technical Summary
The existing generative world model has limited prediction accuracy and reliability in dynamic and complex traffic environments, making it difficult to accurately predict all potential risks and situations, resulting in unstable performance of autonomous driving systems in actual operations.
The autonomous driving decision-making method based on generative world model and multi-step reinforcement learning is adopted, and the trajectory of surrounding traffic participants is predicted through generative world model, and the uncertain behavior is transformed into deterministic behavior, and the decision-making system is guided to learn in a safe and efficient direction through multi-step reinforcement learning, and finally obtain a high-precision behavior prediction autonomous driving decision network.
It improves the prediction accuracy and reliability of the autonomous driving decision system, enhances the stability and robustness of the decision-making, and effectively improves the efficiency and safety of the passing of autonomous driving roads.
Smart Images

Figure CN118790287B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent driving of automobiles, and particularly to an autonomous driving decision-making method and system based on a generative world large model and multi-step reinforcement learning. Background Art
[0002] The goal of autonomous driving vehicles is to make road transportation safer and more efficient. As one of the key modules of autonomous driving vehicles, the autonomous driving decision-making system enables autonomous driving vehicles to select appropriate driving actions in various driving scenarios.
[0003] In the context of automotive intelligence, autonomous driving decision-making methods have evolved into two major categories: 1) Rule-based methods, such as finite state machines, rely on human driving experience and knowledge to manually design driving rules. However, as the complexity of driving scenarios increases, the number of states and state transition parameters will grow exponentially, making it difficult to ensure their reliability. 2) Learning-based methods, which are divided into two types: deep learning-based methods and reinforcement learning-based methods. ① Deep learning-based methods use deep neural networks to learn driving data samples to achieve reasonable decision-making for vehicles. This method has high decision-making accuracy for specific scenarios, but due to its dependence on high-quality data sets, its universality in dynamic scenarios is poor. ② Reinforcement learning-based methods do not require sample data sets and learn optimal strategies through continuous interaction and trial-and-error with the environment, and can easily handle complex and changing traffic scenarios.
[0004] Real-world traffic is highly complex and uncertain. Predicting the future behavior trajectories of surrounding traffic participants can assist the decision-making system in outputting safer and more efficient driving instructions. The generative world large model is an effective means for processing autonomous driving trajectory prediction that has emerged in recent years. The generative world large model learns the general representation and underlying operating laws of the real world, and generates future world states from a series of driving actions, generating high-fidelity multi-view videos in driving scenarios. MILE adopted a model-based imitation learning method to learn the dynamic model and driving behavior in CARLA, verifying the rationality and diversity of the generative world large model in predicting future states and actions. However, in a dynamically complex traffic environment, its prediction accuracy and reliability are still limited, and it may not be able to accurately foresee all potential risks and situations. In addition, existing generative world large models may not be able to make reasonable and robust decisions in the face of sudden traffic conditions, resulting in unstable performance of the autonomous driving system in actual operation. In summary, there are still problems of limited prediction accuracy and insufficient decision-making stability in the application of existing generative world large models in autonomous driving decision-making. Summary of the Invention
[0005] The objective of the present invention is to implement an autonomous driving decision-making system in an interactive scenario, alleviate problems such as difficult system modeling caused by unclear traffic intentions of surrounding vehicles in traditional decision-making systems, and provide an autonomous driving decision-making method and system based on a generative world large model and multi-step reinforcement learning. By predicting the trajectories of surrounding traffic participants through the generative world large model, the uncertain behaviors of surrounding traffic participants are transformed into definite behaviors; then, through multi-step reinforcement learning, the autonomous driving decision-making system is guided to learn towards a safe and efficient decision-making direction, and finally, an autonomous driving decision-making network with high-precision behavior prediction is obtained, which is of great significance for realizing precise decision-making in autonomous driving and can be generalized to various autonomous driving interactive decision-making scenarios for application, effectively improving the road passing efficiency and safety of autonomous driving.
[0006] The objective of the present invention can be achieved through the following technical solutions:
[0007] An autonomous driving decision-making method based on a generative world large model and multi-step reinforcement learning, comprising the following steps:
[0008] Step 1: Establish a driving scenario inference model based on the generative world large model, predict the behaviors of surrounding traffic participants, and output future driving scenario information;
[0009] Step 2: Based on the future driving scenario information, use the reinforcement learning algorithm to perform multi-step look-ahead offline training on the intelligent agent to obtain the optimal value policy network;
[0010] Step 3: Based on the future driving scenario information and the optimal value policy network, use Monte Carlo tree search to online solve the optimal decision sequence and perform rolling optimization;
[0011] Step 4: Establish a trajectory tracking controller for intelligent connected electric vehicles, and control the autonomous driving vehicle to perform real-time trajectory tracking based on the optimal decision sequence.
[0012] The driving scenario inference model integrates multiple heterogeneous inputs using a unified input interface, and the input conditions supported by the interface include:
[0013] 1. Image input: Image input and video input share the same interface. The interface processes the initial context framework and reference view as image input data, encodes and flattens the given image conditions into a d-dimensional embedding sequence, uses ConvexNet as the encoder, and extracts the embedding information from different images and connects them in an n-dimensional feature vector;
[0014] 2. Layout input: The layout input includes 3D boxes, high-definition maps, and BEV segmentation. Project the 3D boxes and high-definition maps into a 2D perspective view, and encode the layout conditions using the same strategy as the image condition encoding to generate an embedding sequence;
[0015] 3. Text Input: Following the convention of diffusion models, a pre-trained CLIP is used as the text encoder to obtain the embedding information of the text input;
[0016] 4. Action Input: Define the action in the time step as (x, y, v), which represents the movement trajectory of the ego vehicle in the future time step. Here, x and y are the position information of the ego vehicle in the Cartesian coordinate system, and v is the speed information of the ego vehicle. Use a multi-layer perceptron to map the action to a d-dimensional embedding.
[0017] The driving scenario inference model introduces a time layer encoding layer to upgrade the pre-trained image diffusion model to a time model, encodes the latent in a frame-by-frame manner, rearranges the time dimension of the latent, and introduces a spatial encoding layer to upgrade the single-viewpoint temporal model to a multi-viewpoint temporal model, and rearranges the latent space to maintain the view dimension, extracts the latent information of the driving scenario in space-time, and uses the auxiliary supervision output from 3D detection and segmentation tasks to handle the partial observability of perception, predicts the future movement trajectories of surrounding traffic participants, and outputs image patches containing spatio-temporal information.
[0018] In step 2, use the driving scenario information inferred by the generative world large model as the input state, define the state space of driving decisions, that is, describe the states of traffic participants in the driving scenario, and define the action space, that is, various actions that can be taken by the autonomous driving system. Use the collected driving data and adopt a multi-step reinforcement learning algorithm for offline training; among them, during the training process, the agent selects actions according to the current state, interacts with the environment and observes the rewards, and updates the policy to maximize the long-term cumulative reward; in the framework of multi-step look-ahead, the agent considers the action sequences at multiple future moments, predicts all actions and state transitions within the next n steps and calculates the expected return reward, continuously calculates the environmental state transition and the probability distribution of action values, and finally obtains the converged value policy network.
[0019] The generative world large model uses Transformer as the main body of the model. Input the last T time steps into Transformer, with a total of 3*T state tokens. Among them, each time step contains three tokens: expected return, state, and action; for non-image inputs, learn a linear layer to project the original input into the embedding dimension, and then perform layer normalization to obtain the token embedding; for image inputs, the state is sent to a convolutional encoder to obtain the embedding; learn the embedding for each time step and add it to each token, and the Transformer model processes the tokens to predict the future action values through autoregressive modeling.
[0020] In step 3, the integration of driving decisions involves the construction and traversal of a tree structure, which represents the possible action sequences that the ego vehicle can take and the associated costs; the tree structure consists of nodes and edges, each node represents a specific state of the environment, and each edge represents the action taken by the ego vehicle. Among them, the nodes include a root node and child nodes. The root node represents the current state of the environment, including the local route, the state of the ego vehicle, and the states of other nearby vehicles; child nodes are generated by considering the possible longitudinal and lateral movements that the ego vehicle can make from the current state. Longitudinal movements include speed acceleration, deceleration with different accelerations, and maintaining the current speed, while lateral movements include lane keeping, left lane change, and right lane change; the tree is traversed by iteratively selecting actions and transitioning to the corresponding child nodes until a terminal state is reached.
[0021] In step 3, the selection of actions is guided by the upper confidence bound value, and the calculation method of the upper confidence bound value is as follows:
[0022]
[0023] where Q(v′) is given by the state-action value function obtained from the reinforcement learning training in step 2, n(v′) is the number of times the child node v′ has been visited, N is the total number of times the parent node v i has been visited, const is a constant, and C(v′) is the total cost associated with the child node v′, that is, the negative of the current value of the action:
[0024]
[0025] where C s (t), C c (t), C p (t), and C o (t) are the safety, comfort, passivity, and other factor costs at time t, respectively; ω s , ω c , ω p , and ω o are the weights associated with safety, comfort, passivity, and other factors, respectively; T is the total time horizon.
[0026] The Monte Carlo tree search includes the following process:
[0027] 1) Look-ahead process: The ego vehicle looks ahead for a preset number of steps, where each step corresponds to a fixed time interval T 1 , and in each step, the Monte Carlo tree search algorithm selects an action from the set of possible actions of the current node and transitions to the corresponding child node;
[0028] 2) Rollout process: The behavior of the ego vehicle is randomly generated with a given movement probability, and the rollout process is executed until the terminal state is reached;
[0029] 3) Terminal state: In the terminal state, calculate the total cost associated with the action sequence taken by the ego vehicle;
[0030] 4) Backpropagation: After simulating reaching the terminal state and calculating the total cost, backpropagate the total cost through the search tree, starting from the leaf nodes and tracing back to the root node, updating the cumulative cost and visit count of each node encountered during this simulation;
[0031] 5) Repeat steps 1)-4) until the termination condition is reached.
[0032] Step 4 includes the following steps:
[0033] Step 41: Establish a dynamic model of the vehicle to describe the motion characteristics of the vehicle at different speeds and accelerations;
[0034] Step 42: Define the state quantity as the error value between the actual trajectory and the reference trajectory, establish a quadratic programming problem for trajectory tracking, and combine the control Lyapunov function to make the sum of trajectory tracking errors approach 0, and combine the control barrier function to ensure that the vehicle state error always remains within a certain range;
[0035] Step 43: Solve the quadratic programming problem to obtain the vehicle control quantity and achieve autonomous driving trajectory tracking.
[0036] An autonomous driving decision-making system based on a generative world large model and multi-step reinforcement learning is used to implement the method as described above. The system includes:
[0037] Driving scenario inference module: Used to establish a driving scenario inference model based on the generative world large model, predict the behaviors of surrounding traffic participants, and output future driving scenario information;
[0038] Reinforcement learning training module: Used to perform multi-step look-ahead offline training on the agent based on the future driving scenario information using the reinforcement learning algorithm to obtain the optimal value policy network;
[0039] Optimal decision sequence solving module: Used to online solve the optimal decision sequence and perform rolling optimization based on the future driving scenario information and the optimal value policy network using Monte Carlo tree search;
[0040] Trajectory tracking control module: Used to establish a trajectory tracking controller for intelligent connected electric vehicles and control the autonomous driving vehicle to perform real-time trajectory tracking based on the optimal decision sequence.
[0041] Compared with the prior art, the present invention has the following beneficial effects:
[0042] 1. Enhance prediction accuracy and reliability: The present application adopts a generative world large model, which can learn and simulate the changes in complex driving scenarios, thereby providing accurate scenario predictions and providing reliable input data for the decision-making system. And compared with existing methods, the autonomous driving system proposed in the present application combines the prediction results of the generative world large model, and can predict multiple future moments through the multi-step look-ahead ability, so as to more accurately predict potential risks and situations, and improve the accuracy and reliability of the prediction.
[0043] 2. Improve the stability and robustness of decision-making: The present application adopts a rolling optimization strategy. By adjusting and optimizing the decision-making strategy in real time, the system can still maintain stable decision-making performance in the face of emergencies and uncertain environments, and improve the robustness of the system. In addition, the present application designs a dynamic controller for intelligent networked electric vehicles in combination with a feedback control mechanism, providing a real-time feedback and adjustment mechanism to ensure the stability and safety of the vehicle during actual driving.
[0044] 3. Improve the utilization efficiency of computing resources: The present application optimizes the computing resource allocation during the deployment of the autonomous driving algorithm on the vehicle side. By transferring the complex training process to the offline stage, the computing resource requirements in the online stage are reduced, and the utilization efficiency of the resources during the system operation is improved. Since the computing burden of the system is relatively light during actual operation, the requirements for hardware devices are relatively reduced, and the overall development cost of autonomous driving vehicles can be saved. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 is a schematic flowchart of the method of the present invention;
[0046] Figure 2 is a schematic internal working principle diagram of the driving scenario inference module;
[0047] Figure 3 is a schematic flowchart of the multi-step reinforcement learning training process combined with inference information;
[0048] Figure 4 is a schematic diagram of the multi-step reinforcement learning training result;
[0049] Figure 5 is a schematic diagram of the optimal decision generated by the host vehicle using Monte Carlo tree search;
[0050] Figure 6 is a schematic diagram of the trajectory tracking effect of the vehicle controller. DETAILED DESCRIPTION OF THE INVENTION
[0051] The present invention will be described in detail below with reference to the drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and gives detailed implementation manners and specific operation processes, but the protection scope of the present invention is not limited to the following embodiments.
[0052] The present invention discloses a generative video large model and a multi-step reinforcement learning-based autonomous driving decision-making system and method. By taking the driving scenario video collected by the vision camera of an intelligent connected vehicle as input, the generative video large model extracts the encoded information in the video and image latent space, understands the current driving scenario and the historical actions of traffic participants, uses the auxiliary supervision output perspective-invariant representation from 3D detection and segmentation tasks to handle the partial observability of perception, predicts the future movement trajectories of surrounding traffic participants, and predicts and outputs image patches containing spatio-temporal information; then, the sequential movement trajectories predicted by the large model are put into a multi-step reinforcement learning framework, and the stochastic Markov process is described as a deterministic Markov process to guide the reinforcement learning agent to learn towards a safe and efficient decision-making direction; finally, an autonomous driving decision-making network based on high-precision behavior prediction is obtained. This decision-making network constructs a Monte Carlo search tree for the action space within multiple future time domains, looks ahead at multi-step future information, and executes the first action, and then cyclically rolls and optimizes the above process.
[0053] This embodiment first provides an autonomous driving decision-making method based on a generative world large model and multi-step reinforcement learning, as Figure 1 shown, including the following steps:
[0054] Step 1: Establish a driving scenario inference model based on the generative world large model, predict the behaviors of surrounding traffic participants, and output future driving scenario information.
[0055] In autonomous driving, predicting future events in advance and evaluating foreseeable risks enable autonomous driving vehicles to better plan their actions and improve safety and efficiency on the road. To this end, the present invention proposes a driving scenario inference model based on a generative world large model, which is a driving world model compatible with existing end-to-end planning models. Through a convenient view decomposition of joint spatio-temporal modeling, this inference model can generate high-fidelity multi-viewpoint videos in the driving scenario, predict future state changes according to the current states of traffic participants and self-behaviors before actual decision-making, help autonomous driving vehicles make more reasonable plans, and enhance the generality and safety of end-to-end autonomous driving.
[0056] Specifically, the autonomous driving system proposed by the present invention first converts the autonomous driving scenario data collected by the perception end into a form that can be understood by a computer. This data includes camera images, lidar scan data, GPS positions, etc. Each data source provides information about different aspects of the scenario. Secondly, feature extraction is performed on this data, including tasks such as road detection, vehicle and pedestrian recognition, and obstacle detection. The purpose is to convert the original data into a higher-level and more abstract representation for subsequent processing. Then, the extracted extensive interaction scenario information in the physical world is sent to the generative world large model. Utilizing the characteristic that the large model can understand the underlying physical laws of the world's operation, it simulates the movement and interaction of traffic participants within future moments to infer the changes in the driving scenario at multiple future moments. Finally, the generative world large model extracts the encoded information in the latent space of the video and image, uses the auxiliary supervised output perspective-invariant representation from 3D detection and segmentation tasks to handle the partial observability of perception, predicts the future movement trajectories of surrounding traffic participants, and outputs image patches containing spatio-temporal information.
[0057] Inspired by the latent video diffusion model, the present invention introduces multi-view and temporal modeling for jointly generating multiple views and video frames. To further improve the consistency of multi-views, a simple and effective unified conditional input interface is also introduced, which can flexibly use heterogeneous conditions such as images, texts, 3D layouts, and actions as the input of the interface, greatly expanding the applicable scope of the model. The internal working principle diagram of the model is as Figure 2 shown.
[0058] Before the driving scenario inference model based on the generative world large model established in this step inputs the current scenario information to the world model for inference, it needs to undergo data processing through the unified conditional input interface and the spatio-temporal encoding layer. The specific implementation methods of these two modules will be introduced in detail below.
[0059] 1. Unified Conditional Input Interface
[0060] Due to the huge complexity of the real world, the world model needs to utilize multiple heterogeneous conditions. The present invention utilizes the initial context framework, text description, self-action, 3D box, BEV map, and reference view. However, developing dedicated interfaces for each input condition is time-consuming and inflexible. To solve this problem, the present invention introduces a unified input interface that simply and effectively integrates multiple heterogeneous inputs. This unified conditional input interface can convert various types of received raw data into a unified higher-level and more abstract representation for subsequent processing. The input conditions supported by this interface include:
[0061] 1) Image input: Since a video is a sequence of frames composed of multiple images, the image input and video input can share the same interface. First, the initial context frame (i.e., the first frame of the clip) and the reference view are processed as image input data. Then, the given image conditions are encoded and flattened into a sequence of d-dimensional embeddings using ConvexNet as the encoder. Finally, the embedding information from different images is extracted in the encoder and concatenated in an n-dimensional feature vector. Here, H and W are the height and width of the image respectively, d is the length of the embedding feature sequence extracted from the image, n is the dimension of each image feature vector, and i n is the image feature sequence of the nth dimension.
[0062] 2) Layout input: The layout input refers to 3D boxes, high-definition maps, and BEV segmentation. For simplicity, the 3D boxes and high-definition maps are projected into a 2D perspective view. In this way, the same strategy as the image condition encoding is used to encode the layout condition, resulting in an embedding sequence where k is the dimension of the feature vector from the projected layout and BEV segmentation, d is the length of the extracted embedding feature sequence, and l k is the layout feature sequence of the kth dimension.
[0063] 3) Text input: Following the convention of the diffusion model, the pre-trained CLIP is adopted as the text encoder. Specifically, the combine harvester view information, weather, and lighting are combined to obtain the text description. The embedding information is represented as where m is the dimension of each text feature vector, and e m is the text feature sequence of the mth dimension.
[0064] 4) Action input: The action input is indispensable for the world model to infer future changes. To be compatible with existing planning methods, we define the action in the time step as (x, y, v), representing the movement trajectory of the ego vehicle in the future time step, and use a multi-layer perceptron to map the action to a d-dimensional embedding where x and y are the position information of the ego vehicle in the Cartesian coordinate system, and v is the speed information of the ego vehicle.
[0065] 2. Spatio-temporal encoding layer
[0066] First, a time layer is introduced to upgrade the pre-trained image diffusion model to a time model. The time encoding layer is attached after the 2D spatial layer in each block to process the latent Encode, rearrange the potential time dimension, expressed as (TK)CHW → KCTHW, to apply 3D convolution in the spatio-temporal dimension THW, arrange the potential (KHW)TC, and apply the time dimension of the standard multi-head self-attention to enhance temporal dependence, where T is the length of the image sequence, K is the number of viewpoints, C is the number of image channels, which are the image height and width respectively. Then, introduce a spatial encoding layer to upgrade the single-view temporal model to a multi-view temporal model, enhance the spatial representation ability, and rearrange the latent space as (KHW)TC → (THW)KC to maintain the view dimension. Then, optimize the objective function by minimizing denoising:
[0067]
[0068] such that all views have similar styles and a consistent overall structure. Among them, c is the input condition, the target y is random noise ∈, p τ is a uniform distribution over the diffusion time τ, z τ is the frame sequence at the diffusion time τ. It should be noted that in view of the powerful image diffusion model, the present invention does not train a new spatio-temporal multi-view network from scratch. Instead, first train a standard image diffusion model with single-view image data and conditions, which corresponds to the parameter θ in Equation (1). Then, fix the parameter θ, and use the video data to fine-tune the additional Temporal Layer (θ) and Multi-View Layer (θ) to extract the latent information of the driving scene in the spatio-temporal domain.
[0069] After the driving scene information passes through the unified conditional input interface and the spatio-temporal encoding layer, it will be input to the world model for understanding and reasoning, predict the behaviors of surrounding traffic participants, and be integrated into future driving scene information. This precise reasoning of the future driving scene can help the ego vehicle agent in Step 2 to perform multi-step reinforcement learning training and the ego vehicle online Monte Carlo tree search in Step 3 to solve the optimal trajectory.
[0070] Step 2: Based on the future driving scene information, use the reinforcement learning algorithm to perform multi-step look-ahead offline training on the agent to obtain the optimal value policy network.
[0071] In the present invention, based on the driving scene information inferred by the generative world large model, a multi-step look-ahead offline training method based on reinforcement learning is designed. This method can consider various possible action sequences at multiple future moments, and through continuous interaction with the environment, obtain the optimal value policy network and save it offline.
[0072] Training process: First, use the driving scenario information inferred by the generative world large model as the input state; Second, define the state space of driving decisions, that is, describe the states of traffic participants in the driving scenario, and define the action space, that is, various actions available to the system, such as accelerating, decelerating, changing lanes, overtaking, etc.; Finally, use the collected driving data and select a suitable reinforcement learning algorithm for offline training. During the training process, the agent selects actions according to the current state, interacts with the environment and observes the rewards, and then updates the policy to maximize the long-term cumulative rewards. In the framework of multi-step look-ahead, the agent needs to consider the action sequences at multiple future moments, continuously calculate the environmental state transition and the probability distribution of action values, and finally obtain the converged value policy network.
[0073] In classical reinforcement learning algorithms for autonomous driving, the sequential decision-making problem is modeled using the Markov decision process formulation defined by a tuple. In this formulation, the agent (both learner and actor) and the environment interact at a series of discrete time steps \(t \geq 0\). At each time step, the agent receives a state \(S\) t \(\in \mathcal{S}\), which encodes information about the environment. Based on this information, the agent selects and executes an action \(A\) t \(\in \mathcal{A}\). As a result of executing the action, the environment sends back to the agent a new state \(S'\) t+1 and a reward \(R\) t+1 \(\in \mathcal{R}\). To make informed decisions, the ego-agent often estimates the action-value function \(q\) π , which maps states and actions to values in \(\mathbb{R}\) and is defined as:
[0074]
[0075] where \(\gamma\) is the discount factor and \(\mathbb{E}\) denotes the expectation.
[0076] Reinforcement learning algorithms for estimating the action-value function are typically characterized by the objective of their update functions. To approximate the action-value function, the agent uses an update rule based on a stochastic approximation algorithm of the following form:
[0077]
[0078] where, is the actual cumulative reward, \(\alpha\) is the learning rate, and \(Q\) t (\(S\) t , \(A\) t ) is the action-value function at time \(t\).
[0079] To reduce the myopic effect caused by single-step value updates, one approach is to take longer trajectories that contain more observations of future rewards, i.e., the multi-step reinforcement learning paradigm. Multi-step reinforcement learning estimates the expectations for all possible driving actions. During the training of the ego vehicle agent, at each step, it combines the future multi-step driving scenario changes inferred by the world large model to predict all actions and state transitions within the next n steps (as Figure 3 shown), and calculates the expected return reward:
[0080]
[0081] where γ is the discount factor, G t:t+n is the expected return reward for the next n steps, R t+n is the reward at time t + n, and Q t+n-1 (S t+n , A t+n ) is the action-value function.
[0082] So far, it has been clear that the training objective of multi-step reinforcement learning is to obtain a function Q(s, a) that can accurately estimate the future expected value of state actions, which is used to help the Monte Carlo tree search update the node state in step 3. Since multi-step reinforcement learning is a typical sequence-to-sequence (seq2seq) problem, the framework of sequence modeling problems can be used to represent multi-step reinforcement learning. In addition, since training autonomous vehicles using reinforcement learning in the real world faces high trial-and-error costs, even with the assistance of future inferences given by the generative world large model, the ego vehicle agent may still attempt dangerous driving actions during training, which may lead to traffic accidents. Therefore, an offline reinforcement learning (Offline RL) training method that collects training data online and then performs value-policy iteration offline is adopted, aiming to converge safely and smoothly to obtain the optimal action-value function and policy function.
[0083] The present invention uses a Transformer as the main body of the algorithm model, and inputs the last T time steps into the Transformer. There are a total of 3*T state tokens (each time step contains three tokens: expected return, state, and action). To obtain token embeddings, we learn a linear layer to project the original input into the embedding dimension and then perform layer normalization. For environments with image or video inputs, the state is fed into a convolutional encoder instead of a linear layer. In addition, the embedding for each time step is learned and added to each token, which is different from the standard positional embeddings used by the Transformer because one time step corresponds to three tokens. Then the tokens are processed by the Transformer model, which predicts future action values through autoregressive modeling. During training, mini-batches with a sequence length of K are sampled from the dataset. The prediction head corresponding to the input state is trained to predict the best action, and it is natural to construct the loss function of the Q network in the form of mean squared error:
[0084]
[0085] where Q ω (s i ,a i ) is the action value network, ω is the parameter to be updated in the Q network, r i is the immediate reward at the i-th step, and N is the sampling scale.
[0086] During the evaluation and deployment, this embodiment specifies the target return according to the desired performance (such as the maximum possible return of the current behavior) and the initial state of the environment to initialize the generation. After performing the generated action, the target return is subtracted from the achieved reward, and the next state is obtained. Repeat this process, generate actions and apply them to obtain the next returnable state until the driving task ends. Taking the condition of multi-lane lane change on a highway as an example, the present invention evaluates the performance of multi-step reinforcement learning on this task, and the training results are as Figure 4 shown. It can be seen that the model loss decreases significantly and the expected return reward shows an upward trend, achieving good training results.
[0087] So far, in the framework of multi-step reinforcement learning, the ego vehicle agent can consider the action sequences at multiple future moments by combining the results inferred from the generative world large model, continuously calculate the environmental state transition and the probability distribution of action values, and finally obtain the converged value network Q and policy network Π.
[0088] Step 3: Based on the future driving scenario information and the optimal value policy network, use Monte Carlo tree search to online solve the optimal decision sequence and perform rolling optimization.
[0089] Since the autonomous driving system ultimately needs to obtain the sequence of path points that the vehicle should track, the present invention combines the prediction and inference information of the world large model in step 1 and the reinforcement learning value and policy network in step 2, and uses Monte Carlo Tree Search (MCTS) to online solve the optimal decision sequence, that is, the optimal feasible trajectory planning, and continuously improves the system performance through a rolling optimization strategy.
[0090] The Monte Carlo tree search algorithm is adopted to online search for the optimal decision. This algorithm can generate a multi-step look-ahead decision tree by using the future state of the driving environment inferred by the generative world large model, sample, evaluate, expand and update nodes in the decision tree, and perform a truncated rollout after a fixed step length, and use the value network obtained by offline training to approximate the value from the state node to the terminal node, so as to find the optimal action sequence within multiple future moments. At each decision time step, search for the optimal action sequence within a fixed step length, and select the action sequence with the highest long-term reward as the current decision strategy, and then execute the first decision instruction in this sequence, and then loop and roll to optimize the above process.
[0091] The integration of driving decisions within the MCTS framework involves the construction and traversal of a tree structure, which represents the possible action sequences that an autonomous driving vehicle (ego vehicle) can take and the associated costs. The tree structure consists of nodes and edges, where each node represents a specific state of the environment and each edge represents the action taken by the ego vehicle. 1) Root node: The root node represents the current state of the environment, which includes the local route (reference line), the state of the ego vehicle, and the states of other nearby vehicles. 2) Child nodes: Child nodes are generated by considering the possible longitudinal and lateral movements that the ego vehicle can make from the current state. Longitudinal movements include speed acceleration, deceleration with different accelerations, and maintaining the current speed. Lateral movements include lane keeping, left lane change, and right lane change.
[0092] The tree is traversed by iteratively selecting actions and transitioning to the corresponding child nodes until a terminal state is reached. The selection of actions is guided by the Upper Confidence Bound (UCB) value, which balances the exploration of new actions and the exploitation of known actions. In the UCB formula:
[0093]
[0094] Q(v′) is given by the state-action value function obtained from the reinforcement learning training in step 2, which measures the future expected value of the action. N(v′) is the number of times the child node v′ has been visited. N is the parent node v iThe total number of times that has been accessed. const is a constant that determines the exploration and exploitation level. C(v′) is the total cost associated with the child node v′, which is the opposite of the current value of the action:
[0095]
[0096] Among them, C s (t), C c (t), C p (t) and C o (t) are the safety, comfort, passivity, and other factor costs at time t respectively; ω s , ω c , ω p and ω o are the weights associated with safety, comfort, passivity, and other factors respectively. These weights determine the relative importance of each cost component in the objective function; T is the total time horizon. This algorithmic approach enables the behavior planner to explore potential action sequences, gradually refine the decision-making, maximize the expected goal, while adapting to safety, kinematic, and environmental constraints.
[0097] The MCTS algorithm is divided into the following four processes:
[0098] 1) Look-ahead process: In this embodiment, the ego vehicle looks ahead several steps, where each step corresponds to a fixed time interval T1. In each step, the MCTS algorithm selects an action from the set of possible actions of the current node and then transitions to the corresponding child node.
[0099] 2) Rollout process: After the look-ahead step, the rollout process begins. In this process, the behavior of this vehicle is randomly generated with a given movement probability until a terminal state is reached.
[0100] 3) Terminal state: In the terminal state, the total cost associated with the action sequence taken by this vehicle is calculated based on the cost function (formula (7)) described in the previous section.
[0101] 4) Backpropagation: After the simulation reaches the terminal state and the cost is calculated, this cost is propagated back through the search tree. Starting from the leaf node and tracing back to the root node, the cumulative cost and visit count of each node encountered during this simulation are updated.
[0102] The specific process of the MCTS algorithm is as follows:
[0103]
[0104] The entire process of tree traversal and expansion is repeated multiple times until a termination condition is reached. The termination condition can be based on a fixed number of iterations, a fixed computation time, or other criteria. At the end of the MCTS process, the action associated with the child node leading from the root node to the one with the highest value (balancing the current cost C and the future expected reward Q) is selected as the best action for the ego vehicle to take. Figure 5 Demonstrates the process of the ego vehicle using Monte Carlo tree search to generate an optimal decision trajectory for the scenario of unprotected left turn.
[0105] After the best action is executed by the vehicle controller, the state of the environment will change due to the action and the movement of other vehicles. Therefore, in the next step, the MCTS process is regenerated and the decision-making plan is redone, i.e., rolling optimization. This method ensures that the behavior planner can adapt to the changing environment and make intelligent decisions in real time. The technical details of the vehicle controller will be described in Step 4.
[0106] Step 4: Establish a trajectory tracking controller for intelligent connected electric vehicles, and control the autonomous vehicle to perform real-time trajectory tracking based on the optimal decision sequence.
[0107] In the present invention, since the decision sequence solved online above is actually the trajectory sequence of the ego vehicle, it is necessary to establish a dynamics controller for intelligent connected electric vehicles to control the autonomous vehicle to perform real-time trajectory tracking.
[0108] Establish a dynamics model of the electric vehicle to describe the motion characteristics of the vehicle at different speeds and accelerations, including considering factors such as vehicle mass, inertia, motor characteristics, and tire characteristics. Define the state quantity as the error value between the actual trajectory and the reference trajectory, and establish a quadratic programming (QP) problem for trajectory tracking. Finally, combine the control Lyapunov function (CLF) to make the sum of the trajectory tracking errors approach 0, that is, the stability problem of the controller; combine the control barrier function (CBF) to ensure that the vehicle state error is always kept within a certain range, that is, the safety problem of the controller.
[0109] This embodiment plans to use a four-wheel drive electric vehicle to model and solve the trajectory tracking problem.
[0110] The dynamics model of the four-wheel drive electric vehicle is as follows:
[0111]
[0112] Among them, M is the vehicle mass, I z is the moment of inertia of the vehicle about the z-axis, l fand l r is the distance from the center of gravity to the longitudinal axis, C f and C r are the lateral stiffnesses of the front and rear wheels respectively. a x is the longitudinal acceleration, δ f is the front wheel steering angle, M z is the yaw moment. d u is the disturbance.
[0113] The four-wheel drive electric vehicle model includes 6 state variables v = [v x , v y , ω r T , and 3 control variables u = [a x , δ f , M z T , and the disturbance d u = [d u1 , d u2 , d u3 T :
[0114]
[0115] Among them, the matrix can be expressed as:
[0116]
[0117]
[0118] Among them, v = [v x , v y , ω r T are state variables, representing position and velocity respectively.
[0119] According to the four-wheel drive electric vehicle model formula, in order to describe the error model for the trajectory tracking problem, the state variables are defined as the error values between the actual trajectory and the reference trajectory, where then the new state variable is e η = η - η r = [e η1 , e η2 , e η3 T , the control variable is still u = [a x , δ f , M z T , and the relative error model can be expressed as:
[0120]
[0121] Among them
[0122] So far, the dynamic error model of the four-wheel drive electric vehicle has been modeled. In order to solve the trajectory tracking problem, the following control objectives are defined:
[0123] Objective 1: The sum of the trajectory tracking errors approaches 0, which is solved by the CLF.
[0124] Objective 2: The control quantity is as small as possible, which is solved by QP.
[0125] Safety constraint: The error values of the three position variables during trajectory tracking are kept within the range, which is solved by HOCBF.
[0126] Control variable constraint: Directly used as a hard constraint.
[0127] Express the control objective as an optimization problem with constraints:
[0128]
[0129] Among them, K u and K l1 are coefficients, ι 1 is a slack variable representing a soft constraint, ψ 0,max :=e η,max -e η (t)=[ψ 0,max1 ,ψ 0,max2 ,ψ 0,max3 T , which is a high-order constraint.
[0130] If using the Python language for online solution, solvers such as OSQP and CVXPY can be called; if using the C++ language, solvers such as CPLEX and CasADi can be called to solve the above problems. The specific solution process belongs to the conventional settings in this field. To avoid obscuring the purpose of this application, it will not be elaborated here. Figure 6 It shows the performance of the vehicle controller designed in this patent for the trajectory tracking problem. It can be seen that the vehicle controller has achieved excellent control effects.
[0131] This embodiment also provides an autonomous driving decision-making system based on a generative large world model and multi-step reinforcement learning for implementing the method as described above. The system includes:
[0132] Driving scenario inference module: used to establish a driving scenario inference model based on the generative world large model, predict the behaviors of surrounding traffic participants, and output future driving scenario information;
[0133] Reinforcement learning training module: used to perform multi-step look-ahead offline training on the agent based on the future driving scenario information by using the reinforcement learning algorithm to obtain the optimal value policy network;
[0134] Optimal decision sequence solving module: used to online solve the optimal decision sequence and perform rolling optimization based on the future driving scenario information and the optimal value policy network by using Monte Carlo tree search;
[0135] Trajectory tracking control module: used to establish a trajectory tracking controller for the intelligent connected electric vehicle, and control the autonomous vehicle to perform real-time trajectory tracking based on the optimal decision sequence.
[0136] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the described modules can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0137] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative work. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field based on the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the prior art should be within the protection scope determined by the claims.
Claims
1. An autonomous driving decision-making method based on a generative world big model and multi-step reinforcement learning, characterized in that: The following steps are involved: Step 1: Establish a driving scenario reasoning model based on the generative world big model, predict the behavior of surrounding traffic participants, and output future driving scenario information; Step 2: Based on the future driving scenario information, use the reinforcement learning algorithm to perform multi-step forward offline training on the agent to obtain the optimal value strategy network; Step 3: Based on the future driving scenario information and the optimal value strategy network, use Monte Carlo tree search to solve the optimal decision sequence online and perform rolling optimization; Step 4: Establish a trajectory tracking controller for intelligent connected electric vehicles to control the autonomous driving vehicle to perform real-time trajectory tracking based on the optimal decision sequence; In step 2, the driving scene information inferred by the generative world big model is used as the input state to define the state space of driving decision-making, that is, to describe the state of traffic participants in the driving scene, and to define the action space, that is, various actions that can be taken by the autonomous driving system. The collected driving data is used to perform offline training using a multi-step reinforcement learning algorithm; wherein, during the training process, the intelligent agent selects actions according to the current state, interacts with the environment and observes rewards, and updates strategies to maximize long-term cumulative rewards; in the framework of multi-step look-ahead, the intelligent agent considers the action sequence at multiple moments in the future, predicts all actions and state transitions in the next n steps and calculates the expected return reward, continuously calculates the probability distribution of environmental state transitions and action values, and finally obtains a converged value strategy network; The Monte Carlo tree search includes the following process: 1) Look-ahead process: The ego vehicle looks ahead for a preset number of steps, where each step corresponds to a fixed time interval T1. In each step, the Monte Carlo tree search algorithm selects an action from the possible action set of the current node and transitions to the corresponding child node; 2) Roll-up process: The behavior of the ego vehicle is randomly generated with a given movement probability, and the roll-up process is performed until the terminal state is reached; 3) Terminal state: In the terminal state, the total cost associated with the sequence of actions taken by the ego vehicle is calculated; 4) Backward propagation: After the simulation reaches the terminal state and the total cost is calculated, the total cost is backpropagated through the search tree, starting from the leaf nodes and tracing back to the root node, updating the cumulative cost and visit count of each node encountered during this simulation; 5) Repeat the process 1)-4) until the termination condition is reached.
2. The autonomous driving decision-making method based on a generative world big model and multi-step reinforcement learning according to claim 1, characterized in that: The driving scenario reasoning model uses a unified input interface to integrate multiple heterogeneous inputs. The input conditions supported by the interface include:
1. Image input: Image input and video input share the same interface, which processes the initial context frame and reference view as image input data, encodes and flattens the given image condition into a d-dimensional embedding sequence, and uses ConvexNet as an encoder, in which the embedding information from different images is extracted and connected in an n-dimensional feature vector; 2. Layout input: The layout input includes 3D boxes, HD maps and BEV segmentation. The 3D boxes and HD maps are projected into 2D perspective images, and the layout conditions are encoded using the same strategy as image conditional encoding to generate an embedded sequence; 3. Text input: Following the convention of the diffusion model, the pre-trained CLIP is used as the text encoder to obtain the embedding information of the text input; 4. Action input: The action in a time step is defined as (x, y, v), which represents the movement trajectory of the ego vehicle in the future time step, where x and y are the position information of the ego vehicle in the Cartesian coordinate system, and v is the speed information of the ego vehicle. A multi-layer perceptron is used to map the action to a d-dimensional embedding.
3. The autonomous driving decision-making method based on a generative world big model and multi-step reinforcement learning according to claim 1, characterized in that: The driving scene reasoning model introduces a temporal encoding layer to upgrade the pre-trained image diffusion model to a temporal model, encodes the potential in a frame-by-frame manner, rearranges the potential time dimension, and introduces a spatial encoding layer to upgrade the single-view temporal model to a multi-view temporal model, and rearranges the potential space to maintain the view dimension, extracts the latent information of the driving scene in time and space, and uses auxiliary supervision from 3D detection and segmentation tasks to output a view-invariant representation to handle the partial observability of perception, predict the future motion trajectory of surrounding traffic participants, and output image blocks containing spatiotemporal information.
4. The autonomous driving decision-making method based on a generative world big model and multi-step reinforcement learning according to claim 1, characterized in that: The generative world model uses Transformer as the main body of the model, and inputs the last T time steps into the Transformer, with a total of 3*T state tags, where each time step contains three tags: expected reward, state, and action; for non-image input, a linear layer is learned to project the original input to the embedding dimension, and then layer normalization is performed to obtain tag embedding; for image input, the state is fed into a convolutional encoder to obtain an embedding; the embedding of each time step is learned and added to each tag, the token is processed by the Transformer model, and the future action value is predicted through autoregressive modeling.
5. The autonomous driving decision-making method based on a generative world big model and multi-step reinforcement learning according to claim 1, characterized in that: In step 3, the integration of driving decisions involves the construction and traversal of a tree structure, the tree structure represents a possible sequence of actions that the ego vehicle can take and the associated costs; the tree structure consists of nodes and edges, each node represents a specific state of the environment, and each edge represents an action taken by the ego vehicle, wherein the node includes a root node and child nodes, the root node represents the current state of the environment, including the local route, the state of the ego vehicle, and the state of other nearby vehicles; child nodes are generated by considering possible longitudinal and lateral movements that the ego vehicle can perform from the current state, the longitudinal movement includes speed acceleration, deceleration with different accelerations, and current speed maintenance, and the lateral movement includes lane keeping, left lane change, and right lane change; the tree is traversed by iteratively selecting actions and switching to corresponding child nodes until a terminal state is reached.
6. The autonomous driving decision-making method based on a generative world big model and multi-step reinforcement learning according to claim 5, characterized in that: In step 3, the action selection is guided by the upper confidence limit value, and the calculation method of the upper confidence limit value is: in, The state-action value function obtained by reinforcement learning training in step 2 is given by n(v′) is the number of times the child node v′ is visited, and N is the number of times the parent node v is visited. i is the total number of times it has been visited, const is a constant, and C(v′) is the total cost associated with the child node v′, which is the opposite of the current value of the action: Among them, C s (t), C c (t), C p (t) and C o (t) are the costs of safety, comfort, passivity and other factors at time t; ω s ,ω c ,ω p and ω o are the weights related to safety, comfort, passivity and other factors respectively; T is the total time frame.
7. The autonomous driving decision-making method based on a generative world big model and multi-step reinforcement learning according to claim 1, characterized in that: The step 4 comprises the following steps: Step 41: Establish a dynamic model of the vehicle to describe the motion characteristics of the vehicle at different speeds and accelerations; Step 42: define the state quantity as the error value between the actual trajectory and the reference trajectory, establish a quadratic programming problem for trajectory tracking, and combine the control Lyapunov function to make the sum of the trajectory tracking errors approach 0, and combine the control obstacle function to ensure that the vehicle state error is always kept within a certain range; Step 43: Solve the quadratic programming problem, obtain the vehicle control quantity, and realize automatic driving trajectory tracking.
8. An autonomous driving decision system based on a generative world model and multi-step reinforcement learning, characterized in that: For implementing the method according to any one of claims 1 to 7, the system comprises: Driving scenario reasoning module: used to establish a driving scenario reasoning model based on the generative world big model, predict the behavior of surrounding traffic participants, and output future driving scenario information; Reinforcement learning training module: used to perform multi-step forward offline training of the intelligent agent based on future driving scenario information using reinforcement learning algorithms to obtain the optimal value strategy network; Optimal decision sequence solving module: used to solve the optimal decision sequence online and perform rolling optimization based on future driving scenario information and the optimal value strategy network using Monte Carlo tree search; Trajectory tracking control module: used to establish a trajectory tracking controller for intelligent connected electric vehicles, and control the autonomous driving vehicle to perform real-time trajectory tracking based on the optimal decision sequence.
Citation Information
Patent Citations
World model-driven learning-type transferable automatic driving method and system
CN116501065A
End-to-end automatic driving method and system based on reinforcement learning driving world model
CN117218618A