End-to-end automatic driving track evaluation method and system based on world model

Through a world model-based method, multimodal sensor data and policy networks are used to generate bird's-eye view state information, combined with a bird's-eye view world model and score prediction network, the problem of insufficient effect of existing trajectory evaluation methods in complex scenarios is solved, and the safety and reliability of the autonomous driving system is improved.

CN120356177APending Publication Date: 2025-07-22INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 0 Cites 10 Cited by

Patent Information

Application Number
CN202510212421.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing end-to-end autonomous driving trajectory evaluation methods rely on perceived results and cannot effectively evaluate trajectory in complex driving scenarios, resulting in insufficient driving safety.

Method used

Using a world model-based method, real-time information on bird's eye view state is generated through multimodal sensor data, real-time trajectory anchor points are output using the policy network, real-time state-action pairs are constructed, and a bird's eye view world model is input to predict future state-actions, and predict predicted driving trajectory is evaluated, and a multi-layer perceptron and score prediction network is combined for comprehensive scoring.

Benefits of technology

It improves the safety and reliability of the autonomous driving system, and can make optimal driving decisions in complex driving environments, ensuring the accuracy and real-time trajectory evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356177A_ABST
    Figure CN120356177A_ABST
Patent Text Reader

Abstract

The invention provides an end-to-end automatic driving track evaluation method and system based on a world model, and the method comprises the steps: generating bird's-eye view state real-time information according to multi-mode sensor data obtained in a current automatic driving scene; inputting the aerial view state real-time information into a strategy network to obtain a plurality of real-time track anchor points output by the strategy network; according to the bird's-eye view state real-time information and the real-time track anchor point, constructing a real-time state-action pair, and inputting the real-time state-action pair into a bird's-eye view world model to obtain a predicted state-action pair output by the bird's-eye view world model; according to the real-time state-action pair and the predicted state-action pair, the predicted driving track corresponding to the predicted state-action pair is evaluated, and an evaluation result of the predicted driving track is obtained. According to the method, the predicted driving track can be evaluated more accurately, and the safety and reliability of the automatic driving system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of autonomous driving, and particularly to an end-to-end autonomous driving trajectory evaluation method and system based on a world model. Background Art

[0002] In recent years, significant progress has been made in end-to-end autonomous driving technology. By integrating perception, prediction, and planning into a fully differentiable framework, automatic control of driverless vehicles has been achieved. However, relying solely on trajectory prediction is not sufficient to ensure driving safety. In practical applications, there will inevitably be some low-quality trajectories, which may lead to unsafe driving results.

[0003] Existing trajectory evaluation methods are mostly rule-based methods that rely on perception results, such as bounding boxes and maps, to evaluate trajectories. However, these methods are sensitive to perception accuracy and cannot be directly optimized in an end-to-end framework. Currently, some end-to-end driving methods have begun to introduce trajectory evaluation, but their evaluation mainly depends on the features of the current state, and the evaluation effect is limited in complex driving scenarios.

[0004] Therefore, there is an urgent need for an end-to-end autonomous driving trajectory evaluation method and system based on a world model to solve the above problems. Summary of the Invention

[0005] Aiming at the problems existing in the prior art, the present invention provides an end-to-end autonomous driving trajectory evaluation method and system based on a world model.

[0006] The present invention provides an end-to-end autonomous driving trajectory evaluation method based on a world model, including: Generating real-time information of the bird's-eye view state according to multi-modal sensor data obtained in the current autonomous driving scenario; Inputting the real-time information of the bird's-eye view state into a policy network to obtain a plurality of real-time trajectory anchor points output by the policy network, wherein each of the real-time trajectory anchor points corresponds to a different driving behavior action; the policy network is trained based on expert driving data and sample trajectory anchor points corresponding to the expert driving data; Constructing a real-time state-action pair according to the real-time information of the bird's-eye view state and the real-time trajectory anchor points, and inputting the real-time state-action pair into a bird's-eye view world model to obtain a predicted state-action pair output by the bird's-eye view world model, wherein the predicted state-action pair includes bird's-eye view state prediction information and driving behavior prediction actions; the bird's-eye view world model is trained based on sample state-action pairs and bird's-eye view state sample information and driving behavior sample actions corresponding to the sample state-action pairs at future moments; Evaluate the predicted driving trajectory corresponding to the predicted state-action pair according to the real-time state-action pair and the predicted state-action pair, to obtain an evaluation result of the predicted driving trajectory.

[0007] According to an end-to-end autonomous driving trajectory evaluation method based on a world model provided by the present invention, the generation of the real-time information of the bird's-eye view state from the multi-modal sensor data obtained in the current autonomous driving scenario includes: Obtain the multi-modal sensor data when the vehicle is driving in the current autonomous driving scenario, where the multi-modal sensor data at least includes multi-view image data and lidar data, and the multi-view image data includes vehicle front-view image data and vehicle left and right-view image data; Perform image cropping preprocessing on the vehicle left and right-view image data in the multi-modal sensor data to obtain preprocessed multi-modal sensor data; Based on a bird's-eye view encoder, convert the preprocessed multi-modal sensor data into the real-time information of the bird's-eye view state.

[0008] According to an end-to-end autonomous driving trajectory evaluation method based on a world model provided by the present invention, the policy network is trained through the following steps: Based on the K-Means clustering algorithm, perform clustering processing on the expert driving data to obtain the sample trajectory anchor points; According to the expert driving data and the sample trajectory anchor points corresponding to the expert driving data, train a neural network to obtain the policy network.

[0009] According to an end-to-end autonomous driving trajectory evaluation method based on a world model provided by the present invention, after inputting the real-time information of the bird's-eye view state into the policy network to obtain a plurality of real-time trajectory anchor points output by the policy network, the method further includes: Based on a multi-layer perceptron, encode the real-time trajectory anchor points into anchor point feature vectors; Perform cross-attention operation on the anchor point feature vectors and the real-time information of the bird's-eye view state to obtain a trajectory anchor point offset; According to the trajectory anchor point offset and the real-time trajectory anchor points, construct optimized real-time trajectory anchor points; The construction of the real-time state-action pair according to the real-time information of the bird's-eye view state and the real-time trajectory anchor points includes: Construct the real-time state-action pair according to the real-time information of the bird's-eye view state and the optimized real-time trajectory anchor points.

[0010] An end-to-end autonomous driving trajectory evaluation method based on a world model provided by the present invention, wherein the bird's-eye view world model is trained through the following steps: Construct a training sample set according to the sample state-action pair, the bird's-eye view state sample information corresponding to the sample state-action pair at a future moment, and the driving behavior sample action; Based on the training sample set, train a preset world model, and when it is determined that the training loss value satisfies a preset loss value, obtain the bird's-eye view world model; Among them, the training loss function of the training loss value is specifically: ; ; ; ; ; Among them, represents the training loss value, represents the loss value based on the traffic simulator score, represents the loss value based on the bird's-eye view state information, represents the loss value based on the trajectory score, represents the loss value based on the trajectory optimization, represents the predicted t + k bird's-eye view state information after the represents the true bird's-eye view state information obtained based on the traffic simulator, represents the simulation score corresponding to the prediction result, represents the true score obtained based on the traffic simulator, represents the imitation score corresponding to the predicted trajectory, represents the score corresponding to the expert driving trajectory, represents the predicted trajectory, represents the expert driving trajectory.

[0011] An end-to-end autonomous driving trajectory evaluation method based on a world model provided by the present invention, wherein according to the real-time state-action pair and the predicted state-action pair, evaluate the predicted driving trajectory corresponding to the predicted state-action pair to obtain the evaluation result of the predicted driving trajectory, including: Input the real-time state-action pair and the predicted state-action pair into a score prediction network, and obtain the evaluation result of the predicted driving trajectory output by the score prediction network; Among them, the scoring prediction network includes a two-dimensional convolutional layer, a global average pooling layer, and a multi-layer perceptron head structure. The preset scoring conditions corresponding to the scoring prediction network at least include a non-responsible collision index, a drivable area compliance index, a collision time index, a driving comfort index, and a vehicle forward distance index.

[0012] According to an end-to-end autonomous driving trajectory evaluation method based on a world model provided by the present invention, after evaluating the predicted driving trajectory corresponding to the predicted state-action pair according to the real-time state-action pair and the predicted state-action pair to obtain the evaluation result of the predicted driving trajectory, the method further includes: Determine the predicted driving trajectory corresponding to the target evaluation result as the target driving path according to the evaluation score corresponding to the evaluation result of the predicted driving trajectory, where the target evaluation result is the evaluation result corresponding to the highest evaluation score.

[0013] The present invention also provides an end-to-end autonomous driving trajectory evaluation system based on a world model, including: An environmental data acquisition module, configured to generate real-time information of the bird's-eye view state according to multi-modal sensor data acquired in the current autonomous driving scenario; A trajectory anchor prediction module, configured to input the real-time information of the bird's-eye view state into a policy network to obtain a plurality of real-time trajectory anchors output by the policy network, where each of the real-time trajectory anchors corresponds to a different driving behavior action; the policy network is trained based on expert driving data and sample trajectory anchors corresponding to the expert driving data; A driving state prediction module, configured to construct a real-time state-action pair according to the real-time information of the bird's-eye view state and the real-time trajectory anchor, and input the real-time state-action pair into a bird's-eye view world model to obtain a predicted state-action pair output by the bird's-eye view world model, where the predicted state-action pair includes bird's-eye view state prediction information and driving behavior prediction actions; the bird's-eye view world model is trained based on sample state-action pairs and bird's-eye view state sample information and driving behavior sample actions corresponding to the sample state-action pairs at future moments; An evaluation module, configured to evaluate the predicted driving trajectory corresponding to the predicted state-action pair according to the real-time state-action pair and the predicted state-action pair to obtain the evaluation result of the predicted driving trajectory.

[0014] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the end-to-end autonomous driving trajectory evaluation method based on a world model as described in any one of the above.

[0015] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the end-to-end autonomous driving trajectory evaluation method based on the world model as described in any one of the above.

[0016] The end-to-end autonomous driving trajectory evaluation method and system based on the world model provided by the present invention generate real-time information of a bird's-eye view through multi-modal sensor data, use a policy network to output real-time trajectory anchor points, and construct real-time state-action pairs and input them into the bird's-eye view world model to predict future state-action pairs, thereby more accurately evaluating the predicted driving trajectory and improving the safety and reliability of the autonomous driving system. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 It is a schematic flowchart of the end-to-end autonomous driving trajectory evaluation method based on the world model provided by the present invention; Figure 2 It is a schematic diagram combining the policy network and the value network provided by the present invention; Figure 3 It is a schematic diagram of the overall architecture of the end-to-end autonomous driving trajectory evaluation method based on the world model provided by the present invention; Figure 4 It is a schematic diagram of the structure of the end-to-end autonomous driving trajectory evaluation system based on the world model provided by the present invention; Figure 5 It is a schematic diagram of the structure of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.

[0020] The world model, as a tool capable of simulating and predicting future states, provides new possibilities for the field of trajectory evaluation. However, when attempting to integrate it into the value network for more accurate trajectory evaluation, a series of technical problems are faced.

[0021] First, how to accurately and efficiently represent future scenarios has become a major challenge. Currently, diffusion models are mostly used in the world models of the driving field to predict the images of future scenarios. However, this method takes a long time in real-time driving environments and is difficult to meet the actual application requirements with strict requirements for response speed. Secondly, when supervising the scoring of future states and trajectories, the problem of insufficient density is encountered. In actual datasets, usually only one future state is known and available, while the world model needs to predict multiple possible future states based on multiple trajectory candidates, which undoubtedly increases the difficulty and complexity of supervision. In addition, even if the world model can successfully predict reasonable future states, due to the lack of corresponding scoring targets or criteria, the training of the value network will be greatly restricted and it is difficult to exert its due potential.

[0022] To address the above existing technical problems in the prior art, the present invention proposes an end-to-end autonomous driving trajectory evaluation method based on a world model. This method combines a policy network and a value network (i.e., a scoring prediction network), and uses a Bird's Eye View (BEV) world model to predict future states, so as to achieve efficient evaluation of driving trajectories in autonomous driving scenarios.

[0023] Figure 1 The flowchart of the end-to-end autonomous driving trajectory evaluation method based on the world model provided by the present invention is as Figure 1 shown. The present invention provides an end-to-end autonomous driving trajectory evaluation method based on a world model, including: Step 101: Generate real-time Bird's Eye View state information according to multi-modal sensor data obtained in the current autonomous driving scenario.

[0024] In an autonomous driving scenario, a vehicle is usually equipped with a variety of sensors, such as cameras, Light Detection and Ranging (LiDAR), millimeter-wave radars, and Inertial Navigation System (INS), etc. These sensors can capture the environmental information around the vehicle in real time. In the present invention, multi-modal sensor data is various types of data information collected based on different sensors.

[0025] In the present invention, these multi-modal sensor data can be preprocessed, including steps such as data alignment and data fusion. Data alignment is to ensure that the data collected by different sensors is consistent in time and space; data fusion is to combine the data of different sensors according to a certain algorithm to obtain more comprehensive and accurate environmental information. The processed multi-modal sensor data is used to generate real-time Bird's Eye View state information. The Bird's Eye View can intuitively reflect the positions and states of roads, vehicles, pedestrians, obstacles, etc. around the vehicle.

[0026] Step 102: Input the real-time information of the bird's-eye view into the policy network to obtain multiple real-time trajectory anchor points output by the policy network, where each of the real-time trajectory anchor points corresponds to a different driving behavior action; the policy network is trained based on expert driving data and the sample trajectory anchor points corresponding to the expert driving data.

[0027] In the present invention, the policy network is a trained neural network model, whose function is to output multiple real-time trajectory anchor points according to the input real-time information of the bird's-eye view state. These trajectory anchor points can be understood as key points on the possible driving paths of the vehicle in the future, and each anchor point corresponds to a specific driving behavior action, such as going straight, turning, accelerating, and decelerating.

[0028] In the present invention, the policy network is trained based on a large amount of expert driving data and the sample trajectory anchor points corresponding to these data. Among them, the expert driving data refers to the data generated by experienced drivers when driving vehicles in real or simulated environments, and these data contain the driving behavior decisions of drivers in various traffic conditions. The sample trajectory anchor points are extracted from these expert driving data, and they represent the key points on the driving paths selected by expert drivers in specific situations. During the training process, the policy network will learn the driving behavior patterns of expert drivers, that is, how to select appropriate driving behavior actions according to the current traffic conditions. Through continuous learning and optimization, the policy network gradually acquires the ability to output reasonable trajectory anchor points according to the real-time information of the bird's-eye view state.

[0029] Step 103: According to the real-time information of the bird's-eye view state and the real-time trajectory anchor points, construct a real-time state-action pair, and input the real-time state-action pair into the bird's-eye view world model to obtain a predicted state-action pair output by the bird's-eye view world model, where the predicted state-action pair includes bird's-eye view state prediction information and driving behavior prediction actions; the bird's-eye view world model is trained based on sample state-action pairs and the bird's-eye view state sample information and driving behavior sample actions corresponding to the sample state-action pairs at future moments.

[0030] In the present invention, for the decision-making and planning of the autonomous driving system, it simulates and predicts the possible states of the vehicle in the future and the corresponding driving behaviors. Specifically, according to the currently obtained real-time information of the bird's-eye view state and the real-time trajectory anchor points output by the policy network, a real-time state-action pair is constructed. Among them, the real-time information of the bird's-eye view state in the real-time state-action pair contains comprehensive perception data of the vehicle's surrounding environment, such as the positions and states of roads, vehicles, pedestrians, obstacles, etc. And the actions refer to the real-time trajectory anchor points, which represent the driving behavior actions that the vehicle may take in the future, such as going straight, turning, accelerating, and decelerating.

[0031] Next, input these real-time state-action pairs into the bird's-eye view world model. In the present invention, the role of the bird's-eye view world model is to predict the bird's-eye view state of the vehicle at a future moment and the corresponding driving behavior actions according to the input state-action pairs, that is, to output predicted state-action pairs. Among them, the predicted state in the predicted state-action pair is the bird's-eye view state prediction information, which describes the environmental state that the vehicle may be in in the future; and the predicted action is the driving behavior prediction action, which represents the driving behavior that the vehicle may take in the future.

[0032] In the present invention, the bird's-eye view world model is trained based on a large number of sample state-action pairs and the corresponding bird's-eye view state sample information and driving behavior sample actions of these samples at a future moment. Among them, the sample state-action pairs can be obtained by extracting from historical driving data, which contain the state and behavior relationships of the vehicle in various past situations, and the corresponding bird's-eye view state sample information and driving behavior sample actions of these samples at a future moment are the actual driving results of the vehicle in these scenarios.

[0033] During the training process, the bird's-eye view world model will learn the state and behavior relationships in these samples, that is, how to predict future states and behaviors based on the current state and actions. Through continuous learning and optimization, the bird's-eye view world model gradually acquires the ability to output reasonable predicted state-action pairs according to real-time state-action pairs.

[0034] Step 104, evaluate the predicted driving trajectory corresponding to the predicted state-action pair according to the real-time state-action pair and the predicted state-action pair, and obtain the evaluation result of the predicted driving trajectory.

[0035] In the present invention, the predicted state-action pairs are obtained through the bird's-eye view world model, which contain the predicted bird's-eye view state information that the vehicle may be in in the future and the corresponding driving behavior prediction actions. In order to verify the rationality and accuracy of these predictions, it is necessary to evaluate the predicted driving trajectory corresponding to the predicted state-action pair.

[0036] In the present invention, the predicted driving trajectory is generated based on the predicted state-action pair, which describes the path that the vehicle may travel in the future for a period of time. The purpose of evaluating the predicted driving trajectory is to determine whether this trajectory is reasonable, feasible, and meets the requirements of safe driving. During the evaluation process, various methods and indicators can be used. For example, the matching degree between the predicted driving trajectory and the real-time environmental information can be calculated to determine whether the trajectory conforms to the actual road, traffic signals, etc.; the smoothness and continuity of the trajectory can also be evaluated to ensure that the vehicle can maintain a stable driving state during driving. In addition, the safety of the trajectory can also be considered, such as avoiding collisions with obstacles and obeying traffic rules.

[0037] The end-to-end autonomous driving trajectory evaluation method based on the world model provided by the present invention generates real-time information of the bird's-eye view through multi-modal sensor data, uses the policy network to output real-time trajectory anchor points, constructs real-time state-action pairs and inputs them into the bird's-eye view world model to predict future state-actions, thereby more accurately evaluating the predicted driving trajectory and improving the safety and reliability of the autonomous driving system.

[0038] Based on the above embodiments, generating real-time information of the bird's-eye view state according to the multi-modal sensor data obtained in the current autonomous driving scenario includes: Obtain the multi-modal sensor data when the vehicle is driving in the current autonomous driving scenario, where the multi-modal sensor data at least includes multi-view image data and lidar data, and the multi-view image data includes vehicle front-view image data and vehicle left and right view image data; Perform image cropping preprocessing on the vehicle left and right view image data in the multi-modal sensor data to obtain preprocessed multi-modal sensor data; Based on the bird's-eye view encoder, convert the preprocessed multi-modal sensor data into the real-time information of the bird's-eye view state.

[0039] In the present invention, the multi-view image data includes image data of the vehicle front view and image data of the vehicle left and right views. Among them, the front view image provides visual information directly in front of the vehicle, while the left and right view images supplement the visual information on both sides of the vehicle, helping to more comprehensively understand the environment around the vehicle.

[0040] Lidar data is obtained by a lidar measuring the distance and position of surrounding objects by emitting and receiving laser pulses to generate point cloud data. Lidar data provides three-dimensional space information of the objects around the vehicle and can be used to judge the distance, shape and size of the objects.

[0041] In the present invention, since the images of the vehicle left and right views may contain some information irrelevant to driving decisions (such as the sky, distant scenery, etc.), it is necessary to perform cropping preprocessing on these images. The purpose of cropping is to retain the parts useful for driving decisions, such as information on the road, vehicles, pedestrians, etc. nearby, while removing irrelevant or interfering information. The cropped image data is more compact and effective, which can reduce the computational amount and complexity of subsequent processing.

[0042] Furthermore, based on the BEV encoder, the preprocessed multi-modal sensor data is converted into real-time bird's-eye view state information. The BEV encoder can convert multi-modal sensor data (including cropped image data and lidar data) into a unified bird's-eye view state representation. This representation shows the environmental information around the vehicle from a top-down perspective, making driving decisions more intuitive and easy to understand. Specifically, the cropped multi-view image data and lidar data are used as the input of the BEV encoder. Among them, the resolution of the multi-view image data is 256×1024 pixels, and the lidar data is a 64m×64m point cloud centered on the vehicle. Then, the BEV encoder processes these data and extracts useful feature information. Finally, these feature information are mapped into the bird's-eye view coordinate system to generate real-time bird's-eye view state information, where the real-time bird's-eye view state information is a three-dimensional tensor B ∈ R h×w×c , h and w represent the height and width of the real-time bird's-eye view state information respectively, c represents the number of channels.

[0043] Based on the above embodiments, the policy network is trained through the following steps: Based on the K-Means clustering algorithm, cluster the expert driving data to obtain the sample trajectory anchor points; According to the expert driving data and the corresponding sample trajectory anchor points of the expert driving data, train the neural network to obtain the policy network.

[0044] In the present invention, the expert driving data is data generated by experienced drivers when driving vehicles in real or simulated environments. These data contain the driving behavior decisions and trajectory information of drivers in various traffic conditions. The present invention uses the K-Means algorithm to cluster the trajectory information in the expert driving data and extracts N representative trajectory anchor points (such as N = 256) in the expert driving data. These trajectory anchor points can summarize the key points on the driving paths selected by expert drivers in specific situations and provide a basis for the subsequent training of the policy network.

[0045] In the present invention, the trajectory anchor points obtained through clustering processing are used as sample trajectory anchor points to train a policy network for simulating and predicting the driving decisions of experts. Further, based on the expert driving data and the corresponding sample trajectory anchor points, a neural network is trained to obtain a policy network. In the present invention, ResNet34 is used as the backbone network for bird's-eye view feature extraction. As a deep residual network, ResNet34 can process multi-modal sensor inputs and extract useful feature information. Through continuous learning and optimization, the neural network gradually acquires the ability to predict multiple trajectory anchor points based on multi-modal sensor inputs. The trained neural network, as a policy network, can output multiple predicted trajectory anchor points according to the current traffic conditions and vehicle surrounding environment information.

[0046] Based on the above embodiments, after inputting the real-time bird's-eye view state information into the policy network to obtain multiple real-time trajectory anchor points output by the policy network, the method further includes: Encoding the real-time trajectory anchor points into anchor feature vectors based on a multi-layer perceptron; Performing cross-attention operation on the anchor feature vectors and the real-time bird's-eye view state information to obtain trajectory anchor point offsets; Constructing optimized real-time trajectory anchor points according to the trajectory anchor point offsets and the real-time trajectory anchor points; The constructing of the real-time state-action pairs according to the real-time bird's-eye view state information and the real-time trajectory anchor points includes: Constructing the real-time state-action pairs according to the real-time bird's-eye view state information and the optimized real-time trajectory anchor points.

[0047] In the present invention, real-time trajectory anchor points are used to represent key points on the possible driving paths of vehicles. A multi-layer perceptron (MLP for short), as a trajectory encoder (TE for short), is used to encode real-time trajectory anchor points into anchor feature vectors. Through the processing of the MLP, real-time trajectory anchor points are converted into high-dimensional feature vectors, which contain key information of the trajectory anchor points, such as position, direction, etc., providing a basis for subsequent cross-attention operations.

[0048] Further, the process of performing cross-attention operation on the anchor feature vectors and the real-time bird's-eye view state information to obtain trajectory anchor point offsets can be specifically expressed as: ; Wherein, represents the optimized real-time trajectory anchor points, represents the real-time trajectory anchor points before optimization, It represents the real-time information of the bird's-eye view state. Through cross-attention operation, the present invention calculates the correlation between the anchor feature vector and the real-time information of the bird's-eye view state, so as to obtain the offset of the trajectory anchor point, which represents the direction and distance that the trajectory anchor point needs to be adjusted according to the current environmental state. Further, applying the trajectory anchor point offset to the real-time trajectory anchor point, the optimized real-time trajectory anchor point is obtained, which is more in line with the actual driving situation and can be used as the input for subsequent trajectory generation and path planning, providing accurate guidance for the vehicle's driving.

[0049] In the present invention, after the neural network corrects the offset of the anchor point, it finally outputs the trajectories of 40 time points (spanning 4 seconds, 10 points per second). Each trajectory point contains the x and y coordinates and the heading angle, and this information can be used to describe the driving path and direction of the vehicle in the future for a period of time.

[0050] Further, based on the current BEV state and the optimized trajectory , N real-time state-action pairs are constructed , where is the action embedding encoded by the trajectory encoder . Then, each real-time state-action pair is input into the BEV world model to predict the corresponding future state and future action .

[0051] In the present invention, the world model performs state prediction in the BEV latent space. By flattening the BEV state and concatenating it with the action embedding and then inputting it into a two-layer Transformer decoder, efficient future state prediction is achieved. Since the BEV world model supports autoregressive prediction, it can efficiently predict the future state within multiple time steps without introducing significant computational overhead. The specific process can be expressed as: ; And the prediction of the BEV world model for the next K steps can be expressed as: .

[0052] Based on the above embodiments, the bird's-eye view world model is trained through the following steps: Construct a training sample set according to the sample state-action pairs and the bird's-eye view state sample information and the driving behavior sample actions corresponding to the sample state-action pairs at future moments; Based on the training sample set, train a preset world model, and when it is determined that the training loss value meets the preset loss value, obtain the bird's-eye view world model; Among them, the training loss function of the training loss value is specifically: ; ; ; ; ; Among them, represents the training loss value, represents the loss value based on the traffic simulator score, represents the loss value based on the bird's-eye view state information, represents the loss value based on the trajectory score, represents the loss value based on the trajectory optimization, represents the predicted t + k bird's-eye view state information after the represents the true bird's-eye view state information obtained from the traffic simulator, represents the simulation score corresponding to the prediction result, represents the true score obtained from the traffic simulator, represents the imitation score corresponding to the predicted trajectory, represents the score corresponding to the expert driving trajectory, represents the predicted trajectory, represents the expert driving trajectory.

[0053] In the present invention, first, a series of sample state-action pairs are collected. These samples include the state of the autonomous vehicle at a specific moment (such as position, speed, surrounding obstacle information, etc.) and the driving actions taken based on these states (such as accelerating, braking, steering, etc.). For each sample state-action pair, the bird's-eye view state information (such as vehicle position, road structure, traffic signals, etc.) and possible driving behavior sample actions corresponding to these sample state-action pairs at a future moment (t + k) are obtained using a traffic simulator or historical data.

[0054] Furthermore, the above-collected sample state-action pairs, the bird's-eye view state information at the future moment, and the driving behavior actions are integrated into a training sample set, which will be used for subsequent world model training. To train this world model, the present invention defines a comprehensive loss function , including four parts: The loss value based on Focal Loss , which is used to monitor the difference between the predicted future bird's-eye view state information and the real bird's-eye view state information provided by the traffic simulator, ensuring that the world model can accurately predict future scenarios.

[0055] The loss value based on the cross-entropy loss function , which is used to monitor the difference between the predicted simulation score and the real score provided by the traffic simulator. This score can be calculated based on a series of rules (such as no collision, drivable area compliance, time to collision, safety distance, driving comfort, and progress, etc.).

[0056] The loss value based on the cross-entropy loss function , which is used to monitor the difference between the predicted trajectory imitation score and the score corresponding to the expert driving trajectory. In the present invention, the imitation score can be obtained by calculating the distance between the predicted trajectory anchor points and the expert trajectory and passing it through the Softmax function of the negative distance.

[0057] The loss value based on the L1 loss , which is used to monitor the difference between the predicted trajectories (including the unoptimized trajectories and the optimized trajectories) and the expert driving trajectory. In the present invention, the winner-takes-all strategy is adopted, that is, only the predicted trajectory closest to the expert trajectory is supervised.

[0058] In the present invention, a preset world model is trained using a training sample set, and the model parameters are continuously adjusted to minimize . When the training loss value meets the preset loss value condition (such as reaching a certain threshold or converging), the training process ends, and a trained bird's-eye view world model is obtained.

[0059] Based on the above embodiments, the evaluating the predicted driving trajectory corresponding to the predicted state-action pair according to the real state-action pair and the predicted state-action pair to obtain an evaluation result of the predicted driving trajectory includes: Inputting the real state-action pair and the predicted state-action pair into a score prediction network to obtain an evaluation result of the predicted driving trajectory output by the score prediction network; Among them, the score prediction network includes a two-dimensional convolutional layer, a global average pooling layer, and a multi-layer perceptron head structure, and the preset score conditions corresponding to the score prediction network at least include a no at-fault collision (NC) index, a drivable area compliance (DAC) index, a time-to-collision (TTC) index, a driving comfort (Comf) index, and an ego progress (EP) index.

[0060] In the present invention, the real-time state-action pair is the state of the autonomous vehicle at the current moment (such as position, speed, direction, etc.) and the driving actions taken based on these states (such as accelerating, braking, steering, etc.). The predicted state-action pair is the state and action at a future moment generated by the BEV world model based on the real-time state-action pair, describing the possible driving trajectory of the vehicle within a certain period in the future.

[0061] In the present invention, the scoring prediction network (ScoreNet), as a value network, mainly includes a two-dimensional convolutional layer, a global average pooling layer, and a multi-layer perceptron (MLP) head structure. Among them, the two-dimensional convolutional layer is used to process the input BEV (bird's-eye view) state information, which contains key information about the vehicle's surrounding environment (such as road structure, obstacle position, etc.). The two-dimensional convolutional layer can extract the features in these images, providing a basis for subsequent scoring prediction. The global average pooling layer is used to perform global average pooling operations on the feature map output by the convolutional layer to reduce the dimension of the feature map and retain key information, which helps to improve the generalization ability of the model and reduce the risk of overfitting. The multi-layer perceptron (MLP) head structure is used to receive the output of the global average pooling layer and the action embedding (a compact representation of the action information processed by the MLP), and generate the final trajectory score through a multi-layer fully connected network. In the present invention, the MLP head structure can comprehensively consider the BEV state information and action information of multiple time steps, thereby generating an accurate score.

[0062] In the present invention, when generating the trajectory score, the scoring prediction network will consider a series of preset scoring conditions. These conditions are designed to ensure that the generated trajectory is both safe and efficient. Specifically, they include: No-collision responsibility index (NC): Evaluates whether the trajectory will cause the vehicle to collide with other obstacles (such as other vehicles, pedestrians, or road facilities).

[0063] Drivable area compliance index (DAC): Evaluates whether the trajectory remains within the drivable area of the road.

[0064] Time to collision index (TTC): Measures the relative speed and distance between the vehicle and other obstacles in the trajectory to predict potential collision risks.

[0065] Driving comfort index (Comf): Evaluates the smoothness and stability of the trajectory to ensure the riding experience of passengers.

[0066] Vehicle forward distance index (EP): Measures the total distance the vehicle advances in the trajectory.

[0067] Finally, the scoring prediction network takes the real-time state-action pairs and the predicted state-action pairs as inputs and outputs the comprehensive scores of each predicted driving trajectory. , and the specific process can be expressed as: ; ; .

[0068] These scores comprehensively consider the above-mentioned preset scoring conditions and reflect the quality and safety of the trajectories. By weighted-averaging multiple scoring weights (which may be obtained based on experience or training data), the trajectory with the highest score can ultimately be selected as the output. Figure 2 For the schematic diagram of the combination of the policy network and the value network provided by the present invention, reference can be made to Figure 2 As shown, the present invention combines a policy network and a value network. Among them, the value network, by introducing the future BEV state predicted by the BEV world model, can comprehensively evaluate the trajectory based on the current state and the predicted future state, compared with the policy network of the existing end-to-end autonomous driving method that mainly focuses on trajectory prediction.

[0069] Figure 3 For the overall architecture schematic diagram of the end-to-end autonomous driving trajectory evaluation method based on the world model provided by the present invention, reference can be made to Figure 3 As shown, in the present invention, the input of the BEV world model includes multi-view images and LiDAR data, and the output is the end-to-end predicted vehicle trajectory. Among them, the architecture is divided into two parts: Figure 3 (a) in it is the policy network, which is used to encode multi-modal perception information into BEV features and generate multiple trajectory candidates; Figure 3 (b) in it is the value network, which scores the trajectory through the future BEV state of each trajectory predicted by the BEV world model, and selects the best trajectory as the final output.

[0070] Based on the above embodiments, after evaluating the predicted driving trajectory corresponding to the predicted state-action pair according to the real-time state-action pair and the predicted state-action pair to obtain the evaluation result of the predicted driving trajectory, the method further includes: Determining the predicted driving trajectory corresponding to the target evaluation result as the target driving path according to the evaluation score corresponding to the evaluation result of the predicted driving trajectory, where the target evaluation result is the evaluation result corresponding to the highest evaluation score.

[0071] In the present invention, among all the predicted driving trajectories, the trajectory with the highest evaluation score (or comprehensive score) is identified, and the predicted driving trajectory with the highest evaluation score is determined as the target driving path. Once the target driving path is determined, the autonomous driving system plans specific driving actions (such as accelerating, braking, steering, etc.) according to this path, and executes these actions through the vehicle control system to achieve safe and efficient autonomous driving.

[0072] The present invention realizes the efficiency and real-time performance of trajectory evaluation by predicting future states in the BEV latent space; at the same time, combined with the dual supervision of the traffic simulator and expert data, it ensures the efficient training and generalization ability of the value network (i.e., the score prediction network), and improves the safety and reliability of the autonomous driving system.

[0073] In summary, by comprehensively evaluating the evaluation scores of multiple predicted driving trajectories and selecting the trajectory corresponding to the highest score as the target driving path, the autonomous driving system can ensure making the optimal driving decision in a complex and changeable driving environment.

[0074] The following describes the end-to-end autonomous driving trajectory evaluation system based on the world model provided by the present invention. The end-to-end autonomous driving trajectory evaluation system based on the world model described below can be correspondingly referred to the end-to-end autonomous driving trajectory evaluation method based on the world model described above.

[0075] Figure 4 is the structural schematic diagram of the end-to-end autonomous driving trajectory evaluation system based on the world model provided by the present invention, as Figure 4As shown in the figure, the present invention provides an end-to-end autonomous driving trajectory evaluation system based on a world model, including: An environmental data acquisition module 401 is used to generate real-time information on the state of a bird's-eye view according to multi-modal sensor data obtained in the current autonomous driving scenario; A trajectory anchor prediction module 402 is used to input the real-time information on the state of the bird's-eye view into a policy network to obtain a plurality of real-time trajectory anchors output by the policy network, where each of the real-time trajectory anchors corresponds to a different driving behavior action; The policy network is trained based on expert driving data and sample trajectory anchors corresponding to the expert driving data; A driving state prediction module 403 is used to construct a real-time state-action pair according to the real-time information on the state of the bird's-eye view and the real-time trajectory anchors, and input the real-time state-action pair into the bird's-eye view world model to obtain a predicted state-action pair output by the bird's-eye view world model, where the predicted state-action pair includes bird's-eye view state prediction information and predicted driving behavior actions; The bird's-eye view world model is trained based on sample state-action pairs and bird's-eye view state sample information and driving behavior sample actions corresponding to the sample state-action pairs at future moments; An evaluation module 404 is used to evaluate the predicted driving trajectory corresponding to the predicted state-action pair according to the real-time state-action pair and the predicted state-action pair to obtain an evaluation result of the predicted driving trajectory.

[0076] The end-to-end autonomous driving trajectory evaluation method provided by the present invention generates real-time information on a bird's-eye view through multi-modal sensor data, uses a policy network to output real-time trajectory anchors, constructs a real-time state-action pair and inputs it into the bird's-eye view world model to predict future state-actions, thereby more accurately evaluating the predicted driving trajectory and improving the safety and reliability of the autonomous driving system.

[0077] The system provided by the embodiments of the present invention is used to execute the above method embodiments. For the specific process and detailed content, please refer to the above embodiments and will not be elaborated here.

[0078] Figure 5 It is a schematic structural diagram of the electronic device provided by the present invention, as Figure 5As shown, the electronic device may include: a processor 501, a communications interface 502, a memory 503, and a communication bus 504. Among them, the processor 501, the communications interface 502, and the memory 503 complete mutual communication through the communication bus 504. The processor 501 may call the logical instructions in the memory 503 to execute an end-to-end autonomous driving trajectory evaluation method based on a world model. The method includes: generating real-time bird's-eye view state information according to multi-modal sensor data obtained in the current autonomous driving scenario; inputting the real-time bird's-eye view state information into a policy network to obtain a plurality of real-time trajectory anchors output by the policy network, where each of the real-time trajectory anchors corresponds to a different driving behavior action; the policy network is trained based on expert driving data and sample trajectory anchors corresponding to the expert driving data; constructing a real-time state-action pair according to the real-time bird's-eye view state information and the real-time trajectory anchors, and inputting the real-time state-action pair into a bird's-eye view world model to obtain a predicted state-action pair output by the bird's-eye view world model, where the predicted state-action pair includes bird's-eye view state prediction information and predicted driving behavior actions; the bird's-eye view world model is trained based on sample state-action pairs and bird's-eye view state sample information and driving behavior sample actions corresponding to the sample state-action pairs at future moments; evaluating the predicted driving trajectory corresponding to the predicted state-action pair according to the real-time state-action pair and the predicted state-action pair to obtain an evaluation result of the predicted driving trajectory.

[0079] In addition, when the logical instructions in the above-mentioned memory 503 can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0080] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the end-to-end autonomous driving trajectory evaluation method based on the world model provided by each of the above methods. The method includes: generating real-time bird's-eye view state information according to multi-modal sensor data obtained in the current autonomous driving scenario; inputting the real-time bird's-eye view state information into a policy network to obtain a plurality of real-time trajectory anchor points output by the policy network, where each of the real-time trajectory anchor points corresponds to a different driving behavior action; the policy network is trained based on expert driving data and sample trajectory anchor points corresponding to the expert driving data; constructing a real-time state-action pair according to the real-time bird's-eye view state information and the real-time trajectory anchor points, and inputting the real-time state-action pair into a bird's-eye view world model to obtain a predicted state-action pair output by the bird's-eye view world model, where the predicted state-action pair includes bird's-eye view state prediction information and predicted driving behavior actions; the bird's-eye view world model is trained based on sample state-action pairs and bird's-eye view state sample information and driving behavior sample actions corresponding to the sample state-action pairs at future times; evaluating the predicted driving trajectory corresponding to the predicted state-action pair according to the real-time state-action pair and the predicted state-action pair to obtain an evaluation result of the predicted driving trajectory.

[0081] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the end-to-end autonomous driving trajectory evaluation method based on a world model provided in the above embodiments. The method includes: generating real-time bird's-eye view state information according to multi-modal sensor data obtained in the current autonomous driving scenario; inputting the real-time bird's-eye view state information into a policy network to obtain a plurality of real-time trajectory anchors output by the policy network, where each of the real-time trajectory anchors corresponds to a different driving behavior action; the policy network is trained based on expert driving data and sample trajectory anchors corresponding to the expert driving data; constructing a real-time state-action pair according to the real-time bird's-eye view state information and the real-time trajectory anchors, and inputting the real-time state-action pair into a bird's-eye view world model to obtain a predicted state-action pair output by the bird's-eye view world model, where the predicted state-action pair includes bird's-eye view state prediction information and predicted driving behavior actions; the bird's-eye view world model is trained based on sample state-action pairs and bird's-eye view state sample information and sample driving behavior actions corresponding to the sample state-action pairs at future times; evaluating the predicted driving trajectory corresponding to the predicted state-action pair according to the real-time state-action pair and the predicted state-action pair to obtain an evaluation result of the predicted driving trajectory.

[0082] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative effort.

[0083] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0084] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An end-to-end autonomous driving trajectory evaluation method based on a world model, characterized in that, Including: Generating real-time information on the bird's-eye view state based on multi-modal sensor data obtained in the current autonomous driving scenario; Inputting the real-time information on the bird's-eye view state into a policy network to obtain multiple real-time trajectory anchors output by the policy network, where each of the real-time trajectory anchors corresponds to a different driving behavior action; the policy network is trained based on expert driving data and sample trajectory anchors corresponding to the expert driving data; Constructing a real-time state-action pair based on the real-time information on the bird's-eye view state and the real-time trajectory anchors, and inputting the real-time state-action pair into a bird's-eye view world model to obtain a predicted state-action pair output by the bird's-eye view world model, where the predicted state-action pair includes predicted information on the bird's-eye view state and a predicted driving behavior action; the bird's-eye view world model is trained based on sample state-action pairs and bird's-eye view state sample information and driving behavior sample actions corresponding to the sample state-action pairs at future times; Evaluating the predicted driving trajectory corresponding to the predicted state-action pair based on the real-time state-action pair and the predicted state-action pair to obtain an evaluation result of the predicted driving trajectory.

2. The end-to-end autonomous driving trajectory evaluation method based on a world model according to claim 1, wherein The generating real-time information on the bird's-eye view state based on multi-modal sensor data obtained in the current autonomous driving scenario includes: Obtaining multi-modal sensor data during vehicle driving in the current autonomous driving scenario, where the multi-modal sensor data at least includes multi-view image data and lidar data, and the multi-view image data includes vehicle front-view image data and vehicle left and right view image data; Performing image cropping preprocessing on the vehicle left and right view image data in the multi-modal sensor data to obtain preprocessed multi-modal sensor data; Based on a bird's-eye view encoder, converting the preprocessed multi-modal sensor data into the real-time information on the bird's-eye view state.

3. The end-to-end autonomous driving trajectory evaluation method based on a world model according to claim 1, wherein The policy network is trained through the following steps: Performing clustering processing on the expert driving data based on the K-Means clustering algorithm to obtain the sample trajectory anchors; Training a neural network based on the expert driving data and the sample trajectory anchors corresponding to the expert driving data to obtain the policy network.

4. The end-to-end autonomous driving trajectory evaluation method based on a world model according to claim 1 or 3, characterized in that After inputting the real-time information on the bird's-eye view state into the policy network to obtain multiple real-time trajectory anchors output by the policy network, the method further includes: Encoding the real-time trajectory anchors into anchor feature vectors based on a multi-layer perceptron; Performing cross-attention operation on the anchor feature vectors and the real-time information on the bird's-eye view state to obtain a trajectory anchor offset; Constructing optimized real-time trajectory anchors based on the trajectory anchor offset and the real-time trajectory anchors; The constructing a real-time state-action pair based on the real-time information on the bird's-eye view state and the real-time trajectory anchors includes: Constructing the real-time state-action pair based on the real-time information on the bird's-eye view state and the optimized real-time trajectory anchors.

5. The end-to-end autonomous driving trajectory evaluation method based on a world model according to claim 1, wherein The bird's-eye view world model is trained through the following steps: Construct a training sample set based on the sample state-action pairs and the corresponding bird's-eye view state samples and driving behavior sample actions of the sample state-action pairs at future moments; Based on the training sample set, train a preset world model, and when it is determined that the training loss value meets the preset loss value, obtain the bird's-eye view world model; Among them, the training loss function of the training loss value is specifically: ; ; ; ; ; Among them, represents the training loss value, represents the loss value based on the traffic simulator score, represents the loss value based on the bird's-eye view state information, represents the loss value based on the trajectory score, represents the loss value based on the trajectory optimization, represents the predicted t + k bird's-eye view state information after the represents the true bird's-eye view state information obtained from the traffic simulator, represents the simulation score corresponding to the prediction result, represents the true score obtained from the traffic simulator, represents the imitation score corresponding to the predicted trajectory, represents the score corresponding to the expert driving trajectory, represents the predicted trajectory, represents the expert driving trajectory.

6. The end-to-end autonomous driving trajectory evaluation method based on a world model according to claim 1, wherein The evaluation of the predicted driving trajectory corresponding to the predicted state-action pair according to the real-time state-action pair and the predicted state-action pair, and obtaining the evaluation result of the predicted driving trajectory includes: Input the real-time state-action pair and the predicted state-action pair into a scoring prediction network, and obtain the evaluation result of the predicted driving trajectory output by the scoring prediction network; Among them, the scoring prediction network includes a two-dimensional convolutional layer, a global average pooling layer, and a multi-layer perceptron head structure. The preset scoring conditions corresponding to the scoring prediction network at least include a non-responsible collision index, a drivable area compliance index, a collision time index, a driving comfort index, and a vehicle forward distance index.

7. The end-to-end autonomous driving trajectory evaluation method based on a world model according to claim 1, wherein After the evaluation of the predicted driving trajectory corresponding to the predicted state-action pair according to the real-time state-action pair and the predicted state-action pair, and obtaining the evaluation result of the predicted driving trajectory, the method further includes: Determine the predicted driving trajectory corresponding to the target evaluation result as the target driving path according to the evaluation score corresponding to the evaluation result of the predicted driving trajectory, where the target evaluation result is the evaluation result corresponding to the highest evaluation score.

8. An end-to-end autonomous driving trajectory evaluation system based on a world model, characterized in that, Including: An environmental data acquisition module, configured to generate real-time bird's-eye view state information according to multi-modal sensor data acquired in the current autonomous driving scenario; A trajectory anchor prediction module, configured to input the real-time bird's-eye view state information into a policy network to obtain a plurality of real-time trajectory anchors output by the policy network, where each of the real-time trajectory anchors corresponds to a different driving behavior action; the policy network is trained based on expert driving data and sample trajectory anchors corresponding to the expert driving data; A driving state prediction module, configured to construct a real-time state-action pair according to the real-time bird's-eye view state information and the real-time trajectory anchors, and input the real-time state-action pair into the bird's-eye view world model to obtain a predicted state-action pair output by the bird's-eye view world model, where the predicted state-action pair includes bird's-eye view state prediction information and driving behavior prediction actions; the bird's-eye view world model is trained based on sample state-action pairs and corresponding bird's-eye view state sample information and driving behavior sample actions of the sample state-action pairs at future moments; An evaluation module, configured to evaluate the predicted driving trajectory corresponding to the predicted state-action pair according to the real-time state-action pair and the predicted state-action pair, and obtain the evaluation result of the predicted driving trajectory.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, wherein, When the processor executes the computer program, it implements the end-to-end autonomous driving trajectory evaluation method based on the world model according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method for end-to-end autonomous driving trajectory evaluation based on a world model according to any one of claims 1 to 7.

Citation Information

Cited By

  • Vehicle driving track prediction method and system and electronic equipment

    CN121106349A

  • End-to-end automatic driving system based on world model and sampling evaluation decision

    CN121143161A

  • An end-to-end autonomous driving system based on world model and sampling evaluation decision-making

    CN121143161B

  • World model-driven automatic driving reinforcement learning double-strategy risk perception control method and system and storage medium

    CN121209264A

  • Driving decision model optimization method and electronic equipment

    CN121209296A