System and method for enhancing end-to-end planning and evaluation in closed loop environment for autonomous driving

By using a Visual Language Planning (VLP) machine learning model, combined with ALP and SLP modules, the difficulty of integrating language understanding and visual planning in autonomous driving systems was solved, enabling more accurate decision-making and closed-loop evaluation, and improving the performance of autonomous driving systems in complex environments.

CN121734430APending Publication Date: 2026-03-27ROBERT BOSCH GMBH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing autonomous driving systems face difficulties in integrating language understanding and vision-based planning, leading to inaccurate decision-making in complex environments. Furthermore, traditional benchmarking methods are mostly open-loop evaluations, which cannot effectively assess the performance of autonomous driving in the real world.

Method used

The Visual Language Planning (VLP) machine learning model is adopted to generate a bird's-eye view (BEV) by receiving images, extracting the visual and textual features of the agent, and using a contrastive learning model to enhance the BEV and planning model to achieve closed-loop evaluation and dynamic response. Combined with ALP and SLP modules, the system's understanding and decision-making capabilities are improved.

Benefits of technology

It improves the decision-making accuracy and safety of autonomous driving systems in complex environments and provides a closed-loop evaluation framework that can effectively evaluate the performance of autonomous driving systems in the real world.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121734430A_ABST
    Figure CN121734430A_ABST
Patent Text Reader

Abstract

Methods and systems for training an end-to-end autonomous driving system in a closed-loop environment using a Visual Language Planning (VLP) machine learning model. An image associated with a vehicle surroundings is generated, and a BEV model is executed to generate a BEV view based on the image. The planning model predicts a navigation trajectory based on the BEV. The VLP model enhances the system by extracting vision-based planning features, generating text cues, and employing a language encoder to create text-based expected features. And comparing a learning model to identify the similarity between visual and text features, and enhancing the performance of the BEV and the planning model. The system experiences closed-loop evaluation in a simulated environment, obtaining metrics to improve the autonomous driving system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to a visual language planning (VLP) foundational model for autonomous driving and a system and method for evaluating it in a closed-loop environment. Background Technology

[0002] Autonomous vehicles, often referred to as self-driving or driverless vehicles, are a type of vehicle capable of navigating and operating on roads and in various environments without direct human control. Autonomous vehicles use a combination of advanced technologies and sensors to perceive their surroundings, make decisions, and perform driving tasks.

[0003] Autonomous vehicles are typically equipped with a variety of sensors, including lidar, radar, cameras, ultrasonic sensors, and sometimes additional technologies such as GPS and IMU (inertial measurement unit). These sensors provide real-time data about the vehicle's surroundings, including the positions of other vehicles, pedestrians, road signs, and road conditions. The vehicle's onboard computer uses the data from the sensors to create detailed environmental maps and perceive objects and obstacles. This information is crucial for navigation and collision avoidance.

[0004] Machine learning (ML) and artificial intelligence (AI) play a crucial role in autonomous vehicles. Deep learning algorithms are used for tasks such as object detection, lane keeping, and decision making, and can rely on image processing to perform these tasks. These algorithms enable vehicles to understand and respond to complex and dynamic traffic conditions. Summary of the Invention

[0005] According to one aspect of the present invention, a method for training an end-to-end autonomous driving system in a closed-loop environment using a visual language planning (VLP) machine learning model includes: receiving images generated from multiple image sensors mounted on a vehicle; performing a BEV machine learning model based on the images to generate a bird's-eye view (BEV) of the environment; performing a planning machine learning model on the BEV to generate a predicted trajectory for navigating the autonomous vehicle in the environment; and performing a VLP machine learning model to improve the end-to-end autonomous driving system, including: extracting vision-based planning features associated with detected agents within the environment, wherein the vision-based planning features include spatiotemporal information associated with the detected agents in the images. The system generates text prompts based on the extracted spatiotemporal information associated with the detected agents in the image; delivers the text prompts through a language encoder to generate text-based expected features associated with the detected agents; performs a contrastive learning model to derive similarity between vision-based planning features and text-based expected features; enhances the BEV model and planning model based on the similarity; performs a closed-loop evaluation of the end-to-end autonomous driving system using the enhanced BEV model and enhanced planning model by interacting with the simulated environment in real time and dynamically responding to actions taken by the vehicle based on predicted trajectories; acquires evaluation metrics during the closed-loop evaluation; and modifies the end-to-end autonomous driving system based on the acquired evaluation metrics.

[0006] According to another aspect, an end-to-end autonomous driving system utilizing a Visual Language Planning (VLP) machine learning model in a closed-loop environment includes a processor and a memory, the memory including instructions that, when executed by the processor, cause the processor to receive an image of the environment surrounding the vehicle; execute a BEV machine learning model based on the image to generate a bird's-eye view (BEV) of the environment; execute a planning machine learning model on the BEV to generate a predicted trajectory for navigating the autonomous vehicle in the environment; and execute the VLP machine learning model to: extract vision-based planning features associated with detected agents within the environment, wherein the vision-based planning features include those related to the image... The system extracts spatiotemporal information associated with the detected agent; generates text prompts based on the extracted spatiotemporal information associated with the detected agent in the image; delivers the text prompts through a language encoder to generate text-based expected features associated with the detected agent; and improves the end-to-end autonomous driving system by performing a contrastive learning model to derive similarities between vision-based planning features and text-based expected features; enhances the BEV model and planning model based on the similarities; and performs closed-loop evaluation of the end-to-end autonomous driving system using the enhanced BEV model and enhanced planning model by interacting with the simulated environment in real time and dynamically responding to actions taken by the vehicle based on predicted trajectories. Attached Figure Description

[0007] Figure 1 A system for training a neural network according to one embodiment is shown.

[0008] Figure 2 A computer-implemented method for training and utilizing a neural network is shown according to one embodiment.

[0009] Figure 3 A schematic diagram of a control system configured for controlling a vehicle, which may be a partially autonomous vehicle, a fully autonomous vehicle, a partially autonomous robot, or a fully autonomous robot, is shown according to one embodiment.

[0010] Figure 4 A schematic overview of an end-to-end autonomous driving system according to one embodiment is shown.

[0011] Figure 5 A schematic overview of a contrastive learning model according to one embodiment is shown.

[0012] Figure 6 A schematic diagram illustrates a method for training an autonomous driving system using a Visual Language Planning (VLP) machine learning model according to one embodiment. Detailed Implementation

[0013] This document describes embodiments of the present disclosure. However, it should be understood that the disclosed embodiments are merely examples, and other embodiments may take different and alternative forms. The drawings are not necessarily drawn to scale; some features may be enlarged or reduced to show details of particular components. Therefore, the specific structural and functional details disclosed herein should not be construed as limiting, but merely as a representative basis for teaching those skilled in the art to adopt the embodiments in various ways. As will be understood by those skilled in the art, the various features illustrated and described with reference to any of the drawings may be combined with features illustrated in one or more other drawings to produce embodiments not explicitly illustrated or described. The combinations of features shown provide representative embodiments for typical applications. However, for a particular application or implementation, various combinations and modifications of features consistent with the teachings of this disclosure may be required.

[0014] As used in this article, “a,” “one,” and “the” refer to both the singular and plural forms, unless the context clearly indicates otherwise. For example, a “processor” programmed to perform various functions means one processor programmed to perform each of those functions, or more than one processor programmed together to perform each of those functions.

[0015] In the context of autonomous vehicles, the term "agent" can refer to an object or entity in the environment surrounding or interacting with an autonomous vehicle. This includes pedestrians, other vehicles, cyclists, road signs, traffic lights, lane markings, and so on. An "agent" can include objects or features detected by the autonomous vehicle's sensors for decision-making in the autonomous vehicle's control.

[0016] In the context of computing devices, the term "model" or "module" refers to a machine learning model (e.g., a neural network), which is a trainable computer program that uses algorithms to learn patterns in data and make predictions or decisions about new data. The various "models" described herein can be executed via one or more processors that execute instructions stored in memory, depending on the specific needs of the model. These one or more processors may collectively be part of a graphics processing unit (GPU), a central processing unit (CPU), an application-specific integrated circuit (ASIC), or the like.

[0017] This disclosure incorporates, by reference in its entirety, U.S. Patent Application No. 18 / 388,601, filed November 10, 2023, entitled “VISION-LANGUAGE-PLANNING (VLP) MODELS WITH AGENT-WISE LEARNING FOR AUTONOMOUSDRIVING”. This disclosure also incorporates, by reference in its entirety, U.S. Patent Application No. 18 / 388,606, filed November 10, 2023, entitled “SYSTEMS AND METHODS FOR VISION-LANGAUGE PLANNING (VLP) FOUNDATION MODELS FOR AUTONOMOUSDRIVING”.

[0018] The rapid development of autonomous driving technology has ushered in a new era of transportation, promising safer and more efficient journeys. Autonomous driving systems generally include three advanced tasks: (1) perception, (2) prediction, and (3) planning. Each of these tasks can be performed on its own machine learning model. Perception involves the vehicle's ability to understand and interpret its environment. This task includes various sub-components such as computer vision, sensor fusion, and localization. Key elements of perception include object detection (e.g., identifying and tracking agents outside the autonomous vehicle), localization (e.g., determining the vehicle's precise location and orientation in the world, typically using GPS and other sensors), and sensor fusion (e.g., combining data from different sensors such as cameras, lidar, radar, and ultrasonic sensors to construct a comprehensive view of the surrounding environment). Prediction involves anticipating how other road users and agents in the environment will behave in the near future. This task typically involves using machine learning models to estimate the trajectories and intentions of agents, including pedestrians, other vehicles, and potential obstacles. Accurate prediction is crucial for making safe driving decisions. Planning involves determining the optimal path and actions for the autonomous vehicle to navigate its environment. This typically includes tasks such as route planning, trajectory planning, and decision making. Planning systems consider information from perception and prediction to make decisions such as when to change lanes, when to stop at an intersection, and how to react to unexpected events.

[0019] Traditional approaches to autonomous driving use independent models where each task (perception, prediction, and planning) is trained and optimized separately. However, this disjoint training and optimization can lead to significant error accumulation. To address this issue, end-to-end autonomous driving systems have been proposed and have gained attention in recent years. End-to-end autonomous driving systems unify all these tasks and perform joint optimization to facilitate and improve planning. In particular, the end-to-end approach fully utilizes bird's-eye view (BEV) representations in all tasks of the perception, prediction, and planning models. The BEV is generated from multi-view camera input and contains spatiotemporal information about the scene. Computer vision systems (e.g., Figure 2 The camera, processor, memory, and machine learning model shown can derive spatiotemporal information about the scene. Joint training and optimization strategies for all tasks across end-to-end autonomous driving have led to state-of-the-art results for autonomous driving. Figure 4(Described in more detail below) An overview of an end-to-end autonomous driving system according to one embodiment of this disclosure is shown. The integration of various modalities that collaboratively enhance perception, decision-making, and planning capabilities is crucial for the success of autonomous vehicles. In the Unified Autonomous Driving (UniAD) model, one of the most recent end-to-end autonomous driving efforts, autonomous driving is constructed by redefining the interactions between key components and tasks. In UniAD, the focus shifts to the optimization pursuit of the ultimate goal: effectively planning for autonomous vehicles. This requires re-examining the core components of perception and prediction and reconfiguring their roles to align with overall planning objectives.

[0020] Despite significant progress in computer vision for autonomous driving, a crucial dimension remains unexplored: the integration of language understanding with vision-based planning systems. Base models, which serve as foundational models for large pre-trained machine learning models trained on open-world data, often involve language as one of the dominant modalities of the data. In base models, there are often connections between language and other modalities. After pre-training, base models can be fine-tuned to adapt to a given task. Base models have demonstrated the importance of incorporating language to achieve state-of-the-art performance and generalization across a wide range of tasks. Despite the tremendous success of base models across different domains, their extension into the field of autonomous driving remains largely unexplored.

[0021] Furthermore, despite advancements in vision-based autonomous driving systems, these approaches often struggle with reasoning, generalization, and handling long-tail scenarios, limiting their deployment in real-world environments. Emerging advances in multimodal large language models (MLLMs) have demonstrated that their common-sense and reasoning capabilities can help address challenges in the embodied AI domain. While most of these approaches are primarily geared towards robotics, existing work utilizing embodied language models (LMs) for autonomous driving tasks is limited. Notably, DiLu and GPT-Driver introduce a GPT-based driver agent for closed-loop simulation tasks. Other systems utilize open-loop driving commentators that combine vision and underlying driving actions with language to interpret and reason about driving behavior. However, it remains unclear how to efficiently refine and fully leverage these approaches to enhance the performance of modular end-to-end autonomous driving tasks.

[0022] To address these challenges, this disclosure proposes a novel Visual Language Planning (VLP) framework that efficiently extracts the power of visual language models for autonomous driving through contrastive learning objectives. Figure 3The VLP framework, illustrated in the diagram and further described below, introduces two components: Agent-Centered Learning Paradigm (ALP) and Self-Driving Vehicle-Centered Learning Paradigm (SLP). ALP enhances the local semantic representation and reasoning capabilities of the BEV feature map by aligning it, which acts as a source memory within the driving system, with human-like reasoning processes. SLP improves the planning process by aligning planning queries with the goals and states of the self-driving vehicle through common-sense reasoning embedded in a language model to guide decision-making. These components collectively enhance the system's ability to understand complex driving environments and make safer, more information-supported decisions.

[0023] Furthermore, performance evaluation is a crucial step in the development of models for autonomous driving. Performance evaluation in a standardized manner using commonly available publicly available datasets and comparisons across publicly reported results from the community are known as “benchmarking.” As in most areas of AI, appropriate benchmarking is important in building confidence in the capabilities of trained models and in expanding the frontiers of innovation. For autonomous driving systems, several limitations currently exist in the available benchmarking methods. Since most benchmarks use open-loop evaluation, which does not assess the various capabilities required for autonomous driving on public roads, such as responding to unseen actions by other agents, this disclosure provides a closed-loop evaluation framework and scenario provided by the Bench2Drive tool—a benchmark designed for evaluating end-to-end autonomous driving systems in a closed-loop environment—according to various embodiments. This disclosure provides details of a novel benchmarking process and results in both open-loop and closed-loop environments.

[0024] Machine learning and neural networks are components of the invention disclosed herein. Figure 1 A system 100 for training a neural network, such as a deep neural network, is shown. System 100 may include an input interface for accessing training data 102 of the neural network. For example, as... Figure 1 As shown, the input interface can be composed of a data storage interface 104, which can access the training data 102 from the data storage device 106. For example, the data storage interface 104 can be a memory interface or a permanent storage interface, such as a hard disk or SSD interface, but it can also be a personal area network (PAN) interface, a local area network (LAN) interface, or a wide area network (WAN) interface, such as a Bluetooth, Zigbee, or Wi-Fi interface, or an Ethernet or fiber optic interface. The data storage device 106 can be an internal data storage device of the system 100, such as a hard disk drive or SSD, but it can also be an external data storage device, such as a network-accessible data storage device.

[0025] In some embodiments, data storage device 106 may further include a data representation 108 of an untrained version of the neural network, which can be accessed by system 100 from data storage device 106. However, it should be understood that the training data 102 and data representation 108 of the untrained neural network may also be accessed from different data storage devices, for example, via different subsystems of data storage interface 104. Each subsystem may be of the type described above for data storage interface 104. In other embodiments, the data representation 108 of the untrained neural network may be generated internally by system 100 based on design parameters for the neural network, and therefore may not be explicitly stored on data storage device 106.

[0026] System 100 may further include a processor subsystem 110 configured to provide an iterative function as a replacement for a stack of layers in a neural network to be trained during operation of system 100. Here, the corresponding layers of the replaced stack may have mutually shared weights and may receive the output of the previous layer as input, or, for the first layer of the stack, receive the initial activation and a portion of the input of the stack as input. Processor subsystem 110 may be further configured to iteratively train the neural network using training data 102. Here, the iteration of training via processor subsystem 110 may include a forward propagation portion and a backward propagation portion. Processor subsystem 110 may be configured to implement the forward propagation portion by: determining an equilibrium point of the iterative function, where the iterative function converges to a fixed point, in addition to defining other operations that define the implementable forward propagation portion, wherein determining the equilibrium point includes using a numerical root-finding algorithm to find a root solution for the iterative function minus its input; and providing the equilibrium point as a replacement for the output of the stack of layers in the neural network. System 100 may further include an output interface for outputting a data representation 112 of the trained neural network; this data may also be referred to as trained model data 112. For example, as also Figure 1 As illustrated in the diagram, the output interface can be comprised of a data storage interface 104, which in these embodiments is an input / output (“IO”) interface through which trained model data 112 can be stored in data storage device 106. For example, the data representation 108 defining an “untrained” neural network can be at least partially replaced by the data representation 112 of the trained neural network during or after training, because the parameters of the neural network, such as the network weights, hyperparameters, and other types of parameters, can be adapted to reflect training on the training data 102. This also... Figure 1The reference numerals 108 and 112 are shown in the diagram, referring to the same data record on data storage device 106. In other embodiments, data representation 112 may be stored separately from data representation 108 defining an "untrained" neural network. In some embodiments, the output interface may be separate from data storage interface 104, but it can generally be of the type described above for data storage interface 104.

[0027] Figure 1 The system 100 shown is an example of a system that can be used to train the machine learning model described in this paper.

[0028] Figure 2 A system 200 is depicted that implements the machine learning models described herein, such as the VLP underlying model. System 200 may include at least one computing system 202. Computing system 202 may include at least one processor 204 operatively connected to memory unit 208. Processor 204 may include one or more integrated circuits implementing the functions of a central processing unit (CPU) 206. CPU 206 may be a commercially available processing unit implementing an instruction set such as x86, ARM, Power, or MIPS instruction set families. During operation, CPU 206 may execute stored program instructions retrieved from memory unit 208. The stored program instructions may include software controlling the operation of CPU 206 to perform the operations described herein. In some examples, processor 204 may be a system-on-a-chip (SoC) that integrates the functions of CPU 206, memory unit 208, network interface, and input / output interface into a single integrated device. Computing system 202 may implement an operating system for managing various aspects of operation. Figure 2 The diagram shows a processor 204, a CPU 206, and a memory 208, but of course more than one of each of these can be used throughout the system.

[0029] Memory cell 208 may include volatile and non-volatile memory for storing instructions and data. Non-volatile memory may include solid-state memory, such as NAND flash memory, magnetic and optical storage media, or any other suitable data storage device that retains data when the computing system 202 is disabled or loses power. Volatile memory may include static and dynamic random access memory (RAM) for storing program instructions and data. For example, memory cell 208 may store machine learning model 210 or algorithm, training dataset 212 for machine learning model 210, and original source dataset 216.

[0030] The computing system 202 may include a network interface device 222 configured to provide communication with external systems and devices. For example, the network interface device 222 may include wired and / or wireless Ethernet interfaces as defined by the Institute of Electrical and Electronics Engineers (IEEE) 802.11 family of standards. The network interface device 222 may include a cellular communication interface for communication with cellular networks (e.g., 3G, 4G, 5G). The network interface device 222 may also be configured to provide a communication interface with an external network 224 or the cloud.

[0031] External network 224 may be referred to as the World Wide Web or the Internet. External network 224 can establish standard communication protocols between computing devices. External network 224 can allow information and data to be easily exchanged between computing devices and the network. One or more servers 230 can communicate with external network 224.

[0032] The computing system 202 may include an input / output (I / O) interface 220, which may be configured to provide digital and / or analog inputs and outputs. The I / O interface 220 is used to transfer information between internal storage devices and external input and / or output devices (e.g., HMI devices). The I / O interface 220 may include associated circuitry or bus networks for transferring information to or between one or more processors and storage devices. For example, the I / O interface 220 may include digital I / O logic lines that can be read or set by one or more processors, handshake lines for monitoring data transfer via the I / O lines, timing and counting facilities, and other known structures for providing such functionality. Examples of input devices include keyboards, mice, sensors, touchscreens, etc. Examples of output devices include monitors, touchscreens, speakers, head-up displays, vehicle control systems, etc. The I / O interface 220 may include additional serial interfaces (e.g., Universal Serial Bus (USB) interfaces) for communicating with external devices. I / O interface 220 can be referred to as an input interface (because it transmits data from external inputs, such as sensors) or an output interface (because it outputs data to external sources, such as displays).

[0033] The computing system 202 may include a human-machine interface (HMI) device 218, which may include any device enabling the system 200 to receive control input. The computing system 202 may include a display 232. The computing system 202 may include hardware and software for outputting graphical and textual information to the display 232. The display 232 may include an electronic display screen, projector, speaker, or other suitable device for displaying information to a user or operator. The computing system 202 may further be configured to allow interaction with a remote HMI and a remote display via a network interface device 222.

[0034] System 200 can be implemented using one or more computing systems. While this example depicts a single computing system 202 implementing all the described features, the intention is that various features and functions can be implemented separately by multiple computing units that communicate with each other. The specific system architecture chosen may depend on a variety of factors.

[0035] System 200 can implement machine learning algorithm 210 configured to analyze raw source dataset 216. Raw source dataset 216 may include raw or unprocessed sensor data, which may represent the input dataset for the machine learning system. Raw source dataset 216 may include video, video clips, images, text-based information, audio or human speech, time-series data, and raw or partially processed sensor data (e.g., radar images of objects). In some examples, machine learning algorithm 210 may be a neural network algorithm (e.g., a deep neural network) designed to perform a predetermined function. For example, a neural network algorithm may be configured in an automotive application to identify street signs or pedestrians in an image. One or more machine learning algorithms 210 may include algorithms configured to operate the machine learning models described herein, including one or more machine learning models among the VLP underlying models.

[0036] The computing system 202 may store a training dataset 212 for a machine learning algorithm 210. The training dataset 212 may represent a set of previously constructed data used to train the machine learning algorithm 210. The training dataset 212 may be used by the machine learning algorithm 210 to learn weighting factors associated with the neural network algorithm. The training dataset 212 may include a set of source data having a corresponding output or result that the machine learning algorithm 210 attempts to replicate through a learning process. In one example, the training dataset 212 may include input images comprising objects (e.g., street signs). The input images may include various scenes in which objects are identified. The training dataset 212 may also include text descriptions corresponding to scenes (e.g., "pedestrians are crossing the road") detected by vehicle sensors.

[0037] Machine learning algorithm 210 can be operated on in learning mode using training dataset 212 as input. Machine learning algorithm 210 can be executed over several iterations using data from training dataset 212. With each iteration, machine learning algorithm 210 can update its internal weighting factors based on the achieved results. For example, machine learning algorithm 210 can compare the output results (e.g., reconstructed or supplemented images in the case of image data as input) with those results included in training dataset 212. Since training dataset 212 includes the expected results, machine learning algorithm 210 can determine when performance is acceptable. After machine learning algorithm 210 achieves a predetermined performance level (e.g., 100% consistency with outputs associated with training dataset 212) or converges, machine learning algorithm 210 can be executed using data not in training dataset 212. It should be understood that in this disclosure, "convergence" can mean that a set (e.g., predetermined) number of iterations has occurred, or the residuals are sufficiently small (e.g., the change in the approximate probability across iterations is less than a threshold change), or other convergence conditions. The trained machine learning algorithm 210 can be applied to new datasets to generate annotated data. Within the context of the VLP model described in this paper, the loss between the predicted trajectory of the autonomous vehicle and its underlying true trajectory can be determined, and the VLP model can be trained to reduce this loss, for example, to achieve convergence.

[0038] Machine learning algorithm 210 can be configured to identify specific features in raw source data 216. Raw source data 216 may include multiple instances or input datasets for which it expects to supplement the results. For example, machine learning algorithm 210 can be configured to identify the presence of an agent in video images, annotate events, and / or command vehicles to take specific actions (planning) based on the agent's location data (perception) and the agent's predicted future movement / location (prediction). Machine learning algorithm 210 can be programmed to process raw source data 216 to identify the presence of specific features. Machine learning algorithm 210 can be configured to identify features in raw source data 216 as predetermined features (e.g., road signs, pedestrians, etc.). Raw source data 216 can be derived from various sources. For example, raw source data 216 can be actual input data collected by a machine learning system. Raw source data 216 can be machine-generated for a test system. As an example, raw source data 216 may include raw video images from a camera. Furthermore, as will be further described below with reference to the VLP basic model, the original source data 216 can be natural language text information associated with a scene (e.g., "a car is entering the intersection from the left").

[0039] Figure 3A schematic diagram is depicted of a control system 302 configured to control a vehicle 300, which may be a partially autonomous or fully autonomous vehicle, a partially autonomous robot, or a fully autonomous robot. The vehicle 300 and / or its control system 302 may incorporate one or more components of system 200, such as computing system 202, to command actuators 304 to perform specific actions based on processed readings from one or more sensors 306. For example, the control system 302 may be configured to utilize a VLP base model disclosed herein to control the movement of the vehicle via the output of the VLP base model, which commands the actions to be taken by actuators 304.

[0040] One or more sensors 306 may include one or more image sensors (e.g., cameras, video sensors, radar sensors, ultrasonic sensors, LiDAR sensors), and / or position sensors (e.g., GPS). Sensor 306 may be configured to generate raw source data 216. One or more of the specific sensors may be integrated into vehicle 300. In the context of agent recognition and processing as described herein, sensor 306 is a camera mounted to or integrated into vehicle 300. Instead of or attached to the one or more specific sensors described above, sensor 306 may include a software module configured to determine the state of actuator 304 upon execution.

[0041] In embodiments where vehicle 300 is a fully or partially autonomous vehicle, actuator 304 may be embodied in the vehicle 300's brakes, accelerator, propulsion system, engine, transmission system, or steering system (e.g., steering wheel). For example, actuator control commands may be determined such that actuator 304 is controlled to cause vehicle 300 to avoid collisions with detected agents. Detected agents may also be classified according to what a classifier deems them most likely to be, such as pedestrians or trees. Actuator control commands may be determined based on the classification.

[0042] In other embodiments where vehicle 300 is a fully or partially autonomous robot, vehicle 300 may be a mobile robot configured to perform one or more functions, such as flying, swimming, diving, and stepping, via actuator 304. The mobile robot may be at least partially autonomous lawnmower or at least partially autonomous cleaning robot. In such embodiments, actuator control commands may be determined such that the propulsion, steering, and / or braking units of the mobile robot can be controlled to prevent collisions with identified objects.

[0043] Figure 4A high-level overview of an end-to-end autonomous driving system 400 according to one embodiment is shown. The system 400 shown herein may be referred to as the VLP model or framework disclosed herein. The end-to-end system 400 may be incorporated into a vehicle 300, such as its computing system 202, to operate the vehicle to avoid objects or otherwise control the vehicle 300 based on sensed surroundings of the vehicle 300.

[0044] Typically, and as will be described in more detail below, the system integrates VLP into both the BEV and the planning model, and evaluates the system's performance through adaptation to a simulation framework (e.g., CARLA) to extract a language-based description of the underlying real-world data and use this information during training. VLP enhances ADS from the aspects of autonomous BEV inference and autonomous driving decision-making through two innovative modules, ALP and SLP. Leveraging LLM and contrastive learning, ALP performs agent-by-agent learning to improve local details on the BEV, while SLP performs sample-by-sample learning to improve the global contextual understanding of ADS. VLP can be activated during training, ensuring that no additional parameters or computations are introduced during inference. To evaluate the pipeline, open-source benchmarking frameworks, such as Bench2Drive, are modified in both their action space (for zero-shot evaluation purposes) and PID controller for more realistic applications.

[0045] At position 402, the system receives images or image data. Images can be generated from one or more image sensors (e.g., cameras) mounted on the vehicle as it travels through its environment. For example, cameras can be mounted to capture images of the environment in front of, behind, and to the sides of the vehicle; however, in this embodiment, the images and associated data are generated from an open-source tool that simulates a real-world driving environment. This open-source tool could be the CARLA (Car Learning to Act) simulator. The CARLA simulator is an open-source tool that helps researchers develop, train, and test autonomous driving systems. It is built on Unreal Engine and uses the ASAM OpenDrive standard to create realistic simulations of the real world. CARLA is scalable; multiple clients can control different actors in the same or different nodes. It also has a flexible API; users can control many aspects of the simulation, including traffic, pedestrians, and weather. CARLA also features realistic environments; CARLA's simulations mimic real-world cities, towns, and highways including vehicles and other objects. CARLA can generate synthetic training data for autonomous driving and other robotic applications. Users can test and evaluate their trained autonomous driving agents within a simulation without putting the hardware or other road users at risk.

[0046] The generated images can be passed through one or more machine learning layers (e.g., BEV encoder 404) to create a BEV representing the environment surrounding the vehicle. The BEV encoder 404 can be a computational model or network designed to process data (such as images from multiple cameras) and transform the data into a bird's-eye view representation of the environment. The primary function of the BEV encoder is to transform perspective images into a top-down, bird's-eye view. This involves implementing complex geometric transformations, where the encoder estimates the 3D geometry of the scene, maps camera images onto planes (such as roads), and combines them into a coherent top-down view. The BEV provides a clear and unobstructed representation of the surrounding environment, facilitating spatial reasoning and navigation.

[0047] Once a BEV image is generated, the BEV encoder extracts relevant spatiotemporal features from the scene. These features can include information about various agents in the field of view. For example, these features can include the position and movement of other vehicles, pedestrians, and cyclists; lane markings, road boundaries, and traffic signs; obstacles or hazards that vehicles need to navigate around; and so on. The encoder can use a convolutional neural network (CNN) or other deep learning architecture to extract these features from the BEV image. Furthermore, in this embodiment, the BEV encoder 404 processes not only camera data but also fuses data from LiDAR, radar, and other image sensors to enhance environmental perception. The combination of these multiple data streams helps create a more robust and accurate BEV representation, particularly for object detection or understanding depth and distance. The output of the BEV encoder can be a vectorized or grid-based representation of the scene that reflects key spatial relationships and agent positions in the environment.

[0048] like Figure 4 As shown, the output of the BEV encoder (i.e., BEV 406) can be used as input to all three models: the perception model, the prediction model, and the planning model. For example, the perception model 408 utilizes computer vision based on input received from an image sensor, such as the BEV, to perform object detection, etc. The prediction model 410 may include a machine learning model configured to estimate the trajectories and intentions of objects detected in the BEV based on their past motions, orientations, and contextual information. The planning model 412 may include route planning, trajectory planning, and decision-making for the vehicle to navigate relative to other objects in the BEV, translating these decisions into actions to be taken by the vehicle in real-world conditions.

[0049] CARLA also possesses underlying real-world data associated with detected agents in the environment. This can include the spatiotemporal information described above for each detected agent in the image. As explained elsewhere in this document, this underlying real-world information can be extracted for various purposes.

[0050] As mentioned above, Figure 4 The VLP framework shown includes ALP and SLP components, which enhance autonomous driving in terms of BEV reasoning and autonomous driving decision-making. The ALP and SLP components focus on improving local details in the BEV source memory and guiding the planning process of the autonomous vehicle, respectively.

[0051] In this embodiment, ALP first aligns the base real-world regions of each agent—the vehicle, foreground objects, and background objects—with the generated BEV map and then clips the regions of interest. Three-dimensional (3D) bounding boxes can be used to clip the regions of the vehicle and foreground objects, and panoramic scene masks can be used to segment lane regions. Figure 4 As shown, a bounding box is displayed above the vehicle, another bounding box is displayed above another foreground object (e.g., another vehicle), and yet another bounding box is displayed above a background object (e.g., lane markings on a road). The system then performs a pooling operation on the obtained local BEV regions to generate a single feature representation for the corresponding agent. After pooling, the local agent features along each sample in the batch are concatenated to form a per-agent BEV feature tensor.

[0052] To ensure that local BEV features express the expected information, the system performs a BEV-expectation alignment process by fully utilizing a Language Model (LM) and contrastive learning. The system defines the expected information as perceptual information derived from the corresponding agent, such as agent labels, bounding boxes, and future trajectories. This driving-related basic real-world information, which can be embedded in the local BEV features, forms text-based prompts. In other words, the LM can generate text-based prompts based on basic real-world information associated with the agent. For example, the generated prompt could include text such as “The agent is {class name}. Its 3D bounding box is located at coordinates {x,y,z}. Its future trajectory is {x1,y1,z1 for time t1}.” This text prompt can be constructed using templates and basic real-world information, such as basic real-world trajectories existing in the training data or high-level commands. It can be used as text input 702. For example, “The autonomous vehicle is turning left, and its future trajectory is (x1,y1), …(x6,y6)” can be generated based on scene-related basic real-world information already existing in the training data; the natural language sentence format can be template-based. Several text entries can be provided for a corresponding number of videos or images. With the help of more relevant and detailed information about autonomous driving included in sentences, the language path can provide higher-level semantic and comprehensive cues for the planning module. Several unrelated text strings can also be provided for training purposes. As explained above, BEV data can be generated by and extracted from open-source simulators (e.g., CARLA), and therefore, these text cues can be generated based on text-based descriptions of underlying real-world data extracted from open-source simulators.

[0053] like Figure 4 As shown, the text prompt is then passed to the Visual Language Model (VLM) to generate the corresponding expected features for the agent. The VLM can implement contrastive learning techniques, such as those introduced in the Contrastive Language-Image Pre-training (CLIP) model. Other contrastive learning models can be employed; CLIP is merely an example and is illustrated in [the diagram]. Figure 5CLIP, developed by OpenAI, is designed to understand and connect images and natural language descriptions in a way that allows it to perform a wide range of visual and language tasks. CLIP employs a dual-encoder architecture, consisting of a visual encoder and a text encoder, along with a shared embedding space. The visual encoder processes images, while the text encoder processes natural language descriptions. The visual encoder, based on a visual model such as a convolutional neural network (CNN), transforms images into fixed-length vector representations. The text encoder processes text descriptions by transforming them into fixed-length vector representations. CLIP is a visual-language foundational model trained on open-world data using contrastive learning. Contrastive learning is a type of machine learning where the model learns to distinguish between positive and negative pairs of data. In the context of CLIP, "positive pairs" consist of semantically related images and text descriptions, while "negative pairs" consist of unrelated images and randomly selected text descriptions. During training, CLIP is designed to encourage features from related text and image pairs to be incorporated into the shared embedding space, while pushing away unrelated pairs.

[0054] CLIP's shared embedding space allows for zero-shot learning. When presented with both image and text cues, CLIP can rank how well an image matches a cue without specific training data for that particular task. CLIP can be used to perform a variety of visual language tasks, including image classification, text-based image retrieval (e.g., retrieving images based on text queries), image captioning, zero-shot object recognition, and more.

[0055] The contrastive learning concept used in CLIP (whose teachings are included in the VLP base model) Figure 5 The model is illustrated in the figure, generally shown as a contrastive learning model at 500. As shown, multiple natural language text descriptions 502 are fed to a text encoder 504, and multiple images 506 are fed to an image encoder 508. Model 500 then performs feature mapping, where the vectors output by the encoders are mapped to a joint embedding space. For example, an image vector (e.g., 1x256 in size) output by the image encoder is matched to a corresponding text vector (e.g., 1x256 in size) output by the text encoder. The model then performs a dot product between batches of image and text features to obtain the similarity between these vectors, generally illustrated at 510.

[0056] refer to Figure 5In the example illustrated, multiple images 506 (one of which in this example is an image of a tiger) are fed into image encoder 508, and multiple text phrases 502 (one of which in this example is a phrase like "photo of a tiger") are fed into text encoder 504. Several unrelated or dissimilar text phrases and images are also fed into the encoder. For example, images of objects other than tigers are fed into image encoder 508, and phrases unrelated to tigers are also fed into text encoder 504. The image encoder generates text with features I1, I2...I... N The image vector, while the text encoder generates images with features T1, T2...T N The text vector. The diagonal of the matrix 510 obtained from the dot product shows paired images and text according to their possible similarity, while the off-diagonal represents unpaired image and text features (e.g., an image of a cat and a text description such as "picture of a dog").

[0057] Therefore, the contrastive learning model brings image and text embeddings together when they correspond and pushes them apart when they don't. In other words, the reference... Figure 5 During training, the contrastive learning model aims to increase the similarity between diagonal elements (i.e., pairs) while decreasing the similarity between off-diagonal elements. As another example, during training, if the model is given an image of a cat and a text description such as "pictures of cats," the model aims to minimize the distance (similarity) between the image and text embeddings in the shared space; conversely, if the model is given an image of a cat and a text description such as "pictures of dogs," the model aims to maximize the distance (dissimilarity) between their embeddings. This contrastive training objective encourages the model to learn to understand the semantic relationships between images and text. This is a way to teach the model to closely associate matching image-text pairs and effectively distinguish mismatched pairs. The result is a shared embedding space where similar pairs cluster together, while dissimilar pairs are far apart.

[0058] Back Figure 4 Generally, text prompts can be passed through a language encoder to generate text-based expected features associated with the detected agent. Furthermore, vision-based planning features associated with the detected agent within the environment can be extracted from the BEV used by planning model 412. Then, contrastive learning and VLM techniques (e.g., CLIP) are applied to system 400, where CLIP performs contrastive learning between (1) the text-based expected features and (2) the vision-based planning features. This derives the similarity between the vision-based planning features and the text-based expected features.

[0059] In other words, in ALP, system 400 is operated to extract underlying real information from BEV data to form cues associated with each agent and the vehicle. This description is then passed to VLM to generate corresponding agent-specific expected features. The system may apply multilayer perceptron (MLP) layers or other types of neural network layers to adapt the expected features to the BEV feature space. The agent-specific expected features are then concatenated batch by batch to generate agent-specific text feature tensors. The system then applies a contrastive learning loss between the agent-specific BEV features and the text-based features for alignment.

[0060] In SLP, System 400 follows a similar process, but specifically focuses on the autonomous vehicle. In other words, textual cues can be generated based on information about the autonomous vehicle. For example, “The autonomous vehicle is turning left. The planned future trajectory with 3 timestamps is (x1, y1), (x2, y2), and (x3, y3).” Therefore, SLP is specifically designed to perform contrastive learning between textual features associated with the autonomous vehicle's planning and image data from the BEV associated with the autonomous vehicle.

[0061] The predicted trajectory of the vehicle is generated by planner model 412, and based on the contrastive learning disclosed herein, the VLP combined with the method described herein can enhance or improve the capabilities of planner model 412 (via SLP) and BEV model (via ALP). The determined similarity between vision-based planning features and text-based expectation features can improve the output of planner model 412.

[0062] The experiments were conducted in a closed-loop environment to determine the performance of the enhanced planner model and the enhanced BEV model. For this purpose, the Bench2Drive framework was used to implement benchmarking. Current benchmarks typically focus on specific tasks or scenarios, neglecting a comprehensive evaluation of the overall performance of autonomous systems. They often lack diversity in driving environments, road conditions, and traffic scenarios, leading to incomplete evaluations. This is why the Bench2Drive framework is designed to evaluate autonomous driving systems across multiple capabilities. The framework includes a wide range of tasks and scenarios for a holistic assessment of perception, planning, and control capabilities. Bench2Drive also introduces novel metrics that consider safety, efficiency, and comfort, providing a more complete description of system capabilities. For implementing closed-loop evaluation, Bench2Drive uses the CARLA simulator as explained above. To address the domain gap between models trained on real datasets and simulated evaluation environments, Bench2Drive also provides a large CARLA dataset for training. This disclosure uses this dataset to train the models described herein and evaluates these models in both open-loop and closed-loop environments.

[0063] Bench2Drive is designed to address several key limitations in current evaluation frameworks for autonomous driving systems, particularly those focused on end-to-end automated driving (E2E-AD) approaches. Traditional methods often rely on open-loop log-replay evaluations, where models are tested based on pre-recorded trajectories and metrics such as L2 error (deviation from the recorded path) or collision rate are used. However, these metrics fail to capture the full complexity of real-world driving, especially in scenarios demanding dynamic decision-making and interactive behavior. Open-loop evaluations do not account for distributional biases or causal confounding, where vehicle actions can influence the environment, as does occur in real-world driving. Furthermore, most benchmarks provide imbalanced datasets, with a significant portion of the scenarios being simple (such as straight-line driving), insufficient to challenge autonomous systems to handle complex and interactive traffic conditions.

[0064] Bench2Drive addresses these issues by introducing a closed-loop evaluation framework that places the vehicle's decisions within a feedback loop with the environment, allowing for a more realistic and comprehensive assessment of driving performance. In this closed-loop scenario, the autonomous vehicle's actions directly impact the environment, which in turn affects the next set of challenges the vehicle must overcome. This effectively replicates real-world conditions where interactions with other road users, traffic signals, and obstacles are unpredictable and require adaptive responses. Bench2Drive implements this closed-loop approach across a diverse set of scenarios, urban environments, and weather conditions, ensuring that the evaluation covers a wide range of driving skills and environments.

[0065] At the heart of Bench2Drive is a large-scale, fully annotated dataset consisting of 2 million frames derived from 10,000 short segments. These segments were collected from 44 interactive driving scenarios, such as cutting in, overtaking, and detouring, all acquired across 12 towns with diverse landscapes (city, village, university scenarios) under varied weather conditions (sunny, foggy, rainy, etc.). The evaluation protocol required the E2E-AD system to complete 220 short routes, each approximately 150 meters long and containing a single interactive scenario. This fine-grained approach to scenario design allows for individual testing of specific driving capabilities, making it easier to identify the strengths and weaknesses of different systems. By focusing on shorter routes, Bench2Drive reduces performance variations that can occur in longer route evaluations and provides more reliable and detailed insights into specific driving capabilities.

[0066] To ensure a fair and algorithmic comparison, Bench2Drive provides a standardized, large-scale training dataset collected using the state-of-the-art expert model Think2Drive. This dataset is annotated with detailed information including 3D bounding boxes, depth, and semantic segmentation. The annotations cover a wide range of sensor configurations, including LiDAR, cameras, radar, and HD maps, allowing the system to be trained and tested on diverse sensor inputs. This eliminates the problem of individual teams using their own training datasets, which previously made direct algorithmic comparisons difficult due to differences in data quality and diversity. Bench2Drive's training data ensures that all autonomous driving models are tested under similar conditions, providing a level playing field for evaluating different approaches.

[0067] Bench2Drive also contributes to benchmarking by implementing several state-of-the-art E2E-AD models, including UniAD, VAD, TCP, and ThinkTwice, and evaluating them using both open-loop and closed-loop metrics. The results demonstrate that traditional open-loop metrics, such as L2 error, may be insufficient for comparing models' driving capabilities, especially in complex scenarios. Instead, closed-loop evaluation metrics, such as driving scores and success rates, provide a more meaningful assessment of how well a model handles interactive and complex traffic conditions. Bench2Drive's fine-grained evaluation framework, combined with its expanded training data and closed-loop testing environment, provides the research community with a comprehensive and equitable platform for advancing the development of E2E-AD systems.

[0068] According to one embodiment, the evaluation of the model using Bench2Drive is described below. Since VLP is a training-only method, basic real-world information needs to be extracted from the CARLA simulator. However, the Bench2Drive framework also provides pre-extracted datasets with various driving scenarios and actions; therefore, this dataset can be used to train System400. For baseline comparison, the same VAD modifications recommended by Bench2Drive are followed. The driving commands are expanded from three to six, including changing lanes to the left, changing lanes to the right, and keeping lanes as commands. For evaluation, Bench2Drive provides new metrics related to multi-capability evaluation in closed loops. The two main metrics obtained are: success rate and driving score. Success rate is the proportion of routes that the vehicle can complete without any traffic violations. Driving score is a comprehensive metric considering both route completion and violation penalties. Traffic violation penalties are used in a multiplicative manner. For open-loop evaluation, metrics such as displacement error and collision rate percentage are obtained. Furthermore, metrics related to non-planning tasks, such as object detection, mapping tracking, and prediction, are available.

[0069] Bench2Drive provides an ideal framework for evaluating System 400 in closed-loop scenarios, where vehicle actions influence the surrounding environment, offering a more realistic assessment. Bench2Drive includes numerous interactive driving scenarios (e.g., cutting in, overtaking, merging), each designed to test specific driving skills. These scenarios, along with diverse weather conditions and environments, enable comprehensive testing of the system's performance under various real-world conditions. Furthermore, the system's ability to generate trajectories based on visual and textual input can be directly tested in closed-loop environments, where vehicle actions influence agents in the scenario (such as pedestrians or other vehicles). This allows for dynamic, real-time evaluation, where Bench2Drive measures how well the vehicle responds to evolving conditions. Because Bench2Drive's evaluations include numerous short routes, each testing a specific driving scenario, system performance can be evaluated individually for its core planning capabilities. For example, the system's ability to plan around pedestrians or cyclists in scenarios involving pedestrian crossings or complex vehicle interactions can be tested. In addition, Bench2Drive's performance metrics (including success rate and driving score) are used to evaluate how effectively System 400 navigates routes while avoiding collisions, obeying traffic rules, and successfully completing tasks. The system's predicted trajectory can also be compared to a baseline real trajectory to measure performance and identify areas for improvement, aligned with loss-based improvement processes (e.g., minimizing the error between the predicted and actual trajectories).

[0070] In the closed-loop evaluation, results were collected from 110 routes (each approximately 150 meters long and containing a single specific scenario) on the Bench2Drive benchmark to demonstrate closed-loop performance compared to the baseline VAD tiny model. The evaluation showed that VLP significantly outperformed VAD in driving score (8% improvement) and route completion (13% improvement).

[0071] Figure 6A method 600 is illustrated for training an end-to-end autonomous driving system (e.g., system 400) in a closed-loop environment using a VLP machine learning model, according to one embodiment. This method can be executed by one or more processors disclosed herein. At 602, image data is generated by multiple image sensors (e.g., cameras, lidar, radar, etc.) mounted to or around the vehicle. The image sensors acquire images of the environment surrounding the vehicle. Image processing is performed on the image data to detect agents in the environment. At 604, a BEV model is executed based on the images or image data. Object recognition and classification, as described above, can be used. A BEV (e.g., BEV 406) is generated based on the image data and the results of object recognition or other object detection. The BEV includes spatiotemporal information associated with the vehicle and detected agents in the environment. At 606, a planning model (e.g., planning model 412) is executed on the BEV to generate a predicted trajectory for the autonomous vehicle to navigate in the environment. For example, the predicted trajectory can be used to issue commands to be taken by actuator 304 to navigate the autonomous vehicle.

[0072] In this context, improvements can be made to the end-to-end autonomous driving system. At 608, a VLP machine learning model is executed to improve the end-to-end autonomous driving system. The VLP model execution at 608 may include the steps taken at 610-616, which are described below. At 610, vision-based planning features associated with detected agents in the environment are extracted. These vision-based planning features may include spatiotemporal information associated with detected agents in an image. As mentioned above, this can be extracted, for example, from simulated data. Then, at 612, a text prompt is generated based on the extracted spatiotemporal information. This text prompt is generated using base ground truth labels from the current time frame, such as position, heading, next position, etc., in the vehicle coordinate system. At 614, this text prompt is passed through a language encoder to generate text-based expected features associated with the detected agents. Finally, at 616, a contrastive learning model can be executed to derive the similarity between the vision-based planning features and the text-based expected features. This contrastive learning can be implemented using a model such as CLIP.

[0073] At point 618, leveraging these similarities generated by the contrastive learning model, the BEV model and planning model can be enhanced or improved. This contrastive learning allows for the generation of modified BEV data and planning model outputs. At point 620, these modified outputs are evaluated via closed-loop evaluation. Here, a closed-loop evaluation of the end-to-end autonomous driving system is implemented, comprising the enhanced BEV model and the enhanced planning model based on similarities derived from the contrastive learning. This closed-loop evaluation is conducted through real-time interaction with the simulated environment and dynamic responses to vehicle actions based on predicted trajectories. During this closed-loop evaluation, evaluation metrics such as success rate and driving score can be acquired. The end-to-end autonomous driving system can then be modified based on the acquired evaluation metrics.

[0074] While exemplary embodiments have been described above, this does not mean that these embodiments describe all possible forms included in the claims. The language used in this specification is descriptive and not limiting, and it should be understood that various changes may be made without departing from the spirit and scope of this disclosure. As previously stated, features of various embodiments may be combined to form further embodiments of the invention that may not be explicitly described or illustrated. While various embodiments may have been described as providing an advantage or superiority over other embodiments or prior art implementations in one or more desired features, those skilled in the art will recognize that one or more features or characteristics may be compromised to achieve desired overall system properties depending on the particular application and implementation. These properties may include, but are not limited to, cost, strength, durability, lifecycle cost, merchantability, appearance, packaging, size, maintainability, weight, manufacturability, ease of assembly, etc. Therefore, in the sense that any embodiment is described as less desirable than other embodiments or prior art in achieving the desired features, such embodiments are not outside the scope of this disclosure and may be desirable for a particular application.

Claims

1. A method for training an end-to-end autonomous driving system using a Visual Language Planning (VLP) machine learning model in a closed-loop environment, the method comprising: Receives images generated from multiple image sensors installed in the vehicle; The BEV machine learning model is executed based on the image to generate a bird's-eye view (BEV) of the environment; A planning machine learning model is executed on the BEV to generate a predicted trajectory for navigating the autonomous vehicle in the environment; Executing a VLP machine learning model to improve the end-to-end autonomous driving system includes: Extract vision-based planning features associated with detected agents in the environment, wherein the vision-based planning features include spatiotemporal information associated with detected agents in the image; Text prompts are generated based on the spatiotemporal information extracted and associated with the detected agents in the image. The text prompt is delivered via a language encoder to generate text-based expected features associated with the detected agent; and A contrastive learning model is executed to derive the similarity between the vision-based planning features and the text-based expected features; The BEV model and the planning model are enhanced based on the aforementioned similarity; Closed-loop evaluation of the end-to-end autonomous driving system is performed by interacting with the simulated environment in real time and responding dynamically to actions taken by the vehicle based on the predicted trajectory, using an enhanced BEV model and an enhanced planning model. The evaluation metrics are obtained during the closed-loop evaluation; and The end-to-end autonomous driving system is modified based on the obtained evaluation metrics.

2. The method of claim 1, wherein the image is generated from an open-source tool that simulates a real-world driving environment.

3. The method of claim 2, wherein generating the text prompt includes a text-based description extracted from the underlying real data of the open-source tool.

4. The method of claim 3, wherein the open-source tool is CARLA.

5. The method of claim 1, wherein the closed-loop evaluation is performed via a Bench2Drive benchmark.

6. The method of claim 1, wherein the closed-loop evaluation includes determining the loss between the predicted trajectory of the vehicle and the basic true trajectory of the vehicle.

7. The method of claim 1, wherein the text prompt is generated from a base real-world label associated with the image.

8. The method according to claim 7, wherein the contrastive learning model comprises: A text encoder configured to output text-based vectors based on the text prompts; and An image encoder is configured to output image-based vectors that represent image-based features associated with the detected agent in the BEV.

9. An end-to-end autonomous driving system utilizing a Visual Language Planning (VLP) machine learning model in a closed-loop environment, the system comprising: processor; and The memory includes instructions that, when executed by the processor, cause the processor to: Receive images of the vehicle's surrounding environment; The BEV machine learning model is executed based on the image to generate a bird's-eye view (BEV) of the environment; A planning machine learning model is executed on the BEV to generate a predicted trajectory for navigation of the autonomous vehicle in the environment; Execute VLP machine learning models to: Extract vision-based planning features associated with detected agents in the environment, wherein the vision-based planning features include spatiotemporal information associated with detected agents in the image; Text prompts are generated based on the spatiotemporal information extracted and associated with the detected agents in the image. The text prompts are delivered via a language encoder to generate text-based expected features associated with the detected agent; and A contrastive learning model is executed to derive the similarity between the vision-based planning features and the text-based expected features. To improve the end-to-end autonomous driving system; The similarity is used to enhance the BEV model and the planning model; and Closed-loop evaluation of the end-to-end autonomous driving system is implemented by interacting with the simulated environment in real time and dynamically responding to actions taken by the vehicle based on the predicted trajectory, using an enhanced BEV model and an enhanced planning model.

10. The system of claim 9, wherein the memory includes additional instructions that, when executed by the processor, cause the processor to: The evaluation metrics are obtained during the closed-loop evaluation; and The end-to-end autonomous driving system is modified based on the obtained evaluation metrics.

11. The system of claim 9, wherein the image is generated from an open-source tool that simulates a real-world driving environment.

12. The system of claim 11, wherein the generated text prompt comprises a text-based description extracted from underlying real data from the open-source tool.

13. The system of claim 12, wherein the open-source tool is CARLA.

14. The system of claim 9, wherein the closed-loop evaluation is performed via a Bench2Drive benchmark.

15. The system of claim 9, wherein the closed-loop evaluation includes determining the loss between the predicted trajectory of the vehicle and the basic true trajectory of the vehicle.

16. The system of claim 9, wherein the text prompt is generated from a base real-world label associated with the image.

17. The system of claim 7, wherein the contrastive learning model comprises: A text encoder configured to output text-based vectors based on the text prompts; and An image encoder is configured to output image-based vectors that represent image-based features associated with the detected agent in the BEV.

18. A method comprising: Receive images associated with the vehicle's surrounding environment; Based on the image, a BEV model is performed to generate a bird's-eye view (BEV) of the environment; A planning model is executed on the BEV to generate a predicted trajectory for navigation of the autonomous vehicle in the environment; Execute VLP machine learning models to: Extract vision-based planning features associated with agents within the environment, wherein the vision-based planning features include spatiotemporal information associated with the agents; Text prompts are generated based on the extracted spatiotemporal information; The text prompt is delivered via a language encoder to generate text-based expected features; and A contrastive learning model is executed to derive the similarity between the vision-based planning features and the text-based expected features. To improve the end-to-end autonomous driving system; The similarity is used to enhance the BEV model and the planning model; and Closed-loop evaluation of the end-to-end autonomous driving system is implemented by interacting with the simulated environment in real time and dynamically responding to actions taken by the vehicle based on the predicted trajectory, using an enhanced BEV model and an enhanced planning model.

19. The method of claim 18, further comprising: During the closed-loop evaluation, the evaluation metrics are obtained; and The end-to-end autonomous driving system is modified based on the obtained evaluation metrics.

20. The method of claim 18, wherein the image is generated from CARLA simulating a real-world driving environment, and wherein the generated text prompt includes a text-based description extracted from underlying real-world data extracted from CARLA.

Citation Information

Patent Citations

  • Systems and methods for vision-language planning (VLP) foundation models for autonomous driving

    US12528507B2

  • Vision-language-planning (VLP) models with agent-wise learning for autonomous driving

    US20250156745A1