Visual language planning (VLP) model with proxy-by-proxy learning for autonomous driving
By adopting the basic model of agency-by-agent learning of visual language planning (VLP) in autonomous driving systems, and using contrast learning technology to combine language knowledge with visual models, the problem of traditional methods performing poorly in open-world scenarios is solved, and higher accuracy, safety and generalization capabilities are achieved.
Patent Information
- Application Number
- CN202411600308.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-10
- Filing Date
- 2024-11-11
- Publication Date
- 2025-05-13
AI Technical Summary
Traditional supervised learning methods rely only on visual input and constrained autonomous datasets during training, resulting in the model that may converge to a suboptimal state when facing open-world scenarios and fail to effectively integrate language understanding with vision-based planning systems.
The basic model of visual language planning (VLP) with agent-by-agent learning is adopted, and language knowledge and visual model information are combined through comparative learning technology to improve the planning and generalization capabilities of autonomous driving systems. During training, the model uses natural language text prompts to compare and learn BEV features, and refines the BEV features to generate a modified predictive trajectory of the vehicle.
By seamlessly incorporating language understanding into the planning process, the accuracy, safety and generalization capabilities of autonomous driving systems are significantly improved, allowing them to better handle a wide range of scenarios and reduce dependence on specific training data.
Smart Images

Figure CN119992487A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to systems and methods for a vision-language planning (VLP) based model with agent-wise learning for autonomous driving. Background Art
[0002] An autonomous vehicle, often referred to as self-driving or driverless vehicle, is a vehicle that is capable of navigating and operating on roads and in various environments without direct human control. Autonomous vehicles use a combination of advanced technologies and sensors to perceive their surroundings, make decisions, and perform driving tasks.
[0003] Autonomous vehicles are typically equipped with a variety of sensors, including lidar, radar, cameras, ultrasonic sensors, and sometimes additional technologies such as GPS and IMU (inertial measurement units). These sensors provide real-time data about the vehicle's surroundings, including the location of other vehicles, pedestrians, road signs, and road conditions. The vehicle's onboard computer uses the data from the sensors to create a detailed map of the environment and perceive objects and obstacles. This information is critical for navigation and collision avoidance.
[0004] Machine learning (ML) and artificial intelligence (AI) play a vital role in autonomous vehicles. Deep learning algorithms are used for tasks such as object detection, lane keeping, and decision making, and can rely on image processing to perform these tasks. These algorithms enable vehicles to understand and respond to complex and dynamic traffic situations. Summary of the invention
[0005] In one embodiment, a method for training an autonomous driving system using a visual-language planning (VLP) machine learning model with agent-by-agent learning includes: receiving image data generated from a camera mounted to a vehicle, wherein the image data includes agents in an environment outside the vehicle; detecting agents in the environment based on the image data via image processing; generating a bird's-eye view (BEV) of the environment based on the image data, wherein the BEV includes BEV features, the BEV features including spatiotemporal information associated with the vehicle and the detected agents; inputting data from the BEV into a perception model, a prediction model, and a planning model of an end-to-end autonomous driving system to generate a predicted trajectory for the vehicle; and executing the agent-centric visual-language planning (VLP) machine learning model. The VLP machine learning model is configured to, when executed: extract agent-by-agent BEV features from the BEV, wherein the agent-by-agent BEV features are associated with corresponding agents in the environment, generate natural language text prompts associated with the agents in the environment, extract agent-by-agent text features from the natural language text prompts, wherein the agent-by-agent text features are associated with corresponding agents in the environment, execute a contrastive learning model to derive similarities between the agent-by-agent BEV features and the agent-by-agent text features, and refine the BEV features for a perception model, a prediction model, and a planning model based on the similarities to generate a modified predicted trajectory for the vehicle.
[0006] In another embodiment, a system utilizing a visual language planning (VLP) machine learning model is provided. The system includes a camera mounted to a vehicle and configured to generate image data associated with an agent in an environment outside the vehicle, a processor, and a memory including instructions that, when executed by the processor, cause the processor to perform the functions described in the preceding paragraphs.
[0007] In another embodiment, an apparatus for training at least one machine learning model includes a processor and a memory containing instructions that, when executed by the processor, cause the processor to perform these functions. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 A system for training a neural network according to one embodiment is shown.
[0009] Figure 2 A computer-implemented method for training and utilizing a neural network is shown according to one embodiment.
[0010] Figure 3 A schematic diagram of a control system configured to control a vehicle, which may be a partially autonomous vehicle, a fully autonomous vehicle, a partially autonomous robot, or a fully autonomous robot, is shown according to one embodiment.
[0011] Figure 4A schematic diagram of an end-to-end autonomous driving system according to one embodiment is shown.
[0012] Figure 5 A schematic diagram of a contrastive learning model according to one embodiment is shown.
[0013] Figure 6 A schematic diagram of a visual-language planning (VLP) machine learning model with agent-by-agent learning for an end-to-end autonomous driving system is shown according to one embodiment.
[0014] Figure 7 A schematic diagram of a method for training an autonomous driving system using a VLP machine learning model according to one embodiment is shown. DETAILED DESCRIPTION
[0015] Embodiments of the present disclosure are described herein. However, it should be understood that the disclosed embodiments are merely examples, and other embodiments may take various and alternative forms. The drawings are not necessarily drawn to scale; some features may be enlarged or reduced to show the details of a particular component. Therefore, the specific structural and functional details disclosed herein should not be interpreted as limiting, but merely as a representative basis for teaching those skilled in the art to adopt the embodiments in various ways. As will be understood by those of ordinary skill in the art, the various features illustrated and described with reference to any of the figures may be combined with the features illustrated in one or more of the other figures to produce embodiments that are not explicitly illustrated or described. The combination of features shown provides representative embodiments of typical applications. However, for specific applications or implementations, various combinations and modifications of features consistent with the teachings of the present disclosure may be required.
[0016] As used herein, "a", "an", and "the" refer to both the singular and the plural, unless the context clearly indicates otherwise. For example, a "processor" programmed to perform various functions refers to one processor programmed to perform each function, or more than one processor programmed together to perform each of the various functions.
[0017] In the context of autonomous vehicles, the term "agent" can refer to objects or entities in the environment that the autonomous vehicle surrounds or interacts with. This includes pedestrians, other vehicles, cyclists, road signs, traffic lights, lane markings, etc. Objects or features detected by the autonomous vehicle's sensors that are used to control the autonomous vehicle's decision making can be collectively referred to as agents.
[0018] The present disclosure incorporates by reference in its entirety U.S. patent application filed on the same day as the present disclosure, having attorney docket number 097182-00294 and entitled “SYSTEMS AND METHODS FOR VISION-LANGUAGEPLANNING (VLP) FOUNDATION MODELS FOR AUTONOMOUSDRIVING”.
[0019] The rapid development of autonomous driving technology has ushered in a new era of transportation, promising safer and more efficient journeys. Autonomous driving systems generally involve three high-level tasks: (1) perception, (2) prediction, and (3) planning. Perception involves the ability of a vehicle to understand and interpret its environment. This task includes various subcomponents, such as computer vision, sensor fusion, and localization. Key elements of perception include object detection (e.g., identifying and tracking agents outside the autonomous vehicle), localization (e.g., determining the precise position and orientation of the vehicle in the world, typically using GPS and other sensors), and sensor fusion (e.g., combining data from different sensors, such as cameras, lidar, radar, and ultrasonic sensors to build a comprehensive view of the surrounding environment). Prediction involves anticipating how other road users and agents in the environment will behave in the near future. This task typically involves using machine learning models to estimate the trajectories and intentions of agents (including pedestrians, other vehicles, and potential obstacles). Accurate prediction is essential for making safe driving decisions. Planning involves determining the best path and actions for the autonomous vehicle to navigate its environment. This typically includes tasks such as route planning, trajectory planning, and decision making. Planning systems consider information from perception and prediction to make decisions such as when to change lanes, when to stop at intersections, how to react to unexpected events, etc.
[0020] In an autonomous driving system, the BEV can be the primary source of information for an end-to-end autonomous driving system, providing a top-down holistic view of the surrounding environment, allowing the autonomous system to capture a comprehensive understanding of the scene. This view includes information about agents, such as road layouts, lanes, intersections, and the location of objects such as vehicles, pedestrians, and obstacles. A detailed BEV view allows the system to identify potential collision risks, anticipate future actions, predict object trajectories, and plan safe and efficient routes. The information and feature space in the BEV graph can lead to a smoother and more reliable driving experience.
[0021] However, traditional supervised learning methods have focused on aligning BEV representations with limited autonomous training data and task-specific supervisory signals. When trained only on visual inputs and constrained autonomous datasets, models tend to converge to suboptimal states when faced with open-world scenarios. This can lead to inconsistencies when compared to human common sense. Furthermore, while significant progress has been made in computer vision for autonomous driving, a key dimension remains unexplored: the fusion of language understanding with vision-based planning systems.
[0022] Therefore, according to various embodiments described herein, the present disclosure proposes a visual language planning (VLP) based model with agent-by-agent learning to bridge this gap. In this VLP approach, language knowledge is exploited during training through comparative learning with visual model information to improve the planning and generalization capabilities of autonomous driving systems. The purpose of this is to revolutionize the landscape of autonomous driving by seamlessly incorporating language understanding into the planning process. By leveraging the power of language-based models in collaboration with advanced computer vision techniques, the accuracy, safety, and generalization capabilities of autonomous driving systems can be significantly improved.
[0023] The present disclosure provides a per-agent learning strategy to integrate a large language model (LLM) with an autonomous driving system, and to improve the generalization capabilities of a BEV feature map. An autonomous driving system may rely on both visual data (such as sensor input) and contextual information (such as natural language commands or descriptions). LLM acquires extensive world knowledge, including common sense reasoning and contextual awareness. Integrating this contextual knowledge with BEV features enhances the system's ability to correctly interpret the environment. For example, it can help the system understand the meaning of terms such as "slow down when pedestrians cross" and adjust the behavior of the vehicle accordingly. In addition, aligning the BEV feature space with the LLM feature space allows for the effective fusion of these different modalities. It enables the system to combine visual perception with the semantic understanding provided by language to create a more comprehensive representation of the environment. By seamlessly integrating LLM with an autonomous driving model during training, the methods and systems described herein ensure that the autonomous driving system can handle a wide range of scenarios, making it safer, more adaptable, less dependent on specific training data, and better generalized on open set autonomous driving scenarios.
[0024] Machine learning and neural networks are integral to the invention disclosed herein. Figure 1 A system 100 for training a neural network (e.g., a deep neural network) is shown. The system 100 may include an input interface for accessing training data 102 for the neural network. For example, Figure 1As shown, the input interface can be constituted by a data storage interface 104, which can access the training data 102 from a data storage device 106. For example, the data storage interface 104 can be a memory interface or a permanent storage interface, such as a hard disk or SSD interface, but can also be a personal area network, a local area network, or a wide area network interface, such as a Bluetooth, Zigbee or Wi-Fi interface or an Ethernet or fiber optic interface. The data storage device 106 can be an internal data storage device of the system 100, such as a hard disk drive or SSD, or an external data storage device, such as a network accessible data storage device.
[0025] In some embodiments, data storage 106 may further include a data representation 108 of an untrained version of the neural network, which may be accessed by system 100 from data storage 106. However, it should be understood that training data 102 and data representation 108 of the untrained neural network may each be accessed from different data storages, for example, via different subsystems of data storage interface 104. Each subsystem may be a type of data storage interface 104 as described above. In other embodiments, data representation 108 of the untrained neural network may be generated internally by system 100 based on design parameters of the neural network, and thus may not be explicitly stored on data storage 106.
[0026] The system 100 may also include a processor subsystem 110 that may be configured to provide an iterated function as a replacement for a stack of layers of a neural network to be trained during operation of the system 100. Here, the respective layers of the stack of layers being replaced may have mutually shared weights and may receive as input the output of a previous layer, or, for a first layer of the stack of layers, an initial activation and a portion of the input of the stack of layers. The processor subsystem 110 may further be configured to iteratively train the neural network using the training data 102. Here, an iteration of the training of the processor subsystem 110 may include a forward propagation portion and a backward propagation portion. The processor subsystem 110 may be configured to perform the forward propagation portion by, among other operations, defining a forward propagation portion that may be performed, determining an equilibrium point of the iterated function at which the iterated function converges to a fixed point, wherein determining the equilibrium point includes using a numerical root finding algorithm to find a root solution of the iterated function minus its input, and by providing the equilibrium point as a replacement for the output of the stack of layers in the neural network. The system 100 may further include an output interface for outputting a data representation 112 of the trained neural network; this data may also be referred to as trained model data 112. Figure 1As shown, the output interface can be constituted by the data storage interface 104, which in these embodiments is an input / output ("IO") interface, through which the trained model data 112 can be stored in the data storage device 106. For example, the data representation 108 defining the "untrained" neural network can be at least partially replaced by the data representation 112 of the trained neural network during or after training, because the parameters of the neural network, such as the weights, hyperparameters and other types of parameters of the neural network, can be adapted to reflect the training of the training data 102. This is also Figure 1 106, which are shown by reference numerals 108, 112, which refer to the same data record on the data storage device 106. In other embodiments, the data representation 112 may be stored separately from the data representation 108 defining the "untrained" neural network. In some embodiments, the output interface may be separate from the data storage interface 104, but may generally be of the type described above for the data storage interface 104.
[0027] Figure 1 The illustrated system 100 is one example of a system that may be used to train the machine learning models described herein.
[0028] Figure 2 A system 200 that implements the machine learning models described herein, such as the VLP base model, is depicted. System 200 may include at least one computing system 202. Computing system 202 may include at least one processor 204 that is operably connected to a memory unit 208. Processor 204 may include one or more integrated circuits that implement the functionality of a central processing unit (CPU) 206. CPU 206 may be a commercially available processing unit that implements an instruction set such as one of the x86, ARM, Power, or MIPS instruction set families. During operation, CPU 206 may execute stored program instructions retrieved from memory unit 208. The stored program instructions may include software that controls the operation of CPU 206 to perform the operations described herein. In some examples, processor 204 may be a system on a chip (SoC) that integrates the functionality of CPU 206, memory unit 208, network interfaces, and input / output interfaces into a single integrated device. Computing system 202 may implement an operating system for managing various aspects of operation. Although in Figure 2 One processor 204, one CPU 206, and one memory 208 are shown, but of course more than one of each may be utilized throughout the system.
[0029] The memory unit 208 may include volatile memory and non-volatile memory for storing instructions and data. Non-volatile memory may include solid-state memory, such as NAND flash memory, magnetic and optical storage media, or any other suitable data storage device that retains data when the computing system 202 is disabled or loses power. Volatile memory may include static and dynamic random access memory (RAM) that stores program instructions and data. For example, the memory unit 208 may store a machine learning model 210 or algorithm, a training data set 212 for the machine learning model 210, and an original source data set 216.
[0030] The computing system 202 may include a network interface device 222 configured to provide communications with external systems and devices. For example, the network interface device 222 may include a wired and / or wireless Ethernet interface defined by the Institute of Electrical and Electronics Engineers (IEEE) 802.11 family of standards. The network interface device 222 may include a cellular communication interface for communicating with a cellular network (e.g., 3G, 4G, 5G). The network interface device 222 may also be configured to provide a communication interface to an external network 224 or a cloud.
[0031] The external network 224 may be referred to as the World Wide Web or the Internet. The external network 224 may establish a standard communication protocol between computing devices. The external network 224 may allow information and data to be easily exchanged between computing devices and the network. One or more servers 230 may communicate with the external network 224.
[0032] The computing system 202 may include an input / output (I / O) interface 220, which may be configured to provide digital and / or analog input and output. The I / O interface 220 is used to transfer information between an internal storage device and an external input and / or output device (e.g., an HMI device). The I / O 220 interface may include an associated circuit or bus network to transfer information between (multiple) processors and storage devices. For example, the I / O interface 220 may include digital I / O logic lines that can be read or set by (multiple) processors, handshake lines for supervising data transmission via I / O lines, timing and counting facilities, and other structures known to provide such functions. Examples of input devices include keyboards, mice, sensors, touch screens, etc. Examples of output devices include monitors, touch screens, speakers, head-up displays, vehicle control systems, etc. The I / O interface 220 may include additional serial interfaces (e.g., universal serial bus (USB) interfaces) for communicating with external devices. The I / O interface 220 may be referred to as an input interface (because it transmits data from an external input (such as a sensor)), or an output interface (because it transmits data to an external output, such as a display).
[0033] The computing system 202 may include a human-machine interface (HMI) device 218, which may include any device that enables the system 200 to receive control inputs. The computing system 202 may include a display device 232. The computing system 202 may include hardware and software for outputting graphical and textual information to the display device 232. The display device 232 may include an electronic display screen, a projector, a speaker, or other suitable device for displaying information to a user or operator. The computing system 202 may also be configured to allow interaction with a remote HMI and a remote display device via a network interface device 222.
[0034] System 200 can be implemented using one or more computing systems. Although this example depicts a single computing system 202 that implements all of the described features, it is intended that the various features and functions can be separated and implemented by multiple computing units that communicate with each other. The specific system architecture selected may depend on a variety of factors.
[0035] The system 200 can implement a machine learning algorithm 210 configured to analyze a raw source data set 216. The raw source data set 216 can include raw or unprocessed sensor data, which can represent an input data set for a machine learning system. The raw source data set 216 can include video, video clips, images, text-based information, audio or human speech, time series data (e.g., a pressure sensor signal that varies over time), and raw or partially processed sensor data (e.g., a radar map of an object). In some examples, the machine learning algorithm 210 can be a neural network algorithm (e.g., a deep neural network) designed to perform a predetermined function. For example, a neural network algorithm can be configured in an automotive application to recognize street signs or pedestrians in an image. (Multiple) machine learning algorithms 210 may include algorithms configured to operate one or more machine learning models described herein, including a VLP base model.
[0036] The computing system 202 may store a training data set 212 for the machine learning algorithm 210. The training data set 212 may represent a set of previously constructed data used to train the machine learning algorithm 210. The machine learning algorithm 210 may use the training data set 212 to learn weighting factors associated with the neural network algorithm. The training data set 212 may include a set of source data having corresponding outputs or results that the machine learning algorithm 210 attempts to replicate through the learning process. In this example, the training data set 212 may include an input image containing an object (e.g., a street sign). The input image may include various scenes in which objects are identified. The training data set 212 may also include a text description of the scene (e.g., "corresponds to an image detected by a vehicle sensor").
[0037] The machine learning algorithm 210 can operate in a learning mode using the training data set 212 as input. The machine learning algorithm 210 can be executed in multiple iterations using data from the training data set 212. With each iteration, the machine learning algorithm 210 can update the internal weighting factors based on the achieved results. For example, the machine learning algorithm 210 can compare the output results (e.g., the reconstructed or supplemented image in the case where the image data is the input) with those included in the training data set 212. Since the training data set 212 includes the expected results, the machine learning algorithm 210 can determine when the performance is acceptable. After the machine learning algorithm 210 achieves a predetermined performance level (e.g., 100% consistent with the output associated with the training data set 212) or converges, the machine learning algorithm 210 can be executed using data that is not in the training data set 212. It should be understood that in the present disclosure, "convergence" can mean that a set (e.g., predetermined) number of iterations have occurred, or the residual is small enough (e.g., the change in the approximate probability with the iteration is less than a threshold), or other convergence conditions. The trained machine learning algorithm 210 can be applied to a new data set to generate annotated data. In the context of the VLP model described herein, a loss between a predicted trajectory of an autonomous vehicle and an underlying true trajectory of the vehicle may be determined, and the VLP model may be trained to reduce that loss, e.g., to converge.
[0038] The machine learning algorithm 210 can be configured to identify specific features in the raw source data 216. The raw source data 216 may include multiple instances or input data sets that need to supplement the results. For example, the machine learning algorithm 210 can be configured to identify the presence of an agent in a video image, annotate an event, and / or command a vehicle to take a specific action (planning) based on the agent's position data (perception) and the agent's predicted future movement / position (prediction). The machine learning algorithm 210 can be programmed to process the raw source data 216 to identify the presence of specific features. The machine learning algorithm 210 can be configured to identify features in the raw source data 216 as predetermined features (e.g., road signs, pedestrians, etc.). The raw source data 216 can be derived from a variety of sources. For example, the raw source data 216 can be actual input data collected by the machine learning system. The raw source data 216 can be machine generated for testing the system. For example, the raw source data 216 can include raw video images from a camera. Also, as will be further described below with reference to the VLP base model, the original source data 216 may be natural language text information associated with the scene (eg, “a car is entering the intersection from the left”).
[0039] Figure 3A schematic diagram of a control system 302 configured to control a vehicle 300, which may be a partially autonomous vehicle or a fully autonomous vehicle, a partially autonomous robot, or a fully autonomous robot, is depicted. The vehicle 300 and / or its control system 302 may incorporate one or more components of the system 200, such as the computing system 202, to command an actuator 304 to perform a particular action based on processed readings from one or more sensors 306. For example, the control system 302 may be configured to utilize the VLP-based model disclosed herein to control the motion of the vehicle via the actuator 304.
[0040] The one or more sensors 306 may include one or more image sensors (e.g., cameras, video sensors, radar sensors, ultrasonic sensors, LiDAR sensors), and / or location sensors (e.g., GPS). The sensors 306 may be configured to generate the raw source data 216. One or more of the one or more specific sensors may be integrated into the vehicle 300. In the context of agent recognition and processing as described herein, the sensor 306 is a camera mounted to or integrated into the vehicle 300. Alternatively or in addition to the one or more specific sensors described above, the sensor 306 may include a software module that is configured to determine the state of the actuator 304 when executed.
[0041] In embodiments where the vehicle 300 is a fully or partially autonomous vehicle, the actuators 304 may be embodied in the brakes, accelerator, propulsion system, engine, transmission, or steering system (e.g., steering wheel) of the vehicle 300. For example, actuator control commands may be determined such that the actuators 304 are controlled such that the vehicle 300 avoids a collision with a detected agent. The detected agents may also be classified according to what the classifier deems them most likely to be, such as a pedestrian or a tree. The actuator control commands may be determined according to the classification.
[0042] In other embodiments where the vehicle 300 is a fully or partially autonomous robot, the vehicle 300 may be a mobile robot configured to perform one or more functions, such as flying, swimming, diving, and stepping, via the actuators 304. The mobile robot may be an at least partially autonomous lawn mower or an at least partially autonomous cleaning robot. In such embodiments, actuator control commands may be determined such that a propulsion unit, a steering unit, and / or a braking unit of the mobile robot may be controlled such that the mobile robot may avoid collision with an identified object.
[0043] Figure 4 A high-level overview of an end-to-end autonomous driving system 400 is shown, according to one embodiment. The end-to-end system 400 may be incorporated into a vehicle 300, such as its computing system 202, to operate the vehicle to avoid objects or otherwise control the vehicle 300 based on the sensed environment around the vehicle 300.
[0044] Image inputs are received and passed through one or more ML (e.g., neural network) layers to create a BEV that represents the environment surrounding the vehicle. The BEV can be used as input to all three perception, prediction, and planning modules. For example, the perception model utilizes computer vision based on input received from an image sensor (e.g., a BEV) to perform object detection, etc. The prediction model may include a machine learning model that is configured to estimate the trajectory and intent of objects detected in the BEV based on past motion, direction, and contextual information of these objects. The planning model may include route planning, trajectory planning, and decisions taken by the vehicle to navigate relative to other objects in the BEV, and translate these decisions into actions taken by the vehicle in real life.
[0045] The present disclosure introduces a visual language planning (VLP) base model for autonomous driving. In an embodiment, the VLP base model uses contrastive learning techniques, such as those introduced in the contrastive language-image pre-training (CLIP) model. Other contrastive learning models can be used. As an example, an introduction to the CLIP model is provided, followed by a further description of the VLP.
[0046] CLIP was developed by OpenAI. It is designed to understand and connect images and natural language descriptions in a way that allows it to perform a wide range of vision and language tasks. CLIP adopts a dual encoder architecture, which includes a visual encoder and a text encoder, as well as a shared embedding space. The visual encoder processes images, while the text encoder processes natural language descriptions. The visual encoder converts images into fixed-length vector representations based on visual models such as convolutional neural networks (CNN). The text encoder processes text descriptions by converting them into fixed-length vector representations. CLIP is a visual language based model trained on open world data using contrastive learning. Contrastive learning is a type of machine learning in which the model learns to distinguish between positive pairs and negative pairs of data. In the context of CLIP, "positive pairs" consist of semantically related images and text descriptions, while "negative pairs" consist of unrelated images and randomly selected text descriptions. During training, CLIP is designed to encourage features from related text and image pairs to be brought into the common embedding space, while pushing unrelated pairs away.
[0047] CLIP's shared embedding space allows for zero-shot learning. When presented with an image and a text prompt, CLIP can rank how well the image matches the prompt without specific training data for that particular task. CLIP can perform a variety of visual-linguistic tasks, including image classification, text-based image retrieval (e.g., retrieving images based on a text query), image captioning, zero-shot object recognition, and more.
[0048] The contrastive learning concept used in CLIP (whose teaching is included in the VLP base model) is Figure 5 , generally shown as a contrastive learning model at 500. As shown, multiple natural language text descriptions 502 are fed to a text encoder 504, and multiple images 506 are fed to an image encoder 508. The model 500 then performs feature mapping, where the vectors output by the encoders are mapped to a joint embedding space. For example, an image vector (e.g., of size 1×256) output by the image encoder is matched to a corresponding text vector (e.g., of size 1×256) output by the text encoder. The model then performs a dot product between a batch of image and text features to obtain the similarity between these vectors, as shown at 510.
[0049] refer to Figure 5 In the example embodied in FIG. 5 , a plurality of images 506 (in this example, one of which is an image of a tiger) are fed into an image encoder 508, and a plurality of text phrases 502 (in this example, one of which is a phrase like “photo of a tiger”) are fed into a text encoder 504. Several unrelated or dissimilar text phrases and images are also fed into the encoder. For example, images of objects that are not tigers are fed into the image encoder 508, and phrases that are not related to tigers are also fed into the text encoder 504. The image encoder generates text phrases with features I1, I2, ... I N The text encoder generates an image vector with features T1, T2, ...T N The diagonal of the matrix 510 resulting from this dot product shows paired images and text according to their possible similarity, while the off-diagonal lines represent unpaired image and text features (e.g., an image of a cat and a text description such as "picture of a dog").
[0050] Therefore, the contrastive learning model pulls the image and text embeddings together when they correspond to each other, and pushes them apart when they do not. In other words, the reference Figure 5During training, the contrastive learning model aims to increase the similarity of diagonal elements (i.e., positive pairs) while reducing the similarity between off-diagonal elements. As another example, during training, if the model is provided with images of cats and text descriptions such as "pictures of cats", the model aims to minimize the distance (similarity) between the image and text embeddings in the shared space; conversely, if the model is provided with images of cats and text descriptions such as "pictures of dogs", the model aims to maximize the distance (dissimilarity) between their embeddings. This contrastive training objective encourages the model to learn to understand the semantic relationship between images and text. This is a way to teach the model to closely associate matching image-text pairs and effectively distinguish between mismatched pairs. The result is a shared embedding space where similar pairs are clustered together and dissimilar pairs are separated.
[0051] Return to reference Figure 4 , in previous end-to-end autonomous driving systems, only visual cues are used for training of BEV feature extraction. The extracted BEV features are treated as source memory and shared among all downstream task modules, including tracking (e.g., TrackHead), mapping (e.g., MapHead), motion prediction (e.g., MotionHead), occupancy prediction (OccupancyHEad), and planning (PlanHead) for specific information prediction. The end-to-end model is trained by a unified loss that summarizes the losses from all tasks. The whole process can be expressed as follows: bev_feat = BEVEncoder(input visual ) track pred =TrackHead(bev_feat), map pred =MapHead(bev_feat), motion pred =MotionHead(bev_feat) occupancy pred =OccupancyHead(bev_feat), plan pred = PlanHead(bev_feat), Loss=Loss track +Loss map +Loss motion +Loss occ +Loss plan , The input visualrepresents the visual input of the system, bev_feat represents the bird's-eye view feature, task pred and Loss task Represents the prediction and loss of each task respectively, and Loss task Represents the loss of each task, such as tracking loss (Loss track ), mapping loss map ). task pred is the prediction from each task (e.g., track prediction pred ) or mapping prediction (map pred )) rather than explicitly stating how each of them is generalized.
[0052] However, unlike this approach, the present disclosure applies both visual and language cues to enrich BEV information in order to produce a better memory source for the autonomous driving system. In an embodiment of the agent-by-agent learning method, an agent-by-agent statement is formulated for each input to describe the environment, the surrounding agent states, and the situation of the autonomous vehicle itself. The agent-by-agent statement may include the ground truth of each agent's task, the surrounding environment (e.g., each lane), high-level navigation commands, the ground truth of the ego vehicle (i.e., the subject autonomous vehicle), and a scene description. This ground truth metadata is available as part of the training data for a specific image / video scene, and can therefore be generated using a template and given this ground truth information. For example, a sentence associated with a particular agent in a BEV may be: “The object is a {building vehicle}. Its 3d bounding box is {cx, cy, cz, w, l, h, rotation, vx, vy}. Its past-future trajectory is {[[x1, y1], [x2, y2], [x3, y3], [x4, y4], [x5, y5], [x6, y6]]}. Its future trajectory will be {[[x1, y1], [x2, y2], [x3, y3], [x4, y4], [x5, y5], [x6, y6]]}. The scene is located in {Singapore Onenorth}. The scene description is {several moving pedestrians, parked cars, and motorcycles}” The information contained within the {} brackets can be generated from the ground truth data as part of the template that constitutes the rest of the sentence. Therefore, using a similar sentence format with different ground truth data associated with this agent, more agent-by-agent sentences can be generated for other agents in the environment.
[0053] Using these agent-by-agent sentences, the system then performs a contrastive learning model (such as the one described above, including the features described in CLIP) between agent-by-agent text features and agent-by-agent BEV image features to push agents with similar situations closer together in feature space, and push agents with different situations farther apart. This aims to build a consistent feature space that is consistent with human common sense.
[0054] The present disclosure provides for adding this text information in the BEV encoder module of the end-to-end autonomous driving system, such as Figure 6 shown. Figure 6 A schematic diagram of a VLP base model 600 with agent-by-agent learning for an end-to-end autonomous driving system is shown. The perception model 602, prediction model 604, and planning model 606 shown are similar to those described above (e.g., Figure 4 ). Here, Figure 6 As shown, these models are improved by implementing a contrastive learning model 608 that compares the per-agent text features 610 with the per-agent BEV features 612 (also referred to as per-agent image features, since the BEVs are populated via image data as described above).
[0055] BEV can be generated from raw image data. Raw image data can be fused together to generate a bird's-eye view representation of the environment. The generated BEV is then used for downstream tasks such as perception, prediction, and planning. Therefore, for our approach, we utilize BEV features, such as features from BEV data instead of raw image data.
[0056] In one embodiment, as described above, the textual prompt 614 is derived from training metadata. For example, a template may be used and populated with textual data from training data associated with a particular scene at a particular time. As a simple example, a text string such as "The subject agent is a {pedestrian}" may be generated. Its current location is {(x1, y1)}, and its future trajectory is {(x2, y2), (x3, y3), (x4, y4)}. The scene is at the intersection of {BroadStreet} and {Milk Street} in {Boston, Massachusetts}.".
[0057] These textual cues 614 are forwarded to a text encoder 616. The text encoder 616 operates similarly to that described above with respect to contrastive learning, such as the text encoder 504. In an embodiment, the text encoder is configured to extract high-level language features as proxy-by-proxy text features. Since the proxy-by-proxy BEV features are used for all downstream tasks, the information in the proxy-by-proxy BEV features should be aligned with the proxy-by-proxy text features. Therefore, using the proxy-by-proxy text features and the proxy-by-proxy BEV features, contrastive learning is performed at 608 between the two modes to push corresponding visual-linguistic pairs closer and other negative pairs further away, aiming to enhance the feature representation capabilities in a more comprehensive manner. During training, the contrastive loss (Loss contra ) are additionally included in the final loss.
[0058] The results of the comparative learning model improve the BEV features that are input to the perception model 602, prediction model 604, and planning model 606.
[0059] With these additions, the overall process differs from the one explained above and can be expressed as follows: bev_feat = BEVEncoder(input visual ), pairs agent =Agent(bev_feat,input gt-text ), contra pred =ContraHead(pairs agent ), track pred =TrackHead(bev_feat), map pred =MapHead(bev_feat), motion pred =MotionHead(bev__feat), occupancy pred =OccupancyHead(bev_feat), plan pred =PlanHead(bev_feat), Loss=Loss contra +Loss track +Loss map +Loss motion +Loss occ +Loss plain , The input gt-text represents text features, which include the ground truth for each agent, and Agent(.) indicates the module for extracting per-agent features and formulating agent pairs for each input.
[0060] With these teachings, the model described in this paper is configured to use text features to supervise BEV image features during training to produce a better source memory for the system and improve the generalization ability of the model.
[0061] Figure 7A method 700 for training an autonomous driving system using a visual language planning (VLP) machine learning model according to one embodiment is shown. The method can be performed by one or more processors disclosed herein. At 702, image data is generated from a camera mounted to a vehicle. For example, the camera can be one of the sensors 306 described above. The generated image data includes an environment or agent in a scene outside the vehicle.
[0062] At 704, image processing is performed on the image data to detect agents in the environment. Object recognition and classification may be used as described above. At 706, a BEV is generated based on the image data and the results of object recognition or other object detection. The BEV includes BEV features, such as spatiotemporal information associated with the vehicle and the detected agents.
[0063] At 708, an agent-centric visual language planning (VLP) machine learning model is executed. During the execution of the VLP model, at 710, agent-by-agent BEV features are extracted from the BEV. The agent-by-agent BEV features focus on and are associated with the agents in the environment. In other words, the extracted agent-by-agent BEV features correspond to a specific known agent, or multiple known agents. At 712, natural language text prompts are generated. These prompts are associated with agents in the environment. For example, one type of natural language prompt may be "This is a pedestrian crossing the road and will arrive at position (x1, y1) in 3.5 seconds at its current speed." At 714, agent-by-agent text features are extracted from the natural language text prompts. These agent-by-agent text features are associated with corresponding agents in the environment. At 716, a contrastive learning model is executed to derive similarities between agent-by-agent text features and agent-by-agent BEV features. At 718, based on the similarities, the BEV features are refined for use in the perception model, the prediction model, and the planning model to generate a new, modified predicted trajectory for the vehicle.
[0064] Although exemplary embodiments are described above, this does not mean that these embodiments describe all possible forms included in the claims. The words used in the specification are descriptive words, not limiting, and it should be understood that various changes can be made without departing from the spirit and scope of the present disclosure. As mentioned above, the features of the various embodiments can be combined to form further embodiments of the present invention that may not be explicitly described or illustrated. Although various embodiments may have been described as providing advantages or being superior to other embodiments or prior art implementations in terms of one or more desired characteristics, it is recognized by those of ordinary skill in the art that one or more features or characteristics can be compromised to achieve the desired overall system properties, depending on the specific application and implementation. These properties may include, but are not limited to, cost, strength, durability, life cycle cost, marketability, appearance, packaging, size, maintainability, weight, manufacturability, ease of assembly, etc. Therefore, in terms of one or more features, in the sense that any embodiment is described as not as good as other embodiments or prior art to achieve the desired, these embodiments are not outside the scope of the present disclosure and may be desired for specific applications.
Claims
1. A method for training an autonomous driving system using a visual language planning (VLP) machine learning model with agent-by-agent learning, the method comprising: receiving image data generated from a camera mounted to the vehicle, wherein the image data includes an agent in an environment external to the vehicle; Detecting agents in the environment based on image data via image processing; generating a bird's eye view (BEV) of an environment based on the image data, wherein the BEV includes BEV features, the BEV features including spatiotemporal information associated with the vehicle and the detected agent; Input data from the BEV into the perception, prediction, and planning models of the end-to-end autonomous driving system to generate a predicted trajectory for the vehicle; as well as Execute an agent-centric vision-language planning (VLP) machine learning model to: extracting agent-by-agent BEV features from the BEV, wherein the agent-by-agent BEV features are associated with corresponding agents in the environment, Generate natural language text prompts associated with agents in the environment, extracting agent-by-agent text features from natural language text prompts, where the agent-by-agent text features are associated with corresponding agents in the environment, Perform a contrastive learning model to derive similarities between per-agent BEV features and per-agent text features, and The BEV features used in the perception model, prediction model, and planning model are refined based on similarities to generate a modified predicted trajectory for the vehicle.
2. The method according to claim 1, further comprising: determining a loss between the predicted trajectory of the vehicle and the ground truth trajectory of the vehicle; as well as The steps of claim 1 are repeated until convergence to minimize the loss.
3. The method according to claim 1, wherein: The natural language text prompt is generated using ground truth information associated with the environment present in the template and training data.
4. The method of claim 1, wherein the output of contrastive learning is used as a training loss for training a VLP machine learning model.
5. The method according to claim 1, wherein the contrastive learning model comprises: a text encoder configured to output a text-based vector representing agent-by-agent text features associated with a natural language text prompt; as well as An image encoder is configured to output an image-based vector representing agent-by-agent BEV features associated with agents in the BEV.
6. The method of claim 5, wherein the contrastive learning model is further configured to perform a dot product to evaluate the similarity between the text-based vector and the image-based vector.
7. The method according to claim 1, wherein: The contrastive learning model is also configured to push aside dissimilarity between the agent-by-agent BEV features and the agent-by-agent text features.
8. The method according to claim 1, further comprising: Repeat the agent-centric VLP machine learning model until convergence; as well as Based on the convergence, a trained agent-centric VLP machine learning model is output.
9. A system utilizing a visual language planning (VLP) machine learning model, the system comprising: a camera mounted to the vehicle and configured to generate image data associated with an agent in an environment external to the vehicle; processor; as well as A memory comprising instructions which, when executed by a processor, cause the processor to: Processing image data to detect agents in the environment, generating a bird's eye view (BEV) of the environment based on the image data, wherein the BEV includes BEV features including spatiotemporal information associated with the vehicle and the detected agent, Input data from the BEV into the perception, prediction, and planning models of the end-to-end autonomous driving system to generate a predicted trajectory for the vehicle, and Execute an agent-centric vision-language planning (VLP) machine learning model to: (i) extracting agent-by-agent BEV features from the BEVs, where the agent-by-agent BEV features are associated with the corresponding agents in the environment, (ii) generate natural language text prompts associated with agents in the environment, (iii) extracting agent-by-agent text features from natural language text cues, where the agent-by-agent text features are associated with corresponding agents in the environment, (iv) performing a contrastive learning model to derive similarities between per-agent BEV features and per-agent text features, and (v) Refine the BEV features used in the perception model, prediction model, and planning model based on similarity to generate a modified predicted trajectory of the vehicle.
10. The system according to claim 9, wherein: The memory includes further instructions which, when executed by the processor, cause the processor to: determining a loss between a predicted trajectory of the vehicle and a ground truth trajectory of the vehicle; and The VLP model is executed until convergence to minimize the loss.
11. The system according to claim 9, wherein: The natural language text prompt is generated using ground truth information associated with the environment present in the template and training data.
12. The system of claim 9, wherein the output of contrastive learning is used as a training loss for training a VLP machine learning model.
13. The system of claim 9, wherein the contrastive learning model comprises: a text encoder configured to output a text-based vector representing agent-by-agent text features associated with a natural language text prompt; as well as An image encoder is configured to output an image-based vector representing agent-by-agent BEV features associated with agents in the BEV.
14. The system of claim 13, wherein the contrastive learning model is further configured to perform a dot product to evaluate similarity between the text-based vector and the image-based vector.
15. The system of claim 9, wherein the contrastive learning model is further configured to push aside dissimilarity between the agent-by-agent BEV features and the agent-by-agent text features.
16. The system of claim 9, wherein: The memory includes further instructions which, when executed by the processor, cause the processor to: Repeatedly executing the agent-centric VLP machine learning model until convergence; and Based on the convergence, a trained agent-centric VLP machine learning model is output.
17. The system of claim 9, wherein the agent in the environment comprises at least one of a pedestrian, another vehicle, or a bicyclist.
18. An apparatus for training at least one machine learning model, the apparatus comprising: processor; as well as A memory comprising instructions which, when executed by a processor, cause the processor to: Processing image data generated from cameras mounted to the vehicle in order to detect agents in the environment, generating a bird's eye view (BEV) of the environment based on the image data, wherein the BEV includes BEV features including spatiotemporal information associated with the vehicle and the detected agent, Input data from the BEV into the perception, prediction, and planning models of the end-to-end autonomous driving system to generate a predicted trajectory for the vehicle, and Execute an agent-centric vision-language planning (VLP) machine learning model to: (i) extracting agent-by-agent BEV features from the BEVs, where the agent-by-agent BEV features are associated with the corresponding agents in the environment, (ii) generate natural language text prompts associated with agents in the environment, (iii) extracting agent-by-agent text features from natural language text cues, where the agent-by-agent text features are associated with corresponding agents in the environment, (iv) performing a contrastive learning model to derive similarities between per-agent BEV features and per-agent text features, and (v) Refine the BEV features used in the perception model, prediction model, and planning model based on similarity to generate a modified predicted trajectory of the vehicle.
19. The device according to claim 18, wherein: The memory includes further instructions which, when executed by the processor, cause the processor to: determining a loss between the predicted trajectory of the vehicle and the ground truth trajectory of the vehicle; as well as The VLP model is executed until convergence to minimize the loss.
20. The apparatus of claim 18, wherein the contrastive learning model comprises: a text encoder configured to output a text-based vector representing agent-by-agent text features associated with a natural language text prompt; as well as An image encoder is configured to output an image-based vector representing agent-by-agent BEV features associated with agents in the BEV.