System and method for visual language planning (VLP) base model for autonomous driving

By introducing the basic model of visual language planning (VLP) in the autonomous driving system, using comparative learning technology to integrate language and visual information, the problems of error accumulation and language understanding fusion in existing systems are solved, and more efficient planning and generalization capabilities are achieved.

CN119992486APending Publication Date: 2025-05-13ROBERT BOSCH GMBH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411589987.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-10
Filing Date
2024-11-08
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing autonomous driving system has the problem of error accumulation in perception, prediction and planning tasks, and lacks a method to effectively integrate language understanding with vision-based planning systems, which affects the accuracy and generalization capabilities of the system.

Method used

The basic model of visual language planning (VLP) is adopted, and through comparative learning technology, language knowledge is combined with visual model information to generate planning features based on vision and text, and then the vehicle's prediction trajectory is derived.

Benefits of technology

It improves the planning and generalization capabilities of autonomous driving systems, enhances rich perception of the environment, leads to more wise and context-conscious planning decisions, and improves the accuracy and safety of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992486A_ABST
    Figure CN119992486A_ABST
Patent Text Reader

Abstract

Systems and methods are provided for a Visual Language Planning (VLP) base model for autonomous driving. Methods and systems for training an autonomous driving system using a Visual Language Planning (VLP) model. Image data is obtained from a camera mounted on a vehicle, encompassing details about behaviorists located within an external environment. Via image processing, the system identifies the behaviours within the environment. A bird's-eye view (BEV) representation of the surroundings is then generated that encapsulates spatio-temporal information associated with the vehicle and the identified actor. Execution of the VLP machine learning model begins by extracting vision-based planning features from the BEV and receiving or generating textual information characterizing various attributes of the vehicle within the environment. Text-based planning features are extracted from the text information. To enhance model performance, a comparative learning model is used to establish similarities between vision-based and text-based planning features, and a predicted trajectory is output based on the similarities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to systems and methods for a vision-language planning (VLP) based model for autonomous driving. Background Art

[0002] Autonomous vehicles—often referred to as self-driving or driverless vehicles—are a class of vehicles that are able to navigate and operate on roads and in various environments without direct human control. Autonomous vehicles use a combination of advanced technologies and sensors to perceive their surroundings, make decisions, and perform driving tasks.

[0003] Autonomous vehicles are typically equipped with a variety of sensors, including lidar, radar, cameras, ultrasonic sensors, and sometimes additional technologies such as GPS and IMU (inertial measurement units). These sensors provide real-time data about the vehicle's surroundings, including the location of other vehicles, pedestrians, road signs, and road conditions. The vehicle's onboard computer uses the data from the sensors to create a detailed map of the environment and perceive objects and obstacles. This information is critical for navigation and collision avoidance.

[0004] Machine learning (ML) and artificial intelligence (AI) play a vital role in autonomous vehicles. Deep learning algorithms are used for tasks such as object detection, lane keeping, and decision making, and can rely on image processing to perform these tasks. These algorithms enable vehicles to understand and respond to complex and dynamic traffic situations. Summary of the invention

[0005] In one embodiment, a method for training an autonomous driving system using a visual language planning (VLP) machine learning model is provided. The method begins by receiving image data obtained from a camera mounted on a vehicle, covering details about agents located within an external environment. Via image processing, the system identifies these agents within the environment. A bird's eye view (BEV) representation of the surrounding environment is then generated, which encapsulates spatiotemporal information associated with the vehicle and the identified agents. The visual language planning (VLP) machine learning model is then executed, which performs several actions. It extracts vision-based planning features from the BEV, covering spatiotemporal data related to the vehicle's location and nearby agents. In addition, the model generates text information that characterizes various attributes of the vehicle within the environment. The text information is then used to derive text-based planning features. In order to enhance model performance, a contrastive learning model is used to establish similarities between vision-based and text-based planning features. A predicted trajectory for the vehicle is generated based on the similarity.

[0006] In another embodiment, a system utilizing a visual language planning (VLP) machine learning model is provided. The system includes a camera mounted to a vehicle and configured to generate image data associated with an actor in an environment external to the vehicle. The system also includes a processor and a memory, the memory including instructions that, when executed by the processor, cause the processor to perform the following operations: process the image data to detect actors in the environment; generate a bird's eye view (BEV) of the environment based on the image data, wherein the BEV includes spatiotemporal information associated with the vehicle and the detected actor; and execute the visual language planning (VLP) machine learning model to perform the following: extract vision-based planning features from the BEV, wherein the vision-based planning features include at least some spatiotemporal information associated with the vehicle; receive text information associated with the environment, wherein the text information describes the characteristics of the vehicle in the environment; extract text-based planning features from the text information; execute a contrastive learning model to derive similarities between vision-based planning features and text-based planning features; and generate a predicted trajectory for the vehicle based on the similarities.

[0007] In another embodiment, a method for training an autonomous driving system includes the following steps: receiving image data generated from a camera mounted to a vehicle, wherein the image data includes actors in an environment external to the vehicle; generating a bird's eye view (BEV) of the environment based on the image data, wherein the BEV includes spatiotemporal information associated with the vehicle and the actors; based on the BEV, executing a perception model to detect actors in the environment and associated information about the detected actors; based on the BEV, executing a prediction model to estimate a trajectory of the detected actors; based on the BEV, executing a visual language planning (VLP) model to output a predicted trajectory of the vehicle, wherein the VLP model is configured to: extract vision-based planning features from the BEV, wherein the vision-based planning features include spatiotemporal information associated with the vehicle, receive text information associated with the environment, wherein the text information describes characteristics of one or more actors in the environment, extract text-based planning features from the text information, perform contrastive learning to derive similarities between the vision-based planning features and the text-based planning features, and output a predicted trajectory based on the similarities. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 A system for training a neural network according to an embodiment is shown.

[0009] Figure 2 A computer-implemented method for training and utilizing a neural network according to an embodiment is shown.

[0010] Figure 3A schematic diagram of a control system configured to control a vehicle, which may be a partially autonomous vehicle, a fully autonomous vehicle, a partially autonomous robot, or a fully autonomous robot, is shown according to an embodiment.

[0011] Figure 4 A schematic overview of an end-to-end autonomous driving system according to an embodiment is shown.

[0012] Figure 5 A schematic overview of a contrastive learning model according to an embodiment is shown.

[0013] Figure 6 Shown is a schematic overview of a planning model used in an end-to-end autonomous driving system according to an embodiment.

[0014] Figure 7 A schematic overview of a Visual Language Planning (VLP) base model according to an embodiment is shown.

[0015] Figure 8 A schematic diagram of a method for training an autonomous driving system using a visual language planning (VLP) machine learning model according to an embodiment is shown. DETAILED DESCRIPTION

[0016] Embodiments of the present disclosure are described herein. However, it should be understood that the disclosed embodiments are merely examples, and other embodiments may take various and alternative forms. The figures are not necessarily to scale; some features may be enlarged or reduced to show the details of a particular component. Therefore, the specific structural and functional details disclosed herein should not be interpreted as restrictive, but merely as a representative basis for teaching those skilled in the art to adopt the embodiments in various ways. As will be understood by those of ordinary skill in the art, the various features illustrated and described with reference to any of the figures may be combined with the features illustrated in one or more other figures to produce embodiments that are not explicitly illustrated or described. The combination of the illustrated features provides representative embodiments of typical applications. However, for a particular application or implementation, various combinations and modifications of features consistent with the teachings of the present disclosure may be desired.

[0017] As used herein, "a", "an", and "the" refer to both singular and plural referents unless the context clearly indicates otherwise. By way of example, a "processor" programmed to perform various functions refers to one processor programmed to perform each function, or more than one processor that are collectively programmed to perform each of the various functions.

[0018] In the context of an autonomous vehicle, the term "actor" may refer to objects or entities in the environment surrounding the autonomous vehicle or with which it interacts. This includes pedestrians, other vehicles, cyclists, road signs, traffic lights, lane markings, etc. Objects or features detected by the autonomous vehicle's sensors that are used to control the autonomous vehicle's decision making may be collectively referred to as actors.

[0019] The present disclosure incorporates by reference in its entirety U.S. patent application Ser. No. 18 / 388,601, filed on Nov. 10, 2023, and entitled “VISON-LANGUAGE-PLANNING (VLP) MODELS WITH AGENT-WISE LEARNING FOR AUTONOMOUSDRIVING,” with agent docket number 097182-00293.

[0020] The rapid development of autonomous driving technology has ushered in a new era of transportation, promising safer and more efficient journeys. Autonomous driving systems generally involve three high-level tasks: (1) perception, (2) prediction, and (3) planning. Perception involves the ability of a vehicle to understand and interpret its environment. This task includes various subcomponents such as computer vision, sensor fusion, and localization. Key elements of perception include object detection (e.g., identification and tracking of actors external to the autonomous vehicle), localization (e.g., determining the precise position and orientation of the vehicle in the world, typically using GPS and other sensors), and sensor fusion (e.g., combining data from different sensors such as cameras, lidar, radar, and ultrasonic sensors to build a comprehensive view of the surrounding environment). Prediction involves anticipating how other road users and actors in the environment will behave in the near future. This task typically involves using machine learning models to estimate the trajectories and intentions of actors (including pedestrians, other vehicles, and potential obstacles). Accurate prediction is essential for making safe driving decisions. Planning involves determining the best paths and actions for an autonomous vehicle to navigate its environment. This typically includes tasks such as route planning, trajectory planning, and decision making. Planning systems consider information from perception and prediction to make decisions such as when to change lanes, when to stop at intersections, how to react to unexpected events, and so on.

[0021] Conventional approaches for autonomous driving use independent models where each task (perception, prediction, and planning) is trained and optimized separately. However, such disjoint training and optimization can lead to severe error accumulation. To address this issue, end-to-end autonomous driving systems have been proposed and gained interest in recent years. End-to-end autonomous driving systems unify all these tasks and perform joint optimization with the goal of facilitating and improving planning. In particular, end-to-end approaches utilize a bird's-eye view (BEV) representation of all tasks in the perception, prediction, and planning modules. The BEV is generated from multi-view camera inputs and contains spatiotemporal information about the scene. Computer vision systems (e.g., Figure 2 The cameras, processors, memory, and machine learning models shown in Figure 1 can derive spatiotemporal information about the scene. Joint training and optimization strategies across all tasks in end-to-end autonomous driving have led to state-of-the-art results for autonomous driving. Figure 4 (More on that below) shows an overview of an end-to-end autonomous driving system. Central to the success of autonomous vehicles is the integration of diverse modalities that synergistically enhance perception, decision making, and planning capabilities. In the Unified Autonomous Driving (UniAD) model—one of the most recent end-to-end autonomous driving works—autonomy is structured by redefining the interactions between basic components and tasks. In UniAD, the focus shifts to the pursuit of optimization towards the ultimate goal: efficient planning of the autonomous vehicle. This requires revisiting the core components of perception and prediction and reconfiguring their roles to sync with the overall planning goal.

[0022] While significant progress has been made in computer vision for autonomous driving, one key dimension remains unexplored: the fusion of language understanding with vision-based planning systems. Base models — which are large pre-trained machine learning models trained on open-world data — typically involve language as one of the main modalities of the data. In base models, there is often a connection between language and other modalities. After pre-training, the base model can be adapted to a given task via fine-tuning. Base models have shown the importance of incorporating language in achieving state-of-the-art performance and generalization across a wide variety of tasks. Despite the great success of base models across different domains, its extension to the autonomous driving domain remains unknown.

[0023] Therefore, according to various embodiments disclosed herein, the present disclosure proposes a visual language planning (VLP) based model to bridge this gap. In this VLP method, language knowledge is exploited during training through comparative learning with visual model information to improve the planning and generalization capabilities of autonomous driving systems. The purpose of this is to revolutionize the prospects of autonomous driving by seamlessly incorporating language understanding into the planning process. By leveraging the capabilities of language-based models in conjunction with advanced computer vision techniques, the accuracy, safety, and generalization capabilities of autonomous driving systems can be significantly improved.

[0024] The integration of language understanding capabilities offers great potential for enhancing future trajectory predictions for autonomous vehicles. Language models—with their extraordinary capabilities in understanding and generating textual content—provide a way to interpret complex and subtle contextual cues. As disclosed herein, this capability is used to extract high-level semantic information from textual instructions, road signs, and other language cues present in the environment. Thus, the disclosed vision-language-planning foundation model will provide autonomous vehicles with a rich perception of their surroundings, leading to more intelligent and context-aware planning decisions.

[0025] Machine learning and neural networks are integral to the invention disclosed herein. Figure 1 A system 100 for training a neural network (e.g., a deep neural network) is shown. The system 100 may include an input interface for accessing training data 102 for the neural network. For example, Figure 1 As shown in FIG. 1 , the input interface may be constituted by a data storage interface 104, which may access the training data 102 from a data storage device 106. For example, the data storage interface 104 may be a memory interface or a permanent storage interface, such as a hard disk or SSD interface, but may also be a personal, local area network or wide area network interface, such as a Bluetooth, Zigbee or Wi-Fi interface or an Ethernet or fiber optic interface. The data storage device 106 may be an internal data storage device of the system 100, such as a hard disk drive or SSD, but may also be an external data storage device, such as a network accessible data storage device.

[0026] In some embodiments, the data storage device 106 may also include a data representation 108 of an untrained version of the neural network, which may be accessed by the system 100 from the data storage device 106. However, it will be appreciated that the training data 102 and the data representation 108 of the untrained neural network may also each be accessed from a different data storage device, such as via a different subsystem of the data storage interface 104. Each subsystem may be of a type as described above with respect to the data storage interface 104. In other embodiments, the data representation 108 of the untrained neural network may be generated internally by the system 100 based on the design parameters of the neural network, and thus may not be explicitly stored on the data storage device 106.

[0027] The system 100 may also include a processor subsystem 110 that may be configured to provide an iterated function as a replacement for a stack of layers of a neural network to be trained during operation of the system 100. Here, the layers of the replaced stack may have mutually shared weights and may receive the output of a previous layer as input, or, for the first layer of the stack, receive an initial activation and a portion of the input of the stack. The processor subsystem 110 may also be configured to iteratively train the neural network using the training data 102. Here, the training iterations of the processor subsystem 110 may include a forward propagation portion and a backward propagation portion. The processor subsystem 110 may be configured to perform the forward propagation portion by, in addition to other operations that may be performed to define the forward propagation portion: determining an iterated function equilibrium point at which the iterated function converges to a fixed point, wherein determining the equilibrium point includes using a numerical root finding algorithm to find a root solution of the iterated function minus its input, and by providing the equilibrium point as a replacement for the output of the stack in the neural network. The system 100 may also include an output interface for outputting a data representation 112 of the trained neural network, which may also be referred to as trained model data 112. For example, Figure 1 , the output interface may be comprised of a data storage interface 104, wherein in these embodiments, the interface is an input / output ("IO") interface via which trained model data 112 may be stored in a data storage device 106. For example, a data representation 108 defining an "untrained" neural network may be at least partially replaced by a data representation 112 of a trained neural network during or after training, since the parameters of the neural network (such as weights, hyperparameters, and other types of parameters of the neural network) may be adapted to reflect the training on the training data 102. This is also Figure 1106, which are illustrated by reference numerals 108, 112, which refer to the same data record on the data storage device 106. In other embodiments, the data representation 112 may be stored separately from the data representation 108 defining the "untrained" neural network. In some embodiments, the output interface may be separate from the data storage interface 104, but may generally be of the type described above for the data storage interface 104.

[0028] Figure 1 The system 100 shown in FIG. 1 is one example of a system that can be used to train the machine learning models described herein.

[0029] Figure 2 A system 200 that implements a machine learning model (e.g., a VLP base model) described herein is depicted. System 200 may include at least one computing system 202. Computing system 202 may include at least one processor 204 operatively connected to a memory unit 208. Processor 204 may include one or more integrated circuits that implement the functionality of a central processing unit (CPU) 206. CPU 206 may be a commercially available processing unit that implements an instruction set such as one of the x86, ARM, Power, or MIPS instruction set families. During operation, CPU 206 may execute stored program instructions retrieved from memory unit 208. Stored program instructions may include software that controls the operation of CPU 206 to perform the operations described herein. In some examples, processor 204 may be a system on a chip (SoC) that integrates the functionality of CPU 206, memory unit 208, network interfaces, and input / output interfaces into a single integrated device. Computing system 202 may implement an operating system for managing various aspects of operation. Although in Figure 2 One processor 204, one CPU 206, and one memory 208 are shown, but of course more than one of each may be utilized throughout the system.

[0030] The memory unit 208 may include volatile memory and non-volatile memory for storing instructions and data. Non-volatile memory may include solid-state memory, such as NAND flash memory, magnetic and optical storage media, or any other suitable data storage device that retains data when the computing system 202 is deactivated or powered off. Volatile memory may include static and dynamic random access memory (RAM) that stores program instructions and data. For example, the memory unit 208 may store a machine learning model 210 or algorithm, a training data set 212 for the machine learning model 210, and an original source data set 216.

[0031] The computing system 202 may include a network interface device 222 configured to provide communications with external systems and devices. For example, the network interface device 222 may include a wired and / or wireless Ethernet interface defined by the Institute of Electrical and Electronics Engineers (IEEE) 802.11 family of standards. The network interface device 222 may include a cellular communication interface for communicating with a cellular network (e.g., 3G, 4G, 5G). The network interface device 222 may also be configured to provide a communication interface to an external network 224 or a cloud.

[0032] The external network 224 may be referred to as the World Wide Web or the Internet. The external network 224 may establish a standard communication protocol between computing devices. The external network 224 may allow information and data to be easily exchanged between computing devices and the network. One or more servers 230 may communicate with the external network 224.

[0033] The computing system 202 may include an input / output (I / O) interface 220, which may be configured to provide digital and / or analog input and output. The I / O interface 220 is used to transmit information between an internal storage device and an external input and / or output device (e.g., an HMI device). The I / O 220 interface may include an associated circuit or bus network to transmit information to (one or more) processors and storage devices and between them. For example, the I / O interface 220 may include a digital I / O logic line that can be read or set by (one or more) processors, a handshake line for supervising data transmission via the I / O line, a timing and counting facility, and other structures known to provide such functions. Examples of input devices include keyboards, mice, sensors, touch screens, etc. Examples of output devices include monitors, touch screens, speakers, head-up displays, vehicle control systems, etc. The I / O interface 220 may include an additional serial interface (e.g., a universal serial bus (USB) interface) for communicating with external devices. The I / O interface 220 may be referred to as an input interface (because it transmits data from an external input such as a sensor), or an output interface (because it transmits data to an external output such as a display).

[0034] The computing system 202 may include a human-machine interface (HMI) device 218, which may include any device that enables the system 200 to receive control inputs. The computing system 202 may include a display device 232. The computing system 202 may include hardware and software for outputting graphical and textual information to the display device 232. The display device 232 may include an electronic display screen, a projector, a speaker, or other suitable device for displaying information to a user or operator. The computing system 202 may also be configured to allow interaction with a remote HMI and a remote display device via a network interface device 222.

[0035] System 200 may be implemented using one or more computing systems. Although the example depicts a single computing system 202 that implements all described features, it is intended that the various features and functions may be separated and implemented by multiple computing units that communicate with each other. The specific system architecture selected may depend on a variety of factors.

[0036] The system 200 may implement a machine learning algorithm 210 configured to analyze a raw source data set 216. The raw source data set 216 may include raw or unprocessed sensor data, which may represent an input data set for a machine learning system. The raw source data set 216 may include video, video clips, images, text-based information, audio or human speech, time series data (e.g., pressure sensor signals over time), and raw or partially processed sensor data (e.g., a radar map of an object). In some examples, the machine learning algorithm 210 may be a neural network algorithm (e.g., a deep neural network) designed to perform a predetermined function. For example, a neural network algorithm may be configured in an automotive application to identify street signs or pedestrians in an image. The machine learning algorithm(s) 210 may include an algorithm configured to operate one or more machine learning models described herein (including a VLP base model).

[0037] The computing system 202 may store a training data set 212 for the machine learning algorithm 210. The training data set 212 may represent a previously constructed data set for training the machine learning algorithm 210. The machine learning algorithm 210 may use the training data set 212 to learn weighting factors associated with the neural network algorithm. The training data set 212 may include a source data set having a corresponding achievement or result that the machine learning algorithm 210 attempts to replicate via the learning process. In this example, the training data set 212 may include an input image containing an object (e.g., a street sign). The input image may include various scenes in which the object is identified. The training data set 212 may also include a text description of the scene corresponding to the image detected by the vehicle sensor (e.g., "a pedestrian is crossing the street").

[0038] The machine learning algorithm 210 can operate in a learning mode using the training data set 212 as input. The machine learning algorithm 210 can perform multiple iterations using data from the training data set 212. With each iteration, the machine learning algorithm 210 can update the internal weighting factors based on the results obtained. For example, the machine learning algorithm 210 can compare the output results (e.g., the reconstructed or supplemented image in the case where the image data is the input) with those included in the training data set 212. Since the training data set 212 includes the expected results, the machine learning algorithm 210 can determine when the performance is acceptable. After the machine learning algorithm 210 reaches a predetermined performance level (e.g., 100% consistent with the results associated with the training data set 212) or converges, the machine learning algorithm 210 can be executed using data that is not in the training data set 212. It should be understood that in the present disclosure, "convergence" can mean that a set (e.g., predetermined) number of iterations have occurred, or the residual is small enough (e.g., the change in the approximate probability in the iteration is changing by less than a threshold), or other convergence conditions. The trained machine learning algorithm 210 can be applied to a new data set to generate annotated data. In the context of the VLP model described herein, a loss between a predicted trajectory of an autonomous vehicle and a ground truth trajectory of the vehicle may be determined, and the VLP model may be trained to reduce the loss, e.g., converge.

[0039] The machine learning algorithm 210 may be configured to identify specific features in the raw source data 216. The raw source data 216 may include multiple instances or input data sets for which supplementary results are desired. For example, the machine learning algorithm 210 may be configured to identify the presence of an actor in a video image, annotate an event, and / or command a vehicle to take a specific action (planning) based on the actor's position data (perception) and the actor's predicted future movement / position (prediction). The machine learning algorithm 210 may be programmed to process the raw source data 216 to identify the presence of specific features. The machine learning algorithm 210 may be configured to identify features in the raw source data 216 as predetermined features (e.g., road signs, pedestrians, etc.). The raw source data 216 may be derived from a variety of sources. For example, the raw source data 216 may be actual input data collected by the machine learning system. The raw source data 216 may be machine generated for testing the system. As an example, the raw source data 216 may include raw video images from a camera. Also, as will be further described below with reference to the VLP base model, the original source data 216 may be natural language text information associated with the scene (eg, “a car is entering the intersection from the left”).

[0040] Figure 3A schematic diagram of a control system 302 configured to control a vehicle 300, which may be a partially autonomous vehicle or a fully autonomous vehicle, a partially autonomous robot, or a fully autonomous robot, is depicted. The vehicle 300 and / or its control system 302 may incorporate one or more components of the system 200, such as the computing system 202, to command an actuator 304 to perform a particular action based on processed readings from one or more sensors 306. For example, the control system 302 may be configured to utilize the VLP basis model disclosed herein to control movement of the vehicle via the actuator 304.

[0041] The one or more sensors 306 may include one or more image sensors (e.g., cameras, video sensors, radar sensors, ultrasonic sensors, lidar sensors), and / or location sensors (e.g., GPS). The sensors 306 may be configured to generate raw source data 216. One or more of the one or more specific sensors may be integrated into the vehicle 300. In the context of actor identification and processing as described herein, the sensor 306 is a camera mounted to or integrated into the vehicle 300. Alternatively or in addition to the one or more specific sensors identified above, the sensor 306 may include a software module that is configured to determine the state of the actuator 304 when executed.

[0042] In embodiments where vehicle 300 is a fully or partially autonomous vehicle, actuator 304 may be embodied in a brake, accelerator, propulsion system, engine, drive train, or steering system (e.g., steering wheel) of vehicle 300. For example, actuator control commands may be determined such that actuator 304 is controlled such that vehicle 300 avoids colliding with a detected actor. Detected actors may also be classified according to what the classifier deems them most likely to be, such as a pedestrian or a tree. Actuator control commands may be determined depending on the classification.

[0043] In other embodiments where the vehicle 300 is a fully or partially autonomous robot, the vehicle 300 may be a mobile robot configured to perform one or more functions, such as flying, swimming, diving, and walking, via the actuators 304. The mobile robot may be an at least partially autonomous lawn mower or an at least partially autonomous cleaning robot. In such embodiments, actuator control commands may be determined such that a propulsion unit, a steering unit, and / or a braking unit of the mobile robot may be controlled such that the mobile robot may avoid colliding with the identified object.

[0044] Figure 4A high-level overview of an end-to-end autonomous driving system 400 is illustrated, according to an embodiment. The end-to-end system 400 may be incorporated into a vehicle 300, such as its computing system 202, to operate the vehicle to avoid objects or otherwise control the vehicle 300 based on the environment sensed around the vehicle 300.

[0045] Image inputs are received and passed through one or more ML (e.g., neural network) layers to create a BEV that represents the environment around the vehicle. The BEV can be used as input to all three perception, prediction, and planning modules. For example, the perception model utilizes computer vision based on input received from an image sensor (e.g., a BEV) to perform object detection, etc. The prediction model may include a machine learning model that is configured to estimate the trajectory and intention of objects detected in the BEV based on the past movement, direction, and contextual information of those objects. The planning model may include route planning, trajectory planning, and decision making for the vehicle to navigate relative to other objects in the BEV, and translate those decisions into actions taken by the vehicle in real life.

[0046] The present disclosure introduces a visual language planning (VLP) base model for autonomous driving. In an embodiment, the VLP base model uses contrastive learning techniques, such as those introduced in the contrastive language-image pre-training (CLIP) model. Other contrastive learning models can be used. As an example, an introduction to the CLIP model is provided, and then followed by a further description of the VLP.

[0047] CLIP was developed by OpenAI. It is designed to understand and connect images and natural language descriptions in a way that allows it to perform a wide range of vision and language tasks. CLIP uses a dual encoder architecture - including a visual encoder and a text encoder, and a shared embedding space. The visual encoder processes images, while the text encoder processes natural language descriptions. The visual encoder converts images into fixed-length vector representations based on visual models such as convolutional neural networks (CNNs). The text encoder processes text descriptions by converting them into fixed-length vector representations. CLIP is a vision-language based model that is trained on open-world data using contrastive learning. Contrastive learning is a class of machine learning in which the model learns to distinguish between positive and negative pairs of data. In the context of CLIP, a "positive pair" consists of a semantically related image and text description, while a "negative pair" consists of an unrelated image and a randomly selected text description. During training, CLIP is designed to encourage features from related text and image pairs to be pulled together into a common embedding space, while pushing unrelated pairs away.

[0048] CLIP's shared embedding space allows for zero-shot learning. When presented with an image and a text prompt, CLIP can rank how well the image matches the prompt without specific training data for that particular task. CLIP can perform a variety of visual-linguistic tasks, including image classification, text-based image retrieval (e.g., retrieving images based on a text query), image captioning, zero-shot object recognition, and more.

[0049] The contrastive learning concept used in CLIP (whose teaching is included in the VLP base model) is Figure 5 , a contrastive learning model is shown generally at 500. As shown, a plurality of natural language text descriptions 502 are fed into a text encoder 504, and a plurality of images 506 are fed into an image encoder 508. The model 500 then performs feature mapping, where the vectors output by the encoder are mapped to a joint embedding space. For example, an image vector (e.g., of size 1×256) output by the image encoder is matched to a corresponding text vector (e.g., of size 1×256) output by the text encoder. The model then performs a dot product between a batch of image and text features to obtain the similarity between these vectors, generally shown at 510.

[0050] refer to Figure 5 In the example embodied in FIG. 5 , a plurality of images 506 (one of which is an image of a tiger in this example) are fed into an image encoder 508, and a plurality of text phrases 502 (one of which is something like “photos of tigers” in this example) are fed into a text encoder 504. Several unrelated or dissimilar text phrases and images are also fed into the encoder. For example, images of objects that are not tigers are fed into the image encoder 508, and phrases that are not related to tigers are also fed into the text encoder 504. The image encoder generates a textual ... N , and the text encoder produces an image vector with features T1, T2, …T N The diagonal of the matrix 510 resulting from this dot product shows paired images and text according to their possible similarity, while the off-diagonal lines represent unpaired image and text features (e.g., an image of a cat and a text description such as "picture of a dog").

[0051] As such, the contrastive learning model pulls image and text embeddings together when they correspond to each other, and pushes them apart when they do not. Figure 5During training, the contrastive learning model aims to increase the similarity of diagonal elements (i.e., positive pairs) while reducing the similarity between non-diagonal elements. As another example, during training, if the model is provided with images of cats and text descriptions such as "pictures of cats", the model aims to minimize the distance (similarity) between the image and text embeddings in the shared space; conversely, if the model is provided with images of cats and text descriptions such as "pictures of dogs", the model aims to maximize the distance (dissimilarity) between their embeddings. This contrastive training objective encourages the model to learn to understand the semantic relationship between images and text. This is a way to teach the model to closely associate matching image-text pairs and effectively distinguish between mismatched pairs. The result is a shared embedding space where similar pairs are clustered together and dissimilar pairs are far apart.

[0052] Figure 6 The diagram shows a Figure 4 A high-level overview of the planning model 600 in the end-to-end autonomous driving system 400 of FIG. 4 is shown. As illustrated, the planning features are used to predict the future trajectory of the autonomous vehicle. For example, the planning query 602 (also referred to as the planning model) may contain both perception information and prediction information, such as Figure 4 . The planning query is a concatenation of information from the perception module, the prediction module, and the high-level command. The planning query is configured to extract planning information by utilizing the BEV feature map. In other words, the planning model plans an appropriate route for the vehicle using the BEV and all the data on it determined from the perception and prediction models. To this end, the features of the BEV are sent to the planning decoder 604 to obtain visual planning features 606, also known as vision-based planning features. For example, a 1×256 vector can be derived. The planning decoder can be a model that extracts visual planning features using planning queries and BEV features. The features extracted from the BEV may include information about the actors in the environment, such as their positions, their trajectories, etc., which, as described above, are inputs to determine how the autonomous vehicle itself should react. The extracted planning features are further sent to the trajectory regression head 608 to plan the future trajectory of the autonomous vehicle in the next P timestamps. The trajectory regression head 608 can be used to transform the vector into a vector of another size. In an embodiment, the trajectory regression head 608 includes a small neural network that maps high-dimensional inputs to the expected P timestamp trajectories. During the training process, it applies two planning losses to optimize the planning module. One is the average distance error (ADE), which aims to reduce the distance or loss between the predicted trajectory 610 and the ground truth trajectory 612, and the other is the collision rate (COL), which aims to ensure the safety of the planned trajectory.

[0053] In general, the planning model 600 can be represented as follows:

[0054] plan_feat visual=PlanDecoder(plan_query, bev_feat),

[0055] plan pred =TrajRegHead(plan_feat visual ),

[0056]

[0057] Where plan_feat visual represents visual planning features, bev_feat represents bird's-eye view features, plan pred Represents the predicted trajectory, TrajRegHead indicates the trajectory regression head, plan gt represents the ground truth trajectory, and Indicates the location and occupancy of actors around future autonomous vehicles.

[0058] As explained above, the present disclosure proposes to apply both visual and language clues to enrich planning information in order to produce better future trajectories for autonomous vehicles. Therefore, according to an embodiment, the proposed VLP base model includes adding formulaic sentences or phrases to describe the environment and / or situation of the autonomous vehicle and its surrounding actors. For example, the phrases or sentences can be text-based. The phrases or sentences may include high-level ground truth metadata, such as navigation commands, ground truth trajectories, scene descriptions, etc., and at this stage, such ground truth metadata may be obtained as part of the training data. For example, the language input may be: text_prompt = "The autonomous autonomous vehicle is {going straight} in the urban area located at {One-North Singapore}. The scene description is {several moving pedestrians, parked cars, and motorcycles}". And, because humans with knowledge of the training data know the future trajectory of the autonomous vehicle, the language input can also include, for example, "The future trajectory of the autonomous self-driving car in the next 6 timestamps will be {[[x1, y1], [x2, y2], [x3, y3], [x4, y4], [x5, y5], [x6, y6]]}", where x and y represent the position coordinates of the vehicle.

[0059] Figure 7 An example of a visual language planning (VLP) base model 700 according to an embodiment is shown. For example, the VLP base model 700 may be incorporated into Figure 4 In the end-to-end model 400 of VLP base model 700, the planning query 602, planning decoder 604, visual planning features 606, trajectory regression head 608, predicted trajectory 610 and basic truth trajectory 612 are similar to those in the above reference. Figure 6 Those described. Figure 6Compared with the planning model in FIG. 1 , the planning model 700 is different in that it includes a text input 702, a text encoder 704, and a text planning feature 706, wherein contrastive learning occurs between the text planning feature 706 and the visual planning feature 606, which can be referred to above. Figure 5 The method described above is performed (e.g., CLIP). For example, a text phrase can be constructed using a template and ground truth information, such as a ground truth trajectory or high-level command present in the training data. This can be used as text input 702. For example, "The autonomous vehicle is turning left, and its future trajectory is (x1, y1), ... (x6, y6)" can be generated based on the ground truth information related to the scene that already exists in the training data; the format of the natural language sentence can be based on the template. Several text entries can be provided for a corresponding number of videos or images. As more autonomous driving-related and detailed information is included in the sentence, the language path can provide more advanced semantics and comprehensive clues to the planning module. And, as mentioned above, Figure 5 As explained, for training purposes, several unrelated text strings may also be provided.

[0060] The text encoder 704 is configured to operate similarly to the text encoder 504. In an embodiment, the text encoder 704 is configured to extract high-level language features as planning text features 706, also referred to as text-based planning features.

[0061] Since the visual planning features are used for future trajectory prediction, the information in the visual planning features 606 should be aligned with the textual planning features 706. Therefore, using the planning text features and the planning visual features, the system performs contrastive learning between the two modes. Contrastive learning is used to compare the visual planning features 606 with the textual planning features 706. As explained above, contrastive learning utilizes a shared embedding space, and the model learns to distinguish between positive and negative pairs of data. In the context of CLIP and other contrastive learning models, a "positive pair" consists of a semantically related image and text description, while a "negative pair" consists of an irrelevant image and a randomly selected text description. During training, the contrastive learning model is designed to encourage features from related text and image pairs to be pulled together into a common embedding space, while pushing irrelevant pairs away. During the training process, the contrastive loss (ContraLoss) is additionally included in the final loss. This results in a closer and improved relationship between the predicted trajectory 610 and the basic truth trajectory 612.

[0062] Figure 7 The process shown in can be expressed as follows:

[0063] plan_feat text =TextEncoder(text_prompt),

[0064] plan_feat visual = PlanDecoder(plan_query,bev_feat), and

[0065]

[0066] Where plan_feat text Represents the text features used for planning. In VLP, text features supervise planning features during training to achieve better prediction trajectories and generalization performance for autonomous vehicles.

[0067] Figure 8 A method 800 for training an autonomous driving system using a visual language planning (VLP) machine learning model according to an embodiment is illustrated. The method may be performed by one or more processors disclosed herein. At 802, image data is generated from a camera mounted to a vehicle. For example, the camera may be one of the sensors 306 described above. The generated image data includes an environment or actors in a scene external to the vehicle.

[0068] At 804, image processing is performed on the image data to detect actors in the environment. As explained above, object recognition and classification may be used. At 806, a BEV is generated based on the image data and the results of object recognition or other object detection. The BEV includes spatiotemporal information associated with the vehicle and the detected actors.

[0069] At 808, a visual language planning (VLP) machine learning model is executed. During execution of the VLP model, at 810, vision-based planning features are extracted from the BEV. These planning features include spatiotemporal information associated with the vehicle. At 812, text information associated with the environment is generated. For example, a template can be used together with basic truth data associated with the environment to fill in text strings, sentences, etc. The text describes the characteristics of the vehicle in the environment, such as "the vehicle is turning left." At 814, text-based planning features are extracted from the text information. Then, at 816, a contrastive learning model is executed to derive similarities between vision-based planning features and text-based planning features. At 818, based on these similarities, a predicted trajectory of the vehicle is generated based on the vision-based planning features. The model is then refined, updated, and trained again, and more feature space is determined based on the similarities.

[0070] Although exemplary embodiments are described above, it is not intended that these embodiments describe all possible forms covered by the claims. The words used in the specification are descriptive rather than restrictive words, and it should be understood that various changes can be made without departing from the spirit and scope of the present disclosure. As previously described, the features of the various embodiments can be combined to form other embodiments of the present invention that may not be explicitly described or illustrated. Although various embodiments may have been described as providing advantages or preferred over other embodiments or prior art implementations in terms of one or more desired characteristics, it is recognized by those of ordinary skill in the art that one or more features or characteristics can be compromised to achieve the desired overall system properties, depending on the specific application and implementation. These properties may include, but are not limited to, cost, strength, durability, life cycle cost, marketability, appearance, packaging, size, applicability, weight, manufacturability, ease of assembly, etc. As such, to the extent that any embodiment is described as less desirable than other embodiments or prior art implementations in terms of one or more characteristics, these embodiments are not outside the scope of the present disclosure and may be desirable for a particular application.

Claims

1. A method for training an autonomous driving system using a visual language planning (VLP) machine learning model, the method comprising: receiving image data generated from a camera mounted to the vehicle, wherein the image data includes actors in an environment external to the vehicle; Detecting actors in the environment based on image data through image processing; generating a bird's eye view (BEV) of an environment based on the image data, wherein the BEV includes spatiotemporal information associated with vehicles and detected actors; as well as Execute a Vision-Language Planning (VLP) machine learning model to: extracting vision-based planning features from the BEV, wherein the vision-based planning features include spatiotemporal information associated with the vehicle, generating text information associated with the environment, wherein the text information describes characteristics of the vehicle in the environment; extracting text-based planning features from the text information, Perform a contrastive learning model to derive similarities between vision-based planning features and text-based planning features, and A predicted trajectory of the vehicle is generated based on the similarity.

2. The method according to claim 1, further comprising: determining a loss between a predicted trajectory of the vehicle and a ground truth trajectory of the vehicle; and The steps of claim 1 are repeated until convergence to minimize the loss.

3. The method according to claim 1, wherein the text information is generated using ground truth information associated with the environment present in the template and training data.

4. The method of claim 1, wherein the output of the contrastive learning is used as a training loss for training a VLP machine learning model.

5. The method according to claim 1, wherein the contrastive learning model comprises: a text encoder configured to output a text-based vector representing text-based features associated with textual information of an environment; and An image encoder is configured to output an image-based vector representing image-based features associated with an actor detected in the BEV.

6. The method of claim 5, wherein the contrastive learning model is further configured to perform a dot product to evaluate the similarity between the text-based vector and the image-based vector.

7. The method of claim 1, wherein the contrastive learning model is further configured to push aside dissimilarity between vision-based planning features and text-based planning features.

8. The method of claim 1, wherein the actors in the environment include at least one of a pedestrian, another vehicle, or a bicyclist.

9. A system utilizing a visual language planning (VLP) machine learning model, the system comprising: a camera mounted to the vehicle and configured to generate image data associated with an actor in an environment external to the vehicle; processor; and A memory comprising instructions that, when executed by a processor, cause the processor to: Processing image data to detect actors in the environment, generating a bird's eye view (BEV) of the environment based on the image data, wherein the BEV includes spatiotemporal information associated with vehicles and detected actors, and Execute a Vision-Language Planning (VLP) machine learning model to: extracting vision-based planning features from the BEV, wherein the vision-based planning features include at least some spatiotemporal information associated with the vehicle, receiving text information associated with an environment, wherein the text information describes characteristics of a vehicle in the environment, extracting text-based planning features from the text information, Performing a contrastive learning model to derive similarities between vision-based planning features and text-based planning features, and A predicted trajectory of the vehicle is generated based on the similarity.

10. The system according to claim 9, wherein: The memory comprises further instructions which, when executed by the processor, cause the processor to: determining a loss between a predicted trajectory of the vehicle and a ground truth trajectory of the vehicle; and The VLP model is executed until convergence to minimize the loss.

11. The system of claim 9, wherein the text information is generated using ground truth information associated with an environment present in a template and training data.

12. The system of claim 9, wherein the output of the contrastive learning is used as a training loss for training a VLP machine learning model.

13. The system of claim 9, wherein the contrastive learning model comprises: a text encoder configured to output a text-based vector representing text-based features associated with textual information of an environment; and An image encoder is configured to output an image-based vector representing image-based features associated with an actor detected in the BEV.

14. The system of claim 13, wherein the contrastive learning model is further configured to perform a dot product to evaluate similarity between the text-based vector and the image-based vector.

15. The system of claim 9, wherein the contrastive learning model is further configured to push aside dissimilarity between vision-based planning features and text-based planning features.

16. The system of claim 9, wherein the actors in the environment include at least one of a pedestrian, another vehicle, or a bicyclist.

17. A method for training an autonomous driving system, the method comprising: receiving image data generated from a camera mounted to the vehicle, wherein the image data includes actors in an environment external to the vehicle; generating a bird's eye view (BEV) of an environment based on the image data, wherein the BEV includes spatiotemporal information associated with vehicles and actors; executing a perception model to detect actors in the environment and associated information about the detected actors based on the BEV; executing a prediction model to estimate a trajectory of a detected actor based on the BEV; Based on the BEV, a visual language planning (VLP) model is executed to output a predicted trajectory of the vehicle, wherein the VLP model is configured to: extracting vision-based planning features from the BEV, wherein the vision-based planning features include spatiotemporal information associated with the vehicle, receiving textual information associated with an environment, wherein the textual information describes characteristics of one or more actors in the environment, extracting text-based planning features from the text information, performing contrastive learning to derive similarities between vision-based planning features and text-based planning features, and A predicted trajectory is output based on the similarity.

18. The method of claim 17, wherein the text information is generated using ground truth information associated with an environment present in a template and training data.

19. The method of claim 17, wherein the output of the contrastive learning is used as a training loss for training a VLP machine learning model.

20. The method of claim 17, wherein the actors in the environment include at least one of a pedestrian, another vehicle, or a bicyclist.

Citation Information

Cited By

  • Vehicle driving track prediction method and system and electronic equipment

    CN121106349A