Contextual awareness optimization of planner for autonomous driving
Through the situation awareness optimization method, the parameters of the planner model are adjusted according to the specific situation of the current environment, and the problem that the planner model parameters in the prior art are difficult to adapt to different driving conditions, improving the performance of the planner model and the safety and efficiency of autonomous vehicles.
Patent Information
- Application Number
- CN202411600925.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-10
- Filing Date
- 2024-11-11
- Publication Date
- 2025-05-13
AI Technical Summary
The planner model parameters in existing autonomous driving systems are usually set at the global level, which is difficult to adapt to various driving conditions and traffic situations, resulting in poor planning performance.
Through the situation awareness optimization method, the parameters of the planner model are adjusted according to the specific situation of the current environment to optimize the performance of the planner model. This method determines situation information through a perceptual model and uses an optimizer to calculate the optimized parameter set for a specific situation.
Improve the performance of the planner model under different driving conditions and traffic situations, ensuring that autonomous vehicles can plan paths more accurately, and improve safety and efficiency.
Smart Images

Figure CN119975387A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to methods and systems for situational awareness optimization of planner models for autonomous driving. Background Art
[0002] An autonomous vehicle, often referred to as a self-driving or driverless vehicle, is a type of vehicle that is able to navigate and operate on roads and in various environments without direct human control. Autonomous vehicles use a combination of advanced technologies and sensors to perceive their surroundings, make decisions, and perform driving tasks.
[0003] Autonomous vehicles are typically equipped with a variety of sensors, including lidar, radar, cameras, ultrasonic sensors, and sometimes additional technologies such as GPS and IMU (inertial measurement unit). These sensors provide real-time data about the vehicle's surroundings, including the location of other vehicles, pedestrians, road signs, and road conditions. The vehicle's onboard computer uses the data from the sensors to create a detailed map of the environment and perceive objects and obstacles. This information is critical for navigation and collision avoidance. BRIEF DESCRIPTION OF THE DRAWINGS
[0004] Figure 1 A system for training a neural network according to an embodiment is shown.
[0005] Figure 2 A computer-implemented method for training and utilizing a neural network is shown according to one embodiment.
[0006] Figure 3 A schematic diagram of a control system configured to control a vehicle, which may be a partially autonomous vehicle, a fully autonomous vehicle, a partially autonomous robot, or a fully autonomous robot, is shown according to an embodiment.
[0007] Figure 4 A schematic diagram of a system implementing an optimizer configured to optimize performance of a planner model according to an embodiment is illustrated.
[0008] Figure 5 is a method for optimizing a planner model according to an embodiment. DETAILED DESCRIPTION
[0009] Embodiments of the present disclosure are described herein. However, it is to be understood that the disclosed embodiments are merely examples, and other embodiments may take different and alternative forms. The figures are not necessarily to scale; some features may be magnified or minimized to show the details of a particular component. Therefore, the specific structural and functional details disclosed herein should not be interpreted as restrictive, but merely as a representative basis for teaching those skilled in the art to use the embodiments differently. As will be understood by those of ordinary skill in the art, the various features illustrated and described with reference to any one of the figures may be combined with the features illustrated in one or more other figures to produce embodiments that are not explicitly illustrated or described. The combination of illustrated features provides representative embodiments of typical applications. However, for a particular application or implementation, various combinations and modifications of features consistent with the teachings of the present disclosure may be required.
[0010] As used herein, "a", "an" and "the" refer to both the singular and the plural, unless the context clearly indicates otherwise. For example, a "processor" programmed to perform various functions refers to one processor programmed to perform each function, or more than one processor programmed together to perform each of the various functions.
[0011] In the context of an autonomous vehicle, the term "agent" may refer to objects or entities in the environment surrounding or interacting with the autonomous vehicle. This includes pedestrians, other vehicles, cyclists, road signs, traffic lights, lane markings, etc. Objects or features detected by the autonomous vehicle's sensors for decision making in controlling the autonomous vehicle may be collectively referred to as agents.
[0012] The present disclosure incorporates by reference in its entirety U.S. Patent Application No. __________, Attorney Docket No. 097182-00294, filed on the same day as the present disclosure and entitled “SYSTEMS AND METHODS FOR VISION-LANGUAGE PLANNING (VLP) FOUNDATION MODELS FOR AUTONOMOUS DRIVING”.
[0013] The present disclosure incorporates by reference in its entirety U.S. Patent Application No. __________, Attorney Docket No. 097182-00293, filed on the same day as the present disclosure and entitled “VISON-LANGUAGE-PLANNING (VLP) MODELS WITH AGENT-WISE LEARNING FOR AUTONOMOUS DRIVING”.
[0014] Rapid advances in autonomous driving technology have ushered in a new era of transportation, promising safer and more efficient journeys. Autonomous driving systems generally include three high-level tasks: (1) perception, (2) prediction, and (3) planning. Perception involves the vehicle's ability to understand and interpret its environment. This task includes various subcomponents such as computer vision, sensor fusion, and localization. Key elements of perception include object detection (e.g., identifying and tracking actors outside the autonomous vehicle), localization (e.g., determining the vehicle's precise position and orientation in the world, typically using GPS and other sensors), and sensor fusion (e.g., combining data from different sensors such as cameras, lidar, radar, and ultrasonic sensors to build a comprehensive view of the surrounding environment). Prediction involves anticipating how other road users and actors in the environment will behave in the near future. This task typically involves using machine learning models to estimate the trajectories and intentions of actors (including pedestrians, other vehicles, and potential obstacles). Accurate prediction is critical to making safe driving decisions. Planning involves determining the optimal path and actions for the autonomous vehicle to navigate its environment. The planner (also called the planner module or planner model) is the autonomous driving software stack that is responsible for planning the trajectory of the autonomous vehicle. This typically includes tasks such as route planning, trajectory planning, and decision making. The planning system considers information from perception and prediction to make decisions such as when to change lanes, when to stop at intersections, how to react to unexpected events, etc.
[0015] Planning the movement of an autonomous vehicle is a crucial component in autonomous driving development. The output trajectory of the planner model depends on a set of configuration parameters that influence the way the planner makes decisions. orig , and environmental information collected by the perception, positioning, and prediction modules of the autonomous driving stack, as illustrated below: trajectory = planner(paramorig, environment) 。
[0016] Such configuration parameters paramorig may include thresholds for certain maneuvers (such as turns or lane changes), the size of the action space to search, acceleration, deceleration, driving length, distance to vehicles in front of the autonomous vehicle (ego vehicle), steering angle, and other characteristics of the autonomous vehicle. These parameters have been set by the software developers on a global level, meaning that a common set of parameters is designed to handle all traffic scenarios or situations, and in the past, these parameters have not changed once put into production.
[0017] However, the planner needs to adapt to various driving conditions and traffic scenarios to operate accurately. Therefore, according to various embodiments disclosed herein, these parameters paramcontex are adjusted based on the scenario or scenario. t, to improve the performance of the planner model. This disclosure proposes a framework for finding an optimized parameter set paramconte x , where context is an expert-defined function of the current environment:
[0018] Here, situational awareness optimization is an optimization routine that takes as input a driving situation and an original parameter set and computes a parameter set paramconte that achieves the best planning performance for a particular situation. x As will be further described below, the context used to adjust the parameters can be based on the agent's sensed movement, as well as map or road characteristics (e.g., whether the road being traveled is a highway, a city street, a roundabout, etc.). The map or road characteristics can be recalled from a storage device based on a vehicle previously traveling on the road, or can be accessed via wireless communication based on another vehicle previously traveling on the road. The context can be determined by the perception model and passed to the optimizer to optimize the planner model.
[0019] Machine learning and neural networks are an integral part of the invention disclosed herein. Figure 1 A system 100 for training a neural network (e.g., a deep neural network) is shown. The system 100 may include an input interface for accessing training data 102 for the neural network. For example, Figure 1 As shown in , the input interface can be constituted by a data storage device interface 104, which can access the training data 102 from a data storage device 106. For example, the data storage device interface 104 can be a memory interface or a permanent storage device interface, such as a hard disk or SSD interface, but can also be a personal area network, local area network or wide area network interface, such as a Bluetooth, Zigbee or Wi-Fi interface or an Ethernet or fiber optic interface. The data storage device 106 can be an internal data storage device of the system 100, such as a hard disk drive or SSD, but can also be an external data storage device, such as a network accessible data storage device.
[0020] In some embodiments, the data storage device 106 may further include a data representation 108 of an untrained version of the neural network, which data representation may be accessed by the system 100 from the data storage device 106. However, it will be appreciated that the training data 102 and the data representation 108 of the untrained neural network may also be accessed from different data storage devices, such as via different subsystems of the data storage device interface 104. Each subsystem may be of a type of data storage device interface 104 as described above. In other embodiments, the data representation 108 of the untrained neural network may be generated internally by the system 100 based on the design parameters of the neural network, and thus may not be explicitly stored on the data storage device 106.
[0021] The system 100 may also include a processor subsystem 110, which may be configured to provide an iterated function as a replacement for a layer stack of a neural network to be trained during operation of the system 100. Here, the corresponding layers of the replaced layer stack may have mutually shared weights and may receive the output of a previous layer as input, or, for the first layer of the layer stack, receive an initial activation and a portion of the input of the layer stack. The processor subsystem 110 may further be configured to iteratively train the neural network using the training data 102. Here, the training iterations by the processor subsystem 110 may include a forward propagation portion and a backward propagation portion. The processor subsystem 110 may be configured to perform the forward propagation portion by determining an equilibrium point of the iterated function among other operations that define the forward propagation portion that may be performed and by providing the equilibrium point as a replacement for the output of the layer stack in the neural network, where the iterated function converges to a fixed point, wherein determining the equilibrium point includes using a numerical root-finding algorithm to find a root solution of the iterated function minus its input. The system 100 may also include an output interface for outputting a data representation 112 of the trained neural network; this data may also be referred to as trained model data 112. Figure 1 , the output interface may be comprised of a data storage device interface 104, which in these embodiments is an input / output ("IO") interface, via which trained model data 112 may be stored in a data storage device 106. For example, data representations 108 defining an "untrained" neural network may be at least partially replaced by data representations 112 of a trained neural network during or after training, as parameters of the neural network (e.g., weights, hyperparameters, and other types of parameters of the neural network) may be adapted to reflect training based on training data 102. This is also Figure 1106 by reference numerals 108, 112, which refer to the same data record on the data store 106. In other embodiments, the data representation 112 may be stored separately from the data representation 108 defining the "untrained" neural network. In some embodiments, the output interface may be separate from the data store interface 104, but may generally be of the type described above for the data store interface 104.
[0022] Figure 1 The illustrated system 100 is one example of a system that may be used to train the machine learning models described herein.
[0023] Figure 2 A system 200 that implements the machine learning models described herein (e.g., planner models) is depicted. System 200 may include at least one computing system 202. Computing system 202 may include at least one processor 204, which may be operably connected to a memory unit 208. Processor 204 may include one or more integrated circuits that implement the functionality of a central processing unit (CPU) 206. CPU 206 may be a commercially available processing unit that implements an instruction set such as one of the x86, ARM, Power, or MIPS instruction set families. During operation, CPU 206 may execute stored program instructions retrieved from memory unit 208. The stored program instructions may include software that controls the operation of CPU 206 to perform the operations described herein. In some examples, processor 204 may be a system on a chip (SoC) that integrates the functionality of CPU 206, memory unit 208, network interfaces, and input / output interfaces into a single integrated device. Computing system 202 may implement an operating system for managing different aspects of operation. Although in Figure 2 One processor 204, one CPU 206, and one memory 208 are shown, but of course more than one of each component may be utilized in the overall system.
[0024] The memory unit 208 may include volatile memory and non-volatile memory for storing instructions and data. The non-volatile memory may include solid-state memory, such as NAND flash memory, magnetic and optical storage media, or any other suitable data storage device that retains data when the computing system 202 is disabled or loses power. The volatile memory may include static and dynamic random access memory (RAM) that stores program instructions and data. For example, the memory unit 208 may store a machine learning model 210 or algorithm, a training data set 212 for the machine learning model 210, and an original source data set 216.
[0025] The computing system 202 may include a network interface device 222 that is configured to provide communications with external systems and devices. For example, the network interface device 222 may include a wired and / or wireless Ethernet interface defined by the Institute of Electrical and Electronics Engineers (IEEE) 802.11 family of standards. The network interface device 222 may include a cellular communication interface for communicating with a cellular network (e.g., 3G, 4G, 5G). The network interface device 222 may also be configured to provide a communication interface to an external network 224 or a cloud.
[0026] The external network 224 may be referred to as the World Wide Web or the Internet. The external network 224 may establish a standard communication protocol between computing devices. The external network 224 may allow information and data to be easily exchanged between computing devices and the network. One or more servers 230 may communicate with the external network 224.
[0027] The computing system 202 may include an input / output (I / O) interface 220, which may be configured to provide digital and / or analog input and output. The I / O interface 220 is used to transfer information between an internal storage device and an external input and / or output device (e.g., an HMI device). The I / O 220 interface may include an associated circuit or bus network to transfer information between (one or more) processors and storage devices. For example, the I / O interface 220 may include digital I / O logic lines that can be read or set by (one or more) processors, handshake lines for supervising data transfer via I / O lines, timing and counting facilities, and other structures known to provide such functions. Examples of input devices include keyboards, mice, sensors, touch screens, etc. Examples of output devices include monitors, touch screens, speakers, head-up displays, vehicle control systems, etc. The I / O interface 220 may include additional serial interfaces (e.g., universal serial bus (USB) interfaces) for communicating with external devices. The I / O interface 220 may be referred to as an input interface because it transfers data from an external input, such as a sensor, or an output interface because it transfers data to an external output, such as a display.
[0028] The computing system 202 may include a human-machine interface (HMI) device 218, which may include any device that enables the system 200 to receive control inputs. The computing system 202 may include a display device 232. The computing system 202 may include hardware and software for outputting graphical and textual information to the display device 232. The display device 232 may include an electronic display screen, a projector, a speaker, or other suitable device for displaying information to a user or operator. The computing system 202 may also be configured to allow interaction with a remote HMI and a remote display device via a network interface device 222.
[0029] System 200 may be implemented using one or more computing systems. Although the example depicts a single computing system 202 that implements all of the described features, it is intended that the various features and functions may be separated and implemented by multiple computing units that communicate with each other. The specific system architecture selected may depend on a variety of factors.
[0030] The system 200 may implement a machine learning algorithm 210 configured to analyze a raw source data set 216. The raw source data set 216 may include raw or unprocessed sensor data that may represent an input data set for a machine learning system. The raw source data set 216 may include video, video clips, images, text-based information, audio or human speech, time series data (e.g., pressure sensor signals over time), and raw or partially processed sensor data (e.g., radar maps of objects). In some examples, the machine learning algorithm 210 may be a neural network algorithm (e.g., a deep neural network) designed to perform a predetermined function. For example, a neural network algorithm may be configured in an automotive application to identify street signs or pedestrians in an image. The (one or more) machine learning algorithms 210 may include algorithms configured to operate one or more machine learning models described herein, including a VLP base model.
[0031] The computing system 202 may store a training data set 212 for the machine learning algorithm 210. The training data set 212 may represent a set of previously constructed data for training the machine learning algorithm 210. The machine learning algorithm 210 may use the training data set 212 to learn weighting factors associated with the neural network algorithm. The training data set 212 may include a set of source data having corresponding achievements or results that the machine learning algorithm 210 attempts to replicate via the learning process. In this example, the training data set 212 may include an input image containing an object (e.g., a street sign). The input image may include various scenes in which the object is identified. The training data set 212 may also include a text description of the scene corresponding to the image detected by the vehicle sensor (e.g., "a pedestrian is crossing the street").
[0032] The machine learning algorithm 210 can operate in a learning mode using the training data set 212 as input. The machine learning algorithm 210 can be executed through multiple iterations using data from the training data set 212. With each iteration, the machine learning algorithm 210 can update the internal weighting factors based on the achieved results. For example, the machine learning algorithm 210 can compare the output results (e.g., in the case where image data is input, the reconstructed or supplemented image) with those included in the training data set 212. Since the training data set 212 includes the expected results, the machine learning algorithm 210 can determine when the performance is acceptable. After the machine learning algorithm 210 reaches a predetermined performance level (e.g., 100% consistent with the results associated with the training data set 212) or converges, the machine learning algorithm 210 can be executed using data that is not in the training data set 212. It should be understood that in the present disclosure, "convergence" can mean that a set (e.g., predetermined) number of iterations have occurred, or the residual is small enough (e.g., the change in the approximation probability within the iteration changes by less than a threshold), or other convergence conditions. The trained machine learning algorithm 210 can be applied to a new data set to generate annotated data. In the context of the planner model described herein, a loss between a predicted trajectory of an autonomous vehicle and a ground truth trajectory of the vehicle may be determined, and the model may be trained with an optimizer to reduce the loss, e.g., to convergence.
[0033] The machine learning algorithm 210 may be configured to identify specific features in the raw source data 216. The raw source data 216 may include multiple instances or input data sets for which supplementary results are desired. For example, the machine learning algorithm 210 may be configured to identify the presence of an actor in a video image, annotate the presence, and / or command the vehicle to take specific actions (planning) based on the actor's location data (perception) and the actor's predicted future movement / position (prediction). The machine learning algorithm 210 may be programmed to process the raw source data 216 to identify the presence of specific features. The machine learning algorithm 210 may be configured to identify features in the raw source data 216 as predetermined features (e.g., road signs, pedestrians, etc.). The raw source data 216 may be obtained from a variety of sources. For example, the raw source data 216 may be actual input data collected by the machine learning system. The raw source data 216 may be machine generated for testing the system. As an example, the raw source data 216 may include raw video images from a camera.
[0034] Figure 3A schematic diagram of a control system 302 configured to control a vehicle 300, which may be a partially autonomous vehicle or a fully autonomous vehicle, a partially autonomous robot or a fully autonomous robot, is depicted. The vehicle 300 and / or its control system 302 may incorporate one or more components of the system 200, such as the computing system 202, to command an actuator 304 to perform an action based on processed readings from one or more sensors 306. For example, the control system 302 may be configured to utilize a planning model disclosed herein to control movement of the vehicle via the actuator 304, wherein the planning model is trained via an optimizer.
[0035] The one or more sensors 306 may include one or more image sensors (e.g., cameras, video sensors, radar sensors, ultrasonic sensors, lidar sensors), and / or location sensors (e.g., GPS). The sensors 306 may be configured to generate raw source data 216. One or more of the one or more specific sensors may be integrated into the vehicle 300. In the context of actor identification and processing as described herein, the sensor 306 is a camera mounted to or integrated into the vehicle 300. Instead of or in addition to the one or more specific sensors described above, the sensor 306 may include a software module configured to determine the state of the actuator 304 when executed. The data generated from these sensors may be fused or otherwise combined to create a bird's eye view (BEV) that provides spatiotemporal information associated with vehicles and detected actors in the environment.
[0036] In embodiments where vehicle 300 is a fully or partially autonomous vehicle, actuator 304 may be embodied in a brake, accelerator, propulsion system, engine, transmission, or steering system (e.g., steering wheel) of vehicle 300. For example, actuator control commands may be determined such that actuator 304 is controlled such that vehicle 300 avoids colliding with a detected actor. Detected actors may also be classified according to what the classifier deems them most likely to be, such as a pedestrian or a tree. Actuator control commands may be determined depending on the classification.
[0037] In other embodiments where the vehicle 300 is a fully or partially autonomous robot, the vehicle 300 may be a mobile robot configured to perform one or more functions, such as flying, swimming, diving, and walking, via the actuators 304. The mobile robot may be an at least partially autonomous lawn mower or an at least partially autonomous cleaning robot. In such embodiments, actuator control commands may be determined such that a propulsion unit, a steering unit, and / or a braking unit of the mobile robot may be controlled such that the mobile robot may avoid a collision with the identified object.
[0038] Figure 4A schematic diagram of a system 400 implementing an optimizer 402 is shown, the optimizer 402 being configured to optimize the performance of a planner model 404. The planning model 404 is a planning model such as the one disclosed in U.S. Patent Application No. __________ (Attorney Docket No. 097182-00294), filed on the same day as the present disclosure and entitled “SYSTEMS AND METODS FOR VISION-LANGUAGE PLANNING (VLP) FOUNDATION MODELS FOR AUTONOMOUS DRIVING”. For example, a planning query may contain perception information and prediction information from a perception model and a prediction model, respectively. The planning query may be configured to extract planning information by utilizing a bird's eye view (BEV) feature map. The BEV is a fusion of information sensed from various vehicle sensors (e.g., cameras, lidars, radars, etc.) corresponding to actors around the vehicle. Using the BEV and all data thereon determined from the perception and prediction models, the planning model plans an appropriate route for the vehicle. The vector may be used with a planning decoder, which extracts visual planning features using the planning query and the BEV features. Features extracted from the BEV can include information about actors in the environment, such as their locations, their trajectories, etc., which, as described above, are inputs used to determine how the autonomous vehicle itself should react. A neural network can map high-dimensional inputs to an expected time-stamped trajectory of the vehicle. Figure 4 The illustrated items on the left side of the optimizer 402 in FIG. 4 may be stored in memory and called out for execution.
[0039] like Figure 4 As shown in the embodiment illustrated in , planning model 404 may include data output from prediction model 406, and data output from perception model 408, including sensed behavior of actors in the environment. Using information extracted from prediction model 406 and perception model 408, the planning model may output a predicted trajectory 410 for the autonomous vehicle. This may be done according to the methods described in the patent references incorporated by reference above.
[0040] The optimizer 402 is configured to optimize the parameters used by the planning model 404 in order to improve the predicted trajectory 410. As explained above, the trajectory 410 output by the planner depends on a set of configuration parameters that affect the way the planner makes decisions. Such parameters may include thresholds for certain maneuvers (such as turns or lane changes), vehicle acceleration, vehicle deceleration, turns, vehicle spacing, etc. These parameters are typically set at a global level. The optimizer 402 is configured to find an optimized set of parameters param for a given set of driving scenarios. contexThe optimizer 402 takes as input the driving scenario(s) and the original parameter set, and computes an optimized or adjusted set of parameters that achieves the best planning performance for a particular scenario or scene. Once optimized, the software can be integrated into the autonomous driving stack. It can be used as a developer tool to improve the development of new autonomous driving features, and ultimately implemented into the autonomous vehicle itself.
[0041] For example, optimizer 402 may be executed using computing system 202. In operation (e.g., during execution of the optimizer), the computing system recalculates the planner output (e.g., predicted trajectory) for a given driving scenario based on a different set of parameters. This is accomplished by re-running the planner on a replay of the drive using data from past real drivers (ground truth) and then using the new set of parameters. The output of this run (e.g., predicted trajectory) is stored in memory (e.g., on disk).
[0042] This process may be repeated, with each iteration producing a new predicted trajectory, which is stored.
[0043] The computer system can then calculate a cost function 412. The cost function is defined based on the target behavior that is desired to be improved. For example, does the predicted trajectory deviate from the actual trajectory by a large margin? Does the predicted trajectory predict a more drastic deceleration than is actually present in the ground truth? The planner model has various planning-related behaviors that are independently configured using different sets of parameters. Therefore, the choice of cost function will also affect which parameters to update or change. The cost function can take various forms, depending on what needs to be improved.
[0044] In one embodiment, cost function 412 evaluates the nominal trajectory difference. As a subtask of planning, the system predicts the trajectory of the autonomous vehicle for the planning horizon, which is called the nominal trajectory or predicted trajectory. However, due to the difference between planning and execution, the nominal trajectory may be different from the executed trajectory. In order to ensure safe and reliable planning, it is important that the nominal trajectory is as close as possible to the executed trajectory. Therefore, the cost function can be defined as the average displacement error between the planned trajectory and the nominal trajectory over the observation window.
[0045] In one embodiment, the cost function 412 evaluates the planning cost. As part of the planning model, the system searches for a sequence of actions that minimize the cost, such as acceleration, deceleration, fuel consumption, sharp turns, etc. The cost can also be used to optimize the planning parameters. For example, when the vehicle is moving towards its goal under the constraints of traffic, safety regulations, and driver comfort, the planning cost is low. A better set of parameters should enable the vehicle to achieve a lower planning cost at each point in time, so the average cost on driving can be considered as the goal of optimization. In short, depending on the driving scenario (e.g., traffic, speed, etc.), the use of different parameters can enable the vehicle to improve the driver's driving experience many times, resulting in an overall better driving experience.
[0046] Regardless of the type of cost function used, the optimizer 402 is configured to generate the next candidate parameters to evaluate the cost function. An efficient optimizer will generate a parameter set with improved cost in fewer iterations. Since the gradient of the cost function for these problems may not be accessible to the provider of the optimizer, in an embodiment, the system can be limited to using a gradient-free optimizer. The framework of the optimizer is compatible with any gradient-free / black-box style optimizer, including Bayesian optimization, genetic algorithms, particle swarms, or any meta-heuristic based algorithm.
[0047] The above-mentioned system 400 can have several training applications. For example, during training, the system can receive context information or context definitions from the perception model. Such context information can include the past and current paths traveled by actors (e.g., other vehicles), their relative speeds, directions, etc. The context information can also include the characteristics of the road on which the vehicle is traveling. Actors usually behave differently on highways than on urban roads, where speed limits are much lower and other pedestrians and / or objects may be crossing the road or adjacent to the vehicle. Therefore, the system can use map information (e.g., from GPS) indicating the type of road on which it is traveling. The optimizer 402 can use this context information to change parameters during training. Training can be completed using several iterations, and the optimizer can output optimized planning performance for a specific situation or scenario. Therefore, when put into production, the operation of the autonomous vehicle can be optimized for different driving environments (e.g., based on how other actors behave and / or differences in the road being traveled). In other embodiments, the system can utilize data used by a localization and mapping system (e.g., a simultaneous localization and mapping model (SLAM)) that stores data associated with the route that the vehicle or other vehicles are traveling and makes that data available when the vehicle is driving through roads that have been previously mapped.
[0048] Figure 5 A method 500 for optimizing a planner model according to an embodiment is illustrated. The method may be implemented or performed using one or more processors described herein. Although only one processor may be described in one or more steps below, it should be understood that more than one processor may be used together to perform the method.
[0049] At 502, image data generated from a camera mounted to a vehicle is received by a processor. The image data includes actors in the environment outside the vehicle (e.g., vehicles, pedestrians, buildings, cyclists, etc.). At 504, a pre-existing perception model is used to detect actors in the environment based on the image data. The perception model can use image processing, neural networks and other machine learning techniques, and sensor fusion, with the goal of determining real-time movement characteristics of actors in the sensed environment. This includes the actor's location, orientation, speed, whether the vehicle is turning, accelerating, braking, etc.
[0050] At 506, the planner model is executed to generate a predicted trajectory for the vehicle. This may include predicted movement, acceleration, deceleration, turning, etc. of the vehicle. The predicted trajectory depends on a set of configuration parameters associated with the vehicle characteristics, param contex For example, parameters may include thresholds for certain maneuvers (eg, turns, lane changes), speeds, and other maneuver instructions.
[0051] At 508, the processor receives information from the perception model regarding the context associated with the detected actor in the environment. The context may include, for example, the speed, orientation, angle, turn, size, type (e.g., truck, van, bus), etc. associated with the actor. In some embodiments, the context may also be associated with the type of road being traveled, such as whether the road is a highway, an on / off ramp, a roundabout, a city street, a parking lot, etc. The optimizer takes this into account because actors behave differently in different scenarios, and therefore the optimizer optimizes the performance of the planner based on the different contexts of the actors and / or roads.
[0052] At 510, an optimizer is executed. The optimizer may be a black box optimizer configured to optimize the performance of the planner model based on the scenario. The execution of the optimizer may include the execution of 512-520, as explained below and described above.
[0053] At 512, the optimizer selects a subset of configuration parameters to be optimized based on the context. At 514, the optimizer selects an objective function based on the context. For example, the objective function can be a cost function configured to optimize a nominal trajectory, or can be a cost function configured to optimize a planning cost. At 516, the optimizer adjusts the selected subset of configuration parameters based on the current value of the objective function. At 518, the optimizer generates a new planned trajectory for the vehicle based on the adjusted configuration parameters. At 520, the optimizer derives the value of the objective function (e.g., cost function) based on the new planned trajectory. Then, steps 516-520 can be repeated through several iterations to determine the optimal configuration parameters that minimize the objective function. This can be based on convergence, for example, when convergence, the new predicted trajectory under iteration is aligned with the actual ground truth trajectory of the vehicle.
[0054] Although exemplary embodiments are described above, these embodiments are not intended to describe all possible forms contained by the claims. The words used in the specification are words of description rather than limitation, and it is understood that various changes can be made without departing from the spirit and scope of the present disclosure. As previously described, the features of the various embodiments can be combined to form more embodiments of the present invention that may not be explicitly described or illustrated. Although various embodiments may have been described as providing advantages in one or more desired characteristics or being preferred over other embodiments or prior art implementations, it is recognized by those of ordinary skill in the art that one or more features or characteristics can be compromised to achieve the desired overall system properties, depending on the specific application and implementation. These properties may include, but are not limited to, cost, strength, durability, life cycle cost, marketability, appearance, packaging, size, applicability, weight, manufacturability, ease of assembly, etc. Therefore, in terms of one or more characteristics, to the extent that any embodiment is described as less desirable than other embodiments or prior art implementations, these embodiments are not outside the scope of the present disclosure and may be ideal for specific applications.
Claims
1. A method for optimizing a planner model for autonomous driving, the method comprising: receiving image data generated from a camera mounted to the vehicle, wherein the image data includes actors in an environment external to the vehicle; Detecting actors in the environment based on image data via pre-existing perception models; executing the planner model to generate a predicted trajectory for the vehicle, wherein the generated predicted trajectory for the vehicle depends on a set of configuration parameters associated with characteristics of the vehicle; receiving information from a perception model regarding a context associated with a detected actor in the environment; as well as Executing an optimizer configured to optimize performance of the planner model based on the scenario, wherein executing the optimizer: (1) selecting a subset of configuration parameters to be optimized based on the scenario, (2) selecting an objective function based on the context, (3) adjusting a selected subset of configuration parameters based on the current value of the objective function, (4) generating a new planned trajectory for the vehicle based on the adjusted configuration parameters, (5) deriving the value of the objective function based on the new planning trajectory, and (6) Repeat (3)-(5) to determine the optimal configuration parameters that minimize the objective function.
2. The method of claim 1, wherein the optimizer is a black box optimizer.
3. The method of claim 1, wherein the subset of configuration parameters is associated with a vehicle controller system; The objective function chosen consists of the average displacement error between the final planned trajectory derived after iterations of (3)-(5) and the nominal trajectory generated by the planner model.
4. The method of claim 3, wherein the optimizer performs: Sampling a subset of configuration parameters, generating a new planning trajectory based on adjustments to the selected subset of configuration parameters, and Calculate the mean displacement error.
5. The method of claim 1 , wherein the planning model comprises a planning algorithm, and the subset of configuration parameters is associated with the planning algorithm; The objective function selected includes the loss function of the planning algorithm.
6. The method of claim 5, wherein the optimizer performs: Sampling a subset of configuration parameters, generating a new planning trajectory based on adjustments to the selected subset of configuration parameters, and Compute the loss function of the planning algorithm.
7. The method of claim 1, wherein the detection of actors in the environment is further based on a pre-existing positioning model.
8. The method according to claim 1, further comprising: receiving sensor data from a sensor mounted to the vehicle, wherein the sensor is not a camera, The detection of actors in the environment is further based on sensor data.
9. The method of claim 1, wherein the context associated with the detected actor is associated with the movement of the detected actor and characteristics of a road on which the vehicle is traveling.
10. The method of claim 1, wherein the characteristics of the road on which the vehicle is traveling include whether the road is a highway, a city road, or a ring road.
11. A system for optimizing a planner model for autonomous driving, the system comprising: processor; as well as A memory comprising instructions which, when executed by a processor, cause the processor to perform the following steps: receiving image data generated from a camera mounted to the vehicle, wherein the image data includes actors in an environment external to the vehicle; Detecting actors in the environment based on image data via pre-existing perception models; executing the planner model to generate a predicted trajectory for the vehicle, wherein the generated predicted trajectory for the vehicle depends on a set of configuration parameters associated with characteristics of the vehicle; receiving information from a perception model regarding a context associated with a detected actor in the environment; as well as Executing an optimizer configured to optimize performance of the planner model based on the scenario, wherein executing the optimizer: (1) selecting a subset of configuration parameters to be optimized based on the scenario, (2) selecting an objective function based on the context, (3) adjusting a selected subset of configuration parameters based on the current value of the objective function, (4) generating a new planned trajectory for the vehicle based on the adjusted configuration parameters, (5) deriving the value of the objective function based on the new planning trajectory, and (6) Repeat (3)-(5) to determine the optimal configuration parameters that minimize the objective function.
12. The system of claim 1, wherein the optimizer is a black box optimizer.
13. The system of claim 1, wherein the subset of configuration parameters is associated with a vehicle controller system; and The objective function chosen consists of the average displacement error between the final planned trajectory derived after iterations of (3)-(5) and the nominal trajectory generated by the planner model.
14. The system of claim 13, wherein the optimizer performs: Sampling a subset of configuration parameters, generating a new planning trajectory based on adjustments to the selected subset of configuration parameters, and Calculate the mean displacement error.
15. The system of claim 11, wherein the planning model comprises a planning algorithm, and the subset of configuration parameters is associated with the planning algorithm; The objective function selected includes the loss function of the planning algorithm.
16. The system of claim 15, wherein the optimizer performs: Sampling a subset of configuration parameters, generating a new planning trajectory based on adjustments to the selected subset of configuration parameters, and Compute the loss function of the planning algorithm.
17. The system of claim 11, wherein the detection of actors in the environment is further based on a pre-existing positioning model.
18. The system of claim 1, wherein the instructions are further configured to, when executed by a processor, cause the processor to: receiving sensor data from a sensor mounted to the vehicle, wherein the sensor is not a camera, The detection of actors in the environment is further based on sensor data.
19. The system of claim 11, wherein the context associated with the detected actor is associated with the movement of the detected actor and characteristics of a road on which the vehicle is traveling.
20. A non-transitory computer-readable storage medium comprising instructions that, when executed by a processor, cause the processor to: receiving image data generated from a camera mounted to the vehicle, wherein the image data includes actors in an environment external to the vehicle; Detecting actors in the environment based on image data via pre-existing perception models; executing the planner model to generate a predicted trajectory for the vehicle, wherein the generated predicted trajectory for the vehicle depends on a set of configuration parameters associated with characteristics of the vehicle; receiving information from a perception model regarding a context associated with a detected actor in the environment; as well as Executing an optimizer configured to optimize performance of the planner model based on the scenario, wherein executing the optimizer: (1) selecting a subset of configuration parameters to be optimized based on the scenario, (2) selecting an objective function based on the context, (3) adjusting a selected subset of configuration parameters based on the current value of the objective function, (4) generating a new planned trajectory for the vehicle based on the adjusted configuration parameters, (5) deriving the value of the objective function based on the new planning trajectory, and (6) Repeat (3)-(5) to determine the optimal configuration parameters that minimize the objective function.