Combination prediction and path planning for autonomous objects using neural networks
By using machine learning and deep learning to predict the responses of other vehicles, more accurate path planning is generated, solving the problem that existing technologies fail to consider the reactions of nearby vehicles and improving the navigation capabilities and path planning accuracy of autonomous vehicles.
Patent Information
- Application Number
- CN202080012391.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-02-05
- Filing Date
- 2020-01-27
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2040-01-27
AI Technical Summary
Existing path planning methods fail to effectively consider the possible reactions of nearby vehicles to the actions taken by the autonomous vehicle, resulting in inaccurate path planning and potentially leading to suboptimal paths or delays.
Machine learning and deep learning methods are used to predict the response actions of other vehicles. More accurate path planning is generated through decision trees and value functions. Vehicle decisions at multiple time scales are considered, and action sequences are optimized using sensor data and vehicle communication information.
It improves the accuracy and efficiency of path planning, reduces vehicle delays and collision risks in complex environments, and enhances the navigation capabilities of autonomous vehicles.
Smart Images

Figure CN113474231B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims priority to U.S. Application No. 16 / 268,188, filed February 5, 2019, entitled “COMBINED PREDICTION AND PATH PLANNING FOR AUTONOMOUS OBJECTS USINGNEURAL NETWORKS”, the entire disclosure of which is incorporated herein by reference for all purposes. Background Technology
[0003] Technological advancements have led to the introduction of autonomous control technologies for many different applications. In the case of autonomous vehicles, for example, this involves determining the path a vehicle should take based on the state of its surroundings, such as using sensor data captured by the individual vehicles. While such methods provide adequate navigation planning in many situations, conventional approaches do not account for the possible reactions of nearby vehicles to actions taken by the navigated vehicle. Therefore, path planning is less accurate than it should be, and may lead to suboptimal paths when nearby vehicles exhibit certain types of reactions or movements. Attached Figure Description
[0004] Various embodiments according to this disclosure will be described with reference to the accompanying drawings, in which:
[0005] Figure 1A , Figure 1B and Figure 1C Example sequences of actions that can be predicted and used for path planning according to various embodiments are shown.
[0006] Figure 2A and Figure 2B Example planning grids that can be used according to various embodiments are shown.
[0007] Figure 3 Example decision trees, according to various embodiments, are shown that can be used to determine the navigation option with the highest value.
[0008] Figure 4 A first example process for determining navigation actions for an object, which can be utilized according to various embodiments, is shown.
[0009] Figure 5 A second example process for determining navigation actions for an object, which can be utilized according to various embodiments, is shown.
[0010] Figure 6 An example environment in which aspects of the various embodiments can be implemented is shown.
[0011] Figure 7 Example systems for training image synthesis networks that can be utilized according to various embodiments are shown.
[0012] Figure 8 The layers of example statistical models that can be utilized according to various embodiments are shown.
[0013] Figure 9 Example components of a computing device that can be used to implement aspects of the various embodiments are shown. Detailed Implementation
[0014] In the following description, various embodiments will be described. Specific configurations and details are set forth for illustrative purposes to provide a thorough understanding of the embodiments. However, it will be apparent to those skilled in the art that these embodiments can be practiced without specific details. Furthermore, well-known features may be omitted or simplified to avoid obscuring the described embodiments.
[0015] Methods according to various embodiments provide navigation of controllable objects, such as autonomous vehicles or robots. These objects may be at least partially autonomous or controllable and are capable of manipulation in part based on determined paths, actions, or goals, as well as other possibilities discussed and suggested herein. One or more sensors may be used to sense information about objects (operable or otherwise) near the current object to be manipulated. Information may also be obtained from nearby objects, if available. This information may be used to determine a possible sequence of actions that the current object may take to achieve a determined goal, such as moving toward a determined destination. For each possible action of the current object, one or more possible response or reaction actions of nearby objects (i.e., actors) may be determined. In some embodiments, this may take the form of a decision tree with alternating levels of nodes corresponding to possible actions of the current vehicle and possible response actions of one or more other vehicles or actors. Machine learning may be used to determine probabilities, as well as project options along the branches and paths (including sequences) of the decision tree. In some embodiments, only actions with at least the minimum probability are considered. In another embodiment, actions may be considered based on factors such as the corresponding risk or loss amount, favorability, occupant comfort, etc. A value function can be used to generate values for each considered sequence or path, and a suggested navigation path with the highest value can be selected. At least a first action of the suggested navigation path can be provided to an optimizer of the control system, which can use the first action to determine how to navigate the current object. The selected path and relevant data can be used to update one or more machine learning models used for this determination, for example, by sending the relevant data to a remote server capable of further training the model, which can then be used for future determinations. Therefore, in at least some embodiments, transfer learning can be used for further learning. Such methods offer significant advantages compared to conventional methods that separate prediction and planning tasks, which often result in predictions that are independent of planning, leading to poor planning.
[0016] Various other functionalities may be implemented in the various embodiments and discussed and suggested elsewhere in this document.
[0017] As previously mentioned, various methods of navigating or maneuvering autonomous (or at least semi-autonomous) vehicles involve some type of path planning for the vehicle. The vehicle may have various sensors, such as those discussed elsewhere herein, including cameras, proximity sensors, depth sensors, motion sensors, position sensors, accelerometers, electronic compasses, etc., which provide data that can be analyzed to determine the state of the world or environment within a determinable distance of the vehicle. For example, in situations such as... Figure 1AIn the illustrated state 100, vehicle 102 may be able to collect and analyze environmental data that enables the vehicle to determine its position on the road. Vehicle 102 is then able to determine a set of navigation actions, at least in part, based on the determined destination or objective, to maneuver the vehicle along the road to that destination. This may include, for example, acceleration or deceleration, lane changing, turning, etc. While such actions can be determined in, for example, a discretizable maneuvering space, in some embodiments, the optimal or preferred trajectory of the vehicle may also be predicted based on factors such as the current state of the environment, the objective, and the predicted response, and then the determined maneuvers take effect to follow the optimal trajectory. This approach may not use a tree-based method but can be part of a generalized decision-making process. In one embodiment, the decision manager may predict a set of discrete actions and a set of continuous optimal values, such as the optimal position within one second, two seconds, etc. In some embodiments, the discretized actions in the maneuvering space may be provided to an optimizer, which may then determine small adjustments that the vehicle actually performs on a short timescale, such as within the next 50 ms.
[0018] However, vehicles must typically consider the presence of other objects in the environment they must consider. This can include, for example, considering other vehicles on road 108 (e.g., vehicles 104, 106, etc.), their relative speeds and directions of travel, etc. In almost all cases, the vehicle's goal is to achieve its objective while avoiding collisions with other vehicles. Various other objectives may also exist, such as maximizing occupant comfort, maximizing visibility or experience, avoiding near-collision or vibrational actions, etc. This approach can include determining a path to the destination that avoids collisions with other vehicles. Conventional path planning algorithms can analyze the relative positions of nearby vehicles 104, 106 on road 108, predict the most likely future paths of these nearby vehicles, and determine the appropriate actions to take. This can include, for example, Figure 1B As shown in example state 120, the vehicle changes lanes to the right between two vehicles 104 and 106, and there is sufficient space between the vehicle and the vehicles in front and behind after the lane change.
[0019] However, such path planning does not take into account the possible or likely actions of other vehicles 104, 106 in response to an attempt to move vehicle 102 into the right lane. For example, the driver of one of the other vehicles 106 might see the turn signal of vehicle 102 attempting to change lanes and accelerate, such as... Figure 1CAs shown in state 140, this prevents vehicle 102 from overtaking other vehicle 106. Another driver might also take other actions, such as slowing down to make room or simply continuing at the current speed and direction, which may not provide enough space for vehicle 102 to move to the right lane as planned. The prediction of vehicle 106's future position is not independent of the actions of another nearby vehicle 102, but at least partially depends on them. Therefore, vehicle 102's planned path may not be successfully executed, and the vehicle may have to determine an alternative path or option after a period of time. This can cause delays, resulting in the vehicle missing an exit or at least not progressing optimally towards its target or destination.
[0020] Therefore, methods according to various embodiments may attempt to include predictions of other actors or objects (e.g., vehicles, pedestrians, cyclists, etc.) during the path planning process. In attempting to determine the optimal action to be taken next by the corresponding vehicle, a range of possible actions for all actors may be considered and taken into account. This may include, for example, using tree search to predict different trajectories for other actors in response to various actions the vehicle to be navigated might take. Such a process can produce a more accurate plan than possible using conventional path planning methods. Furthermore, in at least some embodiments, other actors may be characterized in a variety of ways, such as one of a set of classifications that can be used to more accurately predict their behavior, since aggressive actors may react or move differently than cautious or careless actors. Machine learning can be used to improve the characterization of actors, as well as to improve the prediction of values for each sequence of options used to create a decision tree or determine options. In some embodiments, the characteristics of behavior may not be a set of discrete classifications, but rather determined based on a set of determinable behavioral parameters. This may include having one or more scalar parameters to determine the driver's level of aggression, for example, among other such options.
[0021] In some embodiments, the process may involve predicting or determining the probability of the presence of vehicles or other actors hidden outside the sensor's field of view, such as in a field of view obscured by another vehicle. Therefore, the process may consider the possibility of a car parked outside the line of sight, such as at or near an intersection or lane. This may also include the possibility of pedestrians on a crosswalk, bicycles on the roadside, etc. Future predictions may include the presence of actors not currently present in the scene. If inter-vehicle communication is available, the process may also accept information about other actors or objects that are visible or detectable but nearby but outside the current vehicle's field of view.
[0022] In various embodiments, the vehicle (such as an autonomous vehicle) will make action decisions at multiple time scales. For example, the vehicle may consider the path at the outermost time scale, which may involve the number of hours of travel for the entire route to the destination. Actions determined within this time scale may include making a specific turn to take a specific route and may include lane changes necessary to properly prepare for the turn. Another time scale just below this might be a five- to ten-second time scale, which is comparable to the time required to stop an average autonomous vehicle (although this time scale may vary appropriately depending on factors such as vehicle type, maximum speed, and path conditions). Since this is the time required to stop the vehicle in an emergency, the vehicle can formulate specific plans within this timeframe. The five- to ten-second planner outputs the target location for each "movement" of the current vehicle. A typical movement may take from 250 milliseconds to several seconds. In many embodiments, the vehicle will also make decisions at a third, finer time scale, which is used to make specific adjustments to the vehicle to follow the determined path and take into account changes in the surrounding environment. In some embodiments, a controller or optimizer may take the output of the five- to ten-second planner as input and use this information to determine the action to take over, for example, the next 50-ms time period. In at least some embodiments, these decisions can be made simultaneously. This may include, for example, determining adjustments to steering, acceleration, braking, or deceleration. The optimizer can then use updated information from vehicle sensors, the most recent 5- to 10-second plan, and other such relevant information to make new decisions for 50-millisecond time intervals.
[0023] As mentioned earlier, conventional 5- to 10-second planners use predictions of the actions of other vehicles, independent of the current vehicle's actions, to make decisions about the vehicle's next move. Planners can use simple predictions of the motions of nearby actors, but these predictions are based on factors such as the current direction, position, velocity, and acceleration of nearby actors, without considering that some of these values might change based on the current vehicle's actions, such as whether the current vehicle changes lanes or stops in front of one of the vehicles. For example, conventional robot path planning algorithms perform an A* tree search on the possible moves or actions of the current vehicle, independent of changes that might occur to other actors due to these moves or actions. Other methods have also been tried, such as using reinforcement learning for path planning, but these methods still rely on predictions of the reactions of other nearby or adjacent actors.
[0024] The methods according to various embodiments attempt to consider the predicted trajectories of other actors and utilize those predicted trajectories to determine the path or sequence of actions for the current vehicle to navigate. Furthermore, the prediction of those other actors' trajectories can take into account the current vehicle's movement or actions at various stages, points, or levels of the possible path, thus allowing for consideration of different possible reactions. Therefore, it can be considered that if the current vehicle changes lanes, other vehicles might change their trajectories accordingly, and this then needs to be considered differently. Instead of using conventional methods (such as reinforcement learning or A* search) to manage path prediction, the methods according to various embodiments can view path planning as a multiplayer cooperative game where players may have common goals and essentially take turns achieving those goals, with one player's action at a given point in time depending at least in part on what other players did at a previous point in time. While each actor's ultimate goal may differ, such as reaching different destinations, there may be common goals, such as avoiding collisions and moving as efficiently as possible.
[0025] If path planning is viewed as a multiplayer game where each actor takes a sequence of actions, which may depend at least in part on the actions of others, then the set of possible actions at each stage, point, step, or level can be used to generate a decision tree that includes all possible options for each vehicle. The sequence of possible actions within that time period can each correspond to a path in the tree from the root node to the corresponding leaf node. As discussed herein, a value function can then be used to determine the value of each path or sequence of actions. The path leading to the highest-value leaf node can then be selected as the vehicle's five- to ten-second path. The data for the selected path can then be fed to the optimizer, for example, to determine the vehicle's next action or set of actions, such as turning the wheels, accelerating, decelerating, activating a steering signal, etc. The value function can be used to determine its value using several factors as discussed herein, and may include penalties for options that would result in a collision, a violation, or a rapid acceleration change that could be unpleasant for passengers. The function may also include rewards for moving toward the destination, successfully changing lanes, avoiding collisions, providing smooth driving, maintaining a safe distance, and other such options. Various value functions that consider different value criteria can be used, and different weights can be applied to different value criteria depending on the current situation, which can benefit from machine learning in at least some embodiments.
[0026] However, considering all possible actions of all potentially relevant actors could lead to processing vast amounts of data, potentially requiring significant resources and resulting in longer decision-making times, which may be undesirable in many cases. Therefore, approximation or abandonment options can be considered, with minimal impact on the final decision. For example, the space on the current road could be divided into a grid of 200 or a cell array, such as... Figure 2AAs shown. For example, the grid can be discretized into a fixed number of locations per lane and have a cell size that can be a fraction of the average vehicle size, such as half the average vehicle length. This method can be used to determine which cells a given actor will occupy at a given point in time, requiring far less data than attempting to track actual locations. Furthermore, the set of options for each actor can also be discretized. Figure 2B An example movement option grid 250 is shown, which can be used according to various embodiments. It should be understood that grids of other sizes with other options may also be used within the scope of the various embodiments. In this example, the movement options along the horizontal axis (in the figure) are turn left (L), go straight (S), or turn right (R). The movement options along the vertical axis (in the plane of the figure) are accelerate (A), maintain current speed (M), or decelerate (D). Thus, the actor's possible options can be simplified to a set of nine possible actions, such as accelerate and turn right (AR), maintain current direction and speed (MS), etc. At each level of the decision tree, a given node may have nine branches, each corresponding to a potential movement option. As discussed herein, a probability can be determined for each of those options, which can be a factor in determining the value of a given branch.
[0027] This approach can still result in very large decision trees and a large amount of data. Therefore, in at least some embodiments, a subset of these options can be ignored or discarded, where these options are unlikely to affect the choice of path or action. While analyzing more actors may lead to more accurate predictions and determinations, the decision will be most influenced by nearby or adjacent vehicles, such as those directly in front of or behind the current vehicle, and vehicles in nearby lanes that may be affected by any lane changes by the current vehicle. Therefore, if a grid approach is used, it may be meaningful to consider only up to eight other actors, including other actors in front of, behind, to the sides of, and possibly diagonally opposite the current vehicle in the grid. This approach can significantly reduce the amount of data to be considered and the size of the decision tree, except that only a reasonable amount of data on these actors may be collected due to the relative proximity of the current vehicle. Furthermore, the actions of other actors will more directly affect one or more monitored vehicles and can then be used to adjust the path determination of the current vehicle. In some embodiments, the weight of a car, such as the car in front, can be greater in the determination because its actions may have a greater impact on the path decision than those of cars behind and in different lanes. If data on nearby vehicles (or other actors) is available and their next move can be determined, then no prediction is needed, and a single node at that level in the tree is used, corresponding to the vehicle's next action.
[0028] Similarly, in at least some embodiments, if various path options have very low probability values, such as less than a minimum probability threshold, they can be discarded. However, in other cases, low-probability scenarios may be important for planning because the outcome could be very negative, potentially minimizing the risk objective. In these embodiments, low-probability paths with standard or positive outcomes can be advantageously pruned. In one example, if car 204 in the right lane does not approach from the exit ramp or turning option, the probability that car 204 will turn right along that path at the next time point is very low. Therefore, all options involving right turns can be excluded from consideration, eliminating those branches from the decision tree. Furthermore, since car 202 is currently in front of car 204 and heading to its right, and car 204 has a collision avoidance objective, the probability that car 204 will accelerate and turn left at the next time point is very low. Therefore, this path option can also be excluded from consideration (at least for path planning purposes).
[0029] However, the probability of a given driver choosing any path option can also depend on one or more aspects of the driver. For example, an intoxicated driver is more likely to take any action. An aggressive driver is more likely to increase speed, fail to maintain a safe distance, or try to block any attempt to move in front of them. A cautious driver may be more likely to try to move away from other vehicles, meaning they are more likely to slow down or change lanes away from vehicles. Various other characteristics or behaviors can also be observed. Therefore, in at least some embodiments, it may be possible to attempt to classify the drivers of nearby vehicles to improve probability determination. For example, this could include monitoring data captured for these vehicles over a period of time and using that data to attempt to perform an accurate characterization. In some cases, machine learning can be used to attempt to classify drivers more accurately based on available information, such as speed, distance maintained, lane change frequency, direction changes, etc. If the vehicles are able to communicate, information provided by other vehicles may be helpful for classification.
[0030] The result could then be a tree structure of 300, for example. Figure 3 As shown in the example. In this example, there are many levels of nodes, each non-leaf node having many branches extending from it. Each branch corresponds to the path options discussed herein. In the example, only a single path node is shown. The root node 302 represents the vehicle's current position and may also reference other information such as current speed or acceleration. The action sequence is considered to be a turning-based game, where the first-level node 304 corresponds to the actions that the current vehicle (shadow) can take. For example, this could include a path with up to nine possible motion actions (AR, MS, etc.).
[0031] The next lower-level node 306 will correspond to one or more options that other vehicles (shaded) can take in response to the action taken by the current vehicle in the parent node 304 of the previous level. Thus, if the vehicle moves to the right as shown by parent node 304, the given vehicle (shaded) might take various actions in response, in the illustration being a slight deceleration to give the vehicle more space to change lanes. Other options could include other vehicles accelerating to try to block the lane change, or another lane change, etc. Each of these potential options for other vehicles can then be used as a branch to the corresponding node 306 of that level. The action of the current vehicle can be determined in response to the possible actions of other vehicles as branches to the next level node 308. In some embodiments, the process can continue in multiple levels corresponding to a time scale, for example, in one embodiment each level corresponds to an increment of 0.25 seconds, up to five seconds or ten seconds. The nodes of the final level can then correspond to leaf nodes at the ends of various paths, where path values can be determined. It should be understood that leaf nodes may also exist at other levels, such as when a vehicle might reach its destination, collide, or reach other endpoints along a given path. As described above, the highest value leaf can be determined, and the corresponding sequence of actions for that path is provided as a five- to ten-second plan, which in some embodiments can be provided to an optimizer or controller to determine the action to be taken in the next shorter time interval (e.g., the next 50ms). The optimizer can take the path data, smooth the actions, and generate a trajectory for the vehicle with finer granularity. In some embodiments, the time interval output by the five- to ten-second planner can be variable and can be at least in part based on the type of action to be taken or various environmental factors.
[0032] In some embodiments, deep learning can be used to make trees more efficient, as discussed in more detail elsewhere in this document. In one embodiment, a policy function (part of a policy network) can be used to predict the best or most likely option at each level, thus eliminating the need to explore all options further. This can make the tree narrower, requiring fewer paths to be considered. In another embodiment, deep learning can be used to predict the value function of a path without having to expand the tree to individual leaf nodes. This can make the tree shallower, without having to consider all the data for every path. In some embodiments, optimization can be performed to avoid repetition of nodes, paths, or branches. For example, a sequence of moving right and accelerating 3 times might have the same result as accelerating 3 times and then moving right. Therefore, these paths might be able to collapse into a single path for consideration.
[0033] Furthermore, as mentioned elsewhere in this paper, at least some objects may be able to transmit data related to path planning. This could include, for example, data provided by a vehicle regarding its expected actions over at least one or more time periods. For instance, a vehicle might transmit that it intends to turn right within one mile, intends to move to the right lane within the next half mile, and will begin moving right within the next 50 milliseconds. If this information is available, it can be used to improve the accuracy of vehicle action prediction (since vehicles may not always perfectly follow their intentions), thereby improving the accuracy of path planning for the current object. Any such data can be provided to the action prediction and / or path planning models discussed in this paper.
[0034] Further details about the example implementation are provided in the following examples. Return to Reference Figure 1A The current vehicle 102, being navigated, is in the middle lane, and other vehicles are approaching it on the road. The current vehicle 102 wants to change lanes to the right, as commanded by a higher-level planning system, to leave the road within half a mile. To determine the appropriate action or movement to take, the current vehicle can search a planning range. In this example, the process searches forward over the next five seconds, with a step size of 0.25 seconds for the first two seconds and 0.5 seconds for the next three seconds, for a total of ten steps. A decision tree can be generated, where even-numbered levels correspond to the potential movements of the current vehicle, and odd-numbered levels correspond to the response movements of other nearby actors.
[0035] As previously mentioned, path planning benefits from the prediction of other vehicle movements. The reactions of these other vehicles can be predicted at each time step to calculate the “optimal” path based on currently available sensor data. A series of “movements” that enable a lane change can be determined. In this example, it can be determined that the vehicle should decelerate, move to the right, and signal the expected rightward movement. Predictions may indicate that cars behind and to the right might decelerate but remain on the right. After a successful lane change, the grid can be recentered on the current vehicle 102. This approach involves only about four steps because each step involves two frames for the current vehicle and the corresponding responses of one or more other vehicles. Other vehicles may not be this cooperative. Cars in the right lane might not allow the vehicle to change lanes at the expected location. Cars behind might not decelerate. The current vehicle 102 needs to respond to these possibilities in subsequent planning cycles.
[0036] During each tree search, the planner can generate many possible movements for the current car at each even-numbered level of the tree. In this example, the current vehicle starts from its initial position doing nothing (maintaining speed), can turn left, can turn right while maintaining speed, or can turn right while decelerating. In some embodiments, a "movement generator" deep neural network (DNN) can be used to generate three or four "maximum" movements to explore, which can be trained based on many instances of vehicle motion data. Similarly, each other vehicle or actor can respond in multiple ways. The "action generator" DNN can generate the most likely movements for non-critical actors and up to three or four most likely movements for critical actors (such as actors directly adjacent to the current vehicle). At each step of the tree, a second DNN can be used to assign a value to that position. Large negative values can be associated with any contact, while smaller negative values are associated with being too close or failing to maintain at least a specified distance or separation from another actor or object. Positive values can be associated with achieving their respective goals, such as successfully entering the right lane.
[0037] As previously mentioned, different drivers may react differently. Therefore, methods according to various embodiments can attempt to characterize other nearby vehicles by observing their behavior, which may correspond to the driver's behavior (or the potential navigation system, assuming some may be programmed or cause different behaviors). Some actors may be characterized as "aggressive," e.g., unlikely to cooperate in any desired movement, while others may be characterized differently, e.g., "cautious," "cooperative," "intoxicated," "timid," or "unsteady." There are continuous driving styles, which can be exemplified by a set of scalar or value parameters, although in some embodiments, classifying actors into discrete numbers of types or categories allows the use of a small number of "motion generators," one for each category, to suggest how each other vehicle will respond to the current vehicle's movement. In at least some embodiments, scalar values can be more general and easier to fit. As mentioned above, in some embodiments, these motion generators may correspond to trained machine learning models capable of inferring actions based on determined classifications and current sensor data, etc. With sufficient observation of another vehicle, the motion generator may be able to "fine-tune" that vehicle, e.g., by interpolating between standard types of motion generators. The current vehicle can choose its motion by backing up the values calculated at the leaf nodes to the node directly below the root node. In some embodiments, the highest value can be selected at the current vehicle's node, and a weighted average of the most likely motions can be selected for the nodes of other vehicles.
[0038] At each even-numbered level of the decision tree, a motion generator network can be used to suggest the optimal motion for the current object (e.g., a vehicle). The network can accept the currently occupied grid along with the corresponding speed and objective as input. In some embodiments, the speed can be encoded using multiple copies of the grid: one for an object with a certain vehicle speed, one for an object moving a certain amount faster (e.g., at least 5 mph), one for an object moving a certain amount slower (e.g., at least 5 mph), and so on. The objective can be specified at higher levels of the planning hierarchy. In this example, the objective is to change lanes to the right within a certain objective distance. Other objectives might be changing lanes to the left, leaving (left or right), maximizing speed (in any lane), or turning at an intersection (left or right), and so on.
[0039] In some embodiments, the output of the motion generator network is a set of possible motions. Thresholding the softmax layer of the network can be used to select the best motion M for any M. Motion can be defined as a direction and velocity pair, or a direction and acceleration pair, and so on. For example, (right, brake) could indicate moving to the right and decelerating. Each may also have intensity levels, such as "right++" for a strong right turn, "brake--" for a very light application, etc. In some embodiments, the motion generator network is trained by reinforcement learning to learn the motion most likely to provide the maximum "reward" under a given state and a given objective. Direct reinforcement learning-based solutions to path planning problems can directly use the motion generator to select motions. In some embodiments, the motion generator can be augmented with tree search based on policy networks, as discussed herein, to generate better motions than directly using reinforcement learning schemes. Tree search can be used as a multimodal probability determination method, which can determine the risk and uncertainty of potential response actions. This approach can be beneficial for environments with coupled degrees of freedom, as it can correspond to the reactive motions or actions of actors in the environment. Similar motion generator networks can be used for other vehicles or actors. To limit the search, some embodiments consider only key vehicles or actors, while in other embodiments, to improve the accuracy of the determination, actors or objects that may not be visible can be included. As described above, a single most probable motion can be determined for non-key vehicles, and a small number (e.g., two or three) of possible motions can be determined for "key" vehicles. Multiple motion generator networks (or "personality" inputs can be given to a network) can be trained to model different types of drivers (normal, aggressive, cautious, inattentive, etc.). Other vehicles can be classified based on observations to determine which network (or personality) to model for them.
[0040] As described above, the current vehicle path determination system or manager can select motions or actions that maximize the desired value function. Various value functions can be utilized, and changing these functions can alter the vehicle's behavior. Example value functions can include various terms, such as a goal, a small positive value for achieving the goal (e.g., entering the right lane within 300 meters). A progress term can result in a very small positive value for moving towards the destination. A collision term can be utilized, where a very large negative value reduction can be applied to contact objects. This value can increase by an amount proportional to the square of the contact velocity (approximately the damage) and can be scaled according to the type of object, such that contact with a person results in a very large value reduction, while the penalty for hitting a guardrail might be significantly reduced. A proximity term can result in a negative value applied for getting too close to another object, the magnitude of which corresponds to the proximity to the other object. A smoothing term can result in a small negative value applied to extreme controlled motions, such as rapid acceleration or slamming on the brakes. A legal term can result in different negative values applied to violations of laws, the magnitude of which depends on the law. For example, running a red light might incur a more severe penalty than exceeding the speed limit by 1 mph. Penalties can be selected so that the vehicle violates laws (or at least some laws) when collision avoidance is required. In some embodiments, the laws that can be broken may also depend on the type of collision or the object of the collision.
[0041] The value of a state can be a value calculated for that state plus a discount value for future states. In some embodiments, the value of a state (and thus the state to which it moves) can be calculated by traversing the tree to a selected depth (e.g., to level 10). At the leaves of the tree, a “value network” can be used to estimate the future contribution to the value function, taking the current state and the goal as input and returning a value. The leaf value is added to this future estimate to the calculated value of the leaf state to give the value of the leaf node. At the internal nodes of other vehicle movements, future values can be calculated by summing the values of the sub-states, weighted by the probabilities of those states. In this way, many possible futures are considered based on the movements of other vehicles. This contrasts with adversarial games that employ a minimum value approach.
[0042] Alternatively, the value of a substate that maximizes the value of another vehicle, whose value function is known, can be chosen. However, employing weighted probabilities allows for the avoidance of states where other vehicles might (even if unlikely) take actions that could result in large negative rewards (e.g., collisions). At the internal nodes of the current vehicle's motion, future values can be computed by taking the maximum value among the evaluated motions. In at least some embodiments, the path determination system will always select the motion with the highest value based on the chosen value function. The value function can be manually programmed, indicating a statement of what the vehicle is attempting to accomplish. The value network (for each individual) can be trained using reinforcement learning. As previously mentioned, different "individuals" of the value and motion networks can be trained by varying the terms of the value function. Very aggressive drivers will have high values for "progress" and "goal," while the values of other terms will decrease. Cautious drivers, on the other hand, will do the opposite.
[0043] The example vehicle can utilize a hierarchical planner structure. At the top level, the route planner can select a route to reach the target destination, corresponding to a series of road segments connected by intersections. The time range for this level is from minutes to hours, with each segment in this example taking at least tens of seconds. The next level of planning will involve lane selection and turning, ensuring the vehicle is in the correct lane at each turn along the route before executing the turn. This may have a similar time range. The short-range (i.e., 5s-10s timescale) planner described above applies to lane selection and turning. For example, it determines how the vehicle should perform higher-level objectives with a granularity of 0.25 seconds. A fast-response path optimizer (or other controller) operates at a level below the short-range planner. It can run every 50 milliseconds, with a time range of 1 second, and can quickly respond to emergencies (e.g., applying the brakes when a child is running in front of the vehicle), optimizing the path selected by the short-range planner, drawing smooth paths to avoid obstacles, and determining precise control values.
[0044] Various scenarios can be used to train motion and value networks, as well as to evaluate the effectiveness of these networks. This can include directed and randomized testing. Directed testing involves pre-determined scenarios designed to challenge the system. One such scenario involves uncooperative lane changes, where a vehicle attempts to change lanes to the left or right, with dense, “uncooperative” other vehicles in the target lane. Another example scenario involves merging into dense traffic, where the length of the merging lane and the aggressiveness of other vehicles are variable. Other possible scenarios include merging into dense traffic at a roundabout, turning right into a lane of moving traffic, or considering sudden obstacles, such as cars, bicycles, pedestrians, or animals suddenly entering the road in front of the vehicle from the left or right. In a two-lane scenario, the distance (or time) to the entry point, as well as the presence and distance (or time) of oncoming vehicles, can vary. Another test scenario could involve exposed obstacles, such as a truck suddenly exiting the lane in front of the current vehicle to expose a parked car ahead, which can alter the amount of open space for escape. Another situation involves turning left in oncoming traffic, where the statistics on the spacing between oncoming vehicles and their speeds may vary.
[0045] In some embodiments, the motion and value networks can be trained through a process such as self-play. Self-play can simulate an N-player game instead of a 2-player game. Many simulated vehicles are placed on a simulated road "grid". The weights of the value function can be varied for each vehicle to provide a personalized mix. Each vehicle can operate as if it were currently being controlled, performing its own tree search at each time step and modeling other vehicles without knowing their actual value functions. Simple "courses" can be used for self-play involving merging, roundabouts, turning left in oncoming traffic, etc.
[0046] Figure 4An example process 400 for determining navigation actions for a vehicle, which can be used according to various embodiments, is illustrated. It should be understood that, for this process and other processes discussed herein, unless otherwise stated, additional, alternative, or fewer steps may exist within the scope of the various embodiments, performed in a similar or alternative order or in parallel. In this example, the positions of one or more vehicles in the environment 402 can be sensed. This may include, for example, determining the relative positions of one or more other vehicles relative to the first vehicle to be navigated. Other information may also be determined, such as the speed and direction of movement of other vehicles, the presence of road signals or traffic lights, brake lights, turn signals, the presence of pedestrians, the presence of construction zones, road conditions, or weather conditions, as may be determined using one or more vehicle sensors as discussed herein or obtained from other suitable sources. Based at least in part on the determined vehicle characteristics, one or more sequences of possible actions 404 can be determined, wherein these sequences include alternating levels of possible actions of the first vehicle and possible response actions of other vehicles. As mentioned above, in some embodiments, only the most probable response actions may be considered; in other embodiments, this sequence may be used to determine one or more possible navigation paths for the first vehicle 406, which may include, for example, a sequence of actions with the lowest probability. A value function can be used to calculate path values for at least some of these possible navigation paths. As previously described, the value function may include terms such as collision, proximity, and progress, which can be used to determine their respective path weights. One of the navigation paths can then be selected, having the highest calculated path value. At least a first action can be provided from the selected navigation path to the controller of the first vehicle. This action may be one of a set of discrete possible actions that can be provided to the controller's optimizer to determine the action to be taken to manipulate the first vehicle. The first vehicle can then be manipulated or otherwise make navigation decisions based at least in part on the first action.
[0047] Figure 5Another example process 500 for determining navigation actions of a vehicle, which can be utilized according to various embodiments, is illustrated. In this example, 502 position and motion data of a first vehicle and at least a second vehicle are determined. One or more motion generators (e.g., may include neural networks) can be used to generate 504 a decision tree, wherein the decision tree includes alternating levels of actions of the first vehicle and response actions of at least the second vehicle. 506 Probabilities of response actions can be determined at each level, wherein only actions with at least the minimum probability can be considered. In various embodiments, this is performed using a policy network that uses representations of various other objects to propose one or more relevant policies using the acceptance of representation classification, parameters, or scalars to determine the likely response actions of those vehicles. In various embodiments, this can approximate a Monte Carlo-based approach to determine the probabilities of various actions. 508 A path value for each path in the decision tree can be computed using a selected value function. 510 The path with the highest path value in the decision tree can be selected based on the selected value function. This path is based not only on the goal of the first vehicle but also on the possible response actions of at least the second vehicle. 512 At least the next action from the proposed path can be provided to an optimizer configured to manage vehicle functions. Then the first vehicle, 514, can be manipulated as determined by the optimizer.
[0048] Figure 6An example environment 600 is shown that can be used to implement aspects of the various embodiments. In many embodiments, all components will be contained within the vehicle 602 itself to avoid network or connectivity issues for security-sensitive operations. In other embodiments, at least some components may be in separate systems, but may communicate directly via wired or wireless communication instead of passing communication over a network. In some embodiments, the vehicle 602 may be an autonomous vehicle or other type of vehicle or object that can be at least partially autonomously controlled. The vehicle may be any suitable object capable of at least some type of motion or control, such as including autonomous vehicles, robots, unmanned aerial vehicles, etc. In some embodiments, at least some navigation instructions may be determined using a separate user device, such as including desktop computers, laptops, smartphones, tablets, computer workstations, game consoles, etc. The vehicle may include one or more sensors 604 capable of sensing data about the environment and other vehicles, actors, or objects in the vicinity of the vehicle. These sensors may include, for example, cameras, infrared sensors, motion detectors, accelerometers, electronic compasses, LiDAR devices, radar, computer vision modules, odometers, etc. As mentioned above, data may also be obtained from other sources, such as nearby vehicles that are able and / or permitted to share data. Data can be fed into a control system 606, which can be used to control the vehicle, such as changing direction, accelerating or decelerating, activating a turn signal, honking the horn, or performing another such action. In at least some embodiments, the control system may include a user interface that allows a user, such as a human passenger, to modify one or more aspects of vehicle operation. The vehicle will typically include one or more computer processors 608 and a memory 610, which includes instructions executable by the processors for making decisions about the vehicle and / or formulating those decisions to control the vehicle. In at least some embodiments, data captured by sensors or captured data about the operation of the vehicle 602 may be stored in a local database 612.
[0049] As described above, in some embodiments, all determinations can be made on a controllable object (e.g., an autonomous vehicle). In some embodiments, model training can be performed remotely, and the trained model is provided to the object for use. In some embodiments, long-term planning can be performed remotely, making short-term decisions about the vehicle. In other embodiments, all path planning decisions can be made remotely using data collected from the object and from other objects or sources, and these decisions are fed into the control system on the object. Various other options for partitioning functionality among the object and one or more other computing devices or systems can also be utilized within the scope of the various embodiments.
[0050] In some embodiments, sensor data captured by sensors 604 of vehicle 602 can be processed on a client device to determine navigation actions as discussed herein. In other embodiments, sensor data can be transmitted via at least one network 614 for reception by a remote computing system, such as one that may be part of resource provider environment 616. The software architecture in environment 616 may also be implemented in the vehicle or on a separate computing device, etc. At least one network 614 may include any suitable network, including intranet, Internet, cellular network, local area network (LAN), or any other such network or combination, and communication on the network may be enabled via wired and / or wireless connections. Provider environment 616 may include any suitable components for receiving requests and returning information or performing actions in response to those requests. For example, provider environment may include a web server and / or application server for receiving and processing requests and then returning data or other content or information in response to those requests.
[0051] Received communication to provider environment 616 can be received by interface layer 618. Interface layer 618 may include application programming interfaces (APIs) or other exposed interfaces that enable users to submit requests to provider environment. Interface layer 618 in this example may also include other components, such as at least one web server, routing components, load balancers, etc. Components of interface layer 618 can determine the type of request or communication and can direct the request to the appropriate system or service. For example, if the communication is for training a motion neural network for a specific type of vehicle, the communication can be directed to navigation manager 320, which may be a system or service provided using various resources of provider environment 616. The request can be directed to training manager 622, which can select an appropriate model or network and then train the model using the relevant training data 624. Once the network is trained and successfully evaluated, the network can be stored in model repository 626, which may, for example, store different models or networks for different types of vehicles. If a request including sensor data for vehicle 602 is received, the information in that request can be directed to tree management component 628, which can obtain the corresponding trained network. Tree management component 628 can then generate a decision tree using sequences of possible actions and reactions, and generate a score for each sequence using a selected value function. As discussed elsewhere herein, tree search (including inference for motion generation and value determination) is a different process than training. In many cases, tree search and inference will run on the vehicle, rather than in a separate system or the cloud. A path can be selected using the highest path score, and the next option is provided to optimizer 630, which in at least some embodiments may also be located on the vehicle. Optimizer 630 (which may also be within control system 606) can provide navigation actions that can be used to control the vehicle and guide it along the selected path.
[0052] In various embodiments, processor 608 (or the processor of training manager 622 or tree search module 628) will be the central processing unit (CPU). However, as previously mentioned, resources in such environments can utilize GPUs to process at least some types of requested data. GPUs, with thousands of cores, are designed to handle large amounts of parallel workloads and have therefore become popular in deep learning for training neural networks and generating predictions. While using GPUs for offline building can enable the training of larger and more complex models faster, generating predictions offline means that input features cannot be used at request time, or predictions must be generated for all permutations of features and stored in a lookup table to serve real-time requests. If the deep learning framework supports CPU mode and the model is small and simple enough that feedforward can be performed on the CPU with reasonable latency, then a service on a CPU instance can host the model. In this case, training can be done offline on the GPU, and inference can be done in real time on the CPU. If the CPU approach is not a viable option, then the service can run on a GPU instance. However, because GPUs have different performance and cost characteristics compared to CPUs, running a service that offloads runtime algorithms to the GPU may require designing it differently from a CPU-based service.
[0053] As described above, the various embodiments utilize machine learning. For example, deep neural networks (DNNs) developed on processors have been used in a wide range of use cases, from self-driving cars to faster drug development, from automatic image annotation in online image databases to intelligent real-time language translation in video chat applications. Deep learning is a technique that mimics the neural learning process of the human brain, continuously learning, becoming smarter, and providing more accurate results faster over time. Just as a child is initially taught by an adult how to correctly identify and classify various shapes, eventually becoming able to recognize shapes without any guidance, a deep learning or neural learning system needs to be trained on object recognition and classification as it becomes smarter and more efficient at recognizing basic objects, occluded objects, and so on, while also assigning context to objects.
[0054] In various embodiments, a central model can be trained and propagated outward to various vehicles or objects for path planning and prediction. As described above, in embodiments utilizing continuous learning, data from vehicles can be fed back to edge servers or a central server, for example, to further train one or more central models, which can then be propagated to various vehicles for future determination.
[0055] At its simplest level, neurons in the human brain examine various inputs they receive, assign a hierarchy of importance to each of these inputs, and pass the outputs to other neurons for further processing. Artificial neurons, or perceptrons, are the most basic model of neural networks. In one example, a perceptron might receive one or more inputs representing various features that the perceptron is training to recognize and classify objects, with each feature assigned a specific weight based on its importance in defining the shape of the object.
[0056] Deep neural network (DNN) models consist of multiple layers connecting perceptrons (e.g., nodes), which can be trained with large amounts of input data to solve complex problems quickly and with high accuracy. In one example, the first layer of a DLL model breaks down an input image of a car into its parts and looks for basic patterns such as categories like lines and angles. The second layer assembles the lines to look for higher-level patterns, such as wheels, windshields, and mirrors. The next layer identifies the vehicle category, and the last few layers generate labels for the input image, identifying specific car brand models. Once trained, a DNN can be deployed and used to identify and classify objects or patterns in a process called inference. Examples of inference (the process by which a DNN extracts useful information from a given input) include recognizing handwritten digits on a check deposited into an ATM, recognizing an image of a friend in a photograph, providing movie recommendations to over 50 million users, identifying and classifying cars, pedestrians, and road hazards in different types of self-driving cars, or converting human speech in real time.
[0057] During training, data flows through the DNN in the forward propagation phase until a prediction corresponding to the label of the input is produced. If the neural network does not correctly label the input, the error between the correct label and the predicted label is analyzed, and the weights of each feature are adjusted in the backpropagation phase until the DNN correctly labels the inputs and other inputs in the dataset during training. Training complex neural networks requires significant parallel computational performance, including support for floating-point multiplication and addition. Inference is less computationally intensive than training and is a latency-sensitive process where a trained neural network is applied to new inputs it has never seen before, such as classifying images, converting speech, and typically inferring new information.
[0058] Neural networks heavily rely on matrix mathematical operations, and complex multi-layered networks require significant floating-point performance and bandwidth to improve efficiency and speed. With thousands of processing cores optimized for matrix mathematical operations and delivering tens to hundreds of TFLOPS of performance, computing platforms can provide the performance required for deep neural network-based artificial intelligence and machine learning applications.
[0059] Figure 7An example system 700, according to various embodiments, is illustrated for classifying data or generating inference. It should be apparent from the teachings and suggestions contained herein that various predictions, labels, or other outputs can also be generated for the input data. Furthermore, supervised and unsupervised training can be used in the various embodiments discussed herein. In this example, a set of classified data 702 is provided as input to be used as training data. The classified data may include instances of at least one type of object for which a statistical model is to be trained, as well as information identifying that type of object. For example, the classified data may include a set of images, each containing a representation of an object type, wherein each image also contains labels, metadata, classifications, or other information identifying or associated with the object type represented in the respective image. Various other types of data may also be used as training data and may include text data, audio data, video data, etc. In this example, the classified data 702 is provided as training input to a training manager 704. The training manager 704 may be a system or service including hardware and software, such as one or more computing devices executing a training application for training a statistical model. In this example, the training manager 704 receives instructions or requests indicating the type of model to be used for training. The model can be any suitable statistical model, network, or algorithm applicable to such purposes, such as artificial neural networks, deep learning algorithms, learned classifiers, Bayesian networks, etc. The training manager 704 can select a base model or other untrained models from an appropriate repository 706 and train the model using the classified data 702, generating a trained model 708 that can be used to classify similar types of data. In some embodiments where classified data is not used, it is still possible to select an appropriate base model from the training manager to train the input data.
[0060] Models can be trained in a variety of ways, which may depend in part on the type of model chosen. For example, in one embodiment, a set of training data can be provided to a machine learning algorithm, where the model is an artifact created through a training process. Each instance of the training data contains the correct answer (e.g., a classification), which may be referred to as the target or target attribute. The learning algorithm finds patterns in the training data that map the input data attributes to the target, the answer to be predicted, and outputs a machine learning model that captures these patterns. The machine learning model can then be used to obtain predictions for new data without a specified target.
[0061] In one example, the training manager can select from a set of machine learning models, including binary classification, multi-class, and regression models. The type of model used can depend at least in part on the type of target to be predicted. Machine learning models for binary classification problems predict binary outcomes, such as one of two possible classes. Learning algorithms (such as logistic regression) can be used to train binary classification models. Machine learning models for multi-class classification problems allow predictions to be generated for multiple classes, such as predicting one of more than two outcomes. Multinomial logistic regression can be useful for training multi-class models. Machine learning models for regression problems predict numerical values. Linear regression is useful for training regression models.
[0062] To train a machine learning model according to one embodiment, the training manager must determine the input training data source and other information, such as the names of the data attributes containing the target to be predicted, the required data transformation instructions, and training parameters to control the learning algorithm. During training, in some embodiments, the training manager may automatically select an appropriate learning algorithm based on the target type specified in the training data source. The machine learning algorithm may accept parameters for controlling certain attributes of the training process and the resulting machine learning model. These are referred to herein as training parameters. If no training parameters are specified, the training manager can utilize known default values to handle a wide range of machine learning tasks well. Examples of training parameters for which values can be specified include maximum model size, maximum number of passes on the training data, shuffle type, regularization type, learning rate, and regularization amount. Default settings can be specified, with options for adjusting values to fine-tune performance.
[0063] The maximum model size is the total size (in bytes) of patterns created during model training. By default, a model of a specified size can be created, such as a 100MB model. A smaller model can be created if the training manager cannot determine enough patterns to fill the model size. If the training manager finds that the number of patterns exceeds what the specified size can hold, a maximum cutoff can be enforced by trimming the patterns that have the least impact on the quality of the learned model. Choosing a model size allows control over the trade-off between the model's predictive quality and its cost of use. A smaller model may cause the training manager to remove many patterns to fit the maximum size limit, thus affecting the quality of predictions. On the other hand, a larger model may be more costly to query real-time predictions. A larger input dataset does not necessarily result in a larger model, as the model stores patterns rather than input data. If the patterns are few and simple, the resulting model will be small. Input data with a large number of original attributes (input columns) or derived features (outputs of data transformations) may find and store more patterns during training.
[0064] In some embodiments, the training manager may pass or iterate through the training data multiple times to discover patterns. A default number of passes may exist, such as ten, while in some embodiments, a maximum number of passes may be set, such as up to one hundred passes. In some embodiments, there may be no maximum set, or there may be convergence criteria or other sets of criteria that trigger the termination of the training process. In some embodiments, the training manager may monitor the quality of the patterns (i.e., model convergence) during training and may automatically stop training when there are no more data points or patterns to discover. Datasets with only a few observations may require more data traversal to achieve higher model quality. Larger datasets may contain many similar data points, which can reduce the need for a large number of passes. A potential impact of choosing to pass more data is that model training may take longer and incur higher costs in terms of resources and system utilization.
[0065] In some embodiments, training data is shuffled before training or between training passes. In many embodiments, shuffling is a random or pseudo-random shuffle to generate a truly random ordering, although there may be constraints to ensure that certain types of data are not grouped, or if such grouping exists, the shuffled data can be reshuffled, etc. Shuffling alters the sequence or arrangement in which data is used for training so that the training algorithm does not encounter groupings of similar types of data or a single type of data with too many consecutive observations. For example, a model can be trained to predict product types, where the training data includes movie, toy, and video game product types. Before uploading, the data may be sorted by product type. The algorithm can then process the data alphabetically by product type, initially seeing only data for one type (such as movie). The model will begin to learn patterns for movies. The model will then only encounter data for different product types (e.g., toys) and will attempt to adjust the model to fit that toy product type, which may cause patterns that were suitable for movies to degenerate. This abrupt switch from movie to toy types can result in a model that cannot learn how to accurately predict product types. In some embodiments, shuffling can be performed before dividing the training dataset into training and evaluation subsets to utilize a relatively uniform data type distribution for both stages. In some embodiments, the training manager can use, for example, pseudo-random shuffling techniques to automatically shuffle the data.
[0066] In some embodiments, when creating a machine learning model, the training manager allows users to specify settings or apply custom options. For example, users can specify one or more evaluation settings to indicate which portion of the input data should be retained for evaluating the predictive quality of the machine learning model. Users can specify methods that indicate which attributes and attribute transformations can be used for model training. Users can also specify various training parameters that control the training process and certain attributes of the resulting model.
[0067] Once the training manager determines that model training is complete, for example by using at least one of the final criteria discussed herein, a trained model 708 can be provided for classifier 714 to classify unclassified data 712. However, in many embodiments, the trained model 708 will first be passed to evaluator 710, which may include an application or process executed on at least one computational resource for evaluating the quality (or other aspects) of the trained model. The model is evaluated to determine whether it provides at least a minimum acceptable or threshold level of performance when predicting targets for new and future data. Since future data instances will often have unknown target values, it may be desirable to examine machine learning accuracy metrics on data with known target answers and use that evaluation as a proxy for predicting accuracy for future data.
[0068] In some embodiments, a subset of the classified data 702 provided for training is used to evaluate the model. This subset can be determined using the shuffling and splitting methods described above. This evaluation data subset will be labeled with targets and can therefore serve as a resource for evaluating ground truth. It is useless to use the same data used for training to evaluate the predictive accuracy of the machine learning model, as it may produce a positive evaluation for a model that memorizes the training data rather than generalizes from it. Once training is complete, the trained model 708 is used to process the evaluation data subset, and the evaluator 710 can determine the model's accuracy by comparing the ground truth data with the corresponding output (or prediction / observation) of the model. In some embodiments, the evaluator 710 can provide a summary or performance metric indicating how well the predicted values match the true values. If the trained model does not meet at least a minimum performance criterion or other such accuracy threshold, the training manager 704 can be instructed to perform further training, or in some cases, to try training a new or different model, etc. If the trained model 708 meets the relevant criteria, the trained model can be provided for use by the classifier 714.
[0069] When creating and training machine learning models, in at least some embodiments, it is desirable to specify model settings or training parameters that will result in a model capable of making the most accurate predictions. Example parameters include the number of passes to be performed (forward and / or backward), regularization, model size, and shuffling type. However, as mentioned above, selecting model parameter settings that produce the best predictive performance on evaluation data can lead to model overfitting. Overfitting occurs when a model stores patterns that appear in both training and evaluation data sources but fails to generalize the patterns in the data. Overfitting often occurs when the training data includes all the data used in the evaluation. An overfitted model may perform well during evaluation but may not make accurate predictions on new or other categorized data. To avoid selecting an overfitted model as the best model, the training manager can reserve additional data to validate the model's performance. For example, the training dataset may be divided into 60% for training and 40% for evaluation or validation, which may be divided into two or more phases. After selecting model parameters that best suit the evaluation data, resulting in convergence to a subset of the validation data (e.g., half of the validation data), a second validation can be performed using the remaining validation data to ensure the model's performance. If the model meets the expectations of the validation data, it will not overfit the data. Optionally, a test set or holdout set can be used to test the parameters. Using a second validation or testing step helps in selecting appropriate model parameters to prevent overfitting. However, taking more data from the training process for validation reduces the amount of data available for training. This can be problematic for smaller datasets, as there may not be enough data available for training. One approach in this case is to perform cross-validation, as described elsewhere in this article.
[0070] There are many metrics or insights that can be used to review and evaluate the predictive accuracy of a given model. An example evaluation result includes a predictive accuracy metric to report the overall success of the model, as well as visualizations to help explore instances where the model's accuracy exceeds the predictive accuracy metric. The results may also provide the ability to see the impact of setting score thresholds (such as binary classification) and can generate alerts about the criteria used to check the validity of the evaluation. The choice of metrics and visualizations can depend at least in part on the type of model being evaluated.
[0071] After satisfactory training and evaluation, the trained machine learning model can be used to build or support machine learning applications. In one embodiment, building a machine learning application is an iterative process involving a series of steps. The core machine learning problem can be formulated based on observations and the answer the model aims to predict. Data can then be collected, cleaned, and prepared to suit the use of the algorithm trained on the machine learning model. This data can be visualized and analyzed for integrity checks to verify data quality and understanding. This may be a situation where the raw data (e.g., input variables) and the answer (e.g., the target) are not represented in a way that can be used to train a highly predictive model. Therefore, it may be desirable to construct a more predictive representation or feature from the raw variables. The resulting features can be fed into the learning algorithm to build the model and evaluate its quality based on the data retained from the model construction. The model can then be used to generate predictions of the target answer for new data instances.
[0072] exist Figure 7 In the exemplary system 700, after evaluation, a trained model 710 is provided to or made available to a classifier 714, which is capable of using the trained model to process unclassified data. For example, this might include data received from a user or an unclassified third party, such as a query image seeking information about what is represented in these images. The unclassified data can be processed by the classifier using the trained model, and the resulting output 716 (i.e., classification or prediction) can be sent back to the appropriate source, or further processed or stored. In some embodiments, and where such use is permitted, these currently classified data instances can be stored in a classified data repository, which can be used by the training manager for further training of the trained model 708. In some embodiments, the model is trained continuously as new data becomes available; however, in other embodiments, the models are trained periodically, such as daily or weekly, depending on factors such as the size of the dataset or the complexity of the model.
[0073] A classifier may include appropriate hardware and software for processing unclassified data using a trained model. In some cases, a classifier will include one or more computer servers, each with one or more graphics processing units (GPUs) capable of processing data. The configuration and design of GPUs may make them more suitable than CPUs or other such components for processing machine learning data. In some embodiments, a trained model may be loaded into GPU memory, and received data instances may be fed to the GPU for processing. GPUs can have far more cores than CPUs, and GPU cores can be less complex. Therefore, a given GPU may be able to process thousands of data instances simultaneously through different hardware threads. GPUs can also be configured to maximize floating-point throughput, which can provide a significant additional processing advantage for large datasets.
[0074] Even when using GPUs, accelerators, and other such hardware to accelerate tasks such as model training or data classification using such models, these tasks can still require significant time, resource allocation, and cost. For example, if a machine learning model is to be trained using 100 passes, and the dataset includes 1,000,000 data instances to be used for training, each pass would require processing all millions of instances. Different parts of the architecture can also be supported by different types of devices. For example, training can be performed using a set of servers in a logically centralized location, as can be provided as a service, while the classification of the raw data can be performed by such a service or on client devices, among other such options. In various embodiments, these devices can also be owned, operated, or controlled by the same entity or multiple entities.
[0075] Figure 8An example statistical model 800 that can be utilized according to various embodiments is shown. In this example, the statistical model is an artificial neural network (ANN) comprising multiple layers of nodes, including an input layer 802, an output layer 806, and multiple layers 804 of intermediate nodes, often referred to as “hidden” layers because inner layers and nodes are typically invisible or inaccessible in conventional neural networks. Other types of statistical models, as well as other types of neural networks including other numbers or choices of nodes and layers, can also be used, as discussed elsewhere herein. In this network, all nodes in a given layer are interconnected to all nodes in adjacent layers. As shown, nodes in intermediate layers are then connected to nodes in two adjacent layers, respectively. In some models, nodes are also called neurons or connected units, and the connections between nodes are called edges. Each node can perform a function for the received input, for example, by using a specified function. Nodes and edges can acquire different weights during training, and the individual layers of a node can perform specific types of transformations on the received input, which can also be learned or adjusted during training. Learning can be supervised or unsupervised, which may depend at least in part on the type of information contained in the training dataset. Various types of neural networks can be utilized, including, for example, convolutional neural networks (CNNs), which consist of many convolutional layers and a set of pooling layers, and have proven beneficial for applications such as image recognition. CNNs are also easier to train than other networks because the number of parameters to be determined is relatively small.
[0076] In some embodiments, various tuning parameters can be used to train such complex machine learning models. Selecting parameters, fitting the model, and evaluating the model are part of the model tuning process, often referred to as hyperparameter optimization. In at least some embodiments, this tuning may include introspection of the base model or data. In training or production settings, a robust workflow is crucial to avoid overfitting of hyperparameters, as described elsewhere in this document. Cross-validation and adding Gaussian noise to the training dataset are useful techniques to avoid overfitting to either dataset. For hyperparameter optimization, in some embodiments, it may be necessary to keep the training and validation sets fixed. In some embodiments, hyperparameters can be tuned in certain categories, such as including data preprocessing (in other words, converting words to vectors), CNN architecture definitions (e.g., filter size, number of filters), stochastic gradient descent parameters (e.g., learning rate), regularization (e.g., dropout probability), and other such options.
[0077] In the example preprocessing step, instances of the dataset can be embedded into a lower-dimensional space of a specific size. The size of this space is a parameter to be tuned. The architecture of a CNN contains many tuned parameters. The parameter of the filter size can represent the interpretation of information corresponding to the size of the instances to be analyzed. In computational linguistics, this is called the n-gram size. The example CNN uses three different filter sizes, which represent potentially different n-gram sizes. The number of filters for each filter size can correspond to the depth of the filters. Each filter attempts to learn something different from the structure of the instances, such as the sentence structure of text data. In the convolutional layers, the activation function can be rectified linear units, and the pooling type is set to max pooling. The results can then be concatenated into a one-dimensional vector, with the final layer fully connected to the two-dimensional output. This corresponds to binary classification, to which optimization functions can be applied. One such function is an implementation of the root mean square (RMS) propagation method of gradient descent, where example hyperparameters can include the learning rate, batch size, maximum gradient normal, and epoch. Neural networks, regularization, can be a very important consideration. As mentioned, in some embodiments, the input data can be relatively sparse. In this scenario, the primary hyperparameters can be discarded at the penultimate layer, meaning a certain percentage of nodes will not "trigger" in each training epoch. The example training process can suggest different hyperparameter configurations based on feedback on the performance of previous configurations. The model can be trained using the suggested configurations, evaluated on a specified validation set, and performance can be reported. This process can be repeated, for example, by balancing exploration (learning more about different configurations) and development (leveraging prior knowledge to achieve better results).
[0078] Because CNN training can be parallelized and can leverage GPU-supported computational resources, various optimization strategies can be tried for different scenarios. Complex scenarios allow for tuning of the model architecture, preprocessing, and stochastic gradient descent parameters. This expands the model configuration space. In the basic case, only the preprocessing and stochastic gradient descent parameters are tuned. In complex scenarios, there are many more configuration parameters than in the basic approach. Joint space tuning can be performed using linear or exponential steps, iteratively through the model's optimization loop. Such tuning processes can be significantly less costly than tuning processes such as random search and grid search without incurring any noticeable performance penalty.
[0079] Some embodiments may use backpropagation to compute gradients used to determine the weights of a neural network. Backpropagation is a form of differentiation, and as described above, gradient descent optimization algorithms can be used to adjust the weights applied to various nodes or neurons. In some embodiments, the gradient of a relevant loss function can be used to determine the weights. Backpropagation may utilize the derivative of the loss function with respect to the output generated by the statistical model. As described above, each node may have an associated activation function that defines the output of that node. Various activation functions may be appropriately used, such as radial basis functions (RBF) and sigmoid functions, which can be used by various support vector machines (SVMs) for data transformation. The activation functions of the intermediate layers of a node are referred to herein as inner product kernels. These functions may include, for example, recognition functions, step functions, sigmoid functions, ramp functions, etc. Activation functions may also be linear or nonlinear, among other such options.
[0080] Figure 9 A set of basic components of a computing device 900 is shown, which can be used to implement aspects of various embodiments. In this example, the device includes at least one processor 902 for executing instructions that can be stored in a memory device or element 904. It will be apparent to those skilled in the art that the device can include many types of memory, data storage, or computer-readable media, such as a first data storage for program instructions executed by the processor 902, the same or separate storage for images or data, removable memory for sharing information with other devices, and any number of communication methods for sharing with other devices. The device will typically include some type of display element 906, such as a touchscreen, organic light-emitting diode (OLED), or liquid crystal display (LCD), but the device (such as a portable media player) may convey information via other means, such as through an audio speaker. As discussed, the device in many embodiments will include at least a communication component 908 and / or a network component 910, such as supporting wired or wireless communication via at least one network, such as the Internet, a local area network (LAN). Or cellular networks, etc. These components enable the device to communicate with remote systems or services. The device may also include at least one additional input device 912 capable of receiving regular input from the user. This regular input may include, for example, buttons, touchpads, touchscreens, scroll wheels, joysticks, keyboards, mice, trackballs, keypads, or any other such devices or elements through which the user inputs commands to the device. In some embodiments, these I / O devices may even be connected via infrared or Bluetooth or other links. However, in some implementations, such a device may not include any buttons at all and may only be controlled by a combination of visual and audio commands, allowing the user to control the device without having to physically touch it.
[0081] Various embodiments can be implemented in a wide range of operating environments. In some cases, the operating environment may include one or more user computers or computing devices that can be used to operate any of a number of applications. User or client devices may include any of a variety of general-purpose personal computers, such as desktop or laptop computers running standard operating systems, as well as cellular, wireless, and handheld devices running motion software and capable of supporting a variety of network and messaging protocols. Such a system may also include multiple workstations running a variety of commercial operating systems and other known applications for purposes such as development and database management. These devices may also include other electronic devices, such as virtual terminals, thin clients, gaming systems, and other devices capable of communicating over a network.
[0082] Most embodiments utilize at least one network familiar to those skilled in the art to support communication using any protocol (e.g., TCP / IP or FTP) from a variety of commercial protocols. The network can be, for example, a local area network (LAN), a wide area network (WAN), a virtual private network (VPN), the Internet, an intranet, an extranet, a public switched telephone network (PSTN), an infrared network, a wireless network, and any combination thereof. In embodiments utilizing a network server, the network server can run any of a variety of server or middleware applications, including HTTP servers, FTP servers, CGI servers, data servers, Java servers, and commercial application servers. One or more servers are also capable of executing programs or scripts in response to requests from user devices, for example by executing one or more web applications, which can be implemented in any programming language (e.g., [example language not specified]). One or more scripts or programs written in C, C#, or C++, or any scripting language such as Python, and combinations thereof. One or more servers may also include a database server, including but not limited to those that can access databases from... and Those servers acquired through commercial purchase.
[0083] The environment may include various data storage devices as well as other storage and storage media as described above. These may reside in various locations, such as on storage media local to (and / or residing in) one or more computers, or remotely from any or all computers on a network. In a particular set of embodiments, information may reside in a storage area network (SAN) familiar to those skilled in the art. Similarly, any necessary files for performing functions belonging to a computer, server, or other network device may be appropriately stored locally and / or remotely. Where the system includes computerized devices, each such device may include hardware elements that can be electrically coupled via a bus, including, for example, at least one central processing unit (CPU), at least one input device (e.g., mouse, keyboard, controller, touch-sensitive display element, or keypad), and at least one output device (e.g., display device, printer, or speaker). Such a system may also include one or more storage devices, such as disk drives, optical storage devices, and solid-state storage devices, such as random access memory (RAM) or read-only memory (ROM), as well as removable media devices, memory cards, flash memory cards, etc.
[0084] Such devices may also include computer-readable storage medium readers, communication devices (e.g., modems, network interface cards (wireless or wired), infrared communication devices), and working memory as described above. Computer-readable storage medium readers may be connected to or configured to receive computer-readable storage media representing remote, local, fixed, and / or removable storage devices, as well as storage media for temporarily and / or more permanently containing, storing, transmitting, and retrieving computer-readable information. Systems and various devices will also typically include multiple software applications, modules, services, or other elements residing within at least one working memory device, including operating systems and applications such as client applications or web browsers. It should be understood that alternative embodiments may have many variations different from the embodiments described above. For example, custom hardware and / or specific elements that may be implemented in hardware, software (including portable software, such as applets), or both may also be used. Furthermore, connections to other computing devices (such as network input / output devices) may be employed.
[0085] Storage media and other non-transitory computer-readable media used to contain code or code portions may include any suitable media known or used in the art, such as, but not limited to, volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules or other data), including RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital universal disc (DVD) or other optical storage, magnetic tape cassettes, magnetic tape, disk storage or other magnetic storage devices, or any other media that can be used to store the desired information and is accessible by system devices. Based on the disclosure and teachings provided herein, those skilled in the art will understand other ways and / or methods for implementing the various embodiments.
[0086] Therefore, the specification and drawings are to be considered illustrative rather than restrictive. However, it will be apparent that various modifications and alterations can be made thereto without departing from the broader spirit and scope of the invention as set forth in the claims.
Claims
1. A processor, comprising: One or more circuits that use one or more neural networks to select the one or more possible movements of the autonomous vehicle based at least in part on scoring one or more predictions of one or more possible movements of one or more other vehicles caused by one or more possible movements of the autonomous vehicle.
2. The processor of claim 1, wherein the one or more neural networks are further configured to generate a decision tree for the autonomous vehicle, the decision tree comprising alternating levels of possible actions of the autonomous vehicle and possible reaction actions of at least one of the other vehicles.
3. The processor as described in claim 2, in, The one or more neural networks are also used to determine possible reaction actions at the levels of the decision tree using a policy network, to consider one or more possible navigation paths.
4. The processor of claim 2, further comprising: The one or more neural networks are also used to estimate the values of nodes at one or more levels of the decision tree.
5. The processor of claim 1, further comprising: The one or more neural networks are also used to provide at least a first action of the selected navigation path to the optimizer of the autonomous vehicle; as well as The autonomous vehicle is then operated based on navigation commands generated by the optimizer.
6. A computer-implemented method, comprising: Using one or more neural networks to select the one or more possible movements of the autonomous vehicle based at least in part on scoring one or more predictions of one or more possible movements of one or more other vehicles caused by one or more possible movements of the autonomous vehicle.
7. The computer-implemented method of claim 6, wherein the one or more predictions are based at least in part on one or more features including at least one of position, velocity, acceleration, direction of motion, or motion characteristics.
8. The computer-implemented method as described in claim 6, further comprising: Provide at least the first action of the selected navigation path to the optimizer of the autonomous vehicle; as well as The autonomous vehicle is then operated based on navigation commands generated by the optimizer.
9. The computer-implemented method of claim 6, further comprising: A decision tree is generated for the autonomous vehicle, the decision tree including alternating levels of possible actions of the autonomous vehicle and possible reaction actions of at least one of the other vehicles.
10. The computer-implemented method of claim 9, further comprising: A policy network is used to determine possible response actions at the levels of the decision tree to consider one or more navigation paths.
11. The computer-implemented method of claim 9, further comprising: A trained neural network is used to estimate the values of nodes at one or more levels of the decision tree.
12. The computer-implemented method of claim 6, further comprising: At least one motion generator is used to determine the predicted movement of at least one of the other vehicles, the at least one motion generator comprising a trained neural network for characterizing the at least one of the other vehicles.
13. The computer-implemented method of claim 6, further comprising: Determine the characterization of one or more other vehicles, the characterization determining the probability of possible responses of the one or more other vehicles.
14. The computer-implemented method of claim 6, further comprising: Determine a threshold for one or more other vehicles, wherein the one or more predictions are determined at least in part based on the threshold.
15. A system comprising: At least one processor; as well as Memory, including instructions, which, when executed by the at least one processor, cause the system to: Using one or more neural networks to select one or more possible movements of the autonomous vehicle, at least in part based on scoring one or more predictions of one or more possible movements of one or more other vehicles caused by one or more possible movements of the autonomous vehicle.
16. The system of claim 15, wherein, when the instructions are executed, the system further causes: The autonomous vehicle uses one or more sensors to sense features of the one or more other vehicles. Using the one or more neural networks, determine one or more possible navigation paths for the autonomous vehicle; as well as The selected navigation path is determined from one or more possible navigation paths, at least in part, based on a value function.
17. The system of claim 16, wherein, when the instructions are executed, the system further causes the system to: The autonomous vehicle is then manipulated according to at least a portion of the selected navigation path.
18. The system of claim 15, wherein, when the instructions are executed, the system further causes: The one or more neural networks are used to generate a decision tree for the autonomous vehicle, the decision tree comprising alternating levels of possible actions of the autonomous vehicle and possible reaction actions of the one or more other vehicles.
19. The system of claim 18, wherein the instructions, when executed, further cause the system to: The one or more neural networks are used to leverage the policy network to determine possible response actions at the levels of the decision tree, considering one or more possible navigation paths.
20. The system of claim 18, wherein the instructions, when executed, further cause the system to use one or more neural networks: Receive intent data corresponding to at least one of the other vehicles; and The intent data is used to determine at least a subset of the possible reaction actions.
Citation Information
Patent Citations
Systems and methods for predicting traffic patterns in an autonomous vehicle
US20180374341A1
Distribution decision trees
US9557737B1