Recursive Multi-Fidelity Behavior Prediction

By adopting a recursive multifidelity framework in agent behavior prediction systems such as autonomous vehicles, the problem of handling state and model uncertainty in the prior art is solved, and more efficient and accurate behavior prediction is achieved.

CN112823358BActive Publication Date: 2025-06-20QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN201980064953.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-10-10
Filing Date
2019-10-11
Publication Date
2025-06-20
Estimated Expiration
2039-10-11

AI Technical Summary

Technical Problem

When the prior art predicts the behavior between agents such as autonomous vehicles, it is difficult to accurately deal with the uncertainty of the state and model, resulting in complex behavior prediction and high computing resources.

Method used

Using a recursive multifidelity framework, the behavior prediction system of the self-agent is reduced by assigning the fidelity level to the agent observed in the scene and recursively predicting its future actions by traversing the scene, reducing memory usage and optimizing the behavior prediction system of the self-agent.

Benefits of technology

It improves the accuracy and efficiency of behavior prediction, reduces the consumption of computing resources, and adapts to the uncertainty of environmental information by dynamically adjusting the fidelity level.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112823358B_ABST
    Figure CN112823358B_ABST
Patent Text Reader

Abstract

A method for predicting a future action of an agent in a scenario includes assigning a fidelity level to the agent observed in the scenario. The method further includes recursively predicting the future action of the agent by traversing the scenario. Different forward prediction models are used at each recursive level. The method further includes controlling the action of a self-agent based on the predicted future action of the agent.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit of U.S. Patent Application No. 16 / 599,078, entitled "RECURSIVE MULTI - FIDELITY BEHAVIOR PREDICTION", filed on October 10, 2019, which claims the benefit of U.S. Provisional Patent Application No. 62 / 744,415, entitled "RECURSIVE MULTI - FIDELITY BEHAVIOR PREDICTION", filed on October 11, 2018. The disclosures of the above - mentioned patent applications are hereby incorporated by reference in their entireties.

[0003] Background

[0004] Field

[0005] Aspects of the present disclosure generally relate to behavior prediction, and more particularly to systems and methods for recursive behavior prediction.

[0006] Background

[0007] A convolutional neural network is a feed - forward artificial neural network. A convolutional neural network may include a collection of neurons, where each neuron has a receptive field and together they tile an input space. Convolutional neural networks (CNNs) have numerous applications. In particular, CNNs have been widely used in the fields of pattern recognition and classification.

[0008] CNNs can also be used for behavior prediction. Autonomous and non - autonomous vehicles can use CNNs to predict the behavior between agents (such as other vehicles). For example, autonomous vehicles use behavior prediction for planning and decision - making. There is a desire to improve behavior prediction systems for tasks such as autonomous driving.

[0009] Summary

[0010] In one aspect of the present disclosure, a method is disclosed. The method can predict the future actions of an agent in a scene. The method includes assigning a fidelity level to an agent observed in the scene. The method further includes recursively predicting the future actions of the agent by traversing the scene. The method also includes controlling the actions of a self - agent based on the predicted future actions of the agent.

[0011] Another aspect of the present disclosure relates to an apparatus that includes means for assigning a fidelity level to an agent observed in a scene. The apparatus further includes means for recursively predicting future actions of the agent by traversing the scene. The apparatus further includes means for controlling the actions of a self-agent based on the predicted future actions of the agent.

[0012] In another aspect of the present disclosure, a non-transitory computer-readable medium having non-transitory program code recorded thereon is disclosed. The program code can predict future actions of an agent in a scene. The program code is executed by a processor and includes program code for assigning a fidelity level to an agent observed in the scene. The program code further includes program code for recursively predicting future actions of the agent by traversing the scene. The program code further includes program code for controlling the actions of a self-agent based on the predicted future actions of the agent.

[0013] Another aspect of the present disclosure relates to an apparatus. The apparatus can predict future actions of an agent in a scene. The apparatus has a memory and one or more processors coupled to the memory. The processor(s) is / are configured to assign a fidelity level to an agent observed in the scene. The processor(s) is / are further configured to recursively predict future actions of the agent by traversing the scene. The processor(s) is / are further configured to control the actions of a self-agent based on the predicted future actions of the agent.

[0014] Additional features and advantages of the present disclosure will be described below. Those skilled in the art should appreciate that the present disclosure can be readily used as a basis for modifying or designing other structures for implementing the same purpose as the present disclosure. Those skilled in the art should also recognize that such equivalent constructs do not depart from the teachings of the present disclosure as set forth in the appended claims. The novel features, which are considered to be characteristics of the present disclosure, will be better understood in conjunction with the accompanying drawings when considered in connection with the following description in terms of both their organization and method of operation, together with further objects and advantages. However, it is to be clearly understood that each drawing is provided for the purpose of illustration and description only and is not intended as a definition of the limits of the present disclosure. Brief Description of the Drawings

[0016] The features, nature, and advantages of the present disclosure will become more apparent when the detailed description set forth below is understood in conjunction with the accompanying drawings, in which like reference numerals throughout the drawings identify corresponding parts.

[0017] Figure 1 An example implementation of designing a neural network using a system-on-chip (SOC) including a general-purpose processor in accordance with certain aspects of the present disclosure is illustrated.

[0018] Figure 2A 、 2B2C is a diagram illustrating a neural network according to aspects of the present disclosure.

[0019] Figure 2D is a diagram illustrating an exemplary deep convolutional network (DCN) according to aspects of the present disclosure.

[0020] Figure 3 is a block diagram illustrating an exemplary deep convolutional network (DCN) according to aspects of the present disclosure.

[0021] Figure 4A 、 4B 4C illustrates an example of recursive multi-fidelity behavior prediction according to aspects of the present disclosure.

[0022] Figure 5A 、 5B 5C illustrates an example of the level of reasoning according to aspects of the present disclosure.

[0023] Figure 6 illustrates an example of a model for predicting a trajectory according to aspects of the present disclosure.

[0024] Figure 7A 、 7B 、7C and 7D illustrate examples of recursive multi-fidelity prediction using gap-thread manipulation according to aspects of the present disclosure.

[0025] Figure 8 illustrates an example of determining the most likely strategy according to aspects of the present disclosure.

[0026] Figure 9A and 9B illustrates an example of determining a strategy according to aspects of the present disclosure.

[0027] Figure 10 illustrates a method for predicting future actions of an agent in a scenario according to aspects of the present disclosure.

[0028] DETAILED DESCRIPTION

[0029] The detailed description set forth below in connection with the appended drawings is intended as a description of various configurations and is not intended to represent the only configurations in which the concepts described herein may be practiced. The detailed description includes specific details for the purpose of providing a thorough understanding of the various concepts. However, it will be apparent to those skilled in the art that these concepts may be practiced without these specific details. In some instances, well-known structures and components are shown in block diagram form to avoid obscuring such concepts.

[0030] Based on this teaching, those skilled in the art should appreciate that the scope of the present disclosure is intended to cover any aspect of the present disclosure, whether implemented independently of or in combination with any other aspect of the present disclosure. For example, any number of the aspects described can be used to implement an apparatus or practice a method. Additionally, the scope of the present disclosure is intended to cover such apparatus or methods practiced using other structures, functionality, or a combination of structures and functionality that supplement or are different from the various aspects of the present disclosure described. It should be understood that any aspect of the present disclosure disclosed can be implemented by one or more elements of the claims.

[0031] The term "exemplary" is used herein to mean "serving as an example, instance, or illustration". Any aspect described herein as "exemplary" need not be construed as superior to or better than other aspects.

[0032] Although specific aspects are described herein, numerous variations and permutations of these aspects fall within the scope of the present disclosure. While some benefits and advantages of the preferred aspects are mentioned, the scope of the present disclosure is not intended to be limited to specific benefits, uses, or objectives. Instead, the aspects of the present disclosure are intended to be broadly applicable to different technologies, system configurations, networks, and protocols, some of which are illustrated by way of example in the figures and the following description of the preferred aspects. The detailed description and the figures merely illustrate the present disclosure and do not limit the present disclosure, the scope of the present disclosure being defined by the appended claims and their equivalent technical solutions.

[0033] Behavior prediction can be used in tasks involving interactions between decision-making agents. For example, autonomous vehicles use behavior prediction for planning and decision-making. In this example, the autonomous vehicle uses a behavior prediction system to predict the behavior of agents in the environment around the autonomous vehicle. The autonomous vehicle can be referred to as an ego-agent. The surrounding environment can include dynamic objects, such as autonomous agents. The surrounding environment can also include static objects, such as roads and buildings.

[0034] One or more sensors (such as light detection and ranging (LIDAR) sensors, radio detection and ranging (RADAR) sensors, cameras, and / or another type of sensor) can be used to capture a temporal snapshot (e.g., a scene) of the environment. The temporal snapshot provides information about the state of the environment at a given moment (such as the positions of dynamic and static objects). Based on the temporal snapshot and one or more previous temporal snapshots, the behavior prediction system calculates multiple scenarios (e.g., possible future evolutions of the scene).

[0035] Behavior prediction can be referred to as the task of estimating the posterior distribution on the future trajectories of one or more dynamic agents in a scene. When planning the behavior of a self-agent, it is desirable to accurately predict the behavior of the agents around the self-agent. Predicting the behavior of surrounding agents can be difficult due to state and model uncertainties. State uncertainty refers to the uncertainty in the position and / or velocity of an agent. Model uncertainty refers to the uncertainty in the model of the reasoning process of an agent.

[0036] The future behaviors of both the self-agent and other agents are interdependent. Due to this interdependence, behavior prediction can be difficult, especially in continuous state spaces, action spaces, and / or observation spaces. For example, highway driving can be modeled using continuous state spaces, action spaces, and observation spaces. When traveling on a highway, a prediction model receives information from continuous observations of the actions of other agents while controlling the actions of the self-agent.

[0037] Aspects of the present disclosure relate to improving behavior prediction by incorporating a multi-fidelity framework into a recursive reasoning scheme. Multi-fidelity refers to predicting behavior at various fidelity levels. By using a multi-fidelity framework with a recursive reasoning scheme, aspects of the present disclosure can reduce memory footprint and reduce power consumption for a behavior prediction system of a self-agent.

[0038] A motion hypothesis refers to a representation of the predicted future trajectory of an agent. A motion hypothesis can be a single trajectory or a distribution over trajectories. An atomic prediction model is a model that receives as input a representation of the history of a scene. The atomic prediction model can also receive as input a representation of the predicted future of the scene. A motion hypothesis for a particular target vehicle is generated by the atomic prediction model. The atomic prediction model can vary in its prediction fidelity level. The atomic prediction model can also be referred to as a policy.

[0039] Each agent in a scene can be assigned an inference level and a set of atomic prediction models. The inference level of an agent is an integer greater than or equal to zero. Each atomic prediction model in the assigned set of atomic prediction models for an agent corresponds to a particular inference level greater than or equal to zero and less than or equal to the assigned inference level of the agent. For each level, a given agent can have no more than one assigned atomic prediction model.

[0040] For each agent with an assigned level-0 atomic prediction model, the recursive inference scheme uses the assigned level-0 atomic prediction model to generate level-0 motion hypotheses. For each agent with an assigned level-1 atomic prediction model, the recursive inference scheme uses the assigned level-1 atomic prediction model to generate level-1 motion hypotheses. A subset of the level-0 motion hypotheses of other agents can be used as input to any of the level-1 prediction models. This process can be repeated such that each successive set of motion hypotheses (level k) can be conditioned on the highest-level (up to k-1) previously computed motion hypotheses of each agent in a subset of other agents in the scene.

[0041] The multi-fidelity framework provides the ability to tune the fidelity of predicting the behavior of each agent. In one configuration, the multi-fidelity framework allows customization of the set of atomic policy models assigned to each agent. It may be desirable to tune the fidelity to match the distribution of uncertainty caused by the sensors used to collect environmental information. For example, for an agent whose motion history is not well defined in the scene history available to the model, a lower-fidelity atomic prediction model may be desirable.

[0042] The multi-fidelity framework can also be used to bias the allocation of computing resources towards agents that are considered to be the most significant. That is, more computing resources can be allocated to agents with higher significance compared to other agents in the environment. For example, the behavior of vehicles adjacent to an autonomous vehicle (e.g., the ego-agent) may be considered more important for the planning process of the prediction model compared to the behavior of vehicles that are farther away. As such, more computing resources can be allocated to processing information related to the vehicles adjacent to the autonomous vehicle compared to the computing resources allocated to processing information related to other objects in the environment. Additionally, agents with higher significance can generate increased amounts of information. As such, additional computing resources can be used to process the additional information of agents with higher significance.

[0043] By conditioning on previously computed motion hypotheses, the model explicitly reasons about future interactions between agents. Higher-level motion hypotheses for a target vehicle can be conditioned on predictions for surrounding vehicles, which in turn can be conditioned on lower-level motion hypotheses for the same target vehicle. The conditioning discussed can be referred to as a recursive scheme.

[0044] It is possible to generate predictions corresponding to multiple distinct possible scenarios (which can be encoded by a scenario tree or scenario forest). Multiple scenario trees can share common prediction nodes and fan out into intermediate and leaf nodes represented by higher-level reasoning.

[0045] Figure 1An example implementation of a system-on-chip (SOC) 100 is described, which may include a central processing unit (CPU) 102 or a multi-core CPU configured for recursive multi-fidelity behavior prediction according to certain aspects of the present disclosure. Variables (e.g., neural signals and synaptic weights), system parameters associated with the computing device (e.g., neural network with weights), latencies, frequency slot information, and task information may be stored in a memory block associated with the neural processing unit (NPU) 108, a memory block associated with the CPU 102, a memory block associated with the graphics processing unit (GPU) 104, a memory block associated with the digital signal processor (DSP) 106, memory block 118, or may be distributed across multiple blocks. Instructions executed at the CPU 102 may be loaded from a program memory associated with the CPU 102 or may be loaded from memory block 118.

[0046] The SOC 100 may also include additional processing blocks customized for specific functions, such as the GPU 104, DSP 106, connectivity block 110 (which may include fifth-generation (5G) connectivity, fourth-generation long-term evolution (4G LTE) connectivity, Wi-Fi connectivity, USB connectivity, Bluetooth connectivity, etc.), and a multimedia processor 112 that can detect and recognize poses, for example. In one implementation, the NPU is implemented in the CPU, DSP, and / or GPU. The SOC 100 may also include a sensor processor 114, an image signal processor (ISP) 116, and / or a navigation module 120 (which may include a global positioning system).

[0047] The SOC 100 may be based on the ARM instruction set. In one aspect of the present disclosure, the instructions loaded into the processor 102 may include code for assigning a fidelity level to an agent observed in a scene. The processor 102 may also include code for recursively predicting the future actions of the agent by traversing the scene. The processor 102 may further include code for controlling the actions of a self-agent based on the predicted future actions of the agent.

[0048] Deep learning architectures can perform object recognition tasks by learning to represent the input at successively higher levels of abstraction in each layer, thereby constructing useful feature representations of the input data. In this way, deep learning addresses the main bottlenecks of conventional machine learning. Before the advent of deep learning, machine learning approaches to object recognition problems might have relied heavily on human-engineered features, perhaps combined with shallow classifiers. Shallow classifiers can be two-class linear classifiers, for example, where the weighted sum of the feature vector components is compared to a threshold to predict which class the input belongs to. Human-engineered features can be templates or kernels customized by engineers with domain expertise for a specific problem domain. In contrast, deep learning architectures can learn to represent features similar to those that human engineers might design, but it learns through training. Additionally, deep networks can learn to represent and identify new types of features that humans might not have considered yet.

[0049] Deep learning architectures can learn a hierarchy of features. For example, if visual data is presented to the first layer, the first layer can learn to identify relatively simple features (such as edges) in the input stream. In another example, if auditory data is presented to the first layer, the first layer can learn to identify spectral power in specific frequencies. A second layer that takes the output of the first layer as input can learn to identify combinations of features, such as identifying simple shapes for visual data or combinations of sounds for auditory data. For example, higher layers can learn to represent complex shapes in visual data or words in auditory data. Even higher layers can learn to identify common visual objects or spoken phrases.

[0050] Deep learning architectures may perform particularly well when applied to problems with a natural hierarchical structure. For example, the classification of motor vehicles can benefit from first learning to identify wheels, windshields, and other features. These features can be combined in different ways at higher levels to identify cars, trucks, and airplanes.

[0051] Neural networks can be designed with a variety of connectivity patterns. In a feedforward network, information is passed from lower layers to higher layers, where each neuron in a given layer communicates to neurons in higher layers. As described above, hierarchical representations can be constructed in successive layers of a feedforward network. Neural networks can also have recurrent or feedback (also known as top-down) connections. In a recurrent connection, the output from a neuron in a given layer can be communicated to another neuron in the same layer. Recurrent architectures can help identify patterns that span more than one chunk of input data presented sequentially to the neural network. Connections from neurons in a given layer to neurons in lower layers are called feedback (or top-down) connections. Networks with many feedback connections can be beneficial when the recognition of high-level concepts can assist in discerning specific low-level features of the input.

[0052] The connections between the layers of a neural network can be fully connected or locally connected. Figure 2A An example of a fully connected neural network 202 is illustrated. In the fully connected neural network 202, a neuron in the first layer can convey its output to every neuron in the second layer, such that every neuron in the second layer will receive inputs from every neuron in the first layer. Figure 2B An example of a locally connected neural network 204 is illustrated. In the locally connected neural network 204, a neuron in the first layer can be connected to a limited number of neurons in the second layer. More generally, the locally connected layers of the locally connected neural network 204 can be configured such that every neuron in a layer will have the same or similar connectivity pattern, although the connection strengths can have different values (e.g., 210, 212, 214, and 216). The locally connected connectivity pattern can give rise to spatially distinct receptive fields in higher layers, since the higher layer neurons in a given region can receive inputs that are tuned through training to the nature of a restricted portion of the total input to the network.

[0053] An example of a locally connected neural network is a convolutional neural network. Figure 2C An example of a convolutional neural network 206 is illustrated. The convolutional neural network 206 can be configured such that the connection strengths associated with the inputs for every neuron in the second layer are shared (e.g., 208). Convolutional neural networks can be well-suited for problems where the spatial location of the input is meaningful.

[0054] One type of convolutional neural network is a deep convolutional network (DCN). Figure 2D A detailed example of a DCN 200 designed to identify visual features of an image 226 input from an image capture device 230 (such as an in-vehicle camera) is illustrated. The DCN 200 of the current example can be trained to identify traffic signs and the digits provided on the traffic signs. Of course, the DCN 200 can be trained for other tasks, such as identifying lane markings or identifying traffic signals.

[0055] The DCN 200 can be trained using supervised learning. During training, images (such as image 226 of a speed limit sign) can be presented to the DCN 200, and then a "forward pass" can be computed to produce an output 222. The DCN 200 can include a feature extraction section and a classification section. Upon receiving the image 226, the convolutional layer 232 can apply a convolutional kernel (not shown) to the image 226 to generate a first set of feature maps 218. As an example, the convolutional kernel of the convolutional layer 232 can be a 5x5 kernel that generates 28x28 feature maps. In this example, since four different feature maps are generated in the first set of feature maps 218, four different convolutional kernels are applied to the image 226 at the convolutional layer 232. The convolutional kernel can also be referred to as a filter or a convolutional filter.

[0056] The first set of feature maps 218 can be subsampled by a max pooling layer (not shown) to generate a second set of feature maps 220. The max pooling layer reduces the size of the first set of feature maps 218. That is, the size of the second set of feature maps 220 (such as 14x14) is smaller than the size of the first set of feature maps 218 (such as 28x28). The reduced size provides similar information to subsequent layers while reducing memory consumption. The second set of feature maps 220 can be further convolved via one or more subsequent convolutional layers (not shown) to generate one or more subsequent sets of feature maps (not shown).

[0057] In Figure 2D the example, the second set of feature maps 220 is convolved to generate a first feature vector 224. Additionally, the first feature vector 224 is further convolved to generate a second feature vector 228. Each feature of the second feature vector 228 can include a number corresponding to a possible feature of the image 226 (such as "sign", "60", and "100"). A softmax function (not shown) can convert the numbers in the second feature vector 228 into probabilities. Thus, the output 222 of the DCN 200 is the probability that the image 226 includes one or more features.

[0058] In this example, the probabilities of "sign" and "60" in the output 222 are higher than the probabilities of other features (such as "30", "40", "50", "70", "80", "90", and "100") of the output 222. Before training, the output 222 produced by the DCN 200 is likely to be incorrect. Thus, the error between the output 222 and the target output can be computed. The target output is the ground truth of the image 226 (e.g., "sign" and "60"). The weights of the DCN 200 can then be adjusted so that the output 222 of the DCN 200 is more closely aligned with the target output.

[0059] To adjust the weights, the learning algorithm can compute a gradient vector for the weights. This gradient can indicate the amount by which the error will increase or decrease if the weights are adjusted. At the top layer, this gradient can directly correspond to the value of the weights connecting the activated neurons in the penultimate layer to the neurons in the output layer. In the lower layers, the gradient can depend on the values of the weights and the error gradients computed for the higher layers. The weights can then be adjusted to reduce the error. This way of adjusting the weights can be called "backpropagation" because it involves a "backward pass" through the neural network.

[0060] In practice, the error gradient of the weights may be computed on a small number of examples, so that the computed gradient approximates the true error gradient. This approximation method can be called stochastic gradient descent. Stochastic gradient descent can be repeated until the error rate achievable by the entire system has stopped decreasing or until the error rate has reached a target level. After learning, the DCN can be presented with a new image (e.g., the speed limit sign of image 226) and a forward pass through the network can produce an output 222, which can be considered an inference or prediction of the DCN.

[0061] A deep belief network (DBN) is a probabilistic model that includes multiple layers of hidden nodes. The DBN can be used to extract a hierarchical representation of a training data set. The DBN can be obtained by stacking several layers of restricted Boltzmann machines (RBMs). An RBM is an artificial neural network that can learn a probability distribution over an input set. Since an RBM can learn a probability distribution without information about which class each input should be classified into, RBMs are often used in unsupervised learning. Using a hybrid unsupervised and supervised paradigm, the bottom RBM of the DBN can be trained in an unsupervised manner and can be used as a feature extractor, while the top RBM can be trained in a supervised manner (on the joint distribution of the inputs from the previous layer and the target classes) and can be used as a classifier.

[0062] A deep convolutional network (DCN) is a network of convolutional networks that is configured with additional pooling and normalization layers. The DCN has achieved state-of-the-art performance on many tasks. The DCN can be trained using supervised learning, where both the inputs and the output targets are known for many paradigms and are used to modify the weights of the network using gradient descent.

[0063] The DCN can be a feedforward network. Additionally, as described above, the connections from the neurons in the first layer of the DCN to the groups of neurons in the next higher layer are shared across the neurons in the first layer. The feedforward and shared connections of the DCN can be exploited for fast processing. The computational burden of the DCN can be much smaller than that of, for example, a neural network of similar size that includes recurrent or feedback connections.

[0064] The processing of each layer of a convolutional network can be considered as a spatially invariant template or basis projection. If the input is first decomposed into multiple channels, such as the red, green, and blue channels of a color image, then a convolutional network trained on that input can be considered three-dimensional, having two spatial dimensions along the axes of the image and a third dimension capturing color information. The output of a convolutional connection can be considered to form a feature map in subsequent layers, where each element in the feature map (e.g., 220) receives input from a certain range of neurons in the previous layer (e.g., feature map 218) and from each of the multiple channels. The values in the feature map can be further processed with a non-linearity such as rectified max(0,x). Values from adjacent neurons can be further pooled (which corresponds to downsampling) and can provide additional local invariance and dimensionality reduction. Normalization can also be applied through lateral inhibition between neurons in the feature map, which corresponds to whitening.

[0065] The performance of deep learning architectures can improve as more labeled data points become available or as computing power increases. Modern deep neural networks are routinely trained with thousands of times more computing resources than were available to a typical researcher just fifteen years ago. New architectures and training paradigms can further boost the performance of deep learning. Rectified linear units can reduce a training problem known as vanishing gradients. New training techniques can reduce over-fitting and thus enable larger models to achieve better generalization. Encapsulation techniques can abstract the data within a given receptive field and further improve overall performance.

[0066] Figure 3 is a block diagram illustrating a deep convolutional network 350. The deep convolutional network 350 can include multiple different types of layers based on connectivity and weight sharing. As Figure 3 shown, the deep convolutional network 350 includes convolutional blocks 354A, 354B. Each of the convolutional blocks 354A, 354B can be configured with a convolutional layer (CONV) 356, a normalization layer (LNorm) 358, and a max pooling layer (MAX POOL) 360.

[0067] The convolutional layer 356 can include one or more convolutional filters, which can be applied to input data to generate a feature map. Although only two convolutional blocks 354A, 354B are shown, the present disclosure is not limited thereto, but instead any number of convolutional blocks 354A, 354B can be included in the deep convolutional network 350 according to design preferences. The normalization layer 358 can normalize the output of the convolutional filters. For example, the normalization layer 358 can provide whitening or lateral inhibition. The max pooling layer 360 can provide spatially downsampled aggregation to achieve local invariance and dimensionality reduction.

[0068] For example, the parallel filter banks of the deep convolutional network can be loaded onto the CPU 102 or GPU 104 of the SOC 100 to achieve high performance and low power consumption. In an alternative embodiment, the parallel filter banks can be loaded onto the DSP 106 or ISP 116 of the SOC 100. Additionally, the deep convolutional network 350 can access other processing blocks that may be present on the SOC 100, such as the sensor processor 114 and the navigation module 120 dedicated to sensors and navigation, respectively.

[0069] The deep convolutional network 350 may also include one or more fully connected layers 362 (FC1 and FC2). The deep convolutional network 350 may further include a logistic regression (LR) layer 364. Weights (not shown) to be updated exist between each layer 356, 358, 360, 362, 364 of the deep convolutional network 350. The output of each layer (e.g., 356, 358, 360, 362, 364) can be used as the input to a subsequent layer (e.g., 356, 358, 360, 362, 364) in the deep convolutional network 350 to learn a hierarchical feature representation from the input data 352 (e.g., image, audio, video, sensor data, and / or other input data) supplied at the first convolutional block 354A. The output of the deep convolutional network 350 is a classification score 366 for the input data 352. The classification score 366 can be a set of probabilities, where each probability is the probability that the input data includes a feature from the feature set.

[0070] In a behavior prediction scenario, the environment may exhibit structures (e.g., rules) that can be used to reduce the number of computations. For example, in an environment with many agents, the behavior of each agent may be mainly influenced by a subset of the surrounding agents. In driving, for example, neighboring vehicles may have a greater impact on the behavior of a given vehicle compared to the influence from non-adjacent vehicles.

[0071] The constraints of the environment may also imply an intuitive order of action priorities among the agents. For example, while driving, environmental constraints (such as traffic regulations, local driving habits, and road structure) may affect the order of action priorities. As an example, local driving habits and / or traffic regulations may stipulate that a vehicle in front of other vehicles in the same lane has the right of way. In such a case, the behavior prediction model can be simplified by assuming that the follower vehicle reacts to the actions of the leading vehicle, but the actions of the leading vehicle are largely independent of the actions of the follower vehicle.

[0072] Environmental constraints can be used to construct interaction graphs. The nodes of the graph represent agents. The directed edges of the graph encode one-way dependencies between two nodes. In some instances, it may be appropriate for there to be cyclic relationships between agents. A cyclic relationship is a situation in which the behavior of a given agent is considered to be mutually dependent. Mutual dependence does not privilege one relationship over another.

[0073] Recursive predictions can be generated by traversing the interaction graph in priority order. The predictions are recursive because higher-level predictions for a target vehicle are conditioned on predictions for surrounding vehicles, which may be conditioned on lower-level predictions for the same target vehicle. In one configuration, level 0 predictions are generated for agents that have a higher priority than their direct neighbors in the interaction graph. Level 0 predictions can also be generated for agents that are part of a cyclic dependence relationship (e.g., the action of one agent depends on the action of another agent).

[0074] When there are cyclic relationships between agents, recursive reasoning is used to generate predictions for mutually dependent agents. That is, cyclic relationships are properties of the interaction graph (which is generated based on the structure of the environment as exhibited or inferred). Recursion can be used to manage cyclic dependencies. In the general case, when considering the structure of the environment, it is assumed that the behavior of each agent is conditioned on the behavior of every other agent. In this case, all relationships are cyclic.

[0075] Figure 4A 、 4B and 4C illustrate examples of recursive multi-fidelity behavior prediction in accordance with aspects of the present disclosure. As Figure 4A shown, at a first time step, a prediction model receives information corresponding to an observation of a scene 400 (such as vehicles on a multi-lane road). The scene 400 can be observed via one or more sensors of a self-agent 402 (such as a LIDAR sensor, a RADAR sensor, a camera, and / or another type of sensor). Based on this observation, the prediction model (e.g., an atomic prediction model) identifies the positions of other objects (such as other agents 404) in the scene. This observation can also identify the direction of travel of each agent (identified by the arrows in Figure 4A ). The environmental model is updated based on the observed scene 400.

[0076] As Figure 4BAs shown, after observing scenario 400, the prediction model assigns an inference level and a set of atomic prediction models to each of the observed agents 404. In one configuration, agents 406 adjacent to the ego agent 402 are assigned a higher fidelity level compared to other agents 404. The adjacent agents 406 are a subset of the observed agents 404. The adjacent agents 406 are in a first mode surrounding the ego agent 402. The ego agent 402 is in a second mode. Additionally, agents 408 in the same lane 410 as the ego agent 402 and in front of other agents 404, 406 may be assigned a higher fidelity level because the actions of other agents 404, 406 may depend on the actions of this front agent 408. The front agent 408 is a subset of the observed agents 404.

[0077] As Figure 4C shown, after assigning the fidelity levels, the prediction model generates an interaction graph based on the structure of the environment to encode the relationships between the agents. The interaction graph may also be generated before assigning the fidelity levels. Directed edges ( Figure 4C (not shown in ) encode the direction of the influence between two agents. The connections 412 between agents 404, 406, 408 identify the constraints between two agents. In a directed graph, each edge includes a direction depicted by an arrow tip ( Figure 4C (not shown in ). In the case of a two-way relationship, the edge should have arrow tips at both ends ( Figure 4C (not shown in ). That is, if the action of a first agent 426 has an influence on a second agent 424, the edge should point from the first agent 426 to the second agent 424. Successive subsets of level 0, 1,..., K behavior predictions are selected by traversing the graph from the highest priority to the lowest priority.

[0078] According to aspects of the present disclosure, after generating the interaction graph, the model predicts a level 0 scenario for the set of agents. Each scenario corresponds to a motion hypothesis for each agent in the set of agents. As discussed above, for each agent with an assigned level 0 atomic prediction model, the recursive inference scheme uses the assigned level 0 atomic prediction model to generate a level 0 motion hypothesis. The number of distinct predictions for each agent may be based on the prediction fidelity of the atomic prediction model.

[0079] Figure 5A illustrates an example of two distinct level 0 scenarios according to aspects of the present disclosure. As Figure 5AAs shown, scenarios 500A and 500B are generated for a first set of agents 510 (e.g., vehicles on a road). The first set of agents 510 includes four agents 502, 504, 506, and 508. The fidelity level of the second agent 504 in the set of agents 510 is higher than the fidelity levels of the other agents 502, 506, and 508 in the set of agents 510. Arrows 520 identify the predicted movement of each agent 502, 504, 506, and 508.

[0080] Based on the level 0 scenarios 500A and 500B, the prediction model generates level 1 scenarios for a second set of agents 514. Figure 5B An example of generating level 1 scenarios 550A, 550B, and 550C for a second set of agents 514 conditioned on the level 0 scenarios 500A and 500B is illustrated. The predicted action of each agent in the second set 514 is a response to the predicted actions of the first set of agents 510 in the first level 0 scenarios 500A and 500B. Each level 1 scenario 550A, 550B, and 550C includes a predicted action of each agent within the second set 514. The predicted actions in each level 1 scenario 550A, 550B, and 550C can be generated by a level 1 prediction model (see FIG. 7).

[0081] The prediction model continues to generate scenarios for sets of agents. That is, the prediction model generates level 0 to level K scenarios. Each level k (k = 0 to K) scenario is conditioned on the level k - 1 scenario and the scenarios of prior levels. Specifically, each level k scenario is conditioned on the full node chain from the root node to the parent node in the scenario tree. Additionally, any particular leaf node in the scenario tree includes the highest level prediction for each agent in the ancestor chain all the way back to the root node.

[0082] Figure 5C An example of generating level k scenarios 570A, 570B, and 570C for a fourth set of agents 516 conditioned on the level k - 1 scenarios 560A and 560B is illustrated. The predicted action of each agent in the fourth set 516 is a predicted response to the predicted actions of the third set of agents 518 in the level k - 1 scenarios 560A and 560B. Each level k scenario 570A, 570B, and 570C includes a predicted action of each agent within the fourth set 516. The predicted actions in each level k scenario 570A, 570B, and 550C can be generated by a level k prediction model. One or more scenario trees can be generated based on the level 0 - k scenarios. Each scenario tree encodes a particle - based representation of the joint distribution over the possible future behaviors of the agents in the environment.

[0083] Figure 6 An example of a model 600 for predicting trajectories in accordance with aspects of the present disclosure is illustrated. As Figure 6 shown, the prediction model receives information corresponding to an observation of a scene 604 at a current time step and assigns a fidelity level to each agent 640, 642. Some agents may be high-fidelity agents 642 while other agents are low-fidelity agents 640. Aspects of the present disclosure are not limited to two fidelity levels and two or more fidelity levels may be used. Multiple level 0 scenarios may be generated based on multiple level 0 trajectories of one or more high-fidelity agents 642. A level 0 scenario refers to the potential level 0 trajectories of each observed agent 640, 642 at the next time step. In this example, the prediction model selects one of the high-fidelity agents 642 and determines a level 0 trajectory 610 of the selected agent 606 based on previous observations. A human driver may manually operate the high-fidelity agents 642 and the low-fidelity agents 640.

[0084] To determine the level 0 trajectory 610 of the selected agent 606, the prediction model determines a region of interest 602 of the selected agent 606. The region of interest 602 may be application-dependent. For example, the application may determine the distance and / or position of other agents that can be used for planning. In one example, the range of the region of interest 602 of an emergency vehicle may be greater than that of a personal use vehicle. The range limitation of a sensor may also determine the range of the region of interest 602. The agents in the region of interest 602 include the high-fidelity agents 642, the low-fidelity agents 640, and the ego agent 630. For clarity, the agents 630, 640, 642 adjacent to the selected agent 606 in the region of interest 602 may be referred to as adjacent agents 616. The previous trajectories 608 (e.g., actions) of each agent 606, 616 in the region of interest 602 are known from previous observations. Based on the previous trajectories 608, the model 600 determines the level 0 trajectory 610 of the selected agent 606.

[0085] In one configuration, the previous trajectories 608 of each agent 606, 616 are encoded by a long short-term memory (LSTM) encoder 612. The LSTM encoder 612 may be an LSTM neural network. The output of the LSTM encoder 612 is a recent history tensor 628 that summarizes the most recent history of the behavior of each adjacent agent 616. The LSTM encoder 612 also outputs a vehicle dynamics tensor 618 that encodes the dynamics of the selected agent 606. The recent history tensor 628 may be stored in a three-dimensional (3D) tensor structure 614 that mimics the geometric relationship of the adjacent agents 616 relative to the selected agent 606.

[0086] The most recent history tensor 628 in the 3D tensor structure 614 is processed by multiple layers 624 of a convolutional neural network (CNN) 620. The output of the CNN 620 is a social context tensor representing statistics that describe the state of the local environment. Specifically, the most recent history tensor 628 summarizes the most recent history of the agent 608 as a vector encoded by the LSTM encoder 612. The most recent history tensor 628 is placed in the 3D tensor structure 614 according to its position in the scene, thereby geometrically capturing the interactions between vehicles. The prediction tensor 632 is different from the most recent history tensor 628. The prediction tensor 632 is a vector that encodes the level-k (k = 0 to K) predictions as a vector.

[0087] For the level-0 prediction, the social context output is combined with the vehicle dynamics tensor 618 of the selected agent 606. For the level-1 prediction, three vectors are concatenated: the social context vectors generated from the CNN 620 for levels 0 and 1; and the vehicle dynamics tensor 618 (e.g., a vector). The combination of the social context output and the vehicle dynamics tensor 618 of the selected agent 606 is input to the decoder neural network 622.

[0088] The decoder neural network 622 generates a predictive distribution of future motion over a set of future frames. The inherent multimodality of driver behavior is addressed by predicting the distribution over various maneuver classes and the probability of each maneuver class. In one configuration, the maneuver classes include lateral maneuver classes and longitudinal maneuver classes.

[0089] As Figure 6 shown, the decoder neural network 622 receives a trajectory encoding. The decoder neural network 622 includes two softmax layers (a lateral softmax layer 650 and a longitudinal softmax layer 652). The lateral softmax layer 650 outputs the lateral maneuver probability (P(m i |X)), while the longitudinal softmax layer 652 outputs the longitudinal maneuver probability. The longitudinal maneuver probability and the lateral maneuver probability can be multiplied to determine the maneuver distribution (P(m i |X)). P() is a probability distribution conditioned on the history of the trajectory X and the maneuver m i .

[0090] The LSTM decoder generates t fThe parameters of the bivariate Gaussian distribution on a frame are used to provide a predictive distribution of the vehicle motion. The LSTM decoder generates a distribution that varies depending on the maneuver. That is, the LSTM decoder generates a distribution on the level 0 trajectory 610. This distribution provides the probability for each level 0 trajectory 610. The decoder neural network 622 also generates the shape (e.g., path) of each level 0 trajectory 610. The process for determining the level 0 trajectory 610 is repeated for each high-fidelity agent 642 in the scenario 604. The level 0 trajectory 610 does not provide information about the future interaction between the agents 606, 616 in the region of interest 602. That is, the level 0 trajectory 610 does not provide information about the level 1 trajectory of the neighboring agent 616.

[0091] As Figure 6 shown, the trajectory encoding is concatenated with the maneuver encoding 654 via the concatenator 656. Specifically, the trajectory encoding is concatenated with a vector corresponding to the lateral maneuver category and a vector corresponding to the longitudinal maneuver. The concatenated encoding is input to the LSTM decoder to obtain the maneuver-varying distribution P Θ (Y|m i ,X), where P() is the probability distribution on the predicted trajectory Y (the sequence coordinates of the future positions) conditioned on the history of the trajectory X and the maneuver m i . The LSTM decoder outputs the mean and covariance of the Gaussian distribution (Θ) on t f frames, where t f is the number of future frames.

[0092] The maneuver encoding is obtained from the maneuver category. As discussed, the maneuver category is based on the lateral maneuver and the longitudinal maneuver. The lateral maneuvers include left lane change, right lane change, and lane keeping maneuvers. The left lane change and the right lane change can vary with respect to the actual intersection. Thus, two or more vectors can be defined for each of the left lane change and the right lane change. The longitudinal maneuver can be split into normal driving and braking.

[0093] After determining the level 0 trajectory 610 of each high-fidelity agent 642, the model 600 can be used to determine the level 1 trajectory. Specifically, the level 0 trajectory 610 is determined for each high-fidelity agent 642 in the region of interest 602 of the ego agent 630. Each high-fidelity agent 642 can be used as the selected agent 606. After determining all the level 0 trajectories, the model determines the level 1 trajectories of the agents 641, 642 around each selected agent 606. In the first iteration, the level 1 trajectories are calculated for the agents 641, 642 around each selected agent 606. In the k-th iteration, the level k trajectories are calculated for the agents 641, 642 around each selected agent 606.

[0094] For level 0, the previous trajectories 608 of each neighboring agent 616 are encoded by the LSTM encoder 612. In one configuration, for level 1 trajectories, only the level 0 predicted trajectories are encoded. In another configuration, for level 1 trajectories, instead of using the previous trajectories 608, a constant velocity model is used for the low-fidelity agents 640.

[0095] Similar to determining the level 0 trajectories, for level 1 trajectories, the prediction tensors 632 of the neighboring agents 616 are stored in the 3D tensor structure 614. The prediction tensors 632 in the 3D tensor structure 614 are processed by multiple layers 624 of the CNN 620. The weights of the multiple layers 624 are different between level 1 and level 0. For level 1 trajectories, the social context output of the CNN 620 is combined with the level 0 social context output and the vehicle dynamics tensor 618 of the selected agent 606. This combination is input into the decoder neural network 622 and a distribution over the level 1 trajectories is generated. This distribution provides the probability of each level 1 trajectory. The decoder neural network 622 also generates the shape (e.g., path) of each level 1 trajectory.

[0096] The process of the repeatable model 600 is carried out until level K. The level K predictions are combined with the social context output and the vehicle dynamics tensor 618 of the selected agent 606. The model 600 is not limited to Figure 6 the model 600. Other models can be used for behavior prediction. Other models will generate level 0 predictions and use recursion.

[0097] According to another aspect of the present disclosure, recursive multi-fidelity prediction takes into account gap-thread manipulation. Figure 7A 、 7B 、7C and 7D illustrate examples of recursive multi-fidelity prediction using gap-thread manipulation. According to aspects of the present disclosure. As Figure 7A shown, at the first time step, the prediction model receives information corresponding to an observation of the scene 700. The scene 700 can be observed via one or more sensors of the ego agent 702 (such as LIDAR sensors, RADAR sensors, cameras, and / or another type of sensor). Based on this observation, the prediction model identifies the positions of other objects (such as other agents 704) in the scene. This observation can also identify the direction of travel of each agent 704 (identified by the arrows in Figure 7A ). The environmental model is updated based on the observed scene 700.

[0098] As Figure 7B shown, after observing the scene 700, the prediction model generates an interaction graph based on geometric and map-based pairwise features. Directed edges ( Figure 7B(not shown in the figure) encodes the direction of influence between two agents. The connection 712 between agents 702 and 704 identifies the constraints between the two agents. For example, connections 712 are established between the ego agent 702 and each adjacent agent 706. The connection 712 identifies the relationship between the ego agent 702 and each adjacent agent 706, such that the actions of the ego agent 702 can affect the actions of each adjacent agent 706. Additionally, the actions of the adjacent agent 706 can affect the actions of the ego agent 702.

[0099] After generating the interaction graph, the scenario 700 is partitioned into different fidelity neighborhoods. Figure 7C An example of a fidelity neighborhood is illustrated. The high-fidelity neighborhood can be centered around the ego agent 702. For example, each adjacent agent 706 can be within the high-fidelity neighborhood 710. The non-adjacent agent 708 can be assigned to a low-fidelity neighborhood 722. For clarity, Figure 7C each low-fidelity neighborhood 722 is not illustrated.

[0100] Identify the applicable strategies for each agent 706, 708. Figure 7D An example of the identified strategies 714, 716, 718, 720 for agent 708 is illustrated. Each strategy can be defined by the local neighborhood and the road geometry. For example, as Figure 7D shown, agent 708 can maintain its current trajectory 714, move to the gap 716 in front of the first agent 707A, move to the gap 718 between the first agent 707A and the second agent 707B, or move to the gap 720 behind the second agent 707B. The gaps 716, 718, 720 and the current trajectory 714 can be referred to as strategies.

[0101] In one configuration, a policy likelihood is determined for each policy corresponding to each agent. The policy likelihood determines the likelihood that an agent will execute a policy. The likelihood of executing a policy can be based on the cost of the trajectory, the similarity of the trajectory to the cached trajectory from a previous time step, the agent's previous actions, the map location, the neighbor constellation, etc. In one configuration, the agent's previous actions are used to determine the likelihood of executing a policy. For example, the movement of the agent in a direction or the agent's turn signal can be used to determine the likelihood of executing a policy.

[0102] Figure 8 An example of determining the most likely policy according to aspects of the present disclosure is illustrated. As Figure 8As shown, different policies 802 for the target agent 800 are determined. In this example, the agent 800 may have activated its right turn signal at a previous time step. Based on the activated turn signal, the prediction model may determine that moving into the gap between the first agent 804 and the second agent 806 is the most likely policy.

[0103] Based on the priority order determined from the interaction graph, one or more policies may be sampled for each agent. The number of samples may depend on the fidelity level and the maneuver distribution of the agent. When more than one policy is sampled for an agent, the scenario tree is bifurcated. The bifurcation may change the interaction graph. The change in the interaction graph results in a change in the priority order. Each path from the root node to a leaf node of the scenario tree represents the complete set of sampled policies, where one policy is sampled for each vehicle. Thus, the prediction in the leaf node is conditioned on the complete node chain from the root node to the parent of that leaf.

[0104] Figure 9A An example of determining policies 902 for agents 920, 922, 924, 926, 928 in accordance with aspects of the present disclosure is illustrated. As Figure 9A shown, at the root node 910 of the scenario tree, policies 902 for each of the agents 920, 922, 924, 926, 928 are determined in priority order. For example, the policies 902 may be determined in the order of the numbers (e.g., 1 - 5) corresponding to each of the agents 920, 922, 924, 926, 928. In this example, two policies are generated for the fifth agent 928. The number of policies for the fifth agent 928 may be greater than the number of policies for the other agents 920, 922, 924, 926 because a high fidelity level is assigned to the fifth agent 926.

[0105] In response to the fifth agent 928 having more than one policy 902, the scenario tree is bifurcated. Figure 9B An example of a branch node of a scenario tree in accordance with aspects of the present disclosure is illustrated. As Figure 9B shown, the root node 99 of the scenario tree includes policies 902 for a set of agents 920, 922, 924, 926. Additionally, the first leaf 912 includes policies 902 for a set of agents 920, 922, 924, 926 and a first policy 930 for the fifth agent 928. The second leaf 914 includes policies 902 for a set of agents 920, 922, 924, 926 and a second policy 932 for the fifth agent 928. Policies for other agents 934 may be generated in the first leaf 912 and the second leaf 914. When more than one policy is generated for one of the other agents 934, the first leaf 912 and the second leaf 914 may be bifurcated.

[0106] In accordance with aspects of the present disclosure, a hybrid approach can be used for behavior prediction. The hybrid approach can use a model similar to the model of Figure 6 For the hybrid approach, a level 0 policy is determined for a set of agents in the neighborhood. The actions of this set of agents may not have a substantial impact on the ego agent. Based on the level 0 policy, a level 1 policy for each agent is determined based on priority.

[0107] Figure 10 Method 1000 for predicting the future actions of agents in a prediction scenario in accordance with one aspect of the present disclosure is illustrated. As Figure 10 shown, in a first block 1002, a prediction model assigns a fidelity level to agents observed in the scenario. The fidelity level can refer to the significance of an agent in the scenario. In one configuration, computing resources are biased towards the agents considered to be the most significant. The scenario can be observed via one or more sensors such as LIDAR sensors, RADAR sensors, cameras, and / or another type of sensor. Based on this observation, the prediction model identifies the positions of other objects in the scenario. This observation can also identify the travel direction of each agent.

[0108] In an optional configuration, at block 1004, an inference level and a set of forward prediction models are assigned to each agent in the scenario. The forward prediction models can be referred to as atomic models. The inference level of an agent can be an integer greater than or equal to zero. Each forward prediction model in the assigned set of forward prediction models for an agent corresponds to a specific inference level greater than or equal to zero and less than or equal to the assigned inference level of that agent. For each level, a given agent can have no more than one assigned atomic prediction model.

[0109] For example, if an agent is assigned an inference level of 1, then the agent can include a level 0 forward prediction model and / or a level 1 forward prediction model. The level 0 forward prediction model can generate a level 0 motion hypothesis for the agent (e.g., recursion level 0). The level 1 forward prediction model can generate a level 1 motion hypothesis for the agent (e.g., recursion level 1). That is, each forward prediction model in the set of forward prediction models corresponds to a recursion level determined based on the inference level.

[0110] In an optional configuration, at block 1006, the prediction model divides the scenario into different neighborhoods. Each neighborhood can be assigned a different fidelity. The fidelity can be based on proximity to the ego agent. For example, a high-fidelity neighborhood can be centered on the ego agent. The fidelity of an agent can be based on the fidelity of the corresponding neighborhood.

[0111] At block 1008, the prediction model recursively predicts the future actions of the agent by traversing the scenario. For example, for each agent with an assigned level 0 forward prediction model, the recursive inference scheme uses the assigned level 0 forward prediction model to generate level 0 motion hypotheses. Subsequently, for each agent with an assigned level 1 forward prediction model, the prediction model uses the assigned level 1 forward prediction model to generate level 1 motion hypotheses.

[0112] A subset of the level 0 motion hypotheses of other agents can be used as input to any of the level 1 prediction models. This process can be repeated such that each successive set of motion hypotheses (level k) can be conditioned on the highest level (up to k - 1) previously computed motion hypotheses of each agent in a subset of the other agents in the scenario. As discussed, different forward prediction models (e.g., level 0, level 1, etc.) can be used at each recursive level.

[0113] In one configuration, future actions are recursively predicted based on an initial trajectory that includes historical observations of each agent. That is, the input to the prediction model can be a representation of the history of the scenario. In another configuration, future actions are recursively predicted based on applicable policies for each agent. The policies can be based on the corresponding neighborhood of the agent and the scenario structure. The scenario structure can refer to road geometry.

[0114] Finally, at block 1010, the prediction model controls the actions of the ego agent based on the predicted future actions of the agent. For example, the prediction model can change the route, adjust the speed, or control another action. The prediction model can be a component of the ego agent.

[0115] In some aspects, method 1000 can be performed by SOC 100( Figure 1 )). That is, by way of example and not limitation, each element of method 1000 can be performed by SOC 100 or one or more processors (e.g., CPU 102) and / or other included components.

[0116] The various operations of the methods described above can be performed by any suitable device capable of performing the corresponding functions. These devices can include various hardware and / or software components and / or modules, including but not limited to circuitry, application specific integrated circuits (ASICs), or processors. Generally, where operations are illustrated in the figures, those operations can have corresponding paired device plus functional components with similar numbers.

[0117] As used herein, the term "determine" encompasses a variety of actions. For example, "determine" can include calculating, computing, processing, deriving, researching, looking up (e.g., looking up in a table, database, or other data structure), ascertaining, and the like. Additionally, "determine" can include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory), and similar actions. Further, "determine" can include parsing, selecting, choosing, establishing, and similar actions.

[0118] As used herein, a phrase that recites "at least one" of a list of items refers to any combination of those items, including a single member. As an example, "at least one of a, b, or c" is intended to cover: a, b, c, a - b, a - c, b - c, and a - b - c.

[0119] The various illustrative logical blocks, modules, and circuits described in connection with the present disclosure can be implemented or executed with a general - purpose processor, a digital signal processor (DSP), an application - specific integrated circuit (ASIC), a field - programmable gate array (FPGA) or other programmable logic device (PLD), discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general - purpose processor may be a microprocessor, but in the alternative, the processor may be any commercially available processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.

[0120] The steps of a method or algorithm described in connection with the present disclosure can be implemented directly in hardware, in a software module executed by a processor, or in a combination of the two. The software modules can reside in any form of storage medium known in the art. Some examples of storage media that can be used include random access memory (RAM), read - only memory (ROM), flash memory, erasable programmable read - only memory (EPROM), electrically erasable programmable read - only memory (EEPROM), registers, hard disk, removable disk, CD - ROM, and the like. The software modules can include a single instruction, or many instructions, and can be distributed over several different code segments, distributed among different programs, and across multiple storage media. The storage medium can be coupled to the processor such that the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium can be integrated into the processor.

[0121] The methods disclosed herein include one or more steps or acts for implementing the described methods. These method steps and / or acts may be interchanged with one another without departing from the scope of the claims. In other words, unless a specific order of steps or acts is specified, the order and / or use of the specific steps and / or acts may be altered without departing from the scope of the claims.

[0122] The described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in hardware, an example hardware configuration may include a processing system in a device. The processing system may be implemented with a bus architecture. Depending on the particular application and overall design constraints of the processing system, the bus may include any number of interconnected buses and bridges. The bus may link together various circuits including a processor, a machine-readable medium, and a bus interface. The bus interface may be used to connect, among other things, a network adapter to the processing system via the bus. The network adapter may be used to implement signal processing functions. For some aspects, a user interface (such as a keypad, display, mouse, joystick, etc.) may also be connected to the bus. The bus may also link various other circuits (such as a timing source, peripherals, voltage regulators, power management circuits, etc.), which are well known in the art and will not be described further herein.

[0123] The processor may be responsible for managing the bus and general processing, including the execution of software stored on the machine-readable medium. The processor may be implemented with one or more general and / or special-purpose processors. Examples include a microprocessor, a microcontroller, a DSP processor, and other circuitry capable of executing software. Software should be construed broadly to mean instructions, data, or any combination thereof, whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise. As an example, the machine-readable medium may include random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, magnetic disks, optical disks, hard drives, or any other suitable storage medium, or any combination thereof. The machine-readable medium may be embodied in a computer program product. The computer program product may include packaging materials.

[0124] In a hardware implementation, the machine-readable medium can be a part separate from the processor in a processing system. However, as will be readily appreciated by those skilled in the art, the machine-readable medium or any part thereof can be external to the processing system. As an example, the machine-readable medium can include transmission lines, carriers modulated with data, and / or computer products separate from the device, all of which can be accessed by the processor via a bus interface. Alternatively or additionally, the machine-readable medium or any part thereof can be integrated into the processor, such as may be the case with a cache and / or a general register file. Although the various components discussed may be described as having a specific location, such as local components, they can also be configured in various ways, such as some components being configured as part of a distributed computing system.

[0125] The processing system can be configured as a general-purpose processing system having one or more microprocessors providing processor functionality, and an external memory providing at least a portion of the machine-readable medium, all linked together via an external bus architecture with other support circuitry. Alternatively, the processing system can include one or more neuromorphic processors for implementing the neuron models and nervous system models described herein. As another alternative, the processing system can be implemented with an application-specific integrated circuit (ASIC) having a processor, a bus interface, a user interface, support circuitry, and at least a portion of the machine-readable medium integrated on a single chip, or with one or more field-programmable gate arrays (FPGAs), programmable logic devices (PLDs), controllers, state machines, gated logic, discrete hardware components, or any other suitable circuitry, or any combination of circuits capable of performing the various functions described throughout this disclosure. Depending on the particular application and overall design constraints imposed on the overall system, those skilled in the art will recognize how best to implement the functionality described with respect to the processing system.

[0126] The machine-readable medium can include several software modules. These software modules include instructions that, when executed by the processor, cause the processing system to perform various functions. These software modules can include a transmission module and a receiving module. Each software module can reside in a single storage device or be distributed across multiple storage devices. As an example, when a triggering event occurs, the software module can be loaded from a hard drive into the RAM. During the execution of the software module, the processor can load some instructions into the cache to improve access speed. One or more cache lines can then be loaded into the general register file for execution by the processor. When referring to the functionality of the software modules hereinafter, it will be understood that such functionality is implemented by the processor when the processor executes instructions from the software module. Additionally, it should be appreciated that aspects of the present disclosure result in improvements to the capabilities of the processor, computer, machine, or other system implementing such aspects.

[0127] If implemented in software, each function may be stored on or transmitted via a computer-readable medium as one or more instructions or code. Computer-readable media includes both computer storage media and communication media including any medium that facilitates transfer of a computer program from one place to another. A storage media may be any available media that can be accessed by a computer. By way of example, and not limitation, such computer-readable media can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer. Additionally, any connection is properly termed a computer-readable medium. For example, if the software is transmitted from a web site, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology such as infrared, radio, and microwave is included in the definition of medium. As used herein, the terms "disk" and "disc" include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and discs, where disks typically reproduce data magnetically, while discs reproduce data optically with lasers. Thus, in some aspects, computer-readable media may include non-transitory computer-readable media (e.g., tangible media). Additionally, for other aspects, computer-readable media may include transitory computer-readable media (e.g., signals). Combinations of the above should also be included within the scope of computer-readable media.

[0128] Accordingly, some aspects may include a computer program product for performing the operations presented herein. For example, such a computer program product may include a computer-readable medium having instructions (and / or encoded thereon) that can be executed by one or more processors to perform the operations described herein. For some aspects, the computer program product may include packaging materials.

[0129] In addition, it should be appreciated that modules and / or other suitable means for performing the methods and techniques described herein can be downloaded and / or otherwise obtained by a user terminal and / or a base station where applicable. For example, such devices can be coupled to a server to facilitate transfer of means for performing the methods described herein. Alternatively, the various methods described herein can be provided via a storage device (e.g., RAM, ROM, a physical storage medium such as a compact disc (CD) or a floppy disk, etc.) such that once the storage device is coupled to or provided to the user terminal and / or the base station, the device can obtain the various methods. In addition, any other suitable technique for providing the methods and techniques described herein to a device can be utilized.

[0130] It will be understood that the claims are not limited to the exact configurations and components described above. Various changes, substitutions, and modifications can be made in the layout, operation, and details of the methods and apparatuses described above without departing from the scope of the claims.

Claims

1. A method, comprising: Assign a fidelity level among a plurality of fidelity levels to each of a plurality of agents observed in a scene; Recursively predict future actions of the plurality of agents by traversing the scene, using different forward prediction models at each recursive level to predict the future actions of each agent associated with a fidelity level among the plurality of fidelity levels; and Control the actions of a self-agent based on the predicted future actions of the plurality of agents.

2. The method according to claim 1, further comprising: Assign an inference level and a set of forward prediction models to each agent in the scene.

3. The method according to claim 2, wherein, Each forward prediction model in the set of forward prediction models corresponds to a recursive level determined based on the inference level.

4. The method according to claim 1, further comprising recursively predicting future actions based on an initial trajectory including historical observations of each agent.

5. The method according to claim 1, further comprising partitioning the scene into different neighborhoods.

6. The method according to claim 5, wherein, Each fidelity level is assigned based on the fidelity of the neighborhood.

7. The method according to claim 1, wherein, The future actions are recursively predicted based on a policy associated with each agent, and each policy is based on both the neighborhood of the associated agent and the scene structure.

8. An apparatus, comprising: A memory; And At least one processor coupled to the memory, the at least one processor being configured to: Assign a fidelity level among a plurality of fidelity levels to each of a plurality of agents observed in a scene; Recursively predict future actions of the plurality of agents by traversing the scene, using different forward prediction models at each recursive level to predict the future actions of each agent associated with a fidelity level among the plurality of fidelity levels; and Control the actions of a self-agent based on the predicted future actions of the plurality of agents.

9. The apparatus according to claim 8, wherein, The at least one processor is further configured to assign an inference level and a set of forward prediction models to each agent in the scene.

10. The apparatus according to claim 9, wherein, Each forward prediction model in the set of forward prediction models corresponds to a recursive level determined based on the inference level.

11. The apparatus according to claim 8, wherein, The at least one processor is further configured to recursively predict future actions based on an initial trajectory including historical observations of each agent.

12. The apparatus according to claim 8, wherein, The at least one processor is further configured to partition the scene into different neighborhoods.

13. The apparatus according to claim 12, wherein, Each fidelity level is assigned based on the fidelity of the neighborhood.

14. The apparatus according to claim 8, wherein, The future actions are recursively predicted based on a policy associated with each agent, and each policy is based on both the neighborhood of the associated agent and the scene structure.

15. A non-transitory computer-readable medium having program code recorded thereon, the program code being executed by a processor and comprising: Program code for assigning a fidelity level among a plurality of fidelity levels to each of a plurality of agents observed in a scene; Program code for recursively predicting future actions of the plurality of agents by traversing the scene, using different forward prediction models at each recursive level to predict the future actions of each agent associated with a fidelity level among the plurality of fidelity levels; And Program code for controlling the actions of a self-agent based on the predicted future actions of the plurality of agents.

16. The non-transitory computer-readable medium according to claim 15, wherein, The program code further includes program code for assigning an inference level and a set of forward prediction models to each agent in the scene.

17. The non-transitory computer-readable medium according to claim 16, wherein, Each forward prediction model in the set of forward prediction models corresponds to a recursive level determined based on the inference level.

18. The non-transitory computer-readable medium according to claim 15, wherein, The program code further includes program code for recursively predicting future actions based on an initial trajectory including historical observations of each agent.

19. The non-transitory computer-readable medium according to claim 15, wherein, The program code further includes program code for partitioning the scene into different neighborhoods.

20. The non-transitory computer-readable medium according to claim 19, wherein, Each fidelity level is assigned based on the fidelity of the neighborhood.

21. The non-transitory computer-readable medium according to claim 15, wherein, The future actions are recursively predicted based on policies associated with each agent, and each policy is based on both the corresponding neighborhood and the scene structure of the associated agent.

22. An apparatus, comprising: Means for assigning a fidelity level from a plurality of fidelity levels to each of a plurality of agents observed in a scene; Means for recursively predicting future actions of the plurality of agents by traversing the scene, using different forward prediction models at each recursive level to predict the future actions of each agent associated with a fidelity level from the plurality of fidelity levels; And Means for controlling the actions of a self-agent based on the predicted future actions of the plurality of agents.

23. The apparatus according to claim 22, further comprising means for assigning an inference level and a set of forward prediction models to each agent in the scene.

24. The apparatus according to claim 23, wherein, Each forward prediction model in the forward prediction model set corresponds to a recursive level determined based on the inference level.

25. The apparatus according to claim 22, wherein, The means for recursively predicting future actions includes means for recursively predicting future actions based on an initial trajectory including historical observations of each agent.

26. The apparatus according to claim 22, further comprising means for partitioning the scene into different neighborhoods.

27. The apparatus according to claim 26, wherein, Each fidelity level is assigned based on the fidelity of the neighborhood.

28. The apparatus according to claim 22, wherein, The future actions are recursively predicted based on policies associated with each agent, and each policy is based on both the neighborhood and the scene structure of the associated agent.