Systems, methods, and devices for analyzing interactions between multiple physical objects

Through the combination of recursive neural networks and graph neural networks, the accuracy problem of time-varying interaction modeling and prediction in the prior art is solved, and high accuracy prediction of the future dynamics of multiple interactive physical objects is achieved.

CN112418432BActive Publication Date: 2025-06-24ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202010850503.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-08-23
Filing Date
2020-08-21
Publication Date
2025-06-24
Estimated Expiration
2040-08-21

AI Technical Summary

Technical Problem

The prior art is difficult to effectively model and predict time-varying interaction configurations, resulting in inaccurate prediction of the future state of the power system.

Method used

Recurrent neural networks and trainable graph neural networks are used to classify pairwise interactions between multiple physical objects into multiple interaction types through an encoder model, and the decoder model is used to predict object feature vectors to achieve modeling and prediction of time-varying interactions.

Benefits of technology

It improves the accuracy of prediction of the future dynamics of multiple interactive physical objects, can handle time-varying interactions more effectively, and improves the automatic reasoning capabilities of systems such as autonomous vehicles and manufacturing processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112418432B_ABST
    Figure CN112418432B_ABST
Patent Text Reader

Abstract

Analyze the interactions between multiple physical objects. The present invention relates to a system (100) for predicting object feature vectors of multiple interacting physical objects. The system uses a decoder model, which includes: a set of propagation models and a prediction model for a set of multiple interaction types; and observation data representing a sequence of observed object feature vectors of the interacting objects. For a sequence of object feature vectors to be predicted for a first interacting object, a sequence of corresponding pairwise interaction types is obtained from the set of multiple interaction types. The propagation data from the second object to the first object is determined using the propagation model indicated by the interaction type sequence. The object feature vector is predicted using a given prediction model based at least on the determined propagation data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a system for predicting object feature vectors of multiple interacting physical objects, and a corresponding computer-implemented method. The present invention further relates to a system for classifying pairwise interactions between multiple physical objects into a set of multiple interaction types, and a corresponding computer-implemented method. The present invention also relates to a system for training a model for the above systems, and a corresponding computer-implemented method. The present invention further relates to a computer-readable medium comprising instructions for performing one of the above methods and / or model data for use in the above methods and systems. Background Art

[0002] Various physical systems can be naturally represented as a collection of objects with relational dependencies among them. For example, the traffic situation around an autonomous vehicle can be regarded as a system in which various other actors such as other vehicles, bicycles, and pedestrians interact with each other and thus affect each other's trajectories. As another example, a manufacturing process can be considered as a system in which various devices such as autonomous robots interact with each other according to interaction dynamics. In this case, although the explicit rules of interaction may not be known, the interactions can still exhibit a certain degree of predictability. In various settings, for example, in a control system for an autonomous vehicle or a monitoring system for a manufacturing process, it is desirable to be able to use this predictability to perform automatic reasoning about the state of such a system, for example, to predict its future state. For example, if a dangerous traffic situation is predicted, the autonomous vehicle can brake. To perform accurate reasoning about such a system, it is desirable to model the system based on individual physical objects and the way they interact with each other.

[0003] In "Neural Relational Inference for Interacting Systems" by Kipf et al. (available at https: / / arxiv.org / abs / 1802.04687 and incorporated herein by reference), a model called the Neural Relational Inference (NRI) model was proposed. This is an unsupervised model that learns to infer the interactions between the components of a dynamical system while simultaneously learning the dynamics purely from observational data. The NRI model takes the form of a variational autoencoder that learns to encode the dynamical system by determining the discrete interaction types between each pair of objects in the system. These encodings can then be used to predict their future dynamics.

[0004] Unfortunately, the model of Kipf et al. assumes that the relationships between the objects of the dynamical system are stationary. This assumption often does not hold in practice. For example, in a traffic scenario, two traffic participants may first move away from each other, in which case the behavior of the first traffic participant may not be much affected by the behavior of the other traffic participant. However, when the traffic participants get closer to each other and, for example, there is a risk of colliding with each other, their behaviors will start to influence each other until, for example, they avoid the collision and they start moving away from each other again. Therefore, it would be desirable to be able to model time-varying interaction configurations and use them to more accurately predict the future dynamics of interacting physical objects. Summary of the Invention

[0005] According to a first aspect of the present invention, there is provided a system for predicting object feature vectors of a plurality of interacting physical objects, wherein the object feature vector of each interacting object at discrete time points includes one or more of the position, velocity, and acceleration of each interacting object at respective discrete time points. The system includes: a data interface configured to access observation data representing a sequence of observed object feature vectors of a plurality of interacting objects provided as input data to the system, and to access decoder model data representing a set of propagation models and parameters of a prediction model for a set of a plurality of interaction types; a processor subsystem configured to determine a sequence of predicted object feature vectors of the plurality of interacting objects for extending the sequence of observed object feature vectors by: for a sequence of object feature vectors of a first interacting object to be predicted at discrete time points, based on the observed object feature vectors, determining, by an encoder model for classifying pairwise interactions between a plurality of physical objects into a set of a plurality of interaction types, a sequence of corresponding pairwise interaction types at discrete time points between the first interacting object and a second interacting object, wherein the encoder model includes: a recurrent neural network for determining a hidden state of the observed object feature vectors; and a set of propagation models subsequently provided with a classification model for determining, according to the hidden state of the observed object feature vectors, the pairwise interaction types at the discrete time points between the first interacting object and the second interacting object, wherein the propagation model is a trainable graph neural network for message passing and information sharing between a plurality of interacting objects; predicting, by a decoder model, an object feature vector of the first interacting object at a current time point, wherein the decoder model includes: a recurrent neural network for determining a hidden state of the observed object feature vectors provided as input data, and a propagation model for determining propagation data according to the determined hidden state of the observed object feature vectors, and the prediction model for predicting an object feature vector of the first interacting particle based on the determined propagation data, wherein the propagation data from the second interacting object to the first interacting object includes the contribution of the interaction between the second interacting particle and the first interacting particle to the change in the hidden state of the first interacting particle at the current time point relative to the hidden state of the first interacting particle at a previous time point, wherein the prediction of the object feature vector of the first interacting particle at the current time point by the decoder model includes the following steps: selecting a propagation model from the set of propagation models according to the pairwise interaction type in the sequence of corresponding pairwise interaction types between the first interacting object and the second interacting object determined by the encoder model at the current time point for the object feature vector;Determine propagation data from the second interacting object to the first interacting object by applying a selected propagation model based on a determined hidden state of previous object feature vectors of the first and second interacting objects; predict an object feature vector using a prediction model based at least on the determined propagation data; and determine a control signal based on the predicted object feature vector, provide the control signal to an actuator of an autonomous vehicle, and use the control signal to control the actuator. According to another aspect of the present invention, a corresponding computer-implemented method is provided, wherein the object feature vector of each interacting object at discrete time points includes one or more of the position, velocity, and acceleration of each interacting object at respective discrete time points, and the method includes: accessing decoder model data representing a set of propagation models and parameters of a prediction model for a set of multiple interaction types, and accessing observation data representing a sequence of observed object feature vectors of multiple interacting objects provided as input data to the system; determining a sequence of predicted object feature vectors of the multiple interacting objects that extends the sequence of observed object feature vectors by: for a sequence of object feature vectors to be predicted at discrete time points of a first interacting object, determining, by an encoder model for classifying pairwise interactions between multiple physical objects into a set of multiple interaction types, a sequence of corresponding pairwise interaction types at discrete time points between the first interacting object and a second interacting object based on the observed object feature vectors; wherein the encoder model includes: a recurrent neural network for determining a hidden state of the observed object feature vectors; and a set of propagation models subsequently provided with a classification model for determining pairwise interaction types at the discrete time points between the first interacting object and the second interacting object based on the hidden state of the observed object feature vectors, wherein the propagation model is a trainable graph neural network for message passing and information sharing between multiple interacting objects;Predict an object feature vector of a first interacting object at a current time point by an encoder model, wherein the decoder model comprises: a recurrent neural network for determining a hidden state of the observed object feature vector provided as input data, and a propagation model for determining propagation data based on the determined hidden state of the observed object feature vector, and a prediction model for predicting an object feature vector of a first interacting particle based on the determined propagation data, wherein the propagation data from the second interacting object to the first interacting object includes the contribution of the interaction between the second interacting particle and the first interacting particle to the change of the hidden state of the first interacting particle at the current time point relative to the hidden state of the first interacting particle at a previous time point, wherein the prediction of the object feature vector of the first interacting particle by the decoder model at the current time point comprises the steps of: selecting a propagation model from a set of propagation models according to the pairwise interaction type in the sequence of the corresponding pairwise interaction types between the first interacting object and the second interacting object determined by the encoder model at the current time point for the object feature vector; determining propagation data from the second interacting object to the first interacting object by applying the selected propagation model based on the determined hidden state of the previous object feature vectors of the first and second interacting objects; predicting an object feature vector using a prediction model based at least on the determined propagation data; and - determining a control signal based on the predicted object feature vector, providing the control signal to an actuator of an autonomous vehicle, and using the control signal to control the actuator. According to a further aspect of the present invention, a system for classifying pairwise interactions between a plurality of physical objects into a set of a plurality of interaction types is proposed, wherein the system is configured as the encoder model of the present invention, wherein the object feature vector of each interacting object includes one or more of the position, velocity, and acceleration of each interacting object, and the system comprises: a data interface configured to access: observation data representing a sequence of observed object feature vectors of a plurality of interacting objects; decoder model data representing parameters of a set of propagation models and a classification model of a set of a plurality of interaction types, wherein the propagation model is a trainable graph neural network for message passing and information sharing between a plurality of interacting objects; a processor subsystem configured to determine a sequence of pairwise interaction types of pairs of interacting objects corresponding to the observation data, the determination including determining a sequence of hidden states of the interacting objects corresponding to the observation data, and the processor subsystem is configured to determine the interaction type between the first and second interacting objects by: selecting a propagation model from the set of propagation models according to the immediately preceding interaction type between the first and second interacting objects;Determine the propagation data from the second interaction object to the first interaction object by applying a selected propagation model based on the data of the first and second interaction objects; determine the hidden state of the first interaction object based at least on the previous hidden state of the first interaction object and the determined propagation data; determine the interaction type using a classification model based at least on the hidden states of the first and second interaction objects; and provide the determined interaction type to the behavior planning module of the autonomous vehicle, and select, through the behavior planning module, a maneuver that minimizes the collision risk and / or ensures a socially compliant interaction with the pedestrian, and control the actuators of the autonomous vehicle based on the selected maneuver. According to another aspect of the present invention, a corresponding computer-implemented method is proposed, wherein the object feature vector of each interaction object includes one or more of the position, velocity, and acceleration of each interaction object, and the method is configured to determine the pairwise interaction type between the first interaction object and the second interaction object at discrete time points based on the observed object feature vector. The method includes: accessing encoder model data representing a set of propagation models for a set of multiple interaction types and the parameters of a classification model, and accessing observation data representing a sequence of observed object feature vectors of multiple interaction objects; determining a sequence of pairwise interaction types of interaction object pairs corresponding to the observation data, the determination including determining a sequence of hidden states of the interaction objects corresponding to the observation data, wherein the interaction type between the first and second interaction objects is determined by: selecting a propagation model from the set of propagation models according to the immediately preceding interaction type between the first and second interaction objects, wherein the propagation model is a trainable graph neural network for message passing and information sharing between multiple interaction objects; determining the propagation data from the second interaction object to the first interaction object by applying the selected propagation model based on the data of the first and second interaction objects; determining the hidden state of the first interaction object based at least on the previous hidden state of the first interaction object and the determined propagation data; determining the interaction type using a classification model based at least on the hidden states of the first and second interaction objects, and providing the determined interaction type to the behavior planning module of the autonomous vehicle, and selecting, through the behavior planning module, a maneuver that minimizes the collision risk and / or ensures a socially compliant interaction with the pedestrian, and control the actuators of the autonomous vehicle based on the selected maneuver. According to still another aspect of the present invention, a system for training an encoder model and a decoder model is proposed, the encoder model being used to classify pairwise interactions between multiple physical objects into a set of multiple interaction types, and the decoder model being used to predict the object feature vectors of multiple interacting physical objects. The system includes: a data interface configured to access decoder model data representing the parameters of the decoder model;Encoder model data representing encoder model parameters; and training data representing a plurality of training instances; a processor subsystem configured to optimize the parameters of a decoder model and an encoder model by: selecting at least one training instance from the training data, the training instance including a sequence of observed object feature vectors of a plurality of interacting objects; based on the encoder model data, determining a sequence of pairwise interaction types of interacting object pairs according to the method of the present invention; and minimizing the loss of recovering the observed object feature vectors by predicting object feature vectors according to the method of the present invention based on the determined sequence of pairwise interaction types and the decoder model data. According to another aspect of the present invention, a corresponding computer-implemented method is provided, where an encoder model is used to classify pairwise interactions between a plurality of physical objects into a set of multiple interaction types, and a decoder model is used to predict object feature vectors of multiple interacting physical objects, the method including: accessing decoder model data representing decoder model parameters, encoder model data representing encoder model parameters, and training data representing a plurality of training instances; optimizing the parameters of the decoder model and the encoder model by: selecting at least one training instance from the training data, the training instance including a sequence of observed object feature vectors of a plurality of interacting objects; based on the encoder model data, determining a sequence of pairwise interaction types of interacting object pairs according to the corresponding steps performed by the encoder model in the method of the present invention; and minimizing the loss of recovering the observed object feature vectors by predicting object feature vectors according to the method of the present invention based on the determined sequence of pairwise interaction types and the decoder model data. According to one aspect of the present invention, there is provided a computer-readable medium including transient or non-transient data, the transient or non-transient data representing instructions that, when executed by a processor system, cause the processor system to execute the computer-implemented method according to the present invention; and / or decoder model data for predicting object feature vectors of multiple interacting physical objects according to the method of the present invention; and / or encoder model data for classifying pairwise interactions between a plurality of physical objects into a set of multiple interaction types according to the method of the present invention.;

[0006] The various measures discussed herein relate to interacting physical objects, such as traffic participants, autonomous robots in a manufacturing plant, physical particles, etc. The observations of the separate physical objects may be available in the form of a sequence of observation object feature vectors of the separate objects. The object feature vectors may indicate the corresponding physical quantities of the physical objects. For example, at a given set of discrete measurement times, the object feature vectors of multiple interacting objects may have been determined, such as their positions, velocities, accelerations, etc. Generally, the observation object feature vectors may be determined based on input signals from various sensors (e.g., cameras, radars, LiDARs, ultrasonic sensors, or any combination thereof). Thus, the object feature vectors of the corresponding objects may represent the state of the physical system at a certain point in time.

[0007] Interestingly, in order to model the potential connectivity dynamics of such interacting objects, the inventors envision using interaction types from a set of multiple interaction types to capture the pairwise interactions between objects at a particular point in time. For example, given the observations of objects at multiple time points, the potential connectivity dynamics corresponding to these observations may be formulated in terms of the interaction types between each pair of objects at each time point. Interestingly, the interaction type between a pair of objects may vary over time. For example, in a sequence of interaction types between a pair of objects, at least two different interaction types may occur. For example, two pedestrians may first move independently of each other, but when they get closer to each other, they may start to affect each other's trajectories. Thus, a more accurate model of the connectivity dynamics can be obtained compared to if it is assumed that the interaction type remains static over time.

[0008] The overall concept of modeling potential connectivity dynamics as a sequence of pairwise interaction types as proposed herein can be used in various devices and methods. Based on a sequence of object feature vectors of multiple interacting objects, predictions can be made to extend these sequences according to future pairwise interaction types between the objects, and the future pairwise interaction types themselves can be predicted. For example, a model referred to herein as a decoder model can be trained to perform such predictions most accurately, such as predicting future traffic situations based on a sequence of observed traffic situations.

[0009] In addition to predicting an object feature vector based on future interaction types, it is also possible to determine past interaction types corresponding to an observed sequence of object feature vectors. A model, referred to herein as an encoder model, can be trained to most accurately determine these interaction types, e.g., to accurately select a specific interaction type for a pair of objects or to accurately determine the probability of a corresponding interaction type for a pair of objects. For example, for each pair of objects and each measured time point, such a latent representation can indicate an estimate of the probability that the two objects interact according to the corresponding interaction type at that time point. The interaction types determined in this way can provide a particularly meaningful representation of past observations, e.g., for use in another machine learning model. For example, pairs of traffic participants in a traffic scenario can be classified according to their interaction type at a certain time point, making it possible to train better models for reasoning about traffic scenarios.

[0010] As an example, the encoder model and / or decoder model can include a neural network. For example, the encoder model and / or decoder model can include at least 1000 nodes, at least 10000 nodes, etc. Detailed examples of possible neural network architectures are provided herein.

[0011] Although labeled data can be used to train the encoder model, e.g., the labeled data includes observed data with a given interaction type, interestingly, this is not required. Instead, in various embodiments, the parameters of the encoder model for determining interaction types from observed data and the decoder model for predicting an object feature vector for a given interaction type can be trained in a common training process; optionally in combination with a model (referred to herein as a prior model) for predicting pairwise interaction types from previous observations and / or their interaction types. To perform this training, it may not be necessary to manually assign meanings to specific interaction types; instead, the training process can determine the appropriate interaction types that best fit the training data. For example, it can be automatically determined that it is useful to distinguish between a fast bicycle and a slow bicycle for bicycles, but not for pedestrians. The number of interaction types can be predefined - but even this is not required, e.g., the number can be a hyperparameter.

[0012] Specifically, the decoder model and the encoder model can be jointly trained to minimize the loss of recovering the observed object feature vector, which is the prediction of the decoder model based on the pairwise interaction types determined by the encoder model. Thus, the decoder model and the encoder model as described herein can be related not only because of their co-design but also in the sense that the encoder model can learn the interaction types that best allow the decoder model to predict the object feature vector, and the decoder model can predict the object feature vector for the specific interaction types that the encoder has been trained to determine.

[0013] To predict the object feature vector or determine the interaction types, various embodiments utilize an ensemble of propagation models for multiple interaction types. In particular, both the decoder model and the encoder model can include propagation models for multiple interaction types, although the propagation models used in these two cases are typically different. Propagation models - which are also referred to as "message passing functions" or "message functions" in the context of graph neural networks - provide a particularly effective way of using information about other interacting objects to make derivations about the interacting objects. A propagation model can be a trainable function applied to an interacting object and the corresponding other interacting objects to obtain the corresponding contributions or "messages" for reasoning about the first interacting object. The corresponding contributions can be aggregated into an overall contribution, or "aggregated message". Then, the overall contribution can be used to reason about the interacting object, e.g., to predict the future object feature vector of the interacting object or to determine the hidden state. Generally, propagation models can allow information to be shared among multiple objects in a general and efficient manner because the propagation models can be trained to be independent of specific objects, and aggregation can allow the use of a flexible number of other objects.

[0014] While the conventional use of propagation models in graph neural networks typically involves using a fixed propagation model, interestingly, the encoder and decoder models can include different propagation models for each interaction type, where, furthermore, the propagation model to be applied is dynamically selected based on the current interaction type between a pair of objects, e.g., at a particular point in time. For example, in the decoder model, the propagated data from one or more other interacting objects to the interacting object can thus be used to predict the object feature vector of the interacting object using a trainable prediction model. In the encoder model, the propagated data from one or more interacting objects to the interacting object can be used to determine the hidden state of the interacting object, and then a classification model can use the hidden state of the interacting object to determine the interaction type. In any case, selecting the propagation model according to the current interaction type can allow the potential connectivity dynamics to be used in the encoder and decoder models in a particularly efficient manner, i.e., by allowing separate propagation models for individual interaction types to be trained and used dynamically depending on the inferred relationships between different objects at a particular point in time.

[0015] Accordingly, models are provided that allow the potential connectivity dynamics between sets of interacting objects inferred effectively from a continuous set of observations to evolve, and / or allow the behavior of such interacting objects corresponding to such connectivity dynamics to be predicted. Interestingly, this can be done without directly measuring the pairwise relationships between objects. The decoder model (e.g., a conditional likelihood model) can define how to generate the observed object feature vectors given a latent representation (e.g., a time-dependent pairwise interaction type). Optionally, the decoder model can be combined with a prior model for predicting the pairwise interaction type to together form a generative model that explains the dynamics of a system with a hidden interaction configuration. For example, by optimizing a loss corresponding to a training dataset, e.g., the variational lower bound on the marginal likelihood of the training data, the decoder model and / or the prior can be trained together with the encoder model similar to a variational autoencoder. As a result, a more expressive description of the potential connectivity is obtained, and more accurate predictions of future dynamics can be made.

[0016] Optionally, based on inferred dynamics, e.g., based on a determined sequence of pairwise interaction types, a control signal can be determined. Such a control signal can be provided to one or more actuators to effect an action in the environment where the observations are made. For example, the control signal can be used to control the physical system from which the observations originate. Such a physical system can be, for example, a computer-controlled machine such as a robot, a vehicle, a household appliance, a manufacturing machine, or a personal assistant; or a system for conveying information such as a surveillance system. As an example, an autonomous vehicle can use the inferred interaction configuration of traffic participants, e.g., as a potential scenario representation. For example, the inferred interaction types can be used as an input to the behavior planning module of the autonomous vehicle to select maneuvers that minimize the risk of collisions with other traffic participants and / or ensure socially compliant interactions with pedestrians.

[0017] Optionally, a trained model is used to generate synthetic observation data for use as training and / or test data in training additional machine learning models, such as neural networks. For example, a sequence of predicted object feature vectors can represent the trajectories of fictional interacting objects at discrete time points. The simulated data can be used for data augmentation, e.g., to train a second machine learning on a larger dataset and / or a dataset for which it is difficult to obtain training data (such as dangerous traffic situations, rare combinations of weather and / or traffic conditions, etc.) without performing further real physical measurements.

[0018] Optionally, the current object feature vector includes one or more of the localization, velocity, and acceleration of the first interacting object. In this way, the spatial configuration of the interacting objects, such as traffic participants in a traffic situation, autonomous robots in a warehouse, etc., can be modeled. For example, the object feature vector of the corresponding object can be derived from a camera image or obtained from the position and / or motion sensors of the object itself, etc.

[0019] Optionally, the decoder model data further represents a set of transition probabilities that the previous pairwise interaction type is succeeded by the current pairwise interaction type. For example, the decoder model can include a Markov model of pairwise interaction types. Such transition probabilities can be used to predict pairwise interaction types; in other words, the transition probabilities can form a prior model that, together with the decoder model, forms a generative model. In particular, by sampling the pairwise interaction types according to the set of transition probabilities, the pairwise interaction type between the first and second interacting objects can be predicted based on the previous pairwise interaction type between the first and second interacting objects. By using pairwise interaction types, a prior model can be obtained with a relatively small number of parameters, resulting in efficient learning and a reduced risk of overfitting.

[0020] Optionally, the pairwise interaction type between the first and second interaction objects is predicted by determining a respective representation of the interaction objects based on previous object feature vectors of the respective interaction objects, and determining the pairwise interaction type based on such representations. Determining the representations and / or using them to predict the interaction type can be performed using a trainable model that is part of a decoder model. Interestingly, in this way, the correlations between different interaction objects can be considered for prediction in a relatively efficient manner. As recognized by the present inventors, in order to consider the correlations between interaction objects, it may be possible to model the attributes of different object pairs at a given time step as conditionally dependent on previous interaction types. However, automatically learning such priors may be inefficient, for example, involving estimating K * N^2 free parameters, where K is the number of interaction types and N is the number of objects. However, this can be avoided by determining the representations of the respective objects based on the object feature vectors of these objects; for example, representations can be determined separately for the interaction objects. Still, by using these representations to determine the pairwise interaction type, in other words, by conditioning the prior on the history of the object trajectories, the correlations between the interaction objects can be considered.

[0021] Optionally, the decoder model data further represents the parameters of a recurrent model such as a gated recurrent unit (GRU) or long short-term memory (LSTM). The representation of the interaction objects for predicting the pairwise interaction type can be determined by sequentially applying the recurrent model to the previous object feature vectors of the interaction objects. In this way, the object trajectories can be efficiently aggregated into the representations of the interaction objects in a manner independent of the number of time steps, and it becomes possible to automatically learn a good way to represent the interaction objects for prediction.

[0022] Optionally, the representation of the interaction objects for predicting the pairwise interaction type is determined from the immediately preceding object feature vectors of the respective interaction agents. For example, in various real-life scenarios, the interaction between objects can be assumed to exhibit first-order Markov dynamics, such as to predict behavior at a time step, and only the information from the immediately preceding time step may be relevant. In such cases, the use of a recurrent model can be avoided; for example, the object feature vectors themselves can be directly used as representations, or the representations can be derived from the object feature vectors. Thus, in such cases, a more efficient model with fewer parameters can be obtained.

[0023] Optionally, the object feature vector can be predicted by using a prediction model to determine the parameters of a probability distribution and sampling an object feature vector from the probability distribution. For example, the prediction model can be applied based at least on the determined propagation data. Since the dynamics of interacting objects are typically non-deterministic and can thus vary broadly, using a model that can consider multiple hypotheses using probability distributions to describe the interaction dynamics may be more effective than directly training a model for a specific prediction. For example, such a probability distribution can be used to generate multiple predictions of the future trajectories of the interacting objects, where the multiple predictions represent various possible evolutions of the system.

[0024] Similarly, by applying a classification model to determine the parameters of a probability distribution of the interaction type and sampling an interaction type from the parameters of the probability distribution, the classification model can also be used to determine the interaction type. In other words, the problem of estimating the interaction type can be formulated from a Bayesian perspective as the problem of evaluating a posterior probability distribution over the interaction types of the interacting objects. The classification model can be applied at least to the hidden states of the first and second objects.

[0025] Optionally, when determining the pairwise interaction type based on the observed data, the hidden states of the interacting objects can be determined by recursively applying a recurrent model such as GRU or LTSM. As discussed above, the propagation data is typically used to determine the hidden states. By using a recurrent model, historical data can also be efficiently and generally used in the hidden representation. For example, the propagation data from the second object to the first object can include the hidden state of the object determined by the recurrent model. However, alternatively or additionally, it is also possible to use the currently observed feature vector as the input to the propagation model.

[0026] Optionally, not only the previously observed hidden states but also the additional hidden states observed below are used to determine the interaction type. In particular, the additional hidden state of the first interacting object can be determined based at least on the additional hidden state of the first interacting object immediately following. Similar to the conventional hidden state, a propagation model can be used to integrate the additional hidden states or object feature vectors of other interacting objects into the additional hidden state, and a recurrent model such as GRU or LTSM can be used to determine the additional hidden state. To determine the interaction type between the first and second interacting objects, the additional hidden state can be used as the input to the classification model. By using the additional hidden state, additional knowledge from future observations can be used to more accurately infer the interaction type.

[0027] Optionally, when training the decoder model and the encoder model, the loss includes the difference between the determined sequence of pairwise interaction types and the predicted interaction types of the decoder model. The parameters used to predict the interaction types can be included in the set of parameters of the decoder model. In other words, the loss can be defined to encourage the decoder model to predict interaction types corresponding to the interaction types determined by the encoder model. For example, the decoder model can include the transition probabilities of interaction types as discussed above, in which case the loss function can minimize the difference between the transitions of interaction types derived from the observed data using the encoder and the transition probabilities given by the decoder model. Similarly, the decoder model can include parameters of a model for determining representations of interaction objects and / or for determining pairwise interaction types based on such representations, and the loss function encourages these models to predict the interaction types given by the encoder model. Thus, the decoder model can be trained to accurately predict pairwise interaction types, allowing the decoder model to not only predict interactions of a given interaction type but also predict the interaction types themselves.

[0028] Those skilled in the art will appreciate that the above-mentioned embodiments, implementations, and / or alternative aspects of the present invention can be combined in any manner deemed useful in two or more.

[0029] Any computer-implemented method and / or modification and variation of any computer-readable medium can be implemented by those skilled in the art based on this description, corresponding to the described modifications and variations of the corresponding system. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] These and other aspects of the present invention will be apparent from the embodiments and the drawings, and will be further elucidated by reference to the embodiments and reference to the drawings, which are described as examples in the following description, in which:

[0031] Figure 1 A system for predicting object feature vectors of multiple interacting physical objects is shown;

[0032] Figure 2 A system for classifying pairwise interactions between multiple physical objects into a set of multiple interaction types is shown;

[0033] Figure 3 A system for training an encoder model and a decoder model is shown;

[0034] Figure 4 A detailed example of how to predict object feature vectors of multiple interacting physical objects is shown;

[0035] Figure 5Shows a detailed example of how to classify pairwise interactions between multiple physical objects into a set of multiple interaction types;

[0036] Figure 6 Shows a detailed example of a decoder model;

[0037] Figure 7 Shows a detailed example of an encoder model;

[0038] Figure 8 Shows a computer-implemented method for predicting object feature vectors of multiple interacting physical objects;

[0039] Figure 9 Shows a computer-implemented method for classifying pairwise interactions between multiple physical objects into a set of multiple interaction types;

[0040] Figure 10 Shows a computer-implemented method for training an encoder model and a decoder model;

[0041] Figure 11 Shows a computer-readable medium including data.

[0042] It should be noted that the figures are purely illustrative and not drawn to scale. In the figures, elements corresponding to elements already described may have the same reference numerals. Detailed Description

[0043] Figure 1 Shows a system 100 for predicting object feature vectors of multiple interacting physical objects. The system 100 may include a data interface 120 and a processor subsystem 140, and the processor subsystem 140 may communicate internally via a data communication 121. The data interface 120 may be used to access observation data 030 representing a sequence of observed object feature vectors of multiple interacting objects. The data interface 120 may also be used to access decoder model data 041, which represents the parameters of a set of prediction models and propagation models of a set of multiple interaction types. The decoder model data 041 can be obtained by training a decoder model together with an encoder model according to the methods described herein, for example, by Figure 3 The system 300 of, together with the encoder model.

[0044] The processor subsystem 140 may be configured to access data 030, 041 during operation of the system 100 and using the data interface 120. For example, as Figure 1As shown, data interface 120 can provide access 122 to an external data storage device 021, which can include the data 030, 041. Alternatively, the data 030, 041 can be accessed from an internal data storage device that is part of system 100. Alternatively, the data 030, 041 can be received from another entity via a network. In general, data interface 120 can take various forms, such as a network interface to a local area network or a wide area network (e.g., the Internet), a storage interface to an internal or external data storage device, etc. Data storage device 020 can take any known and suitable form.

[0045] Processor subsystem 140 can be configured to determine, during operation of system 100 and using data interface 120, a sequence of predicted object feature vectors for a plurality of interacting objects, the sequence of predicted object feature vectors for the plurality of interacting objects extending an observed object feature vector sequence. To determine the sequence, processor subsystem 140 can be configured to obtain, for a sequence of object feature vectors to be predicted for a first interacting object, a sequence of corresponding pairwise interaction types between the first interacting object and a second interacting object. To determine the sequence, processor subsystem 140 can further predict an object feature vector of the first interacting object. To predict the object feature vector, processor subsystem 140 can select a propagation model from a set of propagation models according to the pairwise interaction type of the object feature vector. To predict the object feature vector, processor subsystem 140 can further determine propagation data from the second interacting object to the first interacting object by applying the selected propagation model based on previous object feature vectors of the first and second interacting objects. To predict the object feature vector, processor subsystem 140 can further predict the object feature vector using a prediction model based at least on the determined propagation data.

[0046] As an optional component, system 100 can include an image input interface 160 or any other type of input interface for obtaining sensor data 124 from a sensor such as camera 071. Processor subsystem 140 can be configured to determine observed data 030 from sensor data 124. For example, a camera can be configured to capture image data 124, and processor subsystem 140 is configured to determine an object feature vector of observed data 030 based on the image data 124 obtained via data communication 123 from input interface 160.

[0047] As an optional component, system 100 may include an actuator interface (not shown) for providing actuator data to an actuator, the actuator data causing the actuator to effect an action in the environment of system 100. For example, the processor subsystem 140 may be configured to determine the actuator data at least in part based on a sequence of predicted object feature vectors—e.g., by inputting the predicted object feature vectors into a machine learning model for determining actuator data. For example, the machine learning model may classify a sequence of predicted object feature vectors into normal and abnormal situations such as a collision risk, and activate a safety system such as a brake in the event of detecting an abnormal situation.

[0048] Reference will be made to Figures 4 - 7 further elaborate on the various details and aspects of the operation of system 100, including its optional aspects.

[0049] Generally, system 100 may be embodied as a single device or apparatus or be embodied in a single device or apparatus such as a workstation (e.g., laptop- or desktop-based) or a server. The device or apparatus may include one or more microprocessors executing appropriate software. For example, the processor subsystem may be embodied by a single central processing unit (CPU), but also by a system or combination of such a CPU and / or other types of processing units. The software may have been downloaded and / or stored in a corresponding memory such as volatile memory like RAM or non-volatile memory like flash memory. Alternatively, functional units of the system such as a data interface and the processor subsystem may be implemented in the device or apparatus in the form of programmable logic, e.g., as a field programmable gate array (FPGA) and / or a graphics processing unit (GPU). Generally, each functional unit of the system may be implemented in the form of a circuit. Note that system 100 may also be implemented in a distributed manner, e.g., involving different devices or apparatuses such as servers distributed (e.g., in the form of cloud computing).

[0050] Figure 2System 200 is shown for classifying pairwise interactions between multiple physical objects into a set of multiple interaction types. System 200 may include a data interface 220 and a processor subsystem 240, which may communicate internally via a data communication 221. The data interface 220 may be used to access observation data 030 representing a sequence of observation object feature vectors of multiple interacting objects. The data interface 220 may be further used to access encoder model data 042, which represents a set of propagation models of a set of multiple interaction types and the parameters of a classification model. The processor subsystem 240 may be configured to access data 030, 042 during operation of the system 200 and using the data interface 220. Similar to the data interface 120, various implementation options are possible. An access 222 to an external data storage device 022 including data 030, 042 is shown. It may be obtained by training an encoder model together with a decoder model according to the methods described herein, for example, by Figure 3 of system 300 to obtain the encoder model data 042.

[0051] The processor subsystem 240 may be configured to, during operation of the system, determine a sequence of pairwise interaction types of pairs of interacting objects corresponding to the observation data. The determination may include determining a sequence of hidden states of the interacting objects corresponding to the observation data. As part of this, the processor subsystem 240 may be configured to determine the interaction type between a first and a second interacting object. To determine the interaction type, the processor subsystem 240 may select a propagation model from the set of propagation models according to the immediately preceding interaction type between the first and the second interacting objects. To determine the interaction type, the processor subsystem 240 may further determine propagation data from the second interacting object to the first interacting object by applying the selected propagation model based on the data of the first and second interacting objects. To determine the interaction type, the processor subsystem 240 may further determine the hidden state of the first interacting object based at least on the previous hidden state of the first interacting object and the determined propagation data. To determine the interaction type, the processor subsystem 240 may further determine the interaction type using the classification model based at least on the hidden states of the first and second interacting objects.

[0052] As an optional component, system 200 may include an image input interface 260 or any other type of input interface for obtaining sensor data 224 from a sensor such as camera 072. The processor subsystem 240 may be configured to determine observation data 030 from the sensor data 224. For example, the camera may be configured to capture image data 224, and the processor subsystem 240 is configured to determine the object feature vector of the observation data 030 based on the image data 224 obtained via the data communication 223 from the input interface 260.

[0053] As an optional component, system 200 may include an actuator interface 280 for providing actuator data 226 to an actuator, the actuator data 226 causing the actuator to perform an action in the environment 082 of the system 200, where the environment 082 of the system 200 is, for example, the environment from which the observation data is derived. For example, the processor subsystem 140 may be configured to determine a control signal at least in part based on the determined sequence of pairwise interaction types and provide the control signal as control signal 226 to the actuator interface via the data communication 225. For example, a control signal may be used to control a physical system in the environment 082 where the observation is being made.

[0054] Reference will be made to Figures 4 - 7 Various details and aspects of the operation of system 200 will be further elaborated, including its optional aspects. Similar to system 100, system 200 may be embodied as a single device or apparatus or be embodied in a single device or apparatus and implemented in a distributed manner, etc.

[0055] As a specific example, system 200 may be an automotive control system for controlling a vehicle. The sensor 072 may be a camera that provides an image, and the observation data 030 is determined based on the image. The vehicle may be an autonomous or semi-autonomous vehicle, but system 200 may also be a driver assistance system for a non-autonomous vehicle. For example, as discussed, such a vehicle may incorporate system 200 to control the vehicle based on an image obtained from camera 072. In this particular example, system 200 may be configured to control the vehicle by inputting the determined sequence of pairwise interaction types into a behavior planner module, which outputs a signal 226 for controlling the machine actuator. For example, the actuator may be caused to control the steering and / or braking of the vehicle. For example, the control system may control the electric motor of the vehicle to perform (regenerative) braking based on detecting an abnormal situation, where the abnormal situation is based on the determined sequence of pairwise interaction types.

[0056] Figure 3 System 300 for training an encoder model and a decoder model is shown. The encoder model may be used, for example, by Figure 2System 200 classifies pairwise interactions between multiple physical objects into a set of multiple interaction types. A decoder model can be used to predict, for example, object feature vectors of multiple interacting physical objects by Figure 1 System 100. System 300 may include a data interface 320 and a processor subsystem 340, which may communicate internally via a data communication 321. The data interface 320 can be used to access decoder model data 041 representing decoder model parameters. The data interface 320 can be used to further access encoder model data 042 representing encoder model parameters. The data interface 320 can be used to further access training data 050 representing multiple training instances. The processor subsystem 340 may be configured to access data 041, 042, 050 during operation of the system 300 and using the data interface 320. Similar to the data interface 120, various implementation options are possible. An access 322 to an external data storage device 023 including data 041, 042, 050 is shown.

[0057] The processor subsystem 340 may be configured to optimize the parameters of the decoder model and the encoder model during operation of the system. To optimize the parameters, the processor subsystem 340 may select at least one training instance from the training data 050. The training instance may include an observed object feature vector sequence of multiple interacting objects. To optimize the parameters, the processor subsystem 340 may further determine a pairwise interaction type sequence for pairs of interacting objects based on the encoder model data 042 according to methods described herein, such as Figure 9 Method 900. To optimize the parameters 041, 042, the processor subsystem 340 may further minimize the loss of recovering the observed object feature vectors by predicting the object feature vectors based on the determined pairwise interaction type sequence and the decoder model data 041 according to methods described herein, such as Figure 8 Method 800.

[0058] As an optional component, system 300 may further include an input interface (not shown) for obtaining sensor data to, for example, derive the training data 050 from the sensor data, and / or an actuator interface (not shown) for providing actuator data to an actuator, the actuator data causing the actuator to perform an action in the environment of the system. For example, the input and actuator interfaces of system 100 or system 200 may be used. Various details and aspects of the operation of system 300, including its optional aspects, will be further elaborated with reference to Figures 4 - 7 Various details and aspects of the operation of system 300, including its optional aspects, will be further elaborated with reference to. Similar to system 100, system 300 may be embodied as a single device or apparatus or be embodied in a single device or apparatus, implemented in a distributed manner, and so on.

[0059] Figure 4Shows a detailed but non-limiting example of how to use a decoder model to predict object feature vectors of multiple interacting physical objects. The decoder model in this figure is parameterized by a set of parameters DPAR490.

[0060] Mathematically, a set of multiple interacting objects can be modeled by a set v of N interacting objects. For example, the number N of interacting objects can be at most or at least 10, or at most or at least 100. Throughout the text, the index set of objects can be labeled as and the set of objects except the j-th object can be labeled as Multiple interacting objects can be considered to form a fully connected graph This graph has N nodes and directed edges e = (v, v′) ∈ ε without self-loops.

[0061] The object features of multiple interacting objects can be predicted based on observation data representing a sequence of observed feature vectors of the multiple interacting objects. The observed feature vectors can be collectively labeled as a set of measurable feature trajectories where represents a sequence of observed feature vectors of an individual interacting object, also known as a node trajectory, and indicates the set of object feature vectors at time t. Typically, all object feature vectors have the same length, e.g., for some dimension D For example, in various embodiments, the object feature vectors can include one or more of position, velocity, and acceleration. For example, in the case of two-dimensional observations, position, velocity, and / or acceleration can be two-dimensional vectors, and in the case of three-dimensional observations, position, velocity, and / or acceleration can be three-dimensional vectors. Alternatively or additionally, other object features can be used. For example, the number of object features can be one, two, three, or at least four. The number of observed time points can be one, two, at least five, or at least ten. It can be assumed that the number of interacting objects remains constant over time.

[0062] As an example, shown in the figure is the set 400 of observed object feature vectors at time t = τ - 1. In the figure, solid lines are used to indicate that these object feature vectors are known, e.g., included in the observation data given as input to the system. In this case, the set 400 of object feature vectors is schematically represented in box 400 as a set of interacting object positions, as shown by the circles. For example, circle 401 representing the object feature vector of the first interacting object and circle 402 representing the object feature vector of the second interacting object are shown.

[0063] In various embodiments, pairwise interaction types from a set of multiple interaction types are ascribed to interaction objects. When using a decoder model, the set of interaction types is typically fixed. For example, there can be two different interaction types; at least or at most three; or at least or at most five. Although it is possible to assign meanings to the interaction types by hand, e.g., by annotating the training dataset, this is not necessary; instead, during training, the encoder model can be trained to assign interaction types to interaction objects in such a way that the decoder model can optimally predict object feature vectors based on the interaction types.

[0064] Throughout the text, pairwise interaction types can be denoted by the value . The symbol can be used to indicate a set of interaction type trajectories, e.g., the interaction type of a pair of interaction objects in the evolution of an interaction object system. The symbol can denote a single trajectory. Interestingly, typically at certain points t1, t2, e.g., for one or more interaction objects, the interaction type between two objects changes over time. The symbol can be used to denote the interaction type at time t. Here, can be a discrete variable encoding the type of the interaction type (e.g., the type of the edge (v , v i , v j ) in a graph, for example) at time t. The symbol can be used to denote a set of multiple interaction types.

[0065] This figure shows an example of how a decoder model can be used to predict a set 403 of object feature vectors of a set of interacting physical objects at time t = τ, e.g., so as to extend an earlier observation 400 at t = τ - 1 and possibly also earlier additional observations. The dashed lines are used to denote the information to be predicted. As understood by a person skilled in the art, predictions for multiple future time steps can be obtained by repeating the following process multiple times, e.g., at least two or at least five times.

[0066] Mathematically, the decoder model can be described by the following log-likelihood factorization:

[0067]

[0068] where x τ<t = (x 0 ,..., x t-1 ). In other words, the decoder model can assume that the object feature vectors observed at time t are conditionally independent given zt Past relationship graph connectivity z τ<t = (z 0 , ..., z t-1 ) and the observed history x τ<t = (x 0 , ..., x t-1 ). The factors p(z 0 ) and p(z t| z t-1 ) can be regarded as prior models that describe the probabilities of specific interaction types occurring. Instead of p(z t |z t-1 ), other formulations of the prior model are possible, and several of these possibilities are discussed below.

[0069] Given the interaction type and the previous object feature vectors, the factors p(x 0 |z 0 ) and p(x t |x τ<t , z t ) can form a decoder model for predicting object feature vectors. Effectively, the decoder model can capture the parameters that govern the conditional probability distribution of the observed data X for a given interaction type Z by evaluating the distribution as a product of T factors, where T is equal to the number of measurement time points. For each time point t, given the past feature trajectory and the interaction type, the corresponding factor can be the conditional probability density over the features of all interacting objects observed at time t. In general, any differentiable probability density function can be used to model these factors. For example, if the observed features take values in a D-dimensional real coordinate space, a multivariate normal density is a suitable choice.

[0070] In various embodiments, it can be assumed that the observed object feature vectors have Markov dynamics conditional on z t , in which case the conditional likelihood factor can be simplified to:

[0071] p(x t |x τ<t , z t ) = p(x t |x t-1 , z t ).

[0072] Specifically shown are first and second object feature vectors 404 and 405 that respectively form the predicted evolutions of object feature vectors 401 and 402. To predict the first object feature vector 404, a corresponding pairwise interaction type 411 between a first interacting object 404 and a second interacting object 405 can be obtained. Generally speaking, there are several possibilities for obtaining the interaction type 411.

[0073] In some embodiments, the interaction type 411 is obtained as an input, for example, from another unit or from a user. This may be desirable, for example, in order to generate simulation data from scratch. Thus, the interaction type can be manually fixed, and the decoder can generate a trajectory conditioned on this fixed input.

[0074] In other embodiments, the interaction type 411 can be determined by an encoder model as described, for example, in reference Figure 5 and Figure 7 when training the decoder model, as discussed in more detail elsewhere.

[0075] In still other embodiments, the interaction type 411 can be predicted by a prior model, and the decoder model together with this prior model forms a generative model. This may be a suitable choice for predicting the future behavior of interacting objects after one or more time steps have been observed.

[0076] A first possibility for predicting the pairwise interaction type is based on a set of transition probabilities by which a current pairwise interaction type is preceded by a previous pairwise interaction type. In this case, for example, by sampling the pairwise interaction type 411 according to the set of transition probabilities, the pairwise interaction type 411 at time t = τ can be predicted based on the previous interaction type 410 between the same interacting objects at time t = τ - 1. The previous interaction type 410 may have been predicted as described herein, or if the previous interaction type 410 corresponds to a time point at which observed data exists, the encoder model can be used to obtain the previous interaction type 410, as described, for example, in reference Figure 5 and Figure 7 as described.

[0077] For example, mathematically, the transition probabilities can be captured by a Markov state transition prior over graph connectivity as follows:

[0078]

[0079] where, an one - hot encoded representation of the interaction type is adopted, and;

[0080]

[0081] Here, the values and can be included in the set of parameters DPAR. In this particular example, the prior parameters are shared across edges and time steps. For example, the prior on graph connectivity can include a time-homogeneous Markov chain for each graph edge. However, by allowing parameters that depend on time and / or on the edges, the prior model can be made more expressive. It is also possible to use a more powerful prior, where the attributes of different edges at a given time step, such as interaction types, are only conditionally independent given the entire graph configuration at the previous time step. However, such a prior may have more parameters, and thus in most cases, the above prior is desirable.

[0082] A second possibility for predicting pairwise interaction types can involve determining the corresponding representations of the interacting objects based on the previous object feature vectors of the corresponding interacting objects, and determining the pairwise interaction types based on the representations of the interacting objects. In other words, the pairwise interaction types may be conditioned on the history of the node trajectories. Thus, a graph transition prior can be obtained that is more expressive than a graph transition prior using transition probabilities, but has better scaling properties compared to a model where the interaction types are conditionally dependent on the entire graph configuration at the previous time step.

[0083] As a specific example, the interaction types can be conditioned on the node trajectory history as follows:

[0084]

[0085] where

[0086]

[0087] Specifically, through the above operation ρ d By way of example, for instance, the interacting object representations, such as d can be determined by sequentially applying the recurrent model ρ to the previous object feature vectors of the interacting objects In other words, a recurrent operation can be used that, given the last hidden state and the node observations updates the deterministic hidden state for each node j at time t Various recurrent models are known in the art and can be used here. For example, a gated recurrent unit (GRU) or a long short-term memory (LSTM) network can be used. In other embodiments, the representation of the interacting objects It can be determined based on the immediately preceding object feature vector of the corresponding interaction agent, for example For example, as also discussed elsewhere, such a choice may be appropriate in cases where the observed data has conditional Markov dynamics.

[0088] It can be noted that in the above example, the deterministic hidden state is independently updated for each node. In fact, this captures an assumption that the features observed at time t are conditionally independent of the past graph connectivity and the history of observations x t given z τ<t . In particular, it may not be necessary to use the messages from the sender nodes as inputs to the hidden state recursion. Thus, a model with fewer parameters can be obtained, which can better generalize to unseen data and / or prevent overfitting.

[0089] To predict the interaction type based on the representation of the interacting objects in the above specific example, learnable functions and are used, which effectively implement a single-layer graph neural network to share information about past trajectories between nodes. In this example, softmax is used to obtain a probability distribution over the predicted interaction types, from which the interaction type can be sampled. Those skilled in the art will envision several variants for determining the interaction type based on the representation of the interacting objects, such as other types of graph neural networks.

[0090] The interaction type 411 between the first interacting object and the second interacting object has been obtained, and the corresponding object feature vector 404 of the first object can be predicted. Interestingly, as a particularly efficient way to condition the prediction of the interaction type, the decoder model can include a corresponding propagation model for the corresponding interaction type, such as a message passing function. The corresponding propagation model can be parameterized by corresponding parameters included in the parameters DPAR of the decoder model. In operation Fwp 420, the corresponding propagation model can be applied to provide corresponding propagation data PD 430 from the second interacting object to the first interacting object.

[0091] A propagation model can be applied based on the previous object feature vectors of the interacting objects. For example, the propagation model can include a first input calculated as a function of the previous object feature vector of the first interacting object and a second input calculated as a function of the previous object feature vector of the second interacting object. Alternatively, one or more of the previous object feature vectors can be directly used as inputs to the propagation model. In various embodiments, the propagation model can be applied to the representations of the interacting objects as discussed above. More specifically, the application of a propagation model, such as a message function, can be performed as follows to determine propagation data, such as a message. To calculate the propagation data from the i-th interacting object to the j-th interacting object According to the k-th interaction type 411, the message function for that specific interaction type can be applied as follows and

[0092]

[0093] In this case, two message functions specific to the interaction type 411 are used to calculate the respective parts of the propagation data, but there can also be only one message passing function, or more than two message passing functions.

[0094] Typically, multiple propagation models for different interaction types use the same model, such as a neural network architecture, but with different parameters. However, it may be beneficial to define a specific interaction type to indicate the lack of interaction between two objects of the system. For that interaction type, the propagation model can calculate the propagation data based only on the previous object feature vector of the first interacting object. For example, the interaction type k = 0 can be defined using the following unary message operation:

[0095]

[0096] where and denotes the message from node i to node j under the connection of type k = 0. Including such an interaction type may be an efficient way to encourage the model to learn how to handle non-interacting objects.

[0097] To predict the object feature vector 404 of the first interacting object, the propagation data PD of one or more second interacting objects can be aggregated into aggregated propagation data. Preferably, the aggregation is performed according to a permutation-invariant operation to obtain a more general model. For example, a summation operation can be used. For example, the aggregated propagation data can be calculated and

[0098] Similarly, in this specific example, the aggregated propagation data includes the propagation data of two separate slices corresponding to the two message functions given above.

[0099] Finally, to obtain the object feature vector 404 based on the propagation data PD, in operation Pred 440, the object feature vector can be determined by using a prediction model parameterized by a set of parameters DPAR. For example, by aggregating the propagation data into aggregated propagation data as discussed above, the prediction model can be applied based at least on the determined propagation data. For example, the corresponding mapping and can be applied to the aggregated messages and To assist in training the decoder model, the f d and functions are preferably differentiable. For example, the functions can be implemented using small neural networks.

[0100] The prediction model can directly provide the predicted object feature vector. However, as discussed elsewhere, it may be beneficial for the prediction model to instead provide the parameters of a probability distribution from which the predicted object feature vector 404 can be sampled. For example, the above message passing functions can be used to calculate the mean and covariance matrix to sample the object feature vector from it. In other words, in the above example, the corresponding mapping and can be applied to the aggregated messages and to obtain the parameters μ and Σ of the normal distribution.

[0101] In summary, by combining the various options discussed above, a decoder model defined via the following state transitions can be obtained:

[0102]

[0103] Where:

[0104]

[0105] As highlighted in this example, the one-hot encoded interaction type z ij can effectively implement an attention mechanism that controls the influence of node i on node j. For example, if the interaction type is encoded as a missing interaction, setting z ijk = 1 can effectively prevent messages from being passed from i to j.

[0106] As a specific example, Figure 6 is schematically shown, for example, in Figure 1 andFigure 5 Possible neural network architecture 600 of the decoder model used in Figure 5 . As shown in the figure, the neural network 600 can take as input the observation data of the sequences 601, 602, 603 of observation object feature vectors representing the corresponding interacting objects. To predict the object feature vectors of the interacting objects, the representation of the corresponding interacting objects can be determined by sequentially applying a recurrent model 610 - in this case, a gated recurrent unit (GRU) - to the previous object feature vectors of the corresponding interacting objects. Generally, any type of recurrent neural network unit, such as GRU or LSTM, can be used.

[0107] A graph neural network (GNN) 630 can be used to determine the parameters of the probability distribution for predicting the object feature vectors. The GNN can take as input the sequence embedding of the previous object feature vectors together with the current latent graph configuration (e.g., interaction type), and can output a node-level graph embedding, which provides the parameters of the probability distribution for predicting the object feature vectors. The graph neural network 630 can be parameterized with various numbers of layers and various types of aggregation functions, although permutation-invariant and / or differentiable aggregation functions are preferably used.

[0108] For example, in the sense of using the representation of the interacting objects calculated by the GRU to determine the propagation data from the second interacting object to the first interacting object, the GNN 630 can perform edge-wise aggregation 632. The corresponding representation can be input into a message passing function 634, which is implemented as a multi-layer perceptron (MLP) in this case. To use the propagation data to predict the object feature vectors, message aggregation 636 can be used. For example, the propagation data of the corresponding second interacting object according to the corresponding interaction type can be aggregated. As described herein, the corresponding interaction type can be obtained, for example, by applying the encoder model Enc 620; by applying the prior model as described herein; or by generating a specific type of interaction as an input to the model. The aggregated propagation data can be input into a prediction model 638, which is a multi-layer perceptron in this case, to obtain the parameters μ, ∑ of the probability distribution, from which the predicted object feature vectors can be sampled.

[0109] Figure 5 A detailed but non-limiting example of how to use an encoder model to classify pairwise interactions between multiple physical objects into a set of multiple interaction types is shown. The encoder model in this figure is parameterized by a set of parameters EPAR 590.

[0110] In this example, an encoder model is used to classify the type of interaction 580 between a first interacting object and a second interacting object at a specific time point t = τ. This classification can use a sequence of observed feature vectors of multiple interacting objects as input. For example, the feature vectors observed at the following three time points are shown in the figure: at t = τ - 1, 500; at t = τ, 503; and at t = τ + 1, 506. The first interacting object has object feature vectors 501, 504, and 507; the second interacting object has object feature vectors 502, 505, and 508. The number of objects, the type of interaction, etc. are as discussed for Figure 4 Although in this example, a single type of interaction 580 is classified, typically, the type of interaction can be determined for each interacting object at each observed time point.

[0111] The encoder model can determine the interaction type 580 by having the classification model CI 570 output parameters that govern the posterior distribution of the latent interaction type Z, conditioned on the observed data X500, 503, 506. Such a model can be interpreted as an approximation q φ (Z|X) of the true posterior p(z|x) obtained, for example, by variational inference. This distribution can be evaluated as a product of T factors, where T is equal to the number of measurement time points. Such factors can correspond to the conditional probability density of the interaction type at time t, e.g., given the observed data X and the interaction type history z τ<t of z t , for example:

[0112]

[0113] In general, such a distribution can correspond to a complete probability graph over a set of latent variables {z 0 ,..., z T}. The conditional factor q φ (z t |z τ<t , X) can be defined as a function However, this may lead to an undesirably large number of inputs and parameters to be trained.

[0114] However, in various embodiments, the classification model can determine the probability of the corresponding interaction type of a pair of interacting objects based on the hidden state of the interacting objects and optionally also based on an additional hidden state based on future observations, and the hidden state of the interacting objects can be computed relatively efficiently based on past observations. In particular, surrogate forward and backward message passing recursions can be used to compute the hidden state and the additional hidden state. In this way, the relationship history z can be relatively efficiently obtained fromτ<t and calculate variational factors corresponding to each time step from, for example, the node trajectories from time 0 until time T and in X Specifically, the approximate posterior can be modeled as:

[0115]

[0116] where σ(·) denotes the softmax function:

[0117]

[0118] To determine the input to the classification model CI, for each past observation, a hidden state can be calculated for each interacting object and based on these past hidden states, a hidden state can be calculated for the current observation. Optionally, for each future observation, an additional hidden state can be calculated for each interacting object and based on the additional hidden states of the future observations, an additional hidden state can be calculated for the current observation. The classification model CI can use the hidden state and / or the additional hidden state as input. For example, to determine the interaction type 580 between the i-th interacting object and the j-th interacting object, the classification CI can calculate the above as the hidden state of the interacting objects and optionally also the additional hidden state of the function for the probability distribution q φ (z t |z τ<t , X) of the softmax argument For example, a differentiable mapping parameterized by the parameter EPAR can be used to determine

[0119]

[0120] At t = 0, for example, one can use

[0121] For example, the hidden state 541 of the first interacting object at t = τ can be determined in the forward recursion operation Fwr 520 using the hidden state 540 of the interacting object at t = τ - 1. In fact, the forward recursion Fwr can filter information from the past. Thus, the hidden state 541 of the interacting object can be determined based on the previous hidden state 540 of the interacting object and the current input data. For example, a recurrent neural network unit can be used to calculate the hidden state.

[0122] Interestingly, to determine the current input data, information from other interacting objects can be propagated, similar to the decoder model. Specifically, a corresponding propagation model can be defined for a corresponding interaction type by the corresponding parameters included in the set EPAR. To propagate data from the second interacting object to the first interacting object, at the previous time point t = τ - 1, the propagation model of the previous interaction type 510 can be used to determine the propagation data PD530, which can then be used to update the hidden state. For example, the propagation data PD from the i-th interacting object to the j-th interacting object with the interaction type k510 can be calculated using their previous hidden states of the propagation model to be The propagation data PD of other interacting objects can be aggregated into aggregated propagation data for calculating the hidden state For example, the aggregated propagation data of the j-th interacting object can be calculated as:

[0123]

[0124] Instead of using the previous hidden state or in addition to using the previous hidden state, the propagation model can also take the currently observed feature vector as input. For example,

[0125]

[0126] The aggregated propagation data can possibly be used as the current input data in combination with the currently observed object feature vector to calculate the hidden state. For example,

[0127]

[0128] where [·;·] indicates vector concatenation. Initialization can be performed independently for each node, for example, by applying the mapping to the first observed feature

[0129]

[0130] For example, the recurrence for determining the hidden state can be implemented using a GRU or LSTM network. For example, the propagation model can be implemented by a small feed-forward neural network.

[0131] As discussed, in addition to the hidden state 541, another hidden state 561 can be used to determine the interaction type 580. In the backward recursion operation Bwr 550, the hidden state 561 of the interaction object at time t = τ can be determined based at least on, for example, the immediately following another hidden state 560 at time t = τ+1. Thus, effectively, information from the future can be smoothed for determining the interaction type. The hidden state of the interaction object at the last observed time point can be initialized as follows

[0132]

[0133] where is a differentiable mapping from to At other time points, the hidden state 561 can be based on the next hidden state 560 calculated by a recurrent model such as, for example, a GRU or LTSM network ;

[0134]

[0135] For example, the input to the prediction model can include the observed feature vector 504 at that time point, and / or aggregated propagation data from other interaction objects For example,

[0136]

[0137] A common propagation model can be used, where the parameters are included in the encoding model parameters EPAR.

[0138] As a specific example, Figure 7 schematically shows a neural network architecture 700 used as a decoder model in, for example, Figure 2 and Figure 6 As shown in the figure, the neural network 700 can take as input the observation data of sequences 701, 702, 703 of observation object feature vectors representing the corresponding interaction objects.

[0139] To determine the interaction type between the first and second interaction objects, a forward pass 730 can be performed, where the hidden states of the corresponding interaction objects are determined, and optionally, a backward pass 780 can also be performed, where the additional hidden states of the corresponding interaction objects are determined.

[0140] As an input to the forward pass 730, a graph neural network (GNN) 710 can be used to determine the aggregated propagation data of the interaction objects at a specific time point, e.g., The aggregated propagation data can aggregate the propagation data from the second interaction object to the first interaction object, where the first interaction object is determined by applying a propagation model such as and the like, and the propagation model such as is selected according to the type of interaction between the interaction objects, such as . The aggregated propagation data can be concatenated with the current observation data 720 and input into a recurrent model, which is GRU 731 in this case. Applying GRU at the corresponding time point can result in a hidden state at the time point, such as to determine the type of interaction within the time point. Thus, the hidden state can be determined for two interaction objects, such as to determine the type of interaction between the two interaction objects.

[0141] In the backward pass 780, the additional hidden state of the interaction object, such as can be based on the immediately subsequent additional hidden state of the interaction object by applying the recurrent model again, which is GRU 781 in this case.

[0142] Then, the calculated hidden state and optionally the additional hidden state can be subject to edge-by-edge concatenation 740. For example, to determine the type of interaction between the first and the second interaction objects, the hidden state and the additional hidden state can be concatenated, such as a classification model 750, which is a multi-layer perceptron in this case, can be applied to obtain the softmax input, for example, Finally, by applying the softmax function 760, the probabilities 770 of the corresponding relationship types of the first and the second interaction objects can be obtained, and the predicted relationship type can be obtained via sampling from the probabilities 770.

[0143] Returning to Figure 3 , as discussed, for example, with reference to Figures 4 - 7 , the training of the encoder and decoder models is now further elaborated. To train the encoder and decoder models, a stochastic method can be applied to solve inference and learning simultaneously. In particular, the variational lower bound of the marginal likelihood of the training data 050 with respect to the model parameters can be optimized. These parameters can include the parameters 041 of the decoder model and the parameters 042 of the encoder model. In various embodiments, the stochastic estimate of the variational lower bound can be optimized, where q φ (Z|X) is a variational approximation of the posterior of the latent variable applied in the encoder model, parameterized by φ, and p(X) denotes the marginal likelihood, for example, p(X) = ∫p(X, Z)dZ.

[0144] Specifically, as discussed, the training data 050 may represent multiple training instances, where one instance includes an observation object feature vector sequence of multiple interacting objects. For example, at least 1000 or at least 10000 training instances may be available. To train the model, training instances X = {x1,..., x N} can be obtained, for example, randomly selected from the training data 050 and processed by the encoder model to obtain a finite sample set extracted from the variational posterior distribution over interaction types. Such samples Z can provide interaction types z for the set of interacting objects (i, j) ij . The samples returned by the encoder model can then be processed by the decoder model to obtain the parameters of the conditional probability distribution p(X|Z) for the corresponding samples Z. The parameter values calculated by the decoder model can be used to evaluate the reconstruction error loss, such as the expected negative log-likelihood of the training data.

[0145] To obtain the full variational lower bound loss, it is possible to include a regularization term for the reconstruction error. The regularization term can evaluate the divergence, such as the Kullback-Leibler divergence, between the determined pairwise interaction type sequences (e.g., their variational posterior distributions) and the predicted interaction types obtained by applying the decoder model (e.g., the prior model).

[0146] The encoder and decoder models can be trained by minimizing this loss function with respect to the trainable parameters, such as the neural network weights of the encoder model and the decoder model. The trainable parameters can also include the parameters of the prior model for predicting the interaction types. For example, known gradient-based optimization techniques can be used to select the corresponding training instances or batches of training instances, where the parameters are iteratively refined during the process. As is known, such training can be heuristic and / or reach a local optimum.

[0147] For completeness, a detailed example of the training process is now given, understanding that a person skilled in the art can adapt it to the various alternatives discussed above. In particular, a decomposition of the log marginal likelihood is now given, from which an expression for can be derived. In particular, recalling that:

[0148]

[0149] Since the Kullback-Leibler divergence D KL (q||p) is always non-negative, it can follow that

[0150]

[0151] It can also be noted that the variational lower bound can be further decomposed into

[0152]

[0153] where the function

[0154]

[0155] can act as a reconstruction error when sampling from the conditional distribution given samples drawn from the approximate posterior, while

[0156] can be used to regularize the learning objective by penalizing variational distributions that deviate significantly from p(z).

[0157] Known stochastic variational inference methods can be used for the simultaneous learning of both: a generative model, such as a decoder and a prior, and probabilistic inference optimized with respect to the parameters of p(X|Z) and φ for .

[0158] The expression for

[0159] where is can be derived as follows. The reconstruction error term can be evaluated by substitution to obtain:

[0160]

[0161] To improve the efficiency of training, in various embodiments, stochastic estimation methods can be used to optimize the Monte Carlo estimator of the variational learning objective.

[0162] To make the training more amenable to gradient-based optimization, in various embodiments, a continuous relaxation of the learning objective can be used. For example, discrete interaction types in the learning objective can be replaced with a continuous approximation using, for example, the Gumbel-Softmax distribution. Specifically, the discrete, one-hot encoded random variable can be replaced with the continuous random variable where Δ k-1 is the standard simplex defined as Each variational categorical factor

[0163]

[0164] can be replaced with a continuous random variable, for example, approximated using the following continuous Gumbel-Softmax distribution.

[0165]

[0166] For example, for a Gumbel-Softmax distributed variable ζ with a location parameter and a temperature parameter λ, the following density can be used:

[0167]

[0168] The prior term p(Z) can be relaxed in a similar manner, for example:

[0169]

[0170] where n = [π 1 ... π K is a matrix with columns . Note that in the limit as λ d → 0, this latter expression may be equivalent to the expression given earlier for p(z t |z t-1 ). In the case of successive relaxation, the location parameter of the Gumbel-Softmax prior can be computed as the weighted average of , where the weights are given by the elements of .

[0171] In various embodiments, such as may be desirable for applying stochastic variational inference, a low-variance stochastic estimate of the gradient can be obtained by using the reparameterization trick, e.g., sampling from the distribution as follows

[0172]

[0173] where G k are i.i.d. samples and G k ∼ Gumbel(0,1).

[0174] Combining the above gives the following reconstruction error objective:

[0175]

[0176] where G t is a set of KN(N - 1) i.i.d. samples drawn from the standard Gumbel distribution.

[0177] The following regularization term can be obtained:

[0178]

[0179] where, for brevity, the latent variable Regarding φ and G t The deterministic dependence on is implicit. A similar relaxation process can be used for the prior, which depends on the representation of interacting objects by setting the following formula

[0180]

[0181] Figure 8 FIG. shows a block diagram of a computer-implemented method 800 for predicting object feature vectors of multiple interacting physical objects. Method 800 can correspond to the operation of Figure 1 system 100. However, this is not a limitation, as another system, apparatus, or device can also be used to perform method 800.

[0182] Method 800 can include: in an operation entitled "Access decoder, Observe", accessing 810 decoder model data and observation data, where the decoder model data represents parameters of a set of prediction models and propagation models for multiple interaction types, and the observation data represents a sequence of observed object feature vectors of multiple interacting objects.

[0183] Method 800 can include determining a sequence of predicted object feature vectors of multiple interacting objects, where the sequence of predicted object feature vectors of the multiple interacting objects extends the sequence of observed object feature vectors. To determine these sequences, method 800 can include: in an operation entitled "Obtain interaction type", for a sequence of object feature vectors to be predicted of a first interacting object, obtaining 820 a sequence of corresponding pairwise interaction types between the first interacting object and a second interacting object.

[0184] To determine these sequences, method 800 can further include: in an operation entitled "Predict object feature vector", predicting 830 the object feature vector of the first interacting object. To predict the object feature vector, method 800 can include: in an operation entitled "Select propagation model", selecting 832 a propagation model from the set of propagation models according to the pairwise interaction type of the object feature vector. To predict the object feature vector, method 800 can further include: in an operation entitled "Determine propagation data", determining 834 the propagation data from the second interacting object to the first interacting object by applying the selected propagation model based on the previous object feature vectors of the first and second interacting objects. To predict the object feature vector, method 800 can include: in an operation entitled "Predict object feature vector", predicting 836 the object feature vector using a prediction model based at least on the determined propagation data.

[0185] Figure 9FIG. 0 is a block diagram of a computer-implemented method 900 for classifying pairwise interactions between multiple physical objects into a set of multiple interaction types. Method 900 may correspond to the operation of Figure 2 system 200. However, this is not a limitation, as method 900 may also be performed using another system, apparatus, or device.

[0186] Method 900 may include, in an operation entitled "Access Encoder, Observe", accessing 910 encoder model data and observation data, where the encoder model data represents a set of propagation models and parameters of a classification model for a set of multiple interaction types, and the observation data represents a sequence of observed object feature vectors of multiple interacting objects. Method 900 may further include determining a sequence of pairwise interaction types for pairs of interacting objects corresponding to the observation data. This determination may include determining a sequence of hidden states of the interacting objects corresponding to the observation data.

[0187] To determine the interaction type between a first and a second interacting object, method 900 may include, in an operation entitled "Select Propagation Model", selecting 920 a propagation model from the set of propagation models according to the immediately preceding interaction type between the first and the second interacting objects. To determine the interaction type between a first and a second interacting object, method 900 may further include, in an operation entitled "Determine Propagation Data", determining 930 propagation data from the second interacting object to the first interacting object by applying the selected propagation model based on data of the first and second interacting objects. To determine the interaction type between a first and a second interacting object, method 900 may further include, in an operation entitled "Determine Hidden State", determining 940 the hidden state of the first interacting object based at least on the previous hidden state of the first interacting object and the determined propagation data. To determine the interaction type between a first and a second interacting object, method 900 may further include, in an operation entitled "Determine Interaction Type", determining 950 the interaction type using the classification model based at least on the hidden states of the first and second interacting objects.

[0188] Figure 10 FIG. 12 is a block diagram of a computer-implemented method 1000 for training an encoder model and a decoder model, where the encoder model is used to classify pairwise interactions between multiple physical objects into a set of multiple interaction types, and the decoder model is used to predict object feature vectors of multiple interacting physical objects. Method 1000 may correspond to the operation of Figure 3 system 300. However, this is not a limitation, as method 1000 may also be performed using another system, apparatus, or device.

[0189] Method 1000 may include: in an operation entitled "Access Decoder, Encoder, Training Data", accessing 1010 decoder model data representing decoder model parameters; encoder model data representing encoder model parameters; and training data representing a plurality of training instances.

[0190] Method 1000 may include optimizing the parameters of the decoder model and the encoder model. To optimize the parameters, Method 1000 may include: in an operation entitled "Select Training Instances", selecting 1020 at least one training instance from the training data. The training instance may include a sequence of observed object feature vectors of a plurality of interacting objects. To optimize the parameters, Method 1000 may further include: in an operation entitled "Determine Pairwise Interaction Types", based on the encoder model data, for example using Figure 9 Method 900, determining 1900 a sequence of pairwise interaction types for pairs of interacting objects. To optimize the parameters, Method 1000 may further include: in an operation entitled "Minimize Prediction Loss", based on the determined sequence of pairwise interaction types and the decoder model data, by predicting object feature vectors, for example by adapting Figure 8 Method 800, to minimize 1800 the loss of recovering the observed object feature vectors.

[0191] It will be appreciated that, generally speaking, Figure 8 Method 800, Figure 9 Method 900, and Figure 10 the operations of Method 1000 may be performed in any suitable order, such as any suitable order being continuously, simultaneously, or a combination thereof, which in applicable cases is subject to a specific order required, for example, by input / output relationships.

[0192] (One or more) methods may be implemented on a computer as a computer-implemented method, dedicated hardware, or a combination of both. As also illustrated in Figure 11 , instructions for a computer, such as executable code, may be stored, for example, in the form of a series of 1110 machine-readable physical markings and / or as a series of elements having different electrical (e.g., magnetic) or optical properties or values on a computer-readable medium 1100. The executable code may be stored in a transient or non-transient manner. Examples of computer-readable media include memory devices, optical storage devices, integrated circuits, servers, online software, etc. Figure 11A compact disc 1100 is shown. Alternatively, the computer-readable medium 1100 may include transient or non-transient data 1110 that represents: decoder model data, such as described herein for predicting object feature vectors of multiple interacting physical objects; and / or encoder model data, such as described herein for classifying pairwise interactions between multiple physical objects into a set of multiple interaction types.

[0193] Examples, embodiments, or optional features—whether or not indicated as non-limiting—should not be construed as limiting the invention as claimed.

[0194] It should be noted that the above-mentioned embodiments illustrate rather than limit the invention, and those skilled in the art will be able to design many alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The use of the verb “comprise” and its conjugations does not exclude the presence of elements or steps other than those stated in the claim. The article “a” or “an” preceding an element does not exclude the presence of a plurality of such elements. Expressions such as “at least one of...” when preceding a list of elements or groups mean selecting all elements or any subset of elements from the list or group. For example, the expression “at least one of A, B, and C” should be understood to include only A, only B, only C, both A and B, both A and C, both B and C, or all of A, B, and C. The invention may be implemented by means of hardware comprising several distinct elements, and by means of a suitably programmed computer. In a device claim enumerating several components, several of these components may be embodied by the same item of hardware. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage.

Claims

1. A system (100) for predicting object feature vectors of multiple interacting physical objects, wherein the object feature vector of each interacting object at discrete time points includes one or more of the position, velocity, and acceleration of each interacting object at respective discrete time points, the system comprising: A data interface (120) configured to access: Observation data (030) representing a sequence of observed object feature vectors of multiple interacting objects provided as input data to the system; Decoder model data (041) representing a set of propagation models and parameters of a prediction model for a set of multiple interaction types; A processor subsystem (140) configured to determine a sequence of predicted object feature vectors of multiple interacting objects that extends the sequence of observed object feature vectors by: For a sequence of object feature vectors of a first interacting object to be predicted at discrete time points, determining, by an encoder model for classifying pairwise interactions between multiple physical objects into a set of multiple interaction types, a sequence of corresponding pairwise interaction types at discrete time points between the first interacting object and a second interacting object based on the observed object feature vectors; Wherein the encoder model includes: a recurrent neural network for determining a hidden state of the observed object feature vectors; and a set of propagation models followed by a classification model for determining pairwise interaction types at the discrete time points between the first interacting object and the second interacting object based on the hidden state of the observed object feature vectors, Wherein the propagation model is a trainable graph neural network for message passing and information sharing between multiple interacting objects; Predicting, by a decoder model, an object feature vector of the first interacting object at the current time point, Wherein the decoder model includes: A recurrent neural network for determining a hidden state of the observed object feature vectors provided as input data, and a propagation model for determining propagation data based on the determined hidden state of the observed object feature vectors, And the prediction model for predicting an object feature vector of the first interacting particle based on the determined propagation data, Wherein the propagation data from the second interacting object to the first interacting object includes the contribution of the interaction between the second interacting particle and the first interacting particle to the change in the hidden state of the first interacting particle at the current time point relative to the hidden state of the first interacting particle at the previous time point, Wherein the prediction of the object feature vector of the first interacting particle at the current time point by the decoder model includes the steps of: Selecting a propagation model from the set of propagation models according to the pairwise interaction type in the sequence of corresponding pairwise interaction types between the first interacting object and the second interacting object determined by the encoder model at the current time point for the object feature vector; Determine propagation data from the second interacting object to the first interacting object by applying a selected propagation model to a determined hidden state based on previous object feature vectors of the first and second interacting objects; Predict an object feature vector using a prediction model based at least on the determined propagation data; and Determine a control signal based on the predicted object feature vector, provide the control signal to an actuator of an autonomous vehicle, and use the control signal to control the actuator.

2. The system (100) according to claim 1, wherein, The decoder model data further represents a set of transition probabilities that a previous pairwise interaction type is succeeded by a current pairwise interaction type, and the processor subsystem (140) is configured to predict a pairwise interaction type between the first and second interacting objects based on a previous pairwise interaction type between the first and second interacting objects by sampling the pairwise interaction types according to the set of transition probabilities.

3. The system (100) according to any one of the preceding claims 1-2, wherein, The processor subsystem (140) is configured to predict a pairwise interaction type between the first and second interacting objects by determining a respective representation of the interacting objects based on previous object feature vectors of the respective interacting objects and determining the pairwise interaction type based on the representations of the interacting objects.

4. The system (100) according to claim 3, wherein, The decoder model data further represents parameters of a recurrent model, and the processor subsystem (140) is configured to determine a representation of the interacting objects by sequentially applying the recurrent model to previous object feature vectors of the interacting objects.

5. The system (100) according to claim 3, wherein, The processor subsystem (140) is configured to determine a representation of the interacting objects from immediately preceding object feature vectors of the respective interacting agents.

6. The system (100) according to any one of the preceding claims 1-2, wherein, The processor subsystem (140) is configured to determine parameters of a probability distribution for predicting an object feature vector by applying a prediction model based at least on the determined propagation data, and sample an object feature vector from the probability distribution, thereby predicting the object feature vector using the prediction model.

7. A computer-implemented method (800) for predicting object feature vectors of multiple interacting physical objects, wherein the object feature vector of each interacting object at discrete time points includes one or more of the position, velocity, and acceleration of each interacting object at respective discrete time points, the method comprising: Access (810): Decoder model data representing a set of propagation models and parameters of a prediction model for a set of multiple interaction types; Observation data representing a sequence of observed object feature vectors of multiple interacting objects provided as input data to the system; Determine a sequence of predicted object feature vectors of multiple interacting objects that extends the sequence of observed object feature vectors by: For a sequence of object feature vectors of a first interacting object to be predicted at discrete time points, determine a sequence of corresponding pairwise interaction types at discrete time points between the first interacting object and a second interacting object based on the observed object feature vectors by an encoder model for classifying pairwise interactions between multiple physical objects into a set of multiple interaction types; The encoder model includes: a recurrent neural network for determining a hidden state of the observed object feature vector; and a set of propagation models followed by a classification model for determining a pairwise interaction type between the first interacting object and the second interacting object at the discrete time point based on the hidden state of the observed object feature vector, where the propagation model is a trainable graph neural network for message passing and information sharing between multiple interacting objects; Predicting (830) an object feature vector of the first interacting object at the current time point by the encoder model, where the decoder model includes: a recurrent neural network for determining a hidden state of the observed object feature vector provided as input data, a propagation model for determining propagation data based on the determined hidden state of the observed object feature vector, and a prediction model for predicting an object feature vector of the first interacting particle based on the determined propagation data, where the propagation data from the second interacting object to the first interacting object includes the contribution of the interaction between the second interacting particle and the first interacting particle to the change in the hidden state of the first interacting particle at the current time point relative to the hidden state of the first interacting particle at the previous time point, where the prediction of the object feature vector of the first interacting particle at the current time point by the decoder model includes the following steps: selecting (832) a propagation model from the set of propagation models according to the pairwise interaction type in the sequence of the corresponding pairwise interaction types between the first interacting object and the second interacting object determined by the encoder model at the current time point for the object feature vector; determining (834) propagation data from the second interacting object to the first interacting object by applying the selected propagation model based on the determined hidden states of the previous object feature vectors of the first and second interacting objects; predicting (836) an object feature vector using the prediction model based at least on the determined propagation data, and determining a control signal based on the predicted object feature vector, providing the control signal to an actuator of the autonomous vehicle, and using the control signal to control the actuator.

8. The method (800) according to claim 7, further comprising using the predicted target feature vector as training and / or test data to train another machine learning model.

9. A system (200) for classifying pairwise interactions between multiple physical objects into a set of multiple interaction types, where the system is configured according to the encoder model of claim 1, where the object feature vector of each interacting object includes one or more of the position, velocity, and acceleration of each interacting object, the system includes: a data interface (220) configured to access: observation data (030) representing a sequence of observed object feature vectors of multiple interacting objects; Decoder model data (042) representing a set of propagation models and parameters of a classification model for a set of multiple interaction types, where the propagation models are trainable graph neural networks for message passing and information sharing between multiple interacting objects; A processor subsystem (240) configured to determine a sequence of pairwise interaction types for pairs of interacting objects corresponding to observed data, the determination including determining a sequence of hidden states of the interacting objects corresponding to the observed data, the processor subsystem being configured to determine the interaction type between a first and a second interacting object by: Selecting a propagation model from the set of propagation models based on the immediately preceding interaction type between the first and second interacting objects; Determining propagation data from the second interacting object to the first interacting object by applying the selected propagation model based on data of the first and second interacting objects; Determining the hidden state of the first interacting object based at least on the previous hidden state of the first interacting object and the determined propagation data; Determining the interaction type using the classification model based at least on the hidden states of the first and second interacting objects; And Providing the determined interaction type to a behavior planning module of an autonomous vehicle and, via the behavior planning module, selecting a maneuver that minimizes the risk of collision and / or ensures a socially compliant interaction with a pedestrian, and controlling an actuator of the autonomous vehicle based on the selected maneuver.

10. The system (200) according to claim 9, wherein determining the sequence of pairwise interaction types further includes determining a sequence of additional hidden states of the interacting objects, the processor subsystem (240) being configured to determine the additional hidden state of the first interacting object based at least on the immediately succeeding additional hidden state of the first interacting object and additionally applying the classification model based on the additional hidden states of the first and second objects.

11. A computer-implemented method (900) of classifying pairwise interactions between multiple physical objects into a set of multiple interaction types, where the object feature vector of each interacting object includes one or more of the position, velocity, and acceleration of each interacting object, the method being configured to determine the pairwise interaction type at discrete time points between a first and a second interacting object based on the observed object feature vector of method claim 7, the method including: Accessing (910): Encoder model data representing a set of propagation models and parameters of a classification model for a set of multiple interaction types; Observed data representing a sequence of observed object feature vectors of multiple interacting objects; Determining a sequence of pairwise interaction types for pairs of interacting objects corresponding to the observed data, the determination including determining a sequence of hidden states of the interacting objects corresponding to the observed data, wherein the interaction type between a first and a second interacting object is determined by: Select (920) a propagation model from a set of propagation models based on the immediately preceding interaction type between a first and a second interacting object, where the propagation model is a trainable graph neural network for message passing and information sharing among multiple interacting objects; Determine (930) propagation data from the second interacting object to the first interacting object by applying the selected propagation model based on data of the first and second interacting objects; Determine (940) a hidden state of the first interacting object based at least on a previous hidden state of the first interacting object and the determined propagation data; Determine (950) an interaction type using a classification model based at least on the hidden states of the first and second interacting objects, and Provide the determined interaction type to a behavior planning module of an autonomous vehicle and, through the behavior planning module, select a maneuver that minimizes a collision risk and / or ensures a socially compliant interaction with a pedestrian, and control an actuator of the autonomous vehicle based on the selected maneuver.

12. The method (900) according to claim 11, further comprising determining a control signal using the determined sequence of pairwise interaction types, the control signal for implementing an action in an environment from which the observation data is derived.

13. A system (300) for training an encoder model and a decoder model, the encoder model for classifying pairwise interactions between multiple physical objects into a set of multiple interaction types, the decoder model for predicting object feature vectors of multiple interacting physical objects, the system comprising: A data interface (320) configured to access decoder model data (041) representing decoder model parameters; encoder model data (042) representing encoder model parameters; and training data (050) representing multiple training instances; A processor subsystem (340) configured to optimize the parameters of the decoder model and the encoder model by: Selecting at least one training instance from the training data, the training instance including a sequence of observed object feature vectors of multiple interacting objects; Determining, based on the encoder model data, a sequence of pairwise interaction types of pairs of interacting objects according to the method of claim 11; Minimizing a loss of recovering the observed object feature vectors by predicting an object feature vector according to the method of claim 7 based on the determined sequence of pairwise interaction types and the decoder model data.

14. The system (300) according to claim 13, wherein, The loss includes a difference between the determined sequence of pairwise interaction types and a predicted interaction type of the decoder model, and wherein the parameters of the decoder model optimized by the processor subsystem (340) include parameters for predicting interaction types.

15. A computer-implemented method (1000) for training an encoder model and a decoder model, the encoder model for classifying pairwise interactions between multiple physical objects into a set of multiple interaction types, the decoder model for predicting object feature vectors of multiple interacting physical objects, the method comprising: Accessing (1010): Decoder model data representing decoder model parameters; Encoder model data representing encoder model parameters; Training data representing a plurality of training instances; Optimizing the parameters of the decoder model and the encoder model by: Selecting (1020) at least one training instance from the training data, the training instance including an observed object feature vector sequence of a plurality of interacting objects; Based on the encoder model data, determining (1900) a sequence of pairwise interaction types of interacting object pairs according to the corresponding steps performed by the encoder model in the method of claim 11; Minimizing (1800) the loss of recovering the observed object feature vectors by predicting object feature vectors based on the determined sequence of pairwise interaction types and the decoder model data according to the method of claim 7.

16. A computer-readable medium (1100) comprising transient or non-transient data (1110), the transient or non-transient data (1110) representing Instructions which, when executed by a processor system, cause the processor system to perform a computer-implemented method according to any one of claims 7, 11, and 15; and / or Decoder model data for predicting object feature vectors of a plurality of interacting physical objects according to the method of claim 7; and / or Encoder model data for classifying pairwise interactions between a plurality of physical objects into a set of interaction types according to the method of claim 11.