Method and system for predicting the behavior of dynamic objects
By encoding the interactions of different object categories in the sensing environment and generating classified interaction representations, the problem in the prior art is solved that it is difficult to accurately predict the interaction behavior of multiple dynamic objects in complex environments, and efficient and accurate prediction of dynamic object behavior is achieved.
Patent Information
- Application Number
- CN202180072393.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-11-13
- Filing Date
- 2021-06-30
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2041-06-30
AI Technical Summary
The prior art is difficult to accurately predict the interaction behavior between multiple dynamic objects in a complex sensing environment, especially when the number of dynamic objects is large or variable, resulting in prediction accuracy and computing efficiency problems.
By grouping dynamic and static objects in the sensing environment into different categories and encoding object interactions in different categories, a unified classification interaction representation is generated to input a behavior predictor to generate future behavior predictions for dynamic objects.
It realizes accurate prediction of the behavior of target dynamic objects in large or complex environments, reduces the demand for computing time and memory resources, and improves the scalability and efficiency of the system.
Smart Images

Figure CN116348938B_ABST
Abstract
Description
[0001] This application claims the benefit of priority to U.S. patent application No. 17 / 097,840, filed on November 13, 2020, entitled “METHODS AND SYSTEMS FOR PREDICTING DYNAMIC OBJECT BEHAVIOR,” the contents of which are incorporated herein by reference as if reproduced in full. Technical Field
[0002] Examples of the present invention relate to methods and systems for generating dynamic object behavior predictions, including methods and systems for generating dynamic object behavior predictions in a sensed environment of an autonomous vehicle. Background Art
[0003] Behavior prediction units based on machine learning are used in many practical systems, such as autonomous driving systems or advanced driver assistance systems for vehicles. For autonomous driving systems or advanced driver assistance systems, it is important to be able to generate accurate predictions of how dynamic objects will behave (e.g., how pedestrians will move when crossing a street) in a sensing environment (i.e., the environment around the vehicle sensed by the vehicle's sensors) so that the desired path of the vehicle can be planned (e.g., to avoid hitting pedestrians crossing the street), or feedback can be provided to the driver of the vehicle. One factor in generating useful and accurate predictions of the behavior of dynamic objects in the sensing environment is the ability to encode the interactions between the dynamic objects in the sensing environment in order to define how the behavior of the dynamic objects affects each other. Encoding the interactions between dynamic objects in the sensing environment, and specifically encoding the interactions between dynamic objects in the sensing environment, can be challenging due to the complex spatiotemporal nature of the problem.
[0004] Many existing interaction modeling methods usually encode the interaction between a single dynamic object and other objects (also referred to as agents in some contexts) in a sensing environment. Encoding the interaction between the corresponding single dynamic object and other dynamic objects in the sensing environment, and training the model of the prediction unit based on machine learning, to use the encoded data representing the interaction between the corresponding single dynamic object and other dynamic objects to generate an accurate prediction of how the corresponding single dynamic object will behave, may be effective in some contexts. However, this method is usually not scalable to larger or more complex environments, in which there are a large number of dynamic objects or the number of dynamic objects is highly variable (for example, in a busy intersection). Many existing models for accurately predicting how a single dynamic object will behave based on the interaction between a single dynamic object and other dynamic objects in the sensing environment need to estimate the number of other objects (including dynamic objects and static objects) that interact with the single dynamic object, which is not always possible. Another existing interaction modeling method is to use an embedded form to create a fixed-size representation of the interaction between a single dynamic object and other dynamic objects in the environment so that it can be processed by a machine learning model. However, this method of modeling the interaction between dynamic objects will inevitably lose some information about the interaction between dynamic objects. As the number of dynamic objects interacting with each other increases, embedding may lead to unacceptable information loss.
[0005] It would therefore be useful to provide a machine learning based technique for encoding interactions between dynamic objects and for generating accurate predictions of how the dynamic objects will behave in a sensed environment. Summary of the invention
[0006] In various examples, the present invention describes methods and systems for encoding interactions between different classes of dynamic objects and static objects in a sensing environment (i.e., the environment around a vehicle as sensed by various sensors of the vehicle) into a single unified representation of the interactions in the sensing environment. The single unified representation is provided as an input to a behavior predictor to generate predictions of future behavior of one or more dynamic objects of interest (i.e., how the target dynamic object will behave in the sensing environment). The predicted future behavior of the target dynamics in the sensing environment can be further provided as input to a motion planner of an autonomous driving system (ADS) or advanced-assistive driving system (ADAS) of the vehicle to perform motion planning that takes into account the predicted behavior of the target dynamic object in the sensing environment.
[0007] In the example described herein, based on some shared features (e.g., object classes), the dynamic objects and static objects of the sensing environment are grouped into different object categories, and the interactions between the dynamic objects and static objects of different categories are encoded. In the sensing environment, the encoding of the interactions between the dynamic objects and static objects of different categories is called classification encoding. Compared with the traditional modeling for the interactions between the dynamic objects in the sensing environment, the classification encoding method of the present invention is less complex and more computationally efficient, and the traditional modeling requires a complex model to encode a variable number of single dynamic objects in the sensing environment. The object category can be predefined and can be defined by any suitable feature according to the desired application (e.g., the category can be defined by different object categories, where the class of interest depends on the environment of interest).
[0008] Since the categories and number of categories are predefined, the computation time and memory resources required to encode the interactions between dynamic and static object categories can be fixed. Examples of the present invention can be scalable to enable prediction of the behavior of target dynamic objects (e.g., pedestrians or other vehicles) in large or complex environments with little or no increase in the computation time or memory resources required.
[0009] Examples of the present invention may be used in a variety of applications, including autonomous driving, drone piloting, anomaly detection, assisted robotics, or smart traffic management, among others.
[0010] In some exemplary aspects, the present invention describes a method for predicting the behavior of a dynamic object of interest in an environment of a vehicle. The method includes: receiving a plurality of time series feature data, each time series feature data representing a corresponding feature of a plurality of objects in the environment at a plurality of time steps, the plurality of objects including a dynamic object of interest. The method also includes: classifying each time series feature data into one of a plurality of defined object categories to obtain a classification data set for each defined object category, each classification data set including one or more time series feature data, the one or more time series feature data representing one or more objects belonging to the corresponding defined object category; encoding each classification data set into a corresponding classification representation, each classification representation representing a temporal change of a feature in the corresponding defined object category. The method also includes: combining the classification representations into a single shared representation, generating a classification interaction representation, wherein the classification interaction representation is a weighted representation of the effect of the temporal change in each defined object category on the final time step of the shared representation; generating prediction data representing the predicted future behavior of the dynamic object of interest based on the classification interaction representation, data representing the dynamics of the plurality of objects, and data representing the state of the vehicle.
[0011] In any of the above examples, combining the categorical representations may include concatenating the categorical representations according to time steps to generate the single shared representation.
[0012] In any of the above examples, encoding each classification data set may include: for a given classification data set belonging to a given object category, providing the one or more time series feature data to a trained neural network to generate a time series feature vector as a classification representation of the given object category.
[0013] In any of the above examples, the trained neural network can be a recursive neural network, a convolutional neural network, or a combined recursive and convolutional neural network.
[0014] In any of the above examples, at least one defined object class can be specific to the dynamic object of interest.
[0015] In any of the above examples, the method may include: receiving time series sensor data generated by a sensor; preprocessing the time series sensor data into a time series feature data, wherein the time series feature data is included in the received multiple time series feature data.
[0016] In any of the above examples, the vehicle is an autonomous vehicle, and the method may include: providing the predicted data representing the predicted future behavior of the dynamic object of interest to a motion planning subsystem of the autonomous vehicle to generate a planned path for the autonomous vehicle.
[0017] In some exemplary aspects, the present invention describes a computing system for predicting the behavior of a dynamic object of interest in an environment of a vehicle. The computing system includes: a processor system for executing instructions to cause an object behavior prediction subsystem of the computing system to perform the following operations: receiving a plurality of time series feature data, each of which represents a corresponding feature of a plurality of objects in the environment at a plurality of time steps, the plurality of objects including a dynamic object of interest; classifying each of the time series feature data into one of a plurality of defined object categories to obtain a classification data set for each defined object category, each classification data set including one or more time series feature data, the one or more time series feature data representing one or more objects belonging to the corresponding defined object category; encoding each classification data set into a corresponding classification representation, each classification representation representing a temporal change of a feature in the corresponding defined object category; combining the classification representations into a single shared representation; generating a classification interaction representation based on the single shared representation, the classification interaction representation being a weighted representation representing the effect of a temporal change in each defined object category on a final time step of the single shared representation; generating prediction data based on the classification interaction representation, data representing the dynamics of the plurality of objects, and data representing the state of the vehicle, the prediction data representing a predicted future behavior of the dynamic object of interest.
[0018] In some examples, the processing system may be used to execute instructions to perform any of the methods described above.
[0019] In some exemplary aspects, the present disclosure describes a computer-readable medium including computer-executable instructions to implement an object behavior prediction subsystem to predict the behavior of dynamic objects of interest in a sensing environment of a vehicle. The instructions, when executed by a processing system of a computing system, cause the computing system to perform the following operations: receive a plurality of time series feature data, each of which represents corresponding features of a plurality of objects in the environment at a plurality of time steps, the plurality of objects including a dynamic object of interest; classify each of the time series feature data into one of a plurality of defined object categories to obtain a classification data set for each defined object category, each classification data set comprising one or more time series feature data, the one or more time series feature data representing one or more objects belonging to the corresponding defined object category; encode each classification data set into a corresponding classification representation, each classification representation representing a temporal change of a feature in the corresponding defined object category; combine the classification representations into a single shared representation; generate a classification interaction representation based on the single shared representation, the classification interaction representation being a weighted representation representing the effect of the temporal change in each defined object category on a final time step of the single shared representation; generate prediction data based on the classification interaction representation, data representing the dynamics of the plurality of objects, and data representing the state of the vehicle, the prediction data representing the predicted future behavior of the dynamic object of interest.
[0020] In some examples, the instructions may cause the computing system to perform any of the methods described above. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Reference will now be made, by way of example, to the accompanying drawings which show exemplary embodiments of the present application, in which:
[0022] Figure 1 is a schematic diagram of an exemplary environment in which an autonomous driving vehicle operates according to some embodiments of the present invention;
[0023] Figure 2 is a block diagram of some exemplary components of an autonomous driving vehicle provided by some embodiments of the present invention;
[0024] Figure 3 is a block diagram of an exemplary subsystem for encoding classification interactions provided by some embodiments of the present invention;
[0025] Figure 4 is a block diagram of an exemplary subsystem for generating object behavior predictions provided by some embodiments of the present invention;
[0026] Figure 5 is a flowchart of an exemplary method for predicting the behavior of a dynamic object provided by some embodiments of the present invention.
[0027] Similar reference numerals may be used in different drawings to represent similar components. DETAILED DESCRIPTION
[0028] The technical solution of the present invention is described below with reference to the accompanying drawings.
[0029] Some examples of the present invention are described in the context of autonomous vehicles. Although the examples described herein may refer to automobiles as autonomous vehicles, the teachings of the present invention may be implemented in other forms of autonomous or semi-autonomous vehicles, including, for example, trams, subways, trucks, buses, surface and submersible vessels and ships, aircraft, drones (also known as unmanned aerial vehicles (UAVs)), warehouse equipment, manufacturing facility equipment, construction equipment, farm equipment, autonomous service robots such as vacuum cleaners and lawn mowers, and other robotic devices. Autonomous vehicles may include vehicles that do not carry passengers as well as vehicles that carry passengers.
[0030] Examples of the present invention may be applicable to applications other than autonomous vehicles. For example, examples of the present invention may be applied in the following contexts: automated driving systems (ADS) for autonomous vehicles, advanced driver-assistance systems (ADAS) for semi-autonomous vehicles, assistive robots (e.g., enabling assistive robots to predict dynamic object behaviors that stand out from the expected or predicted behaviors of other dynamic objects), anomaly detection (e.g., being able to predict dynamic object behaviors that stand out from the expected or predicted behaviors of other dynamic objects), or intelligent traffic management (e.g., being able to predict traffic behaviors at intersections), etc. The exemplary methods and systems described herein may be applicable to any context in which it is useful or desirable to accurately predict the behavior of one or more dynamic objects of interest (also referred to as target dynamic objects individually and collectively as target dynamic objects) in a sensing environment (also referred to as a sensing scene). A dynamic object is any object in an environment whose state (e.g., position) changes over a period of interest (e.g., more than 10 minutes). A static object is any object in an environment whose state (e.g., position) changes little over a period of interest (e.g., the position change is within a predefined margin, such as 1 meter).
[0031] The methods and systems described in the examples of this article can be used to train an object behavior predictor, which can be deployed to the ADS of an autonomous vehicle or the ADAS of a vehicle after being trained. In the disclosed method and system, feature data representing sensed objects (including static objects and dynamic objects, also referred to as environmental elements or scene elements) and sensed features in the environment are grouped together according to defined object categories. Interactions between object categories are encoded, rather than interactions between individual objects and features. The encoded representation of the classified interactions is used as input to predict the behavior of dynamic objects in the environment. Therefore, the disclosed method and system provide a technical effect, namely, the behavior of dynamic objects can be predicted without having to process the interactions of each individual object in the environment. Another technical effect is that a trained object behavior predictor can be implemented using a less complex machine learning architecture, using less computing time and / or using less memory resources.
[0032] In the context of autonomous or assisted driving, accurate behavior prediction of objects of interest is very important, at least for ensuring the safety of vehicle occupants and other people in the environment. A challenge in designing and implementing behavior prediction models for predicting the behavior of objects of interest in a sensing environment is the problem of how to encode the interactions between different objects in the sensing environment. Accurately representing information about how dynamic objects in the environment interact with each other is very important, because such interactions may affect how the dynamic objects move in the future.
[0033] In the present invention, interaction refers to the behavior of an object of interest due to the presence of other objects in the environment. As will be further discussed below, the interaction can be weighted based on the degree of influence of another object on the behavior of the object of interest. For example, if the behavior of the object of interest is the same, a low (or zero) interaction (which can be assigned a low weight) can be determined, regardless of whether the other objects are in the sensed environment, and if the object of interest changes behavior sharply due to the presence of another object in the environment, a high interaction (which can be assigned a high weight) can be determined. Interaction can be between a dynamic object of interest and one or more other dynamic objects of interest (e.g., a pedestrian walking away from another pedestrian walking), or between a dynamic object of interest and one or more other static objects of interest (e.g., a pedestrian walking away from a fire hydrant in a sidewalk). In some examples, interaction can also be that the object of interest changes behavior due to the semantic meaning represented by another object in the environment (i.e., not just due to the presence of another object) (e.g., if a traffic light turns red, a pedestrian walking may stop walking).
[0034] Some existing technical solutions model the interaction between the corresponding single object and other objects by considering the potential interactions between all single dynamic objects (e.g., road users, such as single vehicles, single pedestrians, and / or single cyclists). Such solutions usually rely on the trajectory and position information of all dynamic objects in the environment, and ignore the interaction between dynamic objects and static objects (e.g., road markings, signs, traffic signals, etc.). Solutions based on modeling the interaction between the object of interest and each single object are usually not scalable (or have limited scalability) and are difficult or impossible to adapt to environments with highly variable numbers of objects. Another disadvantage of the solution based on modeling the interaction with a single object is that when modeling the interaction between a large number of objects (when there are 20 or more dynamic objects in the environment), modeling the interaction between all single objects may be difficult (e.g., requiring an overly complex machine learning model) and / or computationally impossible (e.g., requiring too much computing time, processing power, and / or memory resources). This is because methods based on single object interactions require an estimate of the number of interacting objects, and the computational complexity (and therefore the required memory resources and computing time) increases with the number of interactive elements. Furthermore, representations of a variable number of individual objects and interactions need to be converted to fixed-size representations in order to make them usable in machine learning. This conversion may be difficult to define or may result in significant information loss.
[0035] Existing solutions typically address these challenges in one of two ways. One approach is to define limits on the number of individual objects whose interactions with other objects are to be modeled. However, it can be difficult to define appropriate limits on the number of individual objects whose interactions with other objects are to be modeled in highly variable environments (e.g., traffic densities are mixed and vary widely). Another approach is to use an embedding mechanism to convert a variable-size representation (e.g., representing a variable number of dynamic objects in an environment) into a fixed-size representation. However, such embedding mechanisms can result in significant information loss when encoding fixed-size representations of a large number of dynamic objects.
[0036] In the present invention, exemplary methods and systems are described for modeling interactions by encoding interactions between objects (or object features) of defined classes rather than modeling interactions between corresponding individual objects and other objects in an environment. This may help reduce the computational cost and model complexity for encoding interactions and generating dynamic object behavior predictions (i.e., predictions of the behavior of dynamic objects of interest in a sensing environment) compared to some existing schemes (e.g., as described above).
[0037] Typically, predictions of dynamic object behavior can be generated by a neural network. To facilitate understanding, some concepts related to neural networks and some related terms that may be related to an exemplary neural network that predicts dynamic object behavior of an object of interest disclosed herein are described below.
[0038] Neural networks are composed of neurons. A neuron is a computing unit that uses x s The output of the calculation unit can be:
[0039]
[0040] Where s = 1, 2, ..., n, n is a natural number greater than 1, W s is x s The weight (or coefficient), b is the offset (i.e. bias) of the neuron, and f is the activation function of the neuron, which is used to introduce nonlinear features into the neural network to convert the input of the neuron into output. The output of the activation function can be used as the input of the neurons in the subsequent convolutional layer of the neural network. For example, the activation function can be a sigmoid function. The neural network is composed of multiple single neurons mentioned above. In other words, the output of one neuron can be the input of another neuron. The input of each neuron can be associated with the local receptive area of the previous layer to extract the features of the local receptive area. The local receptive area can be an area composed of several neurons.
[0041] A deep neural network (DNN) is also called a multilayer neural network and can be understood as a neural network that includes a first layer (usually called an input layer), multiple hidden layers, and a final layer (usually called an output layer). There is no special measure for "multiple" in this article. When there is a full connection between two adjacent layers of a neural network, the layer is considered to be a fully connected layer. Specifically, for two adjacent layers to be fully connected (e.g., the i-th layer and the (i+1)-th layer), each neuron in the i-th layer must be connected to each neuron in the (i+1)-th layer.
[0042] The processing of each layer of DNN can be described as follows. In short, the operation of each layer is represented by the following linear relationship expression: in, is the input tensor, is the output tensor, is the bias tensor, W is the weight (also called coefficient), and α(.) is the activation function. In each layer, the input tensor Perform the operation to obtain the output tensor
[0043] Since there are a large number of layers in a DNN, there are also a large number of weights W and bias vectors The definition of these parameters in DNN is as follows, taking weight W as an example. In this example, in a three-layer DNN (i.e., a DNN with three hidden layers), the linear weight from the fourth neuron in the second layer to the second neuron in the third layer is expressed as The superscript 3 indicates the layer of weight W (i.e., the third layer (or layer 3) in this example), and the subscript indicates that the output is at layer 3 index 2 (i.e., the second neuron in the third layer), and the input is at layer 2 index 4 (i.e., the fourth neuron in the second layer). In general, the weight from the kth neuron in the (L–1)th layer to the jth neuron in the Lth layer can be expressed as It should be noted that the input layer has no W parameter.
[0044] In a DNN, a larger number of hidden layers can make the DNN better at modeling complex situations (e.g., real-world situations). In theory, a DNN with more parameters is more complex, has a larger capacity (which may refer to the ability of the learning model to adapt to a variety of possible scenarios), and indicates that the DNN can complete more complex learning tasks. The training of a DNN is the process of learning the weight matrix. The purpose of training is to obtain a trained weight matrix that consists of the learned weights W of all layers of the DNN.
[0045] A convolutional neural network (CNN) is a DNN with a convolutional structure. A CNN includes a feature extractor consisting of a convolution layer and a subsampling layer. The feature extractor can be viewed as a filter. The convolution process can be viewed as performing convolution on a two-dimensional (2D) input image or convolution feature map using a trainable filter.
[0046] A convolutional layer is a layer of neurons where convolution processing is performed on the input in a CNN. In a convolutional layer, a neuron may be connected only to a subset of neurons in an adjacent layer (i.e., not all neurons). That is, a convolutional layer is generally not a fully connected layer. A convolutional layer generally includes several feature maps, each of which may consist of a number of neurons arranged in a rectangular shape. Neurons on the same feature map share weights. Shared weights may be collectively referred to as convolution kernels. Typically, a convolution kernel is a 2D weight matrix. It should be understood that the convolution kernel may be independent of the manner and location of image information extraction. One principle hidden behind the convolutional layer is that the statistical information of one part of the image is the same as the statistical information of another part of the image. This means that image information learned from one part of the image may also be applicable to another part of the image. Multiple convolution kernels may be used in the same convolutional layer to extract different image information. Generally, the more convolution kernels there are, the richer the image information reflected by the convolution operation.
[0047] The convolution kernel can be initialized as a 2D matrix of random values. During the training process of the CNN, the weights of the convolution kernel are learned. One advantage of using convolution kernels to share weights between neurons in the same feature map is that the connections between the convolution layers of the CNN are reduced (compared to the fully connected layers) and the risk of overfitting is reduced.
[0048] In the process of training a DNN, the predicted value output by the DNN can be compared with the expected target value (e.g., the ground truth). The weight vector of each layer of the DNN (which is a vector containing the weight W of a given layer) is updated based on the difference between the predicted value and the expected target value. For example, if the predicted value output by the DNN is too high, the weight vector of each layer can be adjusted to reduce the predicted value. This comparison and adjustment can be performed iteratively until the convergence condition is met (e.g., a predefined maximum number of iterations has been performed, or the predicted value output by the DNN fully converges with the expected target value). A loss function or objective function is defined as a method of quantitatively representing the closeness of the predicted value to the target value. The objective function represents the amount to be optimized (e.g., minimized or maximized) so that the predicted value is as close to the target value as possible. The loss function more specifically represents the difference between the predicted value and the target value, and the goal of training a DNN is to minimize the loss function.
[0049] Backpropagation is an algorithm for training DNNs. Backpropagation is used to adjust (also called update) the values of parameters (e.g., weights) in a DNN so that the error (or loss) in the output becomes smaller. For example, a defined loss function is calculated based on the forward propagation from the input to the output of the DNN. Backpropagation calculates the gradient of the loss function with respect to the parameters of the DNN, and updates the parameters using a gradient algorithm (e.g., gradient descent) to reduce the loss function. Backpropagation is performed iteratively so that the loss function converges or is minimized.
[0050] A recurrent neural network (RNN) is a neural network (typically a DNN) that is typically used to process sequential data where some relationship is expected in the order of the data (e.g., in a temporal dataset containing data on a sequence of time steps, or in a text dataset where information is encoded in the order of words in the text). For example, to predict a word in a sentence, the previous word is usually required because the likelihood of predicting a word depends on the previous word in the sentence. In an RNN, the calculation of the current predicted output of a sequence is also related to the previous output. Conceptually, an RNN can be understood as "memorizing" previous information and applying the previous information to the calculation of the current predicted output. In terms of neural network layers, the nodes between hidden layers are connected so that the input to a given hidden layer includes the output of the previous layer and also includes the output generated by the hidden layer from the previous input. This can be referred to as parameter sharing because parameters (e.g., layer weights) are shared between multiple inputs to the layer. Therefore, the same input to the hidden layer provided at different sequential positions in the sequence data may cause the hidden layer to generate different outputs based on the previous input in the sequence. RNNs can be designed to process sequence data of any desired length.
[0051] The training of RNN can be similar to the training of traditional CNN or DNN. The error back propagation algorithm can also be used. In order to take into account the parameter sharing in RNN, in the gradient descent algorithm, the output of each gradient step is calculated based on the weight of the current step and the weights of several previous steps. The learning algorithm used to train RNN can be called backpropagation through time (BPTT).
[0052] To help understand the present invention, an example of an autonomous vehicle in an environment is now discussed. It should be understood that the present invention is not intended to be limited to implementation in the context of an autonomous vehicle.
[0053] Figure 1 1 is a schematic diagram showing an exemplary environment 100 in which a vehicle 105 operates. Examples of the present invention may be implemented in a vehicle 105, for example, to implement autonomous or semi-autonomous driving. The environment 100 includes a communication system 200 that communicates with the vehicle 105. The vehicle 105 includes a vehicle control system 115. The vehicle control system 115 is connected to a drive control system and a mechanical system of the vehicle 105, as described below with reference to Figure 2 Further described. In various examples, the vehicle control system 115 can enable the vehicle 105 to operate in one or more of fully autonomous, semi-autonomous, or fully user-controlled modes.
[0054] The vehicle 105 may include sensors, shown herein as a plurality of environmental sensors 110 that collect information about the external environment 100 around the vehicle 105 and generate sensor data indicative of such information, and a plurality of vehicle sensors 111 that collect information about operating conditions of the vehicle 105 and generate vehicle data indicative of such information. There may be different types of environmental sensors 110 to collect different types of information about the environment 100, as discussed further below. In an exemplary embodiment, the environmental sensors 110 are mounted and positioned at the front, rear, left, and right sides of the vehicle 105 to collect information about the external environment 100 located at the front, rear, left, and right sides of the vehicle 105. A single unit of the environmental sensor 110 may be mounted or otherwise positioned on the vehicle 105 to have different overlapping or non-overlapping fields of view (FOV) or coverage areas to capture data about the environment 100 around the vehicle 105. The vehicle control system 115 receives the sensor data, which indicates the collected information collected by the environmental sensors 110 about the external environment 100 of the vehicle 105.
[0055] The vehicle sensors 111 provide vehicle data indicative of collected information about the operating conditions of the vehicle 105 to the vehicle control system 115 in real time or near real time. For example, the vehicle control system 115 may use the vehicle data indicative of information about the operating conditions of the vehicle 105 provided by one or more vehicle sensors 111 to determine factors such as the linear speed of the vehicle 105, the angular velocity of the vehicle 105, the acceleration of the vehicle 105, the engine RPM of the vehicle 105, the transmission gear and tire grip of the vehicle 105, etc.
[0056] The vehicle control system 115 may include or be connected to one or more wireless transceivers 130, which enable the vehicle control system 115 to communicate with the communication system 200. For example, the one or more wireless transceivers 130 may include one or more cellular (RF) transceivers for communicating with multiple different wireless access networks (e.g., cellular networks) using different wireless data communication protocols and standards. The one or more wireless transceivers 130 can communicate with any of a plurality of fixed transceiver base stations of a wireless wide area network (WAN) 210 (e.g., a cellular network) within its geographic coverage area. The one or more wireless transceivers 130 can send and receive signals through the wireless WAN 210. The one or more wireless transceivers 130 may include a multi-band cellular transceiver that supports multiple radio frequency bands. The vehicle control system 115 can use the wireless WAN 210 to access a server 240 (e.g., a driving assistance server) through one or more communication networks 220 (e.g., the Internet). The server 240 can be implemented as one or more server modules in a data center and is typically located behind a firewall 230. The server 240 may be connected to network resources 250 , such as a supplemental data source that may provide information used by the vehicle control system 115 .
[0057] The one or more wireless transceivers 130 may also include a wireless local area network (WLAN) transceiver for communicating with a WLAN (not shown) via a WLAN access point (AP). The WLAN may include a wireless local area network (WLAN) that complies with the IEEE 802.11x standard (sometimes referred to as a wireless local area network). ) or other communication protocols. One or more wireless transceivers 130 may also include a short-range wireless transceiver, such as The wireless transceiver 130 may also include other short-range wireless transceivers, including but not limited to near field communication (NFC), IEEE802.15.3a (also known as ultra wideband (UWB)), Z-Wave, ZigBee, ANT / ANT+, or infrared (e.g., Infrared Data Association (IrDA) communication).
[0058] The communication system 100 also includes a satellite network 260 including a plurality of satellites. The vehicle control system 115 can use the signals of the plurality of satellites in the satellite network 260 to determine its position. The satellite network 260 typically includes a plurality of satellites that are part of at least one global navigation satellite system (GNSS). At least one GNSS provides automatic geospatial positioning with global coverage. For example, the satellite network 260 can be a collection of GNSS satellites. Exemplary GNSS include the U.S. NAVSTAR global positioning system (GPS) or the Russian global orbiting navigation satellite system (GLONASS). Other satellite navigation systems that have been deployed or are under development include the European Union's Galileo positioning system, China's BeiDou navigation satellite system (BDS), the Indian regional satellite navigation system, and the Japanese satellite navigation system.
[0059] Figure 2 Selected components of a vehicle 105 provided by some examples described herein are shown. The vehicle 105 includes a vehicle control system 115 connected to a drive control system 150 and an electromechanical system 190. The vehicle control system 115 is also connected to receive data from environmental sensors 110 and vehicle sensors 111.
[0060] The environmental sensors 110 may include, for example, one or more camera units 112, one or more light detection and ranging (LIDAR) units 114, and one or more radar units such as synthetic aperture radar (SAR) units 116, among other possibilities. Each type of sensor unit 112, 114, 116 may collect correspondingly different information about the environment 100 external to the vehicle 105, and may provide sensor data to the vehicle control system 115 in different formats, respectively. For example, the camera unit 112 may provide camera data representing a digital image, the LIDAR unit 114 may provide a two-dimensional or three-dimensional point cloud, and the SAR unit may provide radar data representing a radar image.
[0061] The vehicle sensors 111 may include, for example, an inertial measurement unit (IMU) 118 that senses specific forces and angular rates of the vehicle 105 and provides data about the direction of the vehicle based on the sensed specific forces and angular rates. The vehicle sensors 111 may also include an electronic compass 119 and other vehicle sensors 120, such as a speedometer, a tachometer, a wheel traction sensor, a transmission gear sensor, a throttle and brake position sensor, and a steering angle sensor.
[0062] The vehicle control system 115 may also collect information about the location of the vehicle 105 using signals received from the satellite network 260 via the satellite receiver 132 and generate positioning data representing the location of the vehicle 105 .
[0063] The vehicle 105 also includes various structural elements, such as a frame, doors, panels, seats, windows, mirrors, etc., which are known in the art but have been omitted from the present invention to avoid obscuring the teachings of the present invention. The vehicle control system 115 includes a processor system 102, which is connected to a plurality of components via a communication bus (not shown) that provides a communication path between the components and the processor system 102. The processor system 102 is connected to a drive control system 150, a random access memory (RAM) 122, a read-only memory (ROM) 124, a persistent (non-volatile) memory 126 such as a flash erasable programmable read-only memory (EPROM) (flash memory), one or more wireless transceivers 130, a satellite receiver 132, and one or more input / output (I / O) devices 134 (e.g., a touch screen, a speaker, a microphone, a display screen, a mechanical button, etc.). The processor system 102 may include one or more processing units, including one or more central processing units (CPUs), one or more graphics processing units (GPUs), one or more tensor processing units (TPUs), and other processing units.
[0064] The drive control system 150 provides control signals to the electromechanical system 190 to achieve physical control of the vehicle 105. For example, when in fully automated or semi-automated driving mode, the drive control system 150 receives planned actions from the vehicle control system 115 (as discussed further below) and converts the planned actions into control signals using the steering unit 152, the brake unit 154, and the throttle (or acceleration) unit 156. Each unit 152, 154, 156 can be implemented as one or more software modules or one or more control blocks within the drive control system 150. The drive control system 150 may include other components for controlling other aspects of the vehicle 105 (including controlling turn signals and brake lights, etc.).
[0065] The electromechanical system 190 receives control signals from the drive control system 150 to operate electromechanical components in the vehicle 105. The electromechanical system 190 affects the physical operation of the vehicle 105. The electromechanical system 190 includes an engine 192, a transmission 194, and wheels 196. The engine 192 can be a gasoline-powered engine, a battery-powered engine, a hybrid engine, etc. Other components can be included in the mechanical system 190, including, for example, turn signals, brake lights, fans, and windows.
[0066] The memory 126 of the vehicle control system 115 stores thereon software instructions executable by the processor system 102. The software instructions can be executed by the processor system 102 to implement one or more software systems, software subsystems, and software modules. Generally, it should be understood that the software systems, software subsystems, and software modules disclosed herein can be implemented as instruction sets stored in the memory 126. For example, the memory 126 may include executable instructions for implementing an operating system 160 and an ADS or ADAS including a planning system 300 (also referred to as a path planning system). The planning system 300 may be a machine learning-based system that generates a planned path (which may include a planned subpath and a planned behavior) to be executed by the vehicle 105. The planning system 300 includes a task planning subsystem 305, a behavior planning subsystem 310, a motion planning subsystem 315, and an object behavior prediction subsystem 320. The object behavior prediction subsystem 320 also includes a classification interaction subsystem 400 and a behavior predictor 322. The details of the classification interaction subsystem 400 will be further provided below.
[0067] The planning and decision making of the planning system 300 may be dynamic and repeatedly executed as the environment changes. The changes in the environment may be due to the movement of the vehicle 105 (e.g., the vehicle 105 approaches a newly detected obstacle) and due to the dynamic nature of the environment (e.g., moving pedestrians and other moving vehicles).
[0068] The planning subsystems 305, 310, 315 perform planning at different levels of detail. Task-level planning performed by the task planning subsystem 305 is considered a higher (or more global) planning level, motion-level planning performed by the motion planning subsystem 315 is considered a lower (or more local) planning level, and behavior-level planning performed by the behavior planning subsystem 310 is considered a planning level between the task level and the motion level. Typically, the output from a higher-level planning subsystem may constitute at least a portion of the input to a lower-level planning subsystem.
[0069] Planning of the mission planning subsystem 305 (more simply referred to as mission planning) involves planning the path of the vehicle 105 at a high or global level, such as planning a route of travel from a starting point (e.g., a home address) to a final destination point (e.g., a work address). The behavior planning subsystem 310 can receive the planned route from the mission planning subsystem 305. The behavior planning subsystem 310 is concerned with controlling the behavior of the vehicle 105 on a more local and short-term basis than the mission planning subsystem 305. The behavior planning subsystem 310 can generate behavior decisions based on certain rules (e.g., traffic rules, such as speed limits or signs) or guidance (e.g., guidance for smooth and efficient driving, such as taking a faster lane if possible). The behavior decisions can be provided as part of the input to the motion planning subsystem 315. The motion planning subsystem 315 is concerned with controlling the movement of the vehicle 105 based on the direct environment 100 of the vehicle 105. The motion planning subsystem 315 generates planned vehicle movements to ensure the safety of vehicle occupants and other objects (e.g., pedestrians, other vehicles, cyclists, etc.) in the environment. Because the environment can be highly dynamic (eg, pedestrians and other vehicles are moving), the motion planning subsystem 315 should perform motion planning that can account for expected (or predicted) changes in the environment 100 .
[0070] Sensor data received from the environmental sensors 110 and vehicle data received from the vehicle control sensors 111 (and optionally also positioning data collected from the satellite network 260) can be used by the perception system 178 of the ADS and ADAS to generate processed data (e.g., feature vectors, occupancy grid maps (OGMs), object classifications and bounding boxes, etc.) representing features of the environment 100. The perception system 178 can include one or more machine learning-based systems (e.g., trained neural networks) that generate processed data representing features of the environment 100 at each time step.
[0071] The perception system 178 may include any number of independent or interconnected systems or functions, and may include, for example, rule-based systems, machine learning-based systems, and combinations thereof. Machine learning-based systems may be implemented using neural networks, such as any type of DNN (including CNN or RNN), long short-term memory networks, etc. In some examples, the perception system 178 may include: a fusion subsystem for fusing sensor data and vehicle data from multiple environmental sensors 110 and vehicle sensors 111 to generate fused data; a classification subsystem for processing sensor data or fused data to detect and identify objects in the environment 100 (e.g., detecting and identifying stationary obstacles, pedestrians or another vehicle, lanes and lane boundaries, and traffic lights / signs, etc.); and a positioning and mapping subsystem for building or updating a map of the environment 100 and estimating the position of the vehicle 105 in the map.
[0072] The memory 126 may also have stored thereon instructions for implementing other software systems, subsystems, and modules, such as a navigation system, a climate control system, a media player system, a telephone system, and / or a messaging system, among others.
[0073] Figure 3 is a block diagram of exemplary details of the classification interaction subsystem 400. The classification interaction subsystem 400 may be a subsystem of the object behavior prediction subsystem 320 (and may also be a subsystem of the planning system 300). However, it should be understood that the classification interaction subsystem 400 may also be used outside of the object behavior prediction subsystem 320, outside of the planning system 300, and / or outside of the ADS or ADAS of the autonomous vehicle 105. For example, the classification interaction subsystem 400 may be applicable to any computing system where the encoding of interactions between dynamic objects would be useful (e.g., assisted robotics, anomaly detection, or intelligent traffic management, etc.).
[0074] The classification interaction subsystem 400 receives feature data as input and generates a classification interaction representation as output, which is data (e.g., a tensor) representing the strength of the interaction (also referred to as classification interaction) between defined object categories. The classification interaction representation can be provided to a prediction model, such as a prediction model implemented in the behavior predictor 322, to predict the behavior (e.g., actions and trajectories) of a dynamic object of interest (i.e., a target dynamic object) in the environment 100 at a future time step. The predicted behavior can be a predicted action, a predicted trajectory, or a predicted action-trajectory pair, as well as other possibilities. The predicted action can be a higher-level prediction containing a behavior class (e.g., a predicted class label of the behavior) and can be associated with an expected future event. The predicted action may not provide data on the specific motion (or trajectory) that the dynamic object of interest will perform to perform the predicted action. The predicted trajectory can be a lower-level prediction of the motion or position of the dynamic object of interest in a future time interval. The predicted trajectory does not include a label for understanding the meaning of the trajectory. For example, if the dynamic object of interest is a pedestrian, the predicted action may be a pedestrian walking toward another vehicle (which may be associated with a future event that the other vehicle will begin to move) or a pedestrian crossing a road, then the predicted trajectory may be the path taken by the pedestrian to reach the other vehicle or to the other side of the road. The predicted behavior of the dynamic object of interest may then be provided to the motion planning subsystem 315 to plan the motion to be performed by the vehicle 105.
[0075] The feature data input to the classification interaction subsystem 400 is data representing the environment 100 of the vehicle 105. In some examples, the feature data may include multiple data sets. Each data set may be a corresponding time data set (i.e., data corresponding to multiple time steps) representing one or more specific features of the environment 100 (e.g., a specific object category in the environment 100). Accordingly, the feature data set may also be referred to as time series (or observation sequence) feature data, and the feature data input to the classification interaction subsystem 400 may include multiple time series feature data. All time series feature data input to the classification interaction subsystem 400 may be simply collectively referred to as feature data.
[0076] The feature data may include time series data generated by one or more environmental sensors 110, vehicle sensors 111 (with little or no processing, also referred to as raw sensor data), and may also include time series data generated by the perception system 178. The feature data input to the classification interaction subsystem 400 may include time series data in different formats (e.g., structured data such as matrix data or tensor data, and unstructured data such as point cloud data). For example, the feature data input to the classification interaction subsystem 400 may include one or more feature vectors, object classification data (e.g., encoded as 2D matrix data), object segmentation data (e.g., encoded as 2D matrix data), one or more object bounding boxes, one or more OGMs, one or more 2D images, and / or one or more point clouds, etc. Therefore, the feature data may include time series data representing various features of the environment 100 from an earlier time step to a current time step. The time length of each time series in the feature data may be equal, and the time intervals between time steps may also be equal, but it should be understood that this is not a strict requirement. The time length of the feature data can be any suitable duration (e.g., up to 10 minutes or longer, 10 minutes or less, 5 minutes or less, etc.). The time length of the feature data can be predefined based on the expected environment (e.g., a shorter time length can be suitable for a highly changing dynamic environment, or a longer time length can be suitable for a slowly changing dynamic environment). A moving time window can be used to ensure that the time series of the feature data represents the desired time length, and / or a decay factor (or discount factor) can be used to ensure that older data is gradually reduced from the feature data.
[0077] In some examples, an optional preprocessing subsystem (not shown) may be present in the classification interaction subsystem 400 to preprocess feature data of different formats into a common format. For example, each time step in a given time series feature data may be preprocessed into a corresponding 2D semantic segmentation map (e.g., in the format of a 2D image, where each pixel in the image is assigned a label of an object class in a defined set of object classes) to convert the given time series feature data into a time series 2D image, where each 2D image represents a corresponding time step.
[0078] A plurality of time series feature data (after optional preprocessing, if applicable) is received by a classifier 405 that performs object classification. For example, the classifier 405 may use a trained model (e.g., a model whose parameters are learned using a machine learning algorithm during training) to identify feature data belonging to each object category. The classifier 405 classifies the time series feature data into corresponding object categories based on the shared features of the objects represented in each time series feature data. Each category is defined by one or more shared object features (e.g., object class, dynamic / static properties, etc.). Object categories may be defined according to the expected environment 100 (e.g., in the context of autonomous driving or assisted driving, object categories may be pedestrians, vehicles, cyclists, and static objects). At least one object category may also be defined for a specific object of interest (e.g., a pedestrian of interest, such as a pedestrian physically closest to the vehicle 105; a vehicle of interest, such as an oncoming vehicle; or any other dynamic object of interest in the environment 100). It should be understood that object categories may be defined for any entity or feature in the environment 100 according to the desired application. For example, anything in environment 100 that may affect the behavior of dynamic objects (e.g., objects that can be segmented in an image, such as trees, fire hydrants, people, animals; and objects that may not be easily segmented in an image, such as curbs, lane markings, etc.) can be the basis for defining object categories.
[0079] Any number of object classes may be defined. The defined object classes are typically fixed (e.g., as part of designing a model for classifier 405 of object behavior prediction subsystem 320). The output of classifier 405 is a classification set (or simply a classification data set) of time series data, where the number of classification sets corresponds to the number of defined object classes. Each object class may have one or more time series feature data. Figure 3 In the example of , there are five defined object categories, and classifier 405 outputs five classification data sets.
[0080] The classification data set is received by the spatiotemporal encoder 410, which encodes the classification data set into a corresponding classification representation. For each classification data set, a corresponding time representation is generated by encoding the time changes in each classification data set (which may contain multiple time series feature data) into a single time series representation. For example, if the classification data set contains feature data at m time steps, the classification representation output by the spatiotemporal encoder 410 can be a tensor m×n, where n is the number of feature points (where each feature point is a data point representing a corresponding entry in a feature vector or feature map).
[0081] The spatiotemporal encoder 410 may be implemented using a deterministic approach (e.g., using a deterministic pattern filter, using aggregation, etc.) or a machine learning-based approach (e.g., using Markov chain modeling, RNN, or any other technique capable of encoding temporal relationships). For example, the spatiotemporal encoder 410 may be a single neural network that performs encoding for all object categories. Alternatively or in addition, the spatiotemporal encoder 410 may include a single classification encoder 412a to 412e (generally referred to as classification encoder 412) that performs encoding for the corresponding category (in Figure 3 In the example of , five classification encoders 412 are used to encode five categories). Each classification encoder 412 can independently be a trained neural network, such as an RNN, or a combination of an RNN and a CNN. The classification encoders 412 can operate completely independently of each other, or can share some calculations with each other (e.g., weight calculations of shared convolutional layers) to achieve higher computational efficiency. Each classification encoder 412 receives a corresponding classification data set (which may include different time series feature data that represent different features of the environment 100 and / or have different formats) and maps the classification data set to a set of feature points for each time step (e.g., in the form of a feature vector), where the temporal changes in the features are encoded in the feature points. Therefore, the output from each classification encoder 412 can be a time series feature vector, referred to herein as a classification representation.
[0082] The categorical representations of the corresponding object categories are output by the spatiotemporal encoder 410 and received by the combiner 415. The combiner 412 combines the categorical representations to form a single shared representation, which is a single representation of the spatiotemporal variations in each category. For example, the combiner 415 can perform temporal concatenation of the categorical representations to output a single shared representation. Temporal concatenation means that data belonging to the same time step are concatenated together. For example, if each categorical representation is a corresponding m×n tensor, where m is the number of time steps and n is the number of feature points, the categorical representations can be concatenated together according to the time step indexes from 1 to m to obtain a single shared representation of size m×5n (where 5 corresponds to Figure 3 The combiner 415 may use other techniques to generate shared representations, such as an average operation (e.g., averaging all classification representations at each given time step), or a maximum operation (e.g., taking the maximum value among all classification representations at each given time step), etc.
[0083] The single shared representation is provided to the interactive attention subsystem 420. The interactive attention subsystem 420 generates a classified interactive representation, which is a weighted representation of the environment based on the temporal changes in the weights for each object category in the single shared representation. The interactive attention subsystem 420 can use any suitable attention mechanism to generate attention scores, context vectors, and attention vectors. For example, Luong et al. describe a suitable attention mechanism in "Effective approaches to attention-based neural machine translation" (arXiv:1508.04025, 2015). The attention vector generated by the interactive attention subsystem 420 is called a classified interactive representation because the attention vector is a vector that represents the influence (or interaction) of features in one object category on features in another object category.
[0084] exist Figure 3 In the example of , the interactive attention subsystem 420 includes a sequence attention score generator 422, a context vector generator 424, and a classification interaction encoder 426. In other examples, there may be no separate functional blocks 422, 424, 426, and the interactive attention subsystem 420 may alternatively use more or less functions to implement the operations of the functional blocks 422, 424, 426. The interactive attention subsystem 420 will be described here with respect to the functional blocks 422, 424, 426.
[0085] The sequence attention score generator 422 generates a corresponding weight score for each time step of a single shared representation. The weight scores can be learned by the sequence attention score generator 422 based on the impact of each feature on the prediction task to be performed by the behavior predictor 322. The sequence attention score generator 422 can implement a fully connected layer to perform this learning. If the shared representation is of size m×k (where m is the number of time steps and k is the number of combined feature points), the weight scores can also be stored in a weight score matrix of size m×k. The weight score matrix is multiplied by the k×1 feature points in the last time step of the shared representation, and the resulting m×1 vector is normalized (e.g., using a softmax operation) so that its entries sum to 1. The result of this normalization is an m×1 vector of attention scores, representing the attention (or weight) that should be applied to each feature at each time step.
[0086] The attention scores generated by the sequence attention score generator 422 are provided to the context vector generator 424. At the context vector generator 424, the attention scores generated for each time step are applied to (e.g., multiplied by) the feature points of the single shared representation for each corresponding time step. For example, the m×1 vector of attention scores is transposed and multiplied by the m×k single shared representation. The result of this operation is a context vector of size 1×k, where each entry is a weighted sum of the effects of each feature point representing the single shared representation on the temporal changes in the time series (i.e., the features of the last time step of the single shared representation).
[0087] The context vector generated by the context vector generator 424 is provided to the classification interaction encoder 426. At the classification interaction encoder 426, the context vector is combined (e.g., concatenated) with the features of the last time step of the single shared representation and then processed with a fully connected network. The fully connected network effectively performs an embedding operation to reduce the size of the data and generate a classification interaction representation.
[0088] The classification interaction representation is the output of the classification interaction subsystem 400. The classification interaction representation is a unified representation of size ci, where ci is the output size of the last layer of the fully connected network included in the classification interaction encoder 426 in the classification interaction subsystem 400. The classification interaction representation is a weight vector, where the weight represents the relative influence of the temporal changes of the features in each object category on the features of all object categories at the last time step.
[0089] The classification interaction subsystem 400 may include multiple trained neural networks (e.g., a trained RNN included in the spatiotemporal encoder 410, and a trained fully connected network may be included in the interactive attention subsystem 420). Some neural networks included in the classification interaction subsystem 400 (e.g., a neural network for classifying feature data, or a neural network for extracting feature data) may be trained offline (e.g., before being included in the object behavior prediction subsystem 320), for example, on one or more physical machines or one or more virtual machines, using training data from an external database (e.g., a sequence of images of a sensed environment (commonly referred to as a scene)). Any neural network implemented in the spatiotemporal encoder 410 (e.g., a neural network implemented in the classification encoder 412) should be trained end-to-end. Any suitable training technique may be used for end-to-end training of the neural network included in the object behavior prediction subsystem 320 (e.g., performing forward propagation using samples of labeled training data, calculating loss using the loss function of the object behavior prediction subsystem 320, and performing backpropagation with gradient descent to learn (i.e., optimize) the parameters of the neural network included in the object behavior prediction subsystem 320).
[0090] The classification interaction representation generated by the classification interaction subsystem 400 may be provided to the object behavior prediction subsystem 320).
[0091] Figure 4 is a block diagram of exemplary details of the object behavior prediction subsystem 320. The object behavior prediction subsystem 320 can be part of the planning system 300 of the vehicle 105. In the following discussion, the object behavior prediction subsystem 320 will be described in the context of predicting the behavior of dynamic objects of interest (e.g., pedestrians) in the environment 100 of the vehicle 105. However, it should be understood that the behavior prediction subsystem 320 can also be used outside of the planning system 300 and / or outside of the autonomous vehicle 105 to predict the behavior of any dynamic object of interest in the sensing environment. For example, instructions of the behavior prediction subsystem 320 can be stored in a memory and executed by one or more processing units of any computing system, where the prediction of the behavior of the dynamic object of interest (including the behavior resulting from the interaction of the dynamic object of interest with other objects) will be useful (e.g., assisting robots, anomaly detection, or intelligent traffic management, etc.).
[0092] The input data to the object behavior prediction subsystem 320 may include unprocessed (or only slightly processed) data from the environment sensor 110, the vehicle sensor 111, and may also include processed data generated by the perception system 178 from the environment sensor 110, the vehicle sensor 111. It should be noted that the input data to the object behavior prediction subsystem 320 includes feature data input to the classification interaction subsystem 400. In addition, the input data may include data representing the dynamics of multiple objects in the environment 100 and data representing the state of the vehicle 105 itself (e.g., data representing the current speed, current acceleration, etc.). The data representing the dynamics of multiple objects in the environment 100 may include data representing object trajectories (e.g., trajectories of dynamic objects in the environment 100, which may be represented as a 2D map), data representing object positions (e.g., positions of objects in the environment 100 represented as one or more OGMs), and data representing object motion (e.g., motion of dynamic objects in the environment 100). The input data to the behavior prediction subsystem 320 includes time data (e.g., object trajectories), and may also include non-time data (e.g., the current speed of the vehicle 105). As described above, the classification interaction subsystem 400 receives a portion of the input data (ie, the feature data described above) and generates a classification interaction representation. The classification interaction representation is provided to the behavior predictor 322.
[0093] The behavior predictor 322 also receives at least a portion of the input data (e.g., data representing object trajectories, object positions, vehicle states, etc.). The behavior predictor 322 can be any neural network (e.g., a recursive decoder network, a feedforward generative neural network, etc.) that can be trained to generate prediction data representing the predicted behavior of one or more dynamic objects of interest. Specifically, the behavior predictor 322 can include any neural network that can be trained to receive classified interaction representations as input (among other inputs) and generate prediction data representing the predicted behavior of the dynamic objects of interest. The prediction data can be provided as part of the input to the motion planning subsystem 315. The output of the behavior predictor 322 can include different types of prediction data that can vary depending on a specific application (e.g., depending on the input required by the motion planning subsystem 315). For example, the predicted data may include predicted trajectories of dynamic objects of interest (e.g., predicted trajectories of pedestrians in future time periods, which may be represented as a 2D map), predicted future positions of dynamic objects of interest (e.g., predicted positions of pedestrians at future time steps, which may be represented as a 2D OGM), and / or predicted behaviors of dynamic objects of interest (e.g., predicted crossing actions of pedestrians at future time steps, which may be represented as vector or scalar values).
[0094] The prediction data output by the behavior predictor 322 may also be provided as input to the motion planning subsystem 315. The motion planning subsystem 315 may use the prediction data along with other input data (e.g., processed and / or unprocessed data from the sensors 110, 111) to generate a planned path for the vehicle 105.
[0095] The object behavior prediction subsystem 320 may be trained offline using training data from an external data set (e.g., on one or more physical machines or one or more virtual machines, before being deployed to the ADS or ADAS of the vehicle 105). Specifically, the classification interaction subsystem 400 may have been trained, and the behavior predictor 322 may be trained using the training data (e.g., images, trajectories, motions, etc. of the vehicle 105) and also using the classification interaction representations generated from the training data by the classification interaction subsystem 400. Any suitable training technique may be used depending on the type of neural network implemented in the behavior predictor 322.
[0096] Figure 5 is a flow chart of an exemplary method 500 for generating behavior predictions (eg, predicted behaviors) of a dynamic object of interest based on classified interactions. The method 500 may be performed using the object behavior prediction subsystem 320 (eg, using the classified interactions subsystem 400 ).
[0097] In optional step 502, time series sensor data (i.e., sensor data over a sequence of time steps) may be preprocessed into time series feature data (i.e., feature data over the same sequence of time steps). For example, sensor data generated by environmental sensor 110 over multiple time steps may be preprocessed into multiple 2D maps (e.g., OGMs) to represent the locations of objects classified over the time steps (i.e., objects classified into one of multiple object categories). In some examples, preprocessing may be used to generate at least one time series feature data used in the next step 504, and another time series feature data may not require preprocessing.
[0098] In 504, all feature data (including any time series feature data generated by the optional preprocessing in step 502) are classified into one of a plurality of defined object categories (e.g., using the operations described above with respect to classifier 405). At least one of the object categories may be defined as a specific dynamic object of interest. One or more shared features (e.g., object class, dynamic or static properties of data, etc.) are defined for each object category, and feature data representing objects having one or more shared features defined for a given object category are grouped into the given object category. The classification group is referred to as a classification data set, wherein each object category may have one or more time series feature data.
[0099] At 506, the classification data sets are encoded into corresponding classification representations (e.g., using the operations described above with respect to the spatiotemporal encoder 410). A corresponding neural network (e.g., RNN) can perform encoding on each classification data set. Each classification representation is a representation of the temporal variation of all features belonging to a corresponding object class. For example, each classification representation can be a time series feature vector (or feature map) representing the features of the corresponding object class at each time step, including the temporal variation of the features.
[0100] At 508, the categorical representations from all defined object categories are combined into a single shared representation (eg, using the operations described above with respect to combiner 415). For example, the single shared representation may be a concatenation of the categorical representations, depending on the time step.
[0101] At 510, a classification interaction representation is generated from a single shared representation using an attention mechanism (e.g., using the operations described above with respect to the interaction attention subsystem 420). Typically, the attention mechanism involves learning an attention score (representing the attention (or weight) to be applied to each feature at each time step), computing a context vector (representing a weighted sum representing the influence of the features at each time step on the temporal variation in the time series), and learning an attention vector (representing the relative influence of the temporal variation of the features in each object category on the features of all object categories at the last time step). As previously described, any suitable attention mechanism can be used to generate an attention vector as a classification interaction representation. For example, step 510 can be performed using steps 512, 514, and 516 to generate an attention score, a context vector, and a classification interaction representation, respectively.
[0102] At 512, attention scores are generated (e.g., using the operations described above with respect to the sequence attention score generator 422) representing the influence of the features at each time step in the shared representation on the features at the last time step of the shared representation.
[0103] At 514, the attention scores are applied to the features at each time step of the single shared representation to generate a context vector (e.g., using the operations described above with respect to context vector generator 424).
[0104] At 516, the context vector is combined with the features of the last time step of the single shared representation to generate a categorical interaction representation (e.g., using the operations described above with respect to the categorical interaction encoder 426).
[0105] At 518, the classified interaction representation generated using the attention mechanism is provided as an input to a behavior predictor 322 (which may include any suitable trained neural network) to generate prediction data representing predicted future behavior of the dynamic object of interest (e.g., using the behavior predictor 322). The classified interaction representation is provided as an input to the behavior predictor 322 along with other input data, and the input data required by the behavior predictor 322 may depend on the design of the neural network included in the behavior predictor 322. The prediction data representing the predicted future behavior of the dynamic object of interest may include a trajectory of the dynamic object of interest (e.g., in the context of autonomous driving, a pedestrian of interest, a vehicle of interest, or other road user of interest) at one or more future time steps.
[0106] Prediction data representing predicted future behavior of a dynamic object of interest may be used to generate other prediction and / or planning operations (eg, motion planning for motion planning subsystem 315).
[0107] In various examples, the present invention describes methods and systems that can predict the behavior of dynamic objects of interest in an environment. Specifically, the present invention describes methods and systems that can encode interactions between different object categories. Time series feature data is classified into defined object categories, a single shared representation is generated by combining the individual classification representations, and an attention mechanism is used to transform the single shared representation into a single classification interaction representation, which is a weighted vector that encodes the interactions between the different object categories.
[0108] The present invention can model (i.e., encode) interactions based on interactions of defined object categories rather than interactions between corresponding single objects and other objects in a sensing environment. This achieves better scalability than some existing solutions because the number of object categories is defined rather than variable. Encoding interactions of defined object categories rather than encoding interactions between corresponding single objects and other objects in a sensing environment can also improve computational efficiency and reduce memory resource usage compared to some existing solutions.
[0109] Although the present invention has been described in the context of an autonomous driving system, it should be understood that the present invention can be applied to a variety of applications where multiple interactive objects exist. For example, examples of the present invention can be used to assist robotics, anomaly detection, or intelligent traffic management, where multiple interactive agents and elements are involved. It should be understood that the present invention is not intended to be limited to a specific type of environmental features or a specific type of mission objectives.
[0110] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed in the present invention can be implemented by electrical hardware, or a combination of computer software and electrical hardware. Whether the function is performed by hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions of each specific application, but it should not be considered that the implementation is beyond the scope of the present invention.
[0111] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can be based on the corresponding processes in the above method embodiments and will not be repeated here.
[0112] It should be understood that the disclosed systems and methods may be implemented in other ways. The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, may be located in one location, or may be distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of the embodiment. In addition, the functional units in the various embodiments of the present application may be integrated into a processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0113] When these functions are implemented in the form of software functional units and sold or used as independent products, these functions can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The software product is stored in a storage medium, including several instructions to instruct a computer device (which can be a personal computer, a server or a network device) to perform all or part of the steps of the method described in each embodiment of the present application. The above-mentioned storage medium includes any medium that can store program code, such as a universal serial bus (universal serialbus, USB) flash drive, a removable hard disk, a read-only memory (read-only memory, ROM), a random access memory (random access memory, RAM), a disk or an optical disk, etc.
[0114] The above are only some specific implementations of the present application and are not intended to limit the protection scope of the present application. Within the technical scope disclosed by the present invention, any changes or substitutions that can be thought of by those skilled in the art should be included in the protection scope of the present invention.
Claims
1. A method for predicting the behavior of a dynamic object of interest in an environment of a vehicle, characterized in that The method comprises: Receiving a plurality of time series feature data, each time series feature data represents corresponding features of a plurality of objects in the environment at a plurality of time steps, the plurality of objects including a dynamic object of interest; classifying each time series feature data into one of a plurality of defined object categories to obtain a classification data set for each defined object category, each classification data set comprising one or more time series feature data, the one or more time series feature data representing one or more objects belonging to the corresponding defined object category; encoding each classification dataset into a corresponding classification representation, each classification representation representing a temporal variation of a feature in said corresponding defined object category; combining the categorized representations into a single shared representation; generating a classification interaction representation based on the single shared representation, the classification interaction representation being a weighted representation of the single shared representation representing the impact of temporal changes in each defined object category on a final time step of the single shared representation; Predictive data representing predicted future behavior of the dynamic object of interest is generated based on the classified interaction representations, the data representing the dynamics of the plurality of objects, and the data representing the state of the vehicle.
2. The method according to claim 1, characterized in that Combining the categorical representations includes concatenating the categorical representations according to time steps to generate the single shared representation.
3. The method according to claim 1, characterized in that Encoding each classification dataset involves: For a given classification data set belonging to a given object category, the one or more time series feature data are provided to a trained neural network to generate a time series feature vector as a classification representation of the given object category.
4. The method according to claim 3, characterized in that The trained neural network is a recursive neural network, a convolutional neural network, or a combined recursive and convolutional neural network.
5. The method according to any one of claims 1 to 4, characterized in that At least one defined object class is associated with the dynamic object of interest.
6. The method according to any one of claims 1 to 4, characterized in that The method further comprises: receiving time series sensor data generated by the sensor; The time series sensor data is preprocessed into a time series feature data, wherein the time series feature data is included in the received multiple time series feature data.
7. The method according to any one of claims 1 to 4, characterized in that The method further comprises: The prediction data representing the predicted future behavior of the dynamic object of interest is provided to a motion planning subsystem of the vehicle to generate a planned path for the vehicle.
8. A computing system for predicting the behavior of a dynamic object of interest in an environment of a vehicle, characterized in that The computing system comprises: A processor system is used to execute instructions to enable the object behavior prediction subsystem of the computing system to perform the following operations: Receiving a plurality of time series feature data, each time series feature data represents corresponding features of a plurality of objects in the environment at a plurality of time steps, the plurality of objects including a dynamic object of interest; classifying each time series feature data into one of a plurality of defined object categories to obtain a classification data set for each defined object category, each classification data set comprising one or more time series feature data, the one or more time series feature data representing one or more objects belonging to the corresponding defined object category; encoding each classification dataset into a corresponding classification representation, each classification representation representing a temporal variation of a feature in said corresponding defined object category; combining the categorized representations into a single shared representation; generating a classification interaction representation based on the single shared representation, the classification interaction representation being a weighted vector representing the impact of temporal changes in each defined object category on a final time step of the single shared representation; Predictive data is generated based at least on the classified interaction representation, the data representing the dynamics of the object, and the data representing the state of the vehicle, the predictive data representing a predicted future behavior of the dynamic object of interest.
9. The computing system according to claim 8, characterized in that: The processor system is configured to execute instructions to combine the categorical representations by concatenating the categorical representations according to a time step to generate the single shared representation.
10. The computing system according to claim 8, characterized in that: The processor system is configured to execute instructions to encode each classification data set by: For a given classification data set belonging to a given object category, the one or more time series feature data are provided to a trained neural network to generate a time series feature vector as a classification representation of the given object category.
11. The computing system according to claim 10, characterized in that: The trained neural network is a recursive neural network, a convolutional neural network, or a combined recursive and convolutional neural network.
12. The computing system according to any one of claims 8 to 11, characterized in that: At least one defined object class is associated with the dynamic object of interest.
13. The computing system according to any one of claims 8 to 11, characterized in that: The processor system is used to execute instructions to enable the computing system to perform the following operations: receiving time series sensor data generated by the sensor; The time series sensor data is preprocessed into a time series feature data, wherein the time series feature data is included in the received multiple time series feature data.
14. The computing system according to any one of claims 8 to 11, characterized in that: The vehicle is an autonomous vehicle, wherein the computing system is implemented in the autonomous vehicle, and the processor system is used to execute instructions to cause a motion planning subsystem of the autonomous vehicle to perform the following operations: receiving as input said prediction data representing predicted future behavior of said dynamic object of interest; A planned path is generated for the autonomous vehicle.
15. A computer readable medium, characterized in that The method comprises computer executable instructions which, when executed by a processor system of a computing system, cause the computing system to perform the method according to any one of claims 1 to 7.
16. A computer program, characterized in that The method comprises computer executable instructions which, when executed by a processor system of a computing system, cause the computing system to perform the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Deeply integrated fusion architecture for automated driving systems
CN109291929A
Making object-level predictions of future state of physical system
CN110770760A