Method and system for predicting trajectory of an object
By adopting a recursive convolution one-step feature pyramid network in multi-object tracking of radar data, combining the residual backbone and feature pyramid of recursive components, the problem of target and background distinction and target size and rotation estimation in radar data is solved, and efficient and reliable multi-object tracking is achieved.
Patent Information
- Application Number
- CN202110844313.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-07-16
- Filing Date
- 2021-07-26
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-07-26
AI Technical Summary
Existing radar data multi-object tracking methods are difficult to accurately distinguish targets from backgrounds in noise and sparse environments, estimate target size and rotation, especially when targets are stationary.
A recursive convolution one-step feature pyramid network is adopted, combining the residual backbone and feature pyramid of recursive components, and detection and motion prediction are performed based on radar data to solve the tracking task.
It realizes efficient and reliable multi-object tracking of radar data, and improves the accuracy of object detection and motion prediction in noise and sparse environments.
Smart Images

Figure CN113971433B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to methods and systems for predicting a trajectory of an object, such as methods and systems for determining a parameterized plurality of parameters of a predicted trajectory of an object. Background Art
[0002] Object tracking is an essential function, for example in at least partially autonomously driven vehicles.
[0003] Conventional methods for multi-object tracking via radar data perform clustering to combine single point detections into object proposals, and then use recursive filtering with an appropriate motion model to produce object tracks. The tracking problem is thus split into motion prediction and data association of detected objects with existing object tracks. Due to the noisiness and sparsity of radar data, conventional methods often have problems distinguishing targets from backgrounds and correctly estimating target size and rotation, especially when the target is stationary.
[0004] Therefore, there is a need to provide more reliable and efficient object tracking. Summary of the invention
[0005] The present disclosure provides a computer-implemented method, a computer system, a vehicle, and a non-transitory computer-readable medium. Implementations are presented in the specification and drawings.
[0006] In one aspect, the present disclosure relates to a computer-implemented method for predicting a trajectory of an object, the method comprising the following steps performed (in other words: carried out) by a computer hardware component: acquiring radar data of the object; determining first intermediate data based on the radar data based on a residual backbone using recursive components; determining second intermediate data based on the first intermediate data using a feature pyramid; and predicting the trajectory of the object based on the second intermediate data. According to various embodiments, the radar data of the object includes at least one of radar cube data and radar point data.
[0007] In other words, the method can provide a multi-object tracking method for radar data, which combines multiple methods into a recurrent convolutional one-stage feature pyramid network and jointly performs detection and motion prediction on radar data (e.g., radar point cloud data or radar cube data) to solve the tracking task.
[0008] Radar cube data may also be referred to as radar data cube.
[0009] According to various embodiments, a network may be provided that uses radar data as input, propagates the data through multiple intermediate network layers that partially correlate information from different scans at different time points to generate intermediate data (e.g., using a recursive convolutional neural network). The network may process the intermediate data using a pyramid structure, and generate new intermediate data, and ultimately use the intermediate data to detect objects and predict the trajectory of each object.
[0010] According to various embodiments, the input data may include or may be radar cube data, and the network layers may be used to propagate and transform the input data to a desired output data domain.
[0011] According to various embodiments, the input data may include or may be radar point data as an input representation of radar data.
[0012] This can provide an efficient and reliable way of object detection and motion trajectory prediction.
[0013] According to another aspect, the computer-implemented method further includes the following steps performed by the computer hardware component: using a regression head, determining third intermediate data based on the second intermediate data; and predicting a trajectory of the object based on the third intermediate data.
[0014] According to another aspect, the computer-implemented method further includes the following steps performed by the computer hardware component: using a classification head, determining fourth intermediate data based on the second intermediate data; and predicting a trajectory of the object based on the fourth intermediate data.
[0015] It should be understood that the trajectory of the object may be predicted based on the third intermediate data and the fourth intermediate data.
[0016] The classification head can be used to obtain category information, such as at which of the predetermined anchor positions the actual objects currently reside, and which classes these objects belong to (e.g. pedestrians, vehicles, etc.). The regression head can provide continuous information (such as the exact position, size, and rotation of the object over time), and can therefore only be responsible for performing motion prediction. Therefore, the classification head can only participate in the generation of predictions / tracklets as long as it indicates which positions within the output map of the regression head contain information belonging to actual objects.
[0017] According to another aspect, the step of predicting the trajectory comprises: determining a trajectory segment for a plurality of time steps. A trajectory segment may comprise a predicted trajectory for only a limited or small number of time steps (e.g. 3, or e.g. 5), which may then be used to determine a longer trajectory (starting from a different time step).
[0018] According to another aspect, a residual backbone using a recursive component comprises a residual backbone preceded by a stack of recursive layers.
[0019] According to another aspect, the residual backbone using the recursive component comprises a recursive residual backbone comprising a plurality of recursive layers. It has been found that providing a plurality of recursive layers in the backbone improves performance by allowing the network to fuse temporal information at multiple scales.
[0020] According to another aspect, the plurality of recursive layers comprises: a convolutional long short-term memory network, followed by a convolution, followed by a normalization.
[0021] According to another aspect, the plurality of recursive layers comprises: convolution, followed by normalization, followed by a rectified linear unit, followed by a convolutional long short-term memory network, followed by convolution, followed by normalization.
[0022] According to another aspect, the recursive component comprises a recursive loop that is executed once per time frame. It has been found that this can provide efficient processing.
[0023] According to another aspect, the recurrent component maintains hidden states between time frames. This may provide that the method (or a network used in the method) can learn to use information from arbitrarily distant points in time, and that past sensor readings do not need to be buffered and stacked to operate the method or network.
[0024] According to another aspect, the feature pyramid includes a transposed strided convolution, which can increase the richness of the features.
[0025] According to another aspect, the training of the method includes partitioning the training data set into multiple subsets and / or preserving the hidden state between optimization steps. This can improve the ability of the method to model temporal dependencies within the same memory constraint.
[0026] In one aspect, the present disclosure relates to a computer-implemented method for determining a parameterized plurality of parameters of a predicted trajectory of an object, the method comprising the following steps performed (in other words: carried out) by a computer hardware component: acquiring radar data of the object; determining first intermediate data based on the radar data based on a residual backbone using recursive components; determining second intermediate data based on the first intermediate data using a feature pyramid; and determining the plurality of parameters based on the second intermediate data. According to various embodiments, the radar data of the object comprises at least one of a radar data cube and radar point data.
[0027] A trajectory may be understood as a property (such as position or orientation / direction) that changes over time.
[0028] In other words, the method can provide a multi-object tracking method for radar data. The method can combine multiple methods into a recursive convolutional one-step feature pyramid network and jointly perform detection and motion prediction on radar data (e.g., radar point cloud data or radar data cube) to solve the tracking task.
[0029] According to various embodiments, a neural / neural network may be provided that uses radar data as input, propagates the data through multiple intermediate network layers that partially correlate information from temporally different scans to generate intermediate data (e.g., using a recursive convolutional neural network). The network may use a pyramid structure to process the intermediate data and generate new intermediate data, and ultimately use the intermediate data to detect objects and predict the trajectory of each object.
[0030] According to various embodiments, the input data may include or may be a radar data cube, and a network layer may be used to propagate and transform the input data to a desired output data domain.
[0031] According to various embodiments, the input data may include or may be radar point data as an input representation of radar data.
[0032] This can provide an efficient and reliable way of object detection and motion trajectory prediction.
[0033] According to another aspect, the parameterization comprises a polynomial of a predetermined degree; and the parameters comprise a plurality of coefficients associated with elements of a basis of a polynomial space of the polynomial of the predetermined degree. The method can provide polynomial prediction for DeepTracker.
[0034] According to another aspect, the elements of the basis include monomials. Monomials may include "1", "x", "x 2 ", ..., until the predetermined number of times.
[0035] The polynomial can be a function of time (so x is a variable time value).
[0036] According to another aspect, the parameterization comprises: a first parameterization parameterizing the position trajectory of the object; and a second parameterization parameterizing the orientation of the object. It has been found that not only the position trajectory can be parameterized, but also the orientation trajectory.
[0037] The velocity may be estimated (eg, via differentiation).
[0038] According to various embodiments, a network may be provided that uses radar data as input, propagates the data through multiple intermediate network layers that partially correlate information from different scans at different time points to generate intermediate data (e.g., using a recursive convolutional neural network). The network may process the intermediate data using a pyramid structure, and generate new intermediate data, and ultimately use the intermediate data to detect objects and predict the trajectory of each object.
[0039] According to various embodiments, the input data may include or may be a radar data cube, and a network layer may be used to propagate and transform the input data to a desired output data domain.
[0040] According to various embodiments, the input data may include or may be radar point data as an input representation of radar data.
[0041] According to another aspect, the computer-implemented method may further include the following steps performed by computer hardware components: using a regression head, determining third intermediate data based on the second intermediate data; and determining the parameter based on the third intermediate data.
[0042] The regression head can provide continuous information (such as the exact position, size, and rotation of the object over time) and can therefore only be responsible for performing motion prediction. Therefore, the classification head can only participate in the generation of predictions / tracklets as long as it indicates which locations within the output map of the regression head contain information belonging to actual objects.
[0043] According to another aspect, a residual backbone using a recursive component comprises a residual backbone preceded by a stack of recursive layers.
[0044] According to another aspect, the residual backbone using the recursive component comprises a recursive residual backbone comprising a plurality of recursive layers. It has been found that providing a plurality of recursive layers in the backbone improves performance by allowing the network to fuse temporal information at multiple scales.
[0045] According to another aspect, the plurality of recursive layers comprises: a convolutional long short-term memory network, followed by a convolution, followed by a normalization.
[0046] According to another aspect, the plurality of recursive layers comprises: convolution, followed by normalization, followed by a rectified linear unit, followed by a convolutional long short-term memory network, followed by convolution, followed by normalization.
[0047] According to another aspect, the recursive component comprises a recursive loop that is executed once per time frame. It has been found that this can provide efficient processing.
[0048] According to another aspect, the recurrent component maintains hidden states between time frames. This may provide that the method (or a network used in the method) can learn to use information from arbitrarily distant points in time, and that past sensor readings do not need to be buffered and stacked to operate the method or network.
[0049] According to another aspect, the feature pyramid includes transposed strided convolutions. This can increase the richness of features.
[0050] According to another aspect, the training of the method includes partitioning the training data set into multiple subsets and / or preserving the hidden state between optimization steps. This can improve the ability of the method to model temporal dependencies within the same memory constraint.
[0051] In another aspect, the present disclosure is directed to a computer system comprising a plurality of computer hardware components configured to perform some or all of the steps of the computer-implemented methods described herein.
[0052] The computer system may include multiple computer hardware components (e.g., a processor (e.g., a processing unit or a processing network), at least one memory (e.g., a storage unit or a storage network), and at least one non-transitory data storage device). It should be understood that additional computer hardware components may be provided and used to perform the steps of the computer-implemented method in the computer system. The non-transitory data storage device and / or the storage unit may contain a computer program for instructing a computer, for example, using the processing unit and the at least one storage unit to perform some or all steps or aspects of the computer-implemented method described herein.
[0053] In another aspect, the present disclosure is directed to a vehicle including at least a subset of the computer system described.
[0054] In another aspect, the present disclosure relates to a non-transitory computer-readable medium containing instructions for performing some or all steps or aspects of the computer-implemented methods described herein. The computer-readable medium can be configured as: an optical medium such as a compact disc or a digital versatile disc (DVD); a magnetic medium such as a hard disk drive (HDD); a solid-state drive (SSD); a read-only memory (ROM) such as a flash memory; and the like. In addition, the computer-readable medium can be configured as a data storage device that can be accessed via a data connection such as an Internet connection. The computer-readable medium can be, for example, an online data repository or cloud storage.
[0055] The present disclosure also relates to a computer program for instructing a computer to execute several or all steps or aspects of the computer-implemented method described herein.
[0056] Methods and apparatus according to various aspects or embodiments provide a deep multi-object tracker for radar.
[0057] The tracker is a one-step recursive convolutional neural network that performs dense detection and motion prediction on radar data (e.g., radar point cloud data or radar cube data). The tracker can operate on a single data frame at a time and can continuously update the memory of past information.
[0058] The motion prediction information predicted by the network at each frame is combined to produce the object trajectory.
[0059] The system mainly mixes and improves three recent deep learning systems: RetinaNet (which is an anchor-based feature pyramid detection network), Fast&Furious (which is a prediction-based tracking network), and DeepTracking (which is a recursive occupancy prediction network). BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Exemplary implementations and functions of the present disclosure are described herein in conjunction with the following drawings, which schematically illustrate:
[0061] Figure 1 is an illustrative diagram of a network architecture according to various embodiments;
[0062] Figure 2 is an illustrative diagram of a network architecture according to various embodiments;
[0063] Figure 3 Yes Figure 2 An example diagram of a recursive residual block design of the recursive residual backbone 202 is shown;
[0064] Figure 4 Yes Figure 2 An example diagram 400 of a recursive residual block design for the recursive residual backbone 202 is shown;
[0065] Figure 5 is an example diagram of how to determine a trajectory according to various embodiments;
[0066] Figure 6 is an illustrative diagram of a method of predicting a trajectory of an object according to various embodiments;
[0067] Figure 7 is an illustration of DeepTracker using a polynomial representation in a training configuration according to various embodiments; and
[0068] Figure 8 is a flow chart illustrating a method for determining a plurality of parameters of a parameterization of a predicted trajectory of an object, according to various embodiments. DETAILED DESCRIPTION
[0069] According to various embodiments, deep learning methods (including convolutional (artificial) neural networks) can be applied to detection and tracking problems using radar data. Various embodiments can use or enhance deep learning methods, such as RetinaNet, Fast & Furious tracker, and Deep Tracking.
[0070] The anchor-based one-step detection network, RetinaNet, uses a convolutional feature pyramid for efficient multi-scale detection on top of a ResNet backbone and an extension of binary cross entropy called “focal loss” (which emphasizes hard examples through an intuitive weighting scheme). RetinaNet is designed for vision and does not perform tracking or exploit temporal cues, which are critical for radar data.
[0071] The Fast&Furious tracker is an anchor-based one-stage network designed to operate on LiDAR (Light Detection and Ranging) data. Fast&Furious performs tracking by learning to generate fixed-length motion predictions for all detected targets and then registering and combining the resulting trajectory fragments (i.e., trajectories with only a few time steps) into complete trajectories. Fast&Furious exploits temporal information by performing 3D (three-dimensional) convolutions on stacked representations of multiple data frames. Therefore, the temporal information is limited to the five most recent frames. Moreover, Fast&Furious lacks a feature pyramid, which limits the influence of scene context.
[0072] DeepTracking is a recurrent convolutional network that uses a gated recurrent unit (GRU) to complete and forward predict occupancy and semantic maps based on partial evidence provided by LiDAR scans. However, DeepTracking does not merge information from these maps into object identities in space or time. Moreover, it performs forward prediction by continuing the recursive loop without data, which requires resetting the recursive state and feeding the history of recent frames into the network every time a prediction is to be created.
[0073] According to various embodiments, the aforementioned methods can be mixed into a recursive convolutional one-stage feature pyramid network that jointly performs detection and motion prediction on radar point cloud data to solve tracking tasks. As described in more detail below, multiple improvements have been made to the network details to improve systems used in the radar field. Leveraging the power of convolutional neural networks, classical methods can achieve remarkable results without the need for complex filtering methods or explicit motion models, especially for stationary targets.
[0074] Regarding recurrent layers and backbones, according to various embodiments, the ResNet backbone of RetinaNet can be enhanced by introducing a convolutional recurrent layer similar to DeepTracking as a separate network block before the main part of the backbone (e.g. Figure 1 ), or introduce the recursive layer into the residual block containing it (as shown in Figure 2 ), thereby creating a recursive residual block (e.g., as Figure 3 As shown or Figure 4 shown).
[0075] Figure 1 An example diagram 100 of a network architecture according to various embodiments is shown. A recursive layer stack 102 is followed by a residual backbone 104 and a feature pyramid 106. The feature pyramid 106 is followed by a regression head 108 and a classification head 110.
[0076] Figure 2 An example diagram 200 of a network architecture according to various embodiments is shown. The recursive residual backbone 202 is followed by a feature pyramid 106 (which may be Figure 1 The feature pyramid 106 is similar or identical to the feature pyramid 106 shown in FIG. 1 ). The feature pyramid 106 is followed by a regression head 108 and a classification head 110 (similar to FIG. 1 ). Figure 1 The network architecture shown is similar or identical).
[0077] Figure 2 The network architecture improves performance by allowing the network to incorporate temporal information at multiple scales (however, this may come at the expense of increased running time).
[0078] According to various embodiments, the network can be trained and operated such that the recursive loop must be executed only once per frame while preserving the hidden state between frames. This can provide that the network can learn to use information from arbitrarily far away points in time and that past sensor readings do not need to be buffered and stacked to operate the network.
[0079] Figure 3 Shown as Figure 2 An example diagram 300 of a recursive residual block design for a recursive residual backbone 202 is shown. The input data 302 to the recursive residual block can be a time series of two-dimensional multi-channel feature maps depicting a bird's-eye view of the scene (e.g., a tensor of shape "timesteps × height × width × channels", as it can be used for a recursive convolutional network). Since the backbone is composed of multiple of these blocks, a particular input to a block can be the output of a previous block, or in the case of the first block, the input to the backbone itself.
[0080] According to various embodiments, the input to the backbone itself can be a feature map generated directly from the radar point cloud by assigning each point detection to the nearest position in a two-dimensional grid. Individually for each radar sensor, the channels of the feature map can be populated with information such as the number of points assigned to that position and the average Doppler and RCS (radar cross section) values of the assigned points.
[0081] In an alternative embodiment, the input to the backbone itself can be a raw radar data cube, which can be converted into a suitable feature map via a dedicated sub-network. The use of data cubes as input and the associated sub-networks are described in detail as part of European patent application 20187674.5, the entire contents of which are incorporated herein for all purposes.
[0082] The input data 302 may be provided to a convolutional LSTM (Long Short Term Memory Network) 304 (e.g., with N filters and a kernel size of 3x3, where N may be an integer), followed by a convolution 306 (e.g., with N filters and a kernel size of 3x3), followed by a normalization 308. The input data 302 and the output of the normalization 308 may be provided to a processing block 310 (e.g., to an adder), and the result may be provided to a rectified linear unit (ReLU) 312.
[0083] Figure 4 Shown as Figure 2 An example diagram 400 of a recursive residual block design for the recursive residual backbone 202 is shown. Input data 302 may be provided to a convolution 402 (e.g., with N / 4 filters and a kernel size of 1x1), followed by normalization 404, followed by a rectified linear unit (ReLU) 406, followed by a convolutional LSTM 408 (e.g., with N / 4 filters and a kernel size of 3x3), followed by a convolution 410 (e.g., with N filters and a kernel size of 1x1), followed by normalization 308. The input data 302 and the output of the normalization 308 may be provided to a processing block 310 (e.g., to an adder), and the result may be provided to a rectified linear unit (ReLU) 312.
[0084] With respect to feature pyramids and network heads (e.g., regression and classification heads), the recursive backbone may be followed by feature pyramids and network heads for classification and regression, e.g., similar to RetinaNet. According to various embodiments, the regression head may be modified to perform motion prediction in the manner of Fast&Furious. Thus, the tracking capabilities of Fast&Furious may be combined with the benefits of multi-scale detection and scene context of RetinaNet, while temporal information is provided by the recursive backbone.
[0085] According to various embodiments, forward prediction is handled by a regression head (rather than a recursive loop as in DeepTracking). This way of operating the recursive layers makes it easy and efficient to generate motion predictions at each frame.
[0086] According to various implementations, the GRU layers may be replaced with long short-term memory networks (LSTMs), which may be better at modeling the temporal correlations that are critical in the radar domain.
[0087] According to various implementations, the dilated 3x3 kernel in the recurrent layer may be replaced with a regular 5x5 kernel to counteract the sparsity of radar data.
[0088] According to various embodiments, the nearest neighbor upsampling in the feature pyramid can be replaced with a transposed strided convolution to make the operation learnable, which can increase the richness of the features. The nearest neighbor upsampling can, for example, increase the spatial resolution by replicating pixel information in all directions at the desired upsampling rate. In an alternative approach, upsampling can be provided by a transposed strided convolution, which can distribute information from a lower resolution input map in a higher resolution output map according to a learned convolutional filter. This can provide greater freedom in how to arrange and combine information, and can allow the network to select an upsampling scheme that is more suitable for the type of features being learned on a per-channel basis, which can improve performance.
[0089] According to various embodiments, a custom anchor design using a bounding box prior over multiple rotations may be used, which may make objects easier to separate and rotation estimation more stable.
[0090] According to various embodiments, the focal loss can be modified so that it is normalized by the sum of focal weights (rather than by the number of foreground anchors), which can stabilize training for outlier frames.
[0091] According to various embodiments, a training scheme can be provided that divides the dataset into mini-sequences of data and preserves the hidden states between optimization steps, which can improve the network's ability to model temporal dependencies within the same memory constraint.
[0092] Figure 5 It shows how, according to various embodiments, Figure 1 and Figure 2 The regression head 108 and the classification head 110 are shown in Figure 500 to determine an example trajectory. Figure 5As shown in the left part 502 of , a predicted trajectory fragment (in other words: a predicted trajectory of only a limited or small number of time steps (e.g., three time steps, or, for example, five time steps)) can be determined, and based on multiple trajectory fragments (with different starting time steps), a trajectory (a trajectory of more than the small number of time steps) can be determined or decoded, as shown in Figure 5 As shown in the right portion 514 of .
[0093] For example, a first trajectory segment starting at t=0, a second trajectory segment starting at t=1, and a third trajectory segment starting at t=2 are illustrated in 502. As shown in 504, at time step t=0, only information from the first trajectory segment is available. At time step t=1 (506), information from the first trajectory segment and the second trajectory segment is available. At time step t=2 (508), information from the first trajectory segment, the second trajectory segment, and the third trajectory segment is available. At time step t=3 (510), information from the first trajectory segment is no longer available (because the first trajectory segment is only three time steps in length), but information from the second trajectory segment and the third trajectory segment is available. At time step t=4 (512), information from the second trajectory segment is no longer available (because the second trajectory segment is only three time steps in length), but information from the third trajectory segment is available.
[0094] Figure 6 An example diagram 600 of a method for predicting a trajectory of an object according to various embodiments is shown. At 602, radar data of an object can be obtained. At 604, first intermediate data can be determined based on the radar data based on a residual backbone using recursive components. At 606, second intermediate data can be determined based on the first intermediate data using a feature pyramid. At 608, a trajectory of the object can be predicted based on the second intermediate data.
[0095] According to various embodiments, the radar data of the object includes at least one of radar cube data and radar point data.
[0096] According to various embodiments, a regression head may be used to determine third intermediate data based on the second intermediate data, and a trajectory of the object may be predicted (further) based on the third intermediate data.
[0097] According to various embodiments, a classification head may be used to determine fourth intermediate data based on the second intermediate data, and a trajectory of the object may be predicted (further) based on the fourth intermediate data.
[0098] According to various embodiments, the step of predicting the trajectory may include determining trajectory segments for a plurality of time steps.
[0099] According to various embodiments, a residual backbone using recursive components may include or may be a residual backbone preceded by a stack of recursive layers.
[0100] According to various embodiments, the residual backbone using recursive components may include or may be: a recursive residual backbone including a plurality of recursive layers.
[0101] According to various embodiments, the plurality of recursive layers may include or may be: a convolutional long short-term memory network, followed by a convolution, followed by a normalization.
[0102] According to various embodiments, the plurality of recursive layers may include or may be: a convolution followed by normalization, followed by a rectified linear unit, followed by a convolutional long short-term memory network, followed by a convolution, followed by normalization.
[0103] According to various embodiments, the recursive component may include or may be a recursive loop that is executed once per time frame.
[0104] According to various implementations, the recursive component may remain hidden between time frames.
[0105] According to various implementations, the feature pyramid may include a transposed strided convolution.
[0106] According to various embodiments, training of the method may include partitioning the training dataset into multiple subsets and / or preserving hidden states between optimization steps.
[0107] Each of the above steps 602, 604, 606, 608 and further steps may be performed by computer hardware components.
[0108] In European patent application 20187694.3, which is incorporated herein in its entirety for all purposes and which may be referred to as Deep Multi-Object Tracker for Radar (DeepTracker for short), improvements to the method using a deep neural network that simultaneously performs object detection and short-term (several frames) motion prediction are disclosed. Motion prediction can be used to perform cross-frame correlation and aggregation of object information in a simple post-processing step, allowing efficient object tracking and temporal smoothing.
[0109] DeepTracker can significantly outperform traditional tracking methods, in part because DeepTracker does not rely on explicit motion models, but automatically learns appropriate motion models. It directly outputs forward predictions as a time series of fixed-size bounding boxes, which may mean the following:
[0110] 1) Forecasts are inherently time-discrete and have a limited forecast horizon.
[0111] 2) The space of possible predictions may have more degrees of freedom than any actual motion model. Therefore, the predictions may represent unstable trajectories, making it easy to overfit to noise and errors in the true values.
[0112] 3) Object velocity must be regressed as a separate output, or roughly estimated from a series of discrete position predictions.
[0113] One way to formulate a time-continuous network output is given in European Patent Application 19219051.0, which is incorporated herein in its entirety for all purposes and which may be referred to as continuous polynomial path prediction, where the network parameterizes a polynomial function describing the two-dimensional position of an object over time. This technique is used to perform long-term (e.g., several seconds long) motion prediction for each object in its own coordinate system based on existing detections.
[0114] According to various embodiments, DeepTracker is modified so that, instead of directly outputting a time series of bounding boxes, it provides coefficients to a time-continuous polynomial function that describes the temporal evolution of the bounding boxes in the form of a continuous polynomial path prediction. This may result in the following:
[0115] 1) The forecast is continuous without any specific restrictions on the forecast horizon, so it can be evaluated at any point in time.
[0116] 2) While the network is still not restricted to a specific motion model, it is constrained to produce temporally smooth predictions, thus excluding unrealistic trajectories.
[0117] 3) An estimate of the object's velocity can be easily obtained via differentiation of the polynomial describing the position.
[0118] Figure 7 An example diagram 700 of a Deeptracker using a polynomial representation in a training configuration according to various embodiments is shown. The regression head subnetwork 702 may output a position coefficient 704, a direction coefficient 716, and size information 722 (where the size may refer to the two-dimensional extent (length and width) of the bounding box of the object; the size is not represented as a time series or polynomial because it is not expected to change over time, unlike the position, direction, and velocity that change over time).
[0119] The position coefficients 704 may undergo polynomial evaluation 706 to obtain a position time series 708 , which may be an input to a regression loss 724 .
[0120] The position coefficients 704 may further undergo differentiation 710 (in order to obtain coefficients for velocity), and the differentiated coefficients may then undergo polynomial evaluation 712 to obtain a velocity time series 714 , which may also be an input to a regression loss 724 .
[0121] The directional coefficients 716 may undergo a polynomial evaluation 718 to obtain a directional time series 720 , which may be an input to a regression loss 724 .
[0122] The size information 722 may be another input to the regression loss 724 .
[0123] This change in representation can be computationally efficient and does not complicate network optimization. Since the evaluation of the polynomial function is differentiable, the specific bounding box can be calculated at a point in time when the true value information is available, and the end-to-end network can be trained using standard losses. The bounding boxes can be calculated from the polynomial function (e.g., via a simple polynomial evaluation). These bounding boxes can then be evaluated against the true value bounding boxes, resulting in an error loss term. The gradient of this error loss is propagated through the model to achieve learning. The back-propagation process of the gradient may require a differentiable operator.
[0124] Since a parameterized representation (e.g., using polynomials) may restrict the space of possible trajectories, it may have a strong regularizing effect on the model. The parameterized representation also further improves statistical efficiency for the following two reasons.
[0125] 1) The number of model parameters does not scale with the number of discrete time points in the prediction, but only with the chosen polynomial degree. For example, for a system running at 20Hz, the polynomial degree required to accurately represent a time series may be much lower than the number of frames in the time series.
[0126] 2) Both position and velocity true value information can provide gradient signals for the same basis function.
[0127] Many techniques dedicated to polynomial function optimization (especially in the context of joint detection and tracking networks) are developed or adapted from continuous polynomial path prediction to further improve performance. These techniques are described below.
[0128] These techniques can improve the performance of polynomial-based motion prediction and velocity estimation in neural networks. Therefore, these techniques can be applied to any neural network tracking system that incorporates polynomial-based motion prediction and velocity estimation.
[0129] To ensure favorable behavior of the polynomial function over the relevant time span, the prediction range of the true value during training can be selected to extend beyond the points that are later constructively used for tracking. According to various embodiments, this range can be extended in two ways, which means that the polynomial model can be trained to predict the past relative to the current frame to ensure that the polynomial performs well around its predictions for the present.
[0130] Since velocity is determined as the derivative of position, the same spatial units can be used. Therefore, the time unit of the polynomial can be carefully chosen so that position and velocity have similar absolute scales, allowing joint whitening of the target distribution, which is crucial for neural network optimization. The choice of time unit can depend on the overall standard deviation of position and velocity and is therefore dataset dependent. For example, the time unit can be about 1 / 10 second (or 2 frames). The scale of position and velocity can be determined empirically, and the time unit can be adjusted to equalize the scale of position and velocity so that their distributions can then be whitened without disturbing their relationship.
[0131] According to various embodiments, for each true value prediction during training, a random subset of all available prediction points within the prediction horizon may be selected for regularization purposes.
[0132] Continuous network output according to various embodiments can be applied to synchronization / fusion between multiple sensors. Sensors can transmit at slightly different points in time, which means that environmental scans obtained from these sensors for the same "frame" actually have a small time offset from each other. Therefore, the assumption made by many systems that all frame data is synchronized may introduce small but potentially destructive errors. In the case of time-continuous network output (e.g., when output is generated separately for each sensor), object information can be projected to an arbitrary point in time, allowing more accurate synchronization regardless of the temporally misaligned nature of the sensors.
[0133] Figure 8 A flowchart 800 illustrating a method for determining a parameterized plurality of parameters of a predicted trajectory of an object according to various embodiments is shown. At 802, radar data of an object can be acquired. At 804, first intermediate data can be determined based on the radar data based on a residual backbone using recursive components. At 806, second intermediate data can be determined based on the first intermediate data using a feature pyramid. At 808, a plurality of parameters can be determined based on the second intermediate data.
[0134] According to various embodiments, the parameterization may include or may be a polynomial of a predetermined degree; and the parameters may include or may be a plurality of coefficients associated with elements of a basis of a polynomial space of the polynomial of the predetermined degree.
[0135] According to various embodiments, elements of a basis may include or may be monomials.
[0136] According to various embodiments, the parameterization may include or may be: a first parameterization parameterizing the position trajectory of the object; and a second parameterization parameterizing the orientation of the object.
[0137] According to various embodiments, the radar data of the object may include or may be at least one of a radar data cube and radar point data.
[0138] According to various embodiments, the method may further include: determining third intermediate data based on the second intermediate data using a regression head; and determining the parameter based on the third intermediate data.
[0139] According to various embodiments, the residual backbone using recursive components may include or may be a residual backbone preceded by a recursive layer stack; the residual backbone using recursive components may include or may be a recursive residual backbone including a plurality of recursive layers.
[0140] According to various embodiments, the multiple recursive layers may include or may be: a convolutional long short-term memory network, followed by convolution, followed by normalization; and / or the multiple recursive layers may include or may be: a convolution, followed by normalization, followed by a rectified linear unit, followed by a convolutional long short-term memory network, followed by convolution, followed by normalization.
[0141] According to various embodiments, the recursive component may include or may be a recursive loop that is executed once per time frame.
[0142] According to various embodiments, the recursive component may remain hidden between time frames.
[0143] According to various embodiments, the feature pyramid may include or may be a transposed strided convolution.
[0144] According to various embodiments, training of the method may include partitioning the training dataset into multiple subsets and / or preserving hidden states between optimization steps.
[0145] Each of the above steps 802, 804, 806, 808 and further steps may be performed by computer hardware components.
[0146] Label list
[0147] 100 Example diagram of a network architecture according to various embodiments
[0148] 102 Recursive Layer Stack
[0149] 104 Residual backbone
[0150] 106 Feature Pyramid
[0151] 108 Return Head
[0152] 110 Classification Head
[0153] 200 Example diagram of a network architecture according to various embodiments
[0154] 202 Recursive Residual Backbone
[0155] 300 Figure 2 An example diagram of the recursive residual block design for the recursive residual backbone shown in
[0156] 302 Input Data
[0157] 304 Convolutional LSTM
[0158] 306 Convolution
[0159] 308 Normalization
[0160] 310 Processing Block
[0161] 312 Rectified Linear Unit
[0162] 400 Figure 2 An example diagram of the recursive residual block design for the recursive residual backbone shown in
[0163] 402 Convolution
[0164] 404 Normalization
[0165] 406 Rectified Linear Unit
[0166] 408 Convolutional LSTM
[0167] 410 Convolution
[0168] 500 is an example diagram of how to determine a trajectory according to various embodiments
[0169] 502 Left part
[0170] 504 Trajectory fragment at time step t=0
[0171] 506 Trajectory fragment at time step t=1
[0172] 508 Trajectory fragment at time step t=2
[0173] 510 Trajectory fragment at time step t=3
[0174] 512 Trajectory fragment at time step t=4
[0175] 514 Right part
[0176] 600 is an example diagram of a method for predicting a trajectory of an object according to various embodiments
[0177] 602 Steps to obtain radar data of an object
[0178] 604 A step of determining first intermediate data based on radar data based on a residual backbone using a recursive component
[0179] 606 Using the feature pyramid, determining the second intermediate data based on the first intermediate data
[0180] 608 A step of predicting the trajectory of the object based on the second intermediate data
[0181] 700 is an example diagram of DeepTracker using polynomial representation in a training configuration according to various embodiments
[0182] 702 Return to the leader network
[0183] 704 Position coefficient
[0184] 706 Polynomial Evaluation
[0185] 708 Position Time Series
[0186] 710 Differentiation
[0187] 712 Polynomial Evaluation
[0188] 714 Speed time series
[0189] 716 Directivity coefficient
[0190] 718 Polynomial Evaluation
[0191] 720 Direction Time Series
[0192] 722 Size Information
[0193] 724 Regression Loss
[0194] 800 is a flowchart illustrating a method for determining a plurality of parameterized parameters of a predicted trajectory of an object according to various embodiments.
[0195] 802 Steps to obtain radar data of an object
[0196] 804 A step of determining first intermediate data based on radar data based on a residual backbone using a recursive component
[0197] 806 Using the feature pyramid, the step of determining the second intermediate data based on the first intermediate data
[0198] 808 A step of determining a plurality of parameters based on the second intermediate data
Claims
1. A computer-implemented method for predicting a trajectory of an object, the computer-implemented method comprising the following steps performed by computer hardware components: acquiring radar data of the object; determining first intermediate data based on the radar data based on a residual backbone using recursive components, the residual backbone using recursive components providing time information; Using a feature pyramid, determining second intermediate data based on the first intermediate data; as well as predicting the trajectory of the object based on the second intermediate data using a regression head, wherein the regression head is modified to perform motion prediction using the temporal information, The step of predicting the trajectory of the object comprises: determining a plurality of parameterized parameters of the predicted trajectory of the object based on the second intermediate data, and wherein the parameterization comprises a polynomial of a predetermined degree, wherein the parameters comprise a plurality of coefficients associated with elements of a basis of a polynomial space of the polynomial of the predetermined degree, wherein the residual backbone using the recursive component comprises a residual backbone preceded by a recursive layer stack; and / or The residual backbone using the recursive component comprises a recursive residual backbone comprising a plurality of recursive layers.
2. The computer-implemented method of claim 1 , wherein: The parameterization includes: a first parameterization that parameterizes the position trajectory of the object; and a second parameterization that parameterizes the orientation of the object.
3. The computer-implemented method of claim 1 , wherein: The elements of the basis include monomials.
4. The computer-implemented method of claim 2, wherein: Provides speed estimates.
5. The computer-implemented method of claim 4, wherein: The velocity estimate is provided via differentiation.
6. The computer-implemented method of claim 1 , wherein: The radar data of the object includes at least one of a radar data cube and radar point data.
7. The computer-implemented method of claim 1, further comprising the following steps performed by the computer hardware components: using the regression head, determining third intermediate data based on the second intermediate data; and The trajectory of the object is predicted based on the third intermediate data.
8. The computer-implemented method of claim 1, further comprising the following steps performed by the computer hardware components: using a classification head, determining fourth intermediate data based on the second intermediate data; and The trajectory of the object is predicted based on the fourth intermediate data.
9. The computer-implemented method of claim 1, wherein: The step of predicting the trajectory includes determining trajectory segments for a plurality of time steps.
10. The computer-implemented method of claim 1, wherein: The plurality of recurrent layers comprises: a convolutional long short-term memory network, followed by a convolution, followed by a normalization; and / or The multiple recursive layers include: convolution, followed by normalization, followed by a rectified linear unit, followed by a convolutional long short-term memory network, followed by convolution, and followed by normalization.
11. The computer-implemented method of claim 1 , wherein: The recursive component comprises a recursive loop executed once per time frame; and / or Wherein, the recursive component remains hidden between time frames.
12. The computer-implemented method of claim 1, wherein: The feature pyramid includes transposed strided convolutions.
13. The computer-implemented method of claim 1, wherein: Training of the computer-implemented method may include partitioning the training data set into a plurality of subsets and / or preserving hidden states between optimization steps.
14. A computer system comprising a plurality of computer hardware components configured to perform the computer-implemented method according to any one of claims 1 to 13.
15. A vehicle comprising at least a subset of the computer system of claim 14.
16. A non-transitory computer-readable medium comprising instructions for executing the computer-implemented method according to any one of claims 1 to 13.
Citation Information
Patent Citations
SAR detection method and system based on deep convolutional network
CN111062321A
Employee dressing standard detection method based on improved RetinaNet
CN111401419A
Method and apparatus for tracking target from radar signal using artificial intelligence
US20200057141A1