Method for training neural network for determining features of object for object tracking

By training the neural network, using the sensor data sets in multiple training data elements, the eigenvectors of the object are determined, and the deviation between the eigenvectors is punished through the loss function, the global and persistent problems of object tracking features in the prior art are solved, and more accurate and efficient object tracking is achieved.

CN120202476APending Publication Date: 2025-06-24ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380078191.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-10
Filing Date
2023-10-27
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

In the prior art, the features used by object tracking are usually global and durable, limiting their flexibility and applicability in different scenarios, and affecting the accuracy and efficiency of object tracking.

Method used

By training the neural network, using sensor data sets in multiple training data elements, determine the eigenvectors of the object, and punish the deviations between the eigenvectors through the loss function, the neural network is trained to improve the accuracy of object tracking.

Benefits of technology

More accurate object tracking is achieved, reducing trace interrupts and object loss, and improving the overall performance of object tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120202476A_ABST
    Figure CN120202476A_ABST
Patent Text Reader

Abstract

According to various embodiments, there is provided a method for training a neural network for determining features of an object for object tracking, the method comprising: generating a training data set having a plurality of training data elements, wherein each training data element has a first sensor dataset relating to a first state of an ambient environment having a set of a plurality of objects and a second sensor dataset relating to a second state of the ambient environment wherein a position of the object at least partially changes with respect to the first state; determining a feature of the object by feeding the first sensor data set to a neural network; and determining a feature of the object by feeding the second sensor dataset to the neural network; determining loss according to the generated characteristics; and training the neural network for reducing loss.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a method for training a neural network to determine features of an object for object tracking. Background Art

[0002] Comprehensive recognition of the vehicle environment forms the basis for driver assistance systems and automated driving functions. Detection, classification, and tracking (Tracking) of objects such as other traffic participants are particularly important here. Nowadays, a large number of sensors are used to detect the vehicle environment. For example, cameras, radars, lidars (LiDAR), or ultrasonic systems belong to this.

[0003] Tracking of an object can be performed by features, so-called Re-ID features. However, the features are typically often determined such that the features are "global" and persistently distinct (not only within the scene where object tracking should be performed), which in turn limits the flexibility of their selection and thus affects their use for object tracking.

[0004] Therefore, a more effective way of feature extraction for object tracking is desirable. Summary of the Invention

[0005] According to various embodiments, there is provided a method for training a neural network to determine features of an object for object tracking, the method comprising

[0006] · Configuring a neural network with a network architecture such that when a sensor data set containing sensor data elements for each of a plurality of objects is fed to the neural network, the neural network determines a feature vector for each object based on all the sensor data elements;

[0007] · Generating a training data set having a plurality of training data elements, wherein each training data element has a first sensor data set related to a first state of the surrounding environment of a set of a plurality of objects and a second sensor data set related to a second state of the surrounding environment, wherein in the second state, the position of the object is at least partially changed relative to the first state;

[0008] · For each training data element,

[0009] o Determining the features of the object by feeding the first sensor data set to the neural network, and

[0010] o Determining the features of the object by feeding the second sensor data set to the neural network;

[0011] · Determine a loss that penalizes, for each training data element, the deviation between the features determined by a neural network for an object from a first sensor dataset and the features determined by the neural network for the object from a second sensor dataset, and penalizes the lack of deviation between the features determined by the neural network for an object from a first sensor dataset and the features determined by the neural network for each other object from the second sensor dataset; and

[0012] · Train the neural network to reduce the loss.

[0013] Then, the thus-trained neural network can be used for extracting (Re-ID) features (i.e., as a feature extractor), and the (Re-ID) features are then used within the scope of object tracking. The above-described manner (e.g., in the embodiments described hereinafter) enables an improved association from measurements to object tracking, which advantageously affects the overall performance of object tracking, i.e., objects can be tracked more accurately, there are fewer tracking interruptions, and fewer objects are lost (ignored).

[0014] According to the above-described manner, the use scenario context (especially other objects in the same scenario) is used for training (and, if necessary, for extracting Re-ID features). This allows scene-specific features to be used for the association from measurements to object tracking. In connection therewith, Re-ID features that are better suited for the object tracking task can be extracted, which in turn positively reflects on the overall performance of object tracking. Since the extraction of Re-ID features is tailored to the specific requirements of object tracking, the method can be implemented more computationally efficiently (e.g., with a smaller neural network). A network architecture from the class of network architectures suitable for extracting scene-specific Re-ID features is used.

[0015] Various embodiments are described hereinafter.

[0016] Embodiment 1 is a method for training a neural network for determining features of an object for object tracking, as described above.

[0017] Embodiment 2 is the method according to Embodiment 1, the method comprising: configuring the neural network with a network architecture such that when feeding sensor datasets each containing a sensor data element for each of a plurality of objects to the neural network, the neural network performs a feature vector by processing in a plurality of stages, wherein at least one stage generates a feature component for each of the objects, performs max pooling of the feature components on the objects, and feeds the feature components and the result of the max pooling of the feature components to the next stage.

[0018] Implemented in this way, the neural network determines a feature vector for each object based on all sensor data elements, i.e., the neural network considers other objects when determining a feature vector for an object, i.e., determines the feature vector based on the local context of the object (sensor data elements of a single object) and the global context (the result of max pooling on the object). Alternatively, a Transformer network can be used.

[0019] Embodiment 3 is the method according to Embodiment 1 or 2, wherein each training data element contains information about which object in the first state corresponds to which object in the second state, determines based on the determined features of the objects: which objects in the first state correspond to which objects in the second state, and determines the loss by comparing the information contained in the training data element about which object in the first state corresponds to which object in the second state with the result of determining which objects in the first state correspond to which objects in the second state based on the determined features.

[0020] Therefore, an association of the objects (e.g., in the form of an association matrix) can be set as the ground truth. Thus, the neural network is trained such that the neural network selects features that are particularly well-suited for object tracking. The result of determining which objects in the first state correspond to which objects in the second state based on the determined features is, for example, a soft association matrix.

[0021] Embodiment 4 is a method for tracking objects, the method comprising

[0022] · Detecting sensor data of a first set of objects at a first point in time;

[0023] · Detecting sensor data of a second set of objects at a second point in time;

[0024] · Feeding the sensor data of the first set of objects into a neural network trained according to any one of Embodiments 1 to 3 for generating a first feature;

[0025] · Feeding the sensor data of the second set of objects into the trained neural network for generating a second feature; and

[0026] · Associating the objects of the first set with the objects of the second set pairwise based on the first feature and the second feature.

[0027] Example 5 is the method according to Example 4, the method comprising: detecting sensor data of objects in a scene at a first time point; grouping the objects into a plurality of first clusters according to their spatial proximity; selecting one of the first clusters as a first object set; detecting sensor data of objects in the scene at a second time point; grouping the objects into a plurality of second clusters according to their spatial proximity; and selecting one of the second clusters as a second object set such that the second object set is the second cluster that is closest to the first object set within the scene.

[0028] Thus, the objects in the scene are regarded as (spatial) clusters. This makes it easy to distinguish features from one another, because instead of distinguishing all objects in the scene, only those objects within a cluster need to be distinguished.

[0029] Example 6 is a data processing device (such as a control device) which is configured to perform the method according to any one of Examples 1 to 5.

[0030] Example 7 is a computer program having instructions which, when executed by a processor, cause the processor to perform the method according to any one of Examples 1 to 5.

[0031] Example 8 is a computer-readable medium storing instructions which, when executed by a processor, cause the processor to perform the method according to any one of Examples 1 to 5. Description of the Drawings

[0032] In the drawings, like reference numerals generally refer to the same parts throughout the various views. These drawings are not necessarily to scale, with emphasis instead generally being placed upon illustrating the principles of the invention. In the following description, various aspects are described with reference to the following drawings.

[0033] Figure 1 A vehicle is shown.

[0034] Figure 2 Object tracking by means of Re-ID features is illustrated.

[0035] Figure 3 The association of the tracked object with the currently detected object is illustrated.

[0036] Figure 4 Feature extraction according to one embodiment is illustrated.

[0037] Figure 5 Feature extraction according to another embodiment is illustrated.

[0038] Figure 6 A possible architecture of a neural network for Re-ID feature extraction is shown.

[0039] Figure 7 Shows an example in which the trajectory of a pedestrian is shown, and the trajectory can be used to train a Re-ID feature extractor.

[0040] Figure 8 Shows a set for forming positive Re-ID feature pairs and negative Re-ID feature pairs, and based on the positive Re-ID feature pairs and negative Re-ID feature pairs, various losses (cost functions) can be calculated for training a Re-ID feature extractor.

[0041] Figure 9 Illustrates actions for training a Re-ID feature extractor, where a loss is calculated based on the calculated associations between measurement values at successive measurement time points.

[0042] Figure 10 Shows a flowchart that illustrates a method for training a neural network to determine features of an object for object tracking.

[0043] The following detailed description relates to the accompanying drawings, which illustrate specific details and aspects of the present disclosure for purposes of illustration, in which the invention may be implemented. Other aspects may be used and structural, logical, and electrical changes may be made without departing from the scope of the invention. The various aspects of the present disclosure are not necessarily mutually exclusive, as several aspects of the present disclosure may be combined with one or more other aspects of the present disclosure to form new aspects.

[0044] Various examples are described in more detail below. Detailed Description

[0045] Figure 1 Shows a vehicle 101 (e.g., autonomous).

[0046] In Figure 1 the example, the vehicle 101, such as a passenger car or a truck, is equipped with a vehicle control device 102.

[0047] The vehicle control device 102 has a data processing component, such as a processor (e.g., a CPU (central unit)) 103 and a memory 104 for storing control software according to which the vehicle control device 102 operates and data processed by the processor 103.

[0048] For example, the stored control software (computer program) has instructions that, when executed by the processor, cause the processor 103 to implement one or more neural networks 107.

[0049] The data stored in the memory 104 may, for example, include image data detected by one or more cameras 105. For example, the one or more cameras 105 may record one or more grayscale or color photos of the surroundings of the vehicle 101.

[0050] The vehicle control device 102 may determine based on the image data whether there are and which objects are present in the surroundings of the vehicle 101, such as stationary objects like traffic signs or road markings, or movable objects like pedestrians, animals, and other vehicles.

[0051] Here, the image data is only used as an example of sensor data, and data from other sensors (ultrasound, lidar, etc.) may also be used.

[0052] The vehicle 101 can then be controlled by the vehicle control device 102 based on the result of the object determination. Thus, for example, the vehicle control device 102 can control the actuator 106 (such as a brake) in order to control the speed of the vehicle, for example in order to brake the vehicle.

[0053] Typically, it is desirable for this purpose not only to detect but also to track objects, for example in order to predict the later position of the object by observing how the object has moved in the past and in order to obtain a consistent understanding of the current (traffic) situation.

[0054] For example, the tracking of objects is implemented using a recursive filter (such as a Kalman filter), which performs three processing steps: prediction, association, and update. During prediction, first, the existing objects are predicted for the current measurement time point with the help of a motion model. Subsequently, the current measurement (such as object detection in a camera image) is associated with the predicted objects. Finally, in the update step, the predicted objects are corrected with the help of the associated measurements. In particular, the association from the measurement to object tracking is of great significance because incorrect associations lead to inaccurate object tracking or even loss of object tracking. A distance metric forms the basis for this association, which maps the proximity measured in the feature space to the objects being tracked. The metrics used are, for example, the Euclidean distance or the Mahalanobis distance. Typically, geometric features such as object position, object size, or object speed are used for distance calculation. Subsequently, based on the calculated distances, the measurement values are assigned to the objects being tracked with the help of an association algorithm. Examples of such algorithms are the Hungarian algorithm or various manifestations of the nearest neighbor algorithm. Problems occur mainly when the measurements and objects cannot be clearly associated with each other in these dimensions (e.g., rows of people standing closely together, inaccurate predictions due to long occlusions, etc.) due to the decisive influence of geometric features. The problems mentioned during association can be alleviated or completely prevented by adding suitable features. In this case, a feature is considered suitable if it supports a clear association from the measurement to the object. This can be, for example, a view-based feature that describes the appearance of the object (example: if only one pedestrian in a row of people wears a red jacket, then the red jacket is a good feature for tracking that pedestrian).

[0055] The extraction of so-called re-identification features (abbreviated as Re-ID features) can be applied to support object tracking, for example, in surveillance systems for train stations, airports, or pedestrian areas, so that people can be re-identified across different cameras even over a long period of time.

[0056] Figure 2 Illustrate object tracking with Re-ID features.

[0057] To this end, before associating 204 with the object 205 to be tracked, the measurement value 203 (which is determined from sensor data, such as camera data 201, through preprocessing 202) is enriched with the Re-ID features obtained from feature extraction 206 into an extended measurement value 207, so that subsequent association 204 (and in particular distance calculation) can be performed not only based on geometric features but also based on Re-ID features. In this case, the currently detected object (corresponding to the measurement value 203) is associated with the prediction 208 of the object 205 to be tracked (and subsequently the update 209 of the tracked object 205 is performed). This provides tracking information 210 (e.g., in the form of a trajectory) about the detected object as output.

[0058] Figure 3 It is illustrated that the object 301 to be tracked is associated with the currently detected ("measured") object 302 (from the measurement value 305, e.g., as a detection of an object bounding box) by distance calculation 303 with the aid of an association algorithm 304. The distance calculation can be based on typically used geometric features (position, size, speed, etc.) and / or Re-ID features.

[0059] Re-ID features are typically extracted with the aid of a deep neural network (e.g., one of the neural networks 107). The basis for training the corresponding network is a large annotated data set that contains many objects (e.g., people) with multiple views for each object. The aim of this training is to find suitable weights for the neural network such that a feature space suitable for object re-identification is created (the neural network maps the input (e.g., a view of the object) into this feature space), i.e., different views of the same object (e.g., the same person) should be mapped to features (feature vectors) that have a small distance in this feature space, while the distance to the feature vectors to which views of other objects are mapped should be maximized.

[0060] The feature mapping (feature extraction) trained in this way is suitable for re-identifying objects from a large database and can therefore also contribute to the association during tracking.

[0061] Specifically, the Re-ID features used during object tracking should have the following properties:

[0062] · Scene-specific features: The features should be applicable to distinguish an object from other objects within the same scene. For this it is not necessary that an object can be distinguished from all (also other-scene) objects. However, typical methods for extracting Re-ID features only achieve the latter. Referring to the foregoing example: If a person is the only one wearing a red jacket in a scene, then the red jacket is a good feature for distinguishing that person from the other people in that scene. However, the red jacket is definitely not a good feature for distinguishing a person from all other people in the world. This means that different features may be suitable for tracking applications than in the case of classical Re-ID applications. Additionally, the features for the same person may vary scene-specifically. Depending on the scene context, the feature may be a red jacket in one case and later an umbrella being carried (e.g., if a second person with a red jacket walks into the scene).

[0063] · Consistency over a short time period: When tracking an object, the object must be re-identified only over a small time range. Typically only from one measurement time point to the next (a few milliseconds). In the case of occlusion, it may be necessary to re-identify the object within a few seconds. So it is only necessary that the Re-ID features for the object are almost constant over this small time period, but not constant over minutes, hours, or days. For example, a person must only be re-identified within a scene and distinguished from other people. If the same person appears again in other scenes, it is not necessary to know that this is the same person.

[0064] In terms of the above-described desired characteristics of feature extraction (i.e., requirements for feature extraction), according to various embodiments, a way is provided which can achieve improved object tracking and thus improved object tracing by improving the association from measurements (e.g., object detection) to object tracking (i.e., the object being tracked). The improvement in object tracking is reflected in more accurate object trajectories, fewer tracking interruptions, and fewer object losses.

[0065] Re-ID features are used for improving the association during object tracking (e.g., as described in reference Figures 1 to 3 ). However, the training and extraction of Re-ID features are directed to specific requirements for feature extraction (as described above). In particular, according to various embodiments, a method for extracting Re-ID features during tracking, a specific network architecture for a Re-ID feature extractor, and a method for training a suitable Re-ID feature extractor are provided.

[0066] Common methods for extracting Re-ID features train corresponding deep neural networks based on large annotated datasets for the purpose of being able to delimit different views of an object (e.g., a person) from views of other objects (e.g., all other people). The trained neural network is then applied separately to each detected object in order to generate Re-ID features for that object, i.e., if, for example, five people are detected in a scene, the network is implemented five times independently of each other (once for each detection).

[0067] In contrast, according to various embodiments, a scheme as set forth in Figure 4 is provided.

[0068] Figure 4 Illustrate feature extraction according to one embodiment.

[0069] Here, a Re-ID feature extractor 401 (neural network) is utilized, which takes multiple measurements (e.g., detections) 402 as input and generates the corresponding Re-ID feature vectors 403 as output for the multiple measurements 402. By this action, the feature extractor 401 is enabled to extract scene-specific features - i.e., features that are particularly well-suited for distinguishing or re-identifying the measurements of that scene (in this example for distinguishing the five people).

[0070] This implies that the scene context plays an important role in the extraction of features. Referring to the above example of the red jacket: If only one person in the scene wears a red jacket, then the red jacket will be a good feature for re-identification. If there are multiple people with red jackets in the scene, the feature extractor will consider other features for distinguishing or re-identifying people (e.g., hair color, carried umbrella, etc.).

[0071] Figure 5 Illustrate feature extraction according to another embodiment.

[0072] According to Figure 5 's embodiment, first, the measurements of the scene 501 are grouped (clustered) into clusters 502, 503, and then feature extraction 504, 505 is performed for each cluster 502, 503. This has the advantage that Re-ID feature extraction can focus on the measurements where there is a high risk of false association (e.g., a group of pedestrians standing closely together). Referring to the above example of the red jacket: Even if there are two people with red jackets in a scene, for example, if these two people with red jackets are far from each other (e.g., on different sides of a traffic lane) and thus belong to different clusters 502, 503, the red jacket can also be a suitable feature for distinguishing from and re-identifying (and thus re-identifying) other people.

[0073] According to another embodiment, Re-ID feature extraction takes into account the weighting of measurement values. Specifically, the weights should reflect how strongly each measurement value can be distinguished from other measurement values. Re-ID feature extraction can then focus on the extraction of suitable distinguishing features. For example, weights can be given pairwise (among all measurement values) such that each weight indicates how important the distinction between two measurement values is.

[0074] For example, the weights reflect the separability based on geometric features (such as position, extent, velocity), i.e., poorly separable measurement values receive high weights. Refer to Figure 5 What is described above can be regarded as (or implemented as) a special case of such weighting: the pairwise weights are 1 for all measurement value pairs of the same cluster and 0 otherwise.

[0075] According to various embodiments, an architecture is used for a neural network for Re-ID feature extraction (the neural network should be applied in the manner mentioned above), and the architecture satisfies the following properties:

[0076] · The input can include any number N of measurement values in the form of N input data vectors (the input data vectors can also be regarded as input feature vectors). For example, the measurement values can be the detection of a pedestrian, and its input vector includes, for example, the original sensor data (image pixels, lidar point clouds, radar reflections, etc.) belonging thereto or parameters derived therefrom.

[0077] · The N measurement values can be input in any order, i.e., the calculation of Re-ID features is independent of the order.

[0078] · The output of the network is N Re-ID feature vectors, where N corresponds to the number of input measurement values.

[0079] Figure 6 Shows a possible architecture of a neural network for Re-ID feature extraction.

[0080] The input to the network is measurement value 601 (e.g., pedestrian detection). The measurement values are unordered (since the order should be irrelevant for the extracted features). Each measurement value is characterized by specific features in the form of an input data vector (e.g., the original sensor data for detection, such as image pixels, lidar points, radar reflections, or data derived therefrom). The output 605 of the network corresponds to the calculated Re-ID feature vectors (i.e., one feature vector for each measurement value of input 601).

[0081] The first layer 602 processes the data of each measurement independently by means of a measurement-based fully gridded layer (1D-conv). This layer is characterized by a weight matrix M and a bias vector b. To reach the first hidden layer from the input representation (measurement data), the same weight matrix M and the same bias vector b are used for each measurement (so-called weight-sharing). The output of such a 1D-conv layer is a new representation of the measurement-based data (measurement-specific local feature vectors). Thus, the 1D-conv layer operates locally in such a way that it only combines the features within one measurement with each other (local context). In this example, the 1D-conv layer extracts individual features for each pedestrian (e.g., regarding its shape, color, etc.).

[0082] Following is the pooling layer 603, which reduces the measurement-based features to unique (so-called global) features. In the example shown, the maximum value of the features of the first hidden layer is calculated (global max pooling). By such a pooling layer, information is aggregated across multiple measurements, thereby enabling the mapping of the correlations between the measurements (global context), i.e., the pooling layer 603 allows the features previously extracted individually for each pedestrian to be combined with each other and made relevant.

[0083] To utilize this global relationship between the features in further calculations, the global feature vector can be appended to the measurement-specific local feature vector. Thus, the subsequent 1D-conv layer 604 (similar to the first layer 602, together with the respective max pooling layer 603 if necessary) can also utilize the global context in its measurement-based feature extraction. In this example: the features that distinguish the individual pedestrians can be emphasized.

[0084] As shown in Figure 6 it is possible to arrange multiple 1D-conv + pooling + Append processing blocks in a row in order to finally extract a complex Re-ID feature vector 605.

[0085] In Figure 6 the architecture shown is a specific implementation, and this implementation can be modified in various ways. The general basic architecture can be described as follows:

[0086] · An unordered list of measurements is used as input. The output of the network architecture is a list of Re-ID feature vectors (one output vector for each input measurement).

[0087] · The architecture utilizes any number of layers on which (local) measurement-based feature extraction is performed. In one design, weight-sharing is implemented during measurement-based feature extraction.

[0088] · The architecture includes at least one global pooling layer, which combines the local features extracted by the measurement expressions into a global feature vector. Additionally, depending on the different measurement-expression feature extraction layers, other pooling operators can be used. In one implementation, the pooling layer uses the max pooling operator. However, other pooling operators (such as average pooling) are also conceivable.

[0089] · The architecture uses at least one append module, which merges the global (pooled) feature vector with the local (measurement-expression) feature vector.

[0090] For the described basic architecture, various extensions are possible, such as:

[0091] · Global feature vector: The pooled feature vector can be arbitrarily further processed before being spliced to the local feature vector, for example, by means of a fully convolutional layer.

[0092] · Order: The (local or global) features from earlier layers can also be directly used in later layers, for example, by means of so-called skip connections.

[0093] The training of the Re-ID feature extractor is based on a dataset in which the trajectories of objects are annotated. This means that the objects are annotated at individual measurement time points, and there is an association between these measurement time points (for example, by means of object IDs).

[0094] Figure 7 An example is shown in which the trajectories 701 to 705 of pedestrians are shown, and these trajectories can be used to train the Re-ID feature extractor.

[0095] The pedestrians marked with a '+' are involved in the same sensor measurement time point. Additionally, the pedestrian annotations at earlier and later measurement time points are shown.

[0096] Figure 8 A set for forming positive Re-ID feature pairs 801 and negative Re-ID feature pairs 802 is shown. Based on these positive Re-ID feature pairs 801 and negative Re-ID feature pairs 802, various losses (cost functions) for training the Re-ID feature extractor 803 can be calculated.

[0097] As shown in Figure 8 For training as stipulated in various embodiments, the Re-ID feature extraction is applied to at least two successive measurement time points, and a set of positive pairs 801 and a set of negative pairs 802 are formed from the resulting Re-ID feature vectors.

[0098] Positive pairs respectively correspond to two feature vectors that describe the same object (e.g., a pedestrian) at different measurement time points. Negative pairs respectively correspond to two feature vectors that describe different objects (e.g., two different pedestrians) at the same or different measurement time points. Whether the feature vectors describe different objects or the same object is known for the training data of the training data set (through the corresponding annotations of the input data vectors included in the training data set).

[0099] If this operation is applied to multiple short sequences (i.e., successive measurement time points) of the annotated training data set, a large number of positive and negative pairs can be generated.

[0100] During training, the weights of the neural network are adapted in the following way: reduce (ideally minimize) the distance between the feature vectors of positive pairs, while increase (ideally maximize) the distance between the feature vectors of negative pairs. This can be implemented, for example, by backpropagation and gradient descent of a suitable loss (cost function). For example, contrastive loss, triplet loss, multi-class N-pair loss, or constellation loss (Konstellationsverlust) is used as the cost function.

[0101] Therefore, for example, the neural network is trained as follows:

[0102] 1) Generate multiple sequences consisting of two or more measurement time points

[0103] 2) Calculate the Re-ID feature vectors by applying the neural network

[0104] 3) Generate positive and negative pairs of the calculated Re-ID feature vectors

[0105] 4) Calculate the loss

[0106] 5) Adapt the network weights by backpropagation and gradient descent of the loss.

[0107] The above training steps 1) to 5) can be repeated any number of times.

[0108] According to another embodiment, it is stipulated that the weights of the neural network are adapted based on the loss that reflects the quality of the association between the measurement values at two successive time steps.

[0109] Figure 9 Illustrate the operation for training the Re-ID feature extractor, where the loss is calculated based on the calculated association between the measurement values at successive measurement time points.

[0110] To this end, first, Re-ID feature vectors 901, 902 of measurement values are extracted for two successive time steps. Subsequently, a distance matrix 903 is calculated, and based on this, a soft association matrix 905 is calculated by a differentiable association module 904 (such as a Deep Hungarian Network). In this case, "soft" means that its entries (association values) approximate the binary entries (i.e., 0 or 1) of the true association matrix. Finally, a loss 907 can be calculated by comparing the soft association matrix with the true association matrix (the true association matrix is known through the labels of the training dataset, i.e., the annotations), and based on this loss, the network weights can be adjusted by backpropagation and gradient descent 908.

[0111] It should be noted that if different features are assigned to the same object or different features are assigned to different objects, the loss also penalizes this because then the soft association matrix will accordingly have small association values for the same object or high association values for different objects.

[0112] In summary, according to various embodiments, a method as shown in Figure 10 is provided.

[0113] Figure 10 FIG. 1000 shows a flowchart that illustrates a method for training a neural network for determining features of objects for object tracking.

[0114] In 1001, a neural network is configured with a network architecture (i.e., a neural network with such an architecture is provided) such that when a sensor dataset containing sensor data elements for each of a plurality of objects is fed to the neural network, the neural network determines a feature vector for each object based on all the sensor data elements.

[0115] In 1002, a training dataset with a plurality of training data elements is generated, where each training data element has a first sensor dataset related to a first state of the surroundings of a set of a plurality of objects and a second sensor dataset related to a second state of the surroundings, where in the second state, the positions of the objects are at least partially changed relative to the first state.

[0116] In 1003, for each training data element, the features of the objects are determined by feeding the first sensor dataset to the neural network, and the features of the objects are determined by feeding the second sensor dataset to the neural network.

[0117] In 1004, a loss is determined that penalizes, for each training data element and for each pair of objects, the deviation between the features determined by the neural network for the object from the first sensor dataset and the features determined by the neural network for the object from the second sensor dataset, and penalizes the deviation of the lack between the features determined by the neural network for the object from the first sensor dataset and the features determined by the neural network for each other object from the second sensor dataset for that other object.

[0118] In 1005, the neural network is trained to reduce the loss.

[0119] In this case, 1003, 1004, and 1005 can be repeatedly performed alternately, for example, to determine the loss for batches of training data elements (i.e., the training dataset can have batches, each batch having multiple training data elements of the described form, generating losses for the training data elements respectively, and training the neural network for each batch to reduce the corresponding loss).

[0120] Figure 10 The method can be executed by one or more computers having one or more data processing units. The term "data processing unit" can be understood as any type of entity capable of implementing the processing of data or signals. The data or signals can be processed, for example, according to at least one (i.e., one or more than one) special function, which is executed by the data processing unit. The data processing unit can include an integrated circuit of analog circuits, digital circuits, logic circuits, microprocessors, microcontrollers, central units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), field-programmable gate arrays (FPGAs), or any combination thereof, or can be constructed by these. Any other way of implementing the corresponding functions described in more detail herein can also be understood as a data processing unit or a logic circuit device. One or more of the method steps described in detail here can be implemented (e.g., realized) by the data processing unit through one or more special functions, which are executed by the data processing unit.

[0121] Thus, according to various embodiments, the method is in particular computer-implemented.

[0122] After training, the neural network can be applied to sensor data determined by at least one sensor to obtain features, which are then used for object tracking. The result of the object tracking can then be used to control the robotic device.

[0123] For example, after training, a neural network is used to generate control signals for a robotic device by feeding sensor data related to the robotic device and / or its surrounding environment to the neural network. The term "robotic device" can be understood to refer to any technical system (with mechanical parts whose movement is controlled), such as a computer-controlled machine, vehicle, household appliance, power tool, manufacturing machine, personal assistant, or access control system.

[0124] Various embodiments can receive and use sensor data from various sensors such as, for example, video, radar, lidar, ultrasound, motion, thermal imaging, etc.

[0125] Although specific embodiments have been shown and described herein, it will be recognized by those skilled in the art that the specific embodiments shown and described may be replaced by a variety of alternative and / or equivalent implementations without departing from the scope of the present invention. This application is intended to cover any adaptations or variations of the specific embodiments discussed herein. It is thus intended that the present invention be limited only by the claims and their equivalents.

Claims

1. A method for training a neural network to determine features of an object for object tracking, the method comprising: Configuring a neural network with a network architecture such that when a sensor data set containing sensor data elements for each of a plurality of objects is fed to the neural network, the neural network determines a feature vector for each object based on all the sensor data elements; Generating a training data set having a plurality of training data elements, wherein each training data element has a first sensor data set related to a first state of the surroundings of a set of a plurality of objects and a second sensor data set related to a second state of the surroundings, wherein in the second state, the position of the object is at least partially changed relative to the first state; For each training data element, Determining the features of the object by feeding the first sensor data set to the neural network, and Determining the features of the object by feeding the second sensor data set to the neural network; Determining a loss that penalizes, for each training data element, for each pair of objects, the deviation between the features determined by the neural network for the object from the first sensor data set and the features determined by the neural network for the object from the second sensor data set, and penalizes the insufficient deviation between the features determined by the neural network for the object from the first sensor data set and the features determined by the neural network for each other object from the second sensor data set; and Training the neural network to reduce the loss.

2. The method according to claim 1, the method comprising: Configuring the neural network with a network architecture such that when a sensor data set containing sensor data elements for each of a plurality of objects is fed to the neural network, the neural network performs the feature vector by processing in multiple stages, wherein at least one stage generates feature components for each of the objects, performs max pooling on the feature components on the objects, and feeds the feature components and the results of the max pooling of the feature components to the next stage.

3. The method according to claim 1 or 2, wherein each training data element contains information about which object in the first state corresponds to which object in the second state, determines, based on the determined features of the objects, which objects in the first state correspond to which objects in the second state, and determines the loss by comparing the information about which object in the first state corresponds to which object in the second state contained in the training data element with the result of determining which objects in the first state correspond to which objects in the second state based on the determined features.

4. A method for tracking an object, the method comprising Detecting sensor data of a first set of objects at a first time point; Detecting sensor data of a second set of objects at a second time point; Feed sensor data of the first set of objects into a neural network trained according to any one of claims 1 to 3 to generate first features; Feed sensor data of the second set of objects into the trained neural network to generate second features; and Associate the objects of the first set with the objects of the second set in pairs based on the first features and the second features.

5. The method according to claim 4, the method comprising: Detect sensor data of the objects in the scene at the first time point; Group the objects into a plurality of first clusters according to their spatial proximity, and select one of the first clusters as the first set of objects; Detect sensor data of the objects in the scene at the second time point; Group the objects into a plurality of second clusters according to their spatial proximity; and select one of the second clusters as the second set of objects such that the second set of objects is the second cluster that is closest to the first set of objects within the scene.

6. A data processing apparatus, the data processing apparatus being configured to perform the method according to any one of claims 1 to 5.

7. A computer program having instructions that, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 5.

8. A computer-readable medium storing instructions that, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 5.