Method for object detection from point cloud data by means of transformer having attention model
By using the transformer of the attention model and the backbone network to calculate feature vectors in object detection, combined with the refinement technology of cross attention and anchor point position, the problem of accuracy and waste of computing resources in large-point cloud data is solved, and more accurate and efficient object detection is achieved.
Patent Information
- Application Number
- CN202380068824.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-28
- Filing Date
- 2023-09-12
- Publication Date
- 2025-05-06
AI Technical Summary
Existing object detection technology is difficult to effectively detect multiple objects when processing large point cloud data, especially in sparsely distributed point cloud environments, where traditional methods have the problem of wasted computing resources.
Using a transformer with an attention model, the feature vector is calculated through the backbone network, and the object query is refined using the cross attention mechanism and anchor position, and the anchor position is adjusted to obtain more accurate object detection.
It realizes more precise detection of multiple objects in large point cloud data, reduces the distance between the object query location and the actual object, improves detection accuracy, and saves computing resources in sparse point cloud environments.
Smart Images

Figure CN119948538A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a method for detecting a plurality of objects from point cloud data by means of transformers with an attention model. Background Art
[0002] Object detection is now performed in imaging sensors. There are usually multiple objects in the recorded environment, so that detection is performed on multiple objects. For example, object detection is used in vehicle sensors to detect other vehicles, other traffic participants and infrastructure. This data can be used for (partially) automated or autonomous driving.
[0003] A recently pursued concept is to apply transformers to object detection. Transformers are described in the paper "Attention is all you need" (arXiv preprint arXiv: 1706.03762, 2017) by Ashish Vaswani et al., first in the context of language processing. In object detection, a bounding box (Bounding-Boxen) describing the object and its box parameters (Box-Parameter) are obtained for each object based on measurements, i.e., for example, the position, size, orientation, speed and / or category identifier of the object. Transformers can also be used for downstream applications, such as object tracking, prediction or (path) planning. When using transformers for object detection, the suppression of overlapping detections applied in post-processing in a traditional manner can be ignored. So far, such transformers have been applied to image data, for example. On the contrary, the use of large point clouds, such as those that appear in the context of autonomous and automated driving, is not known. Summary of the invention
[0004] The invention relates to a method for detecting a plurality of objects from point cloud data by means of a transformer with an attention model. The point cloud data are detected, for example, by a lidar. However, the method is not limited to lidar, but other sensor types can also be used. The sensor or the sensor system is preferably arranged on a vehicle so that the point cloud data can be recorded from the vehicle.
[0005] The method comprises the following steps: First, a feature vector is calculated from the point cloud data. This is not performed by the encoder of the transformer as usual, but by the backbone network. The backbone network is a neural network that is used to extract features from the measured data or to convert the input into a specific feature representation that can be processed further. Therefore, the encoder of the transformer can be omitted. Preferably, the output of the backbone network is reformatted to obtain a sequence of feature vectors with a preset length. By using the backbone network to calculate the feature vector, the input sequence in self-attention is less restricted than using the encoder of the transformer, and on the contrary, a sufficiently small cell size can be selected in a grid-based backbone network, such as in point pillars. The feature vector thus calculated is then fed to the transformer and used as a key vector and value vector for obtaining cross-attention.
[0006] In addition, the first anchor position for the first layer of the transformer is calculated from the point cloud data by a sampling method such as farthest point sampling (FPS). A feature vector is obtained based on the first anchor position by means of encoding, such as Fourier encoding. The encoding can be done in particular by a feed-forward network. The feature vector thus calculated is used as an object query for the first layer of the decoder of the transformer. The object query of the anchor position is used as a starting point for searching for objects. However, the search is not limited to these anchor positions, but also detects objects that are at a distance from these anchor positions. The anchor position does not correspond to the anchor box used in other detection schemes. Therefore, the object query for the transformer depends on the data and is not learned as usual. This provides advantages in particular in the case of sparsely distributed point clouds, because otherwise a lot of computing resources will be wasted for finding positions that actually have data. Such sparsely distributed point clouds appear in particular in measurements using lidar. The object queries obtained based on the anchor position are used as slots for possible objects.
[0007] In the first layer, the decoder of the transformer obtains result feature vectors based on the object query (ie, the feature vector mentioned above) and the key vector and the value vector (ie, the feature vector described at the beginning), which are also called decoder output vectors.
[0008] The box parameters of the bounding box describing the object, such as the position of the object or the position difference relative to the anchor point position, size, orientation, speed and / or class identifier, are calculated by means of a feedforward network based on the resulting feature vector. For this purpose, a feedforward network different from the feedforward network described above for determining the object query is preferably used, which feedforward network differs in terms of weights.
[0009] The anchor position is then adjusted with the aid of the obtained frame parameters for processing at least one further layer of the transformer. When adjusting the anchor position, the position difference of the frame parameters calculated based on the result feature vector of the first layer of the transformer is added to the first anchor position. In general, frame parameters that are far from the first anchor position and thus have a large position difference can be obtained for the result feature vector of the first layer. By adjusting the anchor position, an adjusted anchor position that is closer to the actual object can be obtained. With the aid of the encoding as described above, feature vectors are obtained based on the adjusted anchor position, and these feature vectors are used as object queries for at least one further layer of the transformer.
[0010] In order to propagate the information of the high-dimensional result feature vector of the first layer in addition to the adjusted anchor position, a transformation of the result feature vector of the first layer is performed with respect to the adjusted anchor position. Here, the result feature vector is aligned with the adjusted anchor position. Advantageously, this is achieved by a feed-forward network consisting of two layers with a ReLU (Rectified Linear Unit) activation function. This results in low additional overhead, since only a feed-forward network with two layers is used here.
[0011] The above-mentioned steps of adjusting the anchor point position, obtaining the feature vector based on the adjusted anchor point position, and transforming the resulting feature vector are referred to herein as refinement of the object query.
[0012] The transformed result feature vector and the calculated object query, in particular their vector sum, are now fed to the decoder of the transformer as input for at least one further layer and used there as slots for possible objects. The decoder of the transformer determines the result feature vector in at least one further layer based on the transformed result feature vector determined for the preceding layer, based on the calculated object query determined as described above based on the adjusted anchor point positions, and based on the key vector and value vector described at the outset.
[0013] Therefore, feature vectors of at least one further layer and thus also bounding boxes and finally objects found in at least one further layer are determined based on the refined object query of the adjusted anchor point position of the first layer. In this case, the position of the refined object query is usually closer to the actual object than the position of the original object query. It is applicable that the distance between the position of the object query (based on which the detection is performed) and the actual object affects the accuracy of the detection in the corresponding layer. By adjusting the position of the refined object query based on the previous box parameters, the distance between the position of the (refined) object query and the actual object is reduced, thereby achieving a more accurate detection.
[0014] By transforming the resulting feature vectors with respect to the adjusted anchor point positions, these can continue to be used as object queries for evaluation in subsequent layers. The form of the resulting feature vectors does not change here, so that known encoding methods can be used. In particular, when using the above-mentioned feedforward network with only two layers for this purpose, the transformation can be performed with low additional overhead. In addition, the same encoding is used for the anchor point positions as for the first layer, so that no additional parameters need to be used.
[0015] Furthermore, the resulting feature vectors are position-dependent vectors that, when processed by the decoder and adjusted according to the anchor positions, gradually gain more information about the object. Object information is encoded in the latent feature space and not just in low-dimensional box parameters as in the traditional way. In another step, such vectors can then be propagated in time and used, for example, for object tracking and prediction. Thus, the transformer can be used for further downstream applications that presuppose object recognition and work with large point clouds.
[0016] In particular, a significant reduction in the distance is achieved in the first refinement of the object query, so that a refinement of the object query performed only between the first and second layers of the transformer already has a large effect. Preferably, the following steps are performed for at least one further layer of the transformer besides the first layer: calculating frame parameters for the resulting feature vector, adjusting the anchor point positions, determining the feature vector based on the adjusted anchor point positions by means of encoding, and transforming the resulting feature vector with respect to the adjusted anchor point positions, wherein the further layer is used in the above steps instead of the first layer.
[0017] Here, the term "first layer" is to be understood as the first layer of a converter to which the method is applied. Although it is advantageous to apply the method directly to the first layer of a converter, it is also conceivable to first apply the method to a subsequent layer. In this case, the subsequent layer is interpreted as the "first" layer.
[0018] In order to train the transformer or the model of the transformer, the following steps are preferably performed: multiple frame parameter sets are obtained for the decoder output of each layer, preferably as many frame parameter sets as the object query at the input of the decoder are specified. In addition, the frame parameters of the reference truth (English: Ground Truth) are provided, and the frame parameters of the reference truth are assigned to the closest estimated frame parameters. For this purpose, the Hungarian method (die Ungarische Methode) is preferably applied. Unmatched frame parameters are assigned to the "non-object" category and are discarded. Median regression (also known as l1-loss) is applied to the deviation between the frame parameters of the reference truth and the assigned estimated frame parameters. Finally, the transformer is trained with the help of median regression.
[0019] The training of the result feature vectors of the transformation about the adjusted anchor point position, especially when using a feedforward network as described above, can be independent of the transformer or model training and then used with fixed weights. In order to obtain the input data and the reference truth for the transformation, a trained transformer with fixed weights is used, which obtains the result feature vectors from the point cloud data as described above. These result feature vectors are then fed to the transformation, thereby obtaining the transformed result feature vectors. In order to obtain the reference truth, the estimation of the frame parameters is applied not only to the obtained result feature vectors, but also to the transformed result feature vectors. Here, except for the position difference, all frame parameters should remain unchanged. Finally, the transformed result feature vector is adjusted until the position difference of the frame parameters relative to the new anchor point position after the transformation is zero and they thus coincide with each other.
[0020] The computer program is configured to perform each step of the method, in particular when the computer program is executed on a computing device or a control device. The computer program enables the method to be implemented in a conventional electronic control device without making structural changes thereto. The computer program is stored on a machine-readable storage medium for implementation.
[0021] By loading the computer program onto a conventional electronic control device, an electronic control device is obtained which is configured to carry out the detection of a plurality of objects from point cloud data. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Exemplary embodiments of the present invention are shown in the drawings and explained in more detail in the following description.
[0023] FIG. 1 shows a method according to the prior art ( Figure 1a ) and according to an embodiment of the method according to the invention ( Figure 1b ) to obtain a bird’s-eye view visualization of the bounding box.
[0024] Figure 2A flow chart shows one specific embodiment of the method according to the invention.
[0025] Figure 3 A flow chart showing the transformation of the result feature vector with respect to the adjusted anchor position according to the method according to the present invention is shown. DETAILED DESCRIPTION
[0026] FIG1 shows the bounding box B according to the ground truth from a bird's-eye view. gt and the estimated bounding box B obtained by the method for object detection using a transformer e and the position P of the object query y,0 , P y,1 , which is done separately from the positions of these object queries. Figure 1a In the example above, we always query the same position P from the object. y,0 Start to obtain the estimated bounding box B e Since the object query position P y,0 The position of the object, i.e., the bounding box B according to the ground truth gt There is a distance between the positions arranged, which will cause inaccuracy when it is calculated in the decoder of the transformer, making the estimated bounding box B e and the bounding box B according to the ground truth gt Clearly deviating from each other. Figure 1b The results of the method according to the invention are shown. The bounding box B estimated in the first layer of the transformer e The image of Figure 1a As in the original position P of the object query y,0 Start to proceed and Figure 1b As described below, the object query is then refined and adjusted based on the new anchor positions, which depend on the bounding box B obtained in the first layer. e The bounding box B shown here is e The new position P of the query from the refined object y,1 The new positions P of these refined object queries are y,1 Closer to the actual object, i.e., closer to the bounding box B according to the ground truth gt The position is arranged so that the estimated bounding box B can be better determined. e , which in turn allows for more accurate object detection.
[0027] Figure 2A flow chart of the method according to the invention is shown for two layers of a converter. Identical steps are marked with the same reference numerals and are only described in detail once. In the following, s designates the number of the layer of the decoder of the converter. i is used as a loop variable (Laufvariable) for the feature vectors, wherein M feature vectors are provided.
[0028] First, the lidar sensor of vehicle F detects the environment. The visualization of these recorded point cloud data is marked with 1. The backbone network 2 calculates feature vectors from the point cloud data, which are then enhanced by sine and cosine through position encoding 3 and finally used as the key vector k i Sum value vector v i is supplied to the decoder 6 of the converter.
[0029] At the same time, the first anchor point position is obtained from the point cloud data through sampling method 4, such as farthest point sampling These first anchor point positions are then Fourier encoded5:
[0030]
[0031] Here, B is a matrix with normally distributed entries, and FFN is a feed-forward network consisting of two layers with ReLU (Rectified Linear Unit) activation function. is the calculated feature vector, which is supplied as an object query to the decoder 6 of the transformer.
[0032] Directly based on the first anchor point position The first set of feature vectors found is denoted by Y0 and is queried by the object Composition. Each object query Serves as a slot for possible objects (in Figure 2 The decoder 6 of the transformer consists of six layers s, each with eight attention heads. In the first layer s0 (s=0), the decoder 6 is based on the object query and the key vector k i Sum value vector v i To obtain the resulting eigenvector Object Query Key vector k i Sum value vector v i The dimension is, for example, 256.
[0033] Thus, two objects O1 and O2 are detected. The feedforward network 7 is based on the resulting feature vector of the first layer s0 Calculate box parameters for objects O1, O2 Among them, Δx, Δy, and Δz represent the position relative to the anchor point in three dimensions. The difference in position, w, l, h are the sizes of objects O1 and O2 in three dimensions, γ is the orientation of objects O1 and O2, v x 、v y represents the speed of the objects O1 , O2 on the horizontal plane and cls is the class identifier. The objects O1 , O2 have been detected and are here plotted in the visualization indicated by 8 .
[0034] According to the present invention, a refined VQ of an object query is performed. To this end, on the one hand, the anchor point position is adjusted 40 In order to obtain the adjusted anchor point position of the further layer s for the decoder 6 The frame parameters obtained in the first layer s0 of decoder 6 are The position differences Δx, Δy, Δz are added to the first anchor point position Then, we can get the adjusted anchor point position.
[0035]
[0036] From the resulting feature vector The position of the first anchor point can be obtained Frame parameters that are far apart and thus have large position differences Δx, Δy, Δz Adjusted anchor point position by adjusting the anchor point position by 40 to get an adjusted anchor point position closer to the object
[0037] From these adjusted anchor point positions Start to execute the code 50 again, which is equivalent to the code 5 mentioned above, and can refer to the code mentioned above. Thus, a feature vector is obtained, which is used as the object query A further layer s is supplied to a decoder 6 of the converter.
[0038] On the other hand, the result feature vector of the first layer s0 is processed with the help of the anchor alignment module AAM (English: anchor alignment module) Transformation 90, refer to Figure 3 This transformation is described in more detail. Here we get the position of the anchor point after adjustment Aligned transformed result feature vector
[0039]
[0040] The resulting feature vector after transformation and the feature vector obtained by coding 50 above As a set of feature vectors (this set is called Y s ) is fed to another layer s of the decoder.
[0041]
[0042] Per-object query and each transformed resulting feature vector and are used as slots for possible objects (in Figure 2 , which is shown by a single square in the figure). Therefore, a total of M slots are obtained. The decoder 6 performs the above-mentioned operation based on the adjusted anchor point position of the current layer s in the other layer s. Query the object obtained The transformed result feature vector of the previous layer s0 is and the key vector k i Sum value vector v i To obtain the resulting eigenvector The resulting feature vector Subsequently, it is also fed to the feedforward network 7, which calculates the frame parameters for the objects O1, O2. Here, due to the refinement of the object query VQ, the position differences Δx, Δy, and Δz obtained here are small.
[0043] Figure 2 A further refinement QV for an object query of further layers is shown in . In query 100 it is decided whether a further refinement QV should be performed, thereby enabling a further improvement in the detection accuracy in further layers. r Indicate here the level at which the refinement QV should be performed. In the case of Used as object lookup for subsequent layers (not shown here).
[0044] For s∈S r In the case of, the corresponding object query refinement QV is performed. As mentioned above, on the one hand, the anchor position Adjust 140 to obtain the adjusted anchor point position This is done by taking the frame parameters obtained in the current layer s of the decoder 6 The position differences Δx, Δy, Δz are added to the anchor point position From these adjusted anchor point positions Start to execute the encoding 150 again, which is equivalent to the encoding 5 or 50 mentioned above. You can refer to the encoding mentioned above to obtain the feature vector On the other hand, the anchor alignment module AAM (English: anchor alignment module) is used to perform the resulting feature vector The transformation 190 is equivalent to referring to Figure 3 The above-mentioned transformation 90 can be referred to, thereby obtaining the transformed result feature vector
[0045] In general, the set Y of feature vectors of layer s fed to decoder 6 is s Depending on the number of layers and whether to perform a refinement QV of an object query for this layer, the following settings can be made:
[0046]
[0047] Here, j = max{l|l <s∧l∈S r}, which means that in the second case (the second line), the object currently obtained by encoding 5, 50, 150 is always queried is fed to the decoder 6. The last row shows the case for the first layer s0.
[0048] Figure 3 A flow chart of the transformation 90 is shown. The resulting feature vector is fed to a feedforward network consisting of two layers 91 and 92 with a ReLU activation function. The feedforward network is trained so that the two layers 91 and 92 change the resulting feature vector The position differences Δx, Δy, Δz relative to the previous anchor position are set to zero. Layers 91, 92 are themselves deformations of the input with the help of learned weights. After the first layer 91, an intermediate representation of dimension h is obtained. After the second layer 92, the transformed resulting feature vector It has the same resulting feature vector as the input The same dimension d. In addition, the original result feature vector is created The bypass connection 94 is provided to ensure that no information is lost. The above description can also be transferred to the transformation for another layer s, for example to the transformation 190.
Claims
1. A method for detecting a plurality of objects (O1, O2) from point cloud data by means of a transformer with an attention model, wherein: The state of the tracked object (O1, O2) is stored in the feature space within the model. The method has the following steps: - computing a feature vector from the point cloud data by the backbone network (2), wherein the feature vector is used as a key vector (k i ) and the value vector (v i ); - Calculate the first anchor point position for the first layer (s0) of the transformer from the point cloud data by sampling method (4) - Based on the first anchor point position by means of encoding (5) to obtain a feature vector, wherein the feature vector is used as an object query for the first layer (s0) of the transformer - by means of the first layer (s0) of the decoder (6) of the transformer based on the object query and the key vector (k i ) and the value vector (v i ), to obtain the resulting feature vector in the first layer (s0) of the transformer - the resulting feature vector for the first layer (s0) of the transformer Calculate (7) box parameters - adjusting (40, 140) the anchor point position for at least one further layer(s) of the transformer The method is to set the frame parameters The position difference is added to the first anchor point position superior; - Based on the adjusted anchor point position with the help of code (50, 150) Find the eigenvector, wherein the feature vector is used as an object query for the at least one further layer(s) of the transformer -About the adjusted anchor point position Transform (90) the resulting feature vector of the first layer Among them, the transformed result feature vector serving as an object query for said at least one further layer(s) of said transformer; - The decoder (60) of the transformer based on the transformed result feature vector of the previous layer (s0) The object query calculated by the current layer(s) and the key vector (k i ) and the value vector (v i ) obtains the resulting feature vector in the at least one further layer(s) of the transformer 2. The method according to claim 1, characterized in that For at least one additional layer(s), perform the following steps: Compute box parameters for the resulting feature vector Adjust (140) anchor point position Based on the adjusted anchor point position, the encoding (150) Find the eigenvector And about the adjusted anchor point position Transform (190) the resulting feature vector 3. The method according to claim 1 or 2, characterized in that: To train the transformer the following steps are performed: - estimating multiple sets of box parameters for the decoder output of each layer; - Assigning the frame parameters of the ground truth to the frame parameters of the closest estimate; - applying median regression to the deviation between the frame parameters of the ground truth and the assigned estimated frame parameters; - Training the transformer by means of the median regression.
4. The method according to any one of the preceding claims, characterized in that The adjusted anchor point position is realized by a feed-forward network composed of two layers (91, 92) with a ReLU activation function. The resulting feature vector The transformation (90, 190) is performed.
5. The method according to claim 4, characterized in that In order to train the adjusted anchor point position The resulting feature vector The transformation (90, 190) is performed by implementing the following steps: - obtaining a result feature vector from the point cloud data; - transforming the resulting feature vector; - applying estimates of box parameters to the resulting feature vector and the transformed resulting feature vector; - Adjust the resulting feature vector after said transformation until the positional difference of its box parameters is zero.
6. The method according to any one of the preceding claims, characterized in that The point cloud data is detected by a laser radar.
7. The method according to any one of the preceding claims, characterized in that The point cloud data is recorded from a vehicle (F).
8. A computer program, configured to execute each step of the method according to any one of claims 1 to 7.
9. A machine-readable storage medium having stored thereon the computer program according to claim 8. 10 . An electronic control device configured to perform a detection of a plurality of objects from point cloud data by means of a transformer having an attention model by means of a method according to claim 1 .