Three-dimensional target detection method, computer device and computer readable storage medium

By suppressing trivial attention weights in the cross attention module of the Transformer decoder, the problem of low detection accuracy of the existing three-dimensional object detection method is solved, and more accurate and efficient object detection is achieved.

CN120071322APending Publication Date: 2025-05-30ZHENGZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510110882.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing three-dimensional object detection methods have low detection accuracy, especially when dealing with three-dimensional objects in complex scenarios and real-world situations.

Method used

In the cross attention module of the Transformer decoder, attention weights smaller than the threshold are suppressed as trivial attention weights. By reducing these weights, the ratio of their maximum value is smaller than the threshold, thereby highlighting important attention weights.

Benefits of technology

By suppressing trivial attention weights, the model can focus more accurately on key information, output object detection boxes that are more accurate and close to the ground real frame, improving the accuracy of three-dimensional object detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071322A_ABST
    Figure CN120071322A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computer vision, and particularly relates to a three-dimensional target detection method, a computer device and a computer readable storage medium. The method comprises the following steps: S1, acquiring three-dimensional point cloud data of a to-be-detected area; s2, inputting the three-dimensional point cloud data into a pre-trained improved Group-Fre-3D target detection model, and obtaining a spatial position of a target in the to-be-detected region; the improved Group-Free-3D target detection model is obtained by improving Group-Free-3D, and the improvement comprises the following steps: in a cross attention module of a Transform decoder, taking an attention weight smaller than a threshold value as a trivial attention weight, and taking a trivial attention weight as a trivial attention weight; the trivial attention weights are shrunk so that the sum of the trivial attention weights of any row in the shrunk attention weight matrix can be smaller than the maximum value in the attention weights of the row. The technical problem of low detection precision of a three-dimensional target detection method in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a three-dimensional object detection method, a computer device, and a computer-readable storage medium. Background Art

[0002] With the progress of computer vision technology, object detection plays an important role in many fields. Traditional 2D object detection methods perform well in indoor scenes, but there are some limitations in dealing with complex scenes and three-dimensional objects in the real world. Therefore, object detection methods based on 3D data have received increasing attention. Compared with 2D object detection, 3D object detection can not only obtain the complete three-dimensional information of the object, but also has stronger robustness and accurate positioning ability. As a basic technology for 3D scene understanding, it plays an important role in many applications such as autonomous driving, robot operation, and augmented reality.

[0003] Since the development of 3D object detection to date, many excellent methods have emerged. For example, the multi-view based method, such as MV3D (Multi-view 3D Objective Network), projects point clouds onto the bird's-eye view and front view, and combines two-dimensional images as input. There are also voting-based methods, such as VoteNet. The idea of Hough voting is introduced in VoteNet, which makes the scattered points in the scene approach their respective object centers. Then there are attention mechanism-based methods, such as 3DETR (3D Detection Transformer) and MLCVNet (Multi-Level Context VoteNet). The emergence of 3DETR has enabled Transformer to have a general model in the field of three-dimensional object detection. MLCVNet extracts context semantics at multiple levels including points, objects, and the global scene using the attention mechanism. The attention mechanism has received extensive attention and application in the field of computer vision, and it can enable the model to focus on the regions of interest.

[0004] Microsoft Research Asia and the University of Science and Technology of China proposed the Group-Free-3D detection network in the paper "Group-Free-3D Object via Tranformers" published in ICCV2021 in 2021. The core of the method is to achieve adaptive grouping through the network structure of Tranformer, establish long-range association between candidate points and seed points, and use the cascaded Tranformer method for feature fusion, improving the accuracy of object detection.

[0005] However, the traditional attention mechanism in Transformer focuses on global attention and lacks attention to important information. Although the attention mechanism can capture long-range dependencies, in some cases, the model may ignore information that is far away but important for the task, especially when the input sequence is very long. Moreover, it will generate a relatively uniform attention distribution, which may not be conducive to the model focusing on key information. Also, in the attention weights, there are many trivial attention weights, which will generate noise and affect the results of object detection, and finally limit the accuracy of object detection. Summary of the Invention

[0006] The purpose of the present invention is to provide a three-dimensional object detection method, a computer device, and a computer-readable storage medium to solve the technical problem of low detection accuracy in the existing three-dimensional object detection method.

[0007] To solve the above technical problem, the technical solution of a three-dimensional object detection method provided by the present invention is: A three-dimensional object detection method, the method includes:

[0008] S1. Obtain the three-dimensional point cloud data of the area to be detected;

[0009] S2. Input the three-dimensional point cloud data into a pre-trained improved Group-Free-3D object detection model to obtain the spatial position of the object in the area to be detected;

[0010] The improved Group-Free-3D object detection model is obtained by improving Group-Free-3D. The improvement includes: in the cross-attention module of the Transformer decoder, taking the attention weights less than the threshold as trivial attention weights, and shrinking the trivial attention weights so that the sum of the trivial attention weights in any row of the shrunk attention weight matrix is less than the maximum value among the attention weights in that row.

[0011] The beneficial effects of the above technical solution are: The technical solution of a three-dimensional object detection method of the present invention belongs to an improved invention. The present invention suppresses the trivial attention weights in the attention weights of the cross-attention, suppresses the unimportant weights (i.e., trivial attention weights), and highlights the important weights. It highlights the features useful for object detection, outputs more accurate object detection frames that are closer to the ground truth boxes. The present invention solves the technical problem of low detection accuracy in the existing three-dimensional object detection method.

[0012] Further, the trivial attention weights are shrunk according to the following formula:

[0013]

[0014] where, where, xj is the j-th trivial attention weight of any row in the attention weight matrix; x j ′ is the value after suppressing the trivial attention weight x j ; s is the suppression scale, 0 ≤ s ≤ 1; k is the number of trivial attention weights in this row of the attention weight matrix.

[0015] Furthermore, the threshold is determined in the following manner:

[0016] threshold = t × x m

[0017] where threshold is the threshold; t is the scaling scale, 0 ≤ t ≤ 0.1; x m is the maximum value of each row in the attention weight matrix.

[0018] Furthermore, the improvement also includes: in the Transformer decoder, the multi-head self-attention module is a multi-head scale adaptive self-attention module, and each head of the multi-head scale adaptive self-attention module has a different receptive field.

[0019] Furthermore, the calculation formula of the multi-head scale adaptive self-attention module is:

[0020]

[0021] where Q, K, and V are the query vector, key vector, and value vector respectively; d is the dimension of the key vector; D is the all-pair distance between query centers; τ is a coefficient used to control the receptive field, and τ varies with each head.

[0022] Furthermore, the coefficient τ linearly transforms with the query vector of each head.

[0023] The present invention also provides a technical solution for a computer device: a computer device, including a processor, where the processor is used to execute a computer program to implement the steps of the three-dimensional object detection method as described above.

[0024] The present invention also provides a technical solution for a computer-readable storage medium: a computer-readable storage medium, where a computer program is stored inside the computer-readable storage medium, and the computer program is used to be executed by a processor to implement the steps of the three-dimensional object detection method as described above. Brief Description of the Drawings

[0025] Figure 1 is the network structure diagram of the improved Group-Free-3D object detection model in the embodiment of the three-dimensional object detection method of the present invention;

[0026] Figure 2Schematic diagram of the scale adaptive attention mechanism in the embodiment of the 3D object detection method of the present invention;

[0027] Figure 3 Comparison chart before and after suppressing trivial attention weights in the embodiment of the 3D object detection method of the present invention. Detailed implementation manners

[0028] The present invention suppresses the trivial attention weights in the attention weights of cross-attention, suppresses the unimportant weights (i.e., trivial attention weights), highlights the important weights, highlights the features useful for object detection, outputs more accurate object detection frames that are closer to the ground truth boxes. The present invention solves the technical problem of low detection accuracy in the 3D object detection method in the prior art.

[0029] Embodiment of the 3D object detection method:

[0030] A 3D object detection method is implemented based on the improved Group-Free-3D object detection model as shown in Figure 1 and specifically includes the following steps:

[0031] S1. Obtain the 3D point cloud data of the area to be detected;

[0032] S2. Input the 3D point cloud data into the pre-trained improved Group-Free-3D object detection model to obtain the spatial positions of the objects in the area to be detected.

[0033] Specifically, the above S2 specifically includes:

[0034] Step 1: Input the original 3D point cloud data Xin (i.e., the N*3 Input points in Figure 1 ) into the Pointnet++ module dedicated to processing irregular point clouds, process the 3D point cloud data, extract features, and generate seed points that uniformly cover the entire scene;

[0035] The 3D point cloud needs to go through the backbone network Pointnet++ for its processing, which contains four downsampling (SA) layers and two upsampling (FP) layers;

[0036] (1) Downsampling layer: There are a total of four downsampling layers, which are called Set Abstraction (SA) set abstraction sampling layers. Each SA layer is divided into two processes: sampling and grouping. The initial input point cloud has 20,000 points. After the first layer of SA, 2,048 representative points are sampled by the Farthest Point Sampling (FPS) method. Around each point, 64 adjacent points are selected by ball query for grouping, and the information of adjacent points is aggregated for depth feature learning. The second layer of SA continues to downsample to 1,024 points using FPS, and 32 points are selected by ball query around the 1,024 points for grouping and information aggregation. And so on, 512 and 256 points are sampled using FPS in the third and fourth SA downsampling layers respectively, and 16 and 16 points are selected by ball query around them respectively for grouping to aggregate information.

[0037] (2) Upsampling layer: There are a total of two upsampling layers, which are called Feature Propagation (FP) feature propagation layers. The 256 point clouds in the last layer of the SA downsampling layer are input into the first layer of FP. Through the feature backpropagation layer, higher-level point features are attached to the features in the previous layer of SA3 layer by interpolation. Then the newly obtained features are input into the second layer of FP. Similarly, through the feature backpropagation layer, higher-level point features are attached to the features in the previous layer of SA2 layer by interpolation, so as to obtain the final features of the entire backbone network Pointnet++.

[0038] Step 2: Sample the initial candidate objects from the points on the point cloud to generate initial candidate boxes;

[0039] The k-Cloest Point Sampling (KPS) method is used to discriminate the global seed point features inside the target, and K points closest to the target are selected as candidate points (i.e., the Initial Candidate Box Generation Module in Figure 1 );

[0040] (1) According to the seed points and their features, convolution and sigmoid operations are used to predict the probability that the seed points become candidate points;

[0041] (2) Select the 256 points with the highest probability as the initial candidate points;

[0042] (3) The 1,024 seed points are labeled by the KPS (K = 4) method. Those that meet the requirements are assigned a value of 1, otherwise 0;

[0043] (4) Gradually approximate the initial candidate point results to the KPS labels through the loss function and training iterations.

[0044] Step 3: The candidate boxes pass through a Transformer module for refining the candidate boxes, which contains a self-attention processing unit, to further refine the candidate boxes; in order to refine the obtained candidate boxes to get more accurate candidate boxes. Six cascaded Transformers are used as the decoder to perform deeper feature extraction and fusion. The key intermediate input variables of the Transformer structure include query (Q), key (K), and value (V). Q is the candidate box feature, and both K and V are the seed point features. The contribution degree of the feature value V to Q, that is, the weight size, is determined by calculating the similarity between Q and K. Therefore, the encoding layer of the Transfomer considers the correlation between Q and (K, V), especially the contribution degree of (K, V) to Q.

[0045] Therefore, the transformer structure of the model extracts the contribution degree of the seed points to the candidate boxes. Each Transfomer output feature serves as a new candidate box feature, and the Q of the next layer will be updated to the new candidate box feature. Through this cascaded operation, on the one hand, the model can obtain deeper-dimensional features, and on the other hand, it can gradually improve the prediction results, ultimately improving the object detection accuracy.

[0046] Input the point cloud features in the initial candidate boxes into the Transformer. Each layer of the Transformer contains a self-attention module and a cross-attention module. In the improved Group-Free-3D object detection model, the multi-head self-attention module is the multi-head scale-adaptive self-attention module (i.e., Figure 1 Scale-adaptive Self-Attention in

[0047] In this step, the self-attention module in the Transformer is experienced. The formula for ordinary computer global attention is:

[0048]

[0049] As Figure 2 shown, in the scale-adaptive self-attention improved in this embodiment, the following three steps are also included: (1) Calculate the distance between query points, (2) Calculate the receptive field of each head, and (3) Add the distance and receptive field to the ordinary global self-attention to change it to scale-adaptive self-attention;

[0050] (1) Multi-head self-attention has a global receptive field and lacks the ability of local multi-scale aggregation. Therefore, scale-adaptive self-attention is proposed, which learns an appropriate receptive field under the guidance of query Q. First, we calculate the pairwise distances D ∈ R N×N (N is the number of query points), x i , y i , z i represent the center of the i-th query.

[0051]

[0052] (2) Secondly, each head can learn an appropriate receptive field. τ is a scalar used to control the receptive field of each query and is specific to each head. When τ = 0, it degenerates to ordinary self-attention with a global receptive field. As τ increases, the attention weights away from the query become smaller and the receptive field also shrinks. Suppose there are H heads, and a linear transformation is used to generate head-specific τ from a given query q ∈ R d . Each head learns a different receptive field, enabling the network to achieve multi-scale feature aggregation. The calculation formula for τ is:

[0053] [τ 1 , τ 2 ,..., τ H = Linear d→H (q), [τ 1 , τ 2 ,..., τ H ∈ R H

[0054] Linear represents a linear transformation.

[0055] (3) Changing ordinary global self-attention to scale-adaptive self-attention, the calculation formula is:

[0056]

[0057] Among them, Q, K, and V are the query vector, key vector, and value vector respectively; d is the dimension of the key vector; D is the pairwise distance between query centers; τ is the coefficient used to control the receptive field, and τ changes with the query vector of each head.

[0058] After scale-adaptive self-attention, compared with using ordinary global self-attention before, the candidate boxes are refined and the error is smaller.

[0059] Step 4: The candidate boxes then pass through the cross-attention module in the refined Transformer module to obtain more accurate candidate boxes. In the improved Group-Free-3D object detection model, in the cross-attention module of the Transformer decoder (i.e., Figure 1 the cross attention module in

[0060] such as Figure 3 , in the cross-attention layer of the Transformer, the attention weights are adjusted. The attention weights are divided into trivial and non-trivial attention weights through a threshold, and then the accumulated trivial attention weights are suppressed to reduce attention noise and emphasize important weights, so as to highlight the features useful for detection. It is further divided into the following two steps: the first step is division, and the second step is suppression;

[0061] (1) Division. Before suppressing the trivial attention weights, a threshold is designed first. Those below the threshold are considered trivial attention weights, and those above the threshold are considered important attention weights. Among them, xm is the maximum value of each row in the attention weights. t is the scaling factor, and its value is taken as 0.1. The expression of the threshold is:

[0062] threshold=t×x m

[0063] where threshold is the threshold; t is the scaling factor, 0 ≤ t ≤ 0.1; x m is the maximum value of each row in the attention weight matrix.

[0064] (2) Suppression. In each row of attention, x 1 +x 2 +…+x k +x k+1 +…+x n-1 +x n =1. Assume that x 1 to x k are all below the threshold threshold and are all trivial and unimportant attention weights. s is the suppression scale (0 ≤ s ≤ 1), which can make the sum of the suppressed unimportant attention weights less than the maximum value. For x 1 to x k process them according to the following formula:

[0065]

[0066] j = 1, 2, ..., k

[0067] wherein, x j is the j-th trivial attention weight of any row in the attention weight matrix; x j ′ is the value after suppressing the trivial attention weight x j ; s is the suppression scale, 0 ≤ s ≤ 1; k is the number of trivial attention weights in this row of the attention weight matrix.

[0068] As Figure 3 shown, it represents the change of the attention weight before and after suppression. The abscissa in the figure represents the intervals where each attention weight in the attention weight matrix is located, and the ordinate represents the sum of each attention weight in the interval. It can be seen that the trivial attention weights after suppression are well suppressed, and the characteristics of important attention weights can be more prominent.

[0069] (3) Suppress and divide the weights in the cross-attention by the above steps, suppress the unimportant weights, highlight the important weights, highlight the features useful for object detection, output more accurate object detection frames that are closer to the ground truth boxes.

[0070] Embodiment of a computer device:

[0071] A computer device includes a processor, and the processor is configured to execute a computer program to implement the steps of the three-dimensional object detection method as described above. The specific three-dimensional object detection method has been introduced in sufficient detail in the above three-dimensional object detection method embodiment and will not be repeated here.

[0072] Embodiment of a computer-readable storage medium:

[0073] A computer-readable storage medium stores a computer program internally, and the computer program is configured to be executed by a processor to implement the steps of the three-dimensional object detection method as described above. The specific three-dimensional object detection method has been introduced in sufficient detail in the above three-dimensional object detection method embodiment and will not be repeated here.

[0074] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still make modifications to the technical solutions described in the foregoing embodiments without creative efforts, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A three-dimensional target detection method, characterized in that: The method includes: S1. Obtain three-dimensional point cloud data of the area to be detected; S2, inputting the three-dimensional point cloud data into a pre-trained improved Group-Free-3D target detection model to obtain the spatial position of the target in the area to be detected; The improved Group-Free-3D target detection model is obtained by improving Group-Free-3D, and the improvement includes: in the cross-attention module of the Transformer decoder, the attention weights less than the threshold are used as trivial attention weights, and the trivial attention weights are reduced so that the sum of the trivial attention weights of any row in the attention weight matrix after reduction is less than the maximum value among the attention weights in the row.

2. The three-dimensional target detection method according to claim 1, characterized in that: The trivial attention weights are reduced as follows: j=1,2,...,k Among them, x j is the jth trivial attention weight of any row in the attention weight matrix; x j ′ is the trivial attention weight x j The value after suppression; s is the suppression scale, 0≤s≤1; k is the number of trivial attention weights in this row in the attention weight matrix.

3. The three-dimensional target detection method according to claim 1 or 2, characterized in that: The threshold is determined in the following manner: threshold=t×x m Where threshold is the threshold; t is the scaling scale, 0≤t≤0.1; x m is the maximum value of each row in the attention weight matrix.

4. The three-dimensional target detection method according to claim 1, characterized in that: The improvement also includes: in the Transformer decoder, the multi-head self-attention module is a multi-head scale adaptive self-attention module, and each head of the multi-head scale adaptive self-attention module has a different receptive field.

5. The three-dimensional target detection method according to claim 4, characterized in that: The calculation formula of the multi-head scale adaptive self-attention module is: Among them, Q, K and V are query vector, key vector and value vector respectively; d is the dimension of key vector; D is the all-pairs distance between query centers; τ is the coefficient used to control the receptive field, and τ varies with each head.

6. The three-dimensional target detection method according to claim 5, characterized in that: The coefficient τ scales linearly with the query vector in each head.

7. A computer device comprising a processor, characterized in that: The processor is used to execute a computer program to implement the steps of the three-dimensional object detection method according to any one of claims 1 to 6.

8. A computer-readable storage medium, wherein a computer program is stored therein, characterized in that: The computer program is used to be executed by a processor to implement the steps of the three-dimensional object detection method according to any one of claims 1 to 6.