An Offline 3D Object Detection Method and System

By detecting and tracking objects on multi-frame point cloud sequences, combining geometric shape, position and confidence prediction, and optimizing bounding boxes, the problem of offline 3D object detection algorithms in the prior art generate incomplete trajectories and high-cost annotations, achieving efficient and accurate object detection and tracking.

CN116580386BActive Publication Date: 2025-07-11SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310685774.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-09
Publication Date
2025-07-11
Estimated Expiration
2043-06-09

AI Technical Summary

Technical Problem

The existing offline 3D object detection algorithms have shortcomings in generating complete target trajectories and long-term contextual representations of object motion states, resulting in poor detection performance and high cost of manual labeling.

Method used

By detecting and tracking objects on multi-frame point cloud sequences, object bounding boxes and trajectories are generated, combining geometric shape, position and confidence predictions, bounding boxes are optimized, and data augmentation and multi-model fusion are used to use long-time point cloud features to generate accurate object tracking sequences.

Benefits of technology

It improves the detection and tracking performance of 3D objects, reduces the cost of manual labeling, and the generated detection results reach or even exceed the level of manual labeling, which can replace manual labeling for online model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116580386B_ABST
    Figure CN116580386B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of three-dimensional object detection technology, and proposes an offline three-dimensional object detection method and system. The method includes: performing object detection on multiple frames of point cloud sequences of an object to generate object bounding boxes; performing offline object tracking on the object bounding boxes to generate object trajectories; extracting object sequence data from the object trajectories; generating optimized bounding boxes through object attribute prediction according to the object sequence data; and transmitting the optimized bounding boxes back to the coordinate system where the object appears. When applied to the field of autonomous driving, the present invention can greatly improve the 3D object detection performance and 3D object tracking performance. And it can generate high-quality detection results, which can reach or even exceed the level of manual annotation, can replace manual annotation for the ground truth annotation of other online model training, and thus can significantly reduce the investment of labor costs in the training of autonomous driving perception models and promote the development of perception model performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to the technical field of three-dimensional object detection. Specifically, the present invention relates to an offline three-dimensional object detection method and system. Background Art

[0002] Autopilot refers to the technology that enables a vehicle to autonomously drive according to road conditions and destinations without human intervention. To achieve this goal, an autonomous vehicle needs to be able to perceive the surrounding environment, identify objects such as roads, vehicles, pedestrians, obstacles, etc., and make reasonable decisions and controls based on their position, speed, shape, and other information. 3D object detection refers to using sensors (such as cameras, radars, lasers, etc.) to obtain information about objects in three-dimensional space, such as category, position, pose, size, etc. 3D object detection is an important part of autonomous driving perception, which can provide richer and more accurate object information, helping to improve the safety and efficiency of autonomous vehicles.

[0003] However, in practical applications, autonomous driving and 3D object detection face many challenges and requirements. On the one hand, due to the complex and changeable road environment, the variety of objects, and the existence of occlusion and interaction between objects, 3D object detection needs to have high precision and robustness. On the other hand, since autonomous vehicles need to respond to the surrounding situation in real time, 3D object detection needs to have high efficiency and low latency. To meet these requirements, the autonomous driving perception model needs to continuously perform data-driven continuous iteration.

[0004] Existing autonomous driving perception models rely on a data-driven continuous iteration mode. To provide a sufficient amount of high-quality labeled data, the expensive labor cost and slow labeling efficiency cannot be ignored. Therefore, offline 3D object detection algorithms usually follow a modular pipeline design, using the entire sequence of data from sensors (such as video or sequence point cloud data), and are committed to developing high-quality "automatic annotation", aiming to reduce the labor cost of point cloud annotation in 3D detection tasks and promote the performance development of the autonomous driving perception model.

[0005] With the continuous iterative development of technology, many online detection algorithms have emerged that focus on developing complex modules to better utilize the context features of temporal data. These algorithms are much better than the previous online and offline 3D detection algorithms. Compared with these algorithms, the algorithm architectures and modes of the previous offline 3D detection algorithms are too weak to learn the complex representations of long sequence point clouds. However, the following problems still exist in the current state-of-the-art offline 3D object detection algorithms, hindering their full potential: the online multi-object tracker cannot generate sufficiently complete object trajectories; the motion state of objects poses an inevitable challenge to the object-centered optimization model on how to utilize long temporal context representations. Summary of the Invention

[0006] To at least partially solve the above problems in the prior art, the present invention proposes an offline three-dimensional object detection method and system, including:

[0007] Performing object detection on multiple frames of point cloud sequences of an object to generate object bounding boxes;

[0008] Performing offline object tracking on the object bounding boxes to generate object trajectories;

[0009] Extracting object sequence data from the object trajectories;

[0010] Generating optimized bounding boxes through object attribute prediction according to the object sequence data, where object attribute prediction includes geometric shape prediction, position prediction, and confidence prediction of the object; and

[0011] Transmitting the optimized bounding boxes back to the coordinate system where the object appears.

[0012] In an embodiment of the present invention, it is stipulated that performing object detection on the point cloud sequence of an object to generate object bounding boxes includes the following steps:

[0013] Inputting the point cloud sequence into a center point detector, where the point cloud sequence includes combinations of multiple five-frame point clouds;

[0014] Converting the point cloud sequence into a voxel representation in the first stage of object detection and generating candidate initial bounding boxes; and

[0015] Fine-tuning the candidate initial bounding boxes in the second stage of object detection to predict more accurate bounding boxes and confidence levels of the object, where the point cloud density information of the object is used in the second stage of object detection for the fusion of original point cloud features and voxel features;

[0016] Wherein, data augmentation in the inference stage and multi-model result fusion are performed in the above steps.

[0017] In an embodiment of the present invention, it is stipulated that performing offline object tracking on the object bounding boxes to generate object trajectories includes the following steps:

[0018] Dividing the object bounding boxes into high-group bounding boxes and low-group bounding boxes according to the confidence scores of the object bounding boxes;

[0019] Generating a first trajectory in the first stage of offline object tracking, including:

[0020] Associating the existing object trajectories with the high-group bounding boxes;

[0021] Updating the existing object trajectories according to the successfully associated bounding boxes; and

[0022] Associate the unupdated object trajectories with the low-group bounding boxes to generate the first trajectories, where the object bounding boxes that are not successfully associated are removed. The default lifespan of the first trajectories is infinite (i.e., the length of the entire point cloud sequence). After the association process ends, select the moment of the last actual association with the object bounding box as the final length of the first trajectory of the object, and remove subsequent persistent false associations;

[0023] In the second stage of offline object tracking, generate the second trajectories in the reverse chronological order of the first stage;

[0024] Associate the first trajectories with the second trajectories through position-related similarity scores; and

[0025] Fuse the successfully associated first trajectories and second trajectories to generate object trajectories.

[0026] In an embodiment of the present invention, it is stipulated that extracting object sequence data from object trajectories includes:

[0027] Enlarge the region of interest of the object bounding box in three dimensions;

[0028] Extract the point cloud sequence within the region surrounded by the enlarged object bounding box; and

[0029] Save the point cloud sequence within the region surrounded by the enlarged object bounding box, its tracking box sequence, and the confidence score.

[0030] In an embodiment of the present invention, it is stipulated that the geometric shape prediction of an object includes:

[0031] Randomly select n1 candidate objects from different perspectives in the object sequence, randomly select p1 points for each candidate object, extract the corresponding features using the first encoder, and generate n1 geometric query vectors;

[0032] Overlay the point clouds of each object in the entire object sequence together, randomly select p2 points from them, and generate global point cloud dense features using the second encoder;

[0033] Input the geometric query vectors into the multi-head self-attention layer to encode the context relationship and feature dependency between samples, and extract geometric information;

[0034] Send the updated geometric query vectors and the global point cloud dense features into the multi-head cross-attention layer to infer the differences between each geometric query vector and the global dense point cloud features, and further supplement the features of the required perspectives; and

[0035] Input the updated geometric query vectors into the prediction head, output n1 geometric prediction results, and average the n1 geometric prediction results to obtain the final geometric prediction result.

[0036] In one embodiment of the present invention, it is stipulated that the position prediction of an object includes:

[0037] Randomly select p3 points for each object candidate in the object sequence, use the third encoder to extract the features corresponding to each object, and generate p3 position query vectors;

[0038] Use the fourth encoder to generate the point features of the entire object trajectory, where the point features of the entire object trajectory serve as the key features and value features in the calculation of the cross-attention mechanism;

[0039] Send the position query vectors into the self-attention mechanism module to calculate the relative distance between the current position and other positions, where a one-dimensional mask is used to constrain the self-attention near the position of each position query vector;

[0040] Input the local position query vectors and the point features of the entire object trajectory into the cross-attention module to simulate the context relationship from local to global positions; and

[0041] Predict the offset between each ground truth center and the corresponding initial center in the local coordinate system and the heading angle difference.

[0042] In one embodiment of the present invention, it is stipulated that the confidence prediction of an object includes:

[0043] In the first branch of the object confidence prediction, classify the object as a true positive example or a false positive example according to the overlap degree between the tracking box and the ground truth box of the object; and

[0044] In the second branch of the object confidence prediction, predict the overlap degree that an object should have after being optimized, and set the regression target to the overlap degree between the ground truth box and the optimized bounding box after geometric and position optimization.

[0045] The present invention also proposes an offline three-dimensional object detection system, including:

[0046] An object detection module configured to perform object detection on a multi-frame point cloud sequence of an object to generate an object bounding box;

[0047] An offline object tracking module configured to perform offline object tracking on the object bounding box to generate an object trajectory;

[0048] An object sequence data extraction module configured to extract object sequence data from the object trajectory;

[0049] An object optimization module based on attribute prediction configured to generate an optimized bounding box through object attribute prediction according to the object sequence data; and

[0050] A coordinate transformation module, which is configured to transmit the optimized bounding box back to the coordinate system of each frame where the object appears.

[0051] The present invention also provides an offline three-dimensional object detection system and a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it executes the steps according to the method.

[0052] The present invention also provides a computer system, including:

[0053] A processor, which is configured to execute machine-executable instructions; and

[0054] A memory, on which machine-executable instructions are stored. When the machine-executable instructions are executed by the processor, they execute the steps according to the method described in the claims.

[0055] The present invention has at least the following beneficial effects: The present invention provides an offline three-dimensional object detection method and system, which can generate accurate and complete object tracking sequences, and can fully consider the geometric properties, motion constraints and other characteristics of the object, so as to make full use of the effective feature information of the long-time series point cloud. When applied to the field of autonomous driving, the present invention can greatly improve the 3D object detection performance and 3D object tracking performance. On the Waymo dataset, one of the largest existing publicly available autonomous driving datasets, it achieves the optimal 3D object detection performance with 85.15 mAPH (L2) and the optimal 3D object tracking performance with 75.05 MOTA (L2), both of which far exceed the performance of the second place. Using the present invention can generate high-quality detection results, which can reach or even exceed the level of manual annotation, can replace manual annotation for the ground truth annotation of other online model training, and thus can greatly reduce the investment of labor costs in the training of autonomous driving perception models and promote the development of perception model performance. Description of the Drawings

[0056] To further clarify the advantages and features of the embodiments of the present invention, more specific descriptions of the embodiments of the present invention will be presented with reference to the accompanying drawings. It can be understood that these drawings only depict typical embodiments of the present invention and will not be considered as limiting its scope. In the drawings, for clarity, the same or corresponding components will be denoted by the same or similar reference numerals.

[0057] Figure 1 Shows a computer system implementing the system and / or method according to the present invention.

[0058] Figure 2 Shows a schematic flow chart of an offline 3D object detection algorithm in the prior art.

[0059] Figure 3Shows a schematic diagram of a sequence generated by combining an existing 3D detection algorithm and a multi-object tracking algorithm.

[0060] Figure 4 Shows a schematic diagram of predicting a bounding box based on a sliding window-based dynamic object optimization mechanism.

[0061] Figure 5 Shows a schematic flowchart of an offline three-dimensional object detection method in an embodiment of the present invention.

[0062] Figure 6 Shows a schematic framework diagram of an offline three-dimensional object detection method in an embodiment of the present invention.

[0063] Figure 7 Shows a schematic diagram of processing points and their initial bounding boxes during the geometric shape optimization process in an embodiment of the present invention.

[0064] Figure 8 Shows a schematic flowchart of processing a geometric query vector in an embodiment of the present invention.

[0065] Figure 9 Shows a schematic diagram of processing the corners of points and their corresponding boxes during the position optimization process in an embodiment of the present invention.

[0066] Figure 10 Shows a schematic flowchart of processing a position query vector in an embodiment of the present invention. Detailed implementation manners

[0067] It should be noted that the components in the respective drawings may be exaggeratedly shown for illustration purposes and are not necessarily to scale correctly. In the respective drawings, the same or functionally identical components are provided with the same reference numerals.

[0068] In the present invention, unless otherwise specified, "arranged on", "arranged above", and "arranged on top of" do not exclude the existence of an intermediate object therebetween. In addition, "arranged on or above" only represents the relative positional relationship between two components, and in certain cases, such as after reversing the product direction, it can also be converted to "arranged under or below", and vice versa.

[0069] In the present invention, the respective embodiments are only intended to illustrate the solutions of the present invention and should not be construed as restrictive.

[0070] In the present invention, unless otherwise specified, the quantifiers "a" and "one" do not exclude the scenario of multiple elements.

[0071] It should also be noted here that, in the embodiments of the present invention, for the sake of clarity and simplicity, only a part of the components or assemblies may be shown. However, those of ordinary skill in the art can understand that, under the teachings of the present invention, the required components or assemblies can be added according to the specific scenario requirements. In addition, unless otherwise specified, the features in different embodiments of the present invention can be combined with each other. For example, a certain feature in the second embodiment can be used to replace the corresponding or functionally identical or similar feature in the first embodiment, and the obtained embodiment also falls within the scope of disclosure or the scope of recording of this application.

[0072] It should also be noted here that within the scope of the present invention, terms such as "identical", "equal", "equal to" do not mean that the two values are absolutely equal, but allow a certain reasonable error. That is to say, these terms also cover "substantially identical", "substantially equal", "substantially equal to". By analogy, in the present invention, terms indicating directions such as "perpendicular to", "parallel to", etc. also cover the meanings of "substantially perpendicular to", "substantially parallel to".

[0073] In addition, the numbering of the steps of each method of the present invention does not limit the execution order of the method steps. Unless otherwise specified, the method steps can be executed in different orders.

[0074] The present invention will be further described below in conjunction with the specific embodiments with reference to the accompanying drawings.

[0075] Figure 1 A computer system for implementing the system and / or method according to the present invention is shown. Unless otherwise specified, the method and / or system according to the present invention can be executed in Figure 1 the computer system 100 shown to achieve the purpose of the present invention, or the present invention can be distributedly implemented in multiple computer systems 100 according to the present invention through a network, such as a local area network or the Internet. The computer system 100 of the present invention can include various types of computer systems, such as handheld devices, laptop computers, personal digital assistants (PDAs), multi-processor systems, microprocessor-based or programmable consumer electronic devices, network PCs, minicomputers, mainframes, network servers, tablet computers, and so on.

[0076] As Figure 1 shown, the computer system 100 includes a processor 111, a system bus 101, a system memory 102, a video adapter 105, an audio adapter 107, a hard disk drive interface 109, an optical drive interface 113, a network interface 114, and a universal serial bus (USB) interface 112. The system bus 101 can be any of several types of bus structures, such as a memory bus or a memory controller, a peripheral bus, and a local bus using various bus architectures. The system bus 101 is used for communication between various bus devices. In addition to Figure 1In addition to the bus devices or interfaces shown, other bus devices or interfaces are also conceivable. The system memory 102 includes a read-only memory (ROM) 103 and a random access memory (RAM) 104. The ROM 103 can store, for example, basic input / output system (BIOS) data for basic routines for information transfer during startup, while the RAM 104 is used to provide the system with a relatively fast operating memory. The computer system 100 also includes a hard disk drive 109 for reading and writing to the hard disk 110, an optical drive interface 113 for reading and writing to optical media such as CD-ROMs, and so on. The hard disk 110 can store, for example, an operating system and application programs. The drive and its associated computer-readable medium provide non-volatile storage of computer-readable instructions, data structures, program modules, and other data for the computer system 100. The computer system 100 can also include a video adapter 105 for image processing and / or image output, which is used to connect output devices such as the display 106. The computer system 100 can also include an audio adapter 107 for audio processing and / or audio output, which is used to connect output devices such as the speaker 108. In addition, the computer system 100 can also include a network interface 114 for network connection, where the network interface 114 can be connected to the Internet 116 through a network device such as a router 115, and the connection can be wired or wireless. Additionally, the computer system 100 can also include a universal serial bus interface (USB) 112 for connecting peripheral devices, where the peripheral devices include, for example, a keyboard 117, a mouse 118, and other peripheral devices such as a microphone, a camera, etc.

[0077] When the present invention is implemented on Figure 1 the computer system 100 described above, an accurate and complete object tracking sequence can be generated, and the geometric properties, motion constraints, and other characteristics of the object can be fully considered, so as to make full use of the effective feature information of the long-time series point cloud. When applied to the field of autonomous driving, it can greatly improve the 3D object detection performance and 3D object tracking performance, and can generate high-quality detection results, which can reach or even exceed the level of manual annotation, can replace manual annotation for the ground truth annotation of other online model training, and thus can significantly reduce the investment of labor costs in the training of the autonomous driving perception model and promote the development of the perception model performance.

[0078] In addition, each embodiment may be provided as a computer program product that may include one or more machine-readable media having machine-executable instructions stored thereon, which when executed by one or more machines such as a computer, a computer network, or other electronic devices, may cause the one or more machines to perform operations in accordance with the embodiments of the present invention. The machine-readable media may include, but are not limited to, floppy disks, optical disks, CD-ROMs (Compact Disc Read-Only Memory), and magneto-optical disks, ROMs (Read-Only Memory), RAMs (Random Access Memory), EPROMs (Erasable Programmable Read-Only Memory), EEPROMs (Electrically Erasable Programmable Read-Only Memory), magnetic or optical cards, flash memory, or other types of media / machine-readable media suitable for storing machine-executable instructions.

[0079] In addition, each embodiment may be downloaded as a computer program product, where the program may be transmitted from a remote computer (e.g., a server) to a requesting computer (e.g., a client) by one or more data signals implemented and / or modulated by a carrier wave or other propagation medium via a communication link (e.g., a modem and / or a network connection). Thus, the machine-readable media used herein may include such a carrier wave, but this is not necessary.

[0080] In the present invention, the modules of the system according to the present invention may be implemented using software, hardware, firmware, or a combination thereof. When a module is implemented using software, the functions of the module may be implemented through a computer program process. For example, the module may be implemented by a code segment (such as a code segment in languages such as C or C++) stored in a storage device (such as a hard disk, memory, etc.), where when the code segment is executed by a processor, the corresponding functions of the module can be achieved. When a module is implemented using hardware, the functions of the module may be implemented by setting the corresponding hardware structure. For example, the functions of the module may be implemented by hardware programming a programmable device such as a Field Programmable Gate Array (FPGA), or by designing an Application Specific Integrated Circuit (ASIC) including a plurality of electronic devices such as transistors, resistors, and capacitors. When a module is implemented using firmware, the functions of the module may be written in a read-only memory such as an EPROM or EEPROM of the device in the form of program code, and when the program code is executed by a processor, the corresponding functions of the module can be achieved. In addition, certain functions of the module may need to be implemented by separate hardware or in cooperation with the hardware. For example, the detection function is implemented by a corresponding sensor (such as a proximity sensor, an acceleration sensor, a gyroscope, etc.), the signal emission function is implemented by a corresponding communication device (such as a Bluetooth device, an infrared communication device, a baseband communication device, a Wi-Fi communication device, etc.), the output function is implemented by a corresponding output device (such as a display, a speaker, etc.), and so on.

[0081] Traditional offline 3D object detection algorithms usually process point cloud sequence data based on a modular design. Figure 2 The flowchart of an offline 3D object detection algorithm in the prior art is shown, taking the 3DAL proposed by Waymo (Charles R. Qi, Yin Zhou, Mahyar Najibi, Pei Sun, Khoa Vo, Boyang Deng, Dragomir Anguelov. Offboard 3D Object Detection from Point Cloud Sequences. In CVPR, June 2021) as an example. As Figure 2 shown, the existing offline 3D object detection algorithm includes the following steps:

[0082] The detection module (3D Object Detection) takes N consecutive frames of point cloud sequence (Point Cloud Sequence) as input and outputs the 3D object bounding boxes and corresponding categories contained in each frame.

[0083] The multi-object tracking module (3D Multi-Object Tracking) associates the objects detected in each frame to form an object sequence and outputs its corresponding unique object ID. For each object sequence, the object point cloud corresponding to it in the original point cloud frame is extracted, and then the ego-vehicle motion is eliminated and stitched together.

[0084] The motion state estimation module (Track-based Motion State Classification) can determine the motion state (static or dynamic) of an object according to the trajectory characteristics of the object.

[0085] The object-centric auto-labeling module extracts the temporal features of dynamic / static objects respectively according to the motion state predicted by the motion state estimation module to predict accurate bounding boxes. The optimized 3D bounding boxes will finally be transmitted back to the coordinate system of each frame where the object appears through the pose matrix (Pose).

[0086] Although the performance of existing online 3D detection algorithms has reached a relatively high level, when combined with multi-object tracking algorithms based on the "detect-then-track" rule, it is very easy to generate serious problems such as trajectory fragmentation, ID switching, and incorrect association, which will hinder the generation of complete temporal context features corresponding to an object.

[0087] Figure 3Shows a schematic diagram of a sequence generated by combining an existing 3D detection algorithm and a multi-object tracking algorithm. As Figure 3 shown, the first row 301 shows the ground truth sequence of an object. The second row 302 shows the situation of trajectory segmentation. Due to the problem of ID switching, the object is split into 3 sequences, namely T1, T2, and T3. The third row 303 shows the situation where there are more false detection segments (FP1, FP2) at the head and tail of the object sequence T4. The fourth row 304 shows the situation of the incomplete sequence T5, where some detection segments (M1, M2) are lost at the head and tail. When these problematic sequences are input into the motion state estimation module and the automatic annotation module of the offline detection algorithm, unreasonable processing situations will occur: the better-optimized boxes in the T1 sequence cannot be updated and propagated to T2 and T3 (this is because the IDs of T1, T2, and T3 are different and will be considered as different objects); the scores of the false detection segments (FP1, FP2) in the T4 sequence will become higher, so the boxes with these errors cannot be removed; the better-optimized boxes in the T5 sequence also cannot be updated and propagated to the positions where the bounding boxes are not detected.

[0088] In addition, the automatic annotation model based on motion state classification does not fully utilize the commonality of object temporal features. The commonality of object temporal features means that, for example, the size of an object remains consistent over time. By capturing data from different angles, the point cloud of the object can be made denser, thereby achieving more accurate size estimation; and the object trajectory is independent of its shape and size and always follows kinematic constraints within continuous time, which is manifested as the smoothness of the trajectory.

[0089] Figure 4 Shows a schematic diagram of predicting a bounding box by a dynamic object optimization mechanism based on a sliding window. As Figure 4 shown, the dynamic object optimization mechanism based on a sliding window fails to use the complete temporal context information (such as the relationship between local position and global trajectory, the consistency of object geometry, etc.). In the example of 401, it can be seen that for this moving object, the points in the point cloud of several adjacent frames at time t1 are very sparse, resulting in an inaccurate size of the predicted bounding box, and at time t2, a bounding box with a suitable size is output because the point cloud is denser. And in the example of 402, it can be seen that by aggregating all the point clouds of the object together, an accurate-size bounding box can be predicted for each frame.

[0090] To at least partially solve the above problems in the prior art, the present invention proposes an upstream algorithm module including a multi-frame 3D detection algorithm and an offline tracking algorithm to ensure the integrity and continuity of object tracking while maintaining a high recall rate. In addition, the existing automatic annotation model based on sliding windows does not fully utilize the commonalities of object temporal features (for example, the size of an object remains consistent over time; by capturing data from different angles, we can make the point cloud of the object denser, thereby achieving more accurate size estimation; the object trajectory is independent of its shape and size and always follows kinematic constraints over continuous time, which is manifested as the smoothness of the trajectory). And these commonalities are the basis for the present invention to utilize long temporal point clouds through a decomposition-based regression prediction scheme. Specifically, in the present invention, refined geometric sizes, smooth trajectory positions, and updated confidence scores can be performed separately.

[0091] The offline 3D object detection algorithm proposed by the present invention can achieve high-recall object detection and tracking in the upstream module and high-precision optimization based on long-term features in the downstream module. The upstream algorithm module including the multi-frame 3D detection algorithm and the offline tracking algorithm can ensure the integrity and continuity of target tracking while maintaining a high-recall object sequence. Next, the features of long temporal point clouds can be utilized through a decomposition-based regression prediction method to separately predict the refined geometric sizes of objects, smooth their motion trajectory positions, and update confidence scores.

[0092] Figure 5 The flowchart of an offline three-dimensional object detection method in an embodiment of the present invention is shown. As Figure 5 shown, the method includes the following steps:

[0093] Step 501, perform object detection on a multi-frame point cloud sequence of an object to generate an object bounding box.

[0094] Step 502, perform offline object tracking on the object bounding box to generate an object trajectory.

[0095] Step 503, extract object sequence data from the object trajectory.

[0096] Step 504, generate an optimized bounding box through object attribute prediction according to the object sequence data.

[0097] Step 505, transmit the optimized bounding box back to the coordinate system of each frame where the object appears.

[0098] Figure 6 The framework diagram of an offline three-dimensional object detection method in an embodiment of the present invention is shown. As Figure 6 shown, first, an accurate and complete object trajectory is generated through the upstream object detection and offline tracking modules.

[0099] During the object detection process, the existing CenterPoint detector is used as the basic detector. CenterPoint can output the bounding boxes of objects in two stages. The first stage is to generate candidate initial bounding boxes, and the second stage is to predict the bounding boxes of objects based on the center points. Specifically, in the first stage, the point cloud is converted into a voxel representation, then 3D sparse convolution (3DSparseConvolution) and RPN are used to extract features, and a dense output head (CenterHead) is used to generate candidate bounding boxes. In the second stage, Rol-Head is used to classify and regress each candidate bounding box to predict more accurate bounding boxes and confidence levels of objects. Since the output head of the CenterPoint detector without anchor box design will predict dense and redundant object bounding boxes, in order to provide accurate prediction results as much as possible, the present invention strengthens it in the following aspects: using the combination of five-frame point clouds as the input to maximize performance without performance degradation; using a two-stage module that fuses the original point cloud features and voxel features using the point cloud density information to preliminarily optimize the boundary results of the first stage; using techniques such as inference-stage data augmentation (TTA) and multi-model result (different resolutions, network structures, and capacities) fusion to improve the adaptability of the model to complex environments.

[0100] During the process of offline object tracking, since when the detection algorithm focuses on the performance at the bounding box level, the existing online multi-object trackers based on the "detect first then track" principle always perform struggling. To cope with the situation of a large number of redundant boxes, the multi-object tracking algorithm of the present invention proposes to reduce the possibility of incorrect matching through a two-stage data association strategy, which includes:

[0101] In the first stage, the bounding boxes are divided into a high-confidence group and a low-confidence group according to the confidence scores of the detected bounding boxes. The existing object trajectories are associated with the high-confidence group, and the successfully associated bounding boxes are used to update the existing object trajectories. The unupdated object trajectories are further associated with the low-confidence group, and the boxes that are not successfully associated are removed. In the present invention, it is allowed that the life cycle of an object continues indefinitely until the point cloud sequence terminates, and then any redundant boxes that have not been updated will be deleted, which is beneficial to reconnecting the broken object trajectories and can effectively reduce the problem of ID switching.

[0102] In the second stage, the tracking algorithm is executed again in reverse chronological order to generate another set of trajectories. The first set of trajectories is associated with the second set of trajectories through position-related similarity scores. The successfully matched trajectories are fused through the WBF (Weighted Box Fusion) strategy. Through the above steps, the problem of box loss can be further improved, and the motion state of the object can be stabilized. This process is called forward and reverse order tracking fusion. In addition, the boxes of shorter object trajectories and redundant boxes that have not been updated are directly merged into the final output without further downstream optimization.

[0103] For the generated object trajectories, it is necessary to extract object sequence data. For a given object trajectory (distinguished by a unique object ID), first, the region of interest (Rol) of the bounding box is enlarged along three dimensions, which can compensate for certain context information; then, all the point clouds within the regions enclosed by these enlarged bounding boxes are extracted; finally, these point cloud sequences, their corresponding tracking box sequences, and confidence scores are saved.

[0104] The above-extracted object sequence data is input into the object optimization module based on attribute prediction to generate a reconstructed box sequence.

[0105] Traditional object-centered automatic annotation models use motion-state-based strategies to optimize the bounding boxes generated by upstream modules. This method not only passes down the influence of incorrect motion states but also ignores the potential feature similarities between objects (for example, for rigid objects, regardless of their motion states, their geometric shapes do not change significantly over consecutive time periods; in addition, the motion states of objects usually exhibit a regular pattern and maintain strong consistency at adjacent moments). Based on the above findings, the present invention decomposes the traditional bounding box regression task into three different modules to predict the geometric shape, position, and confidence attributes of objects respectively.

[0106] Among them, the geometric shape optimization model can supplement the appearance and shape of the object by obtaining multiple viewpoints of the object. First, a local coordinate transformation operation is performed to align the object point cloud with the local box coordinates at different positions, and then all the point clouds from different frames are directly merged. For example, 4096 points can be randomly selected from them for further processing.

[0107] Figure 7 Fig. shows a schematic diagram of processing points and their initial bounding boxes during the geometric shape optimization process in an embodiment of the present invention. As Figure 7As shown, for each point and its corresponding initial bounding box, the point-to-plane method is used to calculate the projection distances between each point and the six surfaces, and then these distances are concatenated to each point coordinate, so as to better represent the information of the initial bounding box.

[0108] Next, randomly select t samples from the entire object trajectory, and each sample has 256 randomly selected points. These points are then concatenated not only with the distances to the six surfaces of their respective bounding boxes, but also with their respective confidence scores. Then, an encoder is used to extract features for each sample, which are used to initialize the geometric query vector (Query). Then, another encoder can be used to extract the features of 4096 points as the global point cloud dense features.

[0109] Figure 8 Fig. shows a schematic processing flow of a geometric query vector in an embodiment of the present invention. As Figure 8 shown, the geometric query vector is first input into the multi-head self-attention layer to encode the rich context relationships and feature dependencies between the selected samples, and refine their respective geometric information. The updated geometric query vector and the global point cloud dense features are sent to the multi-head cross-attention layer to infer the differences between each geometric query vector and the global dense point cloud features, so as to supplement the features of the required perspective. And for better residual target regression, the size of the initial box can be mapped to the same dimension as the geometric query vector, and then the initial box is added to the geometric query vector.

[0110] During the process of the position optimization model predicting the position of an object, for the jth object, a box position can be randomly selected from its tracking trajectory sequence as the new local coordinate system, and then all other boxes are transformed into this coordinate system, and the corresponding object point cloud is also transferred to this coordinate system. Then, a fixed number of point clouds are randomly selected for each frame in the object sequence.

[0111] Figure 9 Fig. shows a schematic diagram of the processing of points and the corners of their corresponding boxes during the position optimization in an embodiment of the present invention. As Figure 9 shown, for each selected point cloud, in addition to calculating the distance to the center of the box, the relative distances between each point and the eight corners of the corresponding box can also be calculated, which will generate a 27-dimensional feature vector. The final position-aware point cloud feature can be represented as a concatenated vector of the original local coordinates of the point cloud, the point cloud intensity, and the distance feature. For the convenience of training, the trajectories of all objects can be zero-padded to the same length.

[0112] For the trajectory of an object, a structure similar to the encoder in the geometric optimization model can be used to generate the position query vector (Query) corresponding to each frame, and its features include position-aware features and corresponding confidence scores. At the same time, another encoder can be used to extract the point features of the entire object trajectory as the key (Key) feature and value (Value) feature in the calculation of the attention mechanism.

[0113] Figure 10 FIG. shows a schematic diagram of the processing flow of a position query vector in an embodiment of the present invention. As Figure 10 shown, the position query vector is first fed into the self-attention mechanism module to calculate the relative distance between the current position and other positions. In addition, a one-dimensional mask can be applied near the position of each position query vector to constrain the self-attention. Subsequently, the local position query vector and the global point trajectory key feature and value feature are input into the cross-attention module to simulate the context relationship from local to global positions. Finally, the offset between each ground truth center and the corresponding initial center in the local coordinate system and the heading angle difference can be predicted.

[0114] Since the upstream detection and offline tracking modules of the present invention are encouraged to generate sufficient object trajectories, even after being processed by the geometric optimization and position optimization models, these trajectories still contain many incorrect bounding boxes. To solve this problem, the present invention uses a confidence optimization model consisting of two branches to optimize the confidence scores. The first classification branch is similar to a traditional second-stage object detector, which determines true positive examples (TP) or false positive examples (FP) by updating the scores. Among them, negative labels are assigned to the tracking boxes with an overlap degree (IoU, Intersection over Union) lower than the threshold corresponding to the ground truth box, the tracking boxes with an overlap degree higher than the threshold are regarded as positive samples, and the boxes that do not contribute to the classification target are ignored. The second classification branch can predict how much overlap an object should have after being optimized, so the regression target can be set to the overlap degree between the ground truth box and the optimized box after geometric and position optimization.

[0115] In the present invention, the object point cloud can be first processed using the same encoder network as in the geometric optimization model, and the extracted point cloud features are fused by a simple multilayer perceptron (MLP), and then input into the above two branches to predict their respective scores. During the training process, the pre-divided positive and negative target trajectories can be randomly sampled at a ratio of 1:1 in each epoch to achieve better convergence. The final score is the geometric mean of the two branches.

[0116] In one embodiment of the present invention, sufficient experiments and verifications are carried out on the Waymo dataset, one of the largest existing publicly available autonomous driving datasets. This dataset contains a total of 1150 point cloud scenes, among which 798 are used for model training, 202 for verification, and 150 for testing. This dataset provides 20 seconds of point cloud data for each scene, with a sampling frequency of 10Hz, and provides 3D annotations of 4 object categories in a 360-degree field of view. During the experiment, this embodiment follows the evaluation protocol with official metrics, namely mean average precision (AP) and average precision weighted by heading angle (APH), and reports the results for the difficulty levels of LEVEL 1 (L1) and LEVEL 2 (L2). The L1 difficulty includes objects with more than five point clouds, and the L2 difficulty only includes 3D object annotations with at least one and no more than five point clouds. mAPH (L2) is the main metric for ranking in the Waymo 3D detection challenge.

[0117] Table 1 shows the comparison results between the present invention and the existing state-of-the-art detection methods. As shown in Table 1, the present invention has achieved the best results on the Waymo 3D detection challenge leaderboard, with a detection performance of 85.15 mAPH (L2). During the comparison with methods that process long-term continuous point clouds (at least 100 frames), the present invention (Detzero) outperforms 3DAL in the Vehicle category by 5.93 mAPH (L1) and 9.51 mAPH (L2), exceeds INT's mAPH in the Vehicle category by 6.16 (L1) and 7.69 (L2), and has an mAPH advantage of 7.65 (L1) and 9.09 (L2) in the Pedestrian category. It can be seen that the present invention has a powerful ability to perform offline perception using long-time sequential continuous point clouds. In addition, compared with the state-of-the-art multi-modal fusion 3D detectors, the present invention also has strong performance advantages, with at least a lead of 3.43 (L1) and 4.63 (L2) mAPH in the Vehicle category, and at least a lead of 2.93 (L1) and 3.54 (L2) mAPH in the Pedestrian category. From these results, it can be seen that the present invention has explored the great potential of the point cloud sequence.

[0118] Table 1

[0119]

[0120] Table 2 shows the comparison results between the present invention and the existing state-of-the-art 3D detection methods among internal modules on the validation set. As shown in Table 2, the present invention has great advantages over other single-frame and multi-frame based methods in both Vehicle and Pedestrian categories. This is because the upstream module of the present invention can generate high-quality object trajectories, resulting in significant internal improvements in the complete model of the present invention, leading by at least 6.49 (L1) and 7.68 (L2) mAPH in the Vehicle category and at least 3.99 (L1) and 4.67 (L2) mAPH in the Pedestrian category.

[0121] Table 2

[0122]

[0123] Table 3 shows the performance comparison between the present invention and the prior art on the Waymo 3D tracking leaderboard. As shown in Table 3, the present invention ranks first with a leading advantage of 9.97 MOTA (L2).

[0124] Table 3

[0125]

[0126] Table 4 shows the comparison results between the present invention and human annotation capabilities. As shown in Table 4, the average AP performance of the proposed solution in 5 selected sequences is reported according to the experimental settings of 3DAL, using the common 3D AP@0.7 metric. Compared with human annotation and 3DAL, the present invention achieves gains of 3.79 and 4.87, and the gap is even larger in the more stringent 3D AP@0.8 metric. By ignoring the height and using BEV AP, the present invention obtains advantages similar to those of 3D AP. This is the first time that the performance of an offline 3D detection model is better than human annotation.

[0127] Table 4

[0128]

[0129] The present invention can provide high-quality automatic annotation to replace manual annotation for online 3D detection model training. Another in-domain semi-supervised learning experiment was conducted in another embodiment of the present invention. The single-stage CenterPoint was selected as the student model, which uses single-frame point cloud as input and does not use GT-Paste data augmentation during training. First, 10% of the sequences (79) were randomly selected from the Waymo dataset training set to train the entire algorithm framework. Next, "automatic annotations" can be inferred and generated for the remaining 90% of the sequences (719) in the training set. Then, the student model was trained with different combinations of human annotations and "automatic annotations".

[0130] Table 5 shows the results of using the present invention for automatic annotation instead of manual annotation. As shown in Table 5, it can be seen that when the manual annotation is reduced to 10%, the performance of the student model is reduced by 8.53 AP and 8.6 APH in the Vehicle category, and by 10.38 AP and 11.5 APH in the Pedestrian category (the first two rows). When the other 90% of the automatic annotation is added, the performance on Vehicle increases by 7.56 AP and 7.63 APH, and the performance of Pedestrian increases by 11.79 AP and 12.36 APH, and is higher than the performance of 100% manual annotation. In addition, when 10% of the manual annotation is deleted (the third row), the result is predictably slightly lower than that of the fourth row, but still shows a gain of 1.06 AP and 0.23 APH in the Pedestrian category. It can be seen that the "automatic annotation" generated by the present invention can be used to train an online model.

[0131] Although the embodiments of the present invention have been described above, it should be understood that they are presented only as examples and not as limitations. It will be apparent to those skilled in the relevant art that various combinations, variations, and changes can be made thereto without departing from the spirit and scope of the present invention. Therefore, the width and scope of the present invention disclosed herein should not be limited by the exemplary embodiments disclosed above, but should be defined only by the appended claims and their equivalents.

Claims

1. An offline three-dimensional object detection method, characterized in that Including the following steps: Performing object detection on a multi-frame point cloud sequence of an object to generate an object bounding box; Performing offline object tracking on the object bounding box to generate an object trajectory, including the following steps: Dividing the object bounding boxes into a high-confidence group of bounding boxes and a low-confidence group of bounding boxes according to the confidence scores of the object bounding boxes; Generating a first trajectory in the first stage of offline object tracking, which includes: Associating existing object trajectories with the high-confidence group of bounding boxes; Updating the existing object trajectories according to the successfully associated bounding boxes; and Associating the unupdated object trajectories with the low-confidence group of bounding boxes to generate a first trajectory, where the un-successfully associated object bounding boxes are removed, where the default lifespan of the first trajectory is infinite, and after the association process ends, the moment of the last actual association with the object bounding box is selected as the final length of the first trajectory of the object, and subsequent continuous false associations are removed; Generating a second trajectory in the second stage of offline object tracking in the reverse chronological order of the first stage; Associating the first trajectory and the second trajectory through a position-related similarity score; and Fusing the successfully associated first trajectory and the second trajectory to generate an object trajectory; Extracting object sequence data from the object trajectory; Generating an optimized bounding box through object attribute prediction according to the object sequence data, where object attribute prediction includes geometric shape prediction, position prediction, and confidence prediction of the object; and Transmitting the optimized bounding box back to the coordinate system where the object appears.

2. The offline three-dimensional object detection method according to claim 1, wherein Performing object detection on the point cloud sequence of the object to generate an object bounding box includes the following steps: Inputting the point cloud sequence into a center point detector, where the point cloud sequence includes a combination of multiple five-frame point clouds; Converting the point cloud sequence into a voxel representation and generating candidate initial bounding boxes in the first stage of object detection; And Fine-tuning the candidate initial bounding boxes in the second stage of object detection to predict a more accurate bounding box and confidence of the object, where the point cloud density information of the object is used in the second stage of object detection for raw point cloud feature and voxel feature fusion; Wherein inference-stage data augmentation and multi-model result fusion are performed in the above steps.

3. The offline three-dimensional object detection method according to claim 1, wherein Extracting object sequence data from the object trajectory includes: Magnifying the region of interest of the object bounding box in three dimensions; Extracting the point cloud sequence within the region enclosed by the magnified object bounding box; and Saving the point cloud sequence within the region enclosed by the magnified object bounding box, its tracking box sequence, and the confidence score.

4. The offline three-dimensional object detection method according to claim 1, wherein The geometric shape prediction of the object includes: Randomly select candidate objects from different perspectives in the object sequence, randomly select points for each candidate object, extract the corresponding features using the first encoder, and generate geometric query vectors; Overlay the point clouds of each object in the entire object sequence, randomly select points from them, and use a second encoder to generate global point cloud dense features; Inputting a geometric query vector into a multi-head self-attention layer to encode the context relationship and feature dependency between samples and extract geometric information; Feeding the updated geometric query vector and the global point cloud dense feature into a multi-head cross-attention layer to infer the difference between each geometric query vector and the global dense point cloud feature, and further supplementing the features of the required perspective; and Input the updated geometric query vector into the prediction head and output geometric prediction results, and average the geometric prediction results to obtain the final geometric prediction result.

5. The offline three-dimensional object detection method according to claim 4, wherein The position prediction of the object includes: Randomly select points for each object candidate in the object sequence, and use the third encoder to extract the features corresponding to each object to generate position query vectors; ​ Using a fourth encoder to generate point features of the entire object trajectory, where the point features of the entire object trajectory are used as the key features and value features in the cross-attention mechanism calculation; Feed the position query vector into the self-attention mechanism module to calculate the relative distance between the current position and other positions, where a one-dimensional mask is used to constrain the self-attention near the position of each position query vector; Feed the local position query vector and the point features of the entire object trajectory into the cross-attention module to simulate the context relationship from local to global positions; and Predict the offset between each ground truth center and the corresponding initial center in the local coordinate system and the heading angle difference.

6. The offline three-dimensional object detection method according to claim 5, wherein The confidence prediction of the object includes: In the first branch of the object confidence prediction, classify the object as a true positive or a false positive according to the overlap degree between the tracking box and the ground truth box of the object; and In the second branch of the object confidence prediction, predict the overlap degree that an object should have after being optimized, and set the regression target to the overlap degree between the ground truth box and the optimized bounding box after geometric and position optimization.

7. An offline three-dimensional object detection system, characterized in that, Includes: An object detection module configured to perform object detection on a multi-frame point cloud sequence of an object to generate an object bounding box; An offline object tracking module configured to perform offline object tracking on the object bounding box to generate an object trajectory, including the following steps: Classify the object bounding boxes into high-group bounding boxes and low-group bounding boxes according to the confidence scores of the object bounding boxes; Generate a first trajectory in the first stage of offline object tracking, which includes: Perform data association between the existing object trajectories and the high-group bounding boxes; Update the existing object trajectories according to the successfully associated bounding boxes; and Perform data association between the unupdated object trajectories and the low-group bounding boxes to generate a first trajectory, where the unsuccessfully associated object bounding boxes are removed, and the default life cycle of the first trajectory is infinite. After the association process ends, select the moment of the last actual association to the object bounding box as the final length of the first trajectory of the object, and remove subsequent continuous false associations; Generate a second trajectory in the second stage of offline object tracking in the reverse chronological order of the first stage; Associate the first trajectory and the second trajectory through a position-related similarity score; and Fuse the successfully associated first trajectory and the second trajectory to generate an object trajectory; An object sequence data extraction module configured to extract object sequence data from the object trajectory; An object optimization module based on attribute prediction configured to generate an optimized bounding box through object attribute prediction according to the object sequence data; and A coordinate transformation module configured to transmit the optimized bounding box back to the coordinate system where the object appears.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program, when executed by a processor, executes the steps of the method according to one of claims 1-6.

9. A computer system, characterized in that, Includes: A processor configured to execute machine-executable instructions; And A memory having stored thereon machine-executable instructions that, when executed by the processor, execute the steps of the method according to one of claims 1-6.

Citation Information

Patent Citations

  • Target detection and tracking method based on sensor fusion

    CN114926808A

  • Video object tracking

    US20190347806A1