A 3D multi-target tracking method and system based on motion state prediction

By using a 3D multi-target tracking method based on motion state prediction, combined with 3D target detection and data association algorithms, the problem of fast and reliable multi-target tracking in autonomous driving systems is solved, stable tracking and efficient prediction are achieved in complex environments, and the perception capability of autonomous driving systems is improved.

CN120260017BActive Publication Date: 2025-09-09CHONGQING SELIS PHOENIX INTELLIGENT INNOVATION TECH CO LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510754074.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-09-09
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

Existing 3D multi-target tracking methods have difficulty in quickly and reliably tracking multiple surrounding targets in autonomous driving systems, especially in complex dynamic environments, where there are problems of prediction errors and unstable trajectory management.

Method used

A 3D multi-target tracking method based on motion state prediction is adopted. By obtaining the current point cloud information and historical tracking status, combining 3D target detection, motion prediction and data association, and using the matching cost calculation of 3D-CIoU and Mahalanobis distance and the Kuhn-Munkres optimal allocation algorithm, robust association and stable tracking of multiple targets are achieved.

Benefits of technology

It improves the robustness of environmental perception in complex traffic scenarios, enhances the real-time performance and accuracy of the autonomous driving system, provides reliable moving target trajectory prediction capabilities for downstream algorithms, and improves the development and operation efficiency of the autonomous driving system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260017B_ABST
    Figure CN120260017B_ABST
Patent Text Reader

Abstract

The present invention discloses a 3D multi-target tracking method and system based on motion state prediction, which includes: S1, obtaining current point cloud information and historical tracking state information during vehicle driving; S2, processing the current point cloud information to obtain a detection bounding box and its corresponding confidence level; S3, processing the historical tracking state information to obtain a predicted bounding box corresponding to the tracked target; counting the number of detection bounding boxes, and based on the number of detection bounding boxes, the corresponding confidence levels of the detection bounding boxes, and the stored historical matching of the tracked targets, processing the detection bounding boxes and the predicted bounding boxes corresponding to the tracked targets, obtaining and outputting matching pairs and / or unmatched tracking indexes and unmatched detection indexes, and simultaneously updating the historical tracking state information. The present invention can quickly and reliably track multiple surrounding targets, provide a basis for the implementation of downstream algorithms for autonomous driving, and improve the development, testing, and operation efficiency of autonomous driving systems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of autonomous driving, and in particular relates to a 3D multi-target tracking method and system based on motion state prediction. Background Art

[0002] Autonomous driving systems primarily consist of four subsystems: perception, decision-making, planning, and control. The perception subsystem is responsible for accurately collecting and processing information about the surrounding environment. Subsequent decision-making, planning, and control tasks rely on the environmental data it provides. Therefore, the perception system is crucial in autonomous driving. 3D multi-object tracking is a core task of the perception subsystem, aiming to accurately track multiple objects (such as pedestrians and vehicles) in the surrounding environment in real time to support subsequent safety decisions. 3D multi-object tracking includes key components such as object detection, motion prediction, data association, and trajectory management. Object detection uses sensors (such as lidar) to identify and locate dynamic objects; motion prediction predicts the future position of objects based on historical data; data association ensures accurate matching of objects between consecutive frames; and trajectory management is responsible for creating, updating, maintaining, optimizing, and deleting trajectories. Through the collaborative work of these tasks, 3D multi-object tracking enables efficient and accurate object tracking in complex and dynamic environments, providing reliable support for autonomous driving systems. Therefore, how to quickly and reliably track multiple objects in the surrounding environment becomes a key challenge within the perception subsystem. Summary of the Invention

[0003] The purpose of the present invention is to provide a 3D multi-target tracking method and system based on motion state prediction, so as to quickly and reliably track multiple surrounding targets.

[0004] In a first aspect, the 3D multi-target tracking method based on motion state prediction of the present invention comprises the steps of:

[0005] S1. Obtain the current point cloud information and historical tracking status information of the vehicle during driving.

[0006] S2. Process the current point cloud information to obtain the detection bounding box and its corresponding confidence level (i.e., 3D target detection result).

[0007] S3. Process the historical tracking state information to obtain the predicted bounding box corresponding to the tracking target (i.e., the motion prediction result).

[0008] S4. Count the number of detection bounding boxes, and based on the number of detection bounding boxes, the confidence corresponding to the detection bounding boxes, and the stored historical matching of the tracking targets, process the detection bounding boxes and the predicted bounding boxes corresponding to the tracking targets, obtain and output matching pairs and / or unmatched tracking indexes and unmatched detection indexes, and update the historical tracking status information.

[0009] Optionally, step S2 specifically includes:

[0010] S21. Extracting point cloud feature information from the current point cloud information.

[0011] S22: Process the point cloud feature information to obtain a primary bounding box (ie, ROI area, region of interest).

[0012] S23. Process the primary bounding box to obtain point cloud attention-weighted feature information.

[0013] S24. Decode the point cloud attention weighted feature information to obtain a bounding box correction value, superimpose the bounding box correction value on the primary bounding box to obtain the detection bounding box; splice the bounding box correction value with the (original) point cloud attention weighted feature information and then decode it to obtain the confidence level corresponding to the detection bounding box.

[0014] By adopting the above method, high-precision 3D target detection result output is achieved, and the positioning accuracy of the detection bounding box is significantly improved.

[0015] Optionally, in S22, the point cloud feature information is processed to obtain a primary bounding box in a specific manner including:

[0016] S221. Generate multiple anchor boxes of different scales and aspect ratios at each location of the feature map composed of the point cloud feature information. The point cloud feature information can form a series of feature maps, and each anchor box represents a possible target area.

[0017] S222: Perform object confidence probability prediction and anchor frame correction amount prediction for each anchor frame to obtain the anchor frame correction amount.

[0018] S223: Superimpose the anchor frame correction amount on the anchor frame to obtain a primary bounding frame.

[0019] Optionally, in S23, the specific method of processing the primary bounding box to obtain the point cloud attention weighted feature information includes:

[0020] S231 , mapping the primary bounding box to a voxel space, dividing the voxel space into a plurality of voxel grids, locating the center point of each voxel grid, and obtaining a plurality of voxel grid center points.

[0021] S232 , performing feature encoding on all feature points in multiple spaces with different radii from the center point of each voxel grid, and generating multi-scale features for each feature point.

[0022] S233. Perform point cloud density self-attention calculation on the multi-scale features to obtain point cloud attention weighted feature information.

[0023] Optionally, step S3 specifically includes:

[0024] S31. Perform spatial long-term time series modeling on the historical tracking state information based on transformer coding to obtain the historical tracking state coding corresponding to the tracking target.

[0025] S32. Update the hidden state of the GRU model corresponding to the tracking target at the current moment using the historical tracking state code corresponding to the tracking target and the hidden state of the GRU model corresponding to the tracking target at the previous moment, thereby obtaining the hidden state of the GRU model corresponding to the tracking target at the current moment; wherein, the hidden state of the GRU model corresponding to the tracking target at the initial moment (i.e., the initial value of the hidden state of the GRU model corresponding to the tracking target) is an all-zero vector.

[0026] S33: Process the hidden state of the GRU model corresponding to the tracking target at the current moment and the past H time steps to obtain the predicted bounding box corresponding to the tracking target.

[0027] It adopts a dual memory architecture based on transformer encoding (internal memory) and GRU model (external memory), and effectively captures the spatiotemporal correlation of the motion trajectory of the tracked target through the fusion of long-term and short-term state encoding, thereby enhancing the continuity and accuracy of motion prediction and reducing prediction errors in dynamic scenes.

[0028] Optionally, in S33, the specific method of processing the hidden state of the GRU model corresponding to the tracking target at the current moment and the past H time steps to obtain the predicted bounding box corresponding to the tracking target includes:

[0029] S331. Perform softmax classification on the hidden state of the GRU model corresponding to the tracking target at the current moment and the past H time steps to obtain the temporal space weight corresponding to the tracking target at the current moment and the past H time steps.

[0030] S332. Use the temporal spatial weight to perform a weighted sum on the hidden states of the GRU model corresponding to the tracking target at the current moment and the past H time steps, and then obtain the GRU model output corresponding to the tracking target through transformer encoding.

[0031] S333. The GRU model output corresponding to the tracking target is decoded by three layers of MLP to obtain the pose information of the predicted bounding box corresponding to the tracking target. Then, the predicted bounding box corresponding to the tracking target is obtained by combining the shape information of the bounding box corresponding to the most recently matched tracking target.

[0032] Optionally, step S4 specifically includes:

[0033] S41: Count the number of detection bounding boxes, determine the status of each detection bounding box based on the number of detection bounding boxes and the confidence level of the detection bounding boxes, delete the detection bounding boxes in the ignored state, and then execute S42. The status of the detection bounding box is high score, low score, or ignored.

[0034] S42: Determine the status of each tracking target based on the stored tracking target history matching, delete the predicted bounding box corresponding to the tracking target in the destroyed state, and then execute S43. The status of the tracking target can be activated, lost, or destroyed.

[0035] S43. Using the predicted bounding box corresponding to the active tracking target and the detected bounding box in the high-score state, perform a matching cost calculation of the 3D-CIoU and Mahalanobis distance, obtain the corresponding matching cost matrix of the 3D-CIoU and Mahalanobis distance, and then execute S44.

[0036] S44 , using the Kuhn-Munkres optimal allocation algorithm on the matching cost matrix to achieve bipartite matching between the predicted bounding box corresponding to the tracked target in the active state and the detected bounding box in the high-score state, and then executing S45 .

[0037] S45: Determine whether all tracking targets are matched successfully. If so, execute S48; otherwise, execute S46.

[0038] S46. Use the predicted bounding boxes corresponding to the remaining tracking targets (i.e., the tracking targets in the lost state and the unmatched tracking targets in the activated state) and the remaining detection bounding boxes (i.e., the detection bounding boxes in the low-scoring state and the detection bounding boxes in the unmatched high-scoring state) to calculate the matching cost of the 3D-CIoU and Mahalanobis distance, obtain the corresponding matching cost matrix of the 3D-CIoU and Mahalanobis distance, and then execute S47.

[0039] S47 , using the Kuhn-Munkres optimal allocation algorithm on the matching cost matrix to achieve bipartite matching between the predicted bounding boxes corresponding to the remaining tracking targets and the remaining detection bounding boxes, and then executing S48 .

[0040] S48: Output the matching pairs and / or unmatched tracking indexes and unmatched detection indexes, update the historical tracking state information, and then end the tracking.

[0041] During binary matching on a frame, the target matching results for that frame are recorded and stored, forming a historical target matching history. There are two types of target matching scenarios: a successful match, meaning the predicted bounding box corresponding to the target successfully matches a detection bounding box; and an unsuccessful match, meaning the predicted bounding box corresponding to the target does not match any detection bounding boxes.

[0042] By introducing a matching cost calculation strategy that combines 3D-CIoU and Mahalanobis distance, combined with the Kuhn-Munkres optimal allocation algorithm and trajectory lifecycle management (such as deleting the detection bounding boxes in the ignored state and the predicted bounding boxes corresponding to the tracked targets in the destroyed state), we can achieve robust association of multi-target motion trajectories in complex occlusion scenes, improving the stability of cross-frame target tracking and data association efficiency.

[0043] Optionally, in S41, the specific method of determining the status of each detection bounding box is:

[0044] When the number of detection bounding boxes is greater than a preset number threshold, the corresponding detection bounding boxes whose confidence is greater than or equal to a first preset confidence threshold are determined to be in a high score state, the corresponding detection bounding boxes whose confidence is greater than or equal to a second preset confidence threshold and less than the first preset confidence threshold are determined to be in a low score state, and the corresponding detection bounding boxes whose confidence is less than the second preset confidence threshold are determined to be in an ignored state. The second preset confidence threshold is less than the first preset confidence threshold.

[0045] When the number of detection bounding boxes is less than or equal to a preset number threshold, the state of the detection bounding box whose corresponding confidence is greater than or equal to the first preset confidence threshold is determined to be a high-scoring state, and the state of the detection bounding box whose corresponding confidence is less than the first preset confidence threshold is determined to be a low-scoring state.

[0046] Optionally, in S42, the specific manner of determining the status of each tracking target is:

[0047] If the number of consecutive successful matching frames in the target's most recent Fra frame matching (determined by the stored target's historical matching results) is greater than or equal to a first preset frame number threshold, the target is determined to be in an active state. If none of the target's most recent Fra frames have been successfully matched, the target is determined to be in a destroyed state. If the number of unsuccessful matching frames in the target's most recent Fra frame matching is less than Fra and the number of consecutive successful matching frames is less than the first preset frame number threshold, the target is determined to be in a lost state. Fra represents the second preset frame number threshold, which is greater than the first preset frame number threshold.

[0048] Optionally, in S43, the specific method of obtaining the corresponding matching cost matrix of 3D-CIoU and Mahalanobis distance is:

[0049] S431. Calculate the Mahalanobis distance between a predicted bounding box corresponding to an active tracking target and a detected bounding box in a high-score state.

[0050] S432: Project the predicted bounding box corresponding to the active tracking target and the detected bounding box in the high-resolution state to the BEV perspective, and calculate the area and height of the overlapping area of ​​the projections.

[0051] S433: Calculate the intersection-over-union ratio of the corresponding 3D object bounding box based on the area and height of the projection overlap region.

[0052] S434. Calculate the matching cost of the 3D-CIoU and Mahalanobis distance based on the Mahalanobis distance, the intersection over union ratio, the preset scaling factor, the IoU adjustment factor, the angle difference factor, and the diagonal distance of the minimum closed frame area; wherein the minimum closed frame area is the minimum area that simultaneously includes the predicted bounding box corresponding to the active tracking target and the detection bounding box of the high-scoring state.

[0053] S435. Traverse the possible combinations of the predicted bounding box corresponding to the active tracking target and the detection bounding boxes of all high-scoring states, and repeat S431 to S434 to obtain the corresponding matching cost matrix of 3D-CIoU and Mahalanobis distance.

[0054] Optionally, in S46, the specific method of obtaining the corresponding matching cost matrix of 3D-CIoU and Mahalanobis distance is:

[0055] S461: Calculate the Mahalanobis distance between a predicted bounding box corresponding to a remaining tracking target and a remaining detection bounding box.

[0056] S462: Project the predicted bounding box corresponding to the remaining tracking target and the remaining detection bounding box to the BEV perspective, and calculate the area and height of the overlapping area of ​​the projections.

[0057] S463: Calculate the intersection-over-union ratio of the corresponding 3D object bounding box based on the area and height of the projection overlap region.

[0058] S464: Calculate a matching cost of a combination of 3D-CIoU and Mahalanobis distance based on the Mahalanobis distance, the intersection over union (IoU), a preset scaling factor, an IoU adjustment factor, an angle difference factor, and a diagonal distance of a minimum enclosing box area. The minimum enclosing box area is the smallest area that simultaneously contains the predicted bounding box corresponding to the remaining tracked target and the remaining detection bounding box.

[0059] S465: Traverse the possible combinations of the predicted bounding box corresponding to the remaining tracking target and all remaining detection bounding boxes, and repeat S461 to S464 to obtain the corresponding matching cost matrix of 3D-CIoU and Mahalanobis distance.

[0060] In the second aspect, the 3D multi-target tracking system based on motion state prediction described in the present invention includes a memory and a controller, and the memory stores a computer-readable program. When the computer-readable program is called by the controller, it can execute the steps of the above-mentioned 3D multi-target tracking method based on motion state prediction.

[0061] The present invention adopts a full-process closed-loop optimization architecture of detection-prediction-association, integrating historical tracking status with current (real-time) point cloud information to form a 3D multi-target tracking system and method with both real-time performance and precision, thereby quickly and reliably tracking multiple targets around the vehicle, providing the autonomous driving perception module with continuous and reliable motion target trajectory prediction capabilities, significantly improving the robustness of environmental perception in complex traffic scenarios, providing a foundation for the implementation of downstream autonomous driving algorithms, facilitating engineering deployment, and improving the development, testing and operation efficiency of autonomous driving systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 This is a flow chart of a 3D multi-target tracking method based on motion state prediction in an embodiment of the present invention.

[0063] Figure 2 This is an architecture diagram of a 3D multi-target tracking system based on motion state prediction in an embodiment of the present invention.

[0064] Figure 3 3D target detection module in an embodiment of the present invention.

[0065] Figure 4 3D object detection module in an embodiment of the present invention.

[0066] Figure 5 This is a schematic diagram of the improved RPN unit in an embodiment of the present invention.

[0067] Figure 6 This is an execution flow chart of the improved RPN unit in an embodiment of the present invention.

[0068] Figure 7 4 is an execution flow chart of the point cloud density aggregation unit in an embodiment of the present invention.

[0069] Figure 8 Schematic diagram of point cloud density self-attention calculation in an embodiment of the present invention.

[0070] Figure 9 2 is a schematic diagram of a motion prediction module in an embodiment of the present invention.

[0071] Figure 10 FIG. 4 is an execution flow chart of the motion prediction module in an embodiment of the present invention.

[0072] Figure 114 is an execution flow chart of the motion state decoding unit in an embodiment of the present invention.

[0073] Figure 12 4 is an execution flow chart of the data association and trajectory management module in an embodiment of the present invention.

[0074] Figure 13 3D-CIoU and Mahalanobis distance combined matching cost calculation unit in an embodiment of the present invention.

[0075] Figure 14 Schematic diagram of the principle of converting a 3D object projected to a BEV perspective into a 2D object in an embodiment of the present invention. DETAILED DESCRIPTION

[0076] The following embodiments of the present invention are described in further detail with reference to the accompanying drawings and examples. The following examples are used to illustrate the present invention but are not intended to limit the scope of the present invention.

[0077] like Figure 1 As shown, the 3D multi-target tracking method based on motion state prediction in an embodiment of the present invention includes the following steps:

[0078] S1. Obtain the current point cloud information and historical tracking status information of the vehicle during driving.

[0079] S2. Process the current point cloud information to obtain the detection bounding box and its corresponding confidence level (i.e., 3D target detection result).

[0080] S3. Process the historical tracking state information to obtain the predicted bounding box corresponding to the tracking target (i.e., the motion prediction result).

[0081] S4. Count the number of detection bounding boxes, and based on the number of detection bounding boxes, the confidence corresponding to the detection bounding boxes, and the stored historical matching of the tracking targets, process the detection bounding boxes and the predicted bounding boxes corresponding to the tracking targets, obtain and output matching pairs and / or unmatched tracking indexes and unmatched detection indexes, and update the historical tracking status information.

[0082] The 3D multi-target tracking method based on motion state prediction in the embodiment of the present invention can be executed by any computing device with computing, processing and storage functions, such as a PC, a vehicle-mounted computing platform, or a server.

[0083] The 3D multi-target tracking system based on motion state prediction in an embodiment of the present invention includes a memory and a controller. The memory stores a computer-readable program. When the computer-readable program is called by the controller, it can execute the steps of the 3D multi-target tracking method based on motion state prediction in an embodiment of the present invention.

[0084] like Figure 2 As shown, in the embodiment of the present invention, the computer-readable program is divided into functional modules, including: an information acquisition module, a 3D target detection module, a motion prediction module, and a data association and trajectory management module.

[0085] The information acquisition module is used to obtain the current (real-time) point cloud information and historical tracking status information during vehicle driving.

[0086] The 3D object detection module is used to process the current point cloud information to obtain a detection bounding box and its corresponding confidence level. In some embodiments, the 3D object detection module includes: a VoxelNeXt point cloud feature extraction unit, an improved RPN unit, a point cloud density aggregation unit, and a bounding box correction unit.

[0087] The motion prediction module is used to process historical tracking state information to obtain a predicted bounding box corresponding to the tracked target. In some embodiments, the motion prediction module includes: an internal memory unit based on a transformer encoder, an external memory unit based on a GRU, and a motion state decoding unit.

[0088] The data association and trajectory management module is used to count the number of detection bounding boxes, and based on the number of detection bounding boxes, the confidence level corresponding to the detection bounding boxes, and the stored historical matching of the tracking target, process the detection bounding boxes and the predicted bounding boxes corresponding to the tracking target, obtain and output matching pairs and / or unmatched tracking indexes and unmatched detection indexes, and update the historical tracking status information. In some embodiments, the data association and trajectory management module includes: 3D-CIoU and Mahalanobis distance joint matching

[0089] Cost calculation unit, Kuhn-Munkres allocation unit and trajectory lifecycle management unit.

[0090] The following is a detailed description of the various steps of the 3D multi-target tracking method based on motion state prediction.

[0091] S1. Obtain the current point cloud information and historical tracking status information of the vehicle during driving.

[0092] In some embodiments, a vehicle is equipped with a LiDAR and an onboard computing platform. The LiDAR collects point cloud information, while the onboard computing platform stores historical tracking status information during vehicle travel. The embodiments of the present invention do not specifically limit the types of vehicles, LiDAR, and onboard computing platforms.

[0093] S2. Process the current point cloud information to obtain the detection bounding box and its corresponding confidence level (i.e., 3D target detection result).

[0094] like Figure 3 、 Figure 4 As shown, in some embodiments, step S2 specifically includes:

[0095] S21. Extracting point cloud feature information from the current point cloud information. Specifically, extracting features from the current point cloud information using a VoxelNeXt point cloud feature extraction unit to obtain point cloud feature information.

[0096] S22: Process the point cloud feature information to obtain a primary bounding box (i.e., ROI region, region of interest). Specifically, the point cloud feature information is processed by improving the RPN unit to obtain the primary bounding box.

[0097] like Figure 6 As shown, in some embodiments, the improved RPN unit performs the following steps:

[0098] S221. Generate multiple anchor boxes of different scales and aspect ratios at each location of the feature map composed of the point cloud feature information. The point cloud feature information can form a series of feature maps, and each anchor box represents a possible target area.

[0099] S222: Perform object confidence probability prediction and anchor frame correction amount prediction for each anchor frame to obtain the anchor frame correction amount.

[0100] S223: Superimpose the anchor frame correction amount on the anchor frame to obtain a primary bounding frame.

[0101] The prediction of the center point offset is added to the network structure training of the RPN unit to form an improved RPN unit. The improved RPN unit improves the model accuracy and optimizes the final quality of the primary bounding box. The principle of the improved RPN unit is as follows Figure 5 As shown in the figure, represents the anchor frame correction amount, Indicates the correction value on the negative half axis of X, Y, and Z axes. Indicates the correction value on the positive half axis of X, Y, and Z axes. Indicates the rotation correction value of the anchor box around the Z axis.

[0102] S23: Process the primary bounding box to obtain point cloud attention weighted feature information. Specifically, the primary bounding box is processed by a point cloud density aggregation unit to obtain point cloud attention weighted feature information.

[0103] like Figure 7 As shown, in some embodiments, the point cloud density aggregation unit performs the following steps:

[0104] S231 , mapping the primary bounding box to a voxel space, dividing the voxel space into a plurality of voxel grids, locating the center point of each voxel grid, and obtaining a plurality of voxel grid center points.

[0105] For example, the center point positioning of a single voxel grid is expressed as follows:

[0106] (1)

[0107] In formula (1), Represents the coordinates of the center point of the voxel grid, represents the set of all feature points falling within the voxel grid, , Represents the total number of feature points falling within the voxel grid, Represents the feature points in the voxel grid j The coordinates of is the feature point in the laser radar coordinate system j Coordinate values ​​on the X, Y, and Z axes.

[0108] S232 , performing feature encoding on all feature points in multiple spaces with different radii from the center point of each voxel grid, and generating multi-scale features for each feature point.

[0109] For example, the distance The radius is Feature points in the space l (also called original feature points) to perform feature encoding and obtain feature points l The encoding features , which manifests itself as:

[0110] (2)

[0111] In formula (2), Indicates the current feature point l The encoding is composed of feature points l Coordinates After three layers of MLP encoding, , Representing feature points l Relative to The offset characteristics of After three-layer MLP encoding, Represents the feature points l The kernel density feature is determined by the distance The radius is The kernel density estimate of each point in the space Obtained through three-layer MLP encoding.

[0112] Correspondingly, the feature points l Multi-scale features It is expressed as formula (3):

[0113] (3)

[0114] In formula (3), Indicates distance The radius is Feature points in the space l Perform feature encoding and obtain feature points l The encoding features of ; Indicates distance The radius is Feature points in the space l Perform feature encoding and obtain feature points l The encoding features of .

[0115] S233, perform point cloud density self-attention calculation on multi-scale features (see Figure 8 ), and obtain the point cloud attention weighted feature information.

[0116] For example, for feature points l Multi-scale features Perform point cloud density self-attention calculation to obtain point cloud attention weighted feature information The calculation is expressed as the following formulas (4) and (5):

[0117] (4)

[0118] (5)

[0119] In formulas (4) and (5), X is the original input feature of the current self-attention layer, and its dimension is ,Right now B batches to jointly calculate the point cloud density self-attention, each batch contains N feature points, and the feature of each feature point is a multi-scale feature. For example, the feature point l Multi-scale features , multi-scale features The dimension is F (adjusted by a fully connected layer), They are weight matrices that map the original input features to Query, Key, and Value features, respectively. These three weight matrices are all learnable matrices. Q 、 K 、 V They are the calculated Query features, Key features, and Value features respectively. express K The transpose of soft max() represents the normalized exponential function, Indicates the dimension of the key, is an adjustable hyperparameter, for example, .

[0120] S24. Decode the point cloud attention weighted feature information to obtain a bounding box correction value, superimpose the bounding box correction value on the primary bounding box to obtain a detection bounding box; concatenate the bounding box correction value with the (original) point cloud attention weighted feature information and then decode it to obtain the confidence level corresponding to the detection bounding box. Specifically, the bounding box correction unit uses a fully connected layer to decode the point cloud attention weighted feature information to obtain a bounding box correction value (including correction values ​​for the X and Y coordinate values ​​of the center point of the bounding box and correction values ​​for the overall heading angle of the bounding box), superimpose the bounding box correction value on the primary bounding box to obtain a detection bounding box; concatenate the bounding box correction value with the (original) point cloud attention weighted feature information and then decode it through a fully connected layer to obtain the confidence level corresponding to the detection bounding box.

[0121] S3. Process the historical tracking state information to obtain the predicted bounding box corresponding to the tracking target (i.e., the motion prediction result).

[0122] like Figure 9 、 Figure 10 As shown, in some embodiments, step S3 specifically includes:

[0123] S31. Perform spatial long-term time series modeling on the historical tracking state information based on transformer coding to obtain the historical tracking state coding corresponding to the tracking target.

[0124] Specifically, for a single tracking target i ,That t The motion state at the moment is expressed as , Indicates t-1 All the time M The motion state set of the tracking target, That is t The historical tracking status information at each moment. The motion status can be obtained by: t 3D target detection results at each moment , E express t The number of (tracked) targets detected at any moment, where Indicates the i Tracking targets in t The detection bounding box description at the moment, Indicates the i Tracking targets in t The center point coordinates of the detection bounding box at the moment, Respectively expressed in tThe first detected at the moment i The width, height, and length of the detection bounding box of the tracked target, Indicates t Moment i The heading angle of the detection bounding box of the tracked target. and , according to the difference of its coordinate values ​​on the X, Y, and Z axes and the time interval between two frames, its speed and acceleration on the X, Y, and Z axes can be determined, thus obtaining t Motion state at all times .in, Respectively represent i The speed of the tracking target on the X, Y, and Z axes, Respectively represent i The acceleration of the tracked target on the X, Y, and Z axes. Specifically, the historical tracking state information is input into the internal memory unit based on the transformer encoder, and the transformer encoder is used to perform spatial long-term time series modeling on the motion state description to obtain t Historical tracking status code corresponding to the target being tracked at all times .

[0125] S32. Update the hidden state of the GRU model corresponding to the tracking target at the current moment using the historical tracking state code corresponding to the tracking target and the hidden state of the GRU model corresponding to the tracking target at the previous moment, thereby obtaining the hidden state of the GRU model corresponding to the tracking target at the current moment; wherein, the hidden state of the GRU model corresponding to the tracking target at the initial moment (i.e., the initial value of the hidden state of the GRU model corresponding to the tracking target) is an all-zero vector.

[0126] Specifically, through the GRU-based external memory module, according to the current moment (i.e. t The historical tracking state code corresponding to the tracking target at that moment and the previous moment (i.e. t- 1 moment) Tracking the hidden state of the GRU model corresponding to the target , obtain the hidden state of the GRU model corresponding to the current tracking target .

[0127] S33. Process the hidden state of the GRU model corresponding to the tracking target at the current moment and the past H time steps (i.e., the most recent H moments before the current moment) to obtain a predicted bounding box corresponding to the tracking target.

[0128] Specifically, the motion state decoding unit processes the hidden state of the GRU model corresponding to the tracking target at the current moment and the past H time steps to obtain the predicted bounding box corresponding to the tracking target.

[0129] like Figure 11As shown, the motion state decoding unit performs the following steps:

[0130] S331. Perform softmax classification on the hidden state of the GRU model corresponding to the tracking target at the current moment and the past H time steps to obtain the temporal space weight corresponding to the tracking target at the current moment and the past H time steps.

[0131] S332. Use the temporal spatial weight to perform a weighted sum of the hidden states of the GRU model corresponding to the tracking target at the current moment and the past H time steps, and then obtain the GRU model output corresponding to the tracking target through transformer encoding.

[0132] S333. The GRU model output corresponding to the tracking target is decoded by three layers of MLP to obtain the pose information of the predicted bounding box corresponding to the tracking target. Then, the predicted bounding box corresponding to the tracking target is obtained by combining the shape information of the bounding box corresponding to the most recently matched tracking target.

[0133] Specifically, the transformer encoder in the motion state decoding module is used to decode the hidden state After processing, the temporal space weight corresponding to the current tracking target is obtained after processing by the softmax classifier And the temporal space weight corresponding to the tracking target in the past H time steps . use The hidden state of the GRU model corresponding to the tracking target at the current moment and the past H time steps is weighted and then encoded through the transformer to obtain the GRU model output corresponding to the tracking target. . The three-layer MLP decoding obtains the pose information of the predicted bounding box corresponding to the tracking target (i.e. the center coordinates and heading angle of the predicted bounding box), and combines the shape information of the bounding box corresponding to the most recently matched tracking target (i.e. the width, height, and length of the bounding box) to obtain the predicted bounding box corresponding to the tracking target. .

[0134] S4. Count the number of detection bounding boxes, and based on the number of detection bounding boxes, the confidence corresponding to the detection bounding boxes, and the stored historical matching of the tracking targets, process the detection bounding boxes and the predicted bounding boxes corresponding to the tracking targets, obtain and output matching pairs and / or unmatched tracking indexes and unmatched detection indexes, and update the historical tracking status information.

[0135] like Figure 12 As shown, in some embodiments, step S4 specifically includes:

[0136] S41: Count the number of detection bounding boxes, determine the status of each detection bounding box based on the number of detection bounding boxes and the confidence level of the detection bounding boxes, delete the detection bounding boxes in the ignored state, and then execute S42. The status of the detection bounding box is high score, low score, or ignored.

[0137] In some embodiments, the specific manner of determining the status of each detection bounding box is:

[0138] When the number of detection bounding boxes is greater than the preset number threshold, the state of the detection bounding box whose corresponding confidence is greater than or equal to the first preset confidence threshold is determined to be a high-scoring state, the state of the detection bounding box whose corresponding confidence is greater than or equal to the second preset confidence threshold and less than the first preset confidence threshold is determined to be a low-scoring state, and the state of the detection bounding box whose corresponding confidence is less than the second preset confidence threshold is determined to be an ignored state.

[0139] When the number of detection bounding boxes is less than or equal to a preset number threshold, the state of the detection bounding box whose corresponding confidence is greater than or equal to the first preset confidence threshold is determined to be a high-scoring state, and the state of the detection bounding box whose corresponding confidence is less than the first preset confidence threshold is determined to be a low-scoring state.

[0140] As an example, the preset number threshold is 10, the first preset confidence threshold is 0.6, and the second preset confidence threshold is 0.1.

[0141] S42: Determine the status of each tracking target based on the stored tracking target history matching, delete the predicted bounding box corresponding to the tracking target in the destroyed state, and then execute S43. The status of the tracking target can be activated, lost, or destroyed.

[0142] In some embodiments, the specific manner of determining the status of each tracking target is:

[0143] If the number of consecutive successful matching frames in the target's most recent Fra frame matching (determined by stored historical matching results) is greater than or equal to a first preset frame threshold, the target is considered to be in an activated state. If none of the target's most recent Fra frames have been successfully matched, the target is considered to be in a destroyed state. If the number of unsuccessful matching frames in the target's most recent Fra frame matching is less than Fra and the number of consecutive successful matching frames is less than a first preset frame threshold, the target is considered to be in a lost state. Fra represents the second preset frame threshold, which is greater than the first preset frame threshold. For example, the first preset frame threshold is 3 frames, and the second preset frame threshold Fra = 30.

[0144] The steps described in S41 and S42 are executed by the trajectory lifecycle management unit.

[0145] S43. Using the predicted bounding box corresponding to the active tracking target and the detected bounding box in the high-score state, perform a matching cost calculation of the 3D-CIoU and Mahalanobis distance, obtain the corresponding matching cost matrix of the 3D-CIoU and Mahalanobis distance, and then execute S44.

[0146] like Figure 13 As shown in (a), in some embodiments, S43 is performed by a matching cost calculation unit that combines 3D-CIoU and Mahalanobis distance, and the specific steps include:

[0147] S431. Calculate the Mahalanobis distance between a predicted bounding box corresponding to an active tracking target and a detected bounding box in a high-score state.

[0148] For example, the predicted bounding box corresponding to the active tracking target is recorded as: , the high-resolution detection bounding box is recorded as: .in, Respectively represent the coordinates of the center point of the detection bounding box on the X, Y, and Z axes of the laser radar coordinate system, Respectively represent the width, height, and length of the detection bounding box, Indicates the heading angle of the detection bounding box around the Z axis; Respectively represent the coordinates of the center point of the predicted bounding box on the X, Y, and Z axes of the lidar coordinate system, Respectively represent the width, height, and length of the predicted bounding box, Indicates the heading angle of the predicted bounding box around the Z axis, express and The Mahalanobis distance between them.

[0149] S432: Project the predicted bounding box corresponding to the active tracking target and the detected bounding box in the high-resolution state to the BEV perspective, and calculate the area and height of the overlapping area of ​​the projections.

[0150] like Figure 14 As shown in (a) in the figure, the traditional 2D bounding box IoU calculation does not consider the rotation of the box; therefore, the long and short sides of the box are parallel to the coordinate axis, and the common area of ​​the intersection of the two boxes must be rectangular. 3D detection takes the heading angle into account, such as Figure 14 As shown in (b) of the figure, the overlapping portion of the 3D object's bounding box from the BEV perspective is a convex polygon with a maximum of 8 vertices. Therefore, the key to calculating 3D-IoU is to solve the overlapping area and height overlap range of the convex polygons. By projecting the 3D object into the BEV perspective to convert it into a 2D object, the intersection over union (IoU) of the 3D object's bounding box is calculated.

[0151] For example: and The projection under the BEV view is and .calculate and The area of ​​the projected overlapped area is obtained to obtain the intersection from the BEV perspective. .in, They represent the coordinate values ​​of the center point of the detection bounding box projected on the X and Y axes under the BEV view, respectively. Respectively represent the width and height of the detection bounding box projected under the BEV view, Indicates the heading angle of the detection bounding box projected under the BEV view; Respectively represent the coordinate values ​​of the center point of the predicted bounding box projected on the X and Y axes under the BEV view, Respectively represent the width and height of the predicted bounding box projection under the BEV view, Indicates the heading angle of the predicted bounding box projected under the BEV view.

[0152] In the height direction, and Height of the projection overlap area Calculated by formula (6).

[0153] (6)

[0154] In formula (6), min( ) represents the minimum value function, and max( ) represents the maximum value function.

[0155] S433: Calculate the intersection-over-union ratio of the corresponding 3D object bounding box based on the area and height of the projection overlap region.

[0156] Specifically, the intersection-over-union ratio of the corresponding 3D target bounding box is calculated using formula (7): .

[0157] (7)

[0158] S434. Calculate a matching cost of the combination of 3D CIoU and Mahalanobis distance based on the Mahalanobis distance, the intersection over union ratio, a preset scaling factor, an IoU adjustment factor, an angle difference factor, and a diagonal distance of a minimum enclosing box area. The minimum enclosing box area is the smallest area that simultaneously contains the predicted bounding box corresponding to the active tracking target and the detected bounding box of the high-scoring state.

[0159] Specifically, formula (8) is used to calculate the matching cost of 3D-CIoU and Mahalanobis distance. .

[0160] (8)

[0161] In formula (8), λ Indicates the preset zoom factor, c Represents the diagonal distance of the minimum enclosing box area, α represents the IoU adjustment factor, ν Indicates the angle difference factor. Select and The maximum and minimum coordinate values ​​on the X, Y, and Z axes are used to construct the minimum enclosing frame area that contains both the prediction bounding box corresponding to the active tracking target and the detection bounding box of the high-resolution state. The diagonal distance of the minimum enclosing frame area is c . α, ν It is calculated by formula (9).

[0162] (9)

[0163] S435. Traverse the possible combinations of the predicted bounding box corresponding to the active tracking target and the detection bounding boxes of all high-scoring states, and repeat S431 to S434 to obtain the corresponding matching cost matrix of 3D-CIoU and Mahalanobis distance.

[0164] S44. Use the Kuhn-Munkres optimal allocation algorithm on the matching cost matrix to achieve bipartite matching between the predicted bounding box corresponding to the active tracking target and the detected bounding box in the high-score state, and then execute S45.

[0165] Specifically, the Kuhn-Munkres allocation module is used to apply the Kuhn-Munkres optimal allocation algorithm to the matching cost matrix obtained by S43 to achieve bipartite matching between the predicted bounding box corresponding to the active tracking target and the detected bounding box in the high-score state.

[0166] S45: Determine whether all tracking targets are matched successfully. If so, execute S48; otherwise, execute S46.

[0167] S46. Use the predicted bounding boxes corresponding to the remaining tracking targets (i.e., the tracking targets in the lost state and the unmatched tracking targets in the activated state) and the remaining detection bounding boxes (i.e., the detection bounding boxes in the low-scoring state and the detection bounding boxes in the unmatched high-scoring state) to calculate the matching cost of the 3D-CIoU and Mahalanobis distance, obtain the corresponding matching cost matrix of the 3D-CIoU and Mahalanobis distance, and then execute S47.

[0168] like Figure 13 As shown in (b), in some embodiments, S46 is performed by a matching cost calculation unit that combines 3D-CIoU with Mahalanobis distance.

[0169] The specific steps include:

[0170] S461: Calculate the Mahalanobis distance between the predicted bounding box corresponding to the remaining tracking target and the remaining detection bounding box. The specific method is similar to S431 and will not be repeated here.

[0171] S462: Project the predicted bounding box corresponding to the remaining tracking target and the remaining detection bounding box to the BEV perspective, and calculate the area and height of the overlapping area of ​​the projections. The specific method is similar to S432 and will not be repeated here.

[0172] S463: Calculate the intersection-over-union ratio of the corresponding 3D object bounding box based on the area and height of the projected overlapped region. The specific method is similar to S433 and will not be repeated here.

[0173] S464: Calculate a matching cost based on the Mahalanobis distance, the intersection over union (IoU), a preset scaling factor, an IoU adjustment factor, an angle difference factor, and the diagonal distance of the minimum enclosing box area. The minimum enclosing box area is the smallest area that contains both the predicted bounding box corresponding to the remaining tracked target and the remaining detection bounding box. The specific method is similar to S434 and is not further described here.

[0174] S465: Traverse all possible combinations of the predicted bounding box corresponding to the remaining tracked target and all remaining detection bounding boxes, and repeat S461 to S464 to obtain the corresponding matching cost matrix of the 3D-CIoU and Mahalanobis distance. The specific method is similar to S435 and will not be repeated here.

[0175] S47. Use the Kuhn-Munkres optimal allocation algorithm on the matching cost matrix to achieve bipartite matching between the predicted bounding boxes corresponding to the remaining tracking targets and the remaining detection bounding boxes, and then execute S48.

[0176] Specifically, the Kuhn-Munkres allocation module is used to apply the Kuhn-Munkres optimal allocation algorithm to the matching cost matrix obtained by S46 to achieve bipartite matching between the predicted bounding boxes corresponding to the remaining tracking targets and the remaining detection bounding boxes.

[0177] S48: Output matching pairs and / or unmatched tracking indices and unmatched detection indices, update historical tracking status information (add to the current frame), and then terminate tracking. If all tracking targets are successfully matched, matching pairs are obtained and output. If some tracking targets are successfully matched, matching pairs and unmatched tracking indices and unmatched detection indices are obtained and output. If all tracking targets are unmatched, unmatched tracking indices and unmatched detection indices are obtained and output. The judgment step in step S4 is also performed by the trajectory lifecycle management unit.

[0178] During binary matching on a frame, the target matching results for that frame are recorded and stored, forming a historical target matching history. There are two types of target matching scenarios: a successful match, meaning the predicted bounding box corresponding to the target successfully matches a detection bounding box; and an unsuccessful match, meaning the predicted bounding box corresponding to the target does not match any detection bounding boxes.

[0179] In step S4, two levels of matching are performed. The first level of matching, defined by steps S43 and S44, is performed between the predicted bounding boxes corresponding to all active tracked targets and the detected bounding boxes of all high-scoring targets. The second level of matching, defined by steps S46 and S47, is performed between the predicted bounding boxes corresponding to the remaining tracked targets and the remaining detected bounding boxes. This two-level matching approach improves matching efficiency, enabling faster and more reliable tracking of multiple surrounding targets.

[0180] The embodiment of the present invention innovatively introduces a matching cost calculation strategy that combines 3D-CIoU and Mahalanobis distance, combined with the Kuhn-Munkres optimal allocation algorithm and trajectory lifecycle management, to achieve simultaneous tracking of multiple moving objects, realize robust association of multi-target motion trajectories in complex occlusion scenes, and improve the stability of cross-frame target tracking and data association efficiency.

[0181] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A 3D multi-target tracking method based on motion state prediction, characterized in that: Including steps: S1. Obtain the current point cloud information and historical tracking status information of the vehicle during driving; S2. Process the current point cloud information to obtain the detection bounding box and its corresponding confidence level; S3. Process the historical tracking status information to obtain the predicted bounding box corresponding to the tracking target; S4. Count the number of detection bounding boxes, and based on the number of detection bounding boxes, the confidence scores corresponding to the detection bounding boxes, and the stored historical matching of the tracking targets, process the detection bounding boxes and the predicted bounding boxes corresponding to the tracking targets, obtain and output matching pairs and / or unmatched tracking indexes and unmatched detection indexes, and update the historical tracking status information. The step S2 specifically includes: S21, extracting point cloud feature information from the current point cloud information; S22, processing the point cloud feature information to obtain a primary bounding box; S23. Processing the primary bounding box to obtain point cloud attention weighted feature information; the specific method includes: S231. Mapping the primary bounding box to a voxel space, dividing the voxel space into a plurality of voxel grids, locating the center point of each voxel grid, and obtaining a plurality of voxel grid center points; S232. Feature encoding all feature points in a plurality of spaces with different radii from each voxel grid center point, and generating a multi-scale feature for each feature point; S233. Performing point cloud density self-attention calculation on the multi-scale feature to obtain point cloud attention weighted feature information; S24, decoding the point cloud attention weighted feature information to obtain a bounding box correction value, superimposing the bounding box correction value on the primary bounding box to obtain the detection bounding box; splicing the bounding box correction value with the point cloud attention weighted feature information and then decoding it to obtain a confidence level corresponding to the detection bounding box; The step S3 specifically includes: S31, performing spatial long-term time series modeling based on transformer coding on the historical tracking state information to obtain the historical tracking state coding corresponding to the tracking target; S32, using the historical tracking state code corresponding to the tracking target and the hidden state of the GRU model corresponding to the tracking target at the previous moment to update, to obtain the hidden state of the GRU model corresponding to the tracking target at the current moment; wherein, the hidden state of the GRU model corresponding to the tracking target at the initial moment is an all-zero vector; S33. Process the hidden state of the GRU model corresponding to the tracking target at the current moment and the past H time steps to obtain the predicted bounding box corresponding to the tracking target; the specific method includes: S331. Perform softmax classification on the hidden state of the GRU model corresponding to the tracking target at the current moment and the past H time steps to obtain the temporal space weight corresponding to the tracking target at the current moment and the past H time steps; S332. Use the temporal space weight to perform a weighted sum on the hidden state of the GRU model corresponding to the tracking target at the current moment and the past H time steps, and then obtain the GRU model output corresponding to the tracking target through transformer encoding; S333. Decode the GRU model output corresponding to the tracking target through three layers of MLP to obtain the posture information of the predicted bounding box corresponding to the tracking target, and then combine it with the stored bounding box shape information corresponding to the tracking target that was successfully matched last time to obtain the predicted bounding box corresponding to the tracking target.

2. The 3D multi-target tracking method based on motion state prediction according to claim 1, characterized in that: In S22, the point cloud feature information is processed to obtain a primary bounding box in the following manner: S221. Generate multiple anchor frames of different scales and aspect ratios at each position of the feature map composed of the point cloud feature information; S222: Predict the object confidence probability and anchor frame correction amount for each anchor frame to obtain the anchor frame correction amount; S223: Superimpose the anchor frame correction amount on the anchor frame to obtain a primary bounding frame.

3. The 3D multi-target tracking method based on motion state prediction according to claim 1 or 2, characterized in that: The step S4 specifically includes: S41, counting the number of detection bounding boxes, determining the status of each detection bounding box based on the number of detection bounding boxes and the confidence level corresponding to the detection bounding boxes, deleting the detection bounding boxes in the ignored state, and then executing S42; wherein the status of the detection bounding box is a high score state, a low score state, or an ignored state; S42. Determine the status of each tracking target based on the stored tracking target history matching, delete the predicted bounding box corresponding to the tracking target in the destroyed state, and then execute S43; wherein the status of the tracking target is an activated state, a lost state, or a destroyed state; S43, using the predicted bounding box corresponding to the active tracking target and the detected bounding box in the high-scoring state, perform a matching cost calculation of the 3D-CIoU and Mahalanobis distance, obtain the corresponding matching cost matrix of the 3D-CIoU and Mahalanobis distance, and then execute S44; S44, using the Kuhn-Munkres optimal allocation algorithm on the matching cost matrix to achieve bipartite matching between the predicted bounding box corresponding to the active tracking target and the detected bounding box in the high-scoring state, and then executing S45; S45, determine whether all tracking targets are matched successfully, if yes, execute S48, otherwise execute S46; S46, using the predicted bounding boxes corresponding to the remaining tracking targets and the remaining detection bounding boxes, perform a matching cost calculation of the 3D-CIoU and Mahalanobis distance combination to obtain a corresponding matching cost matrix of the 3D-CIoU and Mahalanobis distance combination, and then execute S47; S47, using the Kuhn-Munkres optimal allocation algorithm on the matching cost matrix to achieve bipartite matching between the predicted bounding boxes corresponding to the remaining tracking targets and the remaining detection bounding boxes, and then executing S48; S48, outputting matching pairs and / or unmatched tracking indexes and unmatched detection indexes, updating historical tracking state information, and then ending tracking; When a frame is binary matched, the tracking target matching situation of the frame is recorded and stored to form the tracking target matching history.

4. The 3D multi-target tracking method based on motion state prediction according to claim 3, characterized in that: In S41, the specific method of determining the status of each detection bounding box is: When the number of detection bounding boxes is greater than a preset number threshold, the corresponding detection bounding boxes whose confidence is greater than or equal to the first preset confidence threshold are determined to be in a high score state, the corresponding detection bounding boxes whose confidence is greater than or equal to the second preset confidence threshold and less than the first preset confidence threshold are determined to be in a low score state, and the corresponding detection bounding boxes whose confidence is less than the second preset confidence threshold are determined to be in an ignored state; When the number of detection bounding boxes is less than or equal to a preset number threshold, the corresponding detection bounding boxes whose confidence is greater than or equal to a first preset confidence threshold are determined to be in a high score state, and the corresponding detection bounding boxes whose confidence is less than the first preset confidence threshold are determined to be in a low score state; In S42, the specific method of determining the status of each tracking target is: If the number of consecutive frames successfully matched in the most recent Fra frame matching of the tracking target is greater than or equal to a first preset frame number threshold, then the tracking target is determined to be in an active state; wherein Fra represents a second preset frame number threshold, and the second preset frame number threshold is greater than the first preset frame number threshold; If the latest Fra frame of the tracking target fails to match successfully, the tracking target is judged to be in the destroyed state; If the number of unsuccessful matching frames in the latest Fra frame matching of the tracking target is less than Fra and the number of consecutive successful matching frames is less than a first preset frame number threshold, the state of the tracking target is determined to be a lost state.

5. The 3D multi-target tracking method based on motion state prediction according to claim 3, characterized in that: In the above S43, the specific method of obtaining the corresponding 3D-CIoU and Mahalanobis distance combined matching cost matrix is: S431, calculate the predicted bounding box corresponding to an active tracking target With a high-resolution detection bounding box Mahalanobis distance ; S432: Project the predicted bounding box corresponding to the active tracking target and the detected bounding box in the high-resolution state to the BEV perspective, and calculate the area and height of the overlapping area of ​​the projections; S433, based on the area and height of the projection overlap area, calculate the intersection-over-union ratio of the corresponding 3D object bounding box Specifically: ; in, Respectively The width, height, and length of Respectively The width, height, and length of Represents calculation and The intersection from the BEV perspective is obtained by projecting the area of ​​the overlapping area. express and The height of the projection overlap area, express Projection under BEV view, express Projection under BEV view; S434: Calculate the matching cost of 3D-CIoU and Mahalanobis distance based on the Mahalanobis distance, the intersection over union ratio, the preset scaling factor, the IoU adjustment factor, the angle difference factor, and the diagonal distance of the minimum enclosing frame area. Specifically: ; in, λ Indicates the preset zoom factor, c Represents the diagonal distance of the minimum enclosing box area, which is the minimum area that contains the predicted bounding box corresponding to the active tracking target and the detection bounding box of the high-scoring state. α represents the IoU adjustment factor, ν represents the angle difference factor, ; S435, traverse the possible combinations of the predicted bounding box corresponding to the active tracking target and the detection bounding boxes of all high-scoring states, and repeat S431 to S434 to obtain the corresponding matching cost matrix of 3D-CIoU and Mahalanobis distance; In the above S46, the specific method of obtaining the corresponding 3D-CIoU and Mahalanobis distance combined matching cost matrix is: S461, calculating the Mahalanobis distance between a predicted bounding box corresponding to a remaining tracking target and a remaining detection bounding box; S462: Project the predicted bounding box corresponding to the remaining tracking target and the remaining detection bounding box to the BEV perspective, and calculate the area and height of the overlapping area of ​​the projections; S463: Calculate the intersection-over-union ratio of the corresponding 3D object bounding box based on the area and height of the projection overlap area; S464: Calculate a matching cost value of a combination of 3D-CIoU and Mahalanobis distance based on the Mahalanobis distance, the intersection over union ratio, a preset scaling factor, an IoU adjustment factor, an angle difference factor, and a diagonal distance of a minimum enclosing frame area; wherein the minimum enclosing frame area is the minimum area that simultaneously includes the predicted bounding box corresponding to the remaining tracking target and the remaining detection bounding box; S465: Traverse the possible combinations of the predicted bounding box corresponding to the remaining tracking target and all remaining detection bounding boxes, and repeat S461 to S464 to obtain the corresponding matching cost matrix of 3D-CIoU and Mahalanobis distance.

6. A 3D multi-target tracking system based on motion state prediction, comprising a memory and a controller, wherein the memory stores a computer-readable program, characterized in that: When the computer-readable program is called by the controller, it can execute the steps of the 3D multi-target tracking method based on motion state prediction as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Online multi-target tracking method and device and storage medium

    CN115170602A

  • Pedestrian trajectory prediction method, system and device based on multiple interactions and medium

    CN116654022A

  • 3D multi-target tracking method, system and device based on BEV and medium

    CN117392173A

  • 3D point cloud target detection method based on cascade series attention mechanism

    CN118172537A