Dynamic object tracking prediction method and system based on multi-modal motion model

By adopting a multimodal motion model and region of interest feedback mechanism in dynamic object point cloud tracking, the prediction deviation problem caused by a single motion model is solved, the tracking accuracy and robustness are significantly improved, and the two-way optimization of detection and tracking is realized.

CN119991745AActive Publication Date: 2025-05-13SHANDONG UNIV
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202510479524.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-05-13
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

Existing dynamic object point cloud tracking methods usually use a single motion model to predict different categories of objects, resulting in prediction deviations and affecting tracking accuracy. They fail to feed back the tracking prediction results to the detection system to form closed-loop optimization.

Method used

The dynamic object tracking and prediction method based on the multimodal motion model is adopted. By obtaining the point cloud and semantic information of the dynamic object, candidate object information is generated, and appropriate motion models are selected according to the object category, and state estimation and prediction are used using extended Kalman filtering, and the detection and tracking process is optimized in combination with the feedback mechanism of the region of interest.

Benefits of technology

It significantly improves the accuracy and robustness of dynamic object point cloud tracking, adapts to the changes in object motion patterns through multi-model switching mechanism, enhances the ability to capture complex motion trajectories, and realizes two-way collaborative optimization of detection and tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991745A_ABST
    Figure CN119991745A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of automatic driving environment perception. According to the dynamic object tracking prediction method and system based on the multi-modal motion model, dynamic object point cloud and semantic information data are obtained, candidate object information is generated, a new tracking object motion model is selected, the object state is estimated, candidate objects are matched and managed, and the object state is updated. Updating an object track and object cutting motion model, estimating a future multi-frame object state, and generating a region of interest; according to the invention, a dynamic object tracking prediction method based on a multi-modal motion model is adopted, and by constructing a multi-modal motion model system and a region-of-interest feedback mechanism, the precision of dynamic object point cloud tracking and the system robustness are significantly improved; different from extensive prediction of a traditional single motion model, the method flexibly selects the matched motion model for state estimation according to the motion characteristics of different types of objects, and effectively overcomes the problem of prediction deviation of a unified model to different dynamic characteristic objects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of autonomous driving environment perception technology, and in particular to a dynamic object tracking prediction method and system based on a multimodal motion model. Background Art

[0002] The statements in this section merely provide background art related to the present invention and do not necessarily constitute prior art.

[0003] The environmental perception system is one of the core components of autonomous driving technology. It is responsible for real-time perception of environmental information around the vehicle to support subsequent decision-making and control. Among them, LiDAR, as a key environmental perception sensor, can provide high-precision three-dimensional point cloud data for detecting and tracking dynamic objects. In the autonomous driving system, accurate detection and robust tracking of dynamic objects in the LiDAR point cloud are key links to ensure that the path planning and decision-making system can correctly predict the behavior of traffic participants, reasonably avoid obstacles and optimize driving strategies. A robust dynamic object point cloud tracking method not only needs to accurately identify and track dynamic targets such as pedestrians and vehicles in real time, but also should be combined with a motion prediction model to infer the possible motion trajectory and appearance area of ​​dynamic objects in the next few frames, thereby providing prior information for the dynamic object detection system and further improving the robustness and accuracy of detection.

[0004] At present, the mainstream dynamic object point cloud tracking methods are mainly divided into two categories: tracking-by-detection (TBD) and joint detection and tracking (JDT). Among them, the detection-based tracking method first detects the target on the point cloud data, then matches the same target in consecutive frames through the data association algorithm, and constructs the complete trajectory with the help of motion prediction models (such as Kalman filtering); while the joint detection and tracking method completes target detection and tracking at the same time in the same framework, usually relying on deep learning for global optimization and association. Since the detection-based tracking method can decouple the detection and tracking tasks, the detector and tracker can be optimized separately, and the data association process is more transparent and controllable, this method is still the main method in most dynamic object point cloud tracking tasks to ensure higher accuracy and robustness.

[0005] However, existing detection-based tracking methods usually use a single motion model to predict all targets. In real scenarios, objects of different categories have their own unique motion patterns. Using a unified motion model may lead to prediction bias, which in turn affects tracking accuracy. In addition, existing methods fail to feed back the tracking prediction results to the detection system to form a closed-loop optimization, resulting in the detection and tracking processes being relatively independent, and failing to fully utilize timing information to improve overall performance. Summary of the invention

[0006] In order to address the deficiencies in the prior art, the present invention provides a dynamic object tracking prediction method and system based on a multimodal motion model, which improves the accuracy and robustness of dynamic object point cloud tracking.

[0007] In order to achieve the above object, the present invention adopts the following technical solution: In a first aspect, the present invention provides a dynamic object tracking and prediction method based on a multimodal motion model.

[0008] A dynamic object tracking prediction method based on a multimodal motion model includes the following processes: Obtain dynamic object point cloud and semantic information of each frame of point cloud data; Generate candidate object information based on dynamic object point cloud and semantic information; Determine the category of the candidate object according to the semantic information in the candidate object information of the current frame, and select the motion model according to the category; According to the selected motion model, the extended Kalman filter is used to estimate the state of objects of different categories and predict the state vector and covariance matrix of the next frame of point cloud data; Matching the predicted state vector and covariance matrix with the candidate object received in the next frame to determine the matching relationship between the predicted state vector and the candidate object in the next frame of point cloud data; Update the state vector and covariance matrix of the dynamic object; Append the updated state vector to the trajectory sequence; Estimate the state vector of multiple future frames based on the trajectory sequence; Generate a region of interest based on the state vector estimation results of the dynamic object in the future frame.

[0009] In a second aspect, the present invention provides a dynamic object tracking and prediction system based on a multimodal motion model.

[0010] A dynamic object tracking prediction system based on a multimodal motion model, comprising: The data acquisition unit is configured to: acquire the dynamic object point cloud and semantic information of each frame of point cloud data; The candidate object determination unit is configured to: generate candidate object information according to the dynamic object point cloud and semantic information; The motion model selection unit is configured to: determine the category of the candidate object according to the semantic information in the candidate object information of the current frame, and select the motion model according to the category; The state estimation unit is configured to: perform state estimation on objects of different categories using an extended Kalman filter according to a selected motion model, and predict a state vector and a covariance matrix of a next frame of point cloud data; A matching relationship determination unit is configured to: match the predicted state vector and covariance matrix with the candidate object received in the next frame, and determine the matching relationship between the predicted state vector and the candidate object in the next frame of point cloud data; A state updating unit is configured to: update a state vector and a covariance matrix of a dynamic object; The trajectory sequence updating unit is configured to: append the updated state vector to the trajectory sequence; A future state vector estimation unit is configured to: estimate the state vector of multiple future frames according to the trajectory sequence; The region of interest generation unit is configured to generate a region of interest according to a state vector estimation result of a future frame of the dynamic object.

[0011] Compared with the prior art, the present invention has the following beneficial effects: 1. The present invention adopts a dynamic object tracking and prediction method based on a multimodal motion model. By constructing a multimodal motion model system and a region of interest feedback mechanism, the accuracy of dynamic object point cloud tracking and the system robustness are significantly improved. Different from the extensive prediction of a traditional single motion model, the present invention flexibly selects a matching motion model for state estimation based on the motion characteristics of objects of different categories, effectively overcoming the prediction deviation problem of a unified model for objects with different dynamic characteristics.

[0012] 2. The present invention dynamically adjusts the region of interest based on the tracking prediction results, and feeds back the generated region of interest information to the dynamic object detection system, so that the detection system can reduce missed detection and false detection phenomena, and realizes two-way collaborative optimization of detection and tracking.

[0013] 3. The present invention adapts to the sudden change of the object's motion mode in real time through a multi-model switching mechanism, and combines trajectory update with future multi-frame prediction algorithm to enhance the ability to capture the continuity of complex motion trajectories, thereby showing stronger environmental adaptability in real scenes such as mixed traffic and dense occlusion.

[0014] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0016] Figure 1 A schematic diagram of a flow chart of a dynamic object tracking prediction method based on a multimodal motion model provided in Embodiment 1 of the present invention; Figure 2 A schematic diagram of a data acquisition method provided in Example 1 of the present invention; Figure 3 A schematic diagram of selecting a new tracking object motion model provided in Embodiment 1 of the present invention; Figure 4 A schematic diagram of selecting a vehicle-like object motion model provided in Embodiment 1 of the present invention; Figure 5 A schematic diagram of candidate object matching and management provided in Example 1 of the present invention; Figure 6 A schematic diagram of a fan-shaped and rectangular region of interest generated according to Embodiment 1 of the present invention; Figure 7 A schematic diagram of a dynamic object tracking and prediction system based on a multimodal motion model provided in Example 2 of the present invention. DETAILED DESCRIPTION

[0017] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0018] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.

[0019] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.

[0020] Embodiment 1: This implementation proposes a dynamic object tracking prediction method based on a multimodal motion model. Figure 1 As shown, the following process is included: S1: Obtain dynamic object point cloud and semantic information data.

[0021] More specifically, the data output by the dynamic object point cloud detection system is received at a fixed frequency, such as Figure 2 As shown, each frame of data includes a dynamic object point cloud and its semantic information. The data format of each point in the point cloud is as follows: (1); in, Indicates the first points; Indicates the coordinates of the point in the world coordinate system; Indicates that the point belongs to the frame A point in a dynamic object; Indicates that the point belongs to the semantic information of a dynamic object, including pedestrian class, cyclist class, vehicle class, and other object classes.

[0022] S2: Candidate object information generation.

[0023] More specifically, after receiving the point cloud and semantic information of each frame of the dynamic object, the method processes the point cloud data of the same dynamic object and generates candidate object information; wherein the point cloud data belonging to the same dynamic object refers to the point cloud data with the same semantic label in the frame point cloud. The point set of a dynamic object is defined as follows: (2); in, Indicates A dynamic object; Indicates the first Points, They represent the first, second and third points in the dynamic object respectively.

[0024] Calculate the center of mass and axis-aligned bounding box of all dynamic objects in the current frame. The calculation formula is as follows: (3); in, represents the center of mass of the dynamic object, ; Represents the minimum corner point of the axis-aligned bounding box; Represents the maximum corner point of the axis-aligned bounding box; , , Represents the length, width and height of the axis-aligned bounding box. Represents the total number of points in a dynamic object, For the The coordinates of a point.

[0025] Finally, the candidate object information is obtained: (4); in, Represents the semantic information of dynamic objects.

[0026] S3: Selection of motion model for new tracked objects.

[0027] More specifically, after obtaining the candidate object information, the method identifies the object's category based on its semantic information and selects the corresponding motion model, such as Figure 3 As shown. The pedestrian class object uses the random walk model; the cyclist class object uses the single-track bicycle model; for the vehicle class object, the method first calculates its point cloud density: (5); in, Indicates the total number of points in the dynamic object.

[0028] When the point cloud density Less than or equal to the set threshold When the point cloud density is Greater than the set threshold When , the uniform speed steering model is used; other objects use the constant speed model, such as Figure 4 shown.

[0029] S4: Object state estimation.

[0030] More specifically, after completing the motion model selection of the tracking object, the present invention uses an extended Kalman filter to perform state estimation on objects of different categories. The state equation and observation equation of the extended Kalman filter are as follows: (6); in, Indicates The state vector of the frame point cloud; Indicates The state vector of the frame point cloud; represents the state transition function; Indicates Process noise of frame point cloud; represents the covariance matrix of the process noise; Represents the measurement vector of the point cloud of the current frame; represents the measurement function; Indicates Measurement noise of frame point cloud; represents the covariance matrix of the measurement noise; the subscript Indicates the point cloud of the current frame, subscript Represents the point cloud of the previous frame.

[0031] For different types of dynamic objects, the present invention adopts corresponding state vectors, state transfer functions and process noise modeling: (1) For a pedestrian dynamic object using a random walk model, the state vector, state transfer function, and process noise are as follows: (7); at this time, Indicates The centroid coordinates of human objects in the frame point cloud; Indicates The speed of the frame point cloud in the corresponding direction; Indicates The centroid coordinates of human objects in the frame point cloud; Indicates The speed of the frame point cloud in the corresponding direction; Indicates the time interval between frames; Represents the position noise in three directions respectively; Represent the velocity uncertainty in three directions respectively.

[0032] (2) Using the cyclist-like dynamic object of the single-track bicycle model, the state vector, state transfer function and process noise are as follows: (8); at this time, Indicates The center of mass coordinates of the cyclist object in the frame point cloud; Indicates The XY axis speed of the cyclist object in the frame point cloud; Indicates The heading angle of the cyclist object in the frame point cloud; Indicates The front wheel turning angle of the cyclist object in the frame point cloud; Indicates The center of mass coordinates of the cyclist object in the frame point cloud; Indicates The XY axis speed of the cyclist object in the frame point cloud; Indicates The heading angle of the cyclist-like object in the frame point cloud; Indicates The front wheel turning angle of the cyclist object in the frame point cloud; Represents the wheelbase of the bicycle model, and its approximate calculation formula is as follows: ;in, Indicates the X-axis and Y-axis coordinates of the maximum corner point of the axis-aligned bounding box of the dynamic object. Represents the X-axis and Y-axis coordinates of the minimum corner point of the axis-aligned bounding box of the dynamic object; Indicates the time interval between frames; Represents the position noise in three directions respectively; represents the velocity uncertainty; represents the heading angle uncertainty; represents the uncertainty of the front wheel steering angle.

[0033] (3) For a vehicle-like dynamic object using a constant acceleration model, the state vector, state transfer function, and process noise are as follows: (9); at this time, Indicates The centroid coordinates of vehicle objects in the frame point cloud; Indicates The speed of the frame point cloud in the corresponding direction; Indicates The acceleration of the corresponding direction of the frame point cloud; Indicates The centroid coordinates of vehicle objects in the frame point cloud; Indicates The speed of the frame point cloud in the corresponding direction; Indicates The acceleration of the corresponding direction of the frame point cloud; Indicates the time interval between frames; Represents the position noise in three directions respectively; Represent the velocity uncertainty in three directions respectively; Represent the acceleration uncertainty in three directions respectively.

[0034] (4) For a vehicle-like dynamic object using a uniform steering model, the state vector, state transfer function, and process noise are as follows: (10); at this time, Indicates The centroid coordinates of vehicle objects in the frame point cloud; Indicates The XY axis speed of the vehicle object in the frame point cloud; Indicates The heading angle of the vehicle object in the frame point cloud; Indicates Angular velocity of vehicle objects in frame point cloud; Indicates The centroid coordinates of vehicle objects in the frame point cloud; Indicates The XY axis speed of the vehicle object in the frame point cloud; Indicates The heading angle of the vehicle object in the frame point cloud; Indicates Angular velocity of vehicle objects in frame point cloud; Indicates the time interval between frames; represents position noise; It represents the uncertainty of the velocity of the XY axis of the center of mass; represents the heading angle uncertainty; represents the angular velocity uncertainty.

[0035] (5) For other objects that use the constant velocity model and are dynamic objects, the state vector, state transfer function, and process noise are as follows: (11); at this time, Indicates The centroid coordinates of other object classes in the frame point cloud; Indicates The speed of the frame point cloud in the corresponding direction; Indicates The centroid coordinates of other object classes in the frame point cloud; Indicates The speed of the three directions corresponding to the frame point cloud; Indicates the time interval between frames; represents the variance of acceleration noise; express The identity matrix of .

[0036] Next, according to the motion pattern of the dynamic object, the state vector and covariance matrix of the next frame are predicted. The formula is as follows: (12); at this time, Indicates based on Frame data pair The predicted value of the frame state vector; express The state vector obtained by updating the extended Kalman filter state at the frame time; Indicates based on Frame data pair The predicted value of the frame covariance matrix; express Jacobian matrix of the frame state transfer function; express The covariance matrix obtained by updating the extended Kalman filter state at the frame time; Represents the covariance matrix of the process noise.

[0037] S5: Candidate object matching and management.

[0038] More specifically, after completing the prediction of the state vector and covariance matrix of the dynamic object in the next frame, it is necessary to match the predicted dynamic object with the candidate object received in the next frame.

[0039] For each predicted dynamic object state vector and each measurement candidate object, the innovation amount is calculated as follows: (13); at this time, Indicates A predicted state vector; Indicates measurement candidate objects; represents the measurement vector of the kth frame, ; Indicates the centroid position of the candidate object; represents the measurement function, ; Represents the predicted state vector The center of mass position in .

[0040] Calculate the innovation covariance matrix of the predicted state vector as follows: (14); at this time, Represents the Jacobian matrix of the k-th frame measurement function; represents the measurement noise; Indicates based on Frame data pair Predicted values ​​of the frame covariance matrix.

[0041] Calculate the Mahalanobis distance between each pair of predicted state vector and candidate object, the formula is as follows: (15); Next, construct the cost matrix , the formula is as follows: (16); At this time, the rows of the cost matrix correspond to the predicted state vectors, and the columns correspond to the candidate targets; It means taking a larger value, which represents a very large cost; Indicates the threshold value of the threshold check.

[0042] Use the Hungarian algorithm to solve the cost matrix Perform global optimal allocation to determine the matching relationship between the predicted state vector and the candidate target. After completing the matching of the candidate objects, further manage the unmatched candidate objects and unmatched predicted objects, such as Figure 5 shown.

[0043] For unmatched candidate objects, if the number of their point clouds Exceeding the set threshold , it is regarded as a newly appeared object and its candidate object information is used as the initial value; otherwise, it is determined as a false detection of the dynamic object point cloud detection system.

[0044] For unmatched predicted objects, The state is retained within the time interval of the frame, and the state vector of the future frame is calculated using the step S8. If no candidate target is matched within the frame time interval, it will be discarded.

[0045] S6: Object status update.

[0046] More specifically, after obtaining the matching relationship between the predicted state vector and the candidate target, the method updates the state vector of the object.

[0047] First, calculate the amount of innovation , Innovation Covariance and Kalman gain , the formula is as follows: (17); at this time, represents the measurement vector of the kth frame, ; Indicates the centroid position of the candidate object; represents the measurement function, ; Represents the predicted state vector The center of mass position in ; Represents the Jacobian matrix of the k-th frame measurement function; Indicates based on Frame data pair The predicted value of the frame covariance matrix; Represents the covariance matrix of the measurement noise.

[0048] Then, the Kalman gain is used to update the state vector and covariance matrix, as follows: (18) at this time, express The state vector obtained by updating the frame extension Kalman filter state; express The covariance matrix obtained by updating the frame extension Kalman filter state; Represents the identity matrix.

[0049] S7: Object trajectory update and object motion model switching.

[0050] More specifically, after obtaining the latest state of the object, the method updates the trajectory of the object and converts the latest state vector Append to trajectory sequence ; For vehicle objects, this method uses the latest object information to calculate its point cloud density. , and dynamically adjust the motion model accordingly. When the point cloud density Less than or equal to the set threshold When the point cloud density is Greater than the set threshold , switches to the uniform speed steering model.

[0051] S8: Future multi-frame object state estimation.

[0052] More specifically, after updating the object trajectory, the method estimates the object state for multiple frames in the future.

[0053] For pedestrian-like dynamic objects using the random walk model, vehicle-like dynamic objects using the constant acceleration model, and other object-like dynamic objects using the constant velocity model, in the future The frame state estimation formula is as follows: (19); at this time, is the predicted value of the state vector for the next n frames, is the predicted value of the covariance matrix for the next n frames; The Jacobian matrix representing the state transfer function; Represents the covariance matrix of the process noise.

[0054] For a dynamic object such as a cyclist using a single-track bicycle model, an iterative method is used to estimate the future Frame state vector, the formula is as follows: (20); at this time, Indicates The center of mass coordinates of the cyclist object in the frame point cloud; Indicates The XY axis speed of the cyclist object in the frame point cloud; Indicates The heading angle of the cyclist-like object in the frame point cloud; Indicates The front wheel turning angle of the cyclist object in the frame point cloud; Indicates The center of mass coordinates of the cyclist object in the frame point cloud; Indicates The XY axis speed of the cyclist object in the frame point cloud; Indicates The heading angle of the cyclist-like object in the frame point cloud; Indicates The front wheel turning angle of the cyclist object in the frame point cloud; Indicates the time interval between frames; Represents the wheelbase of the bicycle model.

[0055] For dynamic objects such as vehicles using a constant speed steering model, an iterative method is used to estimate the future Frame state vector, the formula is as follows: (twenty one).

[0056] at this time, Indicates The centroid coordinates of vehicle objects in the frame point cloud; Indicates The XY axis speed of the vehicle object in the frame point cloud; Indicates The heading angle of the vehicle object in the frame point cloud; Indicates Angular velocity of vehicle objects in the frame point cloud, Indicates Angular velocity of vehicle objects in the frame point cloud, Indicates The centroid coordinates of vehicle objects in the frame point cloud; Indicates The XY axis speed of the vehicle object in the frame point cloud; Indicates The heading angle of the vehicle object in the frame point cloud; Indicates the time interval between frames;.

[0057] S9: Region of interest generation.

[0058] More specifically, after completing the estimation of the state of the object in multiple future frames, the method generates a region of interest based on the state vector of the dynamic object in the future frame, and feeds it back to the dynamic object point cloud detection system.

[0059] First, the method uses the latest object information to calculate its point cloud density ; Secondly, for dynamic objects using different motion models, this method generates different regions of interest, including sectors and rectangles. The schematic diagrams of the generated sector and rectangular regions of interest are as follows: Figure 6 shown.

[0060] For pedestrians using the random walk model, cyclists using the monorail model, and vehicles using the constant speed steering model, this method generates a sector-shaped region of interest. The center coordinates of the sector are , the formulas for the outer diameter, inner diameter and central angle of the sector are as follows: (twenty two); (twenty three); (twenty four); at this time, , Represents the length and width of the axis-aligned bounding box; represents the set of integers; Indicates The frame state is updated to obtain the X-axis coordinate of the center of mass of the state vector; Indicates The frame state is updated to obtain the Y-axis coordinate of the center of mass of the state vector; Indicates The X-axis coordinate of the center of mass of the frame prediction state vector; Indicates The Y-axis coordinate of the center of mass of the frame prediction state vector; is the adjustment coefficient.

[0061] For dynamic objects such as vehicles using a constant acceleration model and other dynamic objects using a constant velocity model, this method generates a rectangular region of interest with the center coordinates of the rectangle being , the length of the rectangle He Kuan ; Among them, the adjustment coefficient , when the point cloud density Less than or equal to the set threshold When Take 2.0~3.0, when density Greater than the set threshold When Take 1.0~2.0.

[0062] Finally, the predicted region of interest information is sent to the dynamic object point cloud detection system.

[0063] Embodiment 2: like Figure 7 As shown, this implementation provides a dynamic object tracking prediction system based on a multimodal motion model, including: The data acquisition unit is configured to: acquire the dynamic object point cloud and semantic information of each frame of point cloud data; The candidate object determination unit is configured to: generate candidate object information according to the dynamic object point cloud and semantic information; The motion model selection unit is configured to: determine the category of the candidate object according to the semantic information in the candidate object information of the current frame, and select the motion model according to the category; The state estimation unit is configured to: perform state estimation on objects of different categories using an extended Kalman filter according to a selected motion model, and predict a state vector and a covariance matrix of a next frame of point cloud data; A matching relationship determination unit is configured to: match the predicted state vector and covariance matrix with the candidate object received in the next frame, and determine the matching relationship between the predicted state vector and the candidate object in the next frame of point cloud data; A state updating unit is configured to: update a state vector and a covariance matrix of a dynamic object; The trajectory sequence updating unit is configured to: append the updated state vector to the trajectory sequence; A future state vector estimation unit is configured to: estimate the state vector of multiple future frames according to the trajectory sequence; The region of interest generation unit is configured to generate a region of interest according to a state vector estimation result of a future frame of the dynamic object.

[0064] The specific working process of each of the above units is described in Example 1 and will not be repeated here.

[0065] It is understandable that the above-mentioned units can be separately or completely combined into one or several other units to constitute, or one (some) of the units can be further divided into multiple smaller units in function to constitute, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above-mentioned units are divided based on logical functions. In practical applications, the functions of one unit can also be implemented by multiple units, or the functions of multiple units can be implemented by one unit. In other embodiments of the present application, the system may also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented by the collaboration of multiple units.

[0066] According to another embodiment of the present application, the system described in this embodiment can be constructed, and the method of Example 1 of the present application can be implemented by running a computer program (including program code) capable of executing the steps involved in the corresponding method described in Example 1 on a general-purpose computing device such as a computer, which includes processing elements and storage elements such as a central processing unit (CPU), random access memory (RAM), and read-only memory (ROM). The computer program can be recorded on, for example, a computer-readable recording medium, and loaded into the above-mentioned computing device through the computer-readable recording medium and run therein.

[0067] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A dynamic object tracking prediction method based on a multimodal motion model, characterized in that: The process includes: Obtain dynamic object point cloud and semantic information of each frame of point cloud data; Generate candidate object information based on dynamic object point cloud and semantic information; Determine the category of the candidate object according to the semantic information in the candidate object information of the current frame, and select the motion model according to the category; According to the selected motion model, the extended Kalman filter is used to estimate the state of objects of different categories and predict the state vector and covariance matrix of the next frame of point cloud data; Matching the predicted state vector and covariance matrix with the candidate object received in the next frame to determine the matching relationship between the predicted state vector and the candidate object in the next frame of point cloud data; Update the state vector and covariance matrix of the dynamic object; Append the updated state vector to the trajectory sequence; Estimate the state vector of multiple future frames based on the trajectory sequence; Generate a region of interest based on the state vector estimation results of the dynamic object in the future frame.

2. The dynamic object tracking prediction method based on a multimodal motion model according to claim 1, characterized in that: The data format of each point in the dynamic object point cloud is as follows: ,in, Indicates the first points; Indicates the coordinates of the point in the world coordinate system; Indicates that the point belongs to the frame A point in a dynamic object; Indicates that the point belongs to the semantic information of a dynamic object, including pedestrian class, cyclist class, vehicle class, and other object classes.

3. The dynamic object tracking prediction method based on a multimodal motion model according to claim 1, characterized in that: Candidate object information, including: ,in, is the coordinate value of the center of mass of the dynamic object, , , Represents the length, width, and height of the axis-aligned bounding box. Represents the semantic information of dynamic objects, The number of points representing dynamic objects.

4. The dynamic object tracking prediction method based on a multimodal motion model according to claim 1, characterized in that: The pedestrian class objects use the random walk model, and the cyclist class objects use the single-track bicycle model; When calculating the point cloud density of vehicle objects, when the point cloud density is less than or equal to the set threshold, the constant acceleration model is used; when the point cloud density is greater than the set threshold, the uniform speed steering model is used; The other objects adopt the constant velocity model.

5. The dynamic object tracking prediction method based on multimodal motion model according to claim 1, characterized in that: According to the predicted state vector of the next frame of point cloud data, the innovation amount is calculated for each predicted state vector and each candidate object; According to the predicted covariance matrix of the next frame of point cloud data, the innovative covariance matrix of the predicted state vector is calculated; According to the innovation amount and the innovation covariance matrix, the Mahalanobis distance between the predicted state vector and the candidate object is calculated; The cost matrix is ​​constructed according to the Mahalanobis distance, and the Hungarian algorithm is used to perform global optimal allocation on the cost matrix to determine the matching relationship between the predicted state vector and the candidate object.

6. The dynamic object tracking prediction method based on a multimodal motion model as claimed in claim 5, characterized in that: For unmatched candidate objects, if the number of point clouds exceeds the set threshold, they are regarded as newly appeared objects and their candidate object information is used as the initial value; Otherwise, it is determined to be a false detection of the dynamic object point cloud detection system; For unmatched predicted objects, The state is retained within the frame time interval, and the state vector of the future frame is calculated. If the candidate target is still not matched within the frame time interval, the unmatched predicted object is discarded.

7. The dynamic object tracking prediction method based on a multimodal motion model according to claim 1, characterized in that: After appending the updated state vector to the trajectory sequence, for vehicle objects, the point cloud density is calculated using the latest dynamic object information. When the point cloud density is less than or equal to the set threshold, it is switched to the constant acceleration model; when the point cloud density is greater than the set threshold, it is switched to the uniform steering model.

8. The dynamic object tracking prediction method based on a multimodal motion model according to claim 1, characterized in that: Generate a region of interest based on the state vector estimation results of the dynamic object in the future frame, including: Calculate point cloud density using the latest dynamic object information; For pedestrian-like dynamic objects using a random walk model, cyclist-like dynamic objects using a monorail bicycle model, and vehicle-like dynamic objects using a uniform speed steering model, a sector-shaped region of interest is generated, and the parameters of the sector are determined according to the point cloud density; For dynamic objects such as vehicles using a constant acceleration model and dynamic objects such as other objects using a constant velocity model, a rectangular region of interest is generated, and parameters of the rectangle are determined according to the point cloud density.

9. The dynamic object tracking prediction method based on a multimodal motion model as claimed in claim 8, characterized in that: The coordinates of the center of the sector are , the outer diameter of the sector , inner diameter and the central angle for: ; ; ; The center coordinates of the rectangle are , the length of the rectangle ,Width ; in, , Represents the length and width of the axis-aligned bounding box; represents the set of integers; Indicates The frame state is updated to obtain the X-axis coordinate of the center of mass of the state vector; Indicates The frame state is updated to obtain the Y-axis coordinate of the center of mass of the state vector; Indicates The X-axis coordinate of the center of mass of the frame prediction state vector; Indicates The Y-axis coordinate of the center of mass of the frame prediction state vector; is the adjustment coefficient, when the point cloud density When it is less than or equal to the set threshold, the adjustment coefficient Take 2.0~3.0, when density Greater than threshold When Take 1.0~2.

0.

10. A dynamic object tracking prediction system based on a multimodal motion model, characterized in that: include: The data acquisition unit is configured to: acquire the dynamic object point cloud and semantic information of each frame of point cloud data; The candidate object determination unit is configured to: generate candidate object information according to the dynamic object point cloud and semantic information; The motion model selection unit is configured to: determine the category of the candidate object according to the semantic information in the candidate object information of the current frame, and select the motion model according to the category; The state estimation unit is configured to: perform state estimation on objects of different categories using an extended Kalman filter according to a selected motion model, and predict a state vector and a covariance matrix of a next frame of point cloud data; A matching relationship determination unit is configured to: match the predicted state vector and covariance matrix with the candidate object received in the next frame, and determine the matching relationship between the predicted state vector and the candidate object in the next frame of point cloud data; A state updating unit is configured to: update a state vector and a covariance matrix of a dynamic object; The trajectory sequence updating unit is configured to: append the updated state vector to the trajectory sequence; A future state vector estimation unit is configured to: estimate the state vector of multiple future frames according to the trajectory sequence; The region of interest generation unit is configured to generate a region of interest according to a state vector estimation result of a future frame of the dynamic object.

Citation Information

Patent Citations

  • Image detection method and device, electronic equipment and computer readable storage medium

    CN109948616A

  • Moving target tracking method based on cascade detector

    CN110222585A

  • Target tracking device and method for road monitoring video

    CN113269007A

  • Target detection and tracking method and system based on multi-dimensional point cloud features

    CN114419152A

  • High-stability multi-target tracking method based on optimal motion model trajectory prediction

    CN116612154A