Point cloud target tracking method and device, equipment and storage medium

By adopting a point cloud target tracking method based on lightning vision fusion in the field of intelligent transportation, the problem of low lightning vision fusion accuracy in the existing technology is solved, and more efficient point cloud clustering filtering and target tracking are achieved, reducing computing power demand.

CN119991731APending Publication Date: 2025-05-13PCI TECH GRP CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411760735.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-03
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing Rail Vision fusion technology has problems such as low accuracy and poor fusion effect in the field of intelligent transportation, especially when millimeter-wave radar and visual data are fusion.

Method used

A point cloud target tracking method based on lightning vision fusion is proposed. By acquiring video stream data and radar point cloud data, the visual target structured data is processed, coordinate system conversion, area filtering and point cloud clustering of radar point cloud data is realized, and target tracking is carried out based on the fused point cloud data.

Benefits of technology

The accuracy of point cloud clustering filtering is improved, effectively solving the problem of inaccurate generation of radar false targets and clustering targets, improving the accuracy and real-timeness of target tracking, and reducing computing power requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991731A_ABST
    Figure CN119991731A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent transportation, and discloses a point cloud target tracking method, device and equipment and a storage medium. The point cloud target tracking method comprises the following steps: acquiring video stream data and radar point cloud data; processing the video stream data to obtain visual target structured data; performing thunder-vision fusion on the visual target structured data and the radar point cloud data, and outputting first point cloud frame data subjected to thunder-vision fusion or second point cloud frame data not subjected to thunder-vision fusion; and based on the first point cloud frame data or the second point cloud frame data, tracking a target in a current frame view range. According to the method, the precision of point cloud clustering filtering is improved through Leiyu fusion, and meanwhile, the computing power requirement of an algorithm is effectively reduced through an alternate fusion strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent transportation technology, and in particular to a point cloud target tracking method, device, equipment and storage medium. Background Art

[0002] In the field of intelligent transportation, high-precision sensors are often used to accurately detect traffic participants, such as microwave radar, millimeter-wave radar, and cameras. Among them, millimeter-wave radar and vision are the most widely used. Millimeter-wave radar point cloud has good distance and speed accuracy and good real-time performance, but the target confidence is low, mainly because false targets are prone to appear under multipath reflection. Visual structured data has high accuracy, but high computing power requirements. When there are many traffic targets, the frame rate of stored data is unstable and the real-time performance is poor.

[0003] To solve the problems existing in millimeter-wave radar and vision applications, the industry usually uses radar-vision fusion to improve target tracking accuracy and real-time performance, including target-level radar-vision fusion and point cloud-level radar-vision fusion. Target-level radar-vision fusion loses the information in the original point cloud cluster due to radar target data loss and the mismatch of radar-vision frame rate, which leads to low accuracy of radar-vision fusion. Point cloud-level radar-vision fusion uses millimeter-wave radar, which makes it difficult to achieve the same fusion effect as lidar due to the sparse point cloud of millimeter-wave radar. Summary of the invention

[0004] The main purpose of the present invention is to provide a point cloud target tracking method, device, equipment and storage medium, aiming to solve the technical problems of low accuracy and poor fusion effect of existing radar and vision fusion.

[0005] A first aspect of the present invention provides a point cloud target tracking method, the point cloud target tracking method comprising:

[0006] Obtain video stream data and radar point cloud data;

[0007] Processing the video stream data to obtain visual target structured data;

[0008] Performing radar-visual fusion on the visual target structured data and the radar point cloud data, and outputting first point cloud frame data after radar-visual fusion or second point cloud frame data not after radar-visual fusion;

[0009] Based on the first point cloud frame data or the second point cloud frame data, a target within a field of view of a current frame is tracked.

[0010] Optionally, in a first implementation manner of the first aspect of the present invention, the processing the video stream data to obtain visual target structured data includes:

[0011] Inputting the video stream data into a trained target detection model to perform target detection, and outputting first target structured data;

[0012] Using a preset calibration algorithm, aligning the first target structured data with the radar point cloud data to convert them into second target structured data based on a radar polar coordinate system;

[0013] The second target structured data is input into a trained target tracking model to perform visual target tracking, and the visual target structured data is output.

[0014] Optionally, in a second implementation of the first aspect of the present invention, the performing radar-visual fusion on the visual target structured data and the radar point cloud data, and outputting first point cloud frame data after radar-visual fusion or second point cloud frame data not after radar-visual fusion comprises:

[0015] Mapping the visual target structured data into mapping data for performing radar-vision fusion;

[0016] Based on the mapping data, coordinate system conversion, regional filtering and point cloud clustering are performed on the radar point cloud data in sequence to achieve radar-visual fusion, and first point cloud frame data after radar-visual fusion or second point cloud frame data not after radar-visual fusion is output.

[0017] Optionally, in a third implementation of the first aspect of the present invention, the point cloud target tracking method further includes:

[0018] When performing point cloud clustering on the radar point cloud data after completing the coordinate system conversion and the regional filtering, a preset core point selection operator is used to determine the core points in the radar point cloud data;

[0019] Using a preset neighborhood point selection operator to filter out neighborhood points corresponding to the core point from the radar point cloud data;

[0020] The preset neighborhood maximum value filtering operator in each dimensional direction is used to include the core points and neighborhood points in the corresponding dimensional direction into the neighborhood clustering to obtain the point cloud clustering result;

[0021] The point cloud that has been processed by point cloud clustering in the point cloud clustering result is labeled to obtain first point cloud frame data with labels, wherein the unlabeled point cloud frame data is second point cloud frame data without labels.

[0022] Optionally, in a fourth implementation of the first aspect of the present invention, the calculation formula of the core point selection operator is as follows:

[0023] or

[0024]

[0025] Where P represents the radar point cloud and contains k items of data, MF i Represents the radar point cloud P and the i-th mapping data VYD i The membership degree between n Represents the weight of the nth item of data in the radar point cloud, μ i Represents the mean of each item in the i-th mapping data, M i Represents the covariance matrix of each item in the i-th mapping data.

[0026] Optionally, in a fifth implementation of the first aspect of the present invention, the visual target structured data includes the length, width, height, upper left corner coordinates of the visual target detection box, and the speed and type of the visual target; the radar point cloud data includes the physical position coordinates, speed and radar reflection cross-sectional area of ​​the radar target; and the mapping data includes the ID, physical size, physical position coordinates, speed and radar reflection cross-sectional area of ​​the visual target.

[0027] Optionally, in a sixth implementation manner of the first aspect of the present invention, tracking a target within a current frame field of view based on the first point cloud frame data or the second point cloud frame data includes:

[0028] Based on the first point cloud frame data or the second point cloud frame data, performing target prediction and target association on the point cloud targets within the field of view of the current frame;

[0029] If it is associated with a new point cloud target, the target is updated;

[0030] If it is not associated with a new point cloud target, determining whether the current frame is the first point cloud frame data;

[0031] If the current frame is the first point cloud frame data, initializing a new point cloud target based on the first point cloud frame data and tracking it;

[0032] If the current frame is the second point cloud frame data, the original point cloud target within the field of view of the current frame continues to be tracked.

[0033] A second aspect of the present invention further provides a point cloud target tracking device, the point cloud target tracking device comprising:

[0034] Data acquisition module, used to acquire video stream data and radar point cloud data;

[0035] A visual data processing module, used for processing the video stream data to obtain visual target structured data;

[0036] A radar-visual fusion module, used for performing radar-visual fusion on the visual target structured data and the radar point cloud data, and outputting first point cloud frame data after radar-visual fusion or second point cloud frame data not after radar-visual fusion;

[0037] The target tracking module is used to track the target within the current frame field of view based on the first point cloud frame data or the second point cloud frame data.

[0038] Optionally, in a first implementation of the second aspect of the present invention, the visual data processing module is specifically used to:

[0039] Inputting the video stream data into a trained target detection model to perform target detection, and outputting first target structured data;

[0040] Using a preset calibration algorithm, aligning the first target structured data with the radar point cloud data to convert them into second target structured data based on a radar polar coordinate system;

[0041] The second target structured data is input into a trained target tracking model to perform visual target tracking, and the visual target structured data is output.

[0042] Optionally, in a second implementation of the second aspect of the present invention, the radar and vision fusion module is specifically used to:

[0043] Mapping the visual target structured data into mapping data for performing radar-vision fusion;

[0044] Based on the mapping data, coordinate system conversion, regional filtering and point cloud clustering are performed on the radar point cloud data in sequence to achieve radar-visual fusion, and first point cloud frame data after radar-visual fusion or second point cloud frame data not after radar-visual fusion is output.

[0045] Optionally, in a third implementation of the second aspect of the present invention, the radar and vision fusion module includes:

[0046] The point cloud clustering unit is used to use a preset core point selection operator to determine the core point in the radar point cloud data when performing point cloud clustering on the radar point cloud data after completing the coordinate system conversion and regional filtering; use a preset neighborhood point selection operator to filter out each neighborhood point corresponding to the core point from the radar point cloud data; use a preset neighborhood maximum value filtering operator in each dimensional direction to include the core point and the neighborhood point corresponding to each dimensional direction into the neighborhood clustering respectively, to obtain a point cloud clustering result; mark the point cloud processed by the point cloud clustering in the point cloud clustering result to obtain the first point cloud frame data with a label, wherein the unlabeled point cloud frame data is the second point cloud frame data without a label.

[0047] Optionally, in a fourth implementation of the second aspect of the present invention, the calculation formula of the core point selection operator is as follows:

[0048] or

[0049]

[0050] Where P represents the radar point cloud and contains k items of data, MF i Represents the radar point cloud P and the i-th mapping data VYD i The membership degree between n Represents the weight of the nth item of data in the radar point cloud, μ i Represents the mean of each item in the i-th mapping data, M i Represents the covariance matrix of each item in the i-th mapping data.

[0051] Optionally, in a fifth implementation of the second aspect of the present invention, the visual target structured data includes the length, width, height, upper left corner coordinates of the visual target detection box, and the speed and type of the visual target; the radar point cloud data includes the physical position coordinates, speed and radar reflection cross-sectional area of ​​the radar target; and the mapping data includes the ID, physical size, physical position coordinates, speed and radar reflection cross-sectional area of ​​the visual target.

[0052] Optionally, in a sixth implementation of the second aspect of the present invention, the target tracking module is specifically used to:

[0053] Based on the first point cloud frame data or the second point cloud frame data, performing target prediction and target association on the point cloud targets within the field of view of the current frame;

[0054] If it is associated with a new point cloud target, the target is updated;

[0055] If it is not associated with a new point cloud target, determining whether the current frame is the first point cloud frame data;

[0056] If the current frame is the first point cloud frame data, initializing a new point cloud target based on the first point cloud frame data and tracking it;

[0057] If the current frame is the second point cloud frame data, the original point cloud target within the field of view of the current frame continues to be tracked.

[0058] A third aspect of the present invention provides a computer device, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory so that the computer device executes the above-mentioned point cloud target tracking method.

[0059] A fourth aspect of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the above-mentioned point cloud target tracking method.

[0060] The present invention provides a point cloud target tracking method based on radar-visual fusion, firstly obtaining video stream data and radar point cloud data, then processing the video stream data to obtain visual target structured data, then performing radar-visual fusion on the visual target structured data and the radar point cloud data, outputting the first point cloud frame data after radar-visual fusion or the second point cloud frame data not after radar-visual fusion, and finally tracking the target within the field of view of the current frame based on the first point cloud frame data or the second point cloud frame data. The present invention proposes an algorithm strategy for merging visual detection data into the radar point cloud tracking process, outputting point cloud frame data after radar-visual fusion or point cloud frame data not after radar-visual fusion through radar-visual fusion, thereby effectively solving the problem of radar false target generation and inaccurate clustered targets, and improving the accuracy of point cloud clustering filtering. In addition, the present invention also proposes a target tracking strategy for alternating fused point cloud frames and unfused point cloud frames, thereby effectively reducing computing power requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 A schematic diagram of an embodiment of a point cloud target tracking method in an embodiment of the present invention;

[0062] Figure 2 A schematic diagram of an embodiment of a point cloud target tracking device in an embodiment of the present invention;

[0063] Figure 3 FIG. 1 is a schematic diagram of an embodiment of a computer device in an embodiment of the present invention. DETAILED DESCRIPTION

[0064] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0065] For ease of understanding, the specific process of the embodiment of the present invention is described below. Figure 1, an embodiment of the point cloud target tracking method in an embodiment of the present invention includes:

[0066] 101. Obtain video stream data and radar point cloud data;

[0067] This embodiment is applied to the field of intelligent transportation, and is specifically used for accurate detection of traffic participants, such as digital roads, intelligent intersections, holographic tunnels, intelligent driving, vehicle-road collaboration and other scenarios. For example, in intelligent driving scenarios, cameras and radars can be installed on intelligent vehicles, and in intelligent intersection traffic management, cameras and radars can be installed at intersections to detect traffic participants (such as other vehicles, pedestrians, etc.) that may exist in one or more directions of the vehicle. Among them, the camera can generate and output video stream data in real time, and the radar (such as microwave radar, millimeter wave radar) can generate and output radar point cloud data in real time. Among them, the frame rate of video stream data is usually lower than the frame rate of radar point cloud data. For example, the frame rate of video stream data ranges from 5 to 10FPS (Frames Per Second), and the frame rate of radar point cloud data ranges from 15 to 20FPS.

[0068] In addition, it is necessary to further explain that, in order to realize the radar-visual fusion of this embodiment, it is generally required that the camera and the radar have the same or intersecting collection fields of view, and use the data collected at the same time point for fusion. Of course, in some embodiments, the camera and the radar may have different collection fields of view or no intersection.

[0069] 102. Process the video stream data to obtain visual target structured data;

[0070] In this embodiment, since the video stream data and the radar point cloud data are generated by different sources, data processing is required first. The radar-visual fusion of this embodiment is mainly based on radar point cloud data, and the video stream data is processed so that the processed video stream data can be fused with the radar point cloud data. In this embodiment, the corresponding visual target structured data is obtained by processing the acquired video stream data.

[0071] In an optional embodiment, the above step 102 further includes:

[0072] 1021. Input the video stream data into a trained target detection model to perform target detection, and output first target structured data;

[0073] In this step, the acquired video stream data is input into a trained target detection model (the target detection model can be a computer vision algorithm commonly used in the industry, such as the YOLO series of algorithms). The task of the target detection model is to identify and locate targets of interest (such as vehicles, pedestrians, etc.) in the video frame. After the target detection model processes the video stream data, it will output target structured data. Target structured data usually includes timestamp, pixel-level target detection box, center point, target type (such as vehicle, pedestrian, etc.), confidence and other information, which is presented in a structured form to facilitate subsequent processing and analysis.

[0074] 1022. Using a preset calibration algorithm, align the first target structured data with the radar point cloud data to convert them into second target structured data based on a radar polar coordinate system;

[0075] In this step, since the video stream data and radar point cloud data come from different sensors, there may be a mismatch in time and space. Therefore, it is necessary to use a preset calibration algorithm (such as the internal and external parameter calibration method) to align the two types of data. The calibration algorithm usually includes the determination of parameters such as the relative position relationship between sensors and time synchronization. After the data is aligned, the first target structured data (based on the image coordinate system) will be converted into the second target structured data based on the radar polar coordinate system. The radar polar coordinate system is a coordinate system centered on the radar that describes the target position by distance and angle. This conversion allows the video stream data and radar point cloud data to be subsequently processed in the same coordinate system.

[0076] 1023. Input the second target structured data into a trained target tracking model to perform visual target tracking, and output the visual target structured data.

[0077] In this step, the second target structured data that has been converted to the radar polar coordinate system is input into a trained target tracking model (such as the Bytetrack model). The task of the target tracking model is to track the target of interest in continuous video frames and output dynamic information such as the position and speed of the target in each frame. These dynamic information are presented in the form of visual target structured data, including the trajectory, speed, acceleration, etc. of the target. These data can be used for subsequent decision-making, planning and other tasks, such as path planning and obstacle avoidance of self-driving cars, and intelligent intersection traffic control.

[0078] 103. Perform radar-visual fusion on the visual target structured data and the radar point cloud data, and output first point cloud frame data after radar-visual fusion or second point cloud frame data not after radar-visual fusion;

[0079] In this embodiment, radar-vision fusion refers to combining radar detection data with computer vision technology to analyze, process, and display images, ultimately achieving all-round monitoring and early warning of road traffic conditions. Specifically, radar can work around the clock, has strong penetration and anti-interference capabilities, and can detect information such as the distance and speed of objects; while computer vision technology can capture image information on the road, including vehicles, pedestrians, road signs, etc., and provide richer target feature information. By combining the advantages of both, more accurate and reliable target detection and tracking can be achieved.

[0080] The fusion of radar and vision can be applied to fields such as autonomous driving, intelligent traffic management, and smart cities. For example, at smart intersections, the fusion of radar and vision can accurately collect traffic flow, assist in the dynamic timing optimization of traffic signals and holographic presentation control, scientifically allocate road rights, and improve the efficiency of intersection traffic. For blind spots such as curves, steep slopes, mountainous areas with frequent fog, and tunnel entrances, the fusion of radar and vision can accurately detect the speed and distance of oncoming vehicles, and publish them in real time to the oncoming traffic guidance screen, providing two-way warnings to prevent road accidents.

[0081] In this embodiment, the acquired original radar point cloud data (data frame rate is higher than the data frame rate of the visual target structured data) and the visual target structured data obtained after processing the video stream data are fused and filtered to achieve radar-visual fusion, thereby obtaining fused filtered point cloud data with higher accuracy. It should be noted that after radar-visual fusion is performed in this embodiment, both the first point cloud frame data after radar-visual fusion and the second point cloud frame data without radar-visual fusion can be output.

[0082] In an optional embodiment, the above step 103 further includes:

[0083] 1031. Mapping the visual target structured data into mapping data for performing radar-visual fusion;

[0084] In this optional embodiment, it is assumed that the video stream data collected at time t is processed to generate visual target structured data VD, and there are i targets in total:

[0085] {Length1,Width1,Height1,X1,Y1,Z1,Speed1,Type1},

[0086] {Length2,Width2,Height2,X2,Y2,Z2,Speed2,Type2}.....

[0087] {Lengthi,Widthi,Heighti,Xi,Yi,Zi,Speedi,Typei}

[0088] Among them, Lengthi, Widthi, Heighti are the length, width and height of the visual target detection box, Xi, Yi, Zi are the coordinates of the starting point of the upper left corner of the detection box, Speedi is the target speed indirectly calculated during target tracking, and Typei is the target type, such as trucks, cars, non-motor vehicles, pedestrians, etc. If the visual target detection box is a 2D detection box, the height value Heighti is 0.

[0089] In this optional embodiment, the radar point cloud data generated by radar target detection includes information such as the physical position, size, speed, reflection intensity, etc. of the target. Therefore, according to the radar target detection principle, the visual target structured data VD can be mapped into the mapping data VYD that can be fused with the radar point cloud, as shown in the following table.

[0090] Id NewLength NewWidth NewHeight X Y Z RCS Speed 1 <![CDATA[Length1]]> <![CDATA[Width1]]> <![CDATA[Height1]]> <![CDATA[X1]]> <![CDATA[Y1]]> <![CDATA[Z1]]> <![CDATA[RCS1]]> <![CDATA[Speed1]]> 2 ... i <![CDATA[Length n ]]> <![CDATA[Width n ]]> <![CDATA[Height n ]]> <![CDATA[X n ]]> <![CDATA[Y n ]]> <![CDATA[Z n ]]> <![CDATA[RCS n ]]> <![CDATA[Speed n ]]>

[0091] As shown in the above table, one Id represents one target, NewLength, NewWidth, and NewHeight represent the size of the mapped target, X, Y, and Z represent the physical position of the mapped target (the coordinates of the starting point of the upper left corner of the detection box in the visual target structured data can be taken), RCS represents the radar reflection cross-sectional area of ​​the mapped target, and Speed ​​represents the speed of the mapped target (the speed in the visual target structured data can be taken).

[0092] Although the visual target detection box has size information, it is not directly used due to its detection box accuracy. Instead, the visual target structured data of different target types are mapped to different physical typical sizes and different radar reflection cross-sections (RCS). For example, the visual target structured data of a large truck can be mapped to {15, 2, 3, 20}, where 15, 2, 3 represent the length, width and height of the large truck, and 20 represents the radar reflection cross-section of the large truck; the visual target structured data of a car can be mapped to {5, 1.5, 1, 10} (different models can be further subdivided), the visual target structured data of a non-motor vehicle can be mapped to {1.5, 0.5, 1.2, 5}, and the visual target structured data of a pedestrian can be mapped to {1.5, 0.5, 1.2, 5}.

[0093] 1032. Based on the mapping data, coordinate system conversion, regional filtering and point cloud clustering are performed on the radar point cloud data in sequence to achieve radar-visual fusion, and the first point cloud frame data after radar-visual fusion or the second point cloud frame data not after radar-visual fusion is output.

[0094] In this optional embodiment, radar-visual fusion mainly includes: radar coordinate conversion from polar coordinate system to Cartesian coordinate system, regional filtering, point cloud clustering and other steps, and through radar-visual fusion, high-quality point cloud data is provided to the downstream tracking module.

[0095] (1) Coordinate system conversion

[0096] The polar coordinate system determines the position of a point by two parameters: distance (r) and angle (θ), while the Cartesian coordinate system determines the position of a point by the coordinates (x, y) on the x-axis and y-axis. The coordinate system conversion formula is as follows:

[0097] x=r*cos(θ), y=r*sin(θ);

[0098] Among them, r is the distance from the point to the origin, and θ is the angle between the point and the x-axis.

[0099] (2) Regional filtering

[0100] Regional filtering refers to filtering a specific area of ​​an image. The principle is to suppress noise and highlight the characteristic information of the target area by performing continuous neighborhood processing on the image pixel values. In regional filtering, the user first selects or defines the area to be processed, and then applies a specific filtering algorithm to the area.

[0101] Users can manually select (such as drawing a polygon or rectangle on the displayed image) to determine the area to be processed, or use automatic algorithms (such as image content-based segmentation algorithms) to identify and select specific areas. Once the region of interest is selected, various filtering algorithms can be applied to the region, such as Gaussian filtering, mean filtering, median filtering, bilateral filtering, etc. Filtering algorithms can smooth images, remove noise, enhance edges, etc. The specific effect depends on the selected algorithm and parameter settings.

[0102] (3) Point cloud clustering

[0103] Point cloud data consists of a large number of discrete points, which represent the three-dimensional information of objects or scenes. The goal of point cloud clustering is to divide these discrete points into different groups, each group representing an object or part of a scene. The principle of point cloud clustering is mainly based on the relationship or features between points. Through calculation and analysis, points with similar features or relationships are classified into the same category. Common point cloud clustering algorithms include distance-based clustering algorithms, which determine whether two points belong to the same cluster based on the distance between them. Common distance measurement methods include Euclidean distance, Manhattan distance, and Chebyshev distance. Density-based clustering algorithms determine clustering by calculating the density of points, calculating the number of neighboring points around each point, and then dividing the points into core points, boundary points, and noise points according to the set density threshold.

[0104] In an optional embodiment, in the above step 1032, when performing point cloud clustering on the radar point cloud data after completing the coordinate system conversion and regional filtering, the following processing content is further added to improve the effect of point cloud clustering, specifically including:

[0105] (1) using a preset core point selection operator to determine the core point in the radar point cloud data;

[0106] The core point selection operator is a method used to select key points from a data set. These key points usually have some special characteristics, such as high density, long distance, or significant features. By selecting these core points, we can better understand and analyze the data set, which is convenient for subsequent point cloud clustering.

[0107] In an optional embodiment, the calculation formula of the core point selection operator is as follows:

[0108] Formula 1:

[0109] Formula 2:

[0110] Where P represents the radar point cloud and contains k items of data, MF i Represents the radar point cloud P and the i-th mapping data VYD i The membership degree between n Represents the weight of the nth item of data in the radar point cloud, μ i Represents the mean of each item in the i-th mapping data, M i Represents the covariance matrix of each item in the i-th mapping data. It should be noted that if σ is used n As weights, Formula 1 is used, and if the covariance matrix is ​​used as weights, Formula 2 is used.

[0111] Different from the random selection in general point cloud clustering, this embodiment needs to perform point cloud membership calculation when initializing the core point. The original radar point cloud data belonging to the mapping data VYD can initialize the core point and inherit the attribute information of the corresponding mapping data VYD. For a single point cloud P{PX, PY, PZ, PSpeed, PRCS}, select the corresponding 5 VYD data, and use the above formula 1 or 2 to calculate the point cloud P and each mapping data VYD in each group of mapping data VYD i The membership degree MF i Among them, the visual target structured data at the same time point corresponds to a group of mapping data, and a group of mapping data contains multiple mapping data, and a mapping data contains multiple VYD data. By setting the membership MF threshold, a mapping data within the set threshold and with the smallest membership is used as the membership data corresponding to the point cloud P.

[0112] (2) using a preset neighborhood point selection operator to select neighborhood points corresponding to the core point from the radar point cloud data;

[0113] Neighborhood is used to describe the range or area around a point. For a point in a given mathematical space, the neighborhood is a subset that contains the point, which can be determined by some distance functions defined on the space. In common real number spaces, the neighborhood is usually defined as an open interval centered on the point. For example, for point a in real number space, the neighborhood can be expressed as (ar, a+r), where r is a positive number representing the radius of the neighborhood. The neighborhood point selection operator refers to selecting a central point in a given mathematical space or data set, and selecting other points around the point (i.e., in the neighborhood) according to certain rules or conditions.

[0114] After determining the core point, it is necessary to add the neighboring points to the cluster according to certain rules. The typical method is to search with the minimum radius ε, calculate the Euclidean distance between the core point and the point to be screened, and include the core point if it is less than ε. However, this rule does not take into account the characteristics of point cloud data in different dimensions. In this embodiment, after the membership calculation in the above step, the point cloud inherits the different dimensional information of the membership data (i.e., the mapping data), so the neighborhood point selection can be realized by calculating the Euclidean distance of different dimensions. The new neighborhood calculation function NF is as follows:

[0115]

[0116] Wherein, P{PX, PY, PZ, PSpeed, PRCS} represents the target physical position, target radar reflection cross-sectional area and target speed of the reference point cloud P; Pn{PXn, PYn, PZn, PSpeedn, PRCSn} represents the target physical position, target radar reflection cross-sectional area and target speed of the point cloud Pn of the nth neighborhood point to be calculated; Δ represents the difference in different dimensions (target physical position, target radar reflection cross-sectional area and target speed); XG n , YG n , ZG n ,SpeedG n ,RCSG n are the threshold values ​​of different dimensions of point cloud Pn; NF n is the result of neighborhood calculation. When all conditional equations of the neighborhood calculation function are true, that is, when the differences in different dimensions are all within the corresponding neighborhood threshold values, the point cloud Pn can be determined to be a neighborhood point of the point cloud P.

[0117] (3) Using the preset neighborhood maximum value filtering operator in each dimensional direction, the core points and neighborhood points corresponding to each dimensional direction are included in the neighborhood clustering to obtain the point cloud clustering result;

[0118] The neighborhood maximum filter operator is used to reduce the impact of noise and outliers on clustering results by filtering out points with excessively large values ​​in certain dimensions when processing high-dimensional data.

[0119] General clustering methods usually limit the minimum value to avoid introducing more core points, but after multiple reflections, the radar point cloud will have more reflection points in different dimensions, so it is necessary to limit the maximum number of points in the neighborhood. The radar only returns one point cloud in the resolution unit of different dimensions, so when other dimensions are fixed, the maximum number of point clouds in this dimension is the ratio of the dimension size to the dimension resolution.

[0120] For the three directions X, Y, and Z, there are maximum cluster points NUMx, NUMy, and NUMz according to the three capabilities of distance resolution Dres, horizontal resolution HORres, and vertical resolution VERres corresponding to the overall performance of the radar. For example, suppose the subset of the mapping data VYD to which the point cloud P belongs is {Length n , Width n ,Height n , X n , Y n , Z n , RCS n ,Speed n}, then the maximum number of points in different dimensions is calculated as:

[0121] NUMX=Length n / Dres

[0122] NUMY=Width n / HORres

[0123] NUMZ=Height n / VERres

[0124] Among them, Dres, HORres, and VERres are known numbers and are related to the radar performance. According to the maximum number of points in different dimensions obtained by the above calculation, the core points in each dimension direction are included in the neighborhood clustering. If the number of points exceeds the maximum number of points in the corresponding dimension, they will no longer be included in the neighborhood, thereby achieving false point cloud filtering.

[0125] (4) Labeling the point cloud that has been processed by point cloud clustering in the point cloud clustering result to obtain first point cloud frame data with labels, wherein the unlabeled point cloud frame data is second point cloud frame data without labels.

[0126] After completing the core point selection, neighborhood point selection and neighborhood maximum value filtering, the point cloud clustering result after the radar-visual fusion of the original radar point cloud and the visual target structured data can be obtained. Considering that some original radar point clouds have not achieved radar-visual fusion after the above processing, it is necessary to label each point cloud in the point cloud clustering result. The point cloud that has been fused with radar-visual fusion is the labeled first point cloud frame data, and the point cloud that has not been fused with radar-visual fusion is the unlabeled second point cloud frame data.

[0127] In this optional embodiment, considering that the existing point cloud clustering is very dependent on input parameters, especially when the distribution characteristics of the unknown target point cloud are selected, the clustering parameters (neighborhood radius, minimum cluster point, etc.) lack a priori basis, and the clustering effect is not good, and the false points in the radar point cloud cannot be filtered out, resulting in low target tracking accuracy. Therefore, this optional embodiment uses the visual target detection results as basic data, and maps them into mapping data acceptable to point cloud clustering through the radar target detection principle. On the basis of the traditional clustering algorithm, a core point selection operator, a neighborhood data selection operator, and a maximum cluster value filtering operator of different dimensions in the field are added to perform data fusion filtering, and a fused filtered point cloud (called long-time frame data, i.e., the first point cloud frame data) is output, which can filter out false point clouds and improve clustering accuracy. In addition, considering that the visual data frame rate is lower than the radar point cloud frame rate, the point cloud data that has not been processed by fusion filtering can also be output (called short-time frame data, i.e., the second point cloud frame data), and is also passed to the downstream tracking module.

[0128] 104. Track a target within a field of view of a current frame based on the first point cloud frame data or the second point cloud frame data.

[0129] In this embodiment, through the radar-visual fusion in step 103, two different types of point cloud frame data are output: one is the first point cloud frame data after the radar-visual fusion, and the other is the second point cloud frame data that has not been fused by the radar-visual fusion (that is, the original radar point cloud data that does not meet the radar-visual fusion conditions). In this embodiment, the accuracy of point cloud clustering filtering is improved by the radar-visual fusion, and false point clouds are filtered out at the same time. In this embodiment, the visual and radar alternating fusion strategy is adopted, so that the visual structured data can run at a frame rate several times lower than the radar point cloud, which can not only improve the target tracking accuracy, but also effectively reduce the computing power requirements.

[0130] In an optional embodiment, the above step 104 further includes:

[0131] 1041. Based on the first point cloud frame data or the second point cloud frame data, perform target prediction and target association on the point cloud target within the field of view of the current frame;

[0132] 1042. If associated with a new point cloud target, then update the target;

[0133] 1043. If the new point cloud target is not associated, determine whether the current frame is the first point cloud frame data;

[0134] 1044. If the current frame is the first point cloud frame data, initialize a new point cloud target based on the first point cloud frame data and perform tracking;

[0135] 1045. If the current frame is the second point cloud frame data, continue to track the original point cloud target within the field of view of the current frame.

[0136] For point cloud target tracking algorithms, they are generally divided into target initialization, target prediction, target association, target update and other contents.

[0137] (1) Target initialization

[0138] Target initialization is the starting step of point cloud target tracking. In this stage, the algorithm needs to identify and extract the target of interest from the initial point cloud data. The target can be static (such as buildings) or dynamic (such as vehicles, pedestrians). The initialization process usually involves feature extraction of the target, which can be geometric features (such as size, shape), motion features (such as speed, acceleration) or other attributes that can distinguish the target.

[0139] (2) Target prediction

[0140] In the process of target tracking, due to the delay or lack of sensor data collection, it is necessary to predict the position and state of the target at the next moment. The prediction model can be based on physics (such as Kalman filter, particle filter) or learning (such as deep learning model). These models use the previously observed target state to estimate the future state of the target.

[0141] Target prediction is based on different filter algorithm selections, such as Kalman filter, Bayesian filter, α-β filter, particle filter, etc. The difference lies in the weight of observation and measurement. This optional embodiment does not make any requirements on the filter algorithm. For example, the Kalman filter algorithm can select kinematic models such as uniform acceleration model, uniform speed model, constant torque, etc. for target prediction.

[0142] (3) Target association

[0143] When new point cloud data is collected, the algorithm needs to determine the correspondence between these newly observed points and the previously predicted targets. This usually involves calculating the similarity or distance between the new observation points and the predicted targets, and matching them according to these metrics. The association algorithm can be based on nearest neighbor search, Hungarian algorithm or other optimization methods. In the point cloud association stage, the focus is on the matching relationship between the point cloud and the existing targets, which can be associated through the nearest neighbor algorithm, and the global optimal association result can be obtained through Hungarian matching.

[0144] (4) Target update

[0145] After the target is successfully initialized and associated with the new observation point, the algorithm will update the state of the target (such as position, speed, etc.). This update process may be iterative, that is, as new data arrives, the state of the target will be continuously updated. After associating with the new target, the weights of the observed value and the predicted value are calculated according to the selected tracking algorithm, such as the Kalman filter through the covariance matrix calculation, and the α-β filter through the α and β values.

[0146] This embodiment defines the point cloud data fused by the radar and vision as long-time frame data (i.e., the first point cloud frame data), and the unfused radar point cloud data as short-time frame data (i.e., the second point cloud frame data). In order to achieve the fusion tracking of the video data frame being lower than the radar point cloud frame, this optional embodiment makes full use of the characteristics of the two frames of frame data: the long-time frame has high data accuracy and is responsible for the generation of new point cloud targets, and the short-time frame has high data frame rate and is responsible for the stable tracking of existing targets.

[0147] In the target association stage of the target tracking algorithm, if the new target is not associated, the general tracking algorithm will directly generate a new target, thereby causing a false target. However, this optional embodiment determines the label of the point cloud data frame, and initializes the new target if it is a long-time frame, and continues to track the existing target if it is a short-time data frame. For new targets that need to be initialized, the mean of the point cloud frame data with the same label in the point cloud clustering result is calculated according to the output result of the fusion filter module, and a new target is generated and tracked according to the mean of the point cloud frame data.

[0148] The point cloud target tracking algorithm based on radar and visual fusion proposed in this embodiment proposes an algorithm strategy for merging visual detection data into the radar point cloud tracking process, and applies the radar and visual fusion results to point cloud clustering, which can effectively solve the problems of false radar target generation and inaccurate clustering targets. In addition, the point cloud target tracking algorithm based on radar and visual fusion proposed in this embodiment also proposes a long and short time frame alternating tracking strategy. The point cloud data frame after radar and visual fusion is a long time frame with high accuracy, which is used for new target generation, while the point cloud data frame without radar and visual fusion is a short time frame with high data frame rate, which is used for existing target tracking, which can effectively reduce the algorithm's demand for computing power.

[0149] The above describes the point cloud target tracking method in the embodiment of the present invention. The following describes the point cloud target tracking device in the embodiment of the present invention. Figure 2 In one embodiment of the present invention, a point cloud target tracking device includes:

[0150] Data acquisition module 201, used to acquire video stream data and radar point cloud data;

[0151] A visual data processing module 202 is used to process the video stream data to obtain visual target structured data;

[0152] A radar-visual fusion module 203 is used to perform radar-visual fusion on the visual target structured data and the radar point cloud data, and output first point cloud frame data after radar-visual fusion or second point cloud frame data not after radar-visual fusion;

[0153] The target tracking module 204 is used to track the target within the field of view of the current frame based on the first point cloud frame data or the second point cloud frame data.

[0154] Optionally, in one embodiment, the visual data processing module 202 is specifically used for:

[0155] Inputting the video stream data into a trained target detection model to perform target detection, and outputting first target structured data;

[0156] Using a preset calibration algorithm, aligning the first target structured data with the radar point cloud data to convert them into second target structured data based on a radar polar coordinate system;

[0157] The second target structured data is input into a trained target tracking model to perform visual target tracking, and the visual target structured data is output.

[0158] Optionally, in one embodiment, the radar and visual fusion module 203 is specifically used for:

[0159] Mapping the visual target structured data into mapping data for performing radar-vision fusion;

[0160] Based on the mapping data, coordinate system conversion, regional filtering and point cloud clustering are performed on the radar point cloud data in sequence to achieve radar-visual fusion, and first point cloud frame data after radar-visual fusion or second point cloud frame data not after radar-visual fusion is output.

[0161] Optionally, in one embodiment, the radar and vision fusion module 203 includes:

[0162] The point cloud clustering unit is used to use a preset core point selection operator to determine the core point in the radar point cloud data when performing point cloud clustering on the radar point cloud data after completing the coordinate system conversion and regional filtering; use a preset neighborhood point selection operator to filter out each neighborhood point corresponding to the core point from the radar point cloud data; use a preset neighborhood maximum value filtering operator in each dimensional direction to include the core point and the neighborhood point corresponding to each dimensional direction into the neighborhood clustering respectively, to obtain a point cloud clustering result; mark the point cloud processed by the point cloud clustering in the point cloud clustering result to obtain the first point cloud frame data with a label, wherein the unlabeled point cloud frame data is the second point cloud frame data without a label.

[0163] Optionally, in one embodiment, the calculation formula of the core point selection operator is as follows:

[0164] or

[0165]

[0166] Where P represents the radar point cloud and contains k items of data, MF i Represents the radar point cloud P and the i-th mapping data VYD i The membership degree between n Represents the weight of the nth item of data in the radar point cloud, μ i Represents the mean of each item in the i-th mapping data, M i Represents the covariance matrix of each item in the i-th mapping data.

[0167] Optionally, in one embodiment, the visual target structured data includes the length, width, height, upper left corner coordinates of the visual target detection box, and the speed and type of the visual target; the radar point cloud data includes the physical position coordinates, speed and radar reflection cross-sectional area of ​​the radar target; and the mapping data includes the ID, physical size, physical position coordinates, speed and radar reflection cross-sectional area of ​​the visual target.

[0168] Optionally, in one embodiment, the target tracking module 204 is specifically used for:

[0169] Based on the first point cloud frame data or the second point cloud frame data, performing target prediction and target association on the point cloud targets within the field of view of the current frame;

[0170] If it is associated with a new point cloud target, the target is updated;

[0171] If it is not associated with a new point cloud target, determining whether the current frame is the first point cloud frame data;

[0172] If the current frame is the first point cloud frame data, initializing a new point cloud target based on the first point cloud frame data and tracking it;

[0173] If the current frame is the second point cloud frame data, the original point cloud target within the field of view of the current frame continues to be tracked.

[0174] Since the embodiments of the device part correspond to the embodiments of the above-mentioned method, please refer to the above-mentioned method embodiments for the introduction of the point cloud target tracking device provided by the present invention. The present invention will not be repeated here, and it has the same beneficial effects as the above-mentioned point cloud target tracking method.

[0175] above Figure 2 The point cloud target tracking device in the embodiment of the present invention is described in detail from the perspective of modular functional entities, and the computer device in the embodiment of the present invention is described in detail from the perspective of hardware processing.

[0176] Figure 3 1 is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. The computer device 500 may have relatively large differences due to different configurations or performances, and may include one or more processors (central processing units, CPU) 510 (for example, one or more processors) and a memory 520, and one or more storage media 530 (for example, one or more mass storage devices) storing application programs 533 or data 532. Among them, the memory 520 and the storage medium 530 can be short-term storage or permanent storage. The program stored in the storage medium 530 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations in the computer device 500. Furthermore, the processor 510 may be configured to communicate with the storage medium 530 to execute a series of instruction operations in the storage medium 530 on the computer device 500.

[0177] The computer device 500 may also include one or more power supplies 540, one or more wired or wireless network interfaces 550, one or more input and output interfaces 560, and / or one or more operating systems 531, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. It will be appreciated by those skilled in the art that Figure 3 The illustrated computer device structure does not constitute a limitation on the computer device, and may include more or fewer components than illustrated, or combine certain components, or arrange the components differently.

[0178] The present invention also provides a computer device, which includes a memory and a processor, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the processor executes the steps of the point cloud target tracking method in the above-mentioned embodiments.

[0179] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions are executed on a computer, the computer executes the steps of the point cloud target tracking method.

[0180] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0181] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program codes.

[0182] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A point cloud target tracking method, characterized in that: The point cloud target tracking method comprises: Obtain video stream data and radar point cloud data; Processing the video stream data to obtain visual target structured data; Performing radar-visual fusion on the visual target structured data and the radar point cloud data, and outputting first point cloud frame data after radar-visual fusion or second point cloud frame data not after radar-visual fusion; Based on the first point cloud frame data or the second point cloud frame data, a target within a field of view of a current frame is tracked.

2. The point cloud target tracking method according to claim 1, characterized in that: The step of processing the video stream data to obtain visual target structured data includes: Inputting the video stream data into a trained target detection model to perform target detection, and outputting first target structured data; Using a preset calibration algorithm, aligning the first target structured data with the radar point cloud data to convert them into second target structured data based on a radar polar coordinate system; The second target structured data is input into a trained target tracking model to perform visual target tracking, and the visual target structured data is output.

3. The point cloud target tracking method according to claim 1, characterized in that: The step of fusing the visual target structured data with the radar point cloud data and outputting the first point cloud frame data after the radar point cloud fusion or the second point cloud frame data not after the radar point cloud fusion comprises: Mapping the visual target structured data into mapping data for performing radar-vision fusion; Based on the mapping data, coordinate system conversion, regional filtering and point cloud clustering are performed on the radar point cloud data in sequence to achieve radar-visual fusion, and first point cloud frame data after radar-visual fusion or second point cloud frame data not after radar-visual fusion is output.

4. The point cloud target tracking method according to claim 3, characterized in that: The point cloud target tracking method further includes: When performing point cloud clustering on the radar point cloud data after completing the coordinate system conversion and the regional filtering, a preset core point selection operator is used to determine the core points in the radar point cloud data; Using a preset neighborhood point selection operator to filter out neighborhood points corresponding to the core point from the radar point cloud data; The preset neighborhood maximum value filtering operator in each dimensional direction is used to include the core points and neighborhood points in the corresponding dimensional direction into the neighborhood clustering to obtain the point cloud clustering result; The point cloud that has been processed by point cloud clustering in the point cloud clustering result is labeled to obtain first point cloud frame data with labels, wherein the unlabeled point cloud frame data is second point cloud frame data without labels.

5. The point cloud target tracking method according to claim 4, characterized in that: The calculation formula of the core point selection operator is as follows: or Where P represents the radar point cloud and contains k items of data, MF i Represents the radar point cloud P and the i-th mapping data VYD i The membership degree between n Represents the weight of the nth item of data in the radar point cloud, μ i Represents the mean of each item in the i-th mapping data, M i Represents the covariance matrix of each item in the i-th mapping data.

6. The point cloud target tracking method according to claim 3, characterized in that: The visual target structured data includes the length, width, height, upper left corner coordinates of the visual target detection box, and the speed and type of the visual target; the radar point cloud data includes the physical position coordinates, speed and radar reflection cross-sectional area of ​​the radar target; the mapping data includes the ID, physical size, physical position coordinates, speed and radar reflection cross-sectional area of ​​the visual target.

7. The point cloud target tracking method according to any one of claims 1 to 6, characterized in that: The tracking of the target within the current frame field of view based on the first point cloud frame data or the second point cloud frame data includes: Based on the first point cloud frame data or the second point cloud frame data, performing target prediction and target association on the point cloud targets within the field of view of the current frame; If it is associated with a new point cloud target, the target is updated; If it is not associated with a new point cloud target, determining whether the current frame is the first point cloud frame data; If the current frame is the first point cloud frame data, initializing a new point cloud target based on the first point cloud frame data and tracking it; If the current frame is the second point cloud frame data, the original point cloud target within the field of view of the current frame continues to be tracked.

8. A point cloud target tracking device, characterized in that: The point cloud target tracking device comprises: Data acquisition module, used to acquire video stream data and radar point cloud data; A visual data processing module, used for processing the video stream data to obtain visual target structured data; A radar-visual fusion module, used for performing radar-visual fusion on the visual target structured data and the radar point cloud data, and outputting first point cloud frame data after radar-visual fusion or second point cloud frame data not after radar-visual fusion; The target tracking module is used to track the target within the current frame field of view based on the first point cloud frame data or the second point cloud frame data.

9. A computer device, characterized in that: The computer device comprises: a memory and at least one processor, wherein instructions are stored in the memory; The at least one processor calls the instructions in the memory so that the computer device executes the point cloud target tracking method according to any one of claims 1 to 7.

10. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed by the processor, the point cloud target tracking method as described in any one of claims 1-7 is implemented.