BEV feature-based vehicle tracking method, apparatus and device, and medium
Through the vehicle tracking method based on BEV features, point pillars network and convolutional neural network are used to convert point cloud data into cylinder feature maps. Combined with object detection algorithms and feature matching, the accuracy and adaptability of vehicle tracking in port environments are solved, and stable tracking is achieved in complex environments.
Patent Information
- Application Number
- CN202510765346.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-10
AI Technical Summary
Existing vehicle tracking methods are not accurate and poorly adaptable in port environments, especially in complex and dynamic port scenarios, which are difficult to effectively track vehicles of various shapes and sizes.
The vehicle tracking method based on BEV features is adopted, and the point cloud data is converted into a cylinder feature map through the point pillars network. Combined with the object detection algorithm and the convolutional neural network, the vehicle's target trajectory and detection box features are determined, and the feature matching is used to achieve the tight coupling of detection and tracking.
Improve the accuracy and flexibility of vehicle tracking, especially in complex environments, it can stably track all kinds of vehicles, reduce target loss and ID switching problems in occlusion and dense scenarios, and realize real-time tracking.
Smart Images

Figure CN120279533A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of vehicle tracking, and particularly to a vehicle tracking method, device, equipment and medium based on BEV features. Background Art
[0002] The port environment is a highly complex and dynamic scenario, which contains vehicles of various shapes and sizes, such as container trucks, tractors, trailers, forklifts, etc. In the port autonomous driving system, vehicle tracking is a crucial link, which provides the necessary dynamic scenario understanding for obstacle avoidance, path planning and decision-making.
[0003] Existing vehicle tracking methods generally proceed through the output bounding box information, adopting a two-stage framework of detection and tracking. First, a target detector is used to obtain the target bounding boxes in each frame, and then the detection results between different frames are associated through various association algorithms to form the motion trajectory of the target.
[0004] However, existing vehicle tracking methods have problems of low accuracy and poor adaptability. Summary of the Invention
[0005] The present application provides a vehicle tracking method, device, equipment and medium based on BEV features to solve the problems of low accuracy and poor adaptability existing in existing vehicle tracking methods.
[0006] In a first aspect, the present application provides a vehicle tracking method based on BEV features, and the method includes: Determine target point cloud data according to a collection device, and determine a pillar feature map corresponding to the target point cloud data according to the pointpillars network; Determine target trajectories corresponding to all vehicles in the pillar feature map according to a target detection algorithm; Determine a target detection box and a target feature region corresponding to the target detection box in the pillar feature map, and determine a trajectory feature corresponding to the target trajectory and a detection box feature corresponding to the target detection box according to the target feature region; Determine a target feature matching value according to the detection box feature, all trajectory features and a preset weight; Determine a target vehicle in the target detection box according to the target feature matching value and a preset matching threshold.
[0007] In some embodiments of the present application, determining target point cloud data according to a collection device and determining a pillar feature map corresponding to the target point cloud data according to the pointpillars network includes: Divide the target point cloud data into cylinders with a preset cylinder size, and perform feature encoding on the data in the cylinders according to the pointpillars network to obtain corresponding cylinder data features; Determine a cylinder feature map based on the spatial positions of the cylinders corresponding to the cylinder data features.
[0008] In some embodiments of the present application, determining a cylinder feature map based on the spatial positions of the cylinders corresponding to the cylinder data features includes: Determine a preset convolutional neural network; Determine a high-resolution cylinder feature map, a medium-resolution cylinder feature map, and a low-resolution cylinder feature map based on the preset convolutional neural network and the spatial positions of the cylinders; Determine a cylinder feature map based on the high-resolution cylinder feature map, the medium-resolution cylinder feature map, and the low-resolution cylinder feature map.
[0009] In some embodiments of the present application, determining the target trajectories corresponding to all vehicles in the cylinder feature map according to a target detection algorithm includes: Determine a convolutional region of a preset size in the cylinder feature map; Determine the vehicles in the convolutional region and the target trajectories corresponding to the vehicles according to the target detection algorithm.
[0010] In some embodiments of the present application, determine a target detection box and a target feature region corresponding to the target detection box in the cylinder feature map, and determine the trajectory feature corresponding to the target trajectory and the detection box feature corresponding to the target detection box according to the target feature region, including: Determine the high-resolution feature map, the medium-resolution feature map, and the low-resolution feature map in the cylinder feature map, and the corresponding feature regions respectively; Perform feature map fusion on the feature regions corresponding to the high-resolution feature map, the feature regions corresponding to the medium-resolution feature map, and the feature regions corresponding to the low-resolution feature map to obtain a target feature region; Determine the trajectory feature corresponding to the target trajectory and the detection box feature corresponding to the target detection box according to the target feature region.
[0011] In some embodiments of the present application, determine a target feature matching value according to the detection box feature, all trajectory features, and a preset weight, including: Determine the cosine similarity, intersection over union, and distance metric between the detection box feature and each trajectory feature, as well as the preset weight corresponding to the cosine similarity, the preset weight corresponding to the intersection over union, and the preset weight corresponding to the distance metric; Perform weighted combination on the cosine similarity, intersection over union, and distance metric according to the preset weight to obtain a target feature matching value.
[0012] In some embodiments of the present application, determining the target vehicle in the target detection box according to the target feature matching value and the preset matching threshold includes: Comparing the target feature matching value with the preset matching threshold to obtain a comparison result; If the comparison result is that the target feature matching value is greater than the preset matching threshold, determining that the vehicle in the target detection box is the target vehicle corresponding to the target trajectory; If the comparison result is that the target feature matching value is not greater than the preset matching threshold, determining that the vehicle in the target detection box is an unmatched vehicle.
[0013] In a second aspect, the present application provides a vehicle tracking device based on BEV features. The device includes: A feature map determination module, configured to determine target point cloud data according to a collection device, and determine a column feature map corresponding to the target point cloud data according to the pointpillars network; A trajectory determination module, configured to determine target trajectories corresponding to all vehicles in the column feature map according to a target detection algorithm; A feature determination module, configured to determine a target feature region corresponding to the target detection box and in the column feature map, and determine a trajectory feature corresponding to the target trajectory and a detection box feature corresponding to the target detection box according to the target feature region; A matching value determination module, configured to determine a target feature matching value according to the detection box feature, all trajectory features, and a preset weight; A vehicle determination module, configured to determine the target vehicle in the target detection box according to the target feature matching value and the preset matching threshold.
[0014] In a third aspect, the present application provides a computer device, including: a processor, and a memory communicatively connected to the processor; The memory stores computer execution instructions; The processor executes the computer execution instructions stored in the memory to implement the method of the present application.
[0015] In a fourth aspect, the present application provides a computer-readable storage medium, in which program code is stored, and when the program code is executed by a processor, it is used to implement the method of the present application.
[0016] The vehicle tracking method, device, equipment, and medium based on BEV features provided by this application determine target point cloud data according to the acquisition device, and determine the pillar feature map corresponding to the target point cloud data according to the pointpillars network; determine the target trajectories corresponding to all vehicles in the pillar feature map according to the target detection algorithm; determine the target detection frame and the target feature region corresponding to the target detection frame in the pillar feature map, and determine the trajectory feature corresponding to the target trajectory and the detection frame feature corresponding to the target detection frame according to the target feature region; determine the target feature matching value according to the detection frame feature, all trajectory features, and the preset weight; and determine the target vehicle in the target detection frame according to the target feature matching value and the preset matching threshold.
[0017] In this way, by determining the bounding box information output by the target detection module and introducing the BEV feature map into the tracking process in a tightly coupled manner, especially in the target matching and filter observation stages, accurate and stable tracking of various vehicles in the port environment is achieved. Different from the traditional separated architecture of detection and tracking, the detection module is no longer only used as the pre-input provider for the tracking module, but rather deep information sharing and fusion between the detection and tracking modules are realized. That is, the internal intermediate feature representation (BEV feature map) of the detection module directly participates in the core calculation process of the tracking module, forming a "feature-level" close collaboration, rather than simply being connected in series at the "result level", improving the flexibility and accuracy of vehicle tracking in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application.
[0019] Figure 1 It is a schematic flowchart of a vehicle tracking method based on BEV features provided by an embodiment of this application; Figure 2 It is a system architecture diagram of a vehicle tracking method based on BEV features provided by an embodiment of this application; Figure 3 It is a method schematic diagram of a vehicle tracking method based on BEV features provided by an embodiment of this application; Figure 4 It is a structural schematic diagram of a vehicle tracking device based on BEV features provided by an embodiment of this application; Figure 5 It is a structural block diagram of the equipment for executing the vehicle tracking method based on BEV features according to an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0021] The technical solution of the present application and how the technical solution of the present application solves the above technical problems will be described in detail below with specific embodiments. The following several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below with reference to the accompanying drawings.
[0022] Figure 1 It is a schematic flowchart of a vehicle tracking method based on BEV features provided by an embodiment of the present application. As Figure 1 shown, the vehicle tracking method based on BEV features may include the following steps: S110. Determine target point cloud data according to the acquisition device, and determine a pillar feature map corresponding to the target point cloud data according to the pointpillars network.
[0023] Among them, the acquisition device may be an on-vehicle lidar as the main sensor, and a camera may be optionally configured for auxiliary perception. Point cloud data is a discrete data set used to represent objects or scenes in three-dimensional space, consisting of a large number of three-dimensional coordinate points (x, y, z). Each point may usually also contain additional information such as color, reflection intensity, and normal direction, and records the geometric and attribute features on the object surface or in space through dense sampling.
[0024] The pointpillars network is a deep learning network for three-dimensional object detection, which can be used to process point cloud data and has the characteristics of high efficiency and accuracy.
[0025] The pillar feature map is a key feature representation form in the pointpillars network, which can convert point cloud data into a format suitable for two-dimensional convolution processing, enabling the subsequent two-dimensional convolution backbone network to effectively extract high-level features for target detection and three-dimensional bounding box regression.
[0026] Based on this, the point cloud data is determined by the on-vehicle lidar, and according to the pointpillars network, the three-dimensional point cloud data is converted into a pillar feature map that can be processed two-dimensionally, so as to subsequently determine whether the vehicle to be detected in the detection box is the target vehicle to be tracked according to the pillar feature map.
[0027] S120. Determine the target trajectories corresponding to all vehicles in the pillar feature map according to the target detection algorithm.
[0028] Among them, the target detection algorithm is an algorithm for detecting vehicles, which can be a detection head algorithm. In the pointpillars network, the detection head is a key component responsible for predicting the position, size, direction, and category of three-dimensional objects from the feature map, and can convert the high-level features extracted by the feature encoder and the backbone network into the final detection results.
[0029] The vehicle is all the vehicles included in the point cloud data, and the target trajectory is the driving trajectory corresponding to each vehicle.
[0030] Based on this, through the target detection algorithm, the driving trajectories of all vehicles in the collected point cloud data in the pillar feature map are determined, so as to subsequently determine whether the detection frame trajectory corresponding to the vehicle to be detected matches any vehicle trajectory, thereby determining whether the vehicle to be detected is a target vehicle.
[0031] S130. Determine the target detection frame and the target feature region corresponding to the target detection frame in the pillar feature map, and determine the trajectory feature corresponding to the target trajectory and the detection frame feature corresponding to the target detection frame according to the target feature region.
[0032] Among them, the target detection frame is the detection frame corresponding to the target detection algorithm, which is used to determine the current vehicle to be detected.
[0033] The target feature region is the region including the target detection frame and the target trajectory; the trajectory feature is the feature expression of the target trajectory of the vehicle in the target feature region, and the detection frame feature is the feature expression of the target detection frame in the target feature region. For example, the system is tracking a truck in the port, the target trajectory is the trajectory of the truck in the past 10 frames, the trajectory feature may include position, speed, size, and BEV feature. In a new frame, a bounding box is detected, that is, the target detection frame, which includes the target vehicle to be detected, which may be the same truck (or the position is offset due to occlusion), and the detection frame feature is the feature corresponding to the target vehicle.
[0034] Furthermore, for each detection frame di and the existing trajectory tj, define the region of interest (ROI) on the BEV feature map according to their positions and sizes, use the ROI Align or ROI Pooling method to extract feature map blocks from the multi-scale feature maps F1, F2, and F3, and fuse the feature map blocks of different scales to obtain the feature representation of each target, denoted as fi (detection frame feature) and fj (trajectory feature).
[0035] Based on this, by determining the target feature region in the cylinder feature map and determining the features corresponding to the target detection frame and the target vehicle trajectory in this target feature region, so as to perform feature matching according to the features subsequently, and thus determine whether the vehicle to be detected in the detection frame is the vehicle corresponding to the target trajectory.
[0036] S140. Determine the target feature matching value according to the detection frame feature, all trajectory features and a preset weight.
[0037] Among them, the preset weight is the weight value set in advance and is used to determine the target feature matching value by weighting; the target feature matching value is used to represent whether the detection frame feature and the trajectory feature match.
[0038] Based on this, by performing feature processing on the detection frame feature and the trajectory feature, for example, the cosine similarity, intersection over union and distance metric of the detection frame feature and the trajectory feature can be determined, so as to determine the target feature matching value, and thus further determine whether the vehicle to be detected in the detection frame is the vehicle corresponding to the target trajectory subsequently.
[0039] S150. Determine the target vehicle in the target detection frame according to the target feature matching value and a preset matching threshold.
[0040] Among them, the preset matching threshold is the matching threshold set in advance and is used to determine whether the target feature matching value meets the matching requirements determined by the user.
[0041] Based on this, by determining the preset matching threshold, and thus determining whether the vehicle to be detected is the vehicle corresponding to the target trajectory according to the target feature matching value and the preset matching threshold, the tracking accuracy is improved and real-time tracking is realized.
[0042] On the basis of the feasible implementation manner of the above S110, the present application further provides steps for determining target point cloud data according to a collection device and determining a cylinder feature map corresponding to the target point cloud data according to the pointpillars network, including: Divide the target point cloud data into cylinders with a preset cylinder size, and perform feature encoding on the data in the cylinders according to the pointpillars network to obtain corresponding cylinder data features; Determine the cylinder feature map according to the spatial position of the cylinders corresponding to the cylinder data features.
[0043] Among them, the preset cylinder size is the size of the data cylinder set in advance, so as to divide the point cloud data. For example, the point cloud data is divided into a plurality of vertical cylinders, and the bottom surface size of each cylinder is 0.16m×0.16m.
[0044] Feature encoding refers to the process of converting raw data (such as image pixels, text characters, sensor measurements, etc.) into an abstract feature representation that is more suitable for algorithm processing. The goal is to extract key information from the data and remove redundancy or noise; in the PointPillars network, feature encoding of the points in each pillar is a crucial step, which converts the raw point cloud data into a more expressive feature representation. Feature encoding of the points in each pillar includes features such as the three-dimensional coordinates of the points, density, and the offset of the points relative to the center of the pillar.
[0045] The spatial position is the spatial position arrangement of the pillar features. By processing the point features in each pillar, a pillar-level feature representation is generated, and the pillar features are arranged according to their spatial positions to form a BEV feature map with dimensions C×H×W, where C is the number of feature channels, and H and W are the height and width of the BEV feature map respectively.
[0046] Based on this, by determining the data pillars of a preset size corresponding to the point cloud data and performing feature encoding on the data pillars, pillar data features are obtained, and according to the arrangement order of the spatial positions of the pillars, the pillar feature map is determined.
[0047] Based on the feasible implementation manner of S110 above, the present application further provides steps for determining the pillar feature map according to the spatial position of the pillar corresponding to the pillar data feature, including: Determine a preset convolutional neural network; According to the preset convolutional neural network and the spatial position of the pillar, determine a high-resolution pillar feature map, a medium-resolution pillar feature map, and a low-resolution pillar feature map; According to the high-resolution pillar feature map, the medium-resolution pillar feature map, and the low-resolution pillar feature map, determine the pillar feature map.
[0048] Among them, the preset convolutional neural network is a pre-set convolutional neural network. For example, it can be a 2D convolutional neural network.
[0049] Based on this, through the convolutional neural network, three feature maps of different scales extracted from the BEV backbone network are respectively denoted as F1 (high resolution, low semantics), F2 (medium resolution, medium semantics), and F3 (low resolution, high semantics) feature maps, so as to directly introduce these feature maps into the tracking process subsequently, realizing the tight coupling of the detection and tracking modules.
[0050] Based on the feasible implementation manner of S120 above, the present application further provides steps for determining the target trajectories corresponding to all vehicles in the pillar feature map according to the target detection algorithm, including: Determine a convolutional region of a preset size in the pillar feature map; Determine the vehicles in the convolutional region and the target trajectories corresponding to the vehicles according to the object detection algorithm.
[0051] Among them, the convolutional region of a preset size can be understood as applying a convolutional layer of a preset size on the multi-scale BEV feature map; further, a 3×3 convolutional layer is applied on the multi-scale BEV feature map to generate a classification branch and a regression branch. The classification branch predicts the probabilities of each position belonging to different categories, including common vehicle types in the port such as cars, trucks, and trailers; the regression branch predicts the target geometric attributes of each position, including the center coordinates, dimensions, and orientation angles.
[0052] Based on this, by determining the convolutional region in the pillar feature map, the vehicles in the convolutional layer and their corresponding trajectories are determined, and the target trajectories are obtained.
[0053] On the basis of the feasible implementation manner of the above S130, the present application further provides steps for determining the object detection frame and the target feature region corresponding to the object detection frame in the pillar feature map, and determining the trajectory feature corresponding to the target trajectory and the detection frame feature corresponding to the object detection frame, including: Determine the high-resolution feature map, medium-resolution feature map, and low-resolution feature map in the pillar feature map, and the corresponding feature regions respectively; Perform feature map fusion on the feature regions corresponding to the high-resolution feature map, the feature regions corresponding to the medium-resolution feature map, and the feature regions corresponding to the low-resolution feature map to obtain the target feature region; Determine the trajectory feature corresponding to the target trajectory and the detection frame feature corresponding to the object detection frame according to the target feature region.
[0054] Among them, the high-resolution feature map, medium-resolution feature map, and low-resolution feature map are three feature maps of different scales extracted from the BEV backbone network through a convolutional neural network; the corresponding feature regions are the feature regions corresponding to the three feature maps of different scales respectively, so as to introduce the multi-scale feature maps in the detection network BEV backbone network into the target tracking process to enhance the accuracy of target matching and state estimation.
[0055] Based on this, the feature map blocks of different scales are fused, so as to determine the detection frame feature corresponding to the detection frame and the trajectory feature corresponding to the target trajectory in the pillar feature map, improving the accuracy of feature determination.
[0056] On the basis of the feasible implementation manner of the above S140, the present application further provides steps for determining the target feature matching value according to the detection frame feature, all trajectory features, and a preset weight, including: Determine the cosine similarity, intersection over union (IoU), and distance metric between the detection box features and each trajectory feature, as well as the preset weights corresponding to the cosine similarity, the preset weights corresponding to the intersection over union, and the preset weights corresponding to the distance metric; According to the preset weights, perform weighted combination on the cosine similarity, intersection over union, and distance metric to obtain the target feature matching value.
[0057] Among them, the detection box feature can be represented by fi, the trajectory feature can be represented by fj, the cosine similarity can be represented by The intersection over union can be represented by The distance metric can be represented by Then the cosine similarity , the intersection over union , the distance metric .
[0058] The weight corresponding to the cosine similarity can be represented by The weight corresponding to the intersection over union can be represented by The weight corresponding to the distance metric can be represented by The target feature matching value can be represented by Then , and the preset weights w1, w2, and w3 are weight coefficients, which can be determined by experimental tuning.
[0059] Based on this, by performing weighted combination on the cosine similarity, intersection over union, and distance metric between the detection box features and the trajectory features, the target feature matching value is obtained.
[0060] On the basis of the feasible implementation manner of the above S150, the present application further provides steps for determining the target vehicle in the target detection box according to the target feature matching value and the preset matching threshold, including: Compare the target feature matching value with the preset matching threshold to obtain a comparison result; If the comparison result is that the target feature matching value is greater than the preset matching threshold, determine that the vehicle in the target detection box is the target vehicle corresponding to the target trajectory; If the comparison result is that the target feature matching value is not greater than the preset matching threshold, determine that the vehicle in the target detection box is an unmatched vehicle.
[0061] Based on this, by comparing the target feature matching value with the preset matching threshold, it is determined whether the current target feature matching value exceeds the matching threshold, so as to determine whether the detection box features and the trajectory features match, that is, whether the vehicle corresponding to the target detection box is the vehicle corresponding to the target trajectory, thereby realizing the identification and tracking of the vehicle.
[0062] Please refer to Figure 2 , Figure 2It is the system architecture diagram of a vehicle tracking method based on BEV features provided by an embodiment of the present application; as Figure 2 shown, the overall architecture includes the data flow relationships among modules such as sensor data acquisition, BEV feature extraction, target detection, feature enhancement matching, feature-assisted filtering, and trajectory management; in particular, the figure emphasizes how BEV feature information is tightly coupled from the feature extraction module to the feature enhancement matching and feature-assisted filtering modules, forming a feature-level deep fusion architecture for detection and tracking.
[0063] Among them, the BEV feature information is tightly coupled to the observation stage of the filter, forming a feature-state collaborative estimation mechanism to improve the accuracy of state estimation.
[0064] Please refer to Figure 3 , Figure 3 It is the method schematic diagram of a vehicle tracking method based on BEV features provided by an embodiment of the present application; as Figure 3 shown, it details the conversion process from point cloud data to BEV feature maps, including steps such as point cloud voxelization, column feature encoding, and multi-scale feature map generation. The figure shows how to generate and retain feature maps of different scales (F1, F2, F3), and these feature maps are not only used for detection but also directly participate in the subsequent tracking process, realizing the basis of the tight coupling design.
[0065] In some embodiments of the present application, by introducing the multi-scale feature maps in the BEV backbone network of the detection network into the target tracking process in a tight coupling manner, point cloud data in a complex environment is collected, and according to the pointpillars network, the point cloud data is divided into data columns and feature encoding is performed to generate the column feature map corresponding to the point cloud data. At the same time, through a 2D convolutional neural network, feature maps of different scales are generated, and the feature information is retained at multiple scales and tightly coupled with the tracking module; the feature representation of the target is extracted from the BEV feature map through the ROI Align or ROI Pooling method, and the feature similarity is calculated for data association to realize the tight coupling of the feature and the matching process. Based on the comprehensive matching scores of feature similarity, IoU, and position distance, the Hungarian algorithm can be used for optimal matching to realize the deep fusion of feature information and geometric information; in addition, feature similarity information is introduced in the observation stage of the Kalman filter to dynamically adjust the observation noise covariance matrix, realizing a tight coupling adaptive mechanism of the feature and state estimation.
[0066] Thus, by introducing the BEV feature information in the detection network into the matching and filtering processes of vehicle tracking in a tightly coupled manner, the traditional loose-coupling series mode of detection and tracking is broken, and a tracking framework with feature-level deep fusion is constructed. This tightly coupled design enables the tracking module to directly access and utilize the rich intermediate representations in the detection process, rather than relying solely on the final detection results; an object matching algorithm based on BEV feature similarity is developed, which tightly combines feature similarity with traditional position and size information, improving the accuracy of matching. Especially in occluded and dense scenarios, not only the geometric attributes of the object are considered, but also the semantic information contained in the depth features is fully utilized, significantly enhancing the matching performance in difficult scenarios; the feature similarity is introduced into the observation model of the Kalman filter, and by dynamically adjusting the observation noise, an adaptive evaluation of the observation reliability is achieved, improving the accuracy of state estimation. This tight coupling mechanism between features and filters enables the system to intelligently handle various complex scenarios, especially when the object is partially or completely occluded; in the case where the object is temporarily occluded or the detection fails, the previously saved BEV feature information is used to assist in resuming tracking, reducing the ID switching and object loss problems. Through the temporal transfer and update mechanism of features, the system can still accurately identify and associate objects after a long-term occlusion, maintaining the continuity of tracking.
[0067] Figure 4 FIG. 4 is a schematic structural diagram of a vehicle tracking device 400 based on BEV features provided by an embodiment of the present application. As Figure 4 shown, the vehicle tracking device 400 based on BEV features includes: a feature map determination module 410, a trajectory determination module 420, a feature determination module 430, a matching value determination module 440, and a vehicle determination module 450; wherein: The feature map determination module 410 is configured to determine target point cloud data according to a collection device, and determine a pillar feature map corresponding to the target point cloud data according to the pointpillars network; The trajectory determination module 420 is configured to determine target trajectories corresponding to all vehicles in the pillar feature map according to a target detection algorithm; The feature determination module 430 is configured to determine a target detection box and a target feature region in the pillar feature map corresponding to the target detection box, and determine a trajectory feature corresponding to the target trajectory and a detection box feature corresponding to the target detection box according to the target feature region; The matching value determination module 440 is configured to determine a target feature matching value according to the detection box feature, all trajectory features, and a preset weight; The vehicle determination module 450 is configured to determine a target vehicle in the target detection box according to the target feature matching value and a preset matching threshold.
[0068] In the embodiment of the present application, the feature map determination module 410 may further be specifically configured to: Divide the target point cloud data into cylinders with a preset cylinder size, and perform feature encoding on the data in the cylinders according to the pointpillars network to obtain corresponding cylinder data features; Determine the cylinder feature map according to the spatial positions of the cylinders corresponding to the cylinder data features.
[0069] In the embodiment of the present application, the feature map determination module 410 may further be specifically configured to: Determine a preset convolutional neural network; Determine a high-resolution cylinder feature map, a medium-resolution cylinder feature map, and a low-resolution cylinder feature map according to the preset convolutional neural network and the spatial positions of the cylinders; Determine the cylinder feature map according to the high-resolution cylinder feature map, the medium-resolution cylinder feature map, and the low-resolution cylinder feature map.
[0070] In the embodiment of the present application, the trajectory determination module 420 may further be specifically configured to: Determine a convolutional region of a preset size in the cylinder feature map; Determine the vehicles in the convolutional region and the target trajectories corresponding to the vehicles according to the target detection algorithm.
[0071] In the embodiment of the present application, the feature determination module 430 may further be specifically configured to: Determine the high-resolution feature map, the medium-resolution feature map, and the low-resolution feature map in the cylinder feature map, and the corresponding feature regions respectively; Perform feature map fusion on the feature regions corresponding to the high-resolution feature map, the feature regions corresponding to the medium-resolution feature map, and the feature regions corresponding to the low-resolution feature map to obtain a target feature region; Determine the trajectory features corresponding to the target trajectories and the detection box features corresponding to the target detection boxes according to the target feature region.
[0072] In the embodiment of the present application, the matching value determination module 440 may further be specifically configured to: Determine the cosine similarity, intersection over union, and distance metric between the detection box features and each trajectory feature, as well as the preset weights corresponding to the cosine similarity, the preset weights corresponding to the intersection over union, and the preset weights corresponding to the distance metric; Perform weighted combination on the cosine similarity, intersection over union, and distance metric according to the preset weights to obtain a target feature matching value.
[0073] In the embodiment of the present application, the vehicle determination module 450 may further be specifically configured to: Compare the target feature matching value with a preset matching threshold to obtain a comparison result; If the comparison result shows that the target feature matching value is greater than the preset matching threshold, determine that the vehicle in the target detection box is the target vehicle corresponding to the target trajectory; If the comparison result shows that the target feature matching value is not greater than the preset matching threshold, determine that the vehicle in the target detection box is an unmatched vehicle.
[0074] Figure 5 It is a schematic structural diagram of the device provided by the embodiments of the present application. As Figure 5 shown, the device 500 includes: The device 500 may include a processor 501 with one or more processing cores, a memory 502 with one or more computer-readable storage media, a communication component 503, and other components. Among them, the processor 501, the memory 502, and the communication component 503 are connected through a bus 504.
[0075] In a specific implementation process, at least one processor 501 executes the computer execution instructions stored in the memory 502, so that at least one processor 501 executes the above vehicle tracking method based on BEV features.
[0076] For the specific implementation process of the processor 501, reference can be made to the above method embodiments, and their implementation principles and technical effects are similar, so they will not be elaborated here in this embodiment.
[0077] Furthermore, the processor may be a central processing unit (English: Central Processing Unit, abbreviated as: CPU), or may also be other general-purpose processors, digital signal processors (English: Digital Signal Processor, abbreviated as: DSP), application-specific integrated circuits (English: Application Specific Integrated Circuit, abbreviated as: ASIC), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the present application can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules in the processor.
[0078] The memory may include high-speed memory (Random Access Memory, RAM), and may also include non-volatile memory (Non-volatile Memory, NVM), such as at least one disk memory.
[0079] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the buses in the drawings of this application are not limited to only one bus or one type of bus.
[0080] In some embodiments, a computer program product is also provided, including a computer program or instructions, which, when executed by a processor, implement the steps in any of the above vehicle tracking methods based on BEV features.
[0081] For the specific implementation of each of the above operations, reference can be made to the previous embodiments and will not be elaborated here.
[0082] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by controlling relevant hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0083] Therefore, an embodiment of this application provides a computer-readable storage medium, in which multiple program codes are stored. The program codes can be loaded by a processor to execute the steps in any of the vehicle tracking methods based on BEV features provided by the embodiments of this application.
[0084] Among them, the storage medium can include: Read Only Memory (ROM), Random Access Memory (RAM), a magnetic disk, an optical disc, or the like.
[0085] According to one aspect of this application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium.
[0086] Since the instructions stored in the storage medium can execute the steps in any of the vehicle tracking methods based on BEV features provided by the embodiments of this application, the beneficial effects that can be achieved by any of the vehicle tracking methods based on BEV features provided by the embodiments of this application can be realized. For details, reference can be made to the previous embodiments and will not be elaborated here.
[0087] Other embodiments of the present application will be readily apparent to those skilled in the art in view of the specification and practice of the invention disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and examples are only illustrative, and the true scope and spirit of the present application are pointed out by the claims above.
[0088] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A vehicle tracking method based on BEV feature information, characterized in that, The method includes: Determine target point cloud data according to a collection device, and determine a pillar feature map corresponding to the target point cloud data according to a PointPillars network; Determine target trajectories corresponding to all vehicles in the pillar feature map according to a target detection algorithm; Determine a target detection box and a target feature region in the pillar feature map corresponding to the target detection box, and determine a trajectory feature corresponding to the target trajectory and a detection box feature corresponding to the target detection box according to the target feature region; Determine a target feature matching value according to the detection box feature, all the trajectory features, and a preset weight; Determine a target vehicle in the target detection box according to the target feature matching value and a preset matching threshold.
2. The method according to claim 1, wherein The step of determining target point cloud data according to a collection device and determining a pillar feature map corresponding to the target point cloud data according to a PointPillars network includes: Divide the target point cloud data into pillars with a preset pillar size, and perform feature encoding on the data in the pillars according to the PointPillars network to obtain corresponding pillar data features; Determine the pillar feature map according to the spatial positions of the pillars corresponding to the pillar data features.
3. The method according to claim 2, wherein The step of determining the pillar feature map according to the spatial positions of the pillars corresponding to the pillar data features includes: Determine a preset convolutional neural network; Determine a high-resolution pillar feature map, a medium-resolution pillar feature map, and a low-resolution pillar feature map according to the preset convolutional neural network and the spatial positions of the pillars; Determine the pillar feature map according to the high-resolution pillar feature map, the medium-resolution pillar feature map, and the low-resolution pillar feature map.
4. The method according to claim 1, wherein The step of determining target trajectories corresponding to all vehicles in the pillar feature map according to a target detection algorithm includes: Determine a convolutional region with a preset size in the pillar feature map; Determine vehicles in the convolutional region and the target trajectories corresponding to the vehicles according to the target detection algorithm.
5. The method according to claim 1, wherein The step of determining a target detection box and a target feature region in the pillar feature map corresponding to the target detection box, and determining a trajectory feature corresponding to the target trajectory and a detection box feature corresponding to the target detection box according to the target feature region includes: Determine a high-resolution feature map, a medium-resolution feature map, and a low-resolution feature map in the pillar feature map, and corresponding feature regions respectively; Perform feature map fusion on the feature regions corresponding to the high-resolution feature map, the feature regions corresponding to the medium-resolution feature map, and the feature regions corresponding to the low-resolution feature map to obtain the target feature region; Determine a trajectory feature corresponding to the target trajectory and a detection box feature corresponding to the target detection box according to the target feature region.
6. The method according to claim 1, wherein The step of determining a target feature matching value according to the detection box feature, all the trajectory features, and a preset weight includes: Determine the cosine similarity, intersection over union (IoU), and distance metric between the detected bounding box features and each of the trajectory features, as well as the preset weights corresponding to the cosine similarity, the preset weights corresponding to the IoU, and the preset weights corresponding to the distance metric; According to the preset weights, perform weighted combination on the cosine similarity, IoU, and distance metric to obtain the target feature matching value.
7. The method according to claim 1, characterized in that The determining the target vehicle in the target detection box according to the target feature matching value and the preset matching threshold includes: Compare the target feature matching value with the preset matching threshold to obtain a comparison result; If the comparison result is that the target feature matching value is greater than the preset matching threshold, determine that the vehicle in the target detection box is the target vehicle corresponding to the target trajectory; If the comparison result is that the target feature matching value is not greater than the preset matching threshold, determine that the vehicle in the target detection box is an unmatched vehicle.
8. A vehicle tracking device based on BEV feature information, characterized in that, The device includes: A feature map determination module, configured to determine target point cloud data according to a collection device, and determine a pillar feature map corresponding to the target point cloud data according to the pointpillars network; A trajectory determination module, configured to determine target trajectories corresponding to all vehicles in the pillar feature map according to a target detection algorithm; A feature determination module, configured to determine a target detection box and a target feature region corresponding to the target detection box in the pillar feature map, and determine the trajectory features corresponding to the target trajectory and the detected bounding box features corresponding to the target detection box according to the target feature region; A matching value determination module, configured to determine a target feature matching value according to the detected bounding box features, all the trajectory features, and preset weights; A vehicle determination module, configured to determine the target vehicle in the target detection box according to the target feature matching value and the preset matching threshold.
9. A device, characterized in that, Includes: One or more processors; A memory; One or more programs, where one or more programs are stored in the memory and configured to be executed by one or more processors, and one or more programs are configured to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, Program code is stored in a computer-readable storage medium, and the program code can be called by a processor to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Multi-target tracking method in unmanned driving scene based on deep learning
CN113468950A
Three-dimensional multi-target tracking method fusing point cloud and image information
CN116363171A
Road vehicle tracking method based on lightweight re-recognition network
CN118154642A
Multi-target tracking method and system based on improved Pointpillars network
CN118470061A
Vehicle sensing information acquisition method and device, equipment and storage medium
CN118537834A