Target detection and tracking method and device in automatic driving scene, equipment and medium

By constructing the target dataset and combining rotating target alignment and two-stage matching methods, the detection and tracking challenges of 3D multi-objective tracking in complex environments are solved, and efficient and accurate target recognition and tracking are achieved.

CN120374956APending Publication Date: 2025-07-25SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510519013.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing 3D multi-objective tracking method faces the problems of object rotation, redundant calculation and insufficient sensor adaptability when dealing with large-scale and diverse data sets, especially in complex environments.

Method used

By obtaining historical driving environment information, marking and building target data sets, and using pre-trained models for training, combining rotating target alignment technology and two-stage matching method, efficient tracking of straight and non-linear motion targets is achieved.

Benefits of technology

It improves the target detection accuracy and tracking accuracy, reduces detection errors caused by object rotation, reduces calculation complexity, and improves system efficiency and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374956A_ABST
    Figure CN120374956A_ABST
Patent Text Reader

Abstract

The invention discloses a target detection and tracking method and device in an automatic driving scene, equipment and a medium, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining historical driving environment information, and marking the historical driving environment information to obtain a target data set; training a pre-trained preset detection model by using the target data set to obtain a preset target detection model; according to a preset target detection model, the environment in the automatic driving scene is detected in real time to identify a target object, the target object in linear motion is directly tracked based on a preset target tracking model, and the target object in non-linear motion in the 3D scene is aligned by using the preset target tracking model. And projecting the aligned target object to a preset 2D plane so as to continuously track the target object by calculating the intersection of the target object on the preset 2D plane and comparing the intersection with the target object. According to the invention, the efficient, accurate and automatic multi-target detection and tracking method is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to a method, device, equipment and medium for object detection and tracking in an autonomous driving scenario. Background Technique

[0002] In the context of the rapid development of autonomous driving and robotics, 3D multi-object tracking (MOT, Multiple Object Tracking), as a key link connecting perception and planning tasks, has become increasingly important. However, current 3D MOT solutions face significant challenges when dealing with large-scale and diverse data sets, especially in complex environments such as nuScenes, Waymo, and KITTI. On the one hand, existing filter-based methods such as Fast-Poly (a 3D multi-object tracking method) still have limitations in dealing with object rotation, redundant calculations in the global perspective, and large-scale trajectory management; on the other hand, the adaptability of MCTrack (a 3D multi-object tracking method) to different sensor inputs and environmental conditions still needs to be strengthened. Summary of the Invention

[0003] In view of this, the purpose of this application is to provide a method, device, equipment and medium for object detection and tracking in an autonomous driving scenario, which can achieve more efficient, accurate and automated multi-object tracking to meet the requirements of actual application scenarios. The specific solutions are as follows:

[0004] In the first aspect, this application provides a method for object detection and tracking in an autonomous driving scenario, including:

[0005] Obtain historical driving environment information based on a preset sensor, and annotate the historical driving environment information to obtain a target data set;

[0006] Use the target data set to train a preset detection model that has been pre-trained to obtain a preset object detection model;

[0007] According to the preset object detection model, perform real-time detection on the environment in the autonomous driving scenario to identify target objects, directly track the target objects moving in a straight line based on a preset object tracking model, and use the preset object tracking model to align the target objects moving non-linearly in the 3D scene, and project the aligned target objects onto a preset 2D plane to continuously track the target objects by calculating the intersection over union of the target objects located on the preset 2D plane.

[0008] Optionally, the obtaining historical driving environment information based on a preset sensor, and annotating the historical driving environment information to obtain a target data set includes:

[0009] Obtain the three-dimensional point cloud information of the surrounding driving environment obtained based on the lidar as historical driving environment information;

[0010] and / or, obtain the visual image information of the surrounding driving environment obtained based on the camera as historical driving environment information;

[0011] Perform category annotation and status information annotation on each target object in the historical driving environment information; the status information includes the position information, size information, direction information, and appearance information of each target object.

[0012] Optionally, the training of the pre-trained preset detection model using the target dataset to obtain a preset target detection model includes:

[0013] Train the pre-trained preset detection model using the target dataset, and adjust the hyperparameters of the preset detection model until the performance index of the preset detection model reaches the preset target performance index to obtain a preset target detection model; the hyperparameters include the learning rate, batch size, and loss function;

[0014] and / or, perform transformation operations on the historical driving environment information in the target dataset based on data augmentation techniques to obtain a new target dataset, train the pre-trained preset detection model using the new target dataset, and adjust the hyperparameters of the preset detection model until the performance index of the preset detection model reaches the preset target performance index to obtain a preset target detection model; the transformation operations include information cropping, information flipping, and brightness adjustment.

[0015] Optionally, the alignment of the target objects moving non-linearly in the 3D scene and the projection of the aligned target objects onto a preset 2D plane to continuously track the target objects by calculating the intersection over union of the target objects located on the preset 2D plane includes:

[0016] Monitor the target objects based on the preset target tracking model to obtain target objects moving non-linearly; the target objects moving non-linearly include target objects traveling along a curved road and / or target objects currently performing a flipping motion;

[0017] Perform a preset pose adjustment operation on the target objects moving non-linearly to adjust the pose of the target objects moving non-linearly in the 3D scene at each moment to a preset standard pose to complete the alignment operation of the target objects moving non-linearly at each moment; the preset pose adjustment operation includes rotation and translation;

[0018] Project the aligned target object of the non-linear motion onto a preset 2D plane to obtain 2D bounding boxes corresponding to the target object of the non-linear motion at each moment;

[0019] Determine the intersection over union (IoU) between the 2D bounding box corresponding to the target object of the non-linear motion at the current moment and the 2D bounding box corresponding to the target object of the non-linear motion at the previous moment. Determine the target IoU that exceeds a preset threshold from the IoUs, and identify the target object of the non-linear motion at the current moment and the target object of the non-linear motion at the previous moment corresponding to the target IoU as the same target object at different moments, so as to continuously track the target object.

[0020] Optionally, the target detection and tracking method in the autonomous driving scenario further includes:

[0021] Obtain the state information of each object in the surrounding environment of the autonomous driving vehicle. Based on a preset target tracking model and according to a preset distance range and a preset size range, determine the expected target objects whose distances from the autonomous driving vehicle are within the preset distance range and whose sizes do not exceed the preset size range from the objects corresponding to the state information;

[0022] Match the trajectory information corresponding to the expected target object with the trajectory information of the autonomous driving vehicle to obtain a matching result. Determine the first expected target object whose matching result reaches a preset matching threshold, and track the first expected target object; the trajectory information includes driving direction information and speed change information;

[0023] Determine the second expected target object whose matching result does not reach the preset matching threshold, and perform appearance matching between the appearance information corresponding to the second expected target object and the target objects tracked historically. If the matching is successful, track the second expected target object.

[0024] Optionally, the target detection and tracking method in the autonomous driving scenario further includes:

[0025] Obtain the tracking result obtained by tracking the target object using the preset target tracking model;

[0026] Convert the tracking result into preset structured data, and save the preset structured data;

[0027] Wherein, the preset structured data includes the size information, position coordinate information, rotation direction information, category information, confidence level, and object identifier information uniquely corresponding to the target object of the target object.

[0028] Optionally, the target detection and tracking method in the autonomous driving scenario further includes:

[0029] Determine the model performance metrics corresponding to the preset target tracking model based on the tracking results; the model performance metrics include multi-object tracking average precision and the number of identifier switches; the number of identifier switches is the number of occurrences of object identifier switching events caused by errors in depth information;

[0030] If the model performance metrics do not reach the preset optimal metrics, determine the problems corresponding to the model based on the model performance metrics, and adjust the parameters of the model for the problems to update the preset target tracking model.

[0031] In a second aspect, the present application provides an object detection and tracking device in an autonomous driving scenario, including:

[0032] A data acquisition module, configured to acquire historical driving environment information based on a preset sensor, and annotate the historical driving environment information to obtain a target data set;

[0033] A model training module, configured to train a preset detection model that has been pre-trained using the target data set to obtain a preset target detection model;

[0034] An object recognition and tracking module, configured to perform real-time detection on the environment in the autonomous driving scenario according to the preset target detection model to identify target objects, directly track the target objects moving in a straight line based on a preset target tracking model, and use the preset target tracking model to align the target objects moving non-linearly in a 3D scene, and project the aligned target objects onto a preset 2D plane, so as to continuously track the target objects by calculating the intersection over union of the target objects located on the preset 2D plane.

[0035] In a third aspect, the present application provides an electronic device, including:

[0036] A memory, configured to store a computer program;

[0037] A processor, configured to execute the computer program to implement the foregoing object detection and tracking method in an autonomous driving scenario.

[0038] In a fourth aspect, the present application provides a computer-readable storage medium, configured to store a computer program, wherein the computer program, when executed by a processor, implements the foregoing object detection and tracking method in an autonomous driving scenario.

[0039] In this application, historical driving environment information obtained based on a preset sensor is acquired, and the historical driving environment information is labeled to obtain a target dataset; the preset detection model that has been pre-trained is trained using the target dataset to obtain a preset target detection model; the environment in the autonomous driving scenario is detected in real time according to the preset target detection model to identify target objects, the target objects moving in a straight line are directly tracked based on a preset target tracking model, and the preset target tracking model is used to align the target objects moving non-linearly in the 3D scene, and project the aligned target objects onto a preset 2D plane, so as to continuously track the target objects by calculating the intersection over union of the target objects located on the preset 2D plane. As can be seen from the above, in this application, by labeling the historical driving environment information to obtain a target dataset and using the target dataset to train the preset detection model that has been pre-trained, the trained preset target detection model can better adapt to the complex environment in the autonomous driving scenario and improve the detection accuracy of target objects; when using the preset target tracking model to track target objects moving non-linearly, first, the target objects moving non-linearly in the 3D scene are aligned, and the aligned target objects are projected onto a preset 2D plane, so as to continuously track the target objects by calculating the intersection over union of the target objects located on the preset 2D plane. In this way, target objects in different states can be identified more accurately, the detection errors caused by factors such as object rotation can be reduced, and the target objects can be tracked more precisely, improving the accuracy and robustness of target tracking; at the same time, the calculation in the 3D scene is converted into the calculation in the 2D plane, avoiding redundant calculations and improving the efficiency. Description of the Drawings

[0040] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0041] Figure 1 It is a flowchart of a method for target detection and tracking in an autonomous driving scenario disclosed in the present application;

[0042] Figure 2 It is a schematic diagram of a specific method for target detection and tracking in an autonomous driving scenario disclosed in the present application;

[0043] Figure 3 It is a schematic diagram of the structure of a device for target detection and tracking in an autonomous driving scenario disclosed in the present application;

[0044] Figure 4Schematic diagram of the structure of an electronic device disclosed in this application. Detailed implementation manners

[0045] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the protection scope of this application.

[0046] In the context of the rapid development of autonomous driving and robotics, 3D multi-object tracking, as a key link connecting perception and planning tasks, has become increasingly important. However, current 3D MOT solutions face significant challenges when dealing with large-scale and diverse datasets, especially in complex environments such as nuScenes, Waymo, and KITTI. On the one hand, existing filter-based methods such as Fast-Poly still have limitations in dealing with object rotation, redundant calculations in the global perspective, and large-scale trajectory management; on the other hand, the adaptability of MCTrack to different sensor inputs and environmental conditions still needs to be strengthened. For this reason, this application provides a method for object detection and tracking in an autonomous driving scenario, which can achieve more efficient, accurate, and automated multi-object tracking to meet the requirements of actual application scenarios.

[0047] See Figure 1 As shown, the embodiments of this application disclose a method for object detection and tracking in an autonomous driving scenario, including:

[0048] Step S11: Obtain historical driving environment information based on a preset sensor, and label the historical driving environment information to obtain a target dataset.

[0049] In a specific implementation manner, three-dimensional point cloud information of the surrounding driving environment based on lidar can be obtained as the historical driving environment information.

[0050] In another specific implementation manner, visual image information of the surrounding driving environment based on a camera can be obtained as the historical driving environment information.

[0051] After obtaining the historical driving environment information, each target object in the historical driving environment information can be labeled with a category and status information; where the status information includes, but is not limited to, the position information, size information, direction information, and appearance information of each target object.

[0052] It can be understood that in this embodiment, a large-scale, high-quality, and multi-modal autonomous driving dataset containing rich sensor information (such as point cloud information obtained by lidar and image information obtained by cameras) is first collected and constructed. Then, each data sample in the autonomous driving dataset can be manually and finely annotated to ensure the accuracy of the annotated labels. The target dataset after annotation not only provides a solid data foundation for subsequent model training but also covers information under various weather conditions, time periods, and various road scenarios to ensure that the model has broad generalization capabilities. For example, the target dataset can include various categories such as pedestrians, motor vehicles, and non-motor vehicles, and details such as the position information, size information, direction information, and appearance information of each target object are recorded. The target dataset covers the real situations in various driving environments, thus providing comprehensive and detailed data support for model training.

[0053] Step S12: Use the target dataset to train a pre-trained preset detection model to obtain a preset object detection model.

[0054] In this embodiment, the target dataset obtained in step S11 can be used to train a pre-trained preset detection model to obtain a preset object detection model.

[0055] In a specific implementation manner, the target dataset can be used to train a pre-trained preset detection model, and the hyperparameters of the preset detection model can be adjusted until the performance metrics of the preset detection model reach the preset target performance metrics to obtain a preset object detection model; where the hyperparameters include but are not limited to the learning rate, batch size, and loss function.

[0056] It can be understood that existing large-scale pre-trained models can be used as initialization parameters. These large-scale pre-trained models have usually been trained on a large amount of data and have learned rich general features and patterns, such as the edges and textures of objects. Starting from these models and using the target dataset for fine-tuning avoids the large amount of sample data required for training a neural network from scratch, reduces the model training time, and improves the model's understanding ability and detection accuracy for specific autonomous driving scenarios; at the same time, it can also avoid problems such as overfitting that may occur when training a model from scratch. During the fine-tuning process, the hyperparameters such as the learning rate, batch size, and loss function of the model can be adjusted to ensure that the model reaches the best performance on the target dataset, and finally the trained object detection model is obtained.

[0057] In another specific implementation, the historical driving environment information in the target dataset can be transformed based on data augmentation techniques to obtain a new target dataset. The pre-trained preset detection model is trained using the new target dataset, and the hyperparameters of the preset detection model are adjusted until the performance metrics of the preset detection model reach the preset target performance metrics, so as to obtain the preset target detection model. Among them, the transformation operations include information cropping, information flipping, and brightness adjustment.

[0058] To further enhance the generalization ability of the model, data augmentation techniques such as cropping, flipping, and brightness adjustment can be introduced to transform the original target dataset, generate more diverse training samples, and simulate more diverse driving environments. For example, cropping can change the position and size of the target object in the image, flipping can change the direction of the target object, and brightness adjustment can simulate driving scenarios under different lighting conditions. By using data augmentation techniques to obtain a new target dataset and training the model using the new target dataset to obtain the trained target detection model, the target detection model learns richer features and patterns, and thus can maintain high accuracy and reliability in complex and changing real environments.

[0059] Through this deeply customized fine-tuning process, the model can not only accurately identify static objects, but also accurately predict the motion state of the target object according to the detection information of the target object at different times, combined with the motion patterns and rules learned by itself, providing more sufficient information for the decision-making and planning of autonomous driving, thus greatly improving the overall performance of the system.

[0060] Step S13: Based on the preset target detection model, the environment in the autonomous driving scenario is detected in real time to identify the target object. The target object moving in a straight line is directly tracked based on the preset target tracking model, and the target object moving non-linearly in the 3D scene is aligned using the preset target tracking model, and the aligned target object is projected onto a preset 2D plane, so as to continuously track the target object by calculating the intersection over union of the target object on the preset 2D plane.

[0061] In this embodiment, first, the trained preset target detection model in step S12 can be used to analyze the environmental data obtained by various sensors (such as cameras, lidar) in the autonomous driving scenario in real time and identify the target objects in the scenario, such as pedestrians, vehicles, traffic signs, etc. This is the basis for subsequent tracking operations. Only by accurately detecting the target object can it be tracked.

[0062] Then, the preset target tracking model can be used to track the target object. For a target object moving in a straight line, since its motion law is relatively simple and predictable, the preset target tracking model can directly use the position information of the target object at different times and apply methods such as linear prediction to continuously track it. For example, if a car is detected moving at a constant speed on a straight road ahead, the preset target tracking model can accurately predict the position of the car at the next moment based on the current speed and direction of the car, thus achieving stable tracking.

[0063] For a target object moving in a non - straight line, its motion trajectory is relatively complex and difficult to directly track. Therefore, first, the preset target tracking model can be used to monitor the target object to obtain the target object moving in a non - straight line; among them, the target object moving in a non - straight line includes a target object moving along a curved road and / or a target object currently performing a flipping motion; then, a preset pose adjustment operation is performed on the target object moving in a non - straight line to adjust the pose of the target object moving in a non - straight line at each moment in the 3D scene to a preset standard pose, so as to complete the alignment operation of the target object moving in a non - straight line at each moment; among them, the preset pose adjustment operation includes rotation and translation; further, the aligned target object moving in a non - straight line is projected onto a preset 2D plane to obtain 2D bounding boxes corresponding to the target object moving in a non - straight line at each moment; the intersection - over - union ratio between the 2D bounding box corresponding to the target object moving in a non - straight line at the current moment and the 2D bounding box corresponding to the target object moving in a non - straight line at the previous moment is determined, the target intersection - over - union ratio exceeding a preset threshold is determined from the intersection - over - union ratios, and the target object moving in a non - straight line at the current moment and the target object moving in a non - straight line at the previous moment corresponding to the target intersection - over - union ratio are determined as the same target object at different times, so as to continuously track the target object.

[0064] It can be understood that for a target object moving in a non - straight line, first, the target object can be adjusted to a preset standard pose. For example, the front direction of the car can be made parallel to a fixed reference direction (such as the due - east direction). Through pose adjustment and alignment, the poses of the target object at different times in the 3D scene are comparable, which is convenient for subsequent projection onto the 2D plane for processing. Then, the target object with the adjusted pose is transformed from the similarity calculation in the originally complex 3D space to a similarity calculation similar to that in a 2D plane, similar to the BEV (Bird's - Eye View) perspective, and then the intersection - over - union ratio between the bounding boxes of the target object on the 2D plane is calculated. Calculating the intersection - over - union ratio on the 2D plane is relatively simple, and the bounding boxes of multiple target objects can be calculated in parallel at the same time, which speeds up the calculation speed. In this way, through the rotation - based target alignment technology, the complex calculation that originally needed to consider multiple factors such as the three - dimensional coordinates and rotation angles of the target object in the 3D space is transformed into the calculation of the intersection - over - union ratio on the 2D plane. After that, only the overlapping situation of the bounding boxes needs to be concerned, and the calculation complexity is greatly reduced.

[0065] Furthermore, the target object can also be tracked by a two-stage matching method. Specifically, first, obtain the state information of each object in the surrounding environment of the autonomous vehicle. Based on a preset target tracking model and according to a preset distance range and a preset size range, determine, from the objects corresponding to the state information, the expected target objects whose distances from the autonomous vehicle are within the preset distance range and whose sizes do not exceed the preset size range. Then, match the trajectory information corresponding to the expected target objects with the trajectory information of the autonomous vehicle to obtain a matching result, determine the first expected target object whose matching result reaches the preset matching threshold, and track the first expected target object; wherein, the trajectory information includes driving direction information and speed change information; determine the second expected target object whose matching result does not reach the preset matching threshold, and perform appearance matching between the appearance information corresponding to the second expected target object and the target object tracked historically. If the matching is successful, track the second expected target object.

[0066] It can be understood that in the first-stage matching, the position information of the target object on the bird's-eye view (BEV) plane can be utilized to generate a voxel mask by calculating the low-cost Euclidean distance and combining manually set target size parameters (such as the length and width of a vehicle, etc.). The voxel mask can be regarded as a spatial filter that delimits a range on the bird's-eye view plane, excluding distant targets and only retaining potential matching objects within a certain distance from the autonomous vehicle itself. In an actual traffic scenario, there are numerous target objects. Through the voxel mask, distant target objects can be quickly filtered out, focusing on potentially relevant target objects, reducing the computational amount, improving the matching efficiency, and enabling the system to process data more efficiently.

[0067] For example, suppose there are multiple vehicles such as vehicle A, B, C, D, E, etc. Calculate the Euclidean distance based on the positions of each vehicle and generate a voxel mask in combination with the average size parameters of the vehicles. Taking the autonomous vehicle (AV) as the center, set a voxel area with a radius of 15 meters. Vehicle A is relatively close to the AV and is within the voxel mask range; vehicle B is relatively far from the AV and exceeds the voxel mask range; vehicles C and D are also within the voxel mask range, while vehicle E, although relatively close, is judged to be outside the preset size range by the size parameters and is excluded from the key matching objects. For vehicles A, C, and D within the voxel mask range, perform a preliminary matching analysis based on their trajectory information on the bird's-eye view. It is found that vehicle A has the same driving direction as the autonomous vehicle and similar speeds, vehicle C has a certain angle with the driving direction of the AV and a slower speed, and vehicle D has the same driving direction as the AV but a faster speed. Through this stage of matching, it is preliminarily determined that vehicle A is a potential matching object with a relatively high correlation with the AV, the correlations of vehicles C and D are relatively weak, and at the same time, vehicles B and E and other distant or irrelevant vehicles are excluded, reducing the computational amount.

[0068] For the trajectories that fail to be successfully matched in the first stage, in the second-stage matching, they can be projected onto the image plane. On the image plane, secondary matching can be performed between the appearance information such as the color and texture of the objects that have not been successfully matched and the target objects tracked historically. This is because the first-stage matching is mainly based on position information for matching, which has certain limitations, while the appearance information on the image plane can provide more dimensional information to help further confirm the identity and effectively solve problems such as target identity switching (ID-Switch, misidentifying one target as another target) and tracking fragmentation (Fragmentation, discontinuity of the target during the tracking process) caused by inaccurate depth information (such as lidar measurement errors). Through the second-stage matching, the accuracy of the matching can be improved, ensuring the continuity and stability of the tracking results.

[0069] For example, at an intersection, vehicle A and vehicle C are both traveling in the same direction. From the bird's-eye view plane, their trajectories are relatively similar. In the first stage, vehicle A is identified as an object with a relatively high correlation with the trajectory of the autonomous vehicle, while vehicle C is not successfully matched due to some minor differences. In fact, vehicle C may be an important target vehicle that was detected on another section before, but at the current moment, its trajectory partially overlaps with that of vehicle A. If the second-stage matching is not performed, vehicle C may be ignored, resulting in the interruption of the tracking of this vehicle and the inability to fully grasp its driving state, which may affect the judgment and decision-making of the autonomous vehicle on the overall traffic environment.

[0070] By introducing an efficient rotating target alignment technology and a two-stage matching strategy, the computational resource consumption of the system in dealing with complex scenarios is reduced, and the overall efficiency and economy are further improved.

[0071] In this embodiment, the model can be updated based on the tracking results obtained by the model. Specifically, first, obtain the tracking results obtained by using a preset target tracking model to track the target object; then convert the tracking results into preset structured data and save the preset structured data; where the preset structured data includes but is not limited to the size information, position coordinate information, rotation direction information, category information, confidence level, and object identifier information uniquely corresponding to the target object of the target object.

[0072] Further, determine the model performance metrics corresponding to the preset target tracking model based on the tracking results; wherein, the model performance metrics include, but are not limited to, the average precision of multi-object tracking and the number of identifier switches; the number of identifier switches is the number of occurrences of object identifier switching events caused by errors in depth information; if the model performance metrics do not reach the preset optimal metrics, determine the problems corresponding to the model based on the model performance metrics, and adjust the parameters of the model for the problems to update the preset target tracking model.

[0073] It can be understood that the tracking results can be converted into standardized structured data, such as JSON format (JavaScript Object Notation). Each tracking result record contains in detail key information such as the size information, position coordinate information, rotation direction information, category information, confidence level, and object identifier information uniquely corresponding to the target object of the target object. Save the tracking results to improve the data reusability and interoperability. Then, the model performance metrics corresponding to the preset target tracking model can be determined based on the tracking results. The performance of the current model relative to the advanced model benchmark is measured by a series of performance metrics including the average precision of multi-object tracking (AMOTA, Average Multiple Object Tracking Precision), the number of identifier switches, etc., and it is judged whether the model performance metrics reach the preset optimal metrics. Based on objective data analysis, the problems existing in the model can be quickly located, and the parameters of the model or the improvement method can be adjusted accordingly. By constructing such a closed-loop feedback mechanism, the model can continuously learn the characteristics of multi-batch data and adapt to new challenges, so as to maintain its high-efficiency operation ability in complex environments.

[0074] As can be seen from the above, and referring to Figure 2 As shown, in this embodiment, first, obtain the historical driving environment information based on the preset sensor to construct a multi-modal autonomous driving data set and obtain the target data set; then perform the fine-tuning operation of the pre-trained industry large model to train the pre-trained preset detection model with the target data set to obtain the preset target detection model; further, use the preset target tracking model to track the target object, wherein, introduce the rotation target alignment technology to track the target object moving in a non-linear motion, and improve the matching accuracy based on the two-stage matching method to ensure the continuity and stability of the tracking results; finally, perform performance evaluation and model iteration on the model. In this way, in this embodiment, the pre-trained preset detection model is trained with the target data set, so that the trained preset target detection model can better adapt to the complex environment in the autonomous driving scenario and improve the detection accuracy of the target object; at the same time, it can more accurately identify and track the target objects in different states, reduce the detection errors caused by factors such as object rotation, and improve the accuracy and robustness of target tracking.

[0075] See Figure 3 As shown, the embodiment of the present application also discloses a target detection and tracking device in an autonomous driving scenario, including:

[0076] A data acquisition module 11, configured to acquire historical driving environment information based on a preset sensor, and label the historical driving environment information to obtain a target data set;

[0077] A model training module 12, configured to train a pre-trained preset detection model using the target data set to obtain a preset target detection model;

[0078] An object recognition and tracking module 13, configured to perform real-time detection on the environment in the autonomous driving scenario according to the preset target detection model to identify a target object, directly track the target object moving in a straight line based on a preset target tracking model, and use the preset target tracking model to align the target object moving non-linearly in the 3D scene, and project the aligned target object onto a preset 2D plane, so as to continuously track the target object by calculating the intersection over union of the target object located on the preset 2D plane.

[0079] As can be seen from the above, the present application labels the historical driving environment information to obtain a target data set, and uses the target data set to train a pre-trained preset detection model, so that the trained preset target detection model can better adapt to the complex environment in the autonomous driving scenario and improve the detection accuracy of the target object; when using the preset target tracking model to track the target object moving non-linearly, first align the target object moving non-linearly in the 3D scene, and project the aligned target object onto a preset 2D plane, so as to continuously track the target object by calculating the intersection over union of the target object located on the preset 2D plane. In this way, the target object in different states can be identified more accurately, the detection errors caused by factors such as object rotation can be reduced, and the target object can be tracked more precisely, improving the accuracy and robustness of target tracking; at the same time, the calculation in the 3D scene is converted into the calculation in the 2D plane, avoiding redundant calculations and improving the efficiency.

[0080] In some specific embodiments, the data acquisition module 11 includes:

[0081] A first information acquisition unit, configured to acquire three-dimensional point cloud information of the surrounding driving environment based on a lidar as historical driving environment information;

[0082] A second information acquisition unit, configured to acquire visual image information of the surrounding driving environment based on a camera as historical driving environment information;

[0083] An information annotation unit for classifying and annotating status information for each target object in the historical driving environment information; the status information includes the position information, size information, direction information, and appearance information of each target object.

[0084] In some specific embodiments, the model training module 12 includes:

[0085] A first model training unit for training a pre-trained preset detection model using the target data set and adjusting the hyperparameters of the preset detection model until the performance index of the preset detection model reaches a preset target performance index to obtain a preset target detection model; the hyperparameters include the learning rate, batch size, and loss function.

[0086] A second model training unit for performing transformation operations on the historical driving environment information in the target data set based on data augmentation techniques to obtain a new target data set, training a pre-trained preset detection model using the new target data set, and adjusting the hyperparameters of the preset detection model until the performance index of the preset detection model reaches a preset target performance index to obtain a preset target detection model; the transformation operations include information cropping, information flipping, and brightness adjustment.

[0087] In some specific embodiments, the object recognition and tracking module 13 includes:

[0088] An object monitoring unit for monitoring the target object based on the preset target tracking model to obtain a non-linearly moving target object; the non-linearly moving target object includes a target object traveling along a curved road and / or a target object currently performing a flipping motion.

[0089] An object alignment unit for performing a preset pose adjustment operation on the non-linearly moving target object to adjust the pose of the non-linearly moving target object in the 3D scene at each moment to a preset standard pose to complete the alignment operation of the non-linearly moving target object at each moment; the preset pose adjustment operation includes rotation and translation.

[0090] A projection unit for projecting the aligned non-linearly moving target object onto a preset 2D plane to obtain 2D bounding boxes corresponding to the non-linearly moving target object at each moment.

[0091] A first object tracking unit, configured to determine the intersection over union (IoU) between the 2D bounding box corresponding to a target object moving in a non-linear motion at the current moment and the 2D bounding box corresponding to the target object moving in a non-linear motion at the previous moment, determine a target IoU exceeding a preset threshold from the IoUs, and determine the target object moving in a non-linear motion at the current moment and the target object moving in a non-linear motion at the previous moment corresponding to the target IoU as the same target object at different moments, so as to continuously track the target object.

[0092] In some specific embodiments, the target detection and tracking device in the autonomous driving scenario further includes:

[0093] An object determination unit, configured to obtain the state information of each object in the surrounding environment of the autonomous driving vehicle, and determine, based on a preset target tracking model and according to a preset distance range and a preset size range, an expected target object whose distance from the autonomous driving vehicle is within the preset distance range and whose size does not exceed the preset size range from each object corresponding to the state information;

[0094] A second object tracking unit, configured to match the trajectory information corresponding to the expected target object with the trajectory information of the autonomous driving vehicle to obtain a matching result, determine a first expected target object whose matching result reaches a preset matching threshold, and track the first expected target object; the trajectory information includes driving direction information and speed change information;

[0095] A third object tracking unit, configured to determine a second expected target object whose matching result does not reach the preset matching threshold, and perform appearance matching on the appearance information corresponding to the second expected target object and the target object tracked historically. If the matching is successful, track the second expected target object.

[0096] In some specific embodiments, the target detection and tracking device in the autonomous driving scenario further includes:

[0097] A result acquisition unit, configured to acquire a tracking result obtained by tracking a target object using the preset target tracking model;

[0098] A data storage unit, configured to convert the tracking result into preset structured data and store the preset structured data;

[0099] Wherein, the preset structured data includes size information, position coordinate information, rotation direction information, category information, confidence level, and object identifier information uniquely corresponding to the target object of the target object.

[0100] In some specific embodiments, the target detection and tracking device in the autonomous driving scenario further includes:

[0101] An index determination unit, configured to determine a model performance index corresponding to the preset target tracking model based on the tracking result; the model performance index includes multi-object tracking average precision and the number of identifier switches; the number of identifier switches is the number of occurrences of an object identifier switching event caused by an error in depth information;

[0102] A model update unit, configured to, if the model performance index does not reach a preset optimal index, determine a problem corresponding to the model based on the model performance index, and adjust the parameters of the model for the problem to update the preset target tracking model.

[0103] Furthermore, an embodiment of the present application also discloses an electronic device, Figure 4 which is a structural diagram of an electronic device 20 shown according to an exemplary embodiment. The content in the figure should not be considered as any limitation on the scope of use of the present application.

[0104] Figure 4 This is a schematic structural diagram of an electronic device 20 provided by an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the target detection and tracking method in the autonomous driving scenario disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0105] In this embodiment, the power supply 23 is used to provide a working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and no specific limitation is imposed on it here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application needs, and no specific limitation is made here.

[0106] In addition, as a carrier for resource storage, the memory 22 may be a read-only memory, a random access memory, a disk, or an optical disc, etc. The resources stored thereon may include an operating system 221, a computer program 222, etc., and the storage method may be short-term storage or permanent storage.

[0107] Among them, the operating system 221 is used to manage and control each hardware device and computer program 222 on the electronic device 20, and it can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the object detection and tracking method in the autonomous driving scenario executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 can further include computer programs that can be used to complete other specific tasks.

[0108] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the object detection and tracking method in the autonomous driving scenario disclosed above. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be repeated here.

[0109] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For related parts, reference can be made to the description in the method part.

[0110] Those skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0111] The steps of the methods or algorithms described in combination with the embodiments disclosed in this article can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0112] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.

[0113] The technical solutions provided in this application have been introduced in detail above. Specific examples are used in this text to elaborate on the principles and implementation manners of this application. The descriptions of the above embodiments are only used to help understand the method and its core idea of this application. At the same time, for those of ordinary skill in the art, based on the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.

Claims

1. A method for object detection and tracking in an autonomous driving scenario, characterized in that, Including: Obtain historical driving environment information obtained based on a preset sensor, and annotate the historical driving environment information to obtain a target data set; Use the target data set to train a pre-trained preset detection model to obtain a preset target detection model; According to the preset target detection model, perform real-time detection on the environment in the autonomous driving scenario to identify target objects, directly track the target objects moving in a straight line based on a preset target tracking model, and use the preset target tracking model to align the target objects moving non-linearly in the 3D scene, and project the aligned target objects onto a preset 2D plane, so as to continuously track the target objects by calculating the intersection over union of the target objects located on the preset 2D plane.

2. The target detection and tracking method in the autonomous driving scenario according to claim 1, characterized in that, The obtaining historical driving environment information obtained based on a preset sensor, and annotating the historical driving environment information to obtain a target data set includes: Obtain three-dimensional point cloud information of the surrounding driving environment obtained based on a lidar as historical driving environment information; And / or, obtain visual image information of the surrounding driving environment obtained based on a camera as historical driving environment information; Perform class annotation and status information annotation on each target object in the historical driving environment information; the status information includes the position information, size information, direction information, and appearance information of each target object.

3. The object detection and tracking method in the autonomous driving scenario according to claim 1, characterized in that, The using the target data set to train a pre-trained preset detection model to obtain a preset target detection model includes: Use the target data set to train a pre-trained preset detection model, and adjust the hyperparameters of the preset detection model until the performance index of the preset detection model reaches a preset target performance index to obtain a preset target detection model; the hyperparameters include the learning rate, batch size, and loss function; And / or, perform transformation operations on the historical driving environment information in the target data set based on data augmentation techniques to obtain a new target data set, use the new target data set to train a pre-trained preset detection model, and adjust the hyperparameters of the preset detection model until the performance index of the preset detection model reaches a preset target performance index to obtain a preset target detection model; the transformation operations include information cropping, information flipping, and brightness adjustment.

4. The object detection and tracking method in the autonomous driving scenario according to claim 1, characterized in that, The aligning the target objects moving non-linearly in the 3D scene, and projecting the aligned target objects onto a preset 2D plane, so as to continuously track the target objects by calculating the intersection over union of the target objects located on the preset 2D plane includes: Monitor the target objects based on the preset target tracking model to obtain target objects moving non-linearly; the target objects moving non-linearly include target objects traveling along a curved road and / or target objects currently performing a flipping motion; Perform a preset pose adjustment operation on the target object with non-linear motion to adjust the pose of the target object with non-linear motion in the 3D scene at each moment to the preset standard pose, so as to complete the alignment operation of the target object with non-linear motion at each moment; the preset pose adjustment operation includes rotation and translation; Project the aligned target object with non-linear motion onto a preset 2D plane to obtain 2D bounding boxes corresponding to the target object with non-linear motion at each moment; Determine the intersection over union (IoU) between the 2D bounding box corresponding to the target object with non-linear motion at the current moment and the 2D bounding box corresponding to the target object with non-linear motion at the previous moment, determine the target IoU exceeding the preset threshold from the IoU, and determine the target object with non-linear motion at the current moment and the target object with non-linear motion at the previous moment corresponding to the target IoU as the same target object at different moments, so as to continuously track the target object.

5. The method for target detection and tracking in an autonomous driving scenario according to claim 1, wherein Further include: Obtain the state information of each object in the surrounding environment of the autonomous driving vehicle, and based on a preset target tracking model and according to a preset distance range and a preset size range, determine an expected target object whose distance from the autonomous driving vehicle is within the preset distance range and whose size does not exceed the preset size range from each object corresponding to the state information; Match the trajectory information corresponding to the expected target object with the trajectory information of the autonomous driving vehicle to obtain a matching result, determine a first expected target object whose matching result reaches a preset matching threshold, and track the first expected target object; the trajectory information includes driving direction information and speed change information; Determine a second expected target object whose matching result does not reach the preset matching threshold, and perform appearance matching between the appearance information corresponding to the second expected target object and the target object tracked historically. If the matching is successful, track the second expected target object.

6. The object detection and tracking method in the autonomous driving scenario according to claim 1, wherein Further include: Obtain the tracking result obtained by tracking the target object using the preset target tracking model; Convert the tracking result into preset structured data and save the preset structured data; Wherein, the preset structured data includes the size information, position coordinate information, rotation direction information, category information, confidence level and object identifier information uniquely corresponding to the target object of the target object.

7. The object detection and tracking method in the autonomous driving scenario according to claim 6, characterized in that, Further include: Determine the model performance index corresponding to the preset target tracking model based on the tracking result; The model performance index includes multi-object tracking average precision and the number of identifier switches; The number of identifier switches is the number of occurrences of object identifier switching events caused by incorrect depth information; If the model performance index does not reach the preset optimal index, determine the problem corresponding to the model based on the model performance index, and adjust the parameters of the model for the problem to update the preset target tracking model.

8. An object detection and tracking device in an autonomous driving scenario, characterized in that, Include: A data acquisition module, configured to acquire historical driving environment information obtained based on a preset sensor and label the historical driving environment information to obtain a target data set; A model training module for training a pre-trained preset detection model using the target data set to obtain a preset object detection model; An object recognition and tracking module for performing real-time detection on the environment in an autonomous driving scenario according to the preset object detection model to identify target objects, directly tracking the target objects moving in a straight line based on a preset object tracking model, and using the preset object tracking model to align the target objects moving non-linearly in a 3D scene and project the aligned target objects onto a preset 2D plane, so as to continuously track the target objects by calculating the intersection over union of the target objects located on the preset 2D plane.

9. An electronic device, characterized in that, Comprising: A memory for storing a computer program; A processor for executing the computer program to implement the method for object detection and tracking in an autonomous driving scenario according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, For storing a computer program, which when executed by a processor implements the method for object detection and tracking in an autonomous driving scenario according to any one of claims 1 to 7.