Target tracking method and device, electronic equipment and storage medium
By introducing geometric features and type features of two-dimensional images into the target tracking technology and optimizing the Kalman filtered observations, the problem of ignoring two-dimensional information when filtering in three-dimensional space in the prior art is solved, and the accuracy and stability of the filtering results are improved.
Patent Information
- Application Number
- CN202510356859.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-24
AI Technical Summary
When existing target tracking technology filters in three-dimensional space, it ignores the geometric information in the two-dimensional image, making it difficult to deal with actual scenes such as occlusion and shadow, resulting in insufficient accuracy and stability of the filtering results.
The geometric features and type features of the two-dimensional image of the target are obtained in two-dimensional space, and optimized according to the confidence of the target in three-dimensional space. The target Kalman filter observations are optimized through the comprehensive confidence of different types of targets to obtain the target Kalman filter estimate.
By introducing geometric features and type features of two-dimensional images, the accuracy and stability of filtering results are improved, and the impact of outliers on filtering in scenes such as occlusion and shadow is effectively reduced.
Smart Images

Figure CN120198462A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of target tracking, and particularly to a target tracking method, device, electronic device, and storage medium. Background Art
[0002] In the field of autonomous driving technology, target tracking is an extremely important and fundamental task both in intelligent driving and in the construction of vehicle-road-cloud integration. Using filtering technology for target tracking can reduce sudden changes in the position, speed, and heading angle of the target. Usually, in the autonomous driving technology framework, the position, speed, and heading angle of the target need to be represented in a unified three-dimensional space coordinate system and form a trajectory according to the time series.
[0003] In the related art, the filtering method adopted in target tracking only uses the converted three-dimensional space information for filtering, ignoring the geometric information in the two-dimensional image, thus making it difficult to handle actual scenarios such as occlusion and shadow. Summary of the Invention
[0004] Embodiments of the present application provide a target tracking method, device, electronic device, and storage medium to improve the accuracy and stability of the filtering result in the target tracking process.
[0005] Embodiments of the present application adopt the following technical solutions:
[0006] In a first aspect, embodiments of the present application provide a target tracking method, where the target tracking method includes:
[0007] Obtain the geometric features and type features of the two-dimensional image of the target in the two-dimensional space;
[0008] Determine the comprehensive confidence of different types of targets according to the first confidence of the target at the current moment in the three-dimensional space, the second confidence of the target at the current moment, and the third confidence of the target at consecutive moments. The first confidence of the target is determined according to the geometric features and the type features of the two-dimensional image, the second confidence of the target is related to the geometric features of the two-dimensional image, and the third confidence of the target is determined according to the geometric features of the two-dimensional image of the target at the previous moment and the geometric features of the two-dimensional image of the target at the next moment;
[0009] Optimize the target Kalman filter observation value through the comprehensive confidence of different types of targets to obtain the target Kalman filter estimated value and use it as the current position and current attitude of the target.
[0010] In some embodiments, the determining the comprehensive confidence of different types of targets according to the first confidence of the target at the current moment in the three-dimensional space, the second confidence of the target at the current moment, and the third confidence of the target at consecutive moments includes:
[0011] Obtain the first confidence level of the target at the current moment according to the width-to-height ratio of different types of targets;
[0012] Obtain the second confidence level of the target at the current moment according to the maximum intersection-over-union ratio of the target;
[0013] Determine the anomaly detection result and / or normal detection result according to the width-to-height ratio of the target at the previous moment and the target at the next moment, and obtain the third confidence level of the target at the continuous moments;
[0014] Determine the comprehensive confidence level of different types of targets according to the first confidence level of the target at the current moment, the second confidence level of the target at the current moment, and the third confidence level of the target at the continuous moments.
[0015] In some embodiments, the obtaining of the geometric features and type features of the two-dimensional image of the target in the two-dimensional space includes:
[0016] According to the 2D box in the original target detection, obtain the ratio wh_ratio of the width W and height H of the 2Dbox of each target;
[0017] Determine the IOU of each 2D box with other 2D boxes on the original two-dimensional image, and obtain the maximum intersection-over-union ratio IOU_max.
[0018] In some embodiments, the obtaining of the first confidence level of the target at the current moment according to the width-to-height ratio of different types of targets includes:
[0019] Statistically obtain the threshold of the width-to-height ratio according to the width-to-height ratio of different types of targets;
[0020] If the width-to-height ratio of the target is greater than the threshold of the width-to-height ratio or less than N times the threshold of the width-to-height ratio, the first confidence level of the target at the current moment is the first value; otherwise, the first confidence level of the target at the current moment is the second value;
[0021] The obtaining of the second confidence level of the target at the current moment according to the maximum intersection-over-union ratio of the target includes:
[0022] If the maximum intersection-over-union ratio IOU_max of the target is less than the preset threshold, the second confidence level of the target at the current moment is a fixed value;
[0023] If the maximum intersection-over-union ratio IOU_max of the target is greater than the preset threshold, the second confidence level of the target at the current moment is smaller when the maximum intersection-over-union ratio IOU_max is larger.
[0024] In some embodiments, the determining of the anomaly detection result and / or normal detection result according to the width-to-height ratio of the target at the previous moment and the target at the next moment, and the obtaining of the third confidence level of the target at the continuous moments includes:
[0025] If the error value of the width-to-height ratio between the target at the previous moment and the target at the next moment is less than a preset value, the third confidence level of the target at consecutive moments is a fixed value;
[0026] If the error value of the width-to-height ratio between the target at the previous moment and the target at the next moment is greater than the preset value, the greater the error value of the width-to-height ratio between the target at the previous moment and the target at the next moment, the smaller the third confidence level of the target at consecutive moments.
[0027] In some embodiments, the optimizing the target Kalman filter observation value through the comprehensive confidence level of different types of targets to obtain a target Kalman filter estimated value and using it as the current position and current attitude of the target includes:
[0028] Predicting a prior target three-dimensional space position, target speed, and target angle according to the standard Kalman filter prediction process;
[0029] Optimizing the standard Kalman filter prediction process according to the comprehensive observation value confidence level to update the target angle and target three-dimensional space position;
[0030] Obtaining a posterior target three-dimensional space position, target speed, and target angle according to the updated result.
[0031] In some embodiments, the method further includes:
[0032] Updating the geometric features of the target two-dimensional image at the previous moment using the geometric features of the target two-dimensional image at the current moment according to the standard Kalman filter covariance matrix.
[0033] In a second aspect, an embodiment of the present application further provides a target tracking device, where the target tracking includes:
[0034] A first module, configured to obtain the geometric features and type features of the two-dimensional image of the target in a two-dimensional space;
[0035] A second module, configured to determine the comprehensive confidence level of different types of targets according to the first confidence level of the target at the current moment, the second confidence level of the target at the current moment, and the third confidence level of the target at consecutive moments in a three-dimensional space, where the first confidence level of the target is determined according to the geometric features and type features of the two-dimensional image, the second confidence level of the target is related to the geometric features of the two-dimensional image, and the third confidence level of the target is determined according to the geometric features of the target two-dimensional image at the previous moment and the geometric features of the target two-dimensional image at the next moment;
[0036] A third module, configured to optimize the target Kalman filter observation value through the comprehensive confidence level of different types of targets to obtain a target Kalman filter estimated value and using it as the current position and current attitude of the target.
[0037] In a third aspect, an embodiment of the present application further provides an electronic device, including: a processor; and a memory arranged to store computer-executable instructions, which when executed cause the processor to execute the above method.
[0038] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, which stores one or more programs, and when the one or more programs are executed by an electronic device including a plurality of application programs, the electronic device is caused to execute the above method.
[0039] The above at least one technical solution adopted in the embodiment of the present application can achieve the following beneficial effects: obtaining the geometric features and type features of the two-dimensional image of the target in the two-dimensional space, and determining the comprehensive confidence of different types of targets according to the first confidence of the target at the current moment, the second confidence of the target at the current moment, and the third confidence of the target at consecutive moments in the three-dimensional space. Finally, the observation value of the target Kalman filter is optimized through the comprehensive confidence of different types of targets, and the estimated value of the target Kalman filter is obtained and used as the current position and current attitude of the target. Based on the three-dimensional space filter by the above method, taking the image 2DBOX frame and the timing information of the 2DBOX frame as constraint conditions, the confidence of the target observation value is corrected. Thereby solving the influence of outliers on the filter in scenarios such as occlusion and shadow, and effectively improving the accuracy and stability of the filtering result. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments and descriptions thereof of the present application are used to explain the present application, and do not constitute an improper limitation of the present application. In the drawings:
[0041] Figure 1 It is a schematic flow chart of the target tracking method in the embodiment of the present application;
[0042] Figure 2 It is a schematic structural diagram of the target tracking device in the embodiment of the present application;
[0043] Figure 3 It is a schematic structural diagram of an electronic device in the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0045] Object tracking based on filtering is a method that uses a filter to estimate and predict the position of an object in a video sequence. The filter usually selects the Kalman filter, which is an optimal estimation method for linear systems and is applicable to linear Gaussian systems. It predicts the current state by establishing a motion model of the object and using the previous state and observation values.
[0046] However, during the process of filtering an object in three-dimensional space, if scenes such as occlusion, shadow, and reflection occur in the actual road environment, it is difficult for the deep learning network to detect accurate object detection results. As a result, these abnormal observation values seriously affect the accuracy of the object tracking results after filtering.
[0047] To address the above deficiencies, in the embodiments of the present application, an object tracking method is provided, which adds geometric information in a two-dimensional image to the three-dimensional space object filtering process as a constraint condition. Specifically, while updating the observation value by input, the 2D box corresponding to the observation value is synchronously updated. Further, during the prediction process, the aspect ratio feature of the 2D box itself and the temporal variation feature of the 2D box are used as constraints. In this way, for an object with an obvious mutation in the 2D box, the confidence of the observation value is reduced, and at the same time, the covariance matrix of the measurement value is reduced, obtaining a more accurate and stable filtering tracking result.
[0048] The following will describe in detail the technical solutions provided by the embodiments of the present application with reference to the accompanying drawings.
[0049] The embodiments of the present application provide an object tracking method. As Figure 1 shown, the schematic diagram of the object tracking method flow in the embodiments of the present application is provided. The method at least includes the following steps S110 to step S130:
[0050] Step S110, obtain the geometric features and type features of the two-dimensional image of the object in the two-dimensional space.
[0051] During object detection, the original object detection result, that is, the 2D box in the two-dimensional space, can be obtained. Based on the 2D box, the width and height of each object 2D box, the intersection over union of each 2D box with other 2D boxes, etc. can be calculated as the geometric information of the two-dimensional image. The geometric information of the two-dimensional image can be used as the geometric features of the two-dimensional image.
[0052] It can be understood that during object detection, the type information of the object can also be obtained synchronously, and the type information of the object can be used as the type features of the two-dimensional image.
[0053] Further, on the basis of the three-dimensional information of the three-dimensional filter, the above-mentioned geometric features of the two-dimensional image and the type features of the two-dimensional image are extended and added, thereby expanding the filtering input state quantity.
[0054] It should be noted that the above step S110 is performed in a two-dimensional space and does not involve calculations in a three-dimensional space.
[0055] Step S120: Determine the comprehensive confidence levels of different types of targets according to the target first confidence level at the current moment, the target second confidence level at the current moment, and the target third confidence level at consecutive moments in a three-dimensional space. The target first confidence level is determined according to the geometric features and the type features of the two-dimensional image. The target second confidence level is related to the geometric features of the two-dimensional image. The target third confidence level is determined according to the geometric features of the target two-dimensional image at the previous moment and the geometric features of the target two-dimensional image at the next moment.
[0056] The geometric features and type features of the two-dimensional image, and the temporal variation of the geometric information of the two-dimensional image can be used to determine the target first confidence level at the current moment, the target second confidence level at the current moment, and the target third confidence level at consecutive moments in a three-dimensional space respectively. Through the above target confidence levels, the confidence level of the observation value can be further obtained.
[0057] It can be understood that the "target first confidence level at the current moment", the "target second confidence level at the current moment", and the "target third confidence level at the current moment" can be used as the confidence levels of the target observation values respectively, and the geometric features and type features of the two-dimensional image of the target are used in calculating the confidence level of the target observation value. In this way, it can be used as a constraint condition to correct the confidence level of the observation value. It is used to correspond to situations such as occlusion, shadow, and reflection that occasionally occur in the actual scene, and reduce the impact on filtering.
[0058] The target first confidence level is determined according to the geometric features and the type features of the two-dimensional image. The aspect ratio threshold values of different types of targets are statistically obtained according to the geometric features of the targets in the current trajectory, and the confidence level 1 of the target at the current moment is calculated. It can be understood that if there are multiple targets, it is the confidence levels of multiple targets at the current moment. The target second confidence level is related to the geometric features of the two-dimensional image. Through features such as the intersection over union of the geometric features of the two-dimensional image, the confidence level 2 of the target at the current moment is obtained. The target third confidence level is determined according to the geometric features of the target two-dimensional image at the previous moment and the geometric features of the target two-dimensional image at the next moment. According to the historical trajectory, the change situation of the aspect ratios of the target at the current moment and the target at the next moment is statistically obtained, so as to obtain the confidence level 3 of the target at the current moment and the target at the next moment. Finally, the comprehensive confidence levels of different types of targets are calculated through the confidence levels 1, 2, and 3.
[0059] It should be noted that the above step S120 is calculated in a three-dimensional space, and the geometric features and type features of the two-dimensional image of the target determined in the two-dimensional space are only used for calculating the confidence level.
[0060] Step S130: Optimize the target Kalman filter observation value through the comprehensive confidence levels of different types of targets to obtain the target Kalman filter estimation value, which is used as the current position and current attitude of the target.
[0061] Based on the target observation value obtained by the Kalman filter method, the updated confidence levels, i.e., the comprehensive confidence levels of different types of targets, are used to update the target observation value to obtain the target estimation value. The current position and current attitude (angle) of the target estimation value are output as the results, reducing the sudden changes in the target position, speed, and heading angle.
[0062] Through the above method, based on the position, speed, and heading angle of the target during target tracking in a unified three-dimensional space coordinate system, and forming a trajectory according to the time series. The geometric features and type features of the two-dimensional image of the target are obtained in the two-dimensional space; then, according to the first confidence level of the target at the current moment, the second confidence level of the target at the current moment, and the third confidence level of the target at consecutive moments in the three-dimensional space, the comprehensive confidence levels of different types of targets, i.e., the confidence levels for updating the target observation value, are determined. The target Kalman filter observation value is optimized through the comprehensive confidence levels of different types of targets to obtain the target Kalman filter estimation value, which is used as the current position and current attitude of the target.
[0063] Through the above method, while inputting the updated observation value, the corresponding 2D box is synchronously updated. During the prediction process, the aspect ratio feature of the 2D box itself and the temporal variation feature of the 2D box are used as constraints. For the target with an obvious mutation in the 2D box, the confidence level of the observation value is reduced, and the covariance matrix of the measurement value is reduced to obtain a more accurate and stable filtering tracking result.
[0064] Different from the related technology, in the three-dimensional space coordinate system, the original state quantity matrix is the target position, target speed, and heading angle in the three-dimensional space, and the angular velocity. Through the above method, the input two-dimensional features are synchronously extended, i.e., the geometric features and type features of the two-dimensional image are added. Then, the confidence level of the observation value is calculated using the two-dimensional geometric information and the temporal variation of the two-dimensional geometric information. Finally, the standard Kalman filter prediction process is used to predict the prior three-dimensional space position and angle, and then the posterior three-dimensional space position and angle are calculated by combining the confidence level optimization update process.
[0065] In one embodiment of the present application, determining the comprehensive confidence of different types of targets according to the first confidence of the target at the current moment, the second confidence of the target at the current moment, and the third confidence of the target at consecutive moments in a three-dimensional space includes: obtaining the first confidence of the target at the current moment according to the width-height ratio of different types of targets; obtaining the second confidence of the target at the current moment according to the maximum intersection over union (IoU) of the target; determining the anomaly detection result and / or the normal detection result according to the width-height ratio of the target at the previous moment and the target at the next moment, and obtaining the third confidence of the target at consecutive moments; and determining the comprehensive confidence of different types of targets according to the first confidence of the target at the current moment, the second confidence of the target at the current moment, and the third confidence of the target at consecutive moments.
[0066] According to the ratio of the width W and height H of each target 2D box in the target 2D box in the original target detection result, and according to the intersection over union (IoU) in the original target detection result, the maximum intersection over union (IoU) is calculated. At the same time, the original target detection result also includes target type information.
[0067] Further, the ratio of the width W and height H is used to calculate the observation confidence to obtain the first confidence of the target at the current moment. The intersection over union (IoU) of the target is used to calculate the observation confidence to obtain the second confidence of the target at the current moment. The ratio of the width W and height H of the target at the current moment and the ratio of the width W and height H of the target at the next moment are used to calculate the observation confidence to obtain the third confidence of the target at the current moment and the target at the next moment. It can be understood that "the target at the current moment and the target at the next moment" refers to the targets at consecutive moments on the trajectory.
[0068] In one embodiment of the present application, obtaining the geometric features and type features of the two-dimensional image of the target in a two-dimensional space includes: obtaining the ratio wh_ratio of the width W and height H of each target 2D box according to the 2D box in the original target detection; and determining the IoU of each 2D box with other 2D boxes on the original two-dimensional image to obtain the maximum intersection over union IOU_max.
[0069] According to the 2D box obtained from the original target detection, calculate the ratio wh_ratio of the width W and height H of each target 2D box:
[0070] wh_ratio = (xmax - xmin) / (ymax - ymin)
[0071] where (xmax, ymax) and (xmin, ymin) are the upper left corner point and the lower right corner point of the 2D detection box respectively.
[0072] Then, calculate the Intersection over Union (IOU) between each 2D box and other 2D boxes on the two-dimensional image, and calculate and retain the maximum intersection over union iou_max.
[0073] iou_max = Max(iou_0, iou_1,..., iou_i)
[0074] It can be understood that the intersection over union is only a feasible implementation manner and is not used to limit the protection scope in the embodiments of the present application.
[0075] Finally, synchronously input two-dimensional geometric information when filtering the input state quantity. The original state quantity matrix is
[0076] where p x , p y are the target positions in three-dimensional space, v is the target speed, is the heading angle, is the angular velocity.
[0077] Synchronously expand the input two-dimensional features Adds the geometric features and type features type of the two-dimensional image.
[0078] In an embodiment of the present application, obtaining the first confidence level of the target at the current moment according to the width-to-height ratio of different types of targets includes: statistically obtaining the threshold of the width-to-height ratio according to the width-to-height ratio of different types of targets; if the width-to-height ratio of the target is greater than the threshold of the width-to-height ratio or less than N times the threshold of the width-to-height ratio, the first confidence level of the target at the current moment is the first value; otherwise, the first confidence level of the target at the current moment is the second value; obtaining the second confidence level of the target at the current moment according to the maximum intersection over union of the target includes: if the maximum intersection over union IOU_max of the target is less than the preset threshold, the second confidence level of the target at the current moment is a fixed value; if the maximum intersection over union IOU_max of the target is greater than the preset threshold, the second confidence level of the target at the current moment is smaller when the maximum intersection over union IOU_max is larger.
[0079] For the calculation of the confidence level α1 of the target at the current moment, according to the wh_ratio statistics of different types of targets, the threshold of wh_ratio is statistically obtained, and the threshold after statistics is as follows:
[0080] WH_THRESHOLD = {"person": 1.2, "bicycle": 1.2, "compact car": 2.0, "motorcycle": 1.2, "bus": 1.5, "cyclist": 1.2, "truck": 1.5}
[0081] N takes the value of 0.5,
[0082]
[0083] According to the above formula, for the current target confidence α1 of the target greater than the threshold and less than 0.5 times the threshold, it is assigned 0.5, and the confidence is low; for other cases, the current target confidence α1 is assigned 1, and the confidence is high.
[0084] Secondly, the current target confidence α2 is directly related to iou_max:
[0085]
[0086] The larger iou_max is, the lower the current target confidence α2 is.
[0087] If iou_max is less than 0.7, then α2 is assigned 1.
[0088] In an embodiment of the present application, the determining the abnormal detection result and / or the normal detection result according to the width-to-height ratio of the target at the previous moment and the target at the next moment, and obtaining the third confidence of the target at the continuous moment includes: if the error value of the width-to-height ratio between the target at the previous moment and the target at the next moment is less than a preset value, then the third confidence of the target at the continuous moment is a fixed value; if the error value of the width-to-height ratio between the target at the previous moment and the target at the next moment is greater than the preset value, then the larger the error value of the width-to-height ratio between the target at the previous moment and the target at the next moment is, the smaller the third confidence of the target at the continuous moment is.
[0089] The target confidences α3 at the previous moment and the next moment in the historical trajectory, and the confidence of the target at the continuous moment uses the change of wh_ratio of the target at the previous moment and the next moment to detect abnormal mutations of wh_ratio caused by occlusions, shadows, reflections, etc.
[0090] max_wh = max(wh_ratiopre - wh_ratio)
[0091] diff_wh = fabs(wh_ratiopre - wh_ratio)
[0092]
[0093] If the error of diff_wh between the current moment and the next moment is less than 0.5, then the confidence α3 of the target at the continuous moment is assigned 1, and if the error of diff_wh is greater than 0.5, then the confidence α3 of the target at the continuous moment is equal to max() and fbs() are common functions.
[0094] The larger the diff_wh between the current moment and the next moment, the smaller the target confidence level α3 for consecutive moments. wh_ratio pre Initialize to the WH_THRESHOLD corresponding to the type, and use the target wh_ratio of the previous moment at other moments.
[0095] In an embodiment of the present application, the optimizing the target Kalman filter observation value through the comprehensive confidence level of different types of targets to obtain the target Kalman filter estimated value and using it as the current position and current attitude of the target includes: predicting the prior target three-dimensional space position, target speed, and target angle according to the standard Kalman filter prediction process; optimizing the standard Kalman filter prediction process according to the comprehensive observation value confidence level to update the target angle and target three-dimensional space position; and obtaining the posterior target three-dimensional space position, target speed, and target angle according to the updated result.
[0096] Based on the constant velocity motion model and the constant angular velocity motion model, first use the standard Kalman filter prediction process to predict the prior three-dimensional space position and angle. Then, combine the confidence level optimization update process to calculate the posterior three-dimensional space position and angle. In addition, both the prior and posterior results include target speed information, that is, target speed and target angular velocity.
[0097] Specifically, the standard Kalman filter equation is
[0098]
[0099] Formulas (1) and (2) are the prediction processes, where formula (1) is the prior estimate, A is the state transition matrix, B is the control matrix, u t-1 is the control variable, is the state variable corrected at the previous moment, is the predicted prior estimate variable.
[0100] In formula (2) is the prior error, Q is the system noise covariance matrix, A T is the transpose matrix of the state transition matrix, P t-1 is the estimation error at the previous moment.
[0101] Formulas (3) to (5) are the update processes, K t is the Kalman gain, R is the observation noise covariance matrix, H is the transformation transfer matrix from the state variable to the observation, H T is the transpose matrix of the transformation transfer matrix from the state variable to the observation, z t is the observation variable, is the posterior estimate, is the predicted prior estimate variable. P tis the posterior error, is the prior error.
[0102] Updating the posterior estimate using the confidence only requires optimizing formula (3) to
[0103]
[0104] Confidence optimized update of the heading angle (attitude)
[0105]
[0106] In formula (7), is the predicted heading angle, is the angular velocity, and Δt is the time interval between the previous and current moments.
[0107] In formula (8), is the finally output heading angle after update, α is the confidence of the observation value under two-dimensional geometric condition constraints, is the heading angle component in the Kalman gain.
[0108] Prediction and update of position (position x, y)
[0109]
[0110] Formula (9) is the position update formula including the heading angle. Because in actual situations, the variance of the heading angle is larger in cases such as occlusion, in order to avoid abnormal heading angles affecting the accuracy of position prediction, the angular velocity is not used to update the position.
[0111] Formula (9) degenerates into formula (10).
[0112] In formula (10), are the predicted positions in the x and y directions, are the velocities in the x and y directions, is a constant, and Δt is the time interval between the previous and current moments.
[0113] In formula (11), are the finally output positions in the x and y directions after update, α is the confidence of the observation value under two-dimensional geometric condition constraints, are the position components in the x and y directions in the Kalman gain.
[0114] It can be understood that the constant velocity motion model and the constant angular velocity motion model are both well-known motion models in the art and will not be elaborated here.
[0115] In an embodiment of the present application, the method further includes: updating the geometric features of the target two-dimensional image at the previous moment using the geometric features of the target two-dimensional image at the current moment according to the standard Kalman filter covariance matrix.
[0116] Using the standard Kalman filter covariance matrix, update the target wh_ratio at the previous moment with the current two-dimensional geometric feature information, that is
[0117]
[0118] wh_ratio pre = wh_ratio.
[0119] Where P t is the posterior error, is the prior error, wh_ratio is the current target two-dimensional geometric feature information, and wh_ratio pre is the geometric feature information of the target two-dimensional image at the previous moment.
[0120] The embodiment of the present application also provides a target tracking device 200, as Figure 2 shown, providing a structural schematic diagram of the target tracking device in the embodiment of the present application. The target tracking device 200 at least includes: a first module 210, a second module 220, and a third module 230, where:
[0121] In an embodiment of the present application, the first module 210 is specifically configured to: obtain the geometric feature and type feature of the two-dimensional image of the target in the two-dimensional space.
[0122] During target detection, the original target detection result, that is, the 2D box in the two-dimensional space, can be obtained. Based on the 2D box, the width and height of each target 2D box, the intersection over union of each 2D box with other 2D boxes, etc. can be calculated as the geometric information of the two-dimensional image. The geometric information of the two-dimensional image can be used as the geometric feature of the two-dimensional image.
[0123] It can be understood that during target detection, the type information of the target can also be obtained synchronously, and the type information of the target can be used as the type feature of the two-dimensional image.
[0124] Further, on the basis of the three-dimensional information of the three-dimensional filter, the above-mentioned geometric feature of the two-dimensional image and the type feature of the two-dimensional image are added to expand the filtering input state quantity.
[0125] It should be noted that the above-mentioned first module 210 operates in the two-dimensional space and does not involve three-dimensional space calculations.
[0126] In an embodiment of the present application, the second module 220 is specifically configured to: determine the comprehensive confidence degrees of different types of targets according to the target first confidence degree at the current moment, the target second confidence degree at the current moment, and the target third confidence degree at consecutive moments in the three-dimensional space. The target first confidence degree is determined according to the geometric features and the type features of the two-dimensional image. The target second confidence degree is related to the geometric features of the two-dimensional image. The target third confidence degree is determined according to the geometric features of the target two-dimensional image at the previous moment and the geometric features of the target two-dimensional image at the next moment.
[0127] The geometric features and type features of the two-dimensional image and the temporal variation of the geometric information of the two-dimensional image can be used to determine the target first confidence degree at the current moment, the target second confidence degree at the current moment, and the target third confidence degree at consecutive moments in the three-dimensional space respectively. The confidence degree of the observation value can be further obtained through the above target confidence degrees.
[0128] It can be understood that the "target first confidence degree at the current moment", the "target second confidence degree at the current moment", and the "target third confidence degree at the current moment" can be used as the confidence degrees of the target observation values respectively, and the geometric features and type features of the two-dimensional image of the target are used in calculating the confidence degree of the target observation value, which can be used as constraint conditions to correct the confidence degree of the observation value. This is used to correspond to the occlusions, shadows, and reflections that occasionally occur in the actual scenario, and reduce the impact on filtering.
[0129] The target first confidence degree is determined according to the geometric features and the type features of the two-dimensional image. The aspect ratio threshold values of different types of targets are statistically obtained according to the geometric features of the targets in the current trajectory, and the confidence degree 1 of the target at the current moment is calculated. It can be understood that if there are multiple targets, it is the confidence degrees of multiple targets at the current moment. The target second confidence degree is related to the geometric features of the two-dimensional image. The confidence degree 2 of the target at the current moment is obtained through features such as the intersection over union of the geometric features of the two-dimensional image. The target third confidence degree is determined according to the geometric features of the target two-dimensional image at the previous moment and the geometric features of the target two-dimensional image at the next moment. According to the historical trajectory, the change situation of the aspect ratios of the target at the current moment and the target at the next moment is statistically obtained, so as to obtain the confidence degree 3 of the target at the current moment and the target at the next moment. Finally, the comprehensive confidence degrees of different types of targets are calculated through the confidence degrees 1, 2, and 3.
[0130] It should be noted that the above second module 220 performs calculations in the three-dimensional space, and the geometric features and type features of the two-dimensional image of the target determined in the two-dimensional space are only used for calculating the confidence degree.
[0131] In an embodiment of the present application, the third module 230 is specifically configured to: optimize the target Kalman filter observation value through the target comprehensive confidence degrees of different types to obtain a target Kalman filter estimation value, and use the value as the current position and current attitude of the target.
[0132] Based on the target observation value obtained by the Kalman filter method, the target observation value is updated by using the updated confidence degree, that is, the target comprehensive confidence degrees of different types, to obtain a target estimation value. The current position and current attitude (angle) of the target of the target estimation value are output as results, reducing the sudden changes in the target position, speed, and heading angle.
[0133] It can be understood that the above target tracking device can implement each step of the target tracking method provided in the foregoing embodiment. The relevant explanations of the target tracking method are applicable to the target tracking device and will not be elaborated here.
[0134] Figure 3 is a schematic structural diagram of an electronic device according to an embodiment of the present application. Please refer to Figure 3 , at the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and a memory. Among them, the memory may include a memory, such as a high-speed random access memory (Random-Access Memory, RAM), and may also include a non-volatile memory, such as at least one disk memory, etc. Of course, the electronic device may also include other hardware required for other services.
[0135] The processor, network interface, and memory can be interconnected through an internal bus. The internal bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 3 only a bidirectional arrow is used in [the figure] to represent it, but it does not mean that there is only one bus or one type of bus.
[0136] The memory is used to store a program. Specifically, the program may include program codes, and the program codes include computer operation instructions. The memory may include a memory and a non-volatile memory, and provide instructions and data to the processor.
[0137] The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it, forming a target tracking device at the logical level. The processor executes the program stored in the memory and is specifically used to perform the following operations:
[0138] Obtain the geometric features and type features of the two-dimensional image of the target in the two-dimensional space;
[0139] Determine the comprehensive confidence levels of different types of targets according to the first confidence level of the target at the current moment in the three-dimensional space, the second confidence level of the target at the current moment, and the third confidence level of the target at consecutive moments. The first confidence level of the target is determined according to the geometric features and the type features of the two-dimensional image, the second confidence level of the target is related to the geometric features of the two-dimensional image, and the third confidence level of the target is determined according to the geometric features of the two-dimensional image of the target at the previous moment and the geometric features of the two-dimensional image of the target at the next moment;
[0140] Optimize the target Kalman filter observation value through the comprehensive confidence levels of different types of targets, and obtain the target Kalman filter estimation value and use it as the current position and current attitude of the target.
[0141] The above as in this application Figure 1The method executed by the target tracking device disclosed in the illustrated embodiment can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor or the instructions in the form of software. The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.
[0142] The electronic device can also execute Figure 1 the method executed by the target tracking device in Figure 1 the illustrated embodiment and implement the functions of the target tracking device in
[0143] Embodiments of the present application also propose a computer-readable storage medium that stores one or more programs. The one or more programs include instructions that, when executed by an electronic device including a plurality of application programs, can enable the electronic device to execute Figure 1 the method executed by the target tracking device in the illustrated embodiment, and specifically used to execute:
[0144] Obtain the geometric features and type features of the two-dimensional image of the target in the two-dimensional space;
[0145] Determine the comprehensive confidence levels of different types of targets according to the current target first confidence level, the current target second confidence level, and the consecutive target third confidence level in three-dimensional space. The target first confidence level is determined according to the geometric features and the type features of the two-dimensional image. The target second confidence level is related to the geometric features of the two-dimensional image. The target third confidence level is determined according to the geometric features of the target two-dimensional image at the previous moment and the geometric features of the target two-dimensional image at the next moment.
[0146] Optimize the target Kalman filter observation value through the comprehensive confidence levels of different types of targets to obtain the target Kalman filter estimation value, which is used as the current position and current attitude of the target.
[0147] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can be implemented in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can be implemented in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0148] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0149] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0150] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide for implementing the steps in the process Figure 1 one process or multiple processes and / or blocks Figure 1 steps of the functions specified in one block or multiple blocks.
[0151] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0152] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.
[0153] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0154] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.
[0155] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0156] The above are only the embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A target tracking method, wherein: The target tracking method comprises: Acquire geometric features and type features of a two-dimensional image of a target in two-dimensional space; Determine the comprehensive confidences of different types of targets according to a first confidence of the target at the current moment, a second confidence of the target at the current moment, and a third confidence of the target at consecutive moments in three-dimensional space, wherein the first confidence of the target is determined according to the geometric features of the two-dimensional image and the type features, the second confidence of the target is related to the geometric features of the two-dimensional image, and the third confidence of the target is determined according to the geometric features of the two-dimensional image of the target at a previous moment and the geometric features of the two-dimensional image of the target at a next moment; The target Kalman filter observation value is optimized by using the different types of target comprehensive confidences to obtain the target Kalman filter estimation value and use it as the current position and current posture of the target.
2. The method of claim 1, wherein: Determining different types of comprehensive confidences of targets according to the first confidence of the target at the current moment, the second confidence of the target at the current moment, and the third confidence of the target at consecutive moments in the three-dimensional space includes: Obtaining a first confidence level of the target at the current moment according to the width-to-height ratios of different types of targets; According to the maximum intersection-over-union ratio of the target, a second confidence level of the target at the current moment is obtained; Determine an abnormal detection result and / or a normal detection result according to a width-to-height ratio of a target at a previous moment and a target at a next moment, and obtain a third confidence level of the target at the consecutive moments; According to the first confidence of the target at the current moment, the second confidence of the target at the current moment, and the third confidence of the target at the consecutive moments, the comprehensive confidence of different types of targets is determined.
3. The method of claim 2, wherein: The step of acquiring the geometric features and type features of the two-dimensional image of the target in the two-dimensional space includes: According to the 2D box in the original target detection, get the ratio wh_ratio of the width W and height H of the 2D box of each target; Determine the IOU of each 2D box with other 2D boxes on the original two-dimensional image and obtain the maximum intersection-union ratio IOU_max.
4. The method of claim 2, wherein: The obtaining, according to the width-to-height ratios of different types of targets, the first confidence level of the target at the current moment includes: According to the width-to-height ratios of different types of targets, the threshold of the width-to-height ratio is obtained by statistics; If the width-to-height ratio of the target is greater than the width-to-height ratio threshold or less than the width-to-height ratio threshold N*, the first confidence of the target at the current moment is the first value; otherwise, the first confidence of the target at the current moment is the second value; The step of obtaining the second confidence level of the target at the current moment according to the maximum intersection-over-union ratio of the target includes: If the maximum intersection-over-union ratio IOU_max of the target is less than the preset threshold, the second confidence of the target at the current moment is a fixed value; If the maximum intersection-over-union (IOU_max) of the target is greater than the preset threshold, the larger the maximum intersection-over-union (IOU_max) is, the smaller the second confidence of the target at the current moment is.
5. The method of claim 2, wherein: The determining of the abnormal detection result and / or the normal detection result according to the width-to-height ratio of the target at the previous moment and the target at the next moment, and obtaining the third confidence level of the target at the consecutive moments, includes: If the error value of the width-to-height ratio between the target at the previous moment and the target at the next moment is less than a preset value, the third confidence level of the target at the consecutive moments is a fixed value; If the width-to-height ratio error between the target at the previous moment and the target at the next moment is greater than the preset value, the greater the width-to-height ratio error between the target at the previous moment and the target at the next moment, the smaller the third confidence of the target at the consecutive moments.
6. The method of claim 1, wherein: The optimizing the target Kalman filter observation value by the target comprehensive confidence of different types to obtain the target Kalman filter estimation value as the current position and current posture of the target includes: According to the standard Kalman filter prediction process, the prior target three-dimensional spatial position, target speed and target angle are predicted; Optimizing the standard Kalman filter prediction process according to the comprehensive observation confidence to update the target angle and the target three-dimensional spatial position; According to the updated results, the posterior target three-dimensional spatial position, target speed and target angle are obtained.
7. The method of claim 1, wherein: The method further comprises: According to the standard Kalman filter covariance matrix, the geometric features of the target two-dimensional image at the current moment are used to update the geometric features of the target two-dimensional image at the previous moment.
8. A target tracking device, wherein: The target tracking includes: The first module is used to obtain geometric features and type features of a two-dimensional image of a target in a two-dimensional space; The second module is used to determine the comprehensive confidence of different types of targets according to the first confidence of the target at the current moment, the second confidence of the target at the current moment and the third confidence of the target at consecutive moments in the three-dimensional space, wherein the first confidence of the target is determined according to the geometric features of the two-dimensional image and the type features, the second confidence of the target is related to the geometric features of the two-dimensional image, and the third confidence of the target is determined according to the geometric features of the two-dimensional image of the target at a previous moment and the geometric features of the two-dimensional image of the target at a next moment; The third module is used to optimize the target Kalman filter observation value through the comprehensive confidence of the different types of targets, obtain the target Kalman filter estimation value and use it as the current position and current posture of the target.
9. An electronic device, comprising: processor; as well as A memory arranged to store computer executable instructions, which when executed cause the processor to perform the method of any one of claims 1 to 7.
10. A computer-readable storage medium storing one or more programs, which, when executed by an electronic device including a plurality of application programs, causes the electronic device to execute any one of the methods of claims 1 to 7.