Long-term target tracking method and system integrating depth information
By integrating depth information and CA model to predict the target depth, confidence discriminant and adaptive scale factors are constructed, and the problem of target loss in complex environments is solved, and long-term effective target tracking is achieved.
Patent Information
- Application Number
- CN202111284990.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-01
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2041-11-01
AI Technical Summary
The existing DCF visual target tracking algorithms are prone to loss of targets due to occlusion and background aliasing in complex environments, and lack the confidence evaluation module for target tracking results, resulting in accumulated errors in the tracking model and unable to achieve long-term effective tracking.
By fusion of depth information, a laser rangefinder is used to measure the target depth, and predict the target depth of the next frame based on the CA model, the target confidence discrimination and adaptive scale factor are constructed, the confidence of the tracking results is evaluated, and the filter model is dynamically adjusted.
It effectively reduces the adverse effects of occlusion and background aliasing on tracking, improves the prediction accuracy of target positions and scales, enhances the robustness and real-time nature of the algorithm, and achieves long-term effective target tracking.
Smart Images

Figure CN114359330B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of long-term target tracking, and in particular to a long-term target tracking method and system integrating depth information. Background Art
[0002] Discriminative Correlation Filter (DCF) tracking applies correlation filter theory to target tracking, and can perform classifier learning and tracking in the frequency domain, which greatly speeds up target tracking and has strong robustness. It is one of the key directions of current visual target tracking research. Subsequently, a variety of algorithms were proposed to improve the algorithm framework, the features used, the construction of scale pyramids, and other aspects.
[0003] like Figure 1 As shown in the figure, it is a typical existing DCF tracking algorithm. Its basic principle is that in the training stage, that is, in the current frame, the two-dimensional position filter and the one-dimensional scale filter are learned and trained respectively. The two filters are trained and updated separately without affecting each other. In the target tracking stage, that is, in the next frame, the two-dimensional position-related filter is first used to determine the new position of the target, and then the one-dimensional scale-related filter is used to obtain candidate images of different scales with the current center position as the center point, so as to find the most matching scale. After that, the related filter is retrained and updated, and the new filter is used for target tracking. After that, the process is repeated continuously to achieve target tracking.
[0004] However, whether it is the DSST algorithm or the ECO algorithm, the DCF algorithm will continuously update the position and scale filters online during the tracking process. Although the tracking model is constantly updated, due to the lack of a tracking result confidence assessment module, once the tracking result deviates, it will cause the model trained in the new frame to deviate. As time goes by, the error of the tracking model will continue to accumulate, and eventually the trained related filters will be unable to reflect the true state of the target, and the tracking will eventually fail.
[0005] The overall architecture of the DCF-based visual target tracking algorithm is becoming increasingly perfect, but the algorithm complexity is increasing, and the algorithm running speed cannot meet the real-time requirements. Especially in complex environments, there are complex situations such as background aliasing and target occlusion, and the tracking algorithm is prone to lose the target. Due to the lack of a target tracking result confidence assessment module, DCF-based visual target tracking is a typical short-term target tracking.
[0006] In view of this, there is an urgent need to provide a method based on a DCF-type tracking algorithm that can effectively reduce the adverse effects caused by occlusion and background aliasing, thereby achieving long-term and effective tracking of the target. Summary of the invention
[0007] In order to solve the above technical problems, the technical solution adopted by the present invention is to provide a long-term target tracking method integrating depth information, comprising the following steps:
[0008] Step 1, initialize the position-related filter, scale-related filter and predicted target depth model;
[0009] Step 2: Use a laser rangefinder to measure the target depth information of the current frame;
[0010] Step 3: According to the target depth information of the current frame, the target depth of the next frame is predicted based on the CA model;
[0011] Step 4: According to the input new frame image, the position and scale of the target are estimated using the position correlation filter and the scale correlation filter;
[0012] Step 5: According to the target depth predicted in step 3, a target confidence discriminant is constructed, the output of the position filter is discriminated, and the target position is determined according to the discrimination result;
[0013] Step 6: Based on the predicted target depth, construct an adaptive scale factor to determine the target scale;
[0014] Step 7: Input the estimated target position and scale into the hardware tracking system, aim the gimbal at the target, and acquire a new frame of image;
[0015] Step 8: Repeat steps 2 to 7 until the target tracking is completed.
[0016] In the above method, the step 3 is specifically:
[0017] If the laser rangefinder measurement frequency is set higher than the target motion state change frequency, then within one measurement cycle, the third-order constant acceleration CA model of the target motion state is:
[0018]
[0019] Among them, ω(k) has a mean of zero and a variance of Gaussian white noise;
[0020] The depth information and gimbal attitude angle of the target are measured with a period of T, including the pitch angle γ k and yaw angle θ k , using geometric relationships and motion equations to calculate the target's motion state A(k) at time k:
[0021]
[0022] Combine the discrete coordinate points to get the target path sequence S(k)∈S 1 ,…,Sk ;
[0023] Substitute A(k) in equation (2) into equation (1) to predict the target motion state A(k+1) at the next moment, and finally calculate the distance D of the target at the next moment pred ,
[0024]
[0025] In the above method, in step 5, the reliability of the target position estimated in step 4 is judged according to the preset target position confidence threshold, and whether to update the target or correct the predicted position is specifically as follows:
[0026] When there is an obstruction in front of the target, the depth output by the laser rangefinder is the depth of the obstruction d 2 , d 2 With target depth d 1 Compared with, there is a mutation value δ d =d 2 -d 1 ;
[0027] According to the preset target position confidence threshold δ max , if δ d >δ max , that is, if the mutation value is greater than the threshold, the target is considered to be blocked and interfered, the target is not updated, and the position is corrected using the target position predicted by the CA model; otherwise, the target is updated according to the maximum response of the position-related filter.
[0028] In the above method, step S5 comprises:
[0029] Using acceleration and laser rangefinder noise factor Q sensor Dynamically adjust the confidence threshold δ max , the accelerations of the target in the x, y, and z directions at time k are Dynamically adjust the threshold δ max for:
[0030]
[0031] In the above method, step S6 is specifically:
[0032] The adaptive scale factor is calculated using the predicted target depth and multiplied by the target scale at time k to obtain the predicted target scale s at time k+1 pred , and at the target center, 17 scale samples {s -8 …s 0 …s 8}, after filtering the scale samples, find the scale s that makes the filter respond the largest max , which is the target scale.
[0033] In the above method, the adaptive scale factor calculation process is as follows:
[0034] The target predicted width and target true width are calculated according to the proportional relationship as follows:
[0035]
[0036]
[0037] Where f is the focal length of the camera, u is the pixel size, D is the target depth, and the subscripts last, curr, and pred in the figure represent the previous frame, current frame, and predicted frame respectively;
[0038] D ini and w ini are the initialization target depth and the initialization target image width, respectively.
[0039] The target depth D predicted by the CA model pred Substituting the target actual width w calculated by formula (6) into formula (5), we get the predicted target width:
[0040]
[0041] Introducing the target width adaptive scaling factor A w , target height adaptive scale factor A h The calculation of is as follows:
[0042]
[0043] Where h represents the target height; is the scale factor limiting parameter; β = 1.1;
[0044] The target sample scale selection principle for target scale evaluation is:
[0045]
[0046] Among them, S=17, which is the scale sample.
[0047] The present invention also provides a long-term target tracking system integrating depth information, comprising
[0048] Laser rangefinder, bus servo and boarding position connected to the control system;
[0049] A camera connected to the upper camera station;
[0050] The camera is used to collect images and send them to the upper camera position;
[0051] Laser rangefinder, used to measure the target depth information of the current frame and send it to the control system;
[0052] The bus servo controls the rotation of the laser rangefinder and camera according to the control system instructions;
[0053] The control system includes an initialization unit for initializing a position-related filter, a scale-related filter, and a predicted target depth model;
[0054] The CA model unit predicts the target depth of the next frame based on the target depth information of the current frame measured by the laser rangefinder and the CA model;
[0055] The target position and scale estimation unit estimates the target position and target scale respectively by using the position correlation filter and scale correlation filter in the predicted target depth model according to the input new frame image;
[0056] The target position confirmation unit constructs a target confidence discriminant according to the predicted target depth, discriminates the output of the position filter, and determines the target position according to the discrimination result;
[0057] Target scale confirmation unit: constructs an adaptive scale factor based on the predicted target depth and determines the target scale;
[0058] The command sending unit sends the target position determined by the target position confirmation unit and the target scale determined by the target scale confirmation unit to the bus servo respectively, and the bus servo adjusts the position of the laser rangefinder and the camera; and decides whether to send a control command to the bus servo to control the laser rangefinder and the camera according to the determination result of the target tracking determination unit;
[0059] Target tracking judgment unit: determines whether the target tracking is completed. If not, the instruction sending unit sends an instruction to the camera to obtain a new frame image to restart target tracking; if completed, the work is terminated.
[0060] The present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the long-term target tracking method integrating depth information as described above is implemented.
[0061] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the long-term target tracking method integrating depth information as described above is implemented.
[0062] The present invention integrates the depth information of the target into the entire tracking process, and uses the difference between the target and the background in the depth direction to evaluate the confidence of the tracking result. When the target depth change conforms to the expected law, the tracking result is considered reliable, and the relevant filter model is updated, otherwise it is not updated. With this strategy, the adverse effects caused by occlusion and background aliasing can be effectively reduced, thereby achieving long-term and effective tracking of the target. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0064] Figure 1 The basic principle block diagram of the typical DCF algorithm provided by the present invention;
[0065] Figure 2 A schematic diagram of a target tracking framework provided by the present invention;
[0066] Figure 3 A flow chart of the method provided by the present invention;
[0067] Figure 4 A schematic diagram of the system connection structure provided by the present invention;
[0068] Figure 5 The present invention provides a three-dimensional rectangular coordinate system with the tracking system as the origin;
[0069] Figure 6 A schematic diagram of the principle of constructing a target position confidence discriminant provided by the present invention;
[0070] Figure 7 A schematic diagram of the optimized scale sample extraction process provided by the present invention;
[0071] Figure 8 A schematic diagram of the principle of calculating the adaptive scale factor provided by the present invention;
[0072] Fig. 9 Schematic diagram of the setup (a) and (b) two experiments provided by the present invention;
[0073] Fig.10 The present invention provides two experimental results curves (a) and (b) of the target depth prediction experiment based on FIG. (9);
[0074] Fig.11 A diagram showing the scale accuracy experimental results provided by the present invention;
[0075] Fig.12 6 sets of tracking system performance experimental result diagrams provided by the present invention;
[0076] Fig.13 A schematic diagram of the system structure provided by the present invention;
[0077] Fig.14 A schematic block diagram of a computer device provided by the present invention. DETAILED DESCRIPTION
[0078] The technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0079] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise" and the like indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the referred device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as limiting the present invention.
[0080] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, "multiple" means two or more, unless otherwise clearly and specifically defined. In addition, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal connection of two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0081] The present invention is described in detail below in conjunction with specific implementation methods and the accompanying drawings.
[0082] like Figure 2-3 As shown, the present invention provides a long-term target tracking method integrating depth information, comprising the following steps:
[0083] Step 1: Initialize the position-related filter, scale-related filter and predicted target depth model.
[0084] Step 2: Use a laser rangefinder to measure the target depth information of the current frame.
[0085] This embodiment is preferred because the target tracking method proposed in this embodiment integrates the depth information of the target, while the existing public evaluation sets are mostly based on 2D image sequences and cannot provide depth information; and the target depth information is collected by the laser rangefinder sensor, and the steering gear needs to rotate to ensure that the depth belongs to the tracked target. Therefore, in order to meet the measurement of target depth information in this embodiment, the laser rangefinder is improved and a pan-tilt tracking system is designed. The principle block diagram of the pan-tilt tracking system is shown in FIG. Figure 4 As shown, the system uses STM32F103C8 as the main control chip, and the pan-tilt is equipped with a camera, a laser rangefinder and connected to the main control chip, which can complete the tasks of depth information acquisition, image acquisition, steering gear control, and communication with the host computer. The two-axis pan-tilt is driven by two bus steering gears, which control the rotation in the yaw and pitch directions respectively. The steering gear model is LX-224, which adopts a single bus communication method. It has an internal controller to complete the angle PID and current PID control. Only one serial port instruction is needed to rotate the laser rangefinder to a given angle, and the error does not exceed ±0.24°. The laser rangefinder can measure depth at a frequency of 100Hz, with a resolution of 1cm, which can meet the requirements of measuring the target depth information in this embodiment. The camera is a CMOS camera with a USB interface. This camera can be directly connected to the computer via the USB interface without an image acquisition card, and is easy to develop.
[0086] Step 3: According to the target depth information of the current frame, based on the CA (Constant Acceleration) model, predict the target depth of the next frame; specifically:
[0087] In this embodiment, when establishing a mathematical model for target depth prediction, the general principle is to make the established mathematical model both conform to the actual situation of the target and facilitate real-time processing. When the target is in non-maneuvering motion, the model is easy to establish; but for maneuvering targets, ideal modeling becomes very difficult. Because there is little prior knowledge about the target's maneuvers, it is difficult to accurately express it with mathematical expressions, and it can only be described using approximate methods under various assumptions. The feature of this embodiment is that it uses a camera and a laser rangefinder to continuously measure, and the measurement frequency is set higher than the frequency of change of the target's motion state. Therefore, within a measurement cycle, the visible target performs uniformly accelerated linear motion, and its motion state can be described by the following third-order constant acceleration (CA) motion model:
[0088]
[0089] Assuming that the target moves in three-dimensional space, the target motion state variable A(k) in the above formula can be defined as x, y, z represent the position components of the target in the x, y, and z directions respectively; ω(k) has a mean of zero and a variance of Gaussian white noise is the measurement noise; T is the measurement period.
[0090] The above model is the most basic model in the target motion model, with low computational complexity and suitable for real-time tracking. For uniform, uniformly accelerated linear motion or motion with approximately uniform speed and uniform acceleration, the above model can accurately describe the target motion state and thus predict the target spatial position and depth information.
[0091] This embodiment takes the target doing "S-shaped" motion as an example. Figure 5 As shown in Figure 1, a three-dimensional rectangular coordinate system is established with the tracking system as the origin. First, the system measures the depth information of the target and the gimbal attitude angle (including the pitch angle γ k and yaw angle θ k ), and use the geometric relationship and motion equation to calculate the motion state A(k) of the target at time k:
[0092]
[0093] The discrete coordinate points ( Figure 5 The target path sequence S(k)∈S is obtained by combining the coordinate points in 1 ,…,S k Since the measurement frequency is greater than the target motion state change frequency, the motion path between time k+1 and time k can be approximated as a straight line ( Figure 5 The dashed line in the middle indicates that in the next measurement period T, the target is considered to be moving in a uniformly accelerated linear motion in all three directions. Substituting A(k) in equation (2) into equation (1), the target motion state A(k+1) at the next moment can be predicted, and the distance D of the target at the next moment can be finally calculated. pred ,
[0094]
[0095] Step 4: Based on the input new frame image, the position and scale of the target are estimated using the position-related filter and the scale-related filter.
[0096] Step 5: According to the predicted target depth, a target confidence discriminant is constructed, the output of the position filter is discriminated, and the target position is determined according to the discrimination result.
[0097] In this embodiment, since DCF algorithms (DSST algorithm and ECOHC algorithm) are not robust, the target may be lost when the background is complex or the target is blocked. In order to enhance the robustness of features and improve tracking accuracy, this embodiment fuses the target prediction depth to construct the target position confidence discriminant, such as Figure 6 As shown,
[0098] When there is an obstruction in front of the target, the depth output by the laser rangefinder is the depth of the obstruction d 2 .d 2 With target depth d 1 Compared with, there is a mutation value δ d =d 2 -d 1 .
[0099] This embodiment defines δ max is the target position confidence threshold, if δ d >δ max , that is, if the mutation value is greater than the threshold, the target is considered to be blocked and interfered, the target is not updated, and the position is corrected using the target position predicted by the CA model; otherwise, the target is updated according to the maximum response of the position-related filter.
[0100] The traditional confidence threshold δ max It is determined a priori and is fixed. However, it is very difficult to determine the threshold a priori, and it is always fixed during the tracking process, which will lead to an increase in the tracking failure rate and reduce the tracking accuracy and stability when the actual motion state of the target changes greatly or the laser rangefinder has a large noise. In order to solve this problem, this embodiment uses the acceleration and the laser rangefinder noise factor Q sensor Dynamically adjust the confidence threshold δ max , the accelerations of the target in the x, y, and z directions at time k are Dynamically adjust the threshold δ max for:
[0101]
[0102] The above confidence threshold function has the following characteristics: as the acceleration of the target increases, the threshold increases, indicating that it can tolerate large changes in the depth of the target caused by large changes in the target's motion state; at the same time, after the acceleration increases to a certain value, the threshold tends to remain unchanged, ensuring that it can accurately judge whether the target is obscured. If the target is determined to be obscured, the predicted target space position is used to correct the position to prevent the target from being lost. For the near-linear imaging system in the central area of the field of view, the three-dimensional coordinates of the target space prediction are imaged onto the two-dimensional image plane to generate the two-dimensional predicted image plane coordinates.
[0103] Step 6: Based on the predicted target depth, construct an adaptive scale factor and determine the target scale. The specific steps include:
[0104] The existing DCF algorithm directly extracts 33 samples for scale detection based on the previous scale, which is not very accurate and greatly increases the computational complexity. Figure 7 As shown in the figure, in order to reduce the number of scale samples and improve tracking accuracy and real-time performance, the adaptive scale factor is calculated using the predicted target depth and multiplied by the target scale at time k to obtain the predicted target scale s at time k+1 pre d , and at the target center, 17 scale samples {s -8 …s 0 …s 8}, after filtering the scale samples, find the scale s that makes the filter respond the largest max , which is the target scale.
[0105] In this embodiment, the adaptive scale factor calculation process is as follows: Figure 8 As shown,
[0106] The target predicted width and target true width are calculated according to the proportional relationship as follows:
[0107]
[0108]
[0109] Where f is the focal length of the camera, u is the pixel size, and D is the target depth. last ,curr,pred represent the previous frame, current frame and predicted frame respectively;
[0110] D ini and w ini Initialize the target depth and initialize the target image width respectively.
[0111] The target depth D predicted by the above CA model is pred Substituting the target actual width w calculated by formula (6) into formula (5), we get the predicted target width:
[0112]
[0113] In order to adapt to the scale change caused by target motion, improve scale accuracy and reduce the number of scale sampling, this embodiment introduces a target width adaptive scale factor A based on target depth information. w , for the target height adaptive scale factor A h The calculation of is similar to the above process, and we get:
[0114]
[0115] Where h represents the target height;
[0116] is the scale factor limit parameter, which is used to prevent the sudden change of target depth from causing the scale factor α to oscillate;
[0117] According to the experimental design and analysis results, in this embodiment, β=1.1.
[0118] The target sample scale selection principle for target scale evaluation is:
[0119]
[0120] Among them, S=17, which is the scale sample.
[0121] Based on the above analysis, the introduction of an adaptive scaling factor in this embodiment will effectively improve the scaling accuracy and algorithm real-time performance.
[0122] Step 7: Based on the target position and scale determined in Steps 5 and 6, the image acquisition device is realigned with the target to acquire a new frame of image.
[0123] Step 8: Repeat steps 2 to 7 until the target tracking is completed.
[0124] Aiming at the problem that the traditional discriminant correlation filter (DCF) tracking algorithm is easy to lose the target when it is occluded, this method integrates the target depth information into the DCF framework, uses a laser rangefinder to measure the target depth information of the current frame, and then predicts the target depth of the next frame based on the CA model; then uses the predicted depth to construct the target position confidence discriminant to determine whether to update the target and correct the predicted position, thereby improving the tracking accuracy; finally, an adaptive scale factor based on the predicted depth is introduced to reduce the scale filter level, thereby improving the scale accuracy and real-time performance of the algorithm.
[0125] The present invention predicts the target depth of the next frame based on a constant acceleration model; integrates the target depth information into the DCF algorithm framework, constructs a target position confidence discriminant, and corrects the predicted position, which can enhance the reliability of the algorithm for complex situations such as occlusion and background aliasing; finally, introduces an adaptive scale factor based on the predicted depth to reduce the scale filter level, improve the scale accuracy and real-time performance of the algorithm. The depth information of the target is integrated into the entire tracking process, and the difference between the target and the background in the depth direction is used to evaluate the confidence of the tracking result. When the target depth change conforms to the expected law, the tracking result is considered reliable, and the relevant filter model is updated, otherwise it is not updated. With this strategy, the adverse effects caused by occlusion and background aliasing can be effectively reduced, thereby achieving long-term and effective tracking of the target.
[0126] The effectiveness of the above method is analyzed through experiments as follows.
[0127] This experiment is based on the above-mentioned pan-tilt tracking system framework. In order to test the algorithm tracking performance in the real environment with occlusions, and considering the motion characteristics of ground targets, it is assumed that the target motion has no fluctuations in the z-axis direction. Two experiments are set up as follows: Fig. 9 As shown in Figure (a), the target moves in a curve, and the short line is the occluder; in Figure (b), the target (rightward arrow in the figure) and the occluder (leftward arrow in the figure) move towards each other. The experimental parameters are shown in Table 1.
[0128] Table 1. Experimental parameter settings
[0129]
[0130] (I) Experiment on predicted target depth based on CA model
[0131] The key to the success of this method is the depth information of the target, so the accuracy of the predicted depth determines the accuracy of the algorithm. Fig. 9 The trajectory motion in (a) and (b). The experimental results are shown in Fig.10 As shown in (a)(b).
[0132] The vertical axis in the figure is the depth information, denoted as D, and the horizontal axis is the number of measurements, denoted as n. The error of the predicted depth information is described by the mean difference between the target distance and the predicted target distance:
[0133]
[0134] In stable tracking (no occlusion), that is, Figure (a)n a >40, Figure (b)n b <55: Ba=9.109, B b =12.03. The error is no more than 2% at a distance of 600 cm, indicating that the depth prediction model can accurately predict the target depth. At the same time, when the target is blocked (i.e., n a <40 part, n in Figure (b) b >55 part), the predicted target distance and the actual target distance will vary greatly, which are 105.7cm and 49.8cm respectively, indicating that the model can effectively determine whether the target is blocked.
[0135] (II) Scale Accuracy Experiment
[0136] Too large a template scale will increase background features, while too small a scale will reduce target features and affect tracking accuracy. Therefore, an experiment is designed here to move the target along the optical axis. The results are shown in Figure 2. Fig.11(The number in the lower left corner indicates the coverage rate). When the target moves along the optical axis, the scale of the target will change. As the tracking time increases, the scale error will continue to accumulate. From the results, it can be seen that whether it is DSST, ECOHC, or D-DSST and D-ECOHC that integrate depth information, the overlap rate index between the predicted frame and the true frame is decreasing from about 1. Fig.11 The experimental results shown are the results of tracking for about 30 seconds. It can be seen from the experimental results that the D-DSST and D-ECOHC algorithms that consider depth information have higher scale estimation accuracy and effectively improve the scale accuracy when the target moves along the optical axis.
[0137] (III) Target tracking experiment and analysis
[0138] The host computer CPU uses Inter(R)Core(TM)i5-7300HQ CPU@2.5GHZ; the memory is 12GB; the operating system is Windows10; and the framework is Matlab2017a.
[0139] Evaluate the physical tracking system from two aspects: accuracy and real-time performance. Accuracy:
[0140] 1. Introducing tracking probability P s , within the working time t, when the deviation between the target center and the image center is greater than the fixed threshold, the target tracking is considered successful if the deviation is corrected within 0.5s (counted into the successful tracking time t s ), otherwise it is considered a failure.
[0141] 2. Whether the target is lost during the tracking process; Real-time performance: The algorithm's FPS (frames per second) is used as the basis for judgment, that is, the number of image frames processed per second. Fig.12 In order to compare the tracking performance of the six algorithms in the experiment, the following Table 2 shows the tracking performance (the horizontal line indicates the target is lost). From the experimental results, it can be concluded that: compared with the DSST and ECO algorithms, the improved algorithms D-DSST and D-ECOHC have high tracking probabilities and no target loss, and their frame rates are greater than 30fps, meeting the real-time requirements; although the ECODEEP algorithm can also overcome occlusion interference, it is slow and difficult to apply to actual tracking; the GSF-DCF algorithm is slow, and the target leaves the field of view, resulting in tracking failure.
[0142]
[0143]
[0144] Table 2. Experimental results of tracking system performance
[0145] The experimental results show that in the case of occlusion, complex background and target motion, the tracking probability of the tracking system based on D-DSST and D-ECOHC is 90.01% and 93.66% respectively, and there is no target loss compared with the original algorithm; the average frame rate is 79.5 and 47.2fps respectively, which is 16.7 and 4.3fps higher than the original algorithm. The experimental results show that the above framework can track occluded maneuvering targets and meet the requirements of stability, reliability, high accuracy and good real-time performance.
[0146] like Fig.13 As shown, the present invention also provides a long-term target tracking system integrating depth information, including a laser rangefinder, a bus servo, and an upper position that are communicatively connected to the control system, and also includes a camera that is communicatively connected to the upper position;
[0147] The camera is used to collect images and send them to the upper camera position;
[0148] Laser rangefinder, used to measure the target depth information of the current frame and send it to the control system;
[0149] The bus servo controls the rotation of the laser rangefinder and camera according to the control system instructions;
[0150] in
[0151] The control system includes
[0152] An initialization unit, used to initialize a position-related filter, a scale-related filter, and a predicted target depth model;
[0153] The CA model unit predicts the target depth of the next frame based on the target depth information of the current frame measured by the laser rangefinder and the CA model; the specific method refers to the content in the above method.
[0154] Target position and scale estimation unit: According to the input new frame image, the position and scale of the target are estimated respectively by using the position-related filter and scale-related filter in the predicted target depth model.
[0155] Target position confirmation unit: According to the predicted target depth, a target confidence discriminant is constructed, the output of the position filter is discriminated, and the target position is determined according to the discriminant result. For the specific method, please refer to the content of the above method.
[0156] Target scale confirmation unit: Based on the predicted target depth, an adaptive scale factor is constructed and the target scale is determined; the specific method is referred to the content in the above method.
[0157] The command sending unit sends the target position determined by the target position confirmation unit and the target scale determined by the target scale confirmation unit to the bus servo respectively, and the bus servo adjusts the position of the laser rangefinder and the camera; it also decides whether to send a control command to the bus servo to control the laser rangefinder and the camera according to the judgment result of the target tracking judgment unit.
[0158] Target tracking judgment unit: determines whether the target tracking is completed. If not, the instruction sending unit sends an instruction to the camera to obtain a new frame image to restart target tracking. If completed, the work is terminated.
[0159] like Fig.14 As shown, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the long-term target tracking method integrating depth information in the above-mentioned embodiment, or, when executed by a processor, implements the long-term target tracking method integrating depth information in the above-mentioned embodiment.
[0160] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0161] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0162] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0163] The foregoing is merely a specific embodiment of the present invention, which enables those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features claimed herein.
Claims
1. Long-term target tracking method integrating depth information, It is characterized in that The following steps are involved: Step 1, initialize the position-related filter, scale-related filter and predicted target depth model; Step 2: Use a laser rangefinder to measure the target depth information of the current frame; Step 3: According to the target depth information of the current frame, based on the constant acceleration model, predict the target depth of the next frame; Step 4: According to the input new frame image, the position and scale of the target are estimated using the position correlation filter and the scale correlation filter; Step 5: According to the target depth predicted in step 3, a target confidence discriminant is constructed, and the reliability of the target position output by the position filter is discriminated according to the preset target position confidence threshold, and whether to update the target or correct the predicted position is determined according to the discrimination result; Step 6: Based on the predicted target depth, calculate the adaptive scale factor to obtain the predicted target scale, and find the scale that maximizes the scale-dependent filter response, which is the target scale. Step 7: According to the target position and scale determined in Step 5 and Step 6, the image acquisition device realigns the target to acquire a new frame of image; Step 8: Repeat steps 2 to 7 until the target tracking is completed.
2. The long-term target tracking method integrating depth information as claimed in claim 1, It is characterized in that The step 3 is specifically as follows: If the laser rangefinder measurement frequency is set higher than the target motion state change frequency, then within one measurement cycle, the third-order constant acceleration model of the target motion state is: Among them, ω(k) has a mean of zero and a variance of Gaussian white noise; The depth information and gimbal attitude angle of the target are measured with a period of T, including the pitch angle γ k and yaw angle θ k , using geometric relationships and motion equations to calculate the target's motion state A(k) at time k: Combine the discrete coordinate points to get the target path sequence S(k)∈S 1 ,…,S k ; Substitute A(k) in equation (2) into equation (1) to predict the target motion state A(k+1) at the next moment, and finally calculate the target depth D at the next moment pred , 3. The long-term target tracking method integrating depth information as claimed in claim 1, It is characterized in that In step 5, the reliability of the target position estimated in step 4 is judged according to the preset target position confidence threshold, and whether to update the target or correct the predicted position is specifically as follows: When there is an obstruction in front of the target, the depth output by the laser rangefinder is the depth of the obstruction d 2 , d 2 With target depth d 1 Compared with, there is a mutation value δ d =d 2 -d 1 ; According to the preset target position confidence threshold δ max , if δ d >δ max , that is, if the mutation value is greater than the threshold, the target is considered to be blocked and interfered, the target is not updated, and the position is corrected using the target depth predicted by the constant acceleration model; otherwise, the target is updated according to the maximum response of the position-related filter.
4. The long-term target tracking method integrating depth information as claimed in claim 3, It is characterized in that The step 5 comprises: Using acceleration and laser rangefinder noise factor Q sensor Dynamically adjust the confidence threshold δ max , the accelerations of the target in the x, y, and z directions at time k are Dynamically adjust the threshold δ max for:
5. The long-term target tracking method integrating depth information as claimed in claim 2, It is characterized in that Step 6 is as follows: The adaptive scale factor is calculated using the predicted target depth and multiplied by the target scale at time k to obtain the predicted target scale s at time k+1 pred , and at the target center, 17 scale samples {s -8 …s 0 …s 8 }, after filtering the scale samples, find the scale s that makes the scale correlation filter respond the largest max , which is the target scale.
6. The long-term target tracking method integrating depth information as claimed in claim 5, It is characterized in that The adaptive scale factor calculation process is as follows: The target predicted width and target true width are calculated according to the proportional relationship as follows: Where f is the focal length of the camera, u is the pixel size, D is the target depth, and the subscripts last, curr, and pred in the figure represent the previous frame, current frame, and predicted frame respectively; D ini and w ini are the initialization target depth and the initialization target image width, respectively. The target depth D predicted by the constant acceleration model pred Substituting the target actual width w calculated by formula (6) into formula (5), we get the predicted target width: Introducing the target width adaptive scaling factor A w , target height adaptive scale factor A h The calculation of is as follows: Where h represents the target height; is the scale factor limiting parameter; β=1.1; The target sample scale selection principle for target scale evaluation is: Among them, S=17, which is the scale sample.
7. Long-term target tracking system integrating depth information, It is characterized in that include Laser rangefinder, bus servo and boarding position connected to the control system; A camera connected to the upper camera station; The camera is used to collect images and send them to the upper camera position; Laser rangefinder, used to measure the target depth information of the current frame and send it to the control system; The bus servo controls the rotation of the laser rangefinder and camera according to the control system instructions; The control system includes An initialization unit, used to initialize a position-related filter, a scale-related filter, and a predicted target depth model; The constant acceleration model unit predicts the target depth of the next frame based on the target depth information of the current frame measured by the laser rangefinder and the constant acceleration model; The target position and scale estimation unit estimates the target position and target scale respectively by using the position correlation filter and scale correlation filter in the predicted target depth model according to the input new frame image; The target position confirmation unit constructs a target confidence discriminant according to the predicted target depth, discriminates the reliability of the target position output by the position filter according to a preset target position confidence threshold, and determines whether to update the target or correct the predicted position according to the discriminant result; Target scale confirmation unit: Based on the predicted target depth, the adaptive scale factor is calculated to obtain the predicted target scale, and the scale that makes the scale-related filter respond the largest is found, which is the target scale; The command sending unit sends the target position determined by the target position confirmation unit and the target scale determined by the target scale confirmation unit to the bus servo respectively, and the bus servo adjusts the position of the laser rangefinder and the camera; and decides whether to send a control command to the bus servo to control the laser rangefinder and the camera according to the determination result of the target tracking determination unit; Target tracking judgment unit: determines whether the target tracking is completed. If not, the instruction sending unit sends an instruction to the camera to obtain a new frame image to restart target tracking; if completed, the work is terminated.
8. Computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, It is characterized in that When the processor executes the computer program, the long-term target tracking method integrating depth information as described in any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium storing a computer program. It is characterized in that When the computer program is executed by a processor, the long-term target tracking method integrating depth information as claimed in any one of claims 1 to 6 is implemented.