A target tracking method and device based on target detection

By acquiring the target image and combining multiple sensor information for auxiliary correction, the spatial and temporal information of the target is obtained by using the feature analysis module to solve the problem of high hardware requirements for deep learning target tracking, and the effect of reducing hardware costs and improving tracking efficiency is achieved.

CN116246050BActive Publication Date: 2025-07-25NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211541342.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-02
Publication Date
2025-07-25
Estimated Expiration
2042-12-02

AI Technical Summary

Technical Problem

When the existing target tracking technology adopts deep learning, the hardware requirements are high, the cost-effectiveness of deployment is greatly reduced with the increase in computing power, and the traditional methods are poorly generalized and the deployment and debugging cost is high.

Method used

By acquiring the target image and performing auxiliary correction, combining multiple sensor information to supplement image information, the feature analysis module is used to obtain the target's spatial and temporal information, and combining the space-time context and feature matching to reduce the amount of algorithm calculation and reduce hardware requirements.

Benefits of technology

While meeting real-time and accuracy, the hardware cost is reduced and the stability and efficiency of target tracking are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116246050B_ABST
    Figure CN116246050B_ABST
Patent Text Reader

Abstract

This application relates to the field of target detection technology. Specifically, it relates to a target tracking method and device based on target detection, which can solve the problem that when implementing target detection using deep learning technology, the deep learning-based target detection technology itself has high hardware requirements, and the deployment cost performance drops significantly with the improvement of computing power. The method includes the following steps: obtaining a target image, and performing auxiliary correction on the target image to obtain target image information; based on the target image information, obtaining the spatial information and temporal information of the tracking target; based on the spatial information of the tracking target, obtaining the feature information of the tracking target according to the feature analysis module, and inputting the feature information, temporal information, and spatial information of the tracking target into the tracking resolver to obtain the tracking result of the tracking target.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of object detection. Specifically, it relates to an object tracking method and device based on object detection. Background Art

[0002] Currently, the main methods used by object tracking technologies to utilize detection information are trajectory prediction, extracting appearance features for modeling using convolutional neural networks, and associating and matching detection boxes and trackers in deepsort. The method of trajectory prediction has a low accuracy rate, and the selected object is prone to being lost and cannot be retrieved after being lost, and there are many cases where the selected object changes; the methods of convolutional neural network feature modeling and deepsort have higher requirements for device computing power. To ensure meeting real-time requirements, their requirements for the computing power of hardware increase greatly, and it is easy to use high-performance AI computing devices, significantly increasing the hardware cost.

[0003] In the actual application process of current object tracking algorithms, since the DFT (Detection Free Tracking) method requires manual calibration of the object in the initial frame image for subsequent detection and tracking of the object, the operation difficulty is high, and tasks cannot be processed for objects that appear in non-initial frames or disappear in intermediate frames. Therefore, in actual applications, DBT (Detection Based Tracking) is mostly used to obtain detection information of objects in the image before tracking, which can avoid the influence of human factors on the results and also reduce labor costs. However, traditional object detection algorithms have poor generalization, high deployment and debugging costs, large difficulties, and long time consumption. Therefore, it is best to use current mature deep learning technologies to implement object detection.

[0004] However, when using deep learning technologies to implement object detection, the object detection technology based on deep learning itself has relatively high requirements for hardware, and the deployment cost performance drops significantly with the improvement of computing power. Summary of the Invention

[0005] To solve the problem that when using deep learning technologies to implement object detection, the object detection technology based on deep learning itself has relatively high requirements for hardware, and the deployment cost performance drops significantly with the improvement of computing power, this application provides an object tracking method based on object detection:

[0006] According to one aspect of the embodiments of this application, an object tracking method based on object detection is provided, including the following steps:

[0007] Obtain a target image and perform auxiliary correction on the target image to obtain target image information. The target image is obtained by an image acquisition device collecting the task area. Auxiliary information is obtained through various sensors installed in the image acquisition device, and the target image is auxiliary corrected through the auxiliary information to supplement information and determine the target image information;

[0008] Based on the target image information, obtain the spatial information and time information of the tracking target. After the target image information is processed by an information processing module, it is input into a target detection module to obtain the spatial information of the tracking target, and the time corresponding to the target image information is the time information of the tracking target;

[0009] Based on the spatial information of the tracking target, obtain the feature information of the tracking target according to a feature analysis module, and input the feature information, time information, and spatial information of the tracking target into a tracking resolver to obtain the tracking result of the tracking target.

[0010] In some embodiments, the target tracking method includes a multi-target tracking mode and a single-target tracking mode. The default mode is selected as the multi-target tracking mode. When the single-target tracking mode needs to be selected, the tracking target is manually selected, or the target with the smallest tracking serial number or the highest detection confidence is default selected as the tracking target.

[0011] In some embodiments, in the step of obtaining a target image and performing auxiliary correction on the target image, the following steps are further included:

[0012] Continuously collect images of the task area through the image acquisition device to obtain a target image;

[0013] Obtain compensation information through other devices or relevant parameters preset before debugging;

[0014] Perform auxiliary correction on the target image through the compensation information to supplement information and determine the target image information.

[0015] In some embodiments, in the step of obtaining the spatial information and time information of the target based on the target image, after the target image is processed informatically by an information processing module and input into a target detection module to obtain the spatial information of the target, and the time information corresponding to the target image is the time information of the target, the following steps are further included:

[0016] According to the target image, obtain a first detection result through the target detection module;

[0017] According to the first detection result, obtain various auxiliary information through multiple sensors;

[0018] Based on the first detection result and according to various auxiliary information, determine the three-dimensional spatial position information of the tracking target.

[0019] In some embodiments, after the step of obtaining the first detection result according to the target image and through the target detection module, the following steps are further included:

[0020] If a variety of sensors are not provided on the image acquisition device, the first detection result is the three-dimensional spatial position information of the tracking target.

[0021] In some embodiments, in the step of obtaining the first detection result according to the target image and through the target detection module, the following steps are further included:

[0022] Identify the target image information to obtain the tracking target.

[0023] Detect the tracking target to obtain a target detection result, where the target detection result is the first detection result, and the first detection result includes the position, category, and confidence of the tracking target in the target image.

[0024] In some embodiments, in the step of obtaining the feature information of the tracking target according to the feature analysis module based on the spatial information of the tracking target, the following steps are further included:

[0025] Process the target image information according to the first detection result to obtain the original image information of the tracking target;

[0026] Obtain the feature information based on the original image information of the tracking target.

[0027] In some embodiments, the feature information includes color features, texture features, information entropy, and similarity.

[0028] In some embodiments, in the step of inputting the feature information, temporal information, and spatial information of the tracking target into a tracking resolver to obtain the tracking result of the tracking target, the following steps are further included:

[0029] Based on the feature information of the tracking target and the temporal information of the tracking target, establish a loss calculation mechanism;

[0030] Obtain the losses of different feature information according to the reference frame target and the comparison frame target, and establish a corresponding loss matrix through the loss calculation mechanism;

[0031] Based on the loss matrix, the temporal information, and the spatial information of the tracking target, judge and match the new and old targets to obtain the tracking result of the tracking target.

[0032] According to one aspect of the embodiments of the present application, a target tracking device based on target detection is provided, including:

[0033] An acquisition module: configured to obtain a target image, perform auxiliary correction on the target image to obtain target image information. The target image is acquired by an image acquisition device for a task area, and auxiliary information is obtained through a variety of sensors installed in the image acquisition device. The target image is assisted and corrected through the auxiliary information to supplement information and determine the target image information;

[0034] A processing module: configured to obtain the spatial information and temporal information of a tracking target based on the target image information. After the target image information is processed by an information processing module, it is input into a target detection module to obtain the spatial information of the tracking target, and the time corresponding to the target image information is the temporal information of the tracking target;

[0035] A tracking module: based on the spatial information of the tracking target, obtain the feature information of the tracking target according to a feature analysis module, and input the feature information, temporal information, and spatial information of the tracking target into a tracking resolver to obtain the tracking result of the tracking target.

[0036] Advantages of the present application: By obtaining a target image, performing auxiliary correction on the target image to supplement information, obtaining complete target image information, then processing the target image information to obtain the spatial information and temporal information of the tracking target, and then through a feature analysis module, according to the spatial information of the tracking target, obtaining various feature information of the tracking target, and then through a tracking resolver, according to various feature information of the tracking target, spatial information, and temporal information of the tracking target, obtaining the tracking result of the tracking target, realizing the tracking of the target according to relevant information, reducing the computational amount of the algorithm on the premise of meeting real-time performance and ensuring a certain accuracy, thereby reducing the requirements for hardware during deployment and reducing the cost during final deployment. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0038] Figure 1 Shows a system flow schematic diagram of a target tracking method provided by an embodiment of the present application;

[0039] Figure 2 Shows the schematic diagram of the system process for obtaining target image information provided by an embodiment of the present application;

[0040] Figure 3 Shows the schematic diagram of the system process for obtaining target time information and spatial information provided by an embodiment of the present application;

[0041] Figure 4 Shows the schematic diagram of the system process for obtaining the spatial information of a target provided by an embodiment of the present application;

[0042] Figure 5 Shows the schematic diagram of the system process for obtaining the tracking result of a tracking target provided by an embodiment of the present application;

[0043] Figure 6 Shows the schematic diagram of the system structure of the target tracking device provided by an embodiment of the present application;

[0044] Figure 7 Shows the schematic diagram of the method for calculating the three-dimensional spatial information of the image acquisition device and the target provided by an embodiment of the present application;

[0045] Figure 8 Shows the schematic diagram of the YOLOv6 network structure that can be used for this task provided by an embodiment of the present application;

[0046] Figure 9 Shows the schematic diagram of the backbone network EfficientRepBackbone of the YOLOv6 network provided by an embodiment of the present application;

[0047] Figure 10 Shows the schematic diagram of the utilization and calculation work of relevant information by the tracking solver provided by an embodiment of the present application;

[0048] Figure 11 Shows the schematic diagram of a hierarchical support vector machine structure combining multi-dimensional information provided by an embodiment of the present application;

[0049] Figure 12 Shows the schematic diagram of the support vector machine classification result using target features in a certain color channel provided by an embodiment of the present application;

[0050] Figure 13 Shows the schematic diagram of the training loss curve using methods such as transfer training and frozen training provided by an embodiment of the present application;

[0051] Figure 14 Shows the schematic diagram of the mAP curve during the training process provided by an embodiment of the present application;

[0052] Figure 15Shows the schematic diagram of the detection result of the vehicle target by the detection algorithm provided in an embodiment of the present application;

[0053] Figure 16 Shows the schematic diagram of the actual effect of multi-target tracking of human targets in traffic surveillance videos after applying the method of the present application provided in an embodiment of the present application;

[0054] Figure 17 Shows the schematic diagram of the actual effect of target and candidate target tracking in the single-target tracking mode provided in an embodiment of the present application.

[0055] Description of the drawings: 400: Target tracking device; 410: Acquisition module; 420: Processing module; 430: Tracking module. Detailed implementation manners

[0056] To make the purpose, implementation manners and advantages of the present application clearer, the following will clearly and completely describe the exemplary implementation manners of the present application with reference to the drawings in the exemplary embodiments of the present application. Obviously, the described exemplary embodiments are only a part of the embodiments of the present application, rather than all the embodiments.

[0057] It should be noted that the brief description of the terms in the present application is only for the convenience of understanding the following described implementation manners, rather than intending to limit the implementation manners of the present application. Unless otherwise specified, these terms should be understood in their ordinary and general meanings.

[0058] The terms "first", "second", "third", etc. in the description, claims and above-mentioned drawings of the present application are used to distinguish similar or homogeneous objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise noted. It should be understood that such terms can be interchanged under appropriate circumstances.

[0059] The terms "including" and "having" and any variations thereof are intended to cover but not exclude inclusion. For example, a product or device including a series of components does not necessarily have to be limited to all the clearly listed components, but may include other components not clearly listed or inherent to these products or devices.

[0060] At present, the main methods adopted by target tracking technology for the utilization of detection information are trajectory prediction, extracting appearance features by convolutional neural network for modeling, and associating and matching detection boxes and trackers by deepsort. The method of trajectory prediction has a low accuracy rate, and the selected target is prone to be lost and cannot be retrieved after being lost, and the situation where the selected target is changed often occurs; the methods of convolutional neural network feature modeling and deepsort have higher requirements for the computing power of the device. To ensure meeting the real-time requirements, the requirements for the computing power of the hardware are greatly improved, and it is easy to use high-performance Al computing devices, significantly increasing the hardware cost.

[0061] In the actual application process of current object tracking algorithms, since the DFT (Detection Free Tracking) method requires manual calibration of the object in the initial frame image for subsequent detection and tracking of the object, the operation difficulty is high, and tasks cannot be processed for objects that appear in non-initial frames or disappear in intermediate frames. Therefore, in practical applications, DBT (Detection Based Tracking) is mostly used to obtain the detection information of the object in the image before tracking, which can avoid the influence of human factors on the results and also reduce the labor cost. However, traditional object detection algorithms have poor generalization, high deployment and debugging costs, large difficulty, and long time consumption. Therefore, it is best to use the current mature deep learning technology to implement object detection.

[0062] However, when using deep learning technology to implement object detection, the object detection technology based on deep learning itself has relatively high hardware requirements, and the deployment cost performance drops significantly with the improvement of computing power.

[0063] Therefore, in view of the above problems, the present application proposes an object tracking method based on object detection. Considering the deployment cost and real-time issues of object tracking in practical applications in fields such as monitoring, the detected object is analyzed for features to obtain the feature information of the detected object, such as the color feature, texture feature, information entropy, similarity between new and old objects, etc. Combining the detection itself, relevant auxiliary devices or pre-built parameters can be used to obtain the object position information and its corresponding time information. By using the method of spatio-temporal context and object feature matching based on object detection, the detection results are fully utilized, and analysis is carried out from the spatio-temporal relationship and feature relationship of the object, reducing the hardware requirements of the algorithm and the system deployment cost.

[0064] The following combines Figure 1 to describe the object tracking method provided by some embodiments of the present application.

[0065] Figure 1 shows a schematic system flow diagram of the object tracking method provided by an embodiment of the present application.

[0066] As Figure 1 shown, it includes the following steps:

[0067] Step 100: Obtain a target image, and perform auxiliary correction on the target image to obtain target image information. The target image is obtained by an image acquisition device collecting the task area, and auxiliary information is obtained through various sensors installed in the image acquisition device. The target image is assisted and corrected through the auxiliary information to supplement the information and determine the target image information.

[0068] Among them, the acquisition device mainly acquires image information, and can obtain compensation information by means of other devices or relevant parameters preset before debugging. For example, a laser rangefinder with a motion base equipped with a microcontroller servo motor can be added to the camera to achieve rapid ranging with millimeter-level accuracy point-to-point within 40 meters at a cost of 100 yuan. When the acquisition device has ranging capabilities, it can measure the distance to the target according to the target detection results obtained by the target detection module, and combine the position of the acquisition device itself and the control signal of the servo motor to the laser rangefinder. At this time, the distance and angle information of the target relative to the shooting device can be obtained, and the position calculation of the target in the three-dimensional space can be realized.

[0069] Step 200: Based on the target image information, obtain the spatial information and time information of the tracking target. After the target image information is processed by the information processing module, it is input into the target detection module to obtain the spatial information of the tracking target, and the time corresponding to the target image information is the time information of the tracking target.

[0070] Among them, the information processing module mainly comprehensively manages the relevant information of each module, receives and sends the information of each module according to the task progress, receives the target image information sent by the acquisition device, and passes the target image information into the target detection module to obtain the detection result of the target. After the target detection module completes the detection of the target, it will return the detection result to the information processing module. The returned information mainly includes: the position, type, confidence, etc. of the target in the image. The obtained spatio-temporal information corresponding to the target is passed into the tracking module for subsequent use by the tracking resolver for spatio-temporal context information and target feature information to achieve the tracking of the target.

[0071] Step 300: Based on the spatial information of the tracking target, obtain the feature information of the tracking target according to the feature analysis module, and input the feature information, time information, and spatial information of the tracking target into the tracking resolver to obtain the tracking result of the tracking target.

[0072] Among them, the feature analysis module obtains the feature information of the tracking target by selecting the built-in feature extraction algorithm, and passes the feature information into the tracking analysis module. The feature analysis module consists of a variety of fast algorithms that can be used for image feature analysis, and finally generates a tracking target feature matrix Am×n, where m is the number of targets and n is the number of features extracted for each target. The spatio-temporal information corresponding to the tracking target is passed into the tracking module for subsequent use by the tracking resolver for spatio-temporal context information and target feature information to achieve the tracking of the target.

[0073] In some embodiments, the target tracking method includes a multi-target tracking mode and a single-target tracking mode. The default mode is the multi-target tracking mode. When the single-target tracking mode needs to be selected, the tracking target is manually selected, or the target with the smallest tracking serial number or the highest detection confidence is default selected as the tracking target.

[0074] Among them, generally, the multi-target tracking mode is used. If there are special requirements for a certain target, the single-target tracking mode can be switched. At this time, the loss judgment will use the feature information as the most important basis for tracking, and generate candidate targets corresponding to the tracking target. This situation can better ensure that the tracking target can still be retrieved and stably tracked after being lost for a long time.

[0075] Figure 2 The schematic diagram of the system process for obtaining target image information provided by an embodiment of the present application is shown. Figure 7 The schematic diagram of the method for calculating the three-dimensional space information of the image acquisition device and the target provided by an embodiment of the present application is shown.

[0076] In some embodiments, as Figure 2 shown, in the steps of obtaining the target image and performing auxiliary correction on the target image, the following steps are further included:

[0077] Step 110: Continuously collect images in the task area of the image acquisition device to obtain the target image.

[0078] Step 120: Obtain compensation information through other devices or relevant parameters preset before debugging.

[0079] Step 130: Perform auxiliary correction on the target image through the compensation information to supplement the information and determine the target image information.

[0080] Among them, as Figure 7 shown, R is the distance between the acquisition device point O and the target point P obtained by the ranging device. The azimuth angle θ and elevation angle φ between the acquisition device and the target can be obtained through the servo motor.

[0081] The calculation formula for the corresponding positions of the two in the space coordinate system is as follows:

[0082]

[0083] Among them, (x0, y0, z0) is the initial set coordinate of the acquisition device point O, and (x, y, z) is the coordinate of the calculated target point in the three-dimensional space.

[0084] Figure 3 The schematic diagram of the system process for obtaining the target time information and space information provided by an embodiment of the present application is shown.

[0085] In some embodiments, such as Figure 3 shown, in the step of obtaining the spatial information and temporal information of the target based on the target image, after the target image is processed informatization by the information processing module and input into the target detection module to obtain the spatial information of the target, the temporal information corresponding to the target image is the temporal information of the target, the following steps are further included:

[0086] Step 210: According to the target image, the information passes through the target detection module to obtain a first detection result.

[0087] Step 220: According to the first detection result, multiple sensors are used to obtain various auxiliary information.

[0088] Step 230: Based on the first detection result and according to various auxiliary information, determine the three-dimensional spatial position information of the tracking target.

[0089] Among them, if there is an auxiliary device, the information processing module will send the first detection result information back to the acquisition device. For example, at this time, the laser ranging device can perform point-to-point ranging for each target, and use the distance R, azimuth angle θ, and elevation angle φ between the tracking target and the acquisition device as auxiliary information and send it back to the information processing module, and the information processing module performs calculation to obtain the three-dimensional spatial position information of the tracking target.

[0090] In some embodiments, such as Figure 3 shown, after the step of obtaining the first detection result by passing the information through the target detection module according to the target image, the following steps are further included:

[0091] Step 240: If multiple sensors are not set on the image acquisition device, the first detection result is the three-dimensional spatial position information of the tracking target.

[0092] Among them, if there is no auxiliary device, the information processing module will directly use the position of the tracking target in the image, that is, the result of the target detection module, as the position information of the tracking target. For the corresponding tracking target, the information acquisition module will use the time of the system to stamp the tracking target with a time stamp to obtain the spatio-temporal information corresponding to the tracking target.

[0093] Figure 4 shows a schematic diagram of the system process for obtaining the spatial information of the target provided by an embodiment of the present application, Figure 8 shows a schematic diagram of the YOLOv6 network structure that can be used for this task provided by an embodiment of the present application, Figure 9 shows a schematic diagram of the backbone network EfficientRep Backbone of the YOLOv6 network provided by an embodiment of the present application.

[0094] In some embodiments, such as Figure 4 shown, in the step of obtaining the first detection result according to the target image through the target detection module, the following steps are further included:

[0095] Step 211: Identify the target image information to obtain a tracking target.

[0096] Step 212: Detect the tracking target to obtain a target detection result, where the target detection result is the first detection result, and the first detection result includes the position, category, and confidence of the tracking target in the target image.

[0097] Among them, the target detection module is used to find all interesting targets (objects) in the image and determine their category and location information. Since only the final output result needs to meet the requirements of the information processing module, it is only necessary to ensure that the format of the detection output result matches the interface of the information processing module, and the internal detection network module can be flexibly replaced according to the actual task requirements.

[0098] Such as Figure 8 and Figure 9 shown, YOLOv6 makes full use of the computing power of the hardware with EfficientRep Backbone, enhances the model's representation ability while significantly reducing the inference latency, achieves a good balance between accuracy and speed, constructs a more efficient decoupled Head by adopting a hybrid channel strategy, reduces the number of intermediate 3×3 convolutional layers to only one, further reduces the computational cost, and achieves a lower inference latency.

[0099] Figure 5 Shows a schematic diagram of the system process for obtaining the tracking result of the tracking target provided by an embodiment of the present application.

[0100] In some embodiments, such as Figure 5 shown, in the step of obtaining the feature information of the tracking target according to the feature analysis module based on the spatial information of the tracking target, the following steps are further included:

[0101] Step 310: Process the target image information according to the first detection result to obtain the original image information of the tracking target.

[0102] Step 320: Obtain the feature information based on the original image information of the tracking target.

[0103] Among them, screening and processing are carried out according to requirements such as the categories and tasks to be tracked. The detection results are processed with the original images obtained from the acquisition device, and the images corresponding to the tracking targets are cropped. The corresponding tracking target images and description information such as relevant serial numbers are input into the feature analysis module for subsequent feature calculation of the tracking targets by the feature analysis module.

[0104] In some embodiments, the feature information includes color features, texture features, information entropy, and similarity.

[0105] Among them, available feature information includes color features, texture features, information entropy, similarity, etc. Some of their calculation formulas are as follows:

[0106] (1) Color features

[0107]

[0108]

[0109]

[0110] Among them, i represents the i-th target, (x, 1) represents the first moment of the color component channel x, y represents the y-th pixel value of the corresponding color component channel, and M represents the number of pixels in the image.

[0111] Color channels that can be used include RGB, HSV, CMY, YUV, Lab, etc.

[0112] (2) Shape features

[0113] The calculation of image moments is the discretization of ordinary moments. For a pixel point with intensity f(x, y), the (p + q)-th moment can be defined as:

[0114]

[0115] Among them, C and R represent the number of columns and rows of the image respectively.

[0116] Taking the centroid of the target area as the center to construct the central moment, then the calculation of the moment is always the points in the target area relative to the centroid of the target area, and is independent of the position of the target area, that is, it has translational invariance. Its calculation formula is as follows:

[0117]

[0118] To offset the influence of scale changes on the central moments, the zero-order central moment u00 is used to normalize each order of central moments, construct scale invariance, and obtain the normalized central moment ηpq. The zero-order moment represents the mass (area) of the target region. If the scale of the target region changes, its zero-order central moment will also change accordingly, making the moments scale-invariant.

[0119]

[0120] Seven invariant moment groups can also be derived using the second- and third-order normalized central moments, which remain invariant under image translation, rotation, and scale changes.

[0121] (3) Texture features

[0122] The gray-level co-occurrence matrix starts from the pixels with gray level i in the N×N image f(x, y), in the direction of θ, and statistically calculates the probability P(i, j, δ, θ) that pixels with gray level j appear simultaneously at a distance of from i. Mathematically, it is expressed as:

[0123] P(i, j, δ, θ) = {[(x, y), (x + dx, y + dy)] | f(x, y) = i, f(x + dx, y + dy) = j}

[0124] Subsequently, the statistical attributes of the texture features can be obtained using it. The following lists some commonly used statistical attribute calculation formulas for texture features:

[0125] Mean: ∑i∑jp(i, j) * i;

[0126] Variance: ∑i∑jp(i, j) * (i - Mean)2;

[0127] Contrast: ∑i∑jp(i, j) * (i - j)2;

[0128] Entropy: ∑i∑jp(i, j) * lnp(i, j).

[0129] Figure 10 Shows a schematic diagram of the tracking solver using and solving relevant information in an embodiment of the present application. Figure 11 Shows a schematic diagram of a hierarchical support vector machine structure that combines multi-dimensional information in an embodiment of the present application.

[0130] In some embodiments, as Figure 5 shown, in the step of inputting the feature information, time information, and spatial information of the tracking target into the tracking solver to obtain the tracking result of the tracking target, the following steps are further included:

[0131] Step 330: Establish a loss calculation mechanism based on the feature information and time information of the tracking target.

[0132] Step 340: Obtain the losses of different feature information according to the reference frame target and the comparison frame target, and establish a corresponding loss matrix through the loss calculation mechanism.

[0133] Step 350: Based on the loss matrix, the time information and spatial information of the tracking target, judge and match the new and old targets, and obtain the tracking result of the tracking target.

[0134] Among them, as Figure 10 and Figure 11 shown, the tracking resolver is responsible for implementing the utilization and resolution of relevant information. The tracking resolver combines the spatio-temporal information, feature information, etc. of the tracking target obtained actually, uses the relevant information to establish a loss calculation mechanism, traverses and enumerates the reference frame target and the comparison frame target, obtains the color feature loss, texture feature loss, information entropy, etc., constructs the corresponding loss matrix, and based on the loss determination and the relevant information of the tracking target, uses machine learning or lightweight neural network methods to realize the judgment and matching of the new and old targets, and finally obtains the tracking result of the tracking target.

[0135] Next, the target tracking device provided by some embodiments of the present application will be described in conjunction with Figure 6 As shown in

[0136] As Figure 6 shown, the target tracking device 400 includes:

[0137] The acquisition module 410: is used to acquire a target image, and perform auxiliary correction on the target image to obtain target image information. The target image is acquired by an image acquisition device for the task area, and auxiliary information is acquired through various sensors installed in the image acquisition device. The target image is assisted and corrected through the auxiliary information to supplement the information and determine the target image information;

[0138] The processing module 420: is used to obtain the spatial information and time information of the tracking target based on the target image information. After the target image information is processed by the information processing module and input into the target detection module, the spatial information of the tracking target is obtained, and the time corresponding to the target image information is the time information of the tracking target;

[0139] The tracking module 430: based on the spatial information of the tracking target, obtains the feature information of the tracking target according to the feature analysis module, and inputs the feature information, time information and spatial information of the tracking target into the tracking resolver to obtain the tracking result of the tracking target.

[0140] It can be seen that, compared with the pure deep learning method, on the premise of meeting the actual requirements, the present invention can greatly reduce the amount of calculation, lower the requirements of the algorithm for hardware, and thus reduce the hardware cost in the actual deployment process.

[0141] The simulation experiment of the present invention was carried out under the hardware conditions of a central processing unit of 11th Gen InteI(R)Core(TM)i7-11800H, a graphics card: Nvida GeForce RTX 3050Ti, and a memory of 32G and the software environment of pytorch.

[0142] To verify the effectiveness of the algorithm in this paper, actual surveillance videos were used for testing. The tracking results were displayed in real time during the test. The experimental results show that, compared with related algorithms, the comprehensive effect of the actual surveillance video test in this paper is better, the requirements for hardware are lower, and the real-time performance and tracking effect are better under the same hardware conditions. Among them, when the detection algorithm in this paper adopts the YOLOv6-N network structure model, its mAP on the coCo dataset is improved by 8% compared with YOLOv5s, and the detection speed is increased by 85%. The final tracking test MOTA (Multi-Object Tracking Accuracy) decreased by 2.7% compared with YOLOv5+deepsort in the actual scenario, but the overall tracking speed increased by 125%.

[0143] To verify the effectiveness of the algorithm in this paper, the results of feature information utilization, the results of the detection algorithm, and the actual implementation effect diagram in the final actual application process are shown below. The algorithms for each link can be adjusted according to actual needs to enable it to have multiple capabilities such as detection, single-object tracking, and multi-object tracking.

[0144] Figure 12 It shows a schematic diagram of the support vector machine classification result using target features in a certain color channel provided by an embodiment of the present application.

[0145] As Figure 12 shown, through the combined utilization of different feature information, it is finally possible to realize the use of information matrices and loss calculations. In actual applications, relevant feature extraction methods should be selected according to the characteristics of target features to reduce the calculation amount and improve the stability of the algorithm.

[0146] This paper uses mAP (mean Average Precision) as the evaluation index for object detection. mAP is the average of the APs of various object categories, reflecting the average precision mean of the model; AP is the area under the PR curve, representing the average precision of this category. The PR curve is the Precision-Recall curve, and the calculation formulas for precision Precision and recall Recall are as follows:

[0147]

[0148] Among them, TP represents true positive, which refers to the number of positive samples correctly predicted as positive samples by the model; FP represents false positive, which refers to the number of negative samples mispredicted as positive samples by the model; FN represents false negative, which refers to the number of positive samples mispredicted as negative samples by the model.

[0149] The above network is iteratively trained using a detection dataset. The results of each round of training of the model are saved, and the loss value is visualized. After the model training is completed, it is tested using a test set to obtain the mAP value of the model. Figure 8 and Figure 9 are the loss value and mAP curve corresponding to the training respectively.

[0150] Figure 13 shows a schematic diagram of the training loss curve provided by an embodiment of the present application using methods such as transfer training and freezing training. Figure 14 shows a schematic diagram of the mAP curve during the training process provided by an embodiment of the present application.

[0151] As Figure 13 and Figure 14 shown, the starting generation is set to 300 generations. During the actual deployment process, in the face of changing scenarios, actual data is collected for supplementary training. Combining the above training methods can reduce the training cost and improve the detection effect at the same time.

[0152] Figure 15 shows a schematic diagram of the detection result of a vehicle target by the detection algorithm provided by an embodiment of the present application.

[0153] As Figure 15 shown, during the actual deployment process, the corresponding target can be selected according to the requirements, and the algorithm can be fine-tuned according to the scenario, giving full play to the generalization of deep learning technology and reducing the debugging cost.

[0154] Figure 16 shows a schematic diagram of the actual effect of multi-target tracking of human targets in a traffic surveillance video after applying the method of the present application provided by an embodiment of the present application.

[0155] As Figure 16 shown, in the case of multi-target tracking, the spatio-temporal context information of the target should be utilized more fully, and the result should be adjusted using the target feature information.

[0156] Figure 17 shows a schematic diagram of the actual effect of target and candidate target tracking in the single-target tracking mode provided by an embodiment of the present application.

[0157] As Figure 17As shown, in this mode, more attention should be paid to the design of the target feature utilization algorithm and the strategy of target information update. The re-capture ability of the target feature under the conditions of target occlusion, loss, etc. should be utilized as much as possible, and the stable tracking of a single target should be achieved by combining spatio-temporal context information. When necessary, multiple alternative targets can be listed to provide technical support for subsequent manual intervention.

[0158] For the sake of convenience in explanation, the above description has been made in conjunction with specific embodiments. However, the above discussion in some embodiments is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. According to the above teachings, various modifications and variations can be obtained. The selection and description of the above embodiments are for the purpose of better explaining the principles and practical applications, so that those skilled in the art can better use the embodiments and various different modified embodiments suitable for specific use considerations.

Claims

1. A target tracking method based on target detection, characterized in that The method includes the following steps: Obtain a target image and perform auxiliary correction on the target image to obtain target image information. The target image is obtained by an image acquisition device collecting the task area. Auxiliary information is obtained through a variety of sensors installed in the image acquisition device, and the target image is auxiliary corrected by the auxiliary information to supplement information and determine the target image information; Based on the target image information, obtain the spatial information and time information of the tracking target. After the target image information is processed by an information processing module, it is input into a target detection module to obtain the spatial information of the tracking target, and the time corresponding to the target image information is the time information of the tracking target; Based on the spatial information of the tracking target, obtain the feature information of the tracking target according to a feature analysis module, and input the feature information, time information, and spatial information of the tracking target into a tracking resolver to obtain the tracking result of the tracking target; In the step of obtaining the feature information of the tracking target according to the feature analysis module based on the spatial information of the tracking target, the following steps are further included: Process the target image information according to a first detection result to obtain the original image information of the tracking target. The first detection result is obtained by the target detection module through the target image, the spatial information, and the time information; Based on the original image information of the tracking target, obtain feature information. The feature information includes color features, texture features, information entropy, and similarity; In the step of inputting the feature information, time information, and spatial information of the tracking target into a tracking resolver to obtain the tracking result of the tracking target, the following steps are further included: Based on the feature information of the tracking target and the time information of the tracking target, establish a loss calculation mechanism; According to a reference frame target and a comparison frame target, obtain the losses of different feature information, and establish a corresponding loss matrix through the loss calculation mechanism; Based on the loss matrix, the time information, and the spatial information of the tracking target, judge and match new and old targets to obtain the tracking result of the tracking target.

2. The object tracking method based on object detection according to claim 1, wherein The target tracking method includes a multi-target tracking mode and a single-target tracking mode. The default mode is selected as the multi-target tracking mode. When the single-target tracking mode needs to be selected, the tracking target is manually selected, or the target with the smallest tracking serial number or the highest detection confidence is default selected as the tracking target.

3. The object tracking method based on object detection according to claim 1, characterized in that, In the step of obtaining a target image and performing auxiliary correction on the target image, the following steps are further included: Continuously collect images of the task area through an image acquisition device to obtain a target image; Obtain compensation information through other devices or relevant parameters preset before debugging; Perform auxiliary correction on the target image through the compensation information to supplement information and determine the target image information.

4. The object tracking method based on object detection according to claim 1, wherein In the step of obtaining the spatial information and temporal information of the target based on the target image, after the target image is processed by the information processing module and then input into the target detection module to obtain the spatial information of the target, the temporal information corresponding to the target image is the temporal information of the target, and the following steps are further included: Obtain a first detection result through the target detection module according to the target image, the spatial information, and the temporal information; Obtain various auxiliary information through multiple sensors according to the first detection result; Based on the first detection result and according to various auxiliary information, determine the three-dimensional spatial position information of the tracking target.

5. The object tracking method based on object detection according to claim 4, wherein After the step of obtaining a first detection result through the target detection module according to the target image, the spatial information, and the temporal information, the following steps are further included: If multiple sensors are not set on the image acquisition device, the first detection result is the three-dimensional spatial position information of the tracking target.

6. The object tracking method based on object detection according to claim 4, characterized in that, In the step of obtaining a first detection result through the target detection module according to the target image, the following steps are further included: Identify the target image information to obtain a tracking target; Detect the tracking target to obtain a target detection result, where the target detection result is the first detection result, and the first detection result includes the position, category, and confidence of the tracking target in the target image.

7. An object tracking device based on object detection, characterized in that, Include: Acquisition module: used to obtain a target image, perform auxiliary correction on the target image to obtain target image information. The target image is acquired by the image acquisition device for the task area, and auxiliary information is acquired through multiple sensors installed in the image acquisition device. The target image is assisted and corrected through the auxiliary information to supplement information and determine the target image information; Processing module: used to obtain the spatial information and temporal information of the tracking target based on the target image information. After the target image information is processed by the information processing module and then input into the target detection module, the spatial information of the tracking target is obtained, and the time corresponding to the target image information is the temporal information of the tracking target; Tracking module: Based on the spatial information of the tracking target, obtain the feature information of the tracking target according to the feature analysis module, and input the feature information, temporal information, and spatial information of the tracking target into the tracking resolver to obtain the tracking result of the tracking target; The tracking module is further configured as: Process the target image information according to the first detection result to obtain the original image information of the tracking target; the first detection result is obtained through the target detection module according to the target image, the spatial information, and the temporal information; Obtain feature information based on the original image information of the tracking target; The feature information includes color feature, texture feature, information entropy, and similarity; The tracking module is further configured as: Establish a loss calculation mechanism based on the feature information of the tracking target and the temporal information of the tracking target; Obtain the losses of different said feature information according to the reference frame target and the comparison frame target, and establish a corresponding loss matrix through the said loss calculation mechanism; Based on the said loss matrix, the time information and the space information of the tracking target, judge and match the old and new targets, and obtain the tracking result of the tracking target.

Citation Information

Patent Citations

  • Target tracking method and device

    CN108269269A

  • Target detection and tracking method and device, storage medium and electronic device

    CN112541395A