A lightweight object detection and tracking method and system for micro and small robots

By combining the lightweight object detection and tracking method with YOLOv4-tiny and KCF core-related filtering algorithms, the problem of target tracking accuracy and computational complexity in micro-robots is solved, and efficient target tracking under low computing power conditions is achieved.

CN117291950BActive Publication Date: 2025-07-11HARBIN INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311093365.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-29
Publication Date
2025-07-11
Estimated Expiration
2043-08-29

AI Technical Summary

Technical Problem

The prior art is difficult to reduce computational complexity while improving target tracking accuracy, especially when computing power is limited in micro-robot systems.

Method used

Combining the YOLOv4-tiny object detection algorithm and the KCF core-related filtering algorithm, the object detection is performed through the YOLOv4-tiny model, the category and position information of the initial frame are obtained, and then the KCF core-related filtering algorithm is used for tracking, and the YOLOv4-tiny model is called for detection and matching the tracking results at a fixed frame interval to correct the cumulative error.

Benefits of technology

While reducing the computing power requirements, the accuracy and stability of target tracking are improved, and the calculation amount is reduced. It is suitable for micro robot systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117291950B_ABST
    Figure CN117291950B_ABST
Patent Text Reader

Abstract

The present invention discloses a lightweight object detection and tracking method and system for micro and small robots, which relates to the technical field of object detection and tracking. The technical key points of the present invention include: obtaining a plurality of consecutive video frame sequences; using a pre-trained YOLOv4-tiny model for object detection to obtain the object category and position information in the initial frame containing the object; in subsequent frames, using the KCF kernel correlation filtering algorithm to track the detected object; and further correcting the cumulative error by performing object detection at fixed frame intervals. The present invention can greatly reduce the computing power requirement for object tracking and improve the accuracy and stability of object tracking to a certain extent. The present invention is conducive to being deployed on micro and small robots to realize an autonomous perception system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target detection and tracking, and in particular to a lightweight target detection and tracking method and system for a micro robot. Background Art

[0002] The detection of moving targets is to extract the moving foreground of interest from the image sequence, which is the basis for the subsequent recognition and tracking of moving targets. The quality of the moving foreground detection affects the accuracy of the recognition and tracking of moving targets. Tracking is the process of matching and identifying the effective features of the target one or several frames apart on the basis of detection to obtain the effective position information of the target.

[0003] The moving target detection algorithms in static scenes are mainly divided into two categories: based on motion information and based on feature information. The moving target detection algorithms based on motion information mainly include two-frame difference method, background difference method and optical flow method. After the rise of artificial neural networks, some scholars applied image-specific information to detection, which has achieved good application results. In monitoring scenes, due to changes in the external environment, many factors will interfere with the tracking of moving targets. Common moving target tracking algorithms include tracking based on feature points and feature regions. Feature points are specific identification points of moving targets, which can track different identification targets. Moving target tracking based on regions is usually divided into single target tracking and multi-target tracking. Single target tracking usually manually selects the target template, while manual selection of multi-target tracking is not practical and usually depends on the situation of the moving target. The difficulty of studying region-based target tracking algorithms lies in reducing the complexity of calculation while improving the accuracy of target tracking.

[0004] In summary, how to improve target tracking accuracy while reducing computational complexity needs to be solved urgently. Summary of the invention

[0005] To this end, the present invention proposes a lightweight target detection and tracking method and system for a micro robot, in an effort to solve or at least alleviate at least one of the above problems.

[0006] According to one aspect of the present invention, a lightweight target detection and tracking method for a micro robot is proposed, the method comprising the following steps:

[0007] Step 1: Obtain multiple continuous video frame sequences;

[0008] Step 2: Use the pre-trained YOLOv4-tiny model to detect the target and obtain the target category and location information in the initial frame containing the target;

[0009] Step 3: In subsequent frames, the KCF kernel correlation filter algorithm is used to track the detected target.

[0010] Further, the structure of the YOLOv4-tiny model described in step two includes a feature extraction module, a feature fusion module, and a detection and prediction head module. The feature extraction module includes multiple convolutional layers and multiple pooling layers for performing multi-dimensional feature extraction on the input video frame image. The feature fusion module is used to fuse features of different dimensions, depths, and scales using a feature pyramid. The detection and prediction head module is used to identify and classify the fused features and finally output the target category and location information.

[0011] Further, the specific steps of step three include:

[0012] Initialize the KCF kernel correlation filtering algorithm according to the target category and location information. Initialization means training the target tracking classifier.

[0013] Use the initialized target tracking classifier to track the target in subsequent frames. The tracking formula is as follows:

[0014]

[0015] Among them, represents the target tracking result; represents the correlation matrix between the current image sample and each training sample; represents the key parameters of the target tracking classifier obtained through training;

[0016] Perform the key parameters of the target tracking classifier and the target observation template Update, and the update formula is as follows:

[0017]

[0018]

[0019] Among them, γ represents the template update step size.

[0020] Further, the process of initializing the KCF kernel correlation filtering algorithm according to the target category and location information in step three includes: performing a high-dimensional mapping on the initial frame image containing the target; performing a Fourier transform on the high-dimensional mapped image data; using the transformed data to train the target tracking classifier to obtain the initialized kernel filter, and thus completing the initialization process. Among them, the initialization formula is:

[0021]

[0022] In the formula, represents the correlation matrix of each training sample, It represents the training label, and λ represents the regularization parameter.

[0023] Furthermore, after tracking the target for multiple frames using the KCF kernel correlation filtering algorithm, a pre-trained YOLOv4-tiny model is called to perform target detection on the next frame of the image; a relevant image matching algorithm is adopted to perform similarity matching on the tracking data and the detection data; the accuracy of the target tracking result is verified according to the matching result, and if it is inaccurate, the training target tracking classifier is re-initialized.

[0024] Furthermore, the process of performing similarity matching on the tracking data and the detection data using the relevant image matching algorithm includes: calculating the mean hash similarity, difference hash similarity, and perceptual hash similarity of the tracking data and the detection data, and taking the mean of the three hash similarities as the similarity matching value.

[0025] According to another aspect of the present invention, a lightweight target detection and tracking system for a micro and small robot is proposed, and the system includes:

[0026] An image acquisition module configured to acquire a plurality of consecutive video frame sequences;

[0027] A target detection module configured to perform target detection using a pre-trained YOLOv4-tiny model to obtain the target category and position information in the initial frame containing the target;

[0028] A target tracking module configured to, in subsequent frames, track the detected target using the KCF kernel correlation filtering algorithm.

[0029] Furthermore, the structure of the YOLOv4-tiny model in the target detection module includes a feature extraction module, a feature fusion module, and a detection prediction head module, wherein the feature extraction module includes a plurality of convolutional layers and a plurality of pooling layers for performing multi-dimensional feature extraction on the input video frame image; the feature fusion module is used to fuse features of different dimensions, depths, and scales using a feature pyramid; the detection prediction head module is used to identify and classify the fused features, and finally output the target category and position information.

[0030] Furthermore, the specific process of tracking the detected target using the KCF kernel correlation filtering algorithm in the target tracking module in subsequent frames includes:

[0031] Initialize the KCF kernel correlation filtering algorithm according to the target category and location information. Initialization means training the target tracking classifier. The initialization process includes: performing high-dimensional mapping on the initial frame image containing the target; performing Fourier transform on the high-dimensional mapped image data; training the target tracking classifier using the transformed data to obtain the initialized kernel filter, thus completing the initialization process. Among them, the initialization formula is:

[0032]

[0033] In the formula, represents the correlation matrix of each training sample, represents the training label, and λ represents the regularization parameter;

[0034] Use the initialized target tracking classifier to track the target in subsequent frames. The tracking formula is as follows:

[0035]

[0036] Among them, represents the target tracking result; represents the correlation matrix between the current image sample and each training sample; represents the key parameters of the target tracking classifier obtained through training;

[0037] Perform the key parameters of the target tracking classifier and the target observation template

[0038]

[0039]

[0040] where γ represents the template update step size.

[0041] Furthermore, the system further includes a correction module. The correction module is configured to, after using the KCF kernel correlation filtering algorithm to track the target for multiple frames, call the pre-trained YOLOv4-tiny model to perform target detection on the next frame image; use the relevant image matching algorithm to perform similarity matching on the tracking data and the detection data; verify the accuracy of the target tracking result according to the matching result. If it is inaccurate, re-initialize and train the target tracking classifier. Among them, the process of using the relevant image matching algorithm to perform similarity matching on the tracking data and the detection data includes: calculating the mean hash similarity, difference hash similarity, and perceptual hash similarity of the tracking data and the detection data, and taking the mean of the three hash similarities as the similarity matching value.

[0042] The beneficial technical effects of the present invention are:

[0043] Compared with traditional target detection and tracking algorithms that are either purely for detection or purely for tracking, the lightweight target detection and tracking method that combines the YOLOv4-tiny target detection algorithm and the KCF target tracking algorithm has the following advantages: 1) Compared with traditional pure detection algorithms, by integrating the KCF target tracking algorithm, the target detection algorithm can perform detections every 10-20 frames, which can greatly reduce the computing power requirements for target tracking; 2) Using the target tracking algorithm can avoid the loss of the target caused by missed detections due to the accuracy of the detection algorithm, and to a certain extent improve the accuracy and stability of target tracking; 3) Compared with pure target tracking algorithms, integrating the YOLOv4-tiny neural network avoids the manual initialization of the target tracking algorithm and improves the autonomy and intelligence of algorithm application; 4) Since the target tracking algorithm only tracks pixels and does not involve specific feature information, long-term tracking is prone to cumulative errors resulting in tracking failure. Detecting the target at fixed frame intervals corrects the cumulative errors and improves the tracking accuracy.

[0044] The present invention can, to a certain extent, improve the accuracy of target tracking while reducing the computing power requirements, which is conducive to deploying it on micro and small robots to achieve an autonomous perception system. Description of the Drawings

[0045] The present invention can be better understood by referring to the description given below in conjunction with the accompanying drawings. The accompanying drawings, together with the following detailed description, are included in this specification and form a part of this specification, and are used to further illustrate the preferred embodiments of the present invention and explain the principles and advantages of the present invention.

[0046] Figure 1 is a flowchart of the lightweight target detection and tracking method for micro and small robots according to an embodiment of the present invention.

[0047] Figure 2 is a network structure diagram of the YOLOv4-tiny neural network in an embodiment of the present invention.

[0048] Figure 3 is an example diagram of target detection results in an embodiment of the present invention.

[0049] Figure 4 is an example diagram of target tracking results in a simple scenario in an embodiment of the present invention.

[0050] Figure 5 is an example diagram of target tracking results in a complex scenario in an embodiment of the present invention. Detailed Embodiments

[0051] To enable those skilled in the art to better understand the solution of the present invention, the exemplary embodiments or examples of the present invention will be described below in conjunction with the accompanying drawings. Obviously, the described embodiments or examples are only a part of the embodiments or examples of the present invention, rather than all of them. All other embodiments or examples obtained by those of ordinary skill in the art based on the embodiments or examples in the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0052] To achieve low-computing-power and efficient target tracking for micro-robots, the present invention proposes a target detection and tracking method that combines the YOLOv4-tiny target detection algorithm and the KCF target tracking algorithm. This method first uses the YOLOv4-tiny target detection algorithm to obtain the category and position information of the target, and then uses the KCF algorithm to track each target through the kernel correlation filtering algorithm. The YOLOv4-tiny target detection algorithm is used for target detection every fixed number of frames (set to 20 frames in simple scenarios and 10 frames in complex scenarios), and it is matched with the tracked target and the deviation is compared. If the tracking deviation exceeds the predetermined limit, the YOLOv4-tiny target detection algorithm is used again to initialize and obtain accurate target position and category information; when the target tracking is lost detected by the KCF kernel correlation filtering tracking method, the YOLOv4-tiny target detection algorithm is used to re-initialize; the above steps are repeated until the task is completed.

[0053] The first embodiment of the present invention proposes a lightweight target detection and tracking method for micro-robots, as Figure 1 shown, and this method mainly includes the following steps:

[0054] Step 1: Obtain a sequence of multiple consecutive video frames;

[0055] Step 2: Use the pre-trained YOLOv4-tiny model for target detection to obtain the target category and position information in the initial frame containing the target;

[0056] Step 3: In subsequent frames, use the KCF kernel correlation filtering algorithm to track the detected target.

[0057] According to the embodiment of the present invention, using the pre-trained YOLOv4-tiny model for target detection in Step 2, the main purpose is to obtain the category and position information of the target through target detection, so as to prepare for the initialization of subsequent target tracking. The specific network structure of the YOLOv4-tiny network is as Figure 2 shown.

[0058] The YOLOv4-tiny algorithm is a typical deep learning algorithm. Therefore, it is first necessary to collect data, annotate the data, and use the annotated data to train qualified detection network weights, and then import the weights into relevant micro-robots or the built simulation data platform. Then, the image data obtained by the micro-robot camera is input into the YOLOv4-tiny neural network, and multi-dimensional feature extraction of the picture data is achieved through multiple convolutional layers and pooling layers in the backbone network (i.e., the CSPDarknet-tiny network module in Figure 2 ). In the YOLOv4-tiny neural network, the Leaky Relu function is introduced to replace the Relu function and the structure of the residual network, reducing the training difficulty of the entire network and improving the efficiency and accuracy of feature extraction. After feature extraction, the features of different dimensions, depths, and scales are fused through the feature pyramid, that is, entering the feature fusion module (i.e., the FPN module in Figure 2 ), which will improve the detection accuracy of the network and the detection ability for targets of different scales. Finally, through the detection prediction head (i.e., the YOLO_Head module in Figure 2 ), the fused features extracted are recognized and classified, and finally the category and position information of the target are output.

[0059] The KCF object tracking algorithm is a kernel correlation filtering object tracking algorithm. The KCF algorithm cleverly transforms the object tracking problem into a foreground and background classification problem in the image, introduces the method of ridge regression to train the tracker, and cleverly avoids a large amount of operations caused by dense sampling through the circulant matrix, enabling the KCF algorithm to run efficiently and in real time on the edge processor. And when performing regression training, a kernelization method is introduced, and the accuracy of foreground and background classification is improved through the kernel function, improving the stability of object tracking. The KCF algorithm has a high accuracy and high efficiency in object tracking.

[0060] First, the KCF algorithm is initialized using the image data obtained by the micro-robot camera and the object category and position information obtained by the YOLOv4-tiny algorithm for object detection. Initialization means using the real object position obtained by object detection as the foreground and the rest as the background to train the object tracking classifier. The specific process is as follows:

[0061] Perform a high-dimensional mapping on the obtained image data; perform a Fourier transform on the data after high-dimensional mapping; use the transformed data to quickly train the object tracking classifier, and obtain the initialized kernel filter to complete the initialization process. The core formula for initialization is:

[0062]

[0063] Among them, is the key parameter of the trained object tracking classifier, is the correlation matrix of each training sample, is the training label, and λ is the regularization parameter.

[0064] Then, use the initialized filter to track the next frame of image obtained by the micro robot, that is:

[0065]

[0066] Among them, is the object tracking result, is the key parameter of the trained object tracking classifier, is the correlation matrix between the current image sample and each training sample.

[0067] Finally, update the filter coefficient α and the target observation template x, that is:

[0068]

[0069]

[0070] Among them, γ represents the template update step size.

[0071] In this embodiment, preferably, it further includes step four: call the object detection algorithm once every several frames to check and correct the results of the object tracking algorithm. After the object tracking algorithm tracks the image data obtained by the micro robot camera for several frames, it is necessary to check the accuracy of the object detection results to ensure long-term accurate and stable tracking of the object.

[0072] First, still use the KCF kernel correlation filtering tracking algorithm to track the object in this frame. Then, perform object detection on this frame through the object detection algorithm. Then, use the relevant image matching algorithm to perform similarity matching on the tracking data and the detection data. Here, the average of the mean hash similarity (aHash), difference hash similarity (dHash), and perceptual hash similarity (pHash) of the two objects is used as the basis for similarity matching. Check the accuracy of the object detection and object tracking results after matching. If it is inaccurate, re-initialize the classifier. In fact, during the tracking process, re-initialization is not desired because it will discard a large number of relatively accurate previous tracking training results.

[0073] The specific calculation steps of the mean hash fingerprint are as follows:

[0074] 1) Resize the image: To retain the structure and remove the differences in details and sizes, uniformly resize the image to 8*8 size, 64 pixels;

[0075] 2) Convert the original color image into a grayscale image;

[0076] 3) Calculate the average value of the grayscale image pixels;

[0077] 4) Compare the pixel grayscale values, traverse each pixel of the grayscale image, if it is greater than the average value, record it as 1, otherwise as 0, that is:

[0078]

[0079] 5) Obtain the 64-bit mean hash fingerprint information.

[0080] The specific calculation steps of the difference hash fingerprint are as follows:

[0081] 1) Resize the image: To retain the structure and remove the differences in details and sizes, uniformly resize the image to a size of 9 * 8, 72 pixels;

[0082] 2) Convert the original color image into a grayscale image;

[0083] 3) Different from the calculation method of the mean hash fingerprint, for the difference hash fingerprint, directly compare the pixel values of the left and right pixels here. If the previous pixel is greater than the latter pixel, record it as 1, otherwise as 0. The expression is shown in the following formula. For this reason, the original 9 * 8 pixel grid values become 8 * 8 pixel grid values.

[0084]

[0085] 4) Expand the 8 * 8 grid values into 64-bit difference hash fingerprint information.

[0086] The specific calculation steps of the perceptual hash fingerprint are as follows:

[0087] 1) Resize the image: To retain the structure and remove the differences in details and sizes, uniformly resize the image to a size of 32 * 32;

[0088] 2) Convert the original color image into a grayscale image;

[0089] 3) Calculate the DCT (Discrete Cosine Transform). The DCT transform can convert the image into a set of different frequency components;

[0090]

[0091] where: k1, k2 = 0,..., N - 1;

[0092] 4) For the pixel values of the original image of 32 * 32, only retain the pixel positions of the upper left 8 * 8, which represent the low-frequency features of the image and have an advantageous significance for image matching.

[0093] 5) Calculate the average value of the frequency domain image.

[0094] 6) Compare the pixel frequency domain values, traverse each pixel of the grayscale image, record it as 1 if it is greater than the average value, otherwise 0;

[0095]

[0096] 7) Obtain 64-bit perceptual hash fingerprint information.

[0097] Calculate the mean hash similarity, difference hash similarity, and perceptual hash similarity between the detected target area and the tracked target area, and use their average value as the similarity between the detection and tracking targets;

[0098] The similarity calculation between each type of hash fingerprint adopts the Hamming distance, that is, for the hash fingerprints of two images a = a1, a2,..., a 64 and b = b1, b2,..., b 64 Then the similarity between the two images is:

[0099]

[0100] Through the above formula, calculate the Hamming distances between the mean hash fingerprint, difference hash fingerprint, and perceptual hash fingerprint between two images as similarities, and use the average value of the three fingerprint similarities as the final similarity between the two images, that is:

[0101]

[0102] Use the average value of the three fingerprint similarities to perform similarity matching on the detection results and tracking results. For any one detection result, take the one with the largest fingerprint similarity in the tracking results, that is, d Avg The smallest one is used as the matching result to obtain the similarity matching result for tracking accuracy verification.

[0103]

[0104] According to the matched object detection and object tracking results, perform accuracy verification. The accuracy verification is judged according to the threshold set by the Euler distance between the center of the object detection result and the center of the object tracking result. The calculation steps are as follows:

[0105] 1) Calculate the Euler distance between the object detection center and the object tracking center;

[0106]

[0107] 2) Judge whether to re-initialize the classifier according to the calculated Euler distance

[0108]

[0109] In this embodiment, the threshold d thrSet to 10 pixels.

[0110] It should be noted that if the tracking fails, it needs to be re-initialized. That is, when a scene of target tracking failure occurs, the YOLOv4-tiny algorithm is used to re-detect the target in the next frame of the image, and the KCF target tracking algorithm is re-initialized.

[0111] Furthermore, the technical effects of the present invention are verified through experiments.

[0112] The experiment selects the detection and recognition of micro-unmanned aerial vehicles in a simple environment and a complex environment as the experimental background, and uses the acquisition results of a distributed bionic lens flexible sensing system that can be used for micro-robots as input data to test the accuracy, stability, and real-time requirements of the algorithm. The simulation test software environment is Windows 11+python3.7.10+opencv-python4.5.3.56+opencv-contrib-python 3.4.13.47, and the hardware environment is Intel(R)Core(TMi7-10870H CPU+16.0GB RAM+NVDIA GeForce GTX 1650Ti.

[0113] The experiment first conducts the target detection experiment of the YOLOv4-tiny model, and the experimental results are as Figure 3 shown. It can be seen that whether it is a simple background (see Figure 3 left) or a complex background (see Figure 3 right), the YOLOv4-tiny target detection algorithm can provide accurate position information of the target in the picture with a high confidence level, providing strong support for the initialization of the KCF target tracking algorithm.

[0114] Then, using the results of the YOLOv4-tiny target detection network, the KCF target tracking algorithm is initialized. After the initialization of the KCF target tracking algorithm, the KCF target tracking algorithm is used to detect and track multiple targets in the subsequent video stream. At the same time, every few frames (20 frames in a simple scene and 10 frames in a complex scene), the target detection algorithm is called. If the positioning result between the detection result and the tracking result, that is, the position and size error of the positioning box, is less than the set threshold (set to 10 pixel values here), it is considered that the verification is successful and the target tracking continues. If the positioning error between the detection result and the tracking result is greater than the set threshold, it is considered that there is a large cumulative error in the tracking, so the target detection algorithm needs to be used to re-initialize the target tracking algorithm. If the target is lost during the target tracking, the target detection algorithm also needs to be called to re-initialize the target tracking algorithm. Finally, the target tracking results in the simple scene and the complex scene are as Figure 4 and Figure 5. It can be seen that the target tracking algorithm can track multiple targets in real time, stably and reliably in both simple and complex scenarios.

[0115] Finally, the difference in computational complexity between the lightweight target detection and tracking method for micro-robots that combines the YOLOv4-tiny target detection algorithm and the KCF target tracking algorithm and the pure YOLOv4-tiny target detection method when tracking targets is shown in Table 1. The frame rates of the method of the present invention and the traditional pure YOLOv4-tiny target detection method are compared in simple and complex scenarios respectively. It can be seen that the method of the present invention can nearly double the detection and tracking efficiency, can significantly improve the frame rate of accurately tracking targets under the same hardware conditions and software environment, greatly reduce the computational complexity, and make it more suitable for the low-computing-power and low-power-consumption platform of micro-robots.

[0116] Table 1

[0117]

[0118]

[0119] A second embodiment of the present invention proposes a lightweight target detection and tracking system for micro-robots, which includes:

[0120] An image acquisition module configured to acquire a plurality of consecutive video frame sequences;

[0121] A target detection module configured to perform target detection using a pre-trained YOLOv4-tiny model to obtain target category and position information in an initial frame containing the target;

[0122] A target tracking module configured to track the detected targets using the KCF kernel correlation filtering algorithm in subsequent frames.

[0123] In this embodiment, preferably, the structure of the YOLOv4-tiny model in the target detection module includes a feature extraction module, a feature fusion module, and a detection prediction head module, wherein the feature extraction module includes a plurality of convolutional layers and a plurality of pooling layers for performing multi-dimensional feature extraction on the input video frame image; the feature fusion module is used to fuse features of different dimensions, depths, and scales using a feature pyramid; the detection prediction head module is used to identify and classify the fused features and finally output the target category and position information.

[0124] In this embodiment, preferably, the specific process of using the KCF kernel correlation filtering algorithm to track the detected targets in subsequent frames in the target tracking module includes:

[0125] Initialize the KCF kernel correlation filtering algorithm according to the target category and location information. Initialization means training the target tracking classifier. The initialization process includes: performing a high-dimensional mapping on the initial frame image containing the target; performing a Fourier transform on the image data after high-dimensional mapping; using the transformed data to train the target tracking classifier, and obtaining the initialized kernel filter to complete the initialization process. Among them, the initialization formula is:

[0126]

[0127] In the formula, represents the correlation matrix of each training sample, represents the training label, and λ represents the regularization parameter;

[0128] Use the initialized target tracking classifier to track the target in subsequent frames. The tracking formula is as follows:

[0129]

[0130] Among them, represents the target tracking result; represents the correlation matrix between the current image sample and each training sample; represents the key parameters of the target tracking classifier obtained through training;

[0131] Perform the key parameters of the target tracking classifier and the target observation template The update formula is as follows:

[0132]

[0133]

[0134] Among them, γ represents the template update step size.

[0135] In this embodiment, preferably, the system further includes a correction module. The correction module is configured to, after using the KCF kernel correlation filtering algorithm to track the target for multiple frames, call the pre-trained YOLOv4-tiny model to perform target detection on the next frame image; use the correlation image matching algorithm to perform similarity matching on the tracking data and the detection data; verify the accuracy of the target tracking result according to the matching result. If it is inaccurate, re-initialize and train the target tracking classifier. Among them, the process of using the correlation image matching algorithm to perform similarity matching on the tracking data and the detection data includes: calculating the mean hash similarity, difference hash similarity, and perceptual hash similarity of the tracking data and the detection data, and taking the mean of the three hash similarities as the similarity matching value.

[0136] The function of a lightweight target detection and tracking system for a micro-miniature robot according to an embodiment of the present invention can be illustrated by the aforementioned lightweight target detection and tracking method for a micro-miniature robot. Therefore, for the parts not detailed in the system embodiment, reference can be made to the above method embodiment, which will not be elaborated herein.

[0137] Although the present invention has been described based on a limited number of embodiments, those skilled in the art in this technical field will understand that other embodiments can be envisioned within the scope of the present invention thus described. For the scope of the present invention, the disclosure made herein is illustrative rather than restrictive, and the scope of the present invention is defined by the appended claims.

Claims

1. A lightweight object detection and tracking method for micro-robots, characterized in that, It includes the following steps: Step 1, obtain multiple consecutive video frame sequences; Step 2, use the pre-trained YOLOv4-tiny model for object detection to obtain the object category and location information in the initial frame containing the object; Step 3, in subsequent frames, use the KCF kernel correlation filtering algorithm to track the detected object; the specific steps include: Initialize the KCF kernel correlation filtering algorithm according to the object category and location information, and the initialization is to train the object tracking classifier; specifically: perform a high-dimensional mapping on the initial frame image containing the object; perform a Fourier transform on the high-dimensional mapped image data; use the transformed data to train the object tracking classifier to obtain the initialized kernel filter, and the initialization process is completed; the initialization formula is: In the formula, represents the correlation matrix of each training sample, represents the training label, and λ represents the regularization parameter; Use the initialized object tracking classifier to track the object in subsequent frames, and the tracking formula is as follows: In the formula, represents the target tracking result; represents the correlation matrix between the current image sample and each training sample; represents the key parameters of the target tracking classifier obtained through training; Key parameters of the target tracking classifier and the target observation template are updated, and the update formula is as follows: where γ represents the template update step size.

2. The lightweight object detection and tracking method for micro-robots according to claim 1, wherein, The structure of the YOLOv4-tiny model in Step 2 includes a feature extraction module, a feature fusion module, and a detection prediction head module. The feature extraction module includes multiple convolutional layers and multiple pooling layers for performing multi-dimensional feature extraction on the input video frame image; the feature fusion module is used to fuse features of different dimensions, depths, and scales using a feature pyramid; the detection prediction head module is used to identify and classify the fused features, and finally output the object category and location information.

3. The lightweight object detection and tracking method for micro-robots according to claim 1, characterized in that After tracking the object using the KCF kernel correlation filtering algorithm for multiple frames, call the pre-trained YOLOv4-tiny model to perform object detection on the next frame image; use a relevant image matching algorithm to perform similarity matching on the tracking data and the detection data; verify the accuracy of the object tracking result according to the matching result, and if it is inaccurate, re-initialize and train the object tracking classifier.

4. The lightweight object detection and tracking method for micro-robots according to claim 3, characterized in that, The process of performing similarity matching on the tracking data and the detection data using a relevant image matching algorithm includes: calculating the mean hash similarity, difference hash similarity, and perceptual hash similarity of the tracking data and the detection data, and taking the mean of the three hash similarities as the similarity matching value.

5. A lightweight object detection and tracking system for micro-robots, characterized in that, It includes: An image acquisition module configured to obtain multiple consecutive video frame sequences; An object detection module configured to use the pre-trained YOLOv4-tiny model for object detection to obtain the object category and location information in the initial frame containing the object; An object tracking module configured to use the KCF kernel correlation filtering algorithm to track the detected object in subsequent frames; The specific process includes: Initialize the KCF kernel correlation filtering algorithm according to the object category and location information, and the initialization is to train the object tracking classifier. The initialization process includes: performing a high-dimensional mapping on the initial frame image containing the object; performing a Fourier transform on the high-dimensional mapped image data; using the transformed data to train the object tracking classifier to obtain the initialized kernel filter, and the initialization process is completed; where the initialization formula is: In the formula, represents the correlation matrix of each training sample, represents the training label, and λ represents the regularization parameter; Use the initialized object tracking classifier to track the object in subsequent frames, and the tracking formula is as follows: Among them, represents the target tracking result; represents the correlation matrix between the current image sample and each training sample; represents the key parameters of the target tracking classifier obtained through training; Key parameters of the target tracking classifier and the target observation template are updated, and the update formula is as follows: where γ represents the template update step size.

6. The lightweight object detection and tracking system for micro-robots according to claim 5, characterized in that, The structure of the YOLOv4-tiny model in the target detection module includes a feature extraction module, a feature fusion module, and a detection prediction head module. The feature extraction module includes multiple convolutional layers and multiple pooling layers, which are used to extract multi-dimensional features from the input video frame image. The feature fusion module is used to fuse features of different dimensions, depths, and scales using a feature pyramid. The detection prediction head module is used to identify and classify the fused features, and finally output the target category and location information.

7. The lightweight target detection and tracking system for a micro-robot according to claim 5, wherein The system further includes a correction module. The correction module is configured to use the KCF kernel correlation filtering algorithm to track the target for multiple frames, and then call the pre-trained YOLOv4-tiny model to perform target detection on the next frame image. The relevant image matching algorithm is used to perform similarity matching on the tracking data and the detection data. According to the matching result, the accuracy of the target tracking result is verified. If it is inaccurate, the target tracking classifier is re-initialized for training. Among them, the process of performing similarity matching on the tracking data and the detection data using the relevant image matching algorithm includes: calculating the mean hash similarity, difference hash similarity, and perceptual hash similarity of the tracking data and the detection data, and taking the mean of the three hash similarities as the similarity matching value.

Citation Information

Patent Citations

  • Video semi-automatic target labeling method integrating target detection and tracking

    CN110929560A

  • Target tracking method and system based on artificial intelligence

    CN116385498A