Limb Component Tracking Method Based on Detection of Hierarchical Features of Human Body Components

Through the limb component tracking method based on the hierarchical feature detection of human body parts, the features of human body parts are extracted and tracked by YOLOv7 and BoT-SORT algorithms, combined with the method of task decoupling and data association, the problem of losing local information in the prior art under abnormal situations is solved, and a more efficient and accurate target tracking effect is achieved.

CN115631536BActive Publication Date: 2025-05-30FUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211349251.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-31
Publication Date
2025-05-30
Estimated Expiration
2042-10-31

AI Technical Summary

Technical Problem

When existing target tracking technologies encounter abnormal situations such as occlusion and blur, they are prone to lose a large amount of valuable local information, and are limited to detection and tracking of the overall target, and fail to effectively utilize local semantic features.

Method used

Using the limb component tracking method based on the detection of hierarchical characteristics of human parts, a human part-level object detector is constructed through the YOLOv7 object detection algorithm, and the hierarchical characteristics of limb component are extracted, and the target tracking is tracked using BoT-SORT improved target tracking algorithm. Combining the method of task decoupling and data association, the preliminary tracking results are re-identified and reprocessed to output the final tracking results.

Benefits of technology

Effectively utilize local semantic features, improve the effect and accuracy of target tracking, expand the tracking application scenario, and enable detection and specific processing under abnormal situations such as occlusion and blur to obtain more accurate target tracking sequences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115631536B_ABST
    Figure CN115631536B_ABST
Patent Text Reader

Abstract

The present invention relates to a limb component tracking method based on the detection of hierarchical features of human body components, including the following steps: Step S1: Obtain a human body video or consecutive pictures, and then perform target box annotation on each limb part of the human body to construct a data set; Step S2: Construct a human body component-level target detector based on the target detection algorithm of YOLOv7, and train it. Then, according to the trained human body component-level target detector, detect each frame of the video to be tracked, extract the hierarchical features of the limb components, and output information; Step S3: Use the output result of Step S2 and adopt an improved target tracking algorithm based on BoT-SORT to track the movement trajectories of each limb of the human body; Step S4: Use the method based on limb connectivity and the method based on data association to perform re-identification and reprocessing on the preliminary tracking results of Step S3, and output the final tracking results. The method of the present invention can effectively detect and extract the hierarchical features of each component of the human body in video images and realize the tracking of limb components.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and particularly to a limb part tracking method based on hierarchical feature detection of human body parts. Background Art

[0002] In recent years, with the rapid development of artificial intelligence and multimedia technologies, object tracking has great market potential and academic value, and has always been widely concerned by researchers. With the development of related technologies and the continuous penetration of specific scenario applications into reality, the demand in related fields of object tracking such as pedestrian detection and personnel search has been further expanded. How to obtain more detailed and accurate tracking sequences to obtain more information to assist technical applications has become one of the research focuses and challenges of object tracking.

[0003] Although great progress has been made in object detection technology and object tracking technology, the existing functions are only limited to the detection and tracking of the overall object. In the face of abnormal situations such as occlusion and blur, a large amount of valuable local information is easily lost. Various local information of the object can provide multi-faceted and multi-level motion trajectories to assist in correcting and associating complete motion trajectories. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a limb part tracking method based on hierarchical feature detection of human body parts, which can effectively detect and extract the hierarchical features of each human body part in a video image, realize the tracking of limb parts, improve the tracking effect and expand the tracking application scenarios.

[0005] To achieve the above purpose, the present invention adopts the following technical solutions:

[0006] A limb part tracking method based on hierarchical feature detection of human body parts, comprising the following steps:

[0007] Step S1: Obtain a human body video or consecutive pictures, and then perform target box annotation on each limb part of the human body to construct a data set;

[0008] Step S2: Construct a human body part-level object detector based on the object detection algorithm of YOLOv7, train it according to the data set, and then perform frame-by-frame detection on the human body video or consecutive pictures to be tracked by the trained human body part-level object detector, extract the hierarchical features of limb parts, and output information;

[0009] Step S3: Use the output result of Step S2 to track the motion trajectories of each limb of the human body by using an object tracking algorithm improved based on BoT-SORT;

[0010] Step S4: Re-identify and reprocess the preliminary tracking results in Step S3 using the method based on limb connectivity and the method based on data association, and output the final tracking results.

[0011] Further, the specific content of Step S2 is as follows:

[0012] Step S21: Train the object detection algorithm based on YOLOv7 using the dataset constructed in Step S1 to obtain a human body part-level object detector;

[0013] Step S22: For the human body video or consecutive pictures to be tracked, decompose it into a video frame sequence I according to the frame rate. For each frame image, sequentially detect the position boxes of human body limb parts using the part-level object detection model trained in Step S21;

[0014] Step S23: According to the information of each part position box detected in Step S22, use a 3×3 convolution to extract the hierarchical features of the limb parts within the box, number them by human body unit, retain all detection results, and sequentially splice the hierarchical feature information of each limb part of a person into a feature vector.

[0015] Further, the specific content of Step S3 is as follows:

[0016] Step S31: According to the feature vectors output by the object detector in Step S2, establish a feature sequence for the position change of each limb part object box in the time series;

[0017] Step S32: Use an improved object tracking algorithm based on BoT-SORT to complete the preliminary tracking of each limb part of the human body on the feature sequence in Step S31.

[0018] Further, the specific content of Step S32 is as follows:

[0019] Use the BoT-SORT tracking algorithm based on the detection results of the YOLOv7 object detector to process the feature sequence. Select the target box with the highest confidence in the first frame as the initial position box and use it as the starting point of the tracking sequence. Then, match all candidate target boxes in the next frame with the current target box using the IoU (Intersection over Union) formula, and select the one with the highest matching degree as the target box for the next frame and add it to the tracking sequence;

[0020] Continuously perform the operation in Step S321 on the last target box in the tracking sequence and the candidate target boxes in the next frame until the current tracking sequence is completely matched and ended or the target is lost and the matching is interrupted.

[0021] Further, the IoU formula is as follows:

[0022]

[0023] Among them, L represents the last target box in the current tracking sequence, N represents the candidate target box in the next frame, and area(·) represents the area of the box;

[0024] Furthermore, the method based on limb connectivity is specifically as follows:

[0025] Adopt the concept of task decoupling to decouple the human body according to each limb component to extract the hierarchical features of human body components. According to the articulated structure of the human body, the head, left and right hands, and left and right legs are all connected by the torso. The connection relationship of each component of the human body is used to re-identify and reprocess the preliminary tracking result of step S3 based on limb connectivity, which is formalized as follows:

[0026]

[0027] Among them, and represent the initial tracking sequences of the i-th and j-th limb components, and R 1 (·) represents re-identification and reprocessing based on limb connectivity, and T 1 represents the new tracking sequence after processing.

[0028] Furthermore, the method based on data association is specifically as follows:

[0029] For the problem of target loss in the tracking scenario, utilize the characteristic that the confidence of the target detection box in the extremely short time series changes from high to low and then to high as the target goes from existence to loss and then to appearance. Associate the low-score detection boxes in the part of the time series with matching failures and losses with the normal tracking sequence, simulate the target loss process, and perform re-identification and reprocessing based on the data association method on the tracking sequence T 1 at the detail level, which is formalized as follows:

[0030]

[0031] Among them, T 1 k is the k-th sequence segment lost in T 1 sequence, and T 1 k-1 and T 1 k+1 respectively represent the normal non-lost sequences before and after the k-th lost sequence, and R 2 (·) represents re-identification and reprocessing based on the data association method, and T 2 represents the finally output tracking sequence.

[0032] The present invention has the following beneficial effects compared with the prior art:

[0033] 1. In view of the problem that traditional target tracking methods only use overall information for detection and tracking, which easily lose local semantic feature information, the present invention proposes a method for decoupling human body parts according to the requirements of the application scenario, detecting part-level features and extracting limb information, and performing part-level tracking on each limb part of the human body. The present invention can effectively track from the target part level, effectively utilize local semantic features to assist in improving the overall tracking effect and expanding the tracking application scenario.

[0034] 2. The present invention first adopts the concept of task decoupling to decouple the human body according to each limb part to extract human body part-level features. According to the articulated structure of the human body, the head, left and right hands, and left and right legs are all connected by the torso. Using the connection relationships of these human body parts, the preliminary tracking results are re-identified and reprocessed based on limb connectivity to obtain a more accurate target tracking sequence that integrates the motion information of each limb part of the human body.

[0035] 3. The present invention can detect and perform specific processing on abnormal situations such as occlusion and blur. First, according to the articulated structure of the human body, the visible limb information is used to assist in inferring the information of adjacent occluded limbs; secondly, using the characteristic that the confidence of the target detection box of the target decreases from high to low and then increases from the existence to the loss and then to the appearance of the target, the target box information with low detection confidence is saved, and it is associated and matched with the normal sequence to infer the motion trajectory of abnormal situations such as occlusion. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 is a schematic flowchart of the method in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] The present invention will be further described below with reference to the drawings and embodiments.

[0038] Please refer to Figure 1 , the present invention provides a limb part tracking method based on human body part-level feature detection, including the following steps:

[0039] Step S1: Obtain a human body video or a series of consecutive pictures of the relevant application scenario, and perform target box annotation on each limb part of the human body to construct a data set.

[0040] In this embodiment, step S1 specifically includes the following steps:

[0041] Step S11: Obtain publicly available human body videos or a series of consecutive pictures related to the application scenario from the network, and preliminarily construct a data set;

[0042] Step S12: Perform preprocessing on the data set, screen out background elements that have nothing to do with the human body, and process appropriate image and video segments to complete the construction of the data set;

[0043] Step S13: Perform target box annotation on the limb parts such as the human head, torso, hands, and legs in the dataset according to application requirements.

[0044] In this embodiment, taking the common limb division of the human body as an example, each part is encoded, using p i , i = 0, 1, … 5 to represent, where p represents the torso, head, left hand, right hand, left leg, and right leg respectively. The dataset is divided into a training set and a test set according to a certain ratio.

[0045] Step S2: Based on the dataset constructed in Step S1, train a human body part-level target detector, detect and extract limb part-level features frame by frame, and output information.

[0046] In this embodiment, Step S2 specifically includes the following steps:

[0047] Step S21: Use the human body part-level dataset constructed in Step S1 to train the target detection algorithm based on YOLOv7 to obtain a human body part-level target detector;

[0048] Step S22: For the input video image, decompose the video into a sequence of video frames I according to the frame rate of the video. For each frame image, sequentially use the part-level target detection model trained in Step S21 to detect the position boxes of human body limb parts;

[0049] Step S23: According to the position box information of each part detected in Step S22, use a 3×3 convolution to extract the limb part-level features within the box, number them by human body unit, retain all detection results, and sequentially splice the limb part-level feature information of each person into a feature vector.

[0050] The process of Step S2 is formalized as follows:

[0051]

[0052]

[0053] Among them, n represents the total number of frames in the video frame sequence I, H and W are the height and width of the video image respectively, I t represents the t-th video frame, D(·) represents the part-level target detector, p t i represents the i-th limb part information of the human body in the t-th video frame, and F t represents the part-level feature vector of the human body in the t-th video frame.

[0054] Step S3: Utilize the output result of Step S2 and adopt a detection-based target tracking method to track the movement trajectories of each human body limb.

[0055] In this embodiment, step S3 specifically includes the following steps:

[0056] Step S31: Based on the feature vectors output by the target detector in step S2, establish a feature sequence for the position change of each limb part target box in the time series.

[0057] Step S32: Use an object tracking algorithm improved based on BoT-SORT to complete the preliminary tracking of each limb part of the human body on the feature sequence in step S31.

[0058] Preferably, step S32 specifically includes the following steps:

[0059] Step S321: Use the BoT-SORT tracking algorithm based on the detection results of the YOLOv7 target detector to process the feature sequence. Select the target box with the highest confidence in the first frame as the initial position box and use it as the starting point of the tracking sequence. Then, match all candidate target boxes in the next frame with the current target box using the IoU (Intersection over Union) formula. The candidate target box with the highest matching degree is selected as the target box for the next frame and added to the tracking sequence. The IoU formula is as follows:

[0060]

[0061] where L represents the last target box in the current tracking sequence, N represents the candidate target box in the next frame, and area(·) represents the area of the box.

[0062] Step S322: Continuously perform the operation in step S321 on the last target box in the tracking sequence and the candidate target boxes in the next frame until the current tracking sequence is completely matched or the target is lost and the matching is interrupted.

[0063] Step S4: Use the method based on limb connectivity and the method based on data association to re-identify and reprocess the preliminary tracking results in step S3, and output the final tracking results.

[0064] In this embodiment, step S4 specifically includes the following steps:

[0065] Step S41: Adopt the concept of task decoupling to decouple the human body according to each limb part to extract the hierarchical features of human body parts. According to the articulated structure of the human body, the head, left and right hands, and left and right legs are all connected by the torso. Taking the right hand as an example, the position connection relationship between its tracking sequence and the tracking sequence of the torso should be one-to-one correspondence and conform to the trajectory change of the motion sequence. The relative position between its tracking sequence and the tracking sequence of the left hand and the limb connectivity connected by the torso should also be one-to-one correspondence and conform to the trajectory change of the motion sequence. Based on the connection relationships of the above-mentioned human body parts, re-identify and reprocess the preliminary tracking results in step S3 based on limb connectivity, which is formalized as follows:

[0066]

[0067] Among them, and represent the initial tracking sequences of the i-th and j-th limb components, and R 1 (·) represents re-identification and reprocessing based on limb connectivity, and T 1 represents the new tracking sequence after processing;

[0068] Step S42: For the problem of target loss caused by difficulties such as occlusion and blurring in the tracking scenario, utilize the characteristic that in an extremely short time series, the confidence of the target detection box accompanying the target from existence to loss and then to appearance changes from high to low and then to high. Associate the low-score detection boxes in the partially lost time series with the normal tracking sequence where the matching fails, simulate the target loss process, and perform re-identification and reprocessing based on the data association method on the tracking sequence T 1 output in Step S41 at the detail level, which is formalized as follows:

[0069]

[0070] Among them, T 1 k is the k-th sequence segment lost in the T 1 sequence, and T 1 k-1 and T 1 k+1 respectively represent the normal non-lost sequences before and after the k-th sequence of this loss, and R 2 (·) represents re-identification and reprocessing based on the data association method, and T 2 represents the finally output tracking sequence.

[0071] The above are only the preferred embodiments of the present invention. All equivalent changes and modifications made according to the scope of the patent application of the present invention shall fall within the scope covered by the present invention.

Claims

1. A limb component tracking method based on human body component hierarchical feature detection, characterized in that, it includes the following steps: Step S1: Obtain a human body video or a series of consecutive pictures, and then perform target box annotation on each limb part of the human body to construct a data set; Step S2: Build a human body component-level target detector based on the YOLOv7 object detection algorithm, train it according to the data set, and then perform frame-by-frame detection on the human body video or consecutive pictures to be tracked by the trained human body component-level target detector, extract the limb component hierarchical features, and output information; Step S3: Use the output result of Step S2 and adopt an improved object tracking algorithm based on BoT-SORT to track the movement trajectories of each limb of the human body; Step S4: Use the method based on limb connectivity and the method based on data association to perform re-identification and reprocessing on the tracking result of Step S3, and output the final tracking result; The method based on limb connectivity is specifically: Adopt the concept of task decoupling to decouple the human body according to each limb component to extract the human body component hierarchical features. According to the human articulated structure, the head, left and right hands, and left and right legs are all connected by the torso. The connection relationship of each component of the human body is used to perform re-identification and reprocessing based on limb connectivity on the preliminary tracking result of Step S3, which is formalized as follows: Among them, and represent the initial tracking sequences of the i-th and j-th limb components, and R 1 (·) represents re-identification and reprocessing based on limb connectivity, and T 1 represents the new tracking sequence after processing; The method based on data association is specifically: For the problem of target loss in the tracking scenario, by utilizing the characteristic that in a very short time series, the confidence of the target detection box changes from high to low and then to high as the target goes from existence to loss and then to appearance, the low-score detection boxes in the part of the time series with matching failures and losses are associated with the normal tracking sequence to simulate the target loss process for the tracking sequence T 1 Perform re-identification and reprocessing based on the data association method at the level of details, formalized as follows: Among them, T 1 k is the k-th sequence segment lost in the T 1 sequence. T 1 k-1 and T 1 k+1 respectively represent the normal non-lost sequences on the preamble and postamble of the k-th lost sequence. R 2 (·) represents re-identification and reprocessing based on the data association method. T 2 represents the finally output tracking sequence.

2. The limb component tracking method based on human body component hierarchical feature detection according to claim 1, characterized in that, the specific content of Step S2 is: Step S21: Use the data set constructed in Step S1 to train the YOLOv7 object detection algorithm to obtain a human body component-level target detector; Step S22: For the human body video or consecutive pictures to be tracked, decompose it into a video frame sequence I according to the frame rate. For each frame image, sequentially use the component-level target detection model trained in Step S21 to detect the position boxes of the human body limb components; Step S23: According to the information of each component position box detected in Step S22, use a 3×3 convolution to extract the limb component hierarchical features within the box, number them by human body unit, retain all detection results, and sequentially splice the hierarchical feature information of each limb component of a person into a feature vector.

3. The limb component tracking method based on human body component hierarchical feature detection according to claim 2, characterized in that, the specific content of Step S3 is: Step S31: According to the feature vector output by the target detector in Step S2, establish a feature sequence for the position change of each limb component target box in the time series; Step S32: Adopt an improved object tracking algorithm based on BoT-SORT to complete the preliminary tracking of each limb component of the human body on the feature sequence in Step S31.

4. The limb component tracking method based on human body component hierarchical feature detection according to claim 3, characterized in that, the specific content of Step S32 is: Process the feature sequence using the BoT-SORT tracking algorithm based on the detection results of the YOLOv7 object detector. Select the target box with the highest confidence in the first frame as the initial position box and use it as the starting point of the tracking sequence. Then, match all candidate target boxes in the next frame with the current target box using the IoU (Intersection over Union) formula, and select the one with the highest matching degree as the target box for the next frame and add it to the tracking sequence. Continuously perform the operation in step S321 on the last target box in the tracking sequence and the candidate target boxes in the next frame until the current tracking sequence is completely matched or the target is lost and the matching is interrupted.

5. The limb part tracking method based on human body part hierarchical feature detection according to claim 4, characterized in that the IoU formula is as follows: where L represents the last target box in the current tracking sequence, N represents the candidate target box in the next frame, and area(·) represents the area of the box.

Citation Information

Patent Citations

  • Multi-person concurrent interaction behavior understanding method based on single-frame image

    CN113158782A

  • Three-Dimensional Object Modelling Fitting & Tracking

    US20140334670A1