A multi-object tracking method including visual saliency
By combining visibility and category confidence, the information extraction and matching process of the multi-objective tracking method is optimized, and the problem of inaccurate feature matching in the prior art is solved, and the target appears in other positions after cross-scene occlusion is not possible, thereby achieving higher feature extraction and matching accuracy and cross-scene tracking capabilities.
Patent Information
- Application Number
- CN202310897966.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-20
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2043-07-20
AI Technical Summary
The existing multi-objective tracking methods are insufficient when feature matching is inaccurate, cross-spectrum tracking is not possible, and when the target appears in other locations after being blocked.
By combining visibility and category confidence, the information extraction method is optimized, the accuracy of feature extraction and matching accuracy is improved, and cross-mirror multi-objective tracking and single-mirror multi-objective tracking are supported.
It improves the accuracy of feature extraction and matching, supports cross-mirror tracking, and solves the analysis problem when the target appears in other locations after being blocked.
Smart Images

Figure CN116935276B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence and computer vision tracking, and particularly relates to a multi-object tracking method including visualization degree. Background Art
[0002] Multi-object tracking (MOT) is a key technology in the field of computer vision and is widely applied in directions such as autonomous driving, intelligent monitoring, behavior recognition, smart city, and smart security. In the paper Bot-Sort: Robust Associations Multi-Pedestrian Tracking (abbreviation: Bot-Sort [2]) retrieved from the arxiv website, after adding pedestrian re-identification, the BoT-SORT-ReID model in the paper was proposed. Under a better detector, such as YOLOX [4], the accuracy was only slightly improved compared to the ByteTrack [1] model. This was mainly due to the addition of camera motion compensation. However, the price paid was that the calculation speed decreased by several times and it was simply impossible to be real-time. If the motion compensation was removed, the accuracy would be lower than that of ByteTrack [1]. This was the same as the multiple algorithms verified in the paper Bytetrack [1]. The role of directly adding target features as matching information was becoming increasingly negligible. The main reason was that the target features extracted by the pedestrian re-identification algorithm were no longer accurate and it was impossible to well believe in the correctness of its matching.
[0003] In the current multi-object tracking methods, the model adopting Bot-sort has the following problems:
[0004] 1. Feature matching is inaccurate. When a part of the target is outside the camera lens, the similarity between the full-body features and partial body features of the same target may be low due to feature misalignment, resulting in inaccurate situations. In addition, when being occluded, the extracted features contain interference information. The longer the time, the greater the interference. Because the confidence of the occluded target can also be very high and feature extraction will also be performed, which is also a common practice of many algorithms.
[0005] 2. Cross-camera tracking cannot be performed. When Bot-sort performs trajectory matching, it first calculates the mask matrix of the IoU cost matrix between the target and the trajectory, and then calculates the feature cost matrix between the target and the trajectory. Even if the feature similarity meets the requirements, but does not meet the IoU requirements, it is not feasible. Under multiple cameras, it is difficult to use the position information, and it is easier to use the feature information.
[0006] 3. It is unable to solve the situation where the target is occluded and then appears at other positions. When Bot-sort performs trajectory matching, it first calculates the mask matrix of the IoU cost matrix between the target and the trajectory, and then calculates the feature cost matrix between the target and the trajectory. Even if the feature similarity meets the requirements, but does not meet the IoU requirements, it is not acceptable.
[0007] The invention patent with the application number CN201911261618.2 provides an anti-occlusion target tracking method. First, for the input video or image sequence, an object detector is used to detect each frame of the video to obtain detection-based candidates; according to the output of the object detection result of the current frame, the Kalman filter is used to predict the position of the target in the next frame to obtain tracking-based candidates. The confidence of the candidates is calculated according to the confidence scoring formula, and the non-maximum suppression algorithm is used to obtain the final candidates; the candidates of adjacent frames are input into the feature matching network, and the matching degree between the targets is calculated through the cascade matching algorithm. The detection-based candidates are subjected to feature extraction through a deep neural network to perform the matching of the similarity between the features; the IOU overlap degree matching is performed on the tracking-based candidates. The position of the target in the current frame is determined according to the target matching result of adjacent frames, and then the target motion trajectory is output. This technology cannot solve the above technical problems. Summary of the Invention
[0008] The purpose of the present invention is to provide a multi-target tracking method including visibility. This method uses the combination of visibility and class confidence to optimize the information extraction method, improve the accuracy of feature extraction and the accuracy of matching, and support cross-camera multi-target tracking and single-camera multi-target tracking.
[0009] In order to achieve the above purpose, the technical solution adopted by the present invention is:
[0010] A multi-target tracking method including visibility. In object detection, visibility is used as the target value for learning to predict the visible degree of the overall part of the target. The target visibility degree and the target class confidence are combined as the decision value for whether to perform feature extraction on the target and use the target features for data matching. The specific algorithm process is as follows:
[0011] Step 1: Obtain the video stream and read the pictures in chronological order.
[0012] Step 2: Input the read pictures into the object detection model in chronological order, and then obtain the output of the model, which includes class confidence, target position information, and visibility information.
[0013] Step 3: If it is the first frame containing the target and the class confidence of the target is greater than the initialization new trajectory threshold T newIf so, initialize a new trajectory; if at the same time the visibility of the target is greater than the visibility threshold T v , then crop out the target and input it into the feature extraction model for feature extraction. After processing all the targets in the current image, return to the first step; otherwise, go to the next step;
[0014] Step 4: Determine whether the class confidence of the target in each image is greater than the high class confidence threshold T high ;
[0015] Step 5: If the class confidence of the target is greater than the high class confidence threshold T high , determine whether the visibility of the target is greater than the visibility threshold T v ;
[0016] Step 6: If the class confidence of the target is greater than the high class confidence threshold T high , and the visibility of the target is greater than the visibility threshold T v , then crop out the target and input it into the feature extraction model for feature extraction. At this time, the target contains position information, class confidence information, visibility information, and feature information; then put the extracted target information into the set Set1 of the first-round matching to match the feature information with the trajectory. The trajectory with successful matching updates the trajectory state and related information. The target with successful matching is added to the corresponding trajectory and deleted from the set Set1; the targets and trajectories with unsuccessful matching wait for the second-round matching;
[0017] Step 7: If the class confidence of the target is greater than the high class confidence threshold T high , but the visibility of the target is less than the visibility threshold T v , put it into the set Set2, and merge the remaining targets in the set Set1 with the set Set2 for the second-round matching. This matching only uses the position information for matching; the trajectory with successful matching updates the trajectory state and related information, and the trajectory with unsuccessful matching waits for the third-round matching and deletes the target with successful matching from set2;
[0018] Step 8: If the class confidence of the target is less than the high threshold T high , determine whether the class confidence is greater than the low class confidence threshold T low ;
[0019] Step 9: If the class confidence of the target is less than the high threshold T high , and the class confidence is greater than the low threshold T low , then put it into the set Set3 for the third-round matching to match the position information. The trajectory with successful matching updates the trajectory state and related information. The target with successful matching is added to the corresponding trajectory, updates the trajectory state of the unsuccessful matching, and discards the targets with unsuccessful matching in Set4;
[0020] Step Ten: If the category confidence of the target is less than the low category confidence threshold T low , discard the target information and do not process it;
[0021] Step Eleven: After the above three rounds of matching are completed, delete the trajectories in the Delete state;
[0022] Step Twelve: Initialize a new trajectory and determine whether the category confidence of the targets in Set2 is greater than the new trajectory initialization threshold T new . If the judgment is yes, create a trajectory; otherwise, discard it;
[0023] Step Thirteen: Proceed to the processing of the next image, and then repeat Steps 2 to 12.
[0024] Preferably, in Step Two, the target detection model detects information through the target detection algorithm YOLOX[4]. Based on the output head of the target detection algorithm YOLOX[4], a visibility branch head is added in parallel. The target detection algorithm after adding visibility is named YOLOX_v. The output value of this YOLOX_v branch is only the visibility value, and the final output of the visibility value is H*W*(C + 6); where H is the height of the feature map, W is the width of the feature map, C is equal to the number of target categories to be detected and is the category confidence of each target; 6 includes 1 bit of visibility, 1 bit of confidence of the existence of the target, and 4 bits of position information.
[0025] Preferably, the training method of YOLOX_v is as follows:
[0026] Step 3.1: Obtain a sample data set;
[0027] Step 3.2: Build a detection network training model;
[0028] Step 3.3: Detect and track targets through the training model.
[0029] Preferably, the label information of the targets is collected in the sample data set. The label information includes category information, position coordinate information, and visibility information. The sample data set is divided into a training set and a test set.
[0030] Preferably, the method for building the detection network training model is as follows:
[0031] Step 3.2.1: First, build a YOLOX_v network model;
[0032] Step 3.2.2: After adding a visibility branch to YOLOX, modify the detection head of the YOLOX network to be used to additionally predict the visibility of the targets on the basis of the original prediction of the category confidence, the confidence of the existence of the targets, and the position information;
[0033] Step 3.2.3: Add a corresponding loss function for the increased visibility branch;
[0034] Step 3.2.4: Train the modified model;
[0035] Step 3.2.5: Use the validation set in the improved network model to obtain the optimal weights.
[0036] Preferably, the method for detecting and tracking the target is as follows:
[0037] Step 3.3.1: Use the optimal weights to perform frame-by-frame detection on the video to be tracked;
[0038] Step 3.3.2: Determine whether to input the target into the feature extraction network model for feature extraction according to the detection results;
[0039] Step 3.3.3: Obtain the real-time trajectories of each target.
[0040] Preferably, the loss function is:
[0041]
[0042] In formula (1):
[0043] L cls represents the class loss, L reg represents the localization loss, L obj represents the loss of whether the target is included, L vis represents the visibility loss, λ represents the balance coefficient of the localization loss, ρ represents the balance coefficient of the visibility loss, N pos represents the number of anchor points classified as positive samples.
[0044] Preferably, when the multi-target tracking method realizes state changes, the trajectory state change method is used to complete the display and operation of the trajectory state change during the trajectory and target matching process; among them, during the matching process, the distance between the new target set in the current round and the original trajectory set in the current round is calculated using feature information or position information, and this distance is used as a condition for judging whether the target and the trajectory match, and the Hungarian matching algorithm is used for optimal matching to make the trajectory find the most suitable target for itself;
[0045] There are two situations in the matching process. One is that the new target and the original trajectory meet the matching conditions. The other is that there are new targets that do not match the original trajectories, or the original trajectories that match the new targets. It is necessary to process the target and trajectory states in the new target set and the original trajectories under the two matching states.
[0046] Preferably, in the trajectory state change method, various states existing in the trajectory change are first determined. The various states include New state, Tracked state, Lost_Tracked state, Delete state, Unformed state, and Lost_Unformed state;
[0047] The meanings of each state are as follows:
[0048] New state: The state of a newly created trajectory, which is the first frame of the video that meets the conditions and the target that has not been matched to the trajectory;
[0049] Tracked state: A trajectory that contains feature information and has a new target matched in the previous frame;
[0050] Lost_Tracked state: A trajectory that contains feature information but has not had a new target matched within the previous p frames;
[0051] Delete state: The state of a trajectory to be deleted;
[0052] Unformed state: A trajectory that does not contain feature information and has a new target matched in the previous frame;
[0053] Lost_Unformed state: A trajectory that does not contain feature information but has not had a new target matched within the previous m frames;
[0054] According to the above states, the specific implementation content of the trajectory state change method is as follows:
[0055] Step 6.1: Process the new target set Dets in the current round. If it is the first frame of the video, or the target that has not been matched to the trajectory after three rounds of matching, and its class confidence is greater than the new trajectory creation threshold T new , then set the state to New, otherwise discard it;
[0056] Step 6.2: If the trajectories in the New state have been matched to targets for n consecutive frames, and there is a target with a class confidence greater than T high and a visibility greater than the threshold T v , that is, feature information has been extracted, then set the state to Tracked. Here, n is set to 3;
[0057] Step 6.3: If the trajectories in the New state have been matched to targets for n consecutive frames, and there is no target with a class confidence greater than T high and a visibility greater than the threshold T v , that is, there is no feature information, then set the state to Unformed;
[0058] Step 6.4: When the trajectory in the Unformed state is matched to a target with a visibility greater than Tv When the target is met, update the status to Tracked;
[0059] Step 6.5: When the trajectory in the Tracked state does not match the target, set the trajectory status to Lost_Tracked;
[0060] Step 6.6: When the trajectory in the Lost_Tracked state matches the target, set the trajectory status to Tracked;
[0061] Step 6.7: When the trajectory in the Unformed state does not match the target, set the trajectory status to Lost_Unformed;
[0062] Step 6.8: When the trajectory in the Lost_Unformed state matches the target, if the visibility of the target is greater than the threshold T v and the class confidence is greater than T high , then set the trajectory status to Tracked;
[0063] Step 6.9: When the trajectory in the Lost_Unformed state matches the target, if the visibility of the target is less than or equal to the threshold T v, then set the trajectory status to Unformed;
[0064] Step 6.10: When the trajectory in the Lost_Unformed state continuously loses m frames, set the status to Delete;
[0065] Step 6.11: When the trajectory in the Lost_Tracked state continuously loses p frames, set the status to Delete;
[0066] Step 6.12: When the trajectory in the New state loses one frame, set the status to Delete.
[0067] The beneficial effects of the present invention are:
[0068] In the multi-object tracking method including visibility in the present application, the modified detection algorithm YOLOX_v can output visibility information and can be used in other scenarios where it is necessary to judge the occlusion state of the target. Using the target with higher visibility and class confidence for feature extraction greatly improves the accuracy of feature extraction and thus the accuracy of matching, and supports cross-camera multi-object tracking and single-camera multi-object tracking, solving the situation where it is impossible to accurately judge and analyze when the target is occluded and then appears in other positions. Description of the Drawings
[0069] Figure 1 is a schematic diagram of the algorithm flow of the multi-object tracking method including visibility.
[0070] Figure 2 It is a schematic diagram of the YOLOX_v decoupled head including visibility.
[0071] Figure 3 It is a block diagram of the training method of YOLOX_v.
[0072] Figure 4 It is a schematic diagram of the trajectory state change method.
[0073] Figure 5 It is a schematic diagram shown by the detection box after reading the picture. Specific implementation mode
[0074] The present invention will be described in detail below with reference to the accompanying drawings:
[0075] Example 1, in combination with Figure 1 , a multi-object tracking method including visibility, in object detection, visibility is used as a target value for learning to predict the visible degree of the overall part of the target. The target visibility degree and the target class confidence are combined as a decision value for whether to perform feature extraction on the target and use the target features for data matching.
[0076] The specific algorithm flow of the multi-object tracking method is as follows:
[0077] Step 1: Obtain the video stream and read the pictures in chronological order;
[0078] Step 2: Input the read pictures into the object detection model in chronological order, and then obtain the output of the model, which includes class confidence, target position information, and visibility information; the object detection model detects information through the object detection algorithm YOLOX[4]. The object detection algorithm YOLOX[4] includes a visibility branch head. On the basis of the output head of the object detection algorithm YOLOX[4], a visibility branch head is added in parallel. The object detection algorithm with increased visibility is named YOLOX_v. The output value of the YOLOX_v branch is only the visibility value, and the visibility value is finally output as H*W*(C + 6);
[0079] Among them, H is the height of the feature map, W is the width of the feature map, C is equal to the number of target types to be detected, and is the class confidence of each target; 6 includes 1 bit of visibility, 1 bit of confidence in the existence of the target, and 4 bits of position information.
[0080] Step 3: If it is the first frame containing the target and the class confidence of the existing target is greater than the initialization new trajectory threshold T new then initialize a new trajectory; if at the same time the visibility of the target is greater than the visibility threshold T v, then crop out the target and input it into the feature extraction model for feature extraction. After processing all the targets in the current image, return to the first step; otherwise, proceed to the next step.
[0081] Step 4: Determine whether the class confidence of the target in each image is greater than the high class confidence threshold T high .
[0082] Step 5: If the class confidence of the target is greater than the high class confidence threshold T high , determine whether the visibility of the target is greater than the visibility threshold T v .
[0083] Step 6: If the class confidence of the target is greater than the high class confidence threshold T high , and the visibility of the target is greater than the visibility threshold T v , then crop out the target and input it into the feature extraction model for feature extraction. At this time, the target contains position information, class confidence information, visibility information, and feature information; then put the target information after extraction into the set Set1 of the first-round matching to match the feature information with the trajectory. The trajectory with a successful match updates the trajectory status and related information, the target with a successful match is added to the corresponding trajectory, and is deleted from the set Set1; the targets and trajectories that do not match successfully wait for the second-round matching.
[0084] Step 7: If the class confidence of the target is greater than the high class confidence threshold T high , but the visibility of the target is less than the visibility threshold T v , put it into the set Set2, and merge the remaining targets in the set Set1 with the set Set2 for the second-round matching. Only the position information is used for this matching; the trajectory with a successful match updates the trajectory status and related information, the trajectory that does not match successfully waits for the third-round matching, and the target with a successful match is deleted from set2.
[0085] Step 8: If the class confidence of the target is less than the high threshold T high , determine whether the class confidence is greater than the low class confidence threshold T low .
[0086] Step 9: If the class confidence of the target is less than the high threshold T high , and the class confidence is greater than the low threshold T low , then put it into the set Set3 for the third-round matching to match the position information. The trajectory with a successful match updates the trajectory status and related information, the target with a successful match is added to the corresponding trajectory, update the status of the trajectory that does not match successfully, and discard the target that does not match successfully in Set4.
[0087] Step 10: If the class confidence of the target is less than the low class confidence threshold Tlow , discard the target information and do not process it.
[0088] Step Eleven: After the above three rounds of matching are completed, delete the trajectories in the Delete state.
[0089] Step Twelve: Initialize a new trajectory and determine whether the confidence of the target category in Set2 is greater than the threshold T for initializing a new trajectory new , if the judgment is yes, create a trajectory, otherwise discard it.
[0090] Step Thirteen: Proceed to the processing of the next picture, and then repeat Steps 2 to 12.
[0091] Example 2, based on what is disclosed in Example 1, the following is further disclosed in this example:
[0092] Combined with Figures 1 to 3 , through the object detection algorithm YOLOX[4], the training method of YOLOX_v in this example is as follows:
[0093] Step 3.1: Obtain a sample data set; the label information of the collected targets in the sample data set includes category information, position coordinate information, and visibility information, and the sample data set is divided into a training set and a test set. Step 3.2: Construct a detection network training model; Step 3.3: Perform object detection and tracking through the training model.
[0094] The method for constructing the detection network training model is as follows:
[0095] Step 3.2.1: First, construct the YOLOX_v network model;
[0096] Step 3.2.2: After adding a visibility branch to YOLOX, modify the detection head of the YOLOX network to be used to additionally predict the visibility of the target on the basis of the original prediction of the category confidence, the confidence including the target, and the position information;
[0097] Step 3.2.3: Adding a branch naturally also requires adding a loss function. The present invention selects the mean absolute error in the common regression losses. Since visibility is a measure of the visibility of the target and has great subjectivity, and the mean absolute error has good robustness to noise, this choice is made. The present invention also selects the root mean square error for comparative experiments and finds that the mean absolute error is indeed better than the root mean square error. When selecting the root mean square error, the visibility of all visible targets is still less than 1. The settings and strategies of the remaining parameters are all kept consistent with YOLOX[4], and the visibility loss does not participate in the assignment calculation of positive and negative samples.
[0098] The loss function is:
[0099]
[0100] In formula (1):
[0101] L cls represents the class loss, L reg represents the localization loss, L obj represents the presence loss, L vis represents the visibility loss, λ represents the balance coefficient of the localization loss, ρ represents the balance coefficient of the visibility loss, and N pos represents the number of anchor points classified as positive samples.
[0102] Step 3.2.4: Train the modified model;
[0103] Step 3.2.5: Use the validation set in the improved network model to obtain the optimal weights.
[0104] The method for detecting and tracking the target is as follows: Step 3.3.1: Use the optimal weights to perform frame-by-frame detection on the video to be tracked; Step 3.3.2: Determine whether to input the target into the feature extraction network model for feature extraction according to the detection results; Step 3.3.3: Obtain the real-time trajectories of each target.
[0105] The reasons for the present invention to adopt YOLOX are as follows. First, it has extremely excellent performance itself. More importantly, it adopts a decoupled head, as shown in the YOLOX_v decoupled head. Regarding the principle of decoupling, in the object detection task, it is not beneficial to the detection head to detect the coordinates of the object and the object classification in the same detection head. In the article, this kind of detection head is called a sibling head, which is also the coupled head mentioned in YOLOX. Because the localization task and the classification task focus on different parts, this misalignment can be solved by simple task-aware spatial disentanglement (TSD). The principle of this method is to observe a certain instance. The features of some prominent regions have rich information features for classification, while some boundary features are useful for boundary regression. Therefore, TSD improves the efficiency of the model by decoupling the two tasks. That is, the preferences of the localization task and the classification task are inconsistent, so it is unreasonable to place them in the same detection head.
[0106] Example 3, on the basis of the disclosure of the above example, the present example further discloses the following:
[0107] Combined with Figures 1 to 4, when the multi-object tracking method realizes state changes, it uses the trajectory state change method to complete the display and operation of the trajectory state changes during the trajectory and target matching process; among them, during the matching process, the distance between the new target set in the current round and the original trajectory set in the current round is calculated using feature information or position information, and this distance is used as the condition for judging whether the target and the trajectory match, and the Hungarian matching algorithm is used for optimal matching to make the trajectory find the most suitable target for itself;
[0108] There are two situations in the matching process. One is that the new target and the original trajectory meet the matching conditions. The other is that there are new targets that do not match the original trajectories, or the original trajectories that match the new targets. It is necessary to process the target and trajectory states in the new target set and the original trajectories under the two matching states.
[0109] In the trajectory state change method, multiple states existing in the trajectory change are first determined. The multiple states include New state, Tracked state, Lost_Tracked state, Delete state, Unformed state, and Lost_Unformed state;
[0110] The meanings of each state are as follows:
[0111] New state: The state of a newly created trajectory, which is the first frame of the video that meets the conditions and the target that has not been matched with a trajectory;
[0112] Tracked state: A trajectory that contains feature information and has a new target matched in the previous frame;
[0113] Lost_Tracked state: A trajectory that contains feature information but has not been matched with a new target within the previous p frames;
[0114] Delete state: The state of the trajectory to be deleted;
[0115] Unformed state: A trajectory that does not contain feature information and has a new target matched in the previous frame;
[0116] Lost_Unformed state: A trajectory that does not contain feature information but has not been matched with a new target within the previous m frames;
[0117] According to the above states, the specific implementation content of the trajectory state change method is as follows:
[0118] Step 6.1, process the new target set Dets in the current round. If it is the first frame of the video, or the target that has not been matched with a trajectory after three rounds of matching, and its class confidence is greater than the new trajectory creation threshold T new , then set the state to New, otherwise discard it;
[0119] Step 6.2: If the trajectories in the New state match the target for n consecutive frames, and there is a target with a class confidence greater than T high and a visibility greater than the threshold T v i.e., the feature information has been extracted, then set the state to Tracked, where n is set to 3;
[0120] Step 6.3: If the trajectories in the New state match the target for n consecutive frames, and there is no target with a class confidence greater than T high and a visibility greater than the threshold T v i.e., there is no feature information, then set the state to Unformed;
[0121] Step 6.4: When the trajectory in the Unformed state matches a target with a visibility greater than T v update the state to Tracked;
[0122] Step 6.5: When the trajectory in the Tracked state does not match the target, set the trajectory state to Lost_Tracked;
[0123] Step 6.6: When the trajectory in the Lost_Tracked state matches the target, set the trajectory state to Tracked;
[0124] Step 6.7: When the trajectory in the Unformed state does not match the target, set the trajectory state to Lost_Unformed;
[0125] Step 6.8: When the trajectory in the Lost_Unformed state matches the target, if the visibility of the target is greater than the threshold T v and the class confidence is greater than T high , then set the trajectory state to Tracked;
[0126] Step 6.9: When the trajectory in the Lost_Unformed state matches the target, if the visibility of the target is less than or equal to the threshold T v, then set the trajectory state to Unformed;
[0127] Step 6.10: When the trajectory in the Lost_Unformed state loses m consecutive frames, set the state to Delete;
[0128] Step 6.11: When the trajectory in the Lost_Tracked state loses p consecutive frames, set the state to Delete;
[0129] Step 6.12: When the trajectory in the New state loses one frame, set the state to Delete.
[0130] Example 4. On the basis of the disclosure of the above embodiments, the present embodiment further discloses the following:
[0131] Combined with Figures 1 to 5 , when the actual information is output in the above multi-object tracking method with visual degree, a corresponding detection box is formed on the detected video or picture, and the actual detected value is displayed on the detection box. The first value on the detection box is the class confidence, and the second value is the visibility. It can be observed that when the target is not occluded, the visibility is close to 1; when it is occluded, the visibility is no longer close to 1.
[0132] When detecting the walking information of pedestrians on a certain street, first read the sequence of videos or pictures to be detected, and then perform object detection to output the class confidence, object position information, and visibility information. Determine whether the class confidence of the object is greater than the high threshold T high , if the class confidence of the object is greater than T high , determine whether the visibility of the object is greater than the threshold T v , if the class confidence of the object is greater than T high , and the visibility of the object is greater than T v , then feature extraction is performed on this object, and it is put into the set Set1 for the first round of matching for feature information matching. If the class confidence of the object is greater than T high , but the visibility of the object is less than T v , it is put into the set Set2 for the second round of matching for position information matching. If the class confidence of the object is less than T high , determine whether the class confidence is greater than the low threshold T low , if the class confidence of the object is less than T high , and the class confidence is greater than T low , then it is put into the set Set3 for the third round of matching for position information matching. If the class confidence of the object is less than T low , discard the object information without processing. After three rounds of matching, the trajectory information is processed according to the status of each trajectory.
[0133] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions, or substitutions made by those skilled in the art within the scope of the essence of the present invention should also fall within the protection scope of the present invention.
Claims
1. A multi-object tracking method including visualizability, characterized in that, In object detection, visibility is used as the target value for learning to predict the visibility of the overall part of the target. The target visibility degree and the target category confidence are combined as the decision value for whether to extract features of the target and use the target features for data matching. The specific algorithm process is as follows: Step 1: Obtain the video stream and read the pictures in chronological order. Step 2: Input the read pictures into the object detection model in chronological order, and then obtain the output of the model, which includes category confidence, target position information, and visibility information. Step 3. If the first frame contains the target and the class confidence of the target is greater than the initialized new trajectory threshold T new , then initialize a new trajectory; if at the same time the visibility of the target is greater than the visibility threshold T v , then crop out the target and input it into the feature extraction model for feature extraction. After processing all the targets in the current image, return to the first step, otherwise go to the next step; Step 4: Determine whether the class confidence of the target in each picture is greater than the high class confidence threshold T high ; Step Five: If the category confidence of the target is greater than the high category confidence threshold T high , determine whether the visibility of the target is greater than the visibility threshold T v ; Step 6. If the class confidence of the target is greater than the high class confidence threshold T high , and the visibility of the target is greater than the visibility threshold T v , then crop the target and input it into the feature extraction model for feature extraction. At this time, the target contains position information, class confidence information, visibility information, and feature information; then put the target information after extraction into the set Set1 of the first-round matching to match the feature information with the trajectory. The trajectory with successful matching updates the trajectory status and related information. The target with successful matching is added to the corresponding trajectory and deleted from the set Set1; the target and trajectory with unsuccessful matching wait for the second-round matching; Step 7: If the class confidence of the target is greater than the high class confidence threshold T high , but the visibility of the target is less than the visibility threshold T v , put it into the set Set2, merge the remaining targets in the set Set1 with the set Set2, and perform a second round of matching. Only the position information is used for this matching; for the successfully matched trajectories, update the trajectory status and relevant information. For the trajectories that are not successfully matched, wait for the third round of matching and delete the successfully matched targets from set2; Step VIII. If the class confidence of the target is less than the high threshold T high , determine whether the class confidence is greater than the low class confidence threshold T low ; Step Nine: If the class confidence of the target is less than the high threshold T high , and the class confidence is greater than the low threshold T low , then put it into the set Set3 for the third-round matching of position information. For the successfully matched trajectories, update the trajectory status and related information. For the successfully matched targets, add them to the corresponding trajectories, update the status of the unsuccessfully matched trajectories, and discard the unsuccessfully matched targets in Set4; Step Ten: If the class confidence of the target is less than the low class confidence threshold T low , discard the target information and do not process it; Step 11: After the above three rounds of matching are completed, delete the trajectories in the Delete state. Step Twelve: Initialize a new trajectory and determine whether the target class confidence in Set2 is greater than the new trajectory initialization threshold T new , if the determination is yes, create a trajectory, otherwise discard it; Step 13: Enter the processing of the next picture, and then repeat steps 2 to 12.
2. The multi-object tracking method including visualization degree according to claim 1, characterized in that, In step 2, the object detection model detects information through the object detection algorithm YOLOX[4]. On the basis of the output head of the object detection algorithm YOLOX[4], a visibility branch head is added in parallel. The object detection algorithm with increased visibility is named YOLOX_v. The output value of the YOLOX_v branch is only the visibility value, and the visibility value is finally output as H*W*(C + 6). The object detection algorithm YOLOX[4] contains a visibility branch head, and the visibility branch head is named YOLOX_v. The output value of the YOLOX_v branch is only the visibility value, and the visibility value is finally output as H*W*(C + 6). Among them, H is the height of the feature map, W is the width of the feature map, C is equal to the number of target categories to be detected, which is the category confidence of each target; 6 includes 1 bit of visibility, 1 bit of confidence in the existence of the target, and 4 bits of position information.
3. A multi-target tracking method including visualization degree according to claim 2, characterized in that, The training method of YOLOX_v is as follows: Step 3.1: Obtain the sample data set. Step 3.2: Build a detection network training model. Step 3.3: Perform object detection and tracking through the training model.
4. A multi-object tracking method including visual degree according to claim 3, characterized in that, The label information of the target is collected in the sample data set. The label information includes category information, position coordinate information, and visibility information. The sample data set is divided into a training set and a test set.
5. A multi-object tracking method including visualization degree according to claim 3, characterized in that The method for building the detection network training model is as follows: Step 3.2.1: First, build a YOLOX_v network model. Step 3.2.2: After adding a visibility branch to YOLOX, modify the detection head of the YOLOX network to be used to additionally predict the visibility of the target on the basis of the original prediction of category confidence, including target confidence, and position information. Step 3.2.3: Add a corresponding loss function for the added visibility branch. Step 3.2.4: Train the modified model. Step 3.2.5: Obtain the optimal weights using the validation set in the improved network model.
6. A multi-object tracking method including visual degree according to claim 5, characterized in that, The method for performing object detection and tracking is as follows: Step 3.3.1: Use the optimal weights to perform frame-by-frame detection on the video to be tracked. Step 3.3.2: Judge whether to input the target into the feature extraction network model for feature extraction according to the detection result. Step 3.3.3: Obtain the real-time trajectories of each target.
7. A multi-object tracking method including visualization degree according to claim 5, characterized in that, The loss function is: In formula (1): L cls represents the class loss, L reg represents the localization loss, L obj represents the loss of whether the target is included, L vis represents the visibility loss, λ represents the balance coefficient of the localization loss, ρ represents the balance coefficient of the visibility loss, N pos represents the number of anchor points classified as positive samples.
8. A multi-object tracking method including visualization degree according to claim 1, characterized in that, When the multi-object tracking method realizes state changes, it uses the trajectory state change method to complete the display and operation of the trajectory state changes during the trajectory and target matching process. Among them, during the matching process, the distance between the new target set of the current round and the original trajectory set of the current round is calculated using feature information or position information. This distance is the condition for judging whether the target and the trajectory match, and the Hungarian matching algorithm is used for optimal matching to make the trajectory find the most suitable target for itself. There are two situations in the matching process. One is that the new target and the original trajectory meet the matching conditions. The other is that there are new targets that do not match the original trajectories, or the original trajectories that match the new targets. It is necessary to process the target and trajectory states in the new target set and the original trajectories under the two matching states.
9. A multi-object tracking method including visualizability according to claim 8, characterized in that, In the trajectory state change method, various states existing in the trajectory change are first determined. The various states include New state, Tracked state, Lost_Tracked state, Delete state, Unformed state, and Lost_Unformed state. The meanings of each state are as follows: New state: The state of a newly created trajectory, which is the first frame of the video that meets the conditions and the target that has not been matched with the trajectory. Tracked state: A trajectory that contains feature information and has matched a new target in the previous frame. Lost_Tracked state: A trajectory that contains feature information but has not been matched with a new target within the previous p frames. Delete state: The state of the trajectory to be deleted. Unformed state: A trajectory that does not contain feature information and has matched a new target in the previous frame. Lost_Unformed state: A trajectory that does not contain feature information but has not been matched with a new target within the previous m frames. According to the above states, the specific implementation content of the trajectory state change method is as follows: Step 6.
1. Process the new target set Dets in the current round. If it is the first frame of the video, or the target whose trajectory has not been matched in three rounds of matching and its class confidence is greater than the new trajectory creation threshold T new , set the status to New, otherwise discard it; Step 6.2: If the trajectories in the New state match the target for n consecutive frames, and there exists a target with a class confidence greater than T high and a visibility greater than the threshold T v i.e., the feature information has been extracted, then set the state to Tracked. Here, n is set to 3; Step 6.3: If the trajectories in the New state match the target for n consecutive frames and there is no target with a class confidence greater than T high and a visibility greater than the threshold T v i.e., there is no feature information, then set the state to Unformed; Step 6.
4. When the trajectory in the Unformed state matches a target with a visibility greater than T v , update the state to Tracked; Step 6.5: When the trajectory in the Tracked state does not match the target, set the trajectory state to Lost_Tracked. Step 6.6: When the trajectory in the Lost_Tracked state matches the target, set the trajectory state to Tracked. Step 6.7: When the trajectory in the Unformed state does not match the target, set the trajectory state to Lost_Unformed. Step 6.8: When the trajectory in the Lost_Unformed state matches the target, if the visibility of the target is greater than the threshold T v and the class confidence is greater than T high , then set the trajectory state to Tracked; Step 6.
9. When the trajectory in the Lost_Unformed state matches the target and the visibility of the target is less than or equal to the threshold T v, the trajectory status is set to Unformed; Step 6.10: When the trajectory in the Lost_Unformed state continuously loses m frames, set the state to Delete. Step 6.11: When the trajectory in the Lost_Tracked state continuously loses p frames, set the state to Delete. Step 6.12: When the trajectory in the New state loses one frame, set the state to Delete.
Citation Information
Patent Citations
An occlusion-resistant target tracking method
CN111080673B
Pedestrian multi-target tracking method and device and computer readable storage medium
CN115240130A
Multi-category multi-target tracking method based on triple matching
CN116152292A