Visual tracking method and device, equipment, storage medium and computer product
By determining the type of abnormal event in the visual tracking chain, identifying the target detection box, and using the feature information during initialization and tracking processes for target re-identification, the problem of re-identification of targets that have been occluded for a long time or re-entering from outside the field of view is solved, thus improving the accuracy and efficiency of identification.
Patent Information
- Application Number
- CN202511449478.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-11
- Publication Date
- 2026-02-27
AI Technical Summary
Existing visual tracking technologies have poor re-identification performance for targets that have been occluded for a long time or have moved out of the field of view for a period of time and then re-entered the field of view. They fail to effectively consider the impact of different tracking anomalies on target recognition, resulting in low recognition accuracy and efficiency.
By determining the type of abnormal event in the tracking chain, the target detection box is determined, and the target is re-identified based on the feature information in the detection box and in the chain. The feature information in the initialization and visual tracking process is combined for comprehensive judgment, which can adapt to a variety of abnormal scenarios.
It improves the accuracy and efficiency of target re-identification, reduces the amount of data processed, and enhances the accuracy and efficiency of visual tracking.
Smart Images

Figure CN121582287A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, specifically to a visual tracking method, apparatus, device, storage medium, and computer product. Background Technology
[0002] Visual tracking is the most widely used target tracking method. Real-time and accurate target tracking can provide information such as the position and distance of the target to the robot. For targets that are continuously unobstructed or only briefly occluded within the robot's field of vision, basic trajectory prediction algorithms can achieve the required prediction and update results. Existing visual tracking technologies have poor re-identification performance for targets that are occluded for a long time or have moved out of the field of vision for a period of time and then re-enter it. This is because they use short-term appearance feature information for target re-identification, do not consider the impact of different tracking anomalies on target recognition, and rely solely on appearance feature information for re-identification, which reduces the processing efficiency of the algorithm. Therefore, it is necessary to propose a new visual tracking method to improve the accuracy and efficiency of target recognition, thereby improving the accuracy and efficiency of visual tracking. Summary of the Invention
[0003] This application provides a visual tracking method, apparatus, device, storage medium, and computer product, which can improve the accuracy and efficiency of target recognition, thereby improving the accuracy and efficiency of visual tracking.
[0004] In a first aspect, embodiments of this application provide a visual tracking method, including: In a tracking chain based on a tracking target, if an abnormal tracking event is determined, a target detection box is determined based on the type of the abnormal tracking event; Based on the feature information of the object to be identified within the target detection box and the feature information of the tracking target, the tracking target is re-identified, and / or, based on the information of the target detection box and the information of the detection boxes contained in the tracking chain, the tracking target is re-identified. The feature information of the tracked target includes first feature information obtained after initializing the information of the tracked target and second feature information obtained during the visual tracking process of the tracked target.
[0005] As an example, the tracking anomaly event includes one or more of a first tracking anomaly event, a second tracking anomaly event, and a third tracking anomaly event. The first tracking anomaly event is used to indicate that the tracking target is occluded, the second tracking anomaly event is used to indicate that the tracking target is lost, and the third tracking anomaly event is used to indicate that the tracking target is interfered with.
[0006] As one embodiment, determining the target detection box based on the type of the tracked anomaly event includes: If the tracking anomaly event includes the first tracking anomaly event and / or the second tracking anomaly event, and the duration of the existence of the first tracking anomaly event and / or the second tracking anomaly event is less than a preset duration, the detection box before the occurrence of the tracking anomaly event is inflated, and the inflated detection box is used as the target detection box. If the tracking anomaly event includes the second tracking anomaly event, a detection box is generated as the target detection box; In the case where the tracking anomaly event includes the third tracking anomaly event, the detection box prior to the occurrence of the tracking anomaly event is used as the target detection box.
[0007] As one embodiment, the re-identification of the tracked target based on the feature information of the object to be identified within the target detection box and the feature information of the tracked target includes: Determine the first similarity between the feature information of the object to be identified and the first feature information and / or the second feature information; If the first similarity is greater than the similarity threshold, the object to be identified is determined to be the tracking target.
[0008] As an example, the similarity threshold is determined based on the second feature information.
[0009] As one embodiment, the second feature information is determined based on the following method: If no tracking anomalies are found in the tracking chain based on the tracked target, all reliable frames corresponding to the tracking chain are obtained; the reliable frames are determined based on the intersection-over-union ratio of the images, the confidence of the detection boxes, and the similarity of the feature information. Cluster the trusted frames and extract key frames from the clustered trusted frames; Feature extraction is performed on the detection bounding box of the keyframe to obtain the second feature information.
[0010] As one embodiment, the re-identification of the tracked target based on the feature information of the object to be identified within the target detection box and the feature information of the tracked target includes: Determine a first similarity between the feature information of the object to be identified and the feature information of the tracked target, and a second similarity between the feature information of the object to be identified within the target detection box and the feature information of the non-tracked target; Based on the first similarity and the second similarity, a vote is performed to determine the relationship between the object to be identified and the tracking target.
[0011] As one embodiment, the re-identification of the tracked target based on the information of the target detection box and the information of the detection boxes included in the tracking chain includes: Based on the relative position information of the target detection box and the relative position information of the detection boxes contained in the tracking chain, the tracked target is re-identified.
[0012] As an example, based on the tracking chain, determining the existence of a tracking anomaly event includes: Determine the number, area, and location of the detection frames contained in the tracking chain; Based on the number, area, and position of the detection frames contained in the tracking chain, it is determined whether the tracking anomaly event exists.
[0013] Secondly, embodiments of this application provide a visual tracking device, comprising: The determination module is used to determine a target detection box based on the type of the tracking anomaly event when it is determined that there is a tracking anomaly event in the tracking chain based on the tracking target. The re-identification module is used to re-identify the tracking target based on the feature information of the object to be identified within the target detection box and the feature information of the tracking target, and / or to re-identify the tracking target based on the information of the target detection box and the information of the detection boxes contained in the tracking chain. The feature information of the tracked target includes first feature information obtained after initializing the information of the tracked target and second feature information obtained during the visual tracking process of the tracked target.
[0014] Thirdly, embodiments of this application provide an electronic device, including a processor and a memory storing a computer program, wherein the processor executes the program to implement the steps of the visual tracking method described in the first or second aspect.
[0015] Fourthly, embodiments of this application provide a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the visual tracking method described in the first or second aspect.
[0016] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the visual tracking method described in the first aspect.
[0017] The visual tracking method, apparatus, device, storage medium, and computer product provided in this application, when determining the existence of tracking anomalies based on the tracking chain of the tracking target, determine a target detection box based on the type of the tracking anomaly event; perform re-identification of the tracking target based on the feature information of the object to be identified within the target detection box and the feature information of the tracking target; and / or, perform re-identification of the tracking target based on the information of the target detection box and the information of the detection boxes included in the tracking chain; wherein, the feature information of the tracking target includes first feature information obtained after initializing the tracking target and second feature information obtained during the visual tracking process of the tracking target. This application determines the target detection box through different types of tracking anomalies and performs target re-identification of the object to be identified within the target detection box. Considering the impact of different tracking anomalies on target identification, it is beneficial to improve the accuracy of target re-identification. On this basis, target re-identification is performed using long-term and short-term feature information composed of the first feature information of the initialized tracking target and the second feature information during the visual tracking process. Compared with simple short-term feature information, the re-identification accuracy is higher. By performing re-identification of the tracking target using the information of the target detection box and the information of the detection boxes included in the tracking chain, the amount of data to be processed is greatly reduced, improving the re-identification efficiency. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a flowchart illustrating the visual tracking method provided in the embodiments of this application.
[0020] Figure 2 This is a schematic diagram of the structure of the visual tracking device provided in the embodiments of this application.
[0021] Figure 3 This is a schematic diagram of the structure of the visual tracking system provided in the embodiments of this application.
[0022] Figure 4 This is a schematic diagram of the visual tracking process of the visual tracking system provided in this application embodiment.
[0023] Figure 5 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0025] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same class, not limited in number; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0026] In robot-following scenarios, due to obstacles, occlusion by other people or objects, and environmental changes, it is easy to lose track of or mistrack targets. Existing robot visual tracking technologies typically use historical position or appearance information from the previous frame or a short period to re-identify lost targets. The identification factors and criteria are relatively simple. For example, deep neural network models are used to extract the appearance features of newly entered targets, calculate the similarity between these features and the appearance features of the lost target at a certain point in the past, and use a threshold method to determine whether it is the target that was tracked before it was lost. The application of a single tracking chain and a single threshold method makes the tracking results prone to fluctuations and abrupt changes in certain scenarios. In addition, the addition of appearance feature extraction leads to a decrease in algorithm processing efficiency.
[0027] Specifically, existing robot vision tracking technologies maintain state features through short-term tracking chains, integrating features from the previous frame with momentum into the current features. Taking a frame rate of 30fps and a momentum parameter of 0.9 as an example, calculations show that features from one second ago only have a 0.04 (0.9^30) influence on the current state, insufficient as long-term features for auxiliary judgment. In cases of sudden environmental changes, the features of the tracked target can vary significantly. In such situations, linear tracking chain updates and simple thresholding methods are prone to misjudgment and target loss when similarity parameters fluctuate drastically. Furthermore, when the target is occluded or enters the frame from the edge, only a portion of its appearance is used for feature extraction and recognition, easily leading to index errors and misjudgments.
[0028] To address this, this application provides a visual tracking method, apparatus, device, storage medium, and computer product that, in the event of a target loss incident, takes corresponding re-matching measures for the scenario of the target loss incident to improve the accuracy of target tracking, thereby improving the accuracy of visual tracking and solving the problem of target loss or mismatch during robot visual tracking.
[0029] This application provides a visual tracking method applicable to robot tracking scenarios, which may include steps S110-S120.
[0030] Step S110: If a tracking anomaly event is determined based on the tracking chain of the tracking target, a target detection box is determined based on the type of the tracking anomaly event.
[0031] Optionally, the tracking target is the object being followed in the robot tracking scenario. The robot captures video or images, identifies the objects in the video or images based on a target detection algorithm, and moves with the objects. The tracking target can be a human, vehicle, animal, etc. This application embodiment uses a human as an example for illustration.
[0032] Optionally, the tracking chain is used to characterize the state (such as position, size, appearance, etc.) of the tracked target in different frames of a video sequence or image sequence, and associates them to form a continuous trajectory. When the target detection algorithm identifies objects in a video or image, it generates detection boxes. The tracking chain in this embodiment includes the state information of the detection boxes in different frames of the video sequence or image sequence, including but not limited to the position and size of the detection boxes. The amount of related data for the detection boxes is less, the determination rate is higher, and it is beneficial to improve the efficiency of judging tracking abnormal events.
[0033] Optionally, before step S110, it is necessary to determine whether there are any tracking anomalies based on the tracking chain of the tracked target. Tracking anomalies include, but are not limited to, target loss, target occlusion, and sudden changes in target position. When the tracking chain includes the state information of the detection boxes, the existence of tracking anomalies can be determined by judging whether the detection boxes in the video or image are lost, whether the detection boxes are occluded, and whether the positions of the detection boxes in different frames change abruptly. It is simple and convenient to determine whether there are any anomalies in the tracked target by only judging the tracking chain.
[0034] Optionally, after a tracking anomaly occurs, the tracking target needs to be re-identified, that is, the object in the video or image needs to be identified to determine whether the object is the tracking target. The embodiments of this application determine the target detection box based on the type of tracking anomaly, which can adapt to various tracking anomaly scenarios.
[0035] Step S120: Based on the feature information of the object to be identified within the target detection box and the feature information of the tracking target, perform re-identification of the tracking target, and / or, based on the information of the target detection box and the information of the detection boxes contained in the tracking chain, perform re-identification of the tracking target.
[0036] The feature information of the tracked target includes first feature information obtained after initializing the information of the tracked target and second feature information obtained during the visual tracking process of the tracked target.
[0037] Optionally, in step S120, the re-identification method for the tracked target can also be associated with the type of tracking anomaly event to further adapt to various tracking anomaly scenarios, reduce the amount of data processing and improve the re-identification efficiency while ensuring the accuracy of re-identification.
[0038] Furthermore, in cases of tracking anomalies, including target loss or target occlusion, the tracking target can be re-identified based on the feature information of the object to be identified within the target detection box and the feature information of the tracking target. In cases of tracking anomalies, including sudden changes in the tracking target's position, the tracking target can be re-identified based on the information of the target detection box and the information of the detection boxes contained in the tracking chain; alternatively, the tracking target can be re-identified based on the feature information of the object to be identified within the target detection box and the feature information of the tracking target.
[0039] It should be noted that, regardless of the type of tracking anomaly, the tracking target can be re-identified based on the feature information of the object to be identified within the target detection box, the feature information of the tracking target, the information of the target detection box, and the information of the detection boxes contained in the tracking chain.
[0040] Optionally, the tracking target can be re-identified based on the feature information of the object to be identified within the target detection box and the feature information of the tracking target. Specifically, the tracking target can be re-identified by calculating the similarity between the feature information of the object to be identified within the target detection box and the feature information of the tracking target.
[0041] Optionally, based on the information of the target detection box and the information of the detection boxes contained in the tracking chain, the tracking target is re-identified. Specifically, the position of the target detection box is remapped to the position of the original detection box to achieve the re-identification of the tracking target.
[0042] Understandably, this application determines target detection boxes through different types of tracking anomalies and performs target re-identification on the objects to be identified within the target detection boxes. Considering the impact of different tracking anomalies on target identification, it is beneficial to improve the accuracy of target re-identification. On this basis, target re-identification is performed using long and short-term feature information composed of the first feature information of the initialized tracking target and the second feature information during the visual tracking process. Compared with simple short-term feature information, the re-identification accuracy is higher. By using the information of the target detection box and the information of the detection boxes contained in the tracking chain to perform re-identification of the tracking target, the amount of data to be processed is greatly reduced, thus improving the efficiency of re-identification.
[0043] As an example, the tracking anomaly event includes one or more of a first tracking anomaly event, a second tracking anomaly event, and a third tracking anomaly event. The first tracking anomaly event is used to indicate that the tracking target is occluded, the second tracking anomaly event is used to indicate that the tracking target is lost, and the third tracking anomaly event is used to indicate that the tracking target is interfered with.
[0044] Optionally, the reasons for the loss of the tracked target include occlusion, obstacle avoidance, signal transmission failure, interference, etc. Interference includes environmental interference and signal interference. Furthermore, the third tracking anomaly event is used to characterize the sudden change in the position of the detection box caused by interference to the tracked target.
[0045] It is understood that the embodiments of this application are designed to adapt to various abnormal events, limit the types of abnormal events to be tracked, and provide a basis for improving the accuracy of target re-identification.
[0046] As an example, based on the tracking chain, determining the existence of a tracking anomaly event includes: Determine the number, area, and location of the detection frames contained in the tracking chain; Based on the number, area, and position of the detection frames contained in the tracking chain, it is determined whether the tracking anomaly event exists.
[0047] Optionally, in this embodiment of the application, a coordinate system is constructed for the video or image (i.e., the robot's field of view). The tracking target in the video or image is identified based on the target detection algorithm, and a detection box is generated. The coordinates of the detection box are determined, typically represented as XYXY (a four-dimensional vector, recording the coordinate values of the upper left and lower right corners in a two-dimensional plane, respectively). The area and position of the detection box can be determined by the coordinates of the detection box.
[0048] Optionally, based on the number, area, and position of the detection boxes included in the tracking chain, it is determined whether one or more of the following situations occur: a decrease in the number of detection boxes, a decrease in the area of the detection boxes, an overlap of the position of the detection boxes with the robot's field of view boundary, or loss and reconstruction of detection boxes. If so, a tracking anomaly event is determined. Under normal circumstances, each tracking target corresponds to one detection box. If there are 0 detection boxes, it can be determined that the tracking target is lost. If the detection box of the tracking target continues to decrease and overlaps with other detection boxes, it can be determined that the tracking target is occluded. If the area of the detection box continues to decrease and intersects with the left and right boundaries of the field of view, it can be considered that the tracking target has been lost due to lateral movement away from the field of view. When a large number of non-lateral detection boxes are lost and newly created within the robot's field of view, it can be considered that the tracking target is disturbed, resulting in a sudden change in its position within the field of view.
[0049] It is understood that this application determines whether the tracking anomaly event exists by measuring the number, area, and position of the detection boxes contained in the tracking chain. It is simple and convenient to determine whether the tracking target is abnormal by judging only a single factor of the tracking chain.
[0050] As one embodiment, determining the target detection box based on the type of the tracked anomaly event includes: If the tracking anomaly event includes the first tracking anomaly event and / or the second tracking anomaly event, and the duration of the existence of the first tracking anomaly event and / or the second tracking anomaly event is less than a preset duration, the detection box before the occurrence of the tracking anomaly event is inflated, and the inflated detection box is used as the target detection box. If the tracking anomaly event includes the second tracking anomaly event, a detection box is generated as the target detection box; In the case where the tracking anomaly event includes the third tracking anomaly event, the detection box prior to the occurrence of the tracking anomaly event is used as the target detection box.
[0051] Optionally, when there is only one tracking target, the tracking anomaly event includes either the first tracking anomaly event or the second tracking anomaly event. When there are multiple tracking targets, the tracking anomaly event includes both the first tracking anomaly event and / or the second tracking anomaly event. In Embodiment 1 of this application, when there is only one tracking target and the tracking anomaly event includes either the first tracking anomaly event or the second tracking anomaly event, i.e., when the tracking target is temporarily occluded and / or temporarily lost, the detection box before the occurrence of the tracking anomaly event is inflated, and the inflated detection box is used as the target detection box. There is no need to execute the target detection algorithm, which can reduce the amount of data processing.
[0052] Optionally, if the tracked target is lost due to obstacle avoidance, signal transmission failure, or other reasons, or if the tracked target is lost for a period of time exceeding a preset duration, the detection box is regenerated based on the target detection algorithm.
[0053] Optionally, if the target location changes abruptly, but the number of detection boxes remains the same, and the mapping relationship between the detection boxes and the tracked target changes, remapping can be performed. Furthermore, to improve accuracy, if the duration of a third tracking anomaly event exceeds a preset duration, detection boxes are regenerated based on the target detection algorithm to obtain the target detection boxes.
[0054] It is understood that the embodiments of this application determine the type of tracking anomaly and take corresponding measures to achieve more accurate matching of tracking targets, thus providing a basis for re-identification.
[0055] As one embodiment, the re-identification of the tracked target based on the feature information of the object to be identified within the target detection box and the feature information of the tracked target includes: Determine the first similarity between the feature information of the object to be identified and the first feature information and / or the second feature information; If the first similarity is greater than the similarity threshold, the object to be identified is determined to be the tracking target.
[0056] Optionally, this application does not limit the similarity calculation algorithm, and the first similarity can be calculated using cosine similarity or Euclidean distance.
[0057] Optionally, the first similarity includes the similarity between the feature information of the object to be identified and the first feature information, and / or the similarity between the feature information of the object to be identified and the second feature information.
[0058] It is understood that the embodiments of this application calculate the first similarity between the feature information of the object to be identified and the first feature information and / or the second feature information, and determine whether the object to be identified is a tracking target based on the first similarity. This can realize long-term and / or short-term feature comparison, achieve comprehensive discrimination, and improve the accuracy of feature comparison and re-identification.
[0059] As an example, the similarity threshold is determined based on the second feature information.
[0060] Optionally, the preset single similarity threshold can be adjusted based on the similarity difference between the second feature information in different tracking stages, or the similarity threshold can be determined by weighted summation of the similarity between the second feature information in different tracking stages.
[0061] It is understood that the embodiments of this application determine the similarity threshold of the current tracking task through the second feature information in the current tracking process. Compared with a fixed single threshold, it can handle scenarios where the features of the object to be identified and the tracking target are relatively similar, thereby improving the accuracy of target re-identification.
[0062] As one embodiment, the re-identification of the tracked target based on the feature information of the object to be identified within the target detection box and the feature information of the tracked target includes: Determine a first similarity between the feature information of the object to be identified and the feature information of the tracked target, and a second similarity between the feature information of the object to be identified within the target detection box and the feature information of the non-tracked target; Based on the first similarity and the second similarity, a vote is performed to determine the relationship between the object to be identified and the tracking target.
[0063] Optionally, the feature information of non-tracked targets includes the feature information of objects during the target detection algorithm training phase and / or the feature information of objects appearing in the robot's field of view during tracking. It should be noted that the embodiments of this application do not limit the method for calculating similarity.
[0064] Optionally, the feature information of the tracked target includes first feature information and second feature information. The number of non-tracked targets is greater than 1, so the number of first similarity and second similarity is greater than 1.
[0065] Furthermore, the first similarity and the second similarity are ranked, and a vote is taken based on the results of the higher rankings to determine whether the object to be identified is the tracking target. Combining multiple recall ranking and voting can improve the accuracy of re-identification.
[0066] Understandably, this application can further improve the accuracy of re-identification by calculating the similarity of positive and negative feature information and combining it with voting.
[0067] As one embodiment, the second feature information is determined based on the following method: If no tracking anomalies are found in the tracking chain based on the tracked target, all reliable frames corresponding to the tracking chain are obtained; the reliable frames are determined based on the intersection-over-union ratio of the images, the confidence of the detection boxes, and the similarity of the feature information. Cluster the trusted frames and extract key frames from the clustered trusted frames; Feature extraction is performed on the detection bounding box of the keyframe to obtain the second feature information.
[0068] Optionally, in this embodiment of the application, the product of the intersection-union ratio of the image, the confidence of the detection box, and the similarity of the feature information is used as the tracking score of the image frame, and the image frame with a tracking score greater than a threshold is used as a reliable frame.
[0069] It is understandable that this application clusters trusted frames and extracts one or more trusted frames from each group of trusted frames as key frames. The key frames have a large degree of similarity difference, which is beneficial for determining different feature information of the tracking target.
[0070] As one embodiment, the re-identification of the tracked target based on the information of the target detection box and the information of the detection boxes included in the tracking chain includes: Based on the relative position information of the target detection box and the relative position information of the detection boxes contained in the tracking chain, the tracked target is re-identified.
[0071] It is credible that the relative position of the tracked target within the robot's field of view will not change abruptly in a short period of time. For temporary changes in the position of the tracked target caused by environmental or signal interference, the detection box of the target that is close to the detection box contained in the tracking chain can be used as the detection box of the tracked target.
[0072] It is understood that, by comparing the relative position information of the target detection box and the relative position information of the detection boxes contained in the tracking chain, the embodiments of this application can achieve the re-identification of tracking targets with sudden changes in position in a short period of time, reduce the amount of data processing, and improve the identification efficiency.
[0073] The visual tracking device provided in the embodiments of this application is described below. The visual tracking device described below can be referred to in correspondence with the visual tracking method described above.
[0074] Reference Figure 2 This application provides a visual tracking device, including: The determination module 210 is used to determine a target detection box based on the type of the tracking anomaly event when it is determined that there is a tracking anomaly event in the tracking chain based on the tracking target. The re-identification module 220 is used to re-identify the tracking target based on the feature information of the object to be identified within the target detection box and the feature information of the tracking target, and / or to re-identify the tracking target based on the information of the target detection box and the information of the detection boxes contained in the tracking chain. The feature information of the tracked target includes first feature information obtained after initializing the information of the tracked target and second feature information obtained during the visual tracking process of the tracked target.
[0075] As an example, the tracking anomaly event includes one or more of a first tracking anomaly event, a second tracking anomaly event, and a third tracking anomaly event. The first tracking anomaly event is used to indicate that the tracking target is occluded, the second tracking anomaly event is used to indicate that the tracking target is lost, and the third tracking anomaly event is used to indicate that the tracking target is interfered with.
[0076] As one embodiment, the determining module 210 is further configured to: If the tracking anomaly event includes the first tracking anomaly event and / or the second tracking anomaly event, and the duration of the existence of the first tracking anomaly event and / or the second tracking anomaly event is less than a preset duration, the detection box before the occurrence of the tracking anomaly event is inflated, and the inflated detection box is used as the target detection box. If the tracking anomaly event includes the second tracking anomaly event, a detection box is generated as the target detection box; In the case where the tracking anomaly event includes the third tracking anomaly event, the detection box prior to the occurrence of the tracking anomaly event is used as the target detection box.
[0077] As one embodiment, the re-identification module 220 is further configured to: Determine the first similarity between the feature information of the object to be identified and the first feature information and / or the second feature information; If the first similarity is greater than the similarity threshold, the object to be identified is determined to be the tracking target.
[0078] As an example, the similarity threshold is determined based on the second feature information.
[0079] As one embodiment, the determining module 210 is further configured to: If no tracking anomalies are found in the tracking chain based on the tracked target, all reliable frames corresponding to the tracking chain are obtained; the reliable frames are determined based on the intersection-over-union ratio of the images, the confidence of the detection boxes, and the similarity of the feature information. Cluster the trusted frames and extract key frames from the clustered trusted frames; Feature extraction is performed on the detection bounding box of the keyframe to obtain the second feature information.
[0080] As one embodiment, the re-identification module 220 is further configured to: Determine a first similarity between the feature information of the object to be identified and the feature information of the tracked target, and a second similarity between the feature information of the object to be identified within the target detection box and the feature information of the non-tracked target; Based on the first similarity and the second similarity, a vote is performed to determine the relationship between the object to be identified and the tracking target.
[0081] As one embodiment, the re-identification module 220 is further configured to: Based on the relative position information of the target detection box and the relative position information of the detection boxes contained in the tracking chain, the tracked target is re-identified.
[0082] As one embodiment, the determining module 210 is further configured to: Determine the number, area, and location of the detection frames contained in the tracking chain; Based on the number, area, and position of the detection frames contained in the tracking chain, it is determined whether the tracking anomaly event exists.
[0083] The visual tracking system provided in the embodiments of this application is described below. The visual tracking system described below can be referred to in correspondence with the visual tracking method and apparatus described above.
[0084] Reference Figure 3 This application also provides a visual tracking system, including a data acquisition module, a visual detection and tracking algorithm module, a database module, a motion control module, and a user interaction module.
[0085] The data acquisition module includes at least one camera for acquiring video or image data; the motion control module is used to receive and execute motion commands; and the user interaction module includes at least one display screen for information confirmation display and command reception and processing.
[0086] The database module includes a user database, a system database, and a temporary database. The user database stores feature data related to system users, which refer to users who can be tracked. The system database stores training data for the visual detection and tracking algorithm module. The temporary database stores feature data of both tracked and non-tracked targets during the tracking process. This embodiment uses appearance feature images as an example for illustration. The user database stores appearance feature images of system users entered during user information initialization and appearance feature images of system users entered during tracking.
[0087] The visual detection and tracking algorithm module includes components such as object detection, command detection, trajectory estimation, limb detection, appearance feature extraction, target localization, path planning and navigation, keyframe extraction, target tracking, occlusion (event) detection, and image retrieval. The object detection component contains one or more object detection models, represented by YOLO (or Multitask Cascaded Convolutional Networks, MTCNN), which detects the coordinates of the bounding box of a human body (or face) within the robot's field of vision when it appears. The command detection component includes a gesture detection model to recognize tracking commands or information confirmation commands sent to the robot by the system user through gestures or behaviors such as waving. The trajectory estimation component includes a human motion estimation algorithm, represented by Kalman filtering, which predicts the bounding box position of the target being tracked in the next frame based on the past state of the tracking chain and calculates the intersection-over-union (IoU) value with all bounding boxes identified by the object detection component in the current frame. The limb detection component includes... Limb detection algorithms, such as RTMPose, can extract key points of the human torso to determine whether the detection box covers most of the body. The appearance feature extraction component contains one or more appearance feature extraction models, such as CLIP-ReID and FaceNet, which can map the detection box image into a feature vector of a certain dimension. The similarity of appearance features between different detection boxes is calculated by using cosine similarity or Euclidean distance and other distance measurement methods. A threshold is used to determine whether it is a tracking target. The target localization, path planning and navigation components use the target position information provided by video tracking to locate the tracking target and plan the path. Combined with the motion control module, the robot can move and follow the target.
[0088] Among them, RTMPose is a real-time, high-precision human pose estimation algorithm that can quickly detect human key points; CLIP-ReID is an image re-identification method that uses a vision-language model to enhance feature discrimination through cross-modal learning, without requiring specific text labels; and FaceNet is a deep learning-based face recognition method that maps face images to Euclidean space and directly optimizes the distance between feature vectors to achieve face verification and recognition.
[0089] The target tracking component implements a single-target tracking algorithm based on the multi-target tracking framework BoT-SORT. It detects and classifies tracking anomalies based on the tracking chain of the target. For tracked targets experiencing occlusion events, when the tracking chain state is set to lost, trajectory estimation and prediction box updates are not performed. Instead, image retrieval and target re-matching are performed on the expanded detection box or newly appearing detection boxes within the expanded box, through the slow expansion of the detection box before the loss. For detection boxes that re-enter the field of view from the side when the tracking chain is lost, a new tracking chain is created for them. The limb detection component is used to detect the similarity between the appearance features of the corresponding detection box of the new tracking chain and the appearance features of the original lost tracking chain. For sudden changes in the position of the tracked target caused by environmental changes or signal interference, remapping is performed by combining the relative position of the detection box and the relative position relationship of the original tracking chain, as well as appearance features. Specifically, the detection box can be remapped to the identity ID of the system user.
[0090] The event detection component determines whether the tracking chain has experienced occlusion, lateral loss, or sudden interference events by judging whether the number of human detection boxes generated in a given frame has decreased or whether the detection boxes overlap compared to the previous frame. Specifically, when the area of the matching detection box of the tracking chain continuously decreases and overlaps with another detection box, the tracking chain is considered lost due to occlusion; when the area of the matching detection box of the tracking chain continuously decreases and intersects with the left and right boundaries, the tracking chain is considered lost due to lateral movement out of the field of view; when a large number of non-lateral tracking chain losses and new creations occur within the field of view, the tracking chain is considered lost due to ID abrupt changes caused by interference.
[0091] The keyframe extraction component performs optimal number clustering on the feature vectors of all credible frames within the tracking chain of a tracked target. After compressing the number of frames, it extracts the frames with the largest similarity differences as keyframes for target recognition. A credible frame is a frame with a tracking score greater than a threshold. When the trajectory can be estimated, the tracking score is recorded as the product of the IoU value, the confidence of the detection box, and the similarity of the appearance features.
[0092] The image retrieval component uses the appearance features extracted by the appearance feature component to calculate the similarity between the features of the current object to be identified and all comparable objects in the image library. Based on the retrieval results ranked first by similarity, the identity ID of the object to be identified is determined by majority voting.
[0093] Reference Figure 4 The visual tracking method implemented in this application based on a visual tracking system may include steps 1-8: Step 1: Perform user initialization and trace initialization.
[0094] User initialization refers to inputting the appearance features and identity ID of new system users into the robot. Specifically, the username (identity ID) can be entered through the user interaction module, and facial feature information, body feature information, etc., can be input through the data acquisition module.
[0095] Tracking initialization refers to the process where a system user needs to complete preset actions when using the tracking function for the first time in order to extract the user's appearance features such as hairstyle and clothing.
[0096] Specifically, system users can stand approximately 1.5 meters away from the robot, facing it directly, and position themselves in the center of the robot's field of vision via the screen display, then naturally turn 360 degrees. The keyframe extraction component will collect multiple images into the appearance features folder under the system user directory within the user image library component.
[0097] Step 2: Based on the trajectory estimation component, update the tracking chain of the tracking target during the normal tracking process (i.e., without occlusion, loss, environmental interference, etc.), and use the keyframe extraction component to record the target keyframes during the current tracking process into the tracking target directory in the temporary image library component to extract reference features.
[0098] Step 3: When the occlusion detection component detects an occlusion event of a target within the field of view, within a certain time window, based on the trajectory estimation information, the number of detection boxes near the occlusion range is restored to the level before the occlusion event occurred. The relative position of the tracked target is determined by the IoU between the trajectory estimation detection box and the actual detection box, as well as the similarity of appearance features. When the target is occluded for more than the time window, the trajectory estimation information is lost, and the process proceeds to step 7.
[0099] Step 4: After the tracking target is lost due to obstacle avoidance, signal transmission failure, or other reasons, the robot waits in place for the tracking target to return or actively changes its perspective by rotating and moving so that an unknown human body enters the field of view from the side. The human body detection box should be detected by the target detection component, and then proceed to step 7.
[0100] Step 5: When environmental interference events such as changes in lighting occur, adjust the brightness and contrast of the current frame during processing by changing the background light within a certain time window; if the tracking target position in the field of view changes abruptly due to interference, the number of detection boxes usually does not change, and the changed tracking target position is remapped back to the original position by sorting the similarity correlation; if the interference lasts for too long, the trajectory estimation information is lost after a certain time window, then proceed to step 7.
[0101] Step 6: If the target being tracked is within the field of view and there are no obstructions, scene changes, or other interferences, when a new human detection box appears in the field of view, it is extracted through the keyframe extraction component and entered into the pedestrian directory in the temporary image library component as a reference feature.
[0102] Step 7: When target re-identification is required, the limb detection component is used to detect the limb joints within the detection box. When all limb joints can be detected, the image retrieval component is used to search the detection box in the appearance feature folder, temporary image library, and system image library of the user directory corresponding to the currently tracked target in the user image library component. The results with the highest similarity and confidence are returned. If more than half of the results correspond to the current tracked target, it is determined to be the tracked target and the normal visual tracking process is entered; otherwise, the process continues to find a new target detection box by moving or waiting for a new target detection box to repeat the steps.
[0103] Step 8: When a follow command is received or follow mode is terminated for other reasons, stop data recording and motion tracking, and update the image library according to the following rules.
[0104] Rule 1: Calculate the feature similarity between all appearance features in the target directory of the temporary image library component, and select the images with the greatest differences to be included in the appearance feature folder in the user directory corresponding to the currently tracked target in the user image library component.
[0105] Rule 2: Calculate the similarity between all appearance features in the "Passersby" category of the temporary image library component and all reference features in the system image library component. The images with the largest similarity differences are included in the system image library component as reference features.
[0106] Rule 3: Calculate the similarity between all appearance features in the "Passersby" directory of the temporary image library component and all appearance features in the target directory. The images with the smallest similarity differences are included in the system image library component as reference features.
[0107] Understandably, this application improves the stability of tracking by using event detection to analyze the reasons for target loss during video tracking and taking re-matching measures for different loss scenarios. By introducing user image libraries and temporary image libraries to record long-term features of the tracked target and integrating them into the target re-matching process, the recall rate of target re-identification is improved.
[0108] Figure 5 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 5 As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530, and a communication bus 540, wherein the processor 510, the communication interface 520, and the memory 530 communicate with each other via the communication bus 540. The processor 510 can call a computer program in the memory 530 to execute the steps of the visual tracking method, including: In a tracking chain based on a tracking target, if an abnormal tracking event is determined, a target detection box is determined based on the type of the abnormal tracking event; Based on the feature information of the object to be identified within the target detection box and the feature information of the tracking target, the tracking target is re-identified, and / or, based on the information of the target detection box and the information of the detection boxes contained in the tracking chain, the tracking target is re-identified. The feature information of the tracked target includes first feature information obtained after initializing the information of the tracked target and second feature information obtained during the visual tracking process of the tracked target.
[0109] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0110] On the other hand, embodiments of this application also provide a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can perform the steps of the visual tracking method provided in the above embodiments, including: In a tracking chain based on a tracking target, if an abnormal tracking event is determined, a target detection box is determined based on the type of the abnormal tracking event; Based on the feature information of the object to be identified within the target detection box and the feature information of the tracking target, the tracking target is re-identified, and / or, based on the information of the target detection box and the information of the detection boxes contained in the tracking chain, the tracking target is re-identified. The feature information of the tracked target includes first feature information obtained after initializing the information of the tracked target and second feature information obtained during the visual tracking process of the tracked target.
[0111] On the other hand, embodiments of this application also provide a processor-readable storage medium storing a computer program for causing a processor to perform the steps of the methods provided in the above embodiments, including: In a tracking chain based on a tracking target, if an abnormal tracking event is determined, a target detection box is determined based on the type of the abnormal tracking event; Based on the feature information of the object to be identified within the target detection box and the feature information of the tracking target, the tracking target is re-identified, and / or, based on the information of the target detection box and the information of the detection boxes contained in the tracking chain, the tracking target is re-identified. The feature information of the tracked target includes first feature information obtained after initializing the information of the tracked target and second feature information obtained during the visual tracking process of the tracked target.
[0112] The processor-readable storage medium can be any available medium or data storage device that the processor can access, including but not limited to magnetic memory (e.g., floppy disk, hard disk, magnetic tape, magneto-optical disk (MO)), optical memory (e.g., CD, DVD, BD, HVD), and semiconductor memory (e.g., ROM, EPROM, EEPROM, non-volatile memory (NAND FLASH), solid-state drive (SSD)).
[0113] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0114] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0115] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A visual tracking method, characterized in that, include: In a tracking chain based on a tracking target, if an abnormal tracking event is determined, a target detection box is determined based on the type of the abnormal tracking event; Based on the feature information of the object to be identified within the target detection box and the feature information of the tracking target, the tracking target is re-identified, and / or, based on the information of the target detection box and the information of the detection boxes contained in the tracking chain, the tracking target is re-identified. The feature information of the tracked target includes first feature information obtained after initializing the information of the tracked target and second feature information obtained during the visual tracking process of the tracked target.
2. The visual tracking method according to claim 1, characterized in that, The tracking anomaly event includes one or more of a first tracking anomaly event, a second tracking anomaly event, and a third tracking anomaly event. The first tracking anomaly event is used to indicate that the tracking target is occluded, the second tracking anomaly event is used to indicate that the tracking target is lost, and the third tracking anomaly event is used to indicate that the tracking target is interfered with.
3. The visual tracking method according to claim 2, characterized in that, The step of determining the target detection box based on the type of the tracked anomaly event includes: If the tracking anomaly event includes the first tracking anomaly event and / or the second tracking anomaly event, and the duration of the existence of the first tracking anomaly event and / or the second tracking anomaly event is less than a preset duration, the detection box before the occurrence of the tracking anomaly event is inflated, and the inflated detection box is used as the target detection box. If the tracking anomaly event includes the second tracking anomaly event, a detection box is generated as the target detection box; In the case where the tracking anomaly event includes the third tracking anomaly event, the detection box prior to the occurrence of the tracking anomaly event is used as the target detection box.
4. The visual tracking method according to claim 1, characterized in that, The step of re-identifying the tracked target based on the feature information of the object to be identified within the target detection box and the feature information of the tracked target includes: Determine the first similarity between the feature information of the object to be identified and the first feature information and / or the second feature information; If the first similarity is greater than the similarity threshold, the object to be identified is determined to be the tracking target.
5. The visual tracking method according to claim 4, characterized in that, The similarity threshold is determined based on the second feature information.
6. The visual tracking method according to claim 1, characterized in that, The second feature information is determined based on the following method: If, based on the tracking chain of the tracking target, it is determined that there are no tracking anomalies, all trusted frames corresponding to the tracking chain are obtained. The trusted frame is determined based on the image's intersection-over-union ratio, detection box confidence, and feature information similarity. Cluster the trusted frames and extract key frames from the clustered trusted frames; Feature extraction is performed on the detection bounding box of the keyframe to obtain the second feature information.
7. The visual tracking method according to claim 1, characterized in that, The step of re-identifying the tracked target based on the feature information of the object to be identified within the target detection box and the feature information of the tracked target includes: Determine a first similarity between the feature information of the object to be identified and the feature information of the tracked target, and a second similarity between the feature information of the object to be identified within the target detection box and the feature information of the non-tracked target; Based on the first similarity and the second similarity, a vote is performed to determine the relationship between the object to be identified and the tracking target.
8. The visual tracking method according to any one of claims 1-7, characterized in that, The re-identification of the tracked target based on the information of the target detection box and the information of the detection boxes included in the tracking chain includes: Based on the relative position information of the target detection box and the relative position information of the detection boxes contained in the tracking chain, the tracked target is re-identified.
9. The visual tracking method according to any one of claims 1-7, characterized in that, Based on the tracking chain, an abnormal tracking event is determined to exist, including: Determine the number, area, and location of the detection frames contained in the tracking chain; Based on the number, area, and position of the detection frames contained in the tracking chain, it is determined whether the tracking anomaly event exists.
10. A visual tracking device, characterized in that, include: The determination module is used to determine a target detection box based on the type of the tracking anomaly event when it is determined that there is a tracking anomaly event in the tracking chain based on the tracking target. The re-identification module is used to re-identify the tracking target based on the feature information of the object to be identified within the target detection box and the feature information of the tracking target, and / or to re-identify the tracking target based on the information of the target detection box and the information of the detection boxes contained in the tracking chain. The feature information of the tracked target includes first feature information obtained after initializing the information of the tracked target and second feature information obtained during the visual tracking process of the tracked target.
11. An electronic device comprising a processor and a memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the visual tracking method according to any one of claims 1 to 9.
12. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the visual tracking method as described in any one of claims 1 to 9.
13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the visual tracking method according to any one of claims 1 to 9.