Target tracking method and device, electronic terminal and computer readable storage medium
By performing multi-grained feature extraction and feature matching on the target object, the problem of poor target tracking accuracy and stability in the prior art is solved, and higher matching accuracy and tracking stability are achieved.
Patent Information
- Application Number
- CN202510527341.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-25
AI Technical Summary
In the prior art, the accuracy and stability of target tracking are poor, especially in complex scenarios such as the target being blocked or affected by light.
The multi-grained feature extraction method is used to detect the target object, acquire the location characteristics of each detection part, and match these features with the target track data in the historical video frame to generate the current track data to improve the accuracy and stability of target tracking.
Through multi-grained feature extraction and feature matching, the impact of occluded parts or undetected parts on the target matching results is reduced, and the matching accuracy and tracking stability of the target object are improved.
Smart Images

Figure CN120047492A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target tracking, and particularly to a target tracking method, device, electronic terminal and computer-readable storage medium. Background Art
[0002] Target tracking is one of the most important and fundamental tasks in the field of computer vision. Its purpose is to output the position of the target object in each video frame of the video containing the target object. However, in complex scenarios such as target congestion, frequent changes in target poses, small targets in the distance, and inaccurate target bounding boxes, existing tracking algorithms face many problems. For example, when the target is occluded or affected by light and other interferences, the accuracy and stability of target tracking are not good. Summary of the Invention
[0003] The main technical problem to be solved by the present invention is to provide a target tracking method, device, electronic terminal and computer-readable storage medium to solve the problem of poor accuracy and stability of target tracking in the prior art.
[0004] To solve the above technical problem, the first technical solution adopted by the present invention is: to provide a target tracking method, the target tracking method includes: Performing multi-granularity feature extraction on the target object in the current video frame to obtain detection information of the target object; the detection information includes the part features of each detected part; Matching the detection information of the target object with the historical trajectory data of each corresponding target in the historical video frames before the current video frame to determine the target that matches the target object; Associating the detection information of the target object in the current video frame with the historical trajectory data of the matched target to generate the current trajectory data of the target object.
[0005] Wherein, performing multi-granularity feature extraction on the target object in the current video frame to obtain detection information of the target object includes: Performing target detection on the current video frame to obtain target detection frames containing each target object; Performing multi-granularity feature extraction on the target object in the target detection frame through a multi-granularity feature extraction network to obtain the global feature map corresponding to the target object and the part features of each detected part.
[0006] Wherein, performing multi-granularity feature extraction on the target object in the current video frame to obtain detection information of the target object includes: Performing target detection on the current video frame to obtain target detection frames containing each target object; Performing feature extraction on the target object in the target detection frame to obtain the global feature map of the target object; Perform pixel classification on the global feature map of the target object to obtain the part categories of each pixel point in the global feature map; Aggregate the pixel points of the same part category to obtain the part features of the part category; the part category is used as the part category of the detected part.
[0007] Among them, performing pixel classification on the global feature map of the target object to obtain the part categories of each pixel point in the global feature map includes: Classify each pixel point in the global feature map of the target object to obtain the probability values of each pixel point for each preset category; Select the preset category corresponding to the maximum probability value of the pixel point as the part category of the pixel point.
[0008] Among them, performing multi-granularity feature extraction on the target object in the current video frame to obtain the detection information of the target object further includes: Based on the probability values of the pixel points corresponding to the part category of the detected part, determine whether the detected part is a visible part.
[0009] Among them, based on the probability values of the pixel points corresponding to the category of the detected part, determining whether the detected part is a visible part includes: Select the maximum probability value among the probability values of all pixel points corresponding to the part category of the detected part and compare it with a preset value; In response to the maximum probability value corresponding to the part category of the detected part being greater than the preset value, determine that the detected part is a visible part.
[0010] Among them, matching the detection information of the target object with the historical trajectory data of each corresponding target in the historical video frames before the current video frame to determine the target matching the target object includes: Perform feature matching between the targets corresponding to the historical trajectory data in the first state and each target object to obtain a first matching result; the first matching result includes target object-target matching pairs, unmatched target objects, and unmatched historical trajectory data; the continuous number of historical video frames containing the target in the historical trajectory data in the first state is not less than a preset number of frames; Perform IOU matching between the unmatched target objects in the first matching result and the unmatched historical trajectory data and the historical trajectory data in the second state respectively to obtain a second matching result; the second matching result includes target object-target matching pairs; the continuous number of historical video frames containing the target object in the historical trajectory data in the second state is less than the preset number of frames.
[0011] Among them, the detected part has a part category; Performing feature matching between the targets corresponding to the historical trajectory data in the first state and each target object to obtain a first matching result includes: Determine the part similarity of the part category based on the target object and the part features of the same part category corresponding to the target; the detected part corresponding to the part category is the visible part; Sum and average the part similarities of each part category corresponding between the target object and the target to determine the target similarity between the target object and the target; Determine the target matching degree between the target object and the target based on the target similarity between the target object and the target; Determine whether the target object and the target match based on the target matching degree between the target object and the target.
[0012] Among them, determining the target matching degree between the target object and the target based on the target similarity between the target object and the target includes: In response to the number of part categories of the target object in the current video frame that are the same as those of the target not exceeding the first preset number, determine the target matching degree between the target object and the target only based on the target similarity between the target object and the target.
[0013] Among them, determining the target matching degree between the target object and the target based on the target similarity between the target object and the target includes: In response to the number of part categories of the target object in the current video frame that are the same as those of the target exceeding the first preset number, determine the global similarity between the target object and the target based on the global feature map of the target object and the global feature map of the target; Determine the target matching degree between the target object and the target based on the corresponding global similarity and target similarity between the target object and the target.
[0014] Among them, the target tracking method further includes: In response to the target object matching the historical trajectory data in the second state to form the current trajectory data of the target object, determine whether to update the state of the current trajectory data of the target object based on the number of detected parts of the target object in the current video frame and the continuous appearance times of the target object in the current trajectory data.
[0015] Among them, determining whether to update the state of the current trajectory data of the target object based on the number of detected parts of the target object in the current video frame and the continuous appearance times of the target object in the current trajectory data includes: In response to the number of detected parts of the target object in the current video frame exceeding the second preset number and the continuous appearance times of the target object in the current trajectory data exceeding the preset number, update the state of the current trajectory data of the target object to the first state.
[0016] Among them, the historical trajectory data of the target has a first feature library and a second feature library; the first feature library is composed of the part features of each detected part of the target in the first image; the second feature library is composed of the part features of each detected part of the target in the second image; the first image is the historical video frame corresponding to the historical trajectory data with the largest number of detected parts of the target; the second image is the historical video frame closest to the current video frame and containing the target in the historical trajectory data; The target tracking method further includes: Compare the number of detected parts of the target object in the current video frame with a third preset number to determine the first feature library and the second feature library corresponding to the current trajectory data of the target object.
[0017] Among them, comparing the number of detected parts of the target object in the current video frame with a third preset number to determine the first feature library and the second feature library corresponding to the current trajectory data of the target object includes: In response to the number of detected parts of the target object in the current video frame being greater than the total number of detected parts in the first feature library of the historical trajectory data of the target or not less than the first threshold, update the part features of the detected parts of the target object in the current video frame to the first feature library; In response to the number of detected parts of the target object in the current video frame being not less than the total number of detected parts in the second feature library of the historical trajectory data of the target or not less than the second threshold, update the part features of the detected parts of the target object in the current video frame to the second feature library.
[0018] Among them, the target tracking method further includes: Correct the target detection frame of the target object based on the preset size information to determine the effective detection frame of the target object; Replace the target detection frame of the target object in the current trajectory data with the effective detection frame of the target object; Predict the position information of the target object in the next video frame after the current video frame based on the updated current trajectory data of the target object.
[0019] To solve the above technical problems, the second technical solution adopted by the present invention is: to provide a target tracking device, and the target tracking device includes: A detection module, configured to perform multi-granularity feature extraction on the target object in the current video frame to obtain the detection information of the target object; the detection information includes the part features of each detected part; A matching module, configured to match the detection information of the target object with the historical trajectory data of each corresponding target in the historical video frames before the current video frame to determine the target that matches the target object; An association module, configured to associate the detection information of a target object in a current video frame with the historical trajectory data of a matching target to generate the current trajectory data of the target object.
[0020] To solve the above technical problems, the third technical solution adopted by the present invention is: to provide an electronic terminal, which includes a memory and a processor coupled to each other. The processor is configured to execute program instructions stored in the memory, and the processor is configured to execute program data to implement the steps in the above-mentioned target tracking method.
[0021] To solve the above technical problems, the fourth technical solution adopted by the present invention is: to provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned target tracking method are implemented.
[0022] The beneficial effects of the present invention are as follows: Different from the prior art, the provided target tracking method, device, electronic terminal and computer-readable storage medium, the target tracking method includes performing multi-granularity feature extraction on a target object in a current video frame to obtain the detection information of the target object; the detection information includes the part features of each detected part; matching the detection information of the target object with the historical trajectory data of each corresponding target in the historical video frames before the current video frame to determine the target that matches the target object; associating the detection information of the target object in the current video frame with the historical trajectory data of the matching target to generate the current trajectory data of the target object. By matching the part features of each detected part of the target object with the historical trajectory data of each target in the historical video frames, the present application reduces the influence of occluded parts or undetected parts on the target matching result, thereby improving the matching accuracy of the target object and the tracking stability of the target object. Description of the Drawings
[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0024] Figure 1 is a schematic flowchart of the target tracking method provided by the present invention; Figure 2 is a schematic diagram of a specific embodiment of the key points corresponding to the target provided by the present invention; Figure 3 is Figure 1 a schematic diagram of a specific embodiment of step S1 in the provided target tracking method; Figure 4 isFigure 1 Schematic diagram of another specific embodiment of step S1 in the provided target tracking method; Figure 5 is Figure 1 Schematic diagram of a specific embodiment of step S2 in the provided target tracking method; Figure 6 is Figure 5 Schematic diagram of a specific embodiment of step S21 in the provided target tracking method; Figure 7 Schematic diagram of the framework of an embodiment of the target tracking device provided by the present invention; Figure 8 Schematic diagram of the framework of an embodiment of the electronic terminal provided by the present invention; Figure 9 Schematic diagram of the framework of an embodiment of the computer-readable storage medium provided by the present invention. Detailed implementation manners
[0025] The solutions of the embodiments of the present application will be described in detail below with reference to the accompanying drawings of the specification.
[0026] In the following description, specific details such as specific system architectures, interfaces, and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the present application.
[0027] As used herein, the term "and / or" merely describes an association relationship between associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after. In addition, "multiple" herein means two or more than two.
[0028] To enable those skilled in the art to better understand the technical solutions of the present invention, the target tracking method provided by the present invention will be further described in detail below with reference to the drawings and specific implementation manners.
[0029] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0030] Before further elaborating on the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are described. The nouns and terms involved in the embodiments of the present application are applicable to the following explanations.
[0031] The DeepSORT algorithm is a deep learning-based object tracking algorithm that combines the advantages of the SORT algorithm and deep learning feature extraction. The DeepSORT algorithm extracts features from the object bounding box and uses the Kalman filter to predict the object state, thereby achieving object tracking. The DeepSORT algorithm has good robustness in complex situations such as object occlusion and object disappearance.
[0032] BPBREID is a part-based human feature re-identification technology.
[0033] The object tracking method provided by the embodiments of the present application can be implemented independently by a server or a terminal, or jointly implemented by a server and a terminal. In some embodiments, the terminal or the server can implement the object tracking method provided by the embodiments of the present application by running a computer program. For example, the computer program can be a native program or a software module in an operating system; it can be a local (Native) application (APP, Application), that is, a program that needs to be installed in the operating system to run, such as a client that supports virtual scenarios, such as a game APP; it can also be a small program, that is, a program that only needs to be downloaded to the browser environment to run; it can also be a small program that can be embedded in any APP. In short, the above computer program can be any form of application program, module or plug-in.
[0034] The following takes the implementation by the server as an example to illustrate the object tracking method provided by the embodiments of the present application.
[0035] Please refer to Figure 1 , Figure 1 which is a schematic flow diagram of the object tracking method provided by the present invention.
[0036] In this embodiment, an object tracking method is provided, and the object tracking method includes the following steps.
[0037] S1: Perform multi-granularity feature extraction on the target object in the current video frame to obtain the detection information of the target object; the detection information includes the part features of each detected part.
[0038] S2: Match the detection information of the target object with the historical trajectory data of each corresponding target in the historical video frames before the current video frame to determine the target that matches the target object.
[0039] S3: Associate the detection information of the target object in the current video frame with the historical trajectory data of the matched target to generate the current trajectory data of the target object.
[0040] In one embodiment, a multi-granularity feature extraction network is constructed based on BPBREID, and the constructed multi-granularity feature extraction network is trained according to the actual situation, so that the trained multi-granularity feature extraction network can detect the global features of the target object, the part features of each preset part, and the part visibility type.
[0041] In one embodiment, according to different task accuracy requirements, different granularity preset schemes are determined, that is, the task accuracy corresponds one-to-one with the preset scheme. The number of preset parts included in each preset scheme is different. For example, the finer the task accuracy, the more preset parts are included in the preset scheme, and the more part categories the preset parts correspond to. For example, the granularity selection and feature aggregation are performed according to the classification rules adopted in the target key point detection model based on the openpifpaf library.
[0042] Please refer to Figure 2 , Figure 2 is a schematic diagram of a specific embodiment of the key points corresponding to the target provided by the present invention.
[0043] In one embodiment, all the key points of the target include nose 1, left eye 2, right eye 3, left ear 4, right ear 5, left shoulder 6, right shoulder 7, left elbow 8, right elbow 9, left wrist 10, right wrist 11, left hip 12, right hip 13, left knee 14, right knee 15, left ankle 16, and right ankle 17.
[0044] For example, a preset scheme includes head key points, torso key points, arm and shoulder key points, leg key points, and foot key points. The target is formed by connecting the adjacent key points among the head key points, arm and shoulder key points, torso key points, leg key points, and foot key points in sequence.
[0045] In one embodiment, the training method of the multi-granularity feature extraction network is as follows.
[0046] Obtain a plurality of sample images, where the sample images are images containing the target, and each target has annotation information. The annotation information includes the global annotation features of the target, the part annotation features of each preset part of the target, and the annotation visibility type. Among them, the annotation visibility type is visible parts and occluded parts. The sample images are input into the multi-granularity feature extraction network to obtain the global prediction features of the target, the part prediction features of each preset part of the target, and the predicted visibility probability. The larger the predicted visibility probability value, the smaller the probability that the preset part is occluded and / or the smaller the occluded area of the preset part. Based on the first loss value between the global annotation features and the global prediction features corresponding to the same target, and the second loss value between the part annotation features and the part prediction features corresponding to each preset part of the target, the multi-granularity feature extraction network is iteratively trained, so that the trained multi-granularity feature extraction network can achieve the feature detection of the target task.
[0047] In a specific embodiment, the multi-granularity feature extraction network is iteratively trained based on the following loss function.
[0048] (Formula 1) Where: L GiLt represents the overall loss value of the target; L id represents the first loss value of the target; L id is composed of the common cross-entropy loss L CE ; f g represents the overall embedding, f f represents the foreground embedding, f c represents the part feature; L tri is the second loss value of the target.
[0049] (Formula 2) (Formula 3) Where: dist eucl represents the Euclidean distance; d ij represents the average distance corresponding to all part features of the target; d ap represents the distance between the sample image and the hardest positive sample; d an represents the distance between the sample image and the hardest negative sample; α represents the triplet loss margin.
[0050] Among them, Mobileone based on a lightweight architecture is used as the multi-granularity feature extraction network. The global prediction feature G of the target = R H×W×C .
[0051] Specifically, the specific implementation of extracting multi-granularity features from the target object in the current video frame in step S1 is as follows.
[0052] In one embodiment, the specific steps for determining the detection information of the target object are as follows.
[0053] Please refer to Figure 3 , Figure 3 which Figure 1 is a schematic diagram of a specific embodiment of step S1 in the target tracking method provided.
[0054] S111: Perform target detection on the current video frame to obtain target detection boxes containing each target object.
[0055] Specifically, video stream data containing a target object is obtained. The video stream data can be composed of multiple consecutive video frames. The video stream data can be real-time captured video stream or offline captured video stream. Target detection is performed on the current video frame through a target detection network model to obtain a target detection box containing the target object in the current video frame. The target detection network model can be RCNN series (RCNN, Fast-RCNN, Faster-RCNN), R-FCN, YOLO, SSD, and FPN. Among them, Faster-RCNN is short for Faster Region-based Convolutional Neural Networks; Fast-RCNN is short for Fast Region-based Convolutional Neural Networks; RCNN is short for Region-based Convolutional Neural Networks.
[0056] S112: Multigranularity feature extraction is performed on the target object in the target detection box through a multigranularity feature extraction network to obtain a global feature map corresponding to the target object and part features of each detected part.
[0057] Specifically, multigranularity feature extraction is performed on the target object in the target detection box through the multigranularity feature extraction network trained in the above embodiment to obtain a global feature map of the target object, part features of each detected part corresponding to the target object, and a visibility index. Among them, the visibility index includes two indices, 0 and 1. Among them, 0 indicates that the detected part is occluded and invisible; 1 indicates that the detected part is not occluded and visible.
[0058] In one embodiment, the detected parts of the target object that the multigranularity feature extraction network detects can include the head, arm shoulders, torso, legs, and feet.
[0059] Please refer to Figure 4 , Figure 4 is Figure 1 a schematic diagram of another specific embodiment of step S1 in the provided target tracking method.
[0060] In one embodiment, the specific steps for determining the detection information of the target object are as follows.
[0061] S121: Target detection is performed on the current video frame to obtain target detection boxes containing each target object.
[0062] Specifically, target detection is performed on the current video frame through a target detection network model to obtain a target detection box containing the target object. Each target detection box contains one target object. Among them, the target object can be an animal, a vehicle, etc.
[0063] S122: Extract features from the target object in the target detection box to obtain the global feature map of the target object.
[0064] Specifically, extract features from the target object in the target detection box to obtain the global feature map of the target object, and the size of the global feature map is H×W×C.
[0065] S123: Perform pixel classification on the global feature map of the target object to obtain the part categories of each pixel point in the global feature map.
[0066] Specifically, classify each pixel point in the global feature map of the target object to obtain the probability values of each pixel point for each preset category; select the preset category corresponding to the maximum probability value of the pixel point as the part category of the pixel point.
[0067] Through the above steps, the part categories of each pixel point in the global feature map can be obtained.
[0068] S124: Aggregate the pixel points of the same part category to obtain the part feature of the part category; the part category is used as the part category of the detection part.
[0069] Specifically, aggregate based on all pixel points corresponding to the same part category to obtain the part feature corresponding to the part category. Use the part feature of the part category as the part category of the detection part.
[0070] In one embodiment, the specific steps for determining the detection information of the target object further include the following.
[0071] S125: Based on the probability values of the pixel points corresponding to the part category of the detection part, determine whether the detection part is a visible part.
[0072] In one specific embodiment, select the maximum probability value among the probability values of all pixel points corresponding to the part category of the detection part and compare it with a preset value. In response to the maximum probability value corresponding to the part category of the detection part being greater than the preset value, determine that the detection part is a visible part. In response to the maximum probability value corresponding to the part category of the detection part not being greater than the preset value, determine that the detection part is an invisible part.
[0073] In another specific embodiment, sum up and average the probability values of all pixel points corresponding to the part category of the detection part to obtain the average probability value of the detection part belonging to the part category, and compare the average probability value of the detection part belonging to the part category with the preset value. In response to the average probability value of the detection part belonging to the part category being greater than the preset value, determine that the detection part is a visible part; in response to the average probability value of the detection part belonging to the part category not being greater than the preset value, determine that the detection part is an invisible part.
[0074] Specifically, in step S2, the detection information of the target object is matched with the historical trajectory data of each corresponding target in the historical video frames before the current video frame, and the specific implementation of determining the target that matches the target object is as follows.
[0075] Please refer to Figure 5 , Figure 5 which Figure 1 is a schematic diagram of a specific embodiment of step S2 in the provided target tracking method.
[0076] S21: Match the targets corresponding to the historical trajectory data in the first state with each target object to obtain a first matching result.
[0077] Among them, the first matching result includes target object - target matching pairs, unmatched target objects, and unmatched historical trajectory data; the number of consecutive frames of the historical video frames containing the target in the historical trajectory data in the first state is not less than a preset number of frames. That is, the historical trajectory data in the first state is a confirmed trajectory.
[0078] Please refer to Figure 6 , Figure 6 which Figure 5 is a schematic diagram of a specific embodiment of step S21 in the provided target tracking method.
[0079] In one embodiment, the specific implementation of obtaining the first matching result is as follows.
[0080] S211: Based on the part features of the target object and the target corresponding to the same part category, determine the part similarity of the part category; the detected part corresponding to the part category is a visible part.
[0081] For example, the head of the target object is a visible part, and the head of the target is a visible part. Determine the head similarity between the target object and the target based on the part features corresponding to the head of the target object and the part features corresponding to the head of the target.
[0082] For example, if the foot of the target object is an invisible part and the foot of the target is a visible part, then the foot similarity between the target object and the target does not need to be calculated. Or, when the foot of the target object is a visible part and the foot of the target is an invisible part, the foot similarity between the target object and the target also does not need to be calculated.
[0083] Through this step, the part similarity of the part category where both the target object and the target belong to visible parts can be obtained.
[0084] S212: Sum and average the part similarities of each corresponding part category between the target object and the target to determine the target similarity between the target object and the target.
[0085] In one embodiment, the target similarity between the target object and the target can be obtained through the following formula.
[0086] (Formula 4) In the formula: represents the part feature distance between the target object and the target; the target similarity between the target object and the target is 1 - , the greater the distance between the target object and the target, the smaller the target similarity between the target object and the target; is used to represent the part similarity between the target object and the target that are mutually visible.
[0087] S213: Based on the target similarity between the target object and the target, determine the target matching degree between the target object and the target.
[0088] Specifically, determine the target matching degree between the target object and the target based on the following formula.
[0089] (Formula 5) In the formula: the target matching degree between the target object and the target is 1 - , represents the part difference degree between the target object and the target; represents the global difference degree between the target object and the target; represents a boolean value, taking values of 0 and 1.
[0090] In a specific embodiment, in response to the number of part categories of the target object in the current video frame that are the same as those of the target not exceeding the first preset number, only determine the target matching degree between the target object and the target based on the target similarity between the target object and the target. That is, when the number of part categories of the target object in the current video frame that are the same as those of the target does not exceed the first preset number, the boolean value in Formula 5 takes the value of 0.
[0091] In a specific embodiment, in response to the number of part categories of the target object in the current video frame that are the same as those of the target exceeding the first preset number, determine the global similarity between the target object and the target based on the global feature map of the target object and the global feature map of the target; based on the corresponding global similarity and target similarity between the target object and the target, determine the target matching degree between the target object and the target. That is, when the number of part categories of the target object in the current video frame that are the same as those of the target exceeds the first preset number, the boolean value in Formula 5 takes the value of 1.
[0092] S214: Based on the target matching degree between the target object and the target, determine whether the target object and the target match.
[0093] Specifically, in response to the target matching degree between the target object and the target being greater than the preset matching degree threshold, it is determined that the target object matches the target; in response to the target matching degree between the target object and the target not being greater than the matching degree threshold, it is determined that the target object and the target do not match.
[0094] S22: Perform IOU matching on the unmatched target objects in the first matching result with the unmatched historical trajectory data and each piece of historical trajectory data in the second state respectively to obtain a second matching result.
[0095] Among them, the second matching result includes target object - target matching pairs; the continuous number of frames of the historical video frames containing the target object in the historical trajectory data in the second state is less than the preset number of frames. That is, the historical trajectory data in the second state is non - confirmed trajectory.
[0096] Specifically, in order to improve the matching probability of the target object, perform IOU matching on the unmatched target objects with the unmatched historical trajectory data in the first state and the historical trajectory data in the second state in the first matching result to obtain a second matching result.
[0097] In one embodiment, the target tracking method further includes the following steps.
[0098] In response to the target object matching with the historical trajectory data in the second state to form the current trajectory data of the target object, determine whether to update the state of the current trajectory data of the target object based on the number of detected parts of the target object in the current video frame and the continuous occurrence times of the target object in the current trajectory data.
[0099] In response to the number of detected parts of the target object in the current video frame exceeding the second preset number and the continuous occurrence times of the target object in the current trajectory data exceeding the preset number, update the state of the current trajectory data of the target object to the first state.
[0100] In a specific embodiment, the historical trajectory data of the target has a first feature library and a second feature library; the first feature library is composed of the part features of each detected part of the target in the first image; the second feature library is composed of the part features of each detected part of the target in the second image; the first image is the historical video frame corresponding to the historical trajectory data with the largest number of detected parts of the target; the second image is the historical video frame closest to the current video frame in the historical trajectory data that contains the target.
[0101] In one embodiment, the target tracking method further includes the following steps.
[0102] Compare the number of detected parts of the target object in the current video frame with the third preset number to determine the first feature library and the second feature library corresponding to the current trajectory data of the target object.
[0103] In a specific embodiment, in response to the number of detected parts of the target object in the current video frame being greater than the total number of detected parts in the first feature library in the historical trajectory data of the target or not less than the first threshold, the part features of the detected parts of the target object in the current video frame are updated to the first feature library.
[0104] In response to the number of detected parts of the target object in the current video frame being not less than the total number of detected parts in the second feature library in the historical trajectory data of the target or not less than the second threshold, the part features of the detected parts of the target object in the current video frame are updated to the second feature library.
[0105] In an embodiment, for the case where the target object is partially occluded, due to the jump of the detection frame, even if feature matching realizes the effective association of data, it will still have a great impact on the subsequent Kalman filter update and prediction. Therefore, the target tracking method further includes the following steps.
[0106] Based on the preset size information, correct the target detection frame of the target object to determine the effective detection frame of the target object. Among them, the target detection frame containing the target object is (t, l, w, h), where t represents the x-axis coordinate of the upper left corner of the target detection frame; l represents the y-axis coordinate of the upper left corner of the target detection frame, w represents the length of the target detection frame in the x-axis direction; h represents the length of the target detection frame in the y-axis direction.
[0107] In response to the length of the target detection frame containing the target object in the x-axis direction and the length in the y-axis direction both conforming to the corresponding preset size, determine the target detection frame of the target object as the effective detection frame.
[0108] In a specific embodiment, in response to the length w of the target detection frame of the target object not conforming to the corresponding preset size, correct the length w of the target detection frame in the x-axis direction to obtain the effective length w of the target detection frame in the x-axis direction new , the effective length w of the target detection frame in the x-axis direction new conforms to the corresponding preset size.
[0109] In a specific embodiment, in response to the length h of the target detection frame of the target object in the y-axis direction not conforming to the corresponding preset size, correct the length h of the target detection frame in the y-axis direction to obtain the effective length h of the target detection frame in the y-axis direction new , the effective length h of the target detection frame in the y-axis direction new conforms to the corresponding preset size.
[0110] Replace the target detection frame of the target object in the current trajectory data based on the effective detection frame of the target object.
[0111] Predict the position information of the target object in the next video frame after the current video frame based on the updated current trajectory data of the target object.
[0112] In this application, a multi-granularity feature extraction network and a feature matching method are adopted, which can effectively track the target object when the head or other parts are occluded, and do not rely on the detection results of specific parts, so the stability of the algorithm is better.
[0113] In one embodiment, the detection information of the target object in the current video frame is associated with the historical trajectory data of the matched target to generate the current trajectory data of the target object.
[0114] The target tracking method provided in this embodiment matches the part features of each detected part of the target object with the historical trajectory data of each target in the historical video frames, reduces the influence of occluded parts or undetected parts on the target matching result, and thus improves the matching accuracy of the target object and the tracking stability of the target object.
[0115] Please refer to Figure 7 , Figure 7 , which is a schematic framework diagram of an embodiment of the target tracking device provided by the present invention. This embodiment provides a target tracking device 60, and the target tracking device 60 includes a detection module 61, a matching module 62, and an association module 63.
[0116] The detection module 61 is used to perform multi-granularity feature extraction on the target object in the current video frame to obtain the detection information of the target object; the detection information includes the part features of each detected part.
[0117] The matching module 62 is used to match the detection information of the target object with the historical trajectory data of each corresponding target in the historical video frames before the current video frame to determine the target that matches the target object.
[0118] The association module 63 is used to associate the detection information of the target object in the current video frame with the historical trajectory data of the matched target to generate the current trajectory data of the target object.
[0119] The target tracking device provided in this embodiment matches the part features of each detected part of the target object with the historical trajectory data of each target in the historical video frames, reduces the influence of occluded parts or undetected parts on the target matching result, and thus improves the matching accuracy of the target object and the tracking stability of the target object.
[0120] Please refer to Figure 8 , Figure 8It is a schematic framework diagram of an embodiment of an electronic terminal provided by the present invention. The electronic terminal 80 includes a memory 81 and a processor 82 that are coupled to each other. The processor 82 is configured to execute program instructions stored in the memory 81 to implement the steps of any of the above-mentioned target tracking method embodiments. In a specific implementation scenario, the electronic terminal 80 may include, but is not limited to, a microcomputer and a server. In addition, the electronic terminal 80 may also include mobile devices such as a laptop computer and a tablet computer, which are not limited herein.
[0121] Specifically, the processor 82 is configured to control itself and the memory 81 to implement the steps of any of the above-mentioned target tracking method embodiments. The processor 82 may also be referred to as a CPU (Central Processing Unit). The processor 82 may be an integrated circuit chip with signal processing capabilities. The processor 82 may also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 82 may be implemented jointly by integrated circuit chips.
[0122] In the above solution, the target tracking method includes performing multi-granularity feature extraction on a target object in a current video frame to obtain detection information of the target object; the detection information includes part features of each detected part; matching the detection information of the target object with historical trajectory data of corresponding targets in historical video frames before the current video frame to determine a target that matches the target object; and associating the detection information of the target object in the current video frame with the historical trajectory data of the matched target to generate current trajectory data of the target object.
[0123] Please refer to Figure 9 , Figure 9 It is a schematic framework diagram of an embodiment of a computer-readable storage medium provided by the present invention. The computer-readable storage medium 90 stores program instructions 901 that can be run by a processor. The program instructions 901 are used to implement the steps of any of the above-mentioned target tracking method embodiments.
[0124] In the above solution, the target tracking method includes performing multi-granularity feature extraction on the target object in the current video frame to obtain detection information of the target object; the detection information includes the part features of each detected part; matching the detection information of the target object with the historical trajectory data of each corresponding target in the historical video frames before the current video frame to determine the target that matches the target object; and associating the detection information of the target object in the current video frame with the historical trajectory data of the matched target to generate the current trajectory data of the target object.
[0125] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.
[0126] The descriptions of the above embodiments tend to emphasize the differences between the embodiments. The same or similar parts can be referred to each other. For the sake of brevity, they will not be repeated here.
[0127] In several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation manners described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division manners. For example, units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.
[0128] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0129] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0130] The above are only the embodiments of the present invention, and do not limit the patent protection scope of the present invention. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A target tracking method, characterized in that: The target tracking method comprises: Perform multi-granularity feature extraction on the target object in the current video frame to obtain detection information of the target object; the detection information includes part features of each detection part; Matching the detection information of the target object with the historical trajectory data of each corresponding target in the historical video frame before the current video frame to determine the target that matches the target object; The detection information of the target object in the current video frame is associated with the matched historical trajectory data of the target to generate current trajectory data of the target object.
2. The target tracking method according to claim 1, characterized in that: The step of performing multi-granularity feature extraction on the target object in the current video frame to obtain detection information of the target object includes: Performing target detection on the current video frame to obtain a target detection frame containing each target object; Multi-granularity feature extraction is performed on the target object in the target detection frame through a multi-granularity feature extraction network to obtain a global feature map corresponding to the target object and part features of each detection part.
3. The target tracking method according to claim 1, characterized in that: The step of performing multi-granularity feature extraction on the target object in the current video frame to obtain detection information of the target object includes: Performing target detection on the current video frame to obtain a target detection frame containing each target object; Performing feature extraction on the target object in the target detection frame to obtain a global feature map of the target object; Performing pixel classification on the global feature map of the target object to obtain a part category of each pixel point in the global feature map; Pixel points of the same part category are aggregated to obtain part features of the part category; the part category is used as the part category of the detected part.
4. The target tracking method according to claim 3, characterized in that: The pixel classification of the global feature map of the target object to obtain the location category of each pixel point in the global feature map includes: Classify each pixel point in the global feature map of the target object to obtain a probability value of each pixel point belonging to each preset category; The preset category corresponding to the maximum probability value corresponding to the pixel point is selected as the part category of the pixel point.
5. The target tracking method according to claim 4, characterized in that: The step of performing multi-granularity feature extraction on the target object in the current video frame to obtain detection information of the target object further includes: Whether the detected part is a visible part is determined based on the probability value of the pixel point corresponding to the part category of the detected part.
6. The target tracking method according to claim 5, characterized in that: The determining whether the detected part is a visible part based on the probability value of the pixel point corresponding to the category of the detected part includes: Selecting a maximum probability value among the probability values of all the pixel points corresponding to the part category of the detection part and comparing it with a preset value; In response to the maximum probability value corresponding to the part category of the detected part being greater than the preset value, the detected part is determined to be a visible part.
7. The target tracking method according to claim 1, characterized in that: The step of matching the detection information of the target object with the historical trajectory data of each target in the historical video frame before the current video frame to determine the target that matches the target object includes: Perform feature matching on the target corresponding to each of the historical trajectory data in the first state and each of the target objects to obtain a first matching result; the first matching result includes a target object-target matching pair, an unmatched target object, and an unmatched historical trajectory data; the number of consecutive frames of historical video frames containing the target in the historical trajectory data in the first state is not less than a preset number of frames; The unmatched target object in the first matching result is respectively matched with the unmatched historical trajectory data and each of the historical trajectory data in the second state to obtain a second matching result; the second matching result includes a target object-target matching pair; the number of consecutive frames of historical video frames containing the target object in the historical trajectory data in the second state is less than the preset number of frames.
8. The target tracking method according to claim 7, characterized in that: The detection part has a part category; The performing feature matching between the target corresponding to each of the historical trajectory data in the first state and each of the target objects to obtain a first matching result includes: Based on the part features of the target object and the target corresponding to the same part category, determining the part similarity of the part category; the detected part corresponding to the part category is a visible part; Adding and averaging the part similarities of the part categories corresponding to the target object and the target to determine the target similarity between the target object and the target; Determining a target matching degree between the target object and the target based on the target similarity between the target object and the target; Based on the target matching degree between the target object and the target, it is determined whether the target object matches the target.
9. The target tracking method according to claim 8, characterized in that: The determining the target matching degree between the target object and the target based on the target similarity between the target object and the target includes: In response to the number of the part categories of the target object that are the same as the target in the current video frame not exceeding a first preset number, determining the target matching degree between the target object and the target based only on the target similarity between the target object and the target.
10. The target tracking method according to claim 8, characterized in that: The determining the target matching degree between the target object and the target based on the target similarity between the target object and the target includes: In response to the target object having the same part category as the target in the current video frame as the target, the number of the part categories exceeds a first preset number, determining a global similarity between the target object and the target based on a global feature map of the target object and a global feature map of the target; The target matching degree between the target object and the target is determined based on the global similarity and the target similarity corresponding to the target object and the target.
11. The target tracking method according to claim 7, characterized in that: The target tracking method further comprises: In response to the target object matching the historical trajectory data in the second state to form current trajectory data of the target object, it is determined whether to update the state of the current trajectory data of the target object based on the number of detection parts of the target object in the current video frame and the number of consecutive appearances of the target object in the current trajectory data.
12. The target tracking method according to claim 11, characterized in that: The determining whether to update the state of the current trajectory data of the target object based on the number of the detection parts of the target object in the current video frame and the number of consecutive appearances of the target object in the current trajectory data includes: In response to the number of the detection parts of the target object in the current video frame exceeding a second preset number and the number of consecutive appearances of the target object in the current trajectory data exceeding a preset number, the state of the current trajectory data of the target object is updated to a first state.
13. The target tracking method according to claim 11, characterized in that: The historical trajectory data of the target has a first feature library and a second feature library; the first feature library is composed of the part features of each of the detection parts of the target in the first image; the second feature library is composed of the part features of each of the detection parts of the target in the second image; the first image is the historical video frame corresponding to the detection parts with the largest number of the target in the historical trajectory data; the second image is the historical video frame in the historical trajectory data that contains the target closest to the current video frame; The target tracking method further comprises: The number of the detection parts of the target object in the current video frame is compared with a third preset number to determine a first feature library and a second feature library corresponding to the current trajectory data of the target object.
14. The target tracking method according to claim 13, characterized in that: The step of comparing the number of the detection parts of the target object in the current video frame with a third preset number to determine the first feature library and the second feature library corresponding to the current trajectory data of the target object comprises: In response to the number of the detection parts of the target object in the current video frame being greater than the total number of the detection parts in the first feature library in the historical trajectory data of the target or being not less than a first threshold, updating the part features of the detection parts of the target object in the current video frame to the first feature library; In response to the number of the detection parts of the target object in the current video frame being not less than the total number of the detection parts in the second feature library in the historical trajectory data of the target or not less than a second threshold, the part features of the detection parts of the target object in the current video frame are updated to the second feature library.
15. The target tracking method according to claim 2, characterized in that: The target tracking method further comprises: Correcting the target detection frame of the target object based on preset size information to determine a valid detection frame of the target object; Replace the target detection frame of the target object in the current trajectory data based on the valid detection frame of the target object; Predicting position information of the target object in a next video frame after the current video frame based on the updated current trajectory data of the target object.
16. A target tracking device, characterized in that: The target tracking device comprises: A detection module, used to perform multi-granularity feature extraction on a target object in a current video frame to obtain detection information of the target object; the detection information includes part features of each detection part; A matching module, used to match the detection information of the target object with the historical trajectory data of each corresponding target in the historical video frame before the current video frame, and determine the target that matches the target object; The association module is used to associate the detection information of the target object in the current video frame with the matched historical trajectory data of the target to generate current trajectory data of the target object.
17. An electronic terminal, characterized in that: The electronic terminal includes a memory and a processor coupled to each other, the processor is used to execute program instructions stored in the memory, and the processor is used to execute program data to implement the steps in the target tracking method as described in any one of claims 1 to 15.
18. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the steps in the target tracking method according to any one of claims 1 to 15 are implemented.
Citation Information
Patent Citations
Method and apparatus of detecting image quality
CN107679490A
Target tracking method and device, computer readable storage medium and computer equipment
CN111402294A
Target tracking method and device, electronic equipment and storage medium
CN113160272A
Anti-shielding multi-target tracking method based on trajectory prediction
CN116681729A
Anti-occlusion target tracking method fusing multi-granularity dynamic appearance
CN117036405A