An Automatic Tracking and Re-identification Method for Template Feature Matching
Through the methods of template feature matching and fully connected feature comparison, the problems of target tracking initialization and state determination are solved, and accurate tracking and long-term tracking of fast moving targets are achieved, which improves the robustness and universality of target tracking.
Patent Information
- Application Number
- CN202210252473.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-15
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-03-15
AI Technical Summary
Existing target tracking methods cannot automatically initialize the specified target, tracking status determination is inaccurate and cannot be recaptured after being lost, especially in case of rapid or mutated motion.
The template feature matching method is used to extract convolutional features from the loading template, match and determine the position in the global feature and turn to the tracking process, continuously extract fully connected features and compare them with the feature pool to determine the tracking status. If abnormal, local matching will be performed for re-identification and re-tracking.
Accurate position prediction of targets of different operating states is achieved, the accuracy and robustness of target tracking is improved, and the target can be recaptured after tracking exceptions, enhancing the universality of tracking.
Smart Images

Figure CN114581678B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and particularly relates to an automatic tracking and re-identification method for template feature matching. Background Art
[0002] Object tracking is one of the hotspots in the field of computer vision research, and object tracking has a wide range of applications in many fields such as video surveillance, navigation, military, human-computer interaction, virtual reality, and autonomous driving. Briefly speaking, object tracking is to analyze and track a given object in a video to determine the exact position of the object in the video.
[0003] Currently, the initialization of tracking methods mostly depends on detection and recognition algorithms or manual interaction to select the object to be tracked. For objects that cannot be detected and recognized by the detector or specific designated objects, it is impossible to initialize the tracking quickly and accurately. Moreover, for tracking objects with rapid or mutant movements in the existing object tracking methods, due to the uncontrollability of the movement of the tracking object, it is easy to cause the loss of the object during the object tracking process and it is impossible to determine the object tracking state and recapture it. Summary of the Invention
[0004] The purpose of the present invention is to provide an automatic tracking and re-identification method for template feature matching to solve the problems that currently, for special designated target objects, automatic initialization tracking cannot be performed, the object tracking state cannot be accurately determined, and the lost object cannot be retrieved.
[0005] To solve the above technical problems, the present invention provides an automatic tracking and re-identification method for template feature matching, including a template feature matching method and an object tracking and re-identification method;
[0006] By extracting convolutional features from the loaded template, determining the position in the global features and transferring to the tracking process; during the tracking process, continuously extracting fully connected features of the object and comparing with the feature pool to determine the tracking state. If the tracking continues to be abnormal, local matching is performed for re-identification and then tracking.
[0007] Optionally, the template feature matching method includes:
[0008] Step S101: Obtain the loaded template image and the global image or local image to be matched;
[0009] Step S102: After scaling according to the requirements of the network model, input the template image and the search area image into the siamese network model to obtain the template feature map and the search area feature map;
[0010] Step S103: Perform sliding filtering calculation of the template feature map on the search area feature map to obtain the response feature map, and obtain the maximum response value of the response map and its position;
[0011] Step S104: Globally traverse the response feature map obtained in step S103 to obtain the maximum response value, denoted as P, and record the position of the maximum value as (Px, Py); use the response values above, below, to the left, and to the right of the maximum value point to perform adjacent interpolation to calculate the sub-pixel coordinate positions , and map the sub-pixel coordinates on the response feature map back to the original image.
[0012] Step S105: The maximum value P recorded in the response map obtained after the first global matching is a reference condition for determining whether the matching is successful after meeting the local matching conditions. If the maximum value of the local matching response map is less than P, it is determined that the matching is unsuccessful, and the matching continues. If it fails continuously for a set number of times, it transfers to global matching.
[0013] Optionally, the loaded template image refers to the target image that has been determined to need tracking, and the specific loading method is offline loading or online loading;
[0014] The global image refers to the image captured by the imaging device, and the local image refers to the image within a set size range centered on the target position in the global image.
[0015] Optionally, the siamese network model realizes simultaneously performing inference calculations on the template image and the search area image using a deep convolutional neural network with the same weight parameters to extract their respective feature maps; before the template image and the search area image are input to extract features by the siamese network model, bilinear interpolation image scaling operations need to be performed on the input images respectively according to the input size designed by the network model;
[0016] The siamese network model is an application model that has been pre-trained offline. Obtain the position information of consecutive multiple frames of sample images of different targets as sample data, and label the sample data, and use the labeled sample data for training to obtain this application model.
[0017] Optionally, in step S103, the template feature map is slid and matched with the initially predicted search area feature map to calculate the response feature map of the template feature map in the initially predicted search area feature map; determine the position information of the tracking target in the current frame through the response feature map and use the original size of the template target as the initial size of the matching target; among them, the filtering calculation method involved in the sliding matching is as follows: Correlation(F,X)=O, where F represents the template feature map matrix, X represents the search area feature matrix, and O is calculated as the response feature map.
[0018] Optionally, use the response values above, below, to the left, and to the right of the maximum value point to perform adjacent interpolation to calculate the sub-pixel coordinate positions
[0019] The formula is as follows:
[0020]
[0021] d represents the result of interpolation calculation using adjacent pixel values, represents the response value on the left side of the maximum point, represents the response value on the right side of the maximum point, represents the calculated sub-pixel deviation value, represents the sub-pixel deviation value in the x-axis direction, represents the sub-pixel deviation value in the y-axis direction, represents the response value above the maximum point, represents the response value below the maximum point, represents the abscissa of the sub-pixel, represents the ordinate of the sub-pixel, represents the abscissa of the maximum value, represents the ordinate of the maximum value;
[0022] The calculation formula for mapping the sub-pixel coordinates on the response feature map back to the original image is as follows,
[0023]
[0024] where represents the sub-pixel abscissa mapped back to the original image, represents the sub-pixel ordinate mapped back to the original image, Fw, Fh represent the width and height of the response feature map, Iw, Ih represent the width and height of the original image, Nw, Nh represent the input width and height of the feature extraction network model, and Rw and Rh respectively represent the scaling factors for obtaining the width and height of the response map from the scaled original image through the network model.
[0025] Optionally, the object tracking and re-identification method includes:
[0026] Step S201: For the loaded object template, extract its fully connected features by inputting it into a pre-trained re-identification network model, and input the object's fully connected features into a feature pool with a set capacity of M as the first feature with the highest confidence in the feature pool;
[0027] Step S202: Initialize the tracker with the initial position obtained by the loaded object template according to the template feature matching method and enter the continuous tracking state;
[0028] Step S203: Set the feature pool capacity to M, store it in a double-ended queue, update and delete the fully connected features at any specified position; in the first M - 1 frames after entering the tracking state, the predicted target position of the tracker is regarded as a reliable position, and the predicted target area is input into the re-identification network model to extract its fully connected features and store them in the feature pool in sequence;
[0029] Step S204: When the feature pool reaches the preset capacity, calculate the cosine distance of the weighted average of the fully connected features proposed by the re-identification network for the target region predicted by the tracker and all the features in the feature pool. The cosine distance calculation formula is as follows:
[0030]
[0031] Where represents the feature vector A to be compared, represents the feature item B to be compared with it, represents the included angle between the directions of feature vectors A and B, represents the modulus length of feature vector A, represents the modulus length of feature vector B, represents the eigenvalue at each position of feature vector A, is the eigenvalue at each position representing feature vector B; The calculation formula for the weighted average cosine distance between the feature to be compared and all the features in the feature pool is as follows:
[0032]
[0033] and and... are the cosine distances between the feature to be compared and the features in the feature pool respectively, and and... are the weight coefficients for calculating the similarity between the template features in the feature pool and the feature to be compared respectively;
[0034] Step S205: Compare the weighted average cosine distance value C = obtained by the method in Step S204 with the set threshold = 0.5. If C ≤ then it is determined that the state of the target position predicted by the current tracker is normal. If C > then it is determined that the target state is abnormal;
[0035] Step S206: For the target features that have met the determination conditions in Step S205, further perform an update condition determination. Compare the weighted average cosine distance value C obtained by the method in Step S204 with the set threshold = 0.35. If C ≤ then it is determined that the fully connected features extracted for the target position predicted by the current tracker are reliable and are updated into the feature pool. If C > then no update is performed;
[0036] Step S207: For the target area predicted by the tracker, if it is continuously determined that the determination conditions of Step S205 and Step S206 are not met for N times, then a region with a size 5 times that of the target is intercepted in the global image with the most recent predicted target area as the center for local feature matching, where 5 ≤ N ≤ 10;
[0037] Step S208: For the response extremum of the local feature matching in Step S207 is compared with the extremum P of the first global matching response map; if then it is determined that the local feature matching is successful, and the successful determination transfers to Step S202 for tracker initialization and tracking, otherwise it is determined that the local matching fails;
[0038] Step S209: For the determination condition in Step S208, if it is continuously determined to be a failure for N times, global feature matching is entered with the most recent tracking coordinates as the center, where 5 ≤ N ≤ 10.
[0039] Optionally, in Step S206, when the feature pool is updated, the loaded target template features are always not deleted, the second position features are deleted, and the newly input features are placed at the last position of the feature pool.
[0040] In the automatic tracking and re-identification method for template feature matching provided by the present invention, convolutional features are extracted from the loaded template and matched in the global features to determine the position and transfer to the tracking process. During the tracking process, fully connected features of the target are continuously extracted and compared with the feature pool to determine the tracking state. If the tracking continues to be abnormal, local matching is performed for re-identification and then tracking. The target tracking method provided by the present invention can accurately predict the position of the tracking target for tracking targets in different operating states, improving the accuracy of target tracking; the determination of the target tracking state is strengthened by combining feature matching and comparison, and the target can be re-captured after the target tracking determination is abnormal, realizing long-term tracking of tracking targets in different states, and greatly improving the robustness and universality of the tracking target. Description of the Drawings
[0041] Figure 1 is a schematic flowchart of the target template feature matching method provided by the present invention;
[0042] Figure 2 is a schematic diagram of the Siamese network structure model for template feature matching;
[0043] Figure 3 is a schematic diagram of sliding filter calculation;
[0044] Figure 4 is a schematic diagram of the template feature matching effect;
[0045] Figure 5It is a schematic flowchart of the object tracking and re-identification method provided by the present invention;
[0046] Figure 6 It is a schematic diagram of the re-identification network model for extracting the fully connected features of the object;
[0047] Figure 7 It is a schematic diagram of the feature pool update method;
[0048] Figure 8 It is a schematic framework diagram of an embodiment of the object tracking device;
[0049] Figure 9 It is a schematic structural diagram of an embodiment of the object tracking device;
[0050] Figure 10 It is a schematic framework diagram of an embodiment of the storage device. Detailed implementation manners
[0051] The following further elaborates in detail a method for automatic tracking and re-identification of template feature matching proposed by the present invention in conjunction with the accompanying drawings and specific embodiments. According to the following description and the claims, the advantages and features of the present invention will be clearer. It should be noted that the accompanying drawings are all in a very simplified form and use non-precise scales, only for the purpose of conveniently and clearly assisting in explaining the purpose of the embodiments of the present invention.
[0052] The terms "first" and "second" in this application are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of this application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined. All directional indications (such as up, down, left, right, front, back...) in the embodiments of this application are only used to explain the relative positional relationship and movement conditions between components in a specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indications will also change accordingly. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include steps or units not listed, or may optionally further include other steps or units inherent to these processes, methods, products, or devices.
[0053] Figure 1 It is a schematic flowchart of an embodiment of a method for object template feature matching provided by the present invention, including the following steps:
[0054] Step S101: Obtain the loaded template image and the global image or local image to be matched.
[0055] The loaded template image refers to the target image that has been determined to be tracked, and the specific loading method can be offline loading or online loading. The global image refers to the image captured by the imaging device, and the local image refers to the image within a set size range centered on the target position in the global image.
[0056] It can be understood that when using the method of this embodiment for target tracking, the number of targets is one, that is, there can be one or more targets in each frame of reference image and subsequent images to be tracked, but tracking will only be performed on the target with the maximum response after matching. The type of the tracked target is not limited either, and can be pedestrians, vehicles, various animals, etc.
[0057] Step S102: After scaling the input size according to the requirements of the network model, input the template image and the search area image into the Siamese network model to obtain the template feature map and the search area feature map.
[0058] In this embodiment, the Siamese network model can implement the inference calculation of the template image and the search area image simultaneously using a deep convolutional neural network with the same weight parameters to extract their respective feature maps. Before the template image and the search area image are input to the Siamese network model to extract features, bilinear interpolation image scaling operations need to be performed on the input images to be input according to the input size designed by the network model.
[0059] The Siamese network model described in the present invention is as Figure 2 shown, and can be an application model that has been pre-trained offline. Specifically, obtain different target sample images, where the targets in the sample images are not limited, and can be pedestrians, vehicles, various animals, etc.; obtain the position information of different target consecutive multi-frame sample images as sample data, and label the sample data, and use the labeled sample data for training to obtain the application model.
[0060] Step S103: Perform sliding filtering calculation of the template feature map on the search area feature map to obtain the response feature map, and obtain the maximum response value of the response map and its position.
[0061] Specifically, obtain the template feature map and the initial predicted search area feature map; slide and match the template feature map with the initial predicted search area feature map, and calculate the response feature map of the template feature map in the initial predicted search area feature map; determine the position information of the tracked target in the current frame through the response feature map and use the original size of the template target as the initial size of the matching target. Figure 4It is the effect diagram of global feature matching. The target template is at the upper left corner, and the position of the target in the global image is found through the template feature matching method of the present invention.
[0062] The filtering calculation method involved in sliding matching is as follows: Correlation(F, X) = O, where F represents the template feature map matrix, X represents the search area feature matrix, and O is obtained as the response feature map through calculation. For the specific calculation method, please refer to Figure 3 .
[0063] Step S104: Globally traverse the response feature map calculated in step S103 to obtain the maximum response value and record it as P, and record the position of the maximum value as (Px, Py). Use the response values above, below, left, and right of the maximum value point to perform adjacent interpolation to calculate the sub-pixel coordinate position The formula is as follows:
[0064]
[0065] d represents the result of interpolation calculation using adjacent pixel values, represents the response value on the left side of the maximum value point, represents the response value on the right side of the maximum value point, represents the calculated sub-pixel deviation value, represents the sub-pixel deviation value in the x-axis direction, represents the sub-pixel deviation value in the y-axis direction, represents the response value above the maximum value point, represents the response value below the maximum value point, represents the abscissa of the sub-pixel, represents the ordinate of the sub-pixel, represents the abscissa of the maximum value, represents the ordinate of the maximum value.
[0066] The calculation formula for mapping the sub-pixel coordinates on the response feature map back to the original image is as follows, where represents the sub-pixel abscissa mapped back to the original image, represents the sub-pixel ordinate mapped back to the original image, Fw and Fh represent the width and height of the response feature map, Iw and Ih represent the width and height of the original image, Nw and Nh represent the input width and height of the feature extraction network model, and Rw and Rh respectively represent the scaling coefficients of the width and height of the response map obtained after scaling the original image through the network model:
[0067]
[0068] Step S105: For the maximum value P recorded in the response graph obtained after the first global match, it is a reference condition for determining success or failure of the matching success after subsequent local matching conditions are met. If the maximum value of the local matching response graph is less than P, it is determined that the matching is unsuccessful, and the matching continues. If the set number of consecutive failures occurs, the global matching is entered.
[0069] Please refer to Figure 5 which is a schematic flowchart of an embodiment of the target tracking and re-identification method provided by the present application, including the following steps:
[0070] Step S201: First, for the loaded target template, input it into the pre-trained re-identification network model as shown in Figure 6 to extract its fully connected features, and input the target fully connected features into a feature pool with a set capacity of M as the first feature with the highest confidence in the feature pool.
[0071] Step S202: Initialize the tracker with the initial position obtained by the target template feature matching method of the present application for the loaded target template and enter the continuous tracking state. The tracker here is an arbitrary multi-scale target tracker, such as methods like KCF, SiamRPN, etc.
[0072] Step S203: Set the capacity of the feature pool to M, which is stored in a double-ended queue, and the fully connected features at any specified position can be updated and deleted. In the first M - 1 frames after entering the tracking state, the target positions predicted by the tracker are regarded as reliable positions, and the predicted target regions are input into the re-identification network model to extract their fully connected features and stored in the feature pool in sequence.
[0073] Step S204: When the feature pool reaches the saturation state, calculate the cosine distance by weighted averaging the fully connected features proposed by the re-identification network for the target regions predicted by the tracker and all the features in the feature pool. The cosine distance range is [0, 1.0], and the smaller the cosine distance, the higher the feature similarity. Specifically, the cosine distance calculation formula is as follows:
[0074]
[0075]
[0076] where represents the feature vector A to be compared, represents the feature item B to be compared with it, represents the direction angle between the feature vectors A and B, represents the modulus of the feature vector A, represents the modulus of the feature vector B, represents the feature value at each position of the feature vector A, is the eigenvalue representing each position of the feature vector B; the calculation formula for calculating the weighted average cosine distance between the feature to be compared and all features in the feature pool is as follows:
[0077]
[0078] and ... are respectively the cosine distances between the feature to be compared and the features in the feature pool, and ... are respectively the weight coefficients for calculating the similarity between the template features in the feature pool and the feature to be compared;
[0079] Step S205: Determination of the tracker state: According to the weighted average cosine distance value C = calculated by the method in Step S204, compare it with the set threshold = 0.5. If C ≤ then it is determined that the current tracker predicts that the target position state is normal. If C > then it is determined that the target state is abnormal.
[0080] Step S206: Further determination of the update condition for the target feature that has satisfied the determination condition in Step S205. Specifically, compare the weighted average cosine distance value C calculated by the method in Step S204 with the set threshold = 0.35. If C ≤ then it is determined that the fully connected feature extracted by the current tracker to predict the target position is reliable and is updated into the feature pool. If C > then no update is performed. When updating the feature pool, always keep the loaded target template feature without deletion, delete the second position feature, and put the newly input feature at the last position of the feature pool. The update method is as shown in Figure 7 .
[0081] Step S207: For the target area predicted by the tracker, if the determination conditions in Step S205 and Step S206 are not satisfied continuously for N times, then a region with a size of S = 5 times the target size is intercepted in the global image centered on the most recent predicted target area for local feature matching, where 5 ≤ N ≤ 10.
[0082] Step S208: Compare the response extremum of the local feature matching in Step S207 with the extremum P of the first global matching response map; if then it is determined that the local feature matching is successful, and the successful one transfers to Step S202 for tracker initialization and tracking, otherwise it is determined that the local matching fails.
[0083] Step S209: For the determination condition in step S208, in the case where the determination fails continuously for N times, enter global feature matching with the most recent tracking coordinates as the center, where 5 ≤ N ≤ 10.
[0084] In this embodiment, convolutional features are extracted for the loaded target template. The target tracking device acquires a continuous number of frames of images to be processed and extracts the global convolutional feature map of the first frame. The template features and the global convolutional feature map are used for template feature matching calculation by means of sliding filtering to obtain the extreme value of the response map to determine the target position and size. The tracker is initialized and transferred to a continuous tracking state. During the prediction of the tracker, the fully connected features of the predicted region are extracted using the re-identification network model and compared with the initialized feature pool to determine the state of the tracker. For the case of continuous abnormal tracking, local matching is first performed. If the continuous matching fails, global matching is performed to re-initialize the tracker. For the normal tracking target region, fully connected features are extracted and judged according to the set threshold, and the feature pool is updated to maintain the adaptability of the feature pool to the changes in the target shape and posture. Through the above method, the present application can adapt to the adaptive tracking of specified targets of any type, and can determine the state predicted by the tracker during the tracking process, and can re-identify the abnormal tracking situation to complete the re-initialization of the tracker, thereby greatly improving the robustness and long-term performance of target tracking.
[0085] Please refer to Figure 8 , which is a schematic framework diagram of an embodiment of the target tracking device provided by the present invention. The target tracking device 30 includes:
[0086] An image acquisition module 31, configured to acquire the image collected by the camera and load the loaded target template.
[0087] A feature matching module 32, configured to perform sliding filtering feature matching on the target template features and the convolutional feature map extracted from the global or local image to obtain a response feature map, and calculate the corresponding position of the template in the original image.
[0088] A re-identification feature extraction module 33, configured to extract the fully connected features of the target template and the target region predicted by the tracker, and the features are used for comparison with the feature pool to determine the state of the tracker.
[0089] A feature pool management module 34, which includes functions of cosine distance weighted calculation and feature pool update strategy.
[0090] A tracking module 35, which is a multi-scale tracker, initializes the tracker using the region given by the feature matching, and continuously predicts the target region after entering the tracking.
[0091] Please refer to Figure 9, which is a schematic framework diagram of another embodiment of the target tracking device provided by this application. Specifically, in this embodiment, the target tracking device 40 includes a memory 41, a processor 42, and an imaging device 43 that are coupled to each other. Among them, the memory 41 is used to store program instructions and target templates, and the processor 42 is used to execute the program instructions stored in the memory 41 to implement the steps in any of the above-mentioned target tracking method embodiments. The imaging device 43 captures real-time images. In a specific implementation scenario, the target tracking device 40 may include, but is not limited to: a microcomputer, a server. In addition, the target tracking device 40 may also include mobile devices such as a laptop computer and a tablet computer, which are not limited herein.
[0092] Specifically, the memory 41 may be a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc., which are media that can store program instructions.
[0093] Specifically, the processor 42 is used to control itself and the memory 41 to implement the steps in any of the above-mentioned target tracking method embodiments. The processor 42 may also be referred to as a CPU (Central Processing Unit). The processor 42 may be an integrated circuit chip with signal processing capabilities. The processor 42 may also be a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 42 may be implemented jointly by integrated circuit chips.
[0094] In one embodiment, the target tracking device may further include an imaging device 43, and the processor 42 is further used to control the imaging device 43 so that the imaging device 43 captures the target scene to obtain an image containing the target captured by the imaging device.
[0095] Please refer to Figure 10 , Figure 10 , which is a schematic framework diagram of an embodiment of the computer-readable storage medium 50 of this application.
[0096] In an embodiment of the present application, a computer-readable storage medium 50 is further provided, storing program instructions 51 that can be run by a processor, and the program instructions 51 are used to implement the steps in the embodiments of any of the above target tracking and matching methods.
[0097] In several embodiments provided by the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.
[0098] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0099] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0100] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in each embodiment of the present application.
[0101] The above description is only a description of the preferred embodiments of the present invention, and does not limit the scope of the present invention in any way. Any changes and modifications made by those of ordinary skill in the art of the present invention based on the above disclosure are within the protection scope of the claims.
Claims
1. An automatic tracking and re-identification method for template feature matching, characterized in that It includes a template feature matching method and an object tracking re-identification method; By extracting convolutional features from the loaded template, the position is determined by matching in the global features and transferred to the tracking process; during the tracking process, the fully connected features of the object are continuously extracted and compared with the feature pool to determine the tracking state. If the tracking continues to be abnormal, local matching is performed for re-identification and then tracking; The template feature matching method includes: Step S101: Obtain the loaded template image and the global image or local image to be matched; Step S102: After scaling the input size according to the requirements of the network model, the template image and the search area image are input into the Siamese network model to obtain the template feature map and the search area feature map; Step S103: The template feature map performs sliding filtering calculation on the search area feature map to obtain the response feature map, and the maximum response value and its position of the response map are obtained; Step S104: Globally traverse the response feature map calculated in Step S103 to obtain the maximum response value and record it as P, and record the position where the maximum value is located as (Px, Py); use the response values above, below, left, and right of the maximum point to perform adjacent interpolation to calculate the sub-pixel coordinate position P(x', y'), and map the sub-pixel coordinates on the response feature map back to the original image; Step S105: The maximum value P recorded in the response map obtained after the first global match is a reference condition for determining whether the matching is successful after meeting the local matching conditions; if the maximum value of the local matching response map is less than P, it is determined that the matching is unsuccessful, and continue to match; if it fails continuously for a set number of times, transfer to global matching; The object tracking re-identification method includes: Step S201: Input the loaded object template into the pre-trained re-identification network model to extract its fully connected features, and input the fully connected features of the object into the feature pool with a set capacity of M as the first feature with the highest confidence in the feature pool; Step S202: Initialize the tracker with the initial position obtained by the loaded object template according to the template feature matching method and transfer to the continuous tracking state; Step S203: Set the capacity of the feature pool to M, store it as a double-ended queue, and update and delete the fully connected features at any specified position; in the first M - 1 frames after transferring to the tracking state, the target position predicted by the tracker is regarded as a reliable position, and the predicted target area is input into the re-identification network model to extract its fully connected features and store them in the feature pool in order; Step S204: When the feature pool reaches the preset capacity, calculate the weighted average cosine distance between the fully connected features proposed by the re-identification network for the target area predicted by the tracker and all the features in the feature pool. The cosine distance calculation formula is as follows: Among them represents the feature vector A to be compared represents the feature item B to be compared with it, θ represents the direction angle between the feature vector A and B, ||x|| represents the norm of the feature vector A, ||y|| represents the norm of the feature vector B, and x i represents the eigenvalue of each position of the feature vector A, and y i represents the eigenvalue of each position of the feature vector B; the calculation formula for the weighted average cosine distance q between the feature to be compared and all features in the feature pool is as follows: x1, x2, ... x m are the cosine distances between the features to be compared and the features in the feature pool, respectively. f1, f2, ... f m are the weight coefficients for calculating the similarity between the template features in the feature pool and the features to be compared, respectively; Step S205: Compare the weighted average cosine distance value C = Sim(X, Y) calculated according to the method in Step S204 with the set threshold T0 = 0.
5. If C ≤ T0, it is determined that the current tracker predicts the target position state is normal. If C > T0, it is determined that the target state is abnormal; Step S206: For the target feature that has met the determination condition in step S205, further perform an update condition determination. Compare the weighted average cosine distance value C calculated according to the method in step S204 with the set threshold T1 = 0.
35. If C ≤ T1, it is determined that the fully connected feature extracted by the current tracker to predict the target position is reliable and is updated into the feature pool. If C > T1, no update is performed; Step S207: For the target area predicted by the tracker, if it is continuously determined N times that it does not meet the determination conditions in step S205 and step S206, then a region with a size S = 5 times the target size is intercepted in the global image centered on the most recent predicted target area for local feature matching, where 5 ≤ N ≤ 10; Step S208: Compare the response extreme value P of the local feature matching in step S207 a with the extreme value P of the first global matching response map; if P a ≥P, it is determined that the local feature matching is successful, and if the determination is successful, proceed to step S202 for tracker initialization and tracking, otherwise it is determined that the local matching fails; Step S209: For the determination condition in step S208, if it is continuously determined N times to be a failure, enter global feature matching centered on the most recent tracking coordinates, where 5 ≤ N ≤ 10.
2. The automatic tracking and re-identification method for template feature matching according to claim 1, wherein, The loaded template image refers to the target image that has been determined to need tracking. The specific loading method is offline loading or online loading; The global image refers to the image captured by the imaging device, and the local image refers to the image within a set size range centered on the target position in the global image.
3. The automatic tracking and re-identification method for template feature matching according to claim 1, characterized in that The siamese network model realizes simultaneously performing inference calculations on the template image and the search area image using a deep convolutional neural network with the same weight parameters to extract their respective feature maps; before the template image and the search area image are input to the siamese network model to extract features, they need to be bilinearly interpolated and image scaled according to the input size designed by the network model for the input images respectively; The siamese network model is an application model that has been pre-trained offline. The position information of consecutive multiple frames of sample images of different targets is obtained as sample data, and this sample data is labeled. The application model is trained using the labeled sample data.
4. The automatic tracking and re-identification method for template feature matching according to claim 1, characterized in that In step S103, the template feature map is slid and matched with the initially predicted search area feature map to calculate the response feature map of the template feature map in the initially predicted search area feature map; Determine the position information of the tracking target in the current frame through the response feature map and use the original size of the template target as the initial size of the matching target; among them, the filtering calculation method involved in the sliding match is as follows: Correlation(F,X) = O, where F represents the template feature map matrix, X represents the search area feature matrix, and O is calculated as the response feature map.
5. The automatic tracking and re-identification method for template feature matching according to claim 1, characterized in that The formula for calculating the sub-pixel coordinate position P(x',y') by performing adjacent interpolation using the response values above, below, left, and right of the maximum value point is as follows: d represents the result of interpolation calculation using adjacent pixel values, P left represents the response value on the left side of the maximum point, P right represents the response value on the right side of the maximum point, diff represents the calculated sub-pixel deviation value, diffx represents the sub-pixel deviation value in the x-axis direction, diffy represents the sub-pixel deviation value in the y-axis direction, P up represents the response value above the maximum point, P down represents the response value below the maximum point, Px' represents the abscissa of the sub-pixel, Py' represents the ordinate of the sub-pixel, Px represents the abscissa of the maximum value, Py represents the ordinate of the maximum value; The formula for mapping the sub-pixel coordinates on the response feature map back to the original image is as follows, where Ix represents the sub-pixel abscissa mapped back to the original image, Iy represents the sub-pixel ordinate mapped back to the original image, Fw, Fh represent the width and height of the response feature map, Iw, Ih represent the width and height of the original image, Nw, Nh represent the input width and height of the feature extraction network model, and Rw and Rh respectively represent the scaling coefficients of the width and height of the response map obtained after scaling the original image through the network model.
6. The automatic tracking and re-identification method for template feature matching according to claim 1, characterized in that In the step S206, when the feature pool is updated, the feature of the loaded target template is always kept without deletion, the feature at the second position is deleted, and the newly input feature is placed at the last position of the feature pool.
Citation Information
Patent Citations
Cross-camera pedestrian detection tracking method based on depth learning
CN108875588A
Target tracking method, terminal and computer readable storage medium thereof
CN112037257A