Small target cluster tracking method based on memory enhancement
By adopting a memory-enhanced small-target cluster tracking method in aerial videos, using the deepening memory network and memory module to extract features, and improving tracking performance through memory compensation matching strategies, the problem of non-repeating crowd counting in aerial videos is solved, and efficient small-target tracking and crowd counting are achieved.
Patent Information
- Application Number
- CN202510297787.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-12-30
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art is difficult to achieve non-repetitive population counting in aerial videos, especially when the ability to describe small target features is limited, resulting in repeated statistics on individuals.
Using a small-target cluster tracking method based on memory enhancement, the deepening memory network and memory module are constructed, and the deepening memory characteristics of the human head are extracted, and the deviation of the detection target is reconstructed to determine the trajectory and the matching of the detection target.
The reproduction performance of small target tracking is improved, and the problem of indistinguishability of small target appearance characteristics and trajectory representation is effectively solved, and the goal of not repeating crowd counting is achieved.
Smart Images

Figure CN120147364A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of video processing and artificial intelligence, and particularly relates to a small target cluster tracking method based on memory enhancement. Background Art
[0002] In large-scale events, drones have begun to be used for crowd monitoring tasks due to their advantages such as flexibility and wide surveillance range. However, to ensure the integrity of information, there is usually a large overlapping area between consecutive frames of aerial videos. When counting the number of people, if the number of people in each frame is directly accumulated, a large number of individuals will be double-counted. Therefore, how to achieve non-repetitive crowd counting in aerial videos has become an urgent problem to be solved. Due to the wide shooting range of aerial videos, individuals usually appear as small targets, and their feature description ability is limited. The goal of pedestrian re-identification (Re-Identification, Re-Id) is to distinguish whether different targets are the same individual, which mainly relies on appearance features and is not applicable to small targets with unclear appearance features. The goal of non-repetitive crowd counting in aerial videos is how to effectively obtain the video sequences of each target. In addition, in the panoramic image obtained by stitching, moving targets will appear multiple times, and it is difficult to remove duplicate individuals. Multi-target tracking, as a basic visual perception task, is expected to be a solution to non-repetitive crowd counting. During the tracking process, the detector locates the bounding box of the target frame by frame and assigns a track IDentity (ID) to the detected target according to its similarity to the existing tracks. Through this continuous association, the complete track of each person in the video can be obtained. The number of tracks corresponds to the number of non-repetitive people.
[0003] Currently, most tracking methods consider appearance and location information during the process of associating historical trajectories with detected targets. The calculation of appearance similarity involves the features of the targets detected by the detector and the representation of the trajectories. The extraction of distinguishable target features belongs to the research field of target recognition and classification. The trajectory representation that is extended by continuously associating with detected targets is a core issue unique to target tracking and can be divided into three types. (1) When using the trajectory representation based on the set of detection features to calculate the similarity with the detected target, the minimum value of the distances between the detected object and each element in the set is used as the similarity between the detected object and the trajectory. This essentially represents the entire trajectory with a single target feature. DeepSort, based on SimpleOnline and Realtime Tracking (SORT), adopts an independent Re-Id model to extract appearance features from the detection bounding boxes and add them to the similarity calculation. A cascaded matching strategy is proposed, which preferentially matches the trajectories that have not been continuously tracked and then matches with the lost trajectories. During the frame-by-frame association process, due to interference factors such as occlusion and blur, once individual samples with other ID tags are mixed into the feature set of the current trajectory, that is, the set is contaminated, it is extremely easy to cause confusion in the subsequent association between the trajectory and the detection. (2) The trajectory representation based on the fusion of detection features usually adopts an update strategy based on the dynamic equation or fuses the features of the targets on the trajectory into vectors of equal size during the aggregation process to represent the trajectory. However, the discriminative ability of most small target features is limited, and fusing them will make the trajectory representation homogenized and difficult to distinguish. StrongSORT upgrades DeepSort from multiple perspectives such as detection, embedding, and association, and uses Exponential Moving Average (EMA) to fuse the detection target features and the trajectory features, which makes the representation of the trajectory tend to be the same and is not conducive to the association between the detected target and the existing trajectory. (3) The trajectory representation based on the latest detection features only uses the features of the detected targets in the last one or two frames of the trajectory to represent the trajectory, rather than the features of all the detected targets on the trajectory. This undoubtedly wastes a large amount of valuable information, and due to the very limited description ability of small target features, the trajectory representation is vulnerable to interference. Summary of the Invention
[0004] The object of the present invention is to overcome the defects of the prior art and provide a small target cluster tracking method based on memory enhancement, which supplements memory compensation matching in the association strategy for unreliable position prediction. This strategy reconstructs the deviation of the detected target according to the memory module of the trajectory to determine whether the trajectory and the detected target match, thereby improving the performance of tracking and reproducing the target.
[0005] The object of the present invention is achieved as follows: A small target cluster tracking method based on memory enhancement, comprising the following steps:
[0006] 1) Construct a dataset for non - duplicate crowd counting from aerial photography data into a video frame sequence;
[0007] 2) Detect targets using a pre - trained crowd localization network based on inverse focal length transformation;
[0008] 3) Input the video frame and the localization coordinates into a deep memory network to extract the deep memory features of human heads;
[0009] 4) Construct a loss function to train the deep memory network;
[0010] 5) Track the test dataset and construct memory items for the tracked trajectories;
[0011] 6) Associate the existing trajectories with the detected targets in the current frame;
[0012] 7) Update the trajectories and the detected targets according to the tracking results;
[0013] 8) Perform steps 6) and 7) on subsequent video frames until all video frames are processed.
[0014] As a further limitation of the present invention, step 3) specifically includes: the deep memory network is composed of parallel spatial attention and channel attention; the process of extracting the deep memory features of human heads is as follows:
[0015] From the processed image features, according to the coordinate positions of the crowd localization network, extract fixed - size human head features and transpose them to obtain a human head feature map Fm with an output size of H×W×C; the extracted human head features are simultaneously input into the channel attention and spatial attention modules to calculate the spatial attention weight matrix Wp and the channel attention weight matrix Wc respectively;
[0016] Product of the feature maps D = H×W and Pass through the softmax function to calculate the weight matrix The formula is as follows:
[0017]
[0018] where i,j∈[1,D] respectively represent the row and column indices; ":" represents all elements of a row or a column of the matrix; Ms is weighted to become The calculation formula is Mp = Wp×Ms.
[0019] Product of the feature maps and Sent into the softmax function to calculate the weight matrix The formula is as follows:
[0020]
[0021] where \(i, j\in[1, C]\) represent the indices of rows and columns respectively; \(F_s\) is weighted to become The calculation formula is \(F_c = F_s\times W_c\);
[0022] Deformation operations are performed on the weighted feature map \(M_p\) and \(F_c\) respectively, and then a summation operation is performed with \(F_M\); the fused features undergo convolution and downsampling operations to output deep memory features of size \(H\times W\times C\).
[0023] As a further limitation of the present invention, step 4) specifically includes: inputting the deep memory features into an ID classifier, and the ID classifier includes a batch normalization layer, a correction layer, a ReLu layer, a dropout layer, and two fully connected layers; the ID of the same person remains unchanged during training; assuming that the encoding length of the individual ID is \(L\); the cross-entropy loss between the labeled individual ID encoding distribution \(Y\) and the predicted is calculated as follows:
[0024]
[0025] where \(Y(l)\) represents the value of the \(l\)-th bit of the individual's true (or actual) ID encoding (the ground truth encoding of the individual ID); represents the probability of the \(l\)-th bit predicted by the classifier.
[0026] As a further limitation of the present invention, the tracking of the test data set in step 5) specifically includes: constructing a memory item for the trajectory, online updating of the memory module, and constructing a memory bias;
[0027] Constructing a memory item for the trajectory: Each trajectory creates a memory module \(Mem\) composed of \(M\) memory items \((m = 1,\ldots,M)\), and its initial state is a random value;
[0028] Online updating of the memory module: In the current \(t\)-th frame, the deep memory feature of the detection target successfully associated with the \(k\)-th trajectory is Each feature vector \(d = 1,2,\ldots,D\) is extracted from the deep memory feature matrix \(Det\) t (k) along the channel for updating the memory item; each feature vector \(d = 1,2,\ldots,D\) and the \(m\)-th memory item The product of is fed into the softmax function to calculate the weight \(W_s\) of the feature vector , and the formula is as follows: d
[0029]
[0030] Among them, tr represents the transpose operation of a matrix, and Ws d represents the d-th eigenvector belonging to the appearance information stored in the m-th memory item of the k-th trajectory ; the eigenvector d = 1, 2, …, D is based on the corresponding weight Ws d weighted average is calculated as follows:
[0031]
[0032] The weighted average Δ is fused with the m-th memory item in the (t - 1)-th frame to form the memory item in the t-th frame which is calculated as follows:
[0033]
[0034] The memory module Mem of the k-th trajectory in the t-th frame can be obtained by updating each memory item of the trajectory in the (t - 1)-th frame in the above manner t ;
[0035] Construct the memory deviation: In the (t - 1)-th frame, the M memory items m = 1, …, M form the memory module Mem of the k-th trajectory t-1 ; Assume that a total of N targets are detected in the t-th frame; the enhanced memory feature of the n-th detected target is represented as Each eigenvector is extracted from the feature matrix Det of the target along the channel t (n) d = 1, 2, …, D;
[0036] The process of reconstructing the d-th eigenvector using the memory module Mem of the k-th trajectory t-1 ; The product of each memory item Mem and the d-th eigenvector t-1 is input into the softmax function to calculate the weight Wo of the memory item The formula is as follows: m which is as follows:
[0037]
[0038] Among them, tr represents the transpose operation of a matrix, and Wo m represents the appearance details recorded in the m-th memory item existing in the d-th eigenvector ; The reconstructed feature of the d-th eigenvector is calculated as follows:
[0039]
[0040] wherein the d-th eigenvector has a reconstructed vector of The n-th detected target feature Det t (n) Each eigenvector is reconstructed in the above manner. The cosine distance between is calculated as follows:
[0041]
[0042] The cosine distance is the deviation between the reconstructed memory module Mem t-1 and the enhanced memory feature Det t of the n-th detected target(n).
[0043] As a further limitation of the present invention, step 6) specifically includes: The small target cluster tracking method starts from the second frame of the video and repeatedly associates the existing trajectories with the detected targets in the current frame; Assume that the target detected in the current t-th frame is Det t , and the existing trajectory Trk t-1 consists of a set of trajectories in three states: Active Temporary and ; The detected target Det t and the existing trajectory Trk t-1 need to go through a three-stage matching process;
[0044] First-stage matching: Based on the IoU distance matrix, a distance matrix is constructed using the positions of Det t and the trajectory Trk t-1 in the t-th frame, and the deviation reconstruction Det t-1 of the memory module of Trk t forms an association cost matrix; Then, the Hungarian algorithm is used for optimal assignment; For the successfully matched trajectory Trk t-1 and detection Det t , perform the association success operation; The unmatched detections and trajectories will proceed to the second-stage matching;
[0045] Second-stage matching: The association cost matrix is constructed based on the IoU distance; Perform the association success operation on the successfully matched trajectory Trk t-1 and the detected target Det t ; Perform the association failure operation on the unmatched trajectory Trk t-1 ; The unmatched detected target Det t will proceed to the third-stage matching;
[0046] Third-stage matching: Utilize the sleep state through memory compensation matching The memory module of the trajectory reconstructs the unassociated detected target Det t , and then uses the reconstruction deviation to construct an association cost matrix.
[0047] As a further limitation of the present invention, step 7) specifically includes: The detected target Det with successful association t Updates the corresponding memory module and Kalman filter of the trajectory Trk t-1 ; Trajectories in the Temporary state or Sleep state that are continuously successfully associated 3 times are updated to the Active state;
[0048] The detected target Det with failed association t Is updated to a trajectory Trk with the state of Temporary t ; Trajectories in the Active state that have failed to be associated 30 times in a row will be set to the Sleep state; Trajectories in the Sleep state that have failed to be associated 40 times in a row will be deleted.
[0049] The present invention adopts the above technical solutions. Compared with the prior art, the beneficial effects are as follows: (1) The present invention designs a deep memory network, which only needs to use head point annotations as training data to avoid the huge cost of head bounding box annotations in videos. Here, the feature matrix corresponding to the fixed-scale area around the head point is used to represent the individual; through spatial and channel attention, the target features are focused on optimizing the representation of the foreground head in the head area, which not only improves the ability of the features to describe the same individual, but also enhances the ability of the features to distinguish different targets. (2) This memory-based trajectory representation method integrates the target features on the trajectory into several typical memory items, comprehensively retaining the ability to describe different states of the individual, and improving the problem of extremely difficult to distinguish the appearance features and trajectory representations of different small targets. (3) Factors such as camera movement, individual movement, and occlusion often cause individuals in the field of view to disappear and then reappear. In order to continuously track the reappearing targets, the target trajectory needs to continue to participate in the matching after failing to match with the detected target. The crowd scene captured by the dynamic camera plus the irregular individual movement makes it difficult to use a general motion model to predict the position of the historical trajectory with a complex motion pattern in the current frame. Repeatedly predicting the position of the failed matching trajectory in subsequent frames will accumulate serious errors, resulting in the invalidation of the predicted position. Therefore, the present invention supplements memory compensation matching in the association strategy for unreliable position prediction. This strategy determines whether the trajectory and the detected target match according to the deviation of the detected target reconstructed by the memory module of the trajectory, thereby improving the performance of tracking and reappearing the target. Description of the Drawings
[0050] Figure 1 The architecture of the deep memory network in the present invention.
[0051] Figure 2 The online update flowchart of the trajectory memory module of the present invention.
[0052] Figure 3 The schematic diagram of the calculation process of the trajectory memory deviation of the present invention.
[0053] Figure 4 The overview diagram of the trajectory association strategy of the present invention.
[0054] Figure 5 The visualization diagram of the tracking result of the present invention. Specific implementation manners
[0055] A small target cluster tracking method based on memory enhancement, characterized by comprising the following steps:
[0056] 1) Construct a non-repeating crowd counting dataset into a video frame sequence by using aerial photography data; extract the aerial photography video into a set of single-frame images;
[0057] 2) Detect targets by a pre-trained crowd localization network based on inverse focal length transformation; adopt the crowd localization network model Focal Inverse Distance Transform Maps for CrowdLocalization (FIDTM) pre-trained on the NWPU-Crowd dataset to detect the position coordinates of individuals in the crowd on the video frames;
[0058] 3) Input the video frames and the localization coordinates into the deep memory network to extract the deep memory features of the human heads;
[0059] As Figure 1 shown, the deep memory network is composed of parallel spatial attention and channel attention; the process of extracting the deep memory features of the human heads is as follows:
[0060] According to the coordinate positions of the crowd localization network, extract the human head features with a fixed size of 13×13 from the processed image features, and transpose them to obtain a human head feature map Fm with an output size of H×W×C. The extracted human head features are simultaneously input into the channel attention and spatial attention modules to calculate the spatial attention weight matrix Wp and the channel attention weight matrix Wc respectively;
[0061] The product of the feature maps and pass through the softmax function to calculate the weight matrix The formula is as follows:
[0062]
[0063] where \(i,j\in[1,D]\) respectively represent the indices of rows and columns; ":" represents all elements of a row or a column of the matrix. \(M_s\) is weighted to become The calculation formula is \(M_p = W_p\times M_s\).
[0064] Product of feature maps and are fed into the softmax function to calculate the weight matrix The formula is as follows:
[0065]
[0066] where \(i,j\in[1,C]\) respectively represent the indices of rows and columns; \(F_s\) is weighted to become The calculation formula is \(F_c = F_s\times W_c\);
[0067] The weighted feature maps \(M_p\) and \(F_c\) are respectively subjected to deformation operations, and then a summation operation is performed with \(FM\). The fused features undergo convolution and downsampling operations, and the deepened memory features of size \(H\times W\times C\) are output.
[0068] Table 1 shows the test results of the method (deepened memory network) proposed in the present invention, methods such as ResNet50, bilateral complementary network with temporal kernel selection (BiCNet-TKS), pyramid spatio-temporal aggregation (PSTA), spatio-temporal memory network (STMN), and global guidance-based complementary learning (GRL) on the PRAI-1581 dataset. Although there are interference factors such as occlusion and size differences in this dataset, the method (deepened memory network) proposed in the present invention still has good robustness to these influencing factors. It can be seen from the recognition accuracies from Rank-1 to Rank-20 that the present invention (deepened memory network) has achieved the best recognition performance.
[0069] Table 1 Recognition performance of features on the PRAI-1581 dataset
[0070]
[0071]
[0072] 4) Construct a loss function to train the deepened memory network;
[0073] The deepened memory features are input into an ID classifier, and the ID classifier includes a batch normalization layer, a correction layer, a ReLu layer, a dropout layer, and two fully connected layers; the ID of the same person remains unchanged during training; assume that the encoding length of the individual ID is \(L\); the cross-entropy loss between the labeled individual ID encoding distribution \(Y\) and the predicted is calculated as follows:
[0074]
[0075] Among them, Y(l) represents the value of the l-th bit of the individual's true (or actual) ID code (the ground truth code of the individual ID). represents the probability of the l-th bit predicted by the classifier.
[0076] 5) Track the test data set and construct memory items for the tracked trajectories.
[0077] Construct memory items for the trajectories: Each trajectory creates a memory module Mem consisting of M memory items (m = 1, …, M), and its initial state is a random value.
[0078] As Figure 2 shown, online update of the memory module: In the current t-th frame, the deep memory feature of the detection target successfully associated with the k-th trajectory is each feature vector is extracted from the deep memory feature matrix Det t (k) along the channel for updating the memory item; each feature vector and the m-th memory item 's product is fed into the softmax function to calculate the weight Ws of the feature vector d , and the formula is as follows:
[0079]
[0080] Among them, tr represents the transpose operation of the matrix, and Ws d represents the probability that the d-th feature vector belongs to the appearance information stored in the m-th memory item of the k-th trajectory; the feature vector is weighted averaged d based on the corresponding weight Ws and calculated as follows:
[0081]
[0082] Fuse the weighted average Δ with the m-th memory item in the (t - 1)-th frame to form the memory item in the t-th frame and calculate as follows:
[0083]
[0084] The memory module Mem of the k-th trajectory in the t-th frame can be obtained by updating each memory item of the trajectory in the (t-1)-th frame in the above manner. t ;
[0085] As Figure 3 shown, construct the memory bias: in the (t-1)-th frame, M memory items m = 1, …, M form the memory module Mem of the k-th trajectory t-1 ; Assume that a total of N targets are detected in the t-th frame; the enhanced memory feature of the n-th detected target is denoted as Extract each feature vector t along the channel from the feature matrix Det
[0086] of the target. t-1 Reconstruct the d-th feature vector using the memory module Mem of the k-th trajectory t-1 ; The product of each memory item Mem and the d-th feature vector is input into the softmax function to calculate the weight Wo m of the memory item, and the formula is as follows:
[0087]
[0088] where tr represents the transpose operation of the matrix, and Wo m represents the appearance details recorded in the m-th memory item existing in the d-th feature vector ; The reconstructed feature of the d-th feature vector is calculated as follows:
[0089]
[0090] where the reconstructed vector of the d-th feature vector is The feature Det t (n) of each detected n-th target is reconstructed in the above manner. The cosine distance between is calculated as follows:
[0091]
[0092] The cosine distance is the bias between the reconstructed memory module Mem t-1 and the enhanced memory feature Det t (n) of the n-th detected target.
[0093] 6) Associate the existing trajectory with the detected object in the current frame;
[0094] As Figure 4 shown, the small object cluster tracking method starts from the second frame of the video and repeatedly associates the existing trajectory with the detected object in the current frame; assume that the object detected in the current t-th frame is Det t , and the existing trajectory Trk t-1 consists of a set of trajectories in three states: Active Temporary and Sleep ; the detected object Det t and the existing trajectory Trk t-1 need to go through a three-stage matching process;
[0095] First-stage matching: Based on the Intersection over Union (IoU) distance matrix, use Det t and the position of the trajectory Trk t-1 in the t-th frame to construct a distance matrix, and the deviation of the memory module of Trk t-1 reconstructs Det t to form an association cost matrix; then, use the Hungarian algorithm to optimize the assignment; for the successfully matched trajectory Trk t-1 and the detection Det t , perform the association success operation; the unmatched detections and trajectories will proceed to the second-stage matching;
[0096] Second-stage matching: The association cost matrix is constructed based on the IoU distance; perform the association success operation on the successfully matched trajectory Trk t-1 and the detected object Det t ; perform the association failure operation on the unmatched trajectory Trk t-1 ; the unmatched detected object Det t will proceed to the third-stage matching;
[0097] Third-stage matching: Through memory compensation matching, use the memory module of the sleep state trajectory to reconstruct the unassociated detected object Det t , and then use the reconstructed deviation to construct an association cost matrix.
[0098] 7) Update the trajectory and the detected object according to the tracking result;
[0099] The successfully associated detected object Det t updates the corresponding trajectory Trk t-1The memory module and the Kalman filter; the trajectories in the Temporary state or the Sleep state that are successfully associated 3 times in a row are updated to the Active state;
[0100] The detection target Det with an association failure t Updated to the trajectory Trk with the state of Temporary t ; the trajectories in the Active state with 30 consecutive association failures will be set to the Sleep state; the trajectories in the Sleep state with 40 consecutive association failures will be deleted.
[0101] 8) Perform steps 6) and 7) on subsequent video frames until all video frames are processed; obtain the trajectories of all small targets in the video, that is, the non-repeating number of people. As Figure 5 shown, it is the visualization diagram of the tracking result of the present invention. The left column shows the tracking result on the RiverSide dataset. The right column shows the tracking result on the SquareCorner dataset. In different frames, the same ID label represents the same individual (target) being tracked.
[0102] The present invention provides a method for tracking small target clusters based on memory enhancement. By constructing an independent memory module based on the features of the detection targets on each trajectory and using the deviation of the detection targets reconstructed by the memory module to construct a cost matrix, a memory-based trajectory representation is achieved. This representation integrates the target features on the trajectory into several typical memory items, comprehensively retains the ability to describe individuals in different states, and effectively solves the problem that it is difficult to distinguish the appearance features of small targets from the trajectory expression. In addition, the present invention also designs an association strategy based on memory compensation matching as a supplement to unreliable position prediction. This method is of great significance for analyzing aerial videos and obtaining the population scale of large public scenes.
[0103] The present invention is not limited to the above embodiments. Based on the technical solutions disclosed in the present invention, those skilled in the art can make some substitutions and deformations to some of the technical features without creative labor according to the disclosed technical content, and these substitutions and deformations are all within the protection scope of the present invention.
Claims
1. A small target cluster tracking method based on memory enhancement, characterized in that: The following steps are involved: 1) Using aerial photography data to construct a dataset of non-repetitive crowd counting to video frame sequences; 2) Pre-trained crowd localization network based on inverse focal transform to detect targets; 3) Input the video frame and positioning coordinates into the deep memory network to extract the deep memory features of the head; 4) Construct a loss function to train the deepened memory network; 5) Track the test data set and build memory items for the tracked trajectories; 6) Associate the existing trajectory with the detected target in the current frame; 7) Update the trajectory and detection target according to the tracking results; 8) Perform steps 6) and 7) on subsequent video frames until all video frames are processed.
2. The small target cluster tracking method based on memory enhancement according to claim 1 is characterized in that: The step 3) specifically includes: the deepening memory network is composed of parallel spatial attention and channel attention; the process of extracting the deepening memory feature of the head is: From the processed image features, according to the coordinate position of the crowd positioning network, the fixed-size head features are extracted and transposed to obtain the head feature map Fm with an output size of H×W×C; the extracted head features are simultaneously subjected to the channel attention and spatial attention modules, and the spatial attention weight matrix Wp and the channel attention weight matrix Wc are calculated respectively; Product of feature maps D = H × W and Pass the softmax function to calculate the weight matrix The formula is as follows: Where i, j∈[1,D] represent the row and column indices respectively; ":" represents all elements of a row or column of the matrix; Ms is given a weight and becomes The calculation formula is Mp=Wp×Ms. Product of feature maps and Feed into the softmax function to calculate the weight matrix The formula is as follows: where i, j∈[1,C] represent the row and column indices respectively; Fs is given weights as The calculation formula is Fc = Fs × Wc; The weighted feature maps Mp and Fc are deformed respectively, and then summed with FM; the fused features are convolved and down-sampled to output a deepened memory feature of size H×W×C.
3. The small target cluster tracking method based on memory enhancement according to claim 1 is characterized in that: The step 4) specifically includes: inputting the deepened memory features into an ID classifier, wherein the ID classifier includes a batch normalization layer, a correction layer, a ReLu layer, a dropout layer, and two fully connected layers; the ID of the same person remains unchanged during training; assuming that the encoding length of the individual ID is L; the distribution of the labeled individual ID encoding Y and the predicted The cross entropy loss between is calculated as follows: Where Y(l) represents the value of the lth bit of the ground truth code of the individual ID; Represents the probability of the lth position predicted by the classifier.
4. The small target cluster tracking method based on memory enhancement according to claim 1 is characterized in that: The tracking of the test data set in step 5) specifically includes: constructing memory items for the trajectory, online updating of the memory module and constructing memory bias; Construct memory items for trajectories: Each trajectory creates a memory item consisting of M The memory module Mem composed of the components has an initial state of a random value; Online update of memory module: In the current t-th frame, the deepened memory feature of the detection target successfully associated with the k-th trajectory is Each feature vector is the deep memory feature matrix Det from the target t (k) is extracted along the channel to update the memory item; each feature vector and the mth memory item The product is fed into the softmax function to calculate the feature vector The weight Ws d , the formula is as follows: Among them, tr represents the transpose operation of the matrix, Ws d represents the dth eigenvector belongs to the appearance information stored in the mth memory item of the kth trajectory The possibility of; eigenvector Based on the corresponding weight Ws d The weighted average The calculation is as follows: Add the weighted average Δ to the mth memory item in the t-1th frame Fusion to form the memory item of the tth frame The calculation is as follows: By updating each memory item of the trajectory in the t-1th frame in the above way, the memory module Mem of the kth trajectory in the tth frame can be obtained. t ; Construct memory bias: In the t-1th frame, M memory items The memory module Mem that constitutes the kth trajectory t-1 ; Assume that a total of N targets are detected in the t-th frame; the deepened memory feature of the n-th detected target is expressed as From the target feature matrix Det along the channel t (n) Extract each feature vector Use the memory module Mem of the kth track t-1 Reconstruct the dth eigenvector process; each memory item Mem t-1 and the dth eigenvector The product is input into the softmax function to calculate the memory term The weight of Wo m , the formula is as follows: Among them, tr represents the transpose operation of the matrix, Wo m Represents the appearance details recorded in the mth memory item Exists in the dth eigenvector The probability of ; the dth eigenvector The reconstruction feature of is calculated as follows: The dth eigenvector The reconstruction vector is The nth detected target feature Det t (n) Each eigenvector is reconstructed as described above. and Cosine distance between The calculation is as follows: Cosine distance That is, the reconstructed memory module Mem t-1 And the deepened memory feature Det of the nth detection target t (n) deviation.
5. The small target cluster tracking method based on memory enhancement according to claim 1 is characterized in that: The step 6) specifically includes: starting from the second frame of the video, the small target cluster tracking method repeatedly associates the existing trajectory with the detected target in the current frame; assuming that the target detected in the current t-th frame is Det t , existing track Trk t-1 By Active Temporary and Sleep The trajectory set of three states; detection target Det t and the existing track Trk t-1 There are three stages of matching process required; The first stage of matching: Based on the IoU distance matrix, using Det t and the trajectory of frame t Trk t-1 Location construction distance matrix, Trk t-1 Bias reconstruction of memory module Det t Form the associated cost matrix; then, use the Hungarian algorithm to optimize the allocation; for the successfully matched trajectory Trk t-1 and detection Det t , perform the association success operation; unmatched detections and trajectories will be matched in the second stage; The second stage matching: the associated cost matrix is constructed based on the IoU distance; for the matched trajectory Trk t-1 And the detection target Det t Perform association successfully; for unmatched tracks Trk t-1 Failed to perform association operation; unmatched detection target Det t The third phase of matching will take place; Phase 3 Matching: Exploiting Sleep States via Memory Compensation Matching The memory module of the trajectory reconstructs the unassociated detection target Det t , and then use the reconstruction bias to construct the association cost matrix.
6. The small target cluster tracking method based on memory enhancement according to claim 1 is characterized in that: The step 7) specifically includes: associating the successfully detected target Det t Update the corresponding trajectory Trk t-1 The memory module and Kalman filter are used to update the trajectory of the Temporary state or the trajectory of the Sleep state that has been successfully associated for three times to the Active state; Detection target Det that failed to be associated t Updated to track Trk with status Temporary t ; Active state tracks that have failed to associate for 30 consecutive times will be set to Sleep state; Sleep state tracks that have failed to associate for 40 consecutive times will be deleted.