Cross-view multi-target tracking method and device, electronic equipment and storage medium
By extracting single-view and cross-view features in parallel, and using undirected graphs and minimal multi-slicing problem solving techniques, the problems of feature extraction, target association and trajectory generation in cross-view multi-objective tracking are solved, achieving high-precision and high-rootability tracking effects.
Patent Information
- Application Number
- CN202510464889.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-14
AI Technical Summary
The existing cross-view multi-objective tracking technology has many problems in feature extraction, target association and trajectory generation, and it is difficult to meet the needs of high precision, high robustness and real-time.
Single-view feature extraction network and cross-view feature extraction network are used to extract the features of the target in parallel, and combined with undirected graph and minimal multi-slicing problem solving technology, the target association and trajectory generation are optimized.
Improves the accuracy and robustness of tracking, reduces network complexity, speeds up processing speed, realizes viewing angle and motion alignment, and optimizes target association and trajectory generation.
Smart Images

Figure CN119991739A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the field of target tracking technology, and more specifically, to a cross-view multi-target tracking method, device, electronic device and storage medium. Background Art
[0002] In the field of computer vision and artificial intelligence, cross-view multi-object tracking technology is a key technology for many applications such as intelligent monitoring, traffic management, and robot navigation. Taking urban traffic monitoring as an example, real-time tracking of vehicles and pedestrians through multi-view cameras can effectively solve problems such as target occlusion and view limitation under a single view, improve the accuracy and reliability of the monitoring system, and provide strong support for intelligent traffic management. However, the cross-view multi-object tracking task faces many challenges, especially in terms of feature extraction, target association, and trajectory generation.
[0003] Feature extraction: Feature extraction is one of the key steps in cross-view multi-target tracking. Since the appearance and geometric relationship of targets under different viewpoints are significantly different, traditional feature extraction methods have difficulty adapting to such changes. For example, deep learning-based feature extraction networks are often affected by changes in viewpoints, resulting in large feature distance intervals and difficulty in setting a universal association threshold. In addition, existing methods often rely on stable camera viewpoints and homography matrices of specific datasets, which limits their applicability in open scenarios.
[0004] Target association: Target association is one of the core tasks of cross-view multi-target tracking. Existing methods usually use appearance feature-based matching or geometric constraints to complete the association task. However, in multi-view scenarios, the order of appearance, appearance features, and motion models of the targets will change with different viewpoints, which makes target association more complicated. For example, traditional matching methods based on the Hungarian algorithm have difficulty handling the spatiotemporal constraints and geometric relationships in cross-view target association, resulting in a decrease in association accuracy.
[0005] Trajectory generation: Trajectory generation is the ultimate goal of cross-view multi-target tracking. Existing methods usually regard target association and trajectory generation as independent steps, ignoring the interaction between cross-view information and single-view information. For example, although multi-target tracking methods based on graph models can use graph structures for optimization, how to effectively fuse multi-view information to generate stable trajectories in cross-view scenarios remains an unsolved problem. In addition, existing methods often lack the use of global information in the trajectory generation process, resulting in insufficient continuity and stability of the trajectory.
[0006] In summary, the existing cross-view multi-target tracking technology still has many problems in feature extraction, target association and trajectory generation, and it is difficult to meet the requirements of high precision, high robustness and real-time performance in practical applications. Therefore, developing a cross-view multi-target tracking method that can effectively solve the above problems has important research significance and application value. Summary of the invention
[0007] The present disclosure provides a cross-view multi-target tracking method, device, electronic device and storage medium, which are used to solve at least one of the above problems.
[0008] According to a first aspect of an embodiment of the present disclosure, a cross-view multi-target tracking method is provided, comprising: at each moment in the tracking process, performing the following steps: performing target detection processing on current frames of videos of different viewpoints respectively to obtain target information of each target of each viewpoint; extracting single-view features from the target information using a single-view feature extraction network; extracting cross-view features from the target information using a cross-view feature extraction network, wherein the cross-view feature extraction network is used to extract similar features for the same target of different viewpoints; constructing an undirected graph, wherein the nodes in the undirected graph include target nodes of each target of each viewpoint detected at the current moment, and trajectory nodes of each target trajectory currently determined, wherein the target trajectory is a motion path of the same target in a time series; according to the attributes of each node, segmenting the nodes corresponding to the same target in the undirected graph into the same subgraph, wherein the attributes include cross-view features and single-view features of each viewpoint involved in the corresponding node; and determining the target trajectory corresponding to the subgraph and the attributes of the target trajectory based on the attributes of each node in each subgraph.
[0009] Optionally, the undirected graph also includes an edge for connecting two nodes, wherein the nodes in the undirected graph corresponding to the same target are divided into the same subgraph according to the attributes of each node, including: for each edge in the undirected graph, determining the edge weight of the edge according to the attributes of the two nodes connected by the edge, wherein the edge weight is used to represent the possibility that the corresponding two nodes correspond to the same target; based on the edge weight of each edge in the undirected graph, dividing the undirected graph into multiple non-intersecting subgraphs by solving the minimum multicut problem of the undirected graph.
[0010] Optionally, the attributes also include position information, wherein determining the edge weight of the edge based on the attributes of the two nodes connected by the edge includes: determining the position distance of the two nodes based on the position information of the two nodes connected by the edge as the position weight of the edge; determining the feature similarity of the two nodes based on at least one of the single-view features and cross-view features of the two nodes connected by the edge as the feature weight of the edge; and determining the edge weight of the edge based on the position weight and feature weight of the edge.
[0011] Optionally, determining the edge weight of the edge based on the position weight and feature weight of the edge includes: when the position distance between the two nodes is less than a distance threshold, determining a weighted sum of the position weight and the feature weight of the edge as the edge weight of the edge; and when the position distance between the two nodes is greater than or equal to the distance threshold, determining a difference between the weighted sum and a preset penalty value as the edge weight of the edge.
[0012] Optionally, the method of determining the degree of feature similarity between the two nodes connected by the edge as the feature weight of the edge based on at least one of the single-perspective features and cross-perspective features of the two nodes includes: when the two nodes connected by the edge are target nodes from different perspectives, determining the degree of feature similarity between the two nodes as the feature weight of the edge based on the cross-perspective features of the two nodes; when the two nodes connected by the edge are a target node and a trajectory node respectively, taking the perspective to which the target node belongs as a reference perspective; for the reference perspective, determining the degree of feature similarity between the two nodes in the reference perspective based on the single-perspective features of the two nodes; for each perspective other than the reference perspective, determining the degree of feature similarity between the two nodes in each other perspective based on the cross-perspective features of the two nodes; and determining the feature weight of the edge based on the degree of feature similarity at each perspective.
[0013] Optionally, the single-view feature extraction network is trained by the following steps: obtaining image frames of a video of the same view at two preceding and succeeding moments, recording the image frame at the preceding moment as a preceding sample frame, and recording the image frame at the succeeding moment as a succeeding sample frame, wherein the preceding sample frame carries a true target label; performing target detection processing on the preceding sample frame to obtain target information of the target corresponding to the true target label as the preceding sample target information; using the single-view feature extraction network to be trained to extract the preceding sample single-view features from the preceding sample target information; performing target detection processing on the succeeding sample frame The invention relates to a method for extracting single-view features of the single-view feature extraction network to be trained, and obtaining target information of multiple candidate frames as the subsequent sample target information of each candidate frame; extracting subsequent sample single-view features from the subsequent sample target information of each candidate frame using the single-view feature extraction network to be trained; performing cross-correlation calculation on the previous sample single-view features and the subsequent sample single-view features of each candidate frame to obtain a time clue matrix; determining a loss value according to the time clue matrix and the true target label; and adjusting the parameters of the single-view feature extraction network to be trained according to the loss value to obtain the single-view feature extraction network.
[0014] According to a second aspect of an embodiment of the present disclosure, a cross-view multi-target tracking device is provided, comprising: at each moment in the tracking process, calling the following units: a detection unit, configured to perform target detection processing on current frames of videos of different viewpoints respectively, to obtain target information of each target of each viewpoint; a first extraction unit, configured to use a single-view feature extraction network to extract single-view features from the target information; a second extraction unit, configured to use a cross-view feature extraction network to extract cross-view features from the target information, wherein the cross-view feature extraction network is used to extract similar features for the same target of different viewpoints; a construction unit, The method comprises the following steps: a first step is to construct an undirected graph, wherein the nodes in the undirected graph include a target node for each target of each perspective detected at the current moment, and a trajectory node for each target trajectory currently determined, wherein the target trajectory is a motion path of the same target in a time series; a segmentation unit is configured to segment the nodes corresponding to the same target in the undirected graph into the same subgraph according to the attributes of each node, wherein the attributes include cross-perspective features and single-perspective features of each perspective involved in the corresponding node; and a determination unit is configured to determine the target trajectory corresponding to the subgraph and the attributes of the target trajectory based on the attributes of each node in each subgraph.
[0015] Optionally, the undirected graph also includes an edge for connecting two nodes, and the segmentation unit is further configured to: for each edge in the undirected graph, determine the edge weight of the edge according to the attributes of the two nodes connected by the edge, wherein the edge weight is used to represent the possibility that the corresponding two nodes correspond to the same target; based on the edge weight of each edge in the undirected graph, segment the undirected graph into multiple non-intersecting subgraphs by solving the minimum multi-cut problem of the undirected graph.
[0016] Optionally, the attributes also include position information, and the segmentation unit is further configured to: determine the position distance between the two nodes connected by the edge based on the position information of the two nodes as the position weight of the edge; determine the feature similarity of the two nodes based on at least one of the single-view features and cross-view features of the two nodes connected by the edge as the feature weight of the edge; determine the edge weight of the edge based on the position weight and feature weight of the edge.
[0017] Optionally, the segmentation unit is also configured to: when the position distance between the two nodes is less than a distance threshold, determine the weighted sum of the position weight and the feature weight of the edge as the edge weight of the edge; when the position distance between the two nodes is greater than or equal to the distance threshold, determine the difference between the weighted sum value and a preset penalty value as the edge weight of the edge.
[0018] Optionally, the segmentation unit is also configured to: when the two nodes connected by the edge are target nodes of different perspectives, determine the feature similarity of the two nodes according to the cross-perspective features of the two nodes as the feature weight of the edge; when the two nodes connected by the edge are respectively a target node and a trajectory node, take the perspective to which the target node belongs as a reference perspective; for the reference perspective, determine the feature similarity of the two nodes in the reference perspective according to the single-perspective features of the two nodes; for each perspective other than the reference perspective, determine the feature similarity of the two nodes in each other perspective according to the cross-perspective features of the two nodes; and determine the feature weight of the edge according to the feature similarity of each perspective.
[0019] Optionally, the single-view feature extraction network is trained by the following steps: obtaining image frames of a video of the same view at two preceding and succeeding moments, recording the image frame at the preceding moment as a preceding sample frame, and recording the image frame at the succeeding moment as a succeeding sample frame, wherein the preceding sample frame carries a true target label; performing target detection processing on the preceding sample frame to obtain target information of the target corresponding to the true target label as the preceding sample target information; using the single-view feature extraction network to be trained to extract the preceding sample single-view features from the preceding sample target information; performing target detection processing on the succeeding sample frame The invention relates to a method for extracting single-view features of the single-view feature extraction network to be trained, and obtaining target information of multiple candidate frames as the subsequent sample target information of each candidate frame; extracting subsequent sample single-view features from the subsequent sample target information of each candidate frame using the single-view feature extraction network to be trained; performing cross-correlation calculation on the previous sample single-view features and the subsequent sample single-view features of each candidate frame to obtain a time clue matrix; determining a loss value according to the time clue matrix and the true target label; and adjusting the parameters of the single-view feature extraction network to be trained according to the loss value to obtain the single-view feature extraction network.
[0020] According to a third aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: at least one processor; and at least one memory storing computer executable instructions, wherein the computer executable instructions, when executed by the at least one processor, prompt the at least one processor to execute a cross-view multi-target tracking method according to an exemplary embodiment of the present disclosure.
[0021] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by at least one processor, the at least one processor is prompted to execute a cross-view multi-target tracking method according to an exemplary embodiment of the present disclosure.
[0022] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, comprising computer instructions, which, when executed by at least one processor, prompt the at least one processor to execute a cross-view multi-target tracking method according to an exemplary embodiment of the present disclosure.
[0023] The technical solution provided by the embodiment of the present disclosure brings at least the following beneficial effects: According to the cross-view multi-target tracking method, device, electronic device and storage medium disclosed in the present disclosure, by optimizing feature extraction and target association strategy, the accuracy and robustness of tracking are improved, while the network complexity is reduced and the processing speed is accelerated. Specifically, by extracting the single-view feature and cross-view feature of the target in parallel, it is possible to reduce the feature distance interval of the same target under different view angles, and to extract rich features and improve the feature expression ability. In addition, the single-view feature extraction network and the cross-view feature extraction network are used to respectively realize their respective feature extraction, which reduces the number of network layers and the amount of parameters, helps to reduce the network complexity, and accelerates the training and testing speed of the network. In addition, by constructing the nodes of the undirected graph based on the target detected at the current moment and the target trajectory currently determined, and segmenting the undirected graph according to the attributes of the nodes, the nodes are clustered, and the corresponding relationship between each target under different view angles and the corresponding relationship between the target and the target trajectory can be synchronously determined, and the view alignment and motion alignment can be realized, thereby making full use of the interaction between the cross-view feature and the single-view feature, and synchronously realizing the target association and trajectory generation. At the same time, the target trajectory is gradually expanded over time, which can optimize target association and trajectory generation, and improve tracking accuracy and robustness.
[0024] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute improper limitations on the present disclosure.
[0026] Figure 1 is a flowchart of a cross-view multi-target tracking method according to an exemplary embodiment of the present disclosure.
[0027] Figure 2 is a structural diagram of a single-view feature extraction network or a cross-view feature extraction network according to an exemplary embodiment of the present disclosure.
[0028] Figure 3 is a schematic diagram of partitioning an undirected graph according to an exemplary embodiment of the present disclosure.
[0029] Figure 4 It is a schematic flow chart of a cross-view multi-target tracking method according to a specific embodiment of the present disclosure.
[0030] Figure 5 is a block diagram of a cross-view multi-target tracking apparatus according to an exemplary embodiment of the present disclosure.
[0031] Figure 6 is a block diagram of an electronic device according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION
[0032] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings.
[0033] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The implementation methods described in the following examples do not represent all implementation methods consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the attached claims.
[0034] It should be noted that the phrase "at least one of the items" in the present disclosure includes three types of parallel situations: "any one of the items", "a combination of any number of the items", and "all of the items". For example, "including at least one of A and B" includes the following three types of parallel situations: (1) including A; (2) including B; (3) including A and B. Another example is "executing at least one of step 1 and step 2" which means the following three types of parallel situations: (1) executing step 1; (2) executing step 2; (3) executing step 1 and step 2.
[0035] Hereinafter, a cross-view multi-target tracking method, apparatus, electronic device, and storage medium according to exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings.
[0036] Figure 1 is a flowchart of a cross-view multi-target tracking method according to an exemplary embodiment of the present disclosure. The method can be executed on an electronic device with sufficient computing power, and refers to Figure 1 The entire process can be executed cyclically during the tracking process, specifically, the entire process is executed once at each execution moment, and the information of the target trajectory obtained by the previous execution can be continuously accumulated for use at subsequent moments. It should be understood that Figure 1 If the drawing is not finished, the tracking can be ended in response to the end instruction.
[0037] Reference Figure 1 In step S101, target detection processing is performed on the current frames of videos of different viewing angles respectively to obtain target information of each target of each viewing angle.
[0038] As an example, the videoCapture function of opencv can be used to extract the image frame information of the video to obtain the current frame of each perspective. Then, target detection is performed on each image frame obtained, for example, using the open source YOLOX-X (You Only Look Once eXtreme) target detection algorithm to extract the target. This method can generate multi-scale high-dimensional features with dimensions of H×W×C, where H and W represent the height and width of the image frame respectively, and C represents the feature size of the high-dimensional feature, for example, 512. The obtained high-dimensional features are passed through the head network of the target detection model to obtain the target's position information (x, y, w, h), classification prediction, and confidence prediction. The target information obtained in step S101 can specifically be the high-dimensional feature, that is, the intermediate feature obtained during the target detection process.
[0039] In step S102, a single-view feature extraction network is used to extract single-view features from target information.
[0040] This step models the target motion features under a single view through a single view feature extraction network.
[0041] Optionally, the single-view feature extraction network is trained by the following steps: obtaining image frames of a video of the same view at two preceding and succeeding moments, recording the image frame at the preceding moment as a prior sample frame, and recording the image frame at the succeeding moment as a subsequent sample frame, wherein the prior sample frame carries a true target label; performing target detection processing on the prior sample frame to obtain target information of the target corresponding to the true target label as the prior sample target information; using the single-view feature extraction network to be trained, extracting prior sample single-view features from the prior sample target information; performing target detection processing on the subsequent sample frame to obtain target information of multiple candidate frames as the subsequent sample target information of each candidate frame; using the single-view feature extraction network to be trained, extracting subsequent sample single-view features from the subsequent sample target information of each candidate frame; performing cross-correlation calculation on the prior sample single-view features and the subsequent sample single-view features of each candidate frame to obtain a temporal cue matrix; determining a loss value based on the temporal cue matrix and the true target label; adjusting the parameters of the single-view feature extraction network to be trained based on the loss value to obtain a single-view feature extraction network. By calculating the time cue matrix, the features of the image frames at the previous and next moments can be fused, and the single-view feature extraction network can be trained based on this, which enables the network to learn how to capture the dynamic information of the target, making it easier to effectively associate the target with its trajectory in target tracking.
[0042] As an example, the structure of a single-view feature extraction network is as follows: Figure 2As shown, a deep learning network structure based on a 2D convolution (Conv2d)-normalization (Batch Normalization, referred to as BN)-activation function (such as a rectified linear unit, referred to as ReLU) structure can be adopted, and specifically can include two of the deep learning network structures and a 2D convolution structure.
[0043] As an example, during the training process, target detection processing can be performed on the previous sample frame, and several candidate frames and their corresponding high-dimensional features can be obtained through the detector. According to the candidate frame position of the real target label carried by the previous sample frame, the detected candidate frame corresponding to the real target and its corresponding high-dimensional features of size 1×C are obtained as the previous sample target information. After passing through the single-view feature extraction network to be trained, the previous sample single-view features of size 1×C' can be obtained. For the subsequent sample frame, the anchor-free detector can be used to output a high-dimensional feature matrix of size H×W×C as the subsequent sample target information of multiple candidate frames as a whole. Each element in the matrix represents a candidate frame, and the feature size is C. After the high-dimensional feature matrix is processed by the single-view feature extraction network, a subsequent sample single-view feature matrix of size H×W×C' can be obtained. Since the dimension of the single-view feature of the prior sample is 1×C', the cross-correlation calculation operation in a single target is reduced to matrix multiplication, and the matrix multiplication is performed by the transpose of the single-view feature matrix of the subsequent sample and the single-view feature of the prior sample to obtain a temporal cue matrix of size H×W×1.
[0044] As an example, regarding the loss function in training, the Logistic-MSE loss function can be used, and the formula is as follows.
[0045]
[0046] Since the detector will obtain feature maps of multiple scales, the loss function needs to be calculated for each feature map. In the above formula, Represents the total number of candidate boxes detected in the subsequent sample frame, represents the calculated time cue matrix, represents the Gaussian function matrix generated using the true target label, Represents the coordinate point in the image frame, represents the t-th frame image, The response map represents the true target label, Represents the coordinate point in the calculated time clue matrix Response map of the true target label Generated by Gaussian function, the formula is as follows.
[0047]
[0048] in Represents the standard deviation of the Gaussian function, which can be selected based on experience, such as 0.75. , ) represents the true value position of the target, specifically the true center position of the target, which is recorded in the true target label. This formula indicates that at the true center position of the target, the response value is 1, and the farther away from the true center position other positions are, the lower their response values are.
[0049] Return to reference Figure 1 ,In step S103, a cross-view feature extraction network is used to extract cross-view features from the target information.
[0050] The cross-view feature extraction network is used to extract similar features for the same target at different viewpoints, which can reduce the feature distance interval of the same target at different viewpoints.
[0051] Optionally, the cross-view feature extraction network is trained by the following steps: obtaining image frames of videos of different viewpoints at the same time as sample frames; performing target detection processing on each sample frame to obtain target information of each target in each sample frame as sample target information; using the cross-view feature extraction network to be trained to extract sample cross-view features from each sample target information; determining the value of the cross-view consistency loss function according to the extracted sample cross-view features as the loss value; adjusting the parameters of the cross-view feature extraction network to be trained according to the loss value to obtain the cross-view feature extraction network. By introducing the cross-view consistency loss function to train the cross-view feature extraction network, the similarity of features of the same target under different viewpoints can be enhanced, while the similarity of features between different targets can be suppressed, thereby reducing the distance interval of cross-view features extracted by the network under different viewpoints.
[0052] As an example, the structure of the cross-view feature extraction network can also be Figure 2 As shown, no further details are given here.
[0053] As an example, the formula of the cross-view consistency loss function is as follows.
[0054]
[0055] In the above formula, N and M are the image frames of view 1 (hereinafter referred to as view Figure 1 ) and image frames of view 2 (hereinafter referred to as view Figure 2 ) is the number of objects detected in the image. is an indicator function. Figure 1 The i-th target and view Figure 2The value is 1 if the jth target in represents the same object, otherwise it is 0. Respectively represent the Figure 1 The cross-view features of the i-th target, Figure 2 The cross-view features of the jth target. D represents the distance function, which is the cosine distance here. It should be noted that the cross-view consistency loss function is calculated based on the view pair (i.e., two views). If the number of views is greater than 2, multiple view pairs can be determined, and the cross-view consistency loss is calculated for each view pair. Then, multiple cross-view consistency losses are fused, for example, including but not limited to calculating the sum value, weighted sum, etc., as the final cross-view consistency loss.
[0056] According to the cross-view multi-target tracking method of the exemplary embodiment of the present disclosure, by extracting the single-view features and cross-view features of the target in parallel, it is possible to reduce the feature distance interval of the same target under different viewpoints, extract rich features, and improve the feature expression ability. In addition, the single-view feature extraction network and the cross-view feature extraction network are used to respectively realize their respective feature extraction, which reduces the number of network layers and the amount of parameters, helps to reduce the complexity of the network, and speeds up the training and testing speed of the network.
[0057] Return to reference Figure 1 , in step S104, an undirected graph is constructed.
[0058] The nodes in the undirected graph include the target node of each target detected at each perspective at the current moment, and the trajectory node of each target trajectory that has been determined. The target node represents the independent target detected at the current moment, and the target trajectory is the motion path of the same target in the time series. It should be noted that at the initial moment, no target trajectory has been determined, so the constructed undirected graph only contains target nodes, not trajectory nodes. At the second moment, the target detected at the initial moment can be used as a target trajectory.
[0059] In step S105 , nodes corresponding to the same target in the undirected graph are divided into the same subgraph according to the attribute of each node.
[0060] The attributes include the cross-view features and single-view features of each view involved in the corresponding node. Specifically, since the target node has a clear view, it only involves the cross-view features and single-view features under this view. Since the target trajectory of the trajectory node is obtained based on the target of different viewpoints, it will involve multiple different viewpoints, and each viewpoint contains cross-view features and single-view features.
[0061] When segmenting an undirected graph, each target node corresponds to a clear target, and each target trajectory also corresponds to a clear target. Therefore, the nodes in the undirected graph can be segmented by predicting whether two nodes correspond to the same target.
[0062] In step S106, based on the attributes of each node in each subgraph, the target trajectory corresponding to the subgraph and the attributes of the target trajectory are determined.
[0063] Specifically, each subgraph contains at most one trajectory node and at most one target node under each viewing angle. For each subgraph, if the subgraph contains only trajectory nodes, it means that the target trajectory has not found the target in the target association at the current moment, so its trajectory state can be set to lost, and the target trajectory can be discarded after a certain disappearance time. If the subgraph contains only target nodes, it means that the target is a newly appeared target, and it can be determined whether the target can be used as a new target trajectory based on the confidence predicted during the target detection process. That is, if the confidence of the target is large enough, for example, greater than a preset confidence threshold, the target is used as a new target trajectory, and the attributes of the target under each viewing angle are used as the attributes of the new target trajectory, otherwise the target is discarded. If the subgraph contains both target nodes and trajectory nodes, it means that the tracked target trajectory can be updated, and the information of the target node is added on the basis of the trajectory node as the updated target trajectory. In the case of updating the target trajectory, the properties of the target trajectory also need to be updated. At this time, for each perspective, the cross-perspective features of the target node under that perspective can be fused with the cross-perspective features of the trajectory node at that perspective, and the single-perspective features of the target node under that perspective can be fused with the single-perspective features of the trajectory node at that perspective, for example, by using a linear fusion method, which is not limited in the present disclosure. Figure 3 A specific example of partitioning an undirected graph is shown. Figure 3 In , the hollow circles represent target nodes, and different line types (i.e., thick solid lines and thin solid lines) are used to distinguish different viewpoints, i.e., hollow circles of the same line type represent targets detected under the same viewpoint, e.g. Figure 3 The three target nodes obj1, obj2, and obj3 under the same perspective are marked in the figure; the solid circles represent trajectory nodes, which are marked as Tra1, Tra2, Tra3, and Tra4 respectively. Among them, trajectory node Tra4 (in order to distinguish it from other trajectory nodes, Figure 3 Tra4 is colored black in the figure. After executing the algorithm, it is not connected to other nodes, indicating that there is no target associated with it in this frame. The dotted edge connects two trajectory nodes. Based on the introduction of edge weights in the previous article, the edge weights between trajectory nodes will introduce a preset penalty value, resulting in its edge weight being negative infinity. Therefore, the edge will be deleted after executing the algorithm, so it is represented by a dotted line before executing the algorithm.
[0064] According to the cross-view multi-target tracking method of the exemplary embodiment of the present disclosure, by constructing nodes of an undirected graph based on the target detected at the current moment and the target trajectory that has been determined so far, and segmenting the undirected graph according to the attributes of the nodes, the nodes are clustered, and the correspondence between each target under different viewpoints and the correspondence between the target and the target trajectory can be determined synchronously, and the viewpoint alignment and motion alignment can be achieved, thereby making full use of the interaction between cross-view features and single-view features, and synchronously achieving target association and trajectory generation. At the same time, the target trajectory is gradually expanded as time goes by, which can optimize the target association and trajectory generation, and improve the tracking accuracy and robustness.
[0065] It should be noted that the order of the steps in the various processes introduced in this disclosure is only for the convenience of distinguishing different steps, and is not used to limit the execution order of the steps. The execution order of different steps can be adjusted if it is logically reasonable. Figure 1 Step S102 and step S103 in the embodiment may be executed successively or in parallel in any order, and the present disclosure does not impose any limitation on this.
[0066] Next, step S105 is further introduced.
[0067] Optionally, the undirected graph also includes an edge for connecting two nodes, and step S105 includes: for each edge in the undirected graph, determining the edge weight of the edge according to the attributes of the two nodes connected by the edge, wherein the edge weight is used to represent the possibility that the corresponding two nodes correspond to the same target; based on the edge weight of each edge in the undirected graph, the undirected graph is divided into multiple non-intersecting subgraphs by solving the minimum multi-cut problem of the undirected graph. By modeling the target association task as an optimized solution of the minimum multi-cut problem in the tracking stage, a stable and accurate target trajectory can be generated, providing a more accurate and robust solution for cross-view target tracking. Specifically, the minimum multi-cut problem is to divide the undirected graph into multiple non-intersecting subgraphs by cutting the edges in the undirected graph, and the goal of cutting is to find a cutting scheme that minimizes the sum of the edge weights of the cut edges.
[0068] The minimum multi-cut problem can be expressed using the following formulas.
[0069]
[0070]
[0071]
[0072]
[0073] In the above formulas, G represents an undirected graph, V represents a node in the undirected graph, E represents an edge in the undirected graph, and w is the edge weight. The lifted_multicut function of the nifty library is used to implement this. is the edge that needs to be deleted. The edge to be deleted The following two formulas indicate that for any circle in an undirected graph , assuming If it is 1 (meaning it needs to be deleted), then in addition to this edge, at least one more edge should be deleted from the circle. This restriction is to ensure that when an edge is deleted, the nodes at both ends of the edge are no longer connected by other edges, so as to achieve the clustering effect.
[0074] Further optionally, the attribute also includes position information, which can be obtained based on the position information obtained during the aforementioned target detection process. In this regard, when updating the attributes of the target trajectory in step S106, for example, the predicted position and the observed position under the same viewing angle can be passed through a Kalman filter to update the position information. At this time, the operation of determining the edge weight of the edge according to the attributes of the two nodes connected by the edge in step S105 can include the following three steps.
[0075] The first step is to determine the position distance of the two nodes based on the position information of the two nodes connected by the edge, which is used as the position weight of the edge.
[0076] Specifically, the position information is used to describe the spatial position of the target in the image, which usually includes the coordinates of the candidate box or the projection position of the target in the plane view. When executing the first step, the position weight can be measured by calculating the proximity of the two nodes in space.
[0077] As an example, for targets within the same perspective, the position information may include the center point coordinates and size of the candidate box; for targets across perspectives, the position information can be used to project the target into a unified planar view coordinate system through a homography matrix and calculate the center point distance.
[0078] In the second step, based on at least one of the single-view features and cross-view features of the two nodes connected by the edge, the feature similarity of the two nodes is determined as the feature weight of the edge.
[0079] Specifically, single-view features and cross-view features are used to characterize the appearance characteristics of the target and are obtained by a single-view feature extraction network and a cross-view feature extraction network.
[0080] Optionally, the second step includes: when the two nodes connected by the edge are target nodes of different perspectives, according to the cross-perspective features of the two nodes, determine the feature similarity of the two nodes as the feature weight of the edge; when the two nodes connected by the edge are the target node and the trajectory node respectively, take the perspective to which the target node belongs as the reference perspective; for the reference perspective, according to the single-perspective features of the two nodes, determine the feature similarity of the two nodes in the reference perspective; for each perspective other than the reference perspective, according to the cross-perspective features of the two nodes, determine the feature similarity of the two nodes in each other perspective; according to the feature similarity of each perspective, determine the feature weight of the edge. By reasonably selecting single-perspective features or multi-perspective features for different node types and specific perspectives to calculate the similarity between the feature vectors of the two nodes, and measuring the feature weight accordingly, a reliable feature weight can be obtained. As an example, cosine similarity can be used as a measure of feature similarity.
[0081] The third step is to determine the edge weight based on the edge’s position weight and feature weight.
[0082] By combining the position weight and feature weight to determine the edge weight, the calculation of the edge weight can comprehensively consider the spatial position and appearance characteristics of the target, thereby comprehensively evaluating the similarity and correlation between nodes, which helps to optimize target association and trajectory generation.
[0083] Optionally, the third step includes: when the position distance between the two nodes is less than the distance threshold, determining the weighted sum of the position weight and the feature weight of the edge as the edge weight of the edge; when the position distance between the two nodes is greater than or equal to the distance threshold, determining the difference between the weighted sum and the preset penalty value as the edge weight of the edge. In the case where the position distance between the two nodes is large, the two nodes are very likely to correspond to different targets. By further subtracting the preset penalty value to correct the edge weight in this case, the edge between the two nodes can be preferentially cut when solving the minimum multi-cut problem, thereby improving the segmentation quality and efficiency.
[0084] As an example, the edge weight is calculated by weighted summing the feature weight and the position weight as follows.
[0085]
[0086] In the above formula, represents the edge weight, i and j represent the two nodes connected by the calculated edge respectively, Represents the weight used to control the influence of feature weight and position weight. The feat function uses cosine distance to calculate the distance between the feature vectors of two nodes. The IoU function is used to calculate the position weight. Penalty is a preset penalty value, which is set when the distance between the estimated position center point on the ground plane exceeds the distance threshold. The distance threshold can be set to the square root of the area of the candidate box projected to the ground perspective multiplied by 2, which is a rough empirical value. Of course, it can also be set to other values; when the distance threshold is exceeded, it is considered that the two candidate boxes point to different targets, and the penalty can be set to positive infinity.
[0087] Figure 4 It is a schematic flow chart of a cross-view multi-target tracking method according to a specific embodiment of the present disclosure.
[0088] In this specific embodiment, the hardware and programming language for the specific operation of the method are not limited, and the method can be written in any language. This specific embodiment uses a GPU with a 3.2 GHz CPU and 16 GB memory and at least one 24 GB video memory, and uses the python 3.7 programming language and the pytorch 1.9.0 version deep learning framework to implement the method.
[0089] Reference Figure 4 , the cross-view multi-target tracking method of this specific embodiment mainly includes four steps.
[0090] Step 1: Extract the feature information and location information of the target from the current frame of the multi-view video.
[0091] Step 2: Model the single-view features through the single-view feature extraction network and capture the dynamic information of the target through the response matrix.
[0092] Step 3: Introduce cross-view consistency loss through the cross-view feature extraction network to reduce the feature distance interval under different viewpoints and improve the multi-view target association performance.
[0093] Step 4: Convert the target association phase into a minimum multi-cut problem for solving an undirected graph and model the tracking task as a global optimization problem.
[0094] Figure 5 is a block diagram of a cross-view multi-target tracking apparatus according to an exemplary embodiment of the present disclosure. Figure 5 The cross-view multi-target tracking device 500 includes a detection unit 501, a first extraction unit 502, a second extraction unit 503, a construction unit 504, a segmentation unit 505, and a determination unit 506. It should be understood that Figure 1 The various steps in the tracking process can be executed cyclically. Similarly, the various units in the cross-view multi-target tracking device 500 can also be called cyclically during the tracking process.
[0095] The detection unit 501 may perform target detection processing on current frames of videos of different viewing angles respectively, and obtain target information of each target of each viewing angle.
[0096] The first extraction unit 502 may use a single-view feature extraction network to extract single-view features from the target information.
[0097] The second extraction unit 503 may use a cross-view feature extraction network to extract cross-view features from the target information, wherein the cross-view feature extraction network is used to extract similar features for the same target at different views.
[0098] The construction unit 504 can construct an undirected graph, wherein the nodes in the undirected graph include a target node for each target detected at each perspective at the current moment, and a trajectory node for each target trajectory currently determined, where the target trajectory is a motion path of the same target in a time series.
[0099] The segmentation unit 505 may segment the nodes corresponding to the same target in the undirected graph into the same subgraph according to the attributes of each node, wherein the attributes include cross-view features and single-view features of each view involved in the corresponding node.
[0100] The determination unit 506 may determine the target trajectory corresponding to the subgraph and the attributes of the target trajectory based on the attributes of each node in each subgraph.
[0101] Optionally, the undirected graph also includes an edge for connecting two nodes, and the segmentation unit 505 can also: determine the edge weight of each edge in the undirected graph according to the attributes of the two nodes connected by the edge, wherein the edge weight is used to represent the possibility that the corresponding two nodes correspond to the same target; based on the edge weight of each edge in the undirected graph, the undirected graph is segmented into multiple non-intersecting subgraphs by solving the minimum multi-cut problem of the undirected graph.
[0102] Optionally, the attributes also include position information, and the segmentation unit 505 can also: determine the position distance of the two nodes connected by the edge based on the position information of the two nodes as the position weight of the edge; determine the feature similarity of the two nodes based on at least one of the single-view features and cross-view features of the two nodes connected by the edge as the feature weight of the edge; determine the edge weight of the edge based on the position weight and feature weight of the edge.
[0103] Optionally, the segmentation unit 505 may also: when the position distance between two nodes is less than a distance threshold, determine the weighted sum of the position weight and the feature weight of the edge as the edge weight of the edge; when the position distance between two nodes is greater than or equal to the distance threshold, determine the difference between the weighted sum and a preset penalty value as the edge weight of the edge.
[0104] Optionally, the segmentation unit 505 may also: when the two nodes connected by the edge are target nodes of different perspectives, determine the feature similarity of the two nodes according to the cross-perspective features of the two nodes as the feature weight of the edge; when the two nodes connected by the edge are a target node and a trajectory node respectively, take the perspective to which the target node belongs as the reference perspective; for the reference perspective, determine the feature similarity of the two nodes in the reference perspective according to the single-perspective features of the two nodes; for each perspective other than the reference perspective, determine the feature similarity of the two nodes in each other perspective according to the cross-perspective features of the two nodes; and determine the feature weight of the edge according to the feature similarity of each perspective.
[0105] Optionally, the single-view feature extraction network is trained by the following steps: obtaining image frames of a video of the same view at two preceding and succeeding moments, recording the image frame at the preceding moment as a prior sample frame, and recording the image frame at the succeeding moment as a subsequent sample frame, wherein the prior sample frame carries a true target label; performing target detection processing on the prior sample frame to obtain target information of the target corresponding to the true target label as the prior sample target information; using the single-view feature extraction network to be trained, extracting prior sample single-view features from the prior sample target information; performing target detection processing on the subsequent sample frame to obtain target information of multiple candidate frames as the subsequent sample target information of each candidate frame; using the single-view feature extraction network to be trained, extracting subsequent sample single-view features from the subsequent sample target information of each candidate frame; performing cross-correlation calculation on the prior sample single-view features and the subsequent sample single-view features of each candidate frame to obtain a temporal cue matrix; determining a loss value based on the temporal cue matrix and the true target label; adjusting the parameters of the single-view feature extraction network to be trained based on the loss value to obtain a single-view feature extraction network.
[0106] Regarding the device in the above embodiment, the specific manner in which each unit performs the operation has been described in detail in the embodiment of the method, and will not be elaborated here.
[0107] Figure 6 A structural block diagram of an electronic device 600 according to an exemplary embodiment of the present disclosure is shown.
[0108] Reference Figure 6 The electronic device 600 includes: at least one memory 601 and at least one processor 602, wherein the at least one memory 601 stores computer executable instructions, and when the computer executable instructions are executed by the at least one processor 602, the at least one processor is prompted to execute the cross-view multi-target tracking method as described in the above exemplary embodiment.
[0109] As an example, the electronic device 600 may be a PC, a tablet device, a personal digital assistant, a smart phone, or other device capable of executing the above instruction set. Here, the electronic device 600 is not necessarily a single electronic device 600, but may also be any device or circuit collection capable of executing the above instruction (or instruction set) individually or in combination. The electronic device 600 may also be part of an integrated control system or system manager, or may be configured as a portable electronic device 600 that is interconnected with a local or remote (e.g., via wireless transmission) interface.
[0110] In the electronic device 600, the processor 602 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller or a microprocessor. As an example and not limitation, the processor 602 may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.
[0111] The processor 602 may execute instructions or codes stored in the memory 601, which may also store data. Instructions and data may also be sent and received over a network via a network interface device, which may employ any known transmission protocol.
[0112] The memory 601 may be integrated with the processor 602, for example, by placing RAM or flash memory within an integrated circuit microprocessor or the like. In addition, the memory 601 may include a separate device, such as an external disk drive, a storage array, or any other storage device that can be used by a database system. The memory 601 and the processor 602 may be operatively coupled, or may communicate with each other, such as through an I / O port, a network connection, etc., so that the processor 602 can read files stored in the memory.
[0113] In addition, the electronic device 600 may further include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.) All components of the electronic device 600 may be connected to each other via a bus and / or a network.
[0114] According to an exemplary embodiment of the present disclosure, a computer-readable storage medium storing instructions may also be provided, wherein the instructions, when executed by at least one processor, cause the at least one processor to execute the cross-view multi-target tracking method as described in the above exemplary embodiment. Examples of computer-readable storage media here include: read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disk storage, hard disk drive (HDD), solid state drive (SSD), card storage (such as, multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk and any other device, any other device is configured to store computer programs and any associated data, data files and data structures in a non-transitory manner and provide the computer programs and any associated data, data files and data structures to a processor or computer so that the processor or computer can execute the computer program. The computer program in the above-mentioned computer-readable storage medium can be run in an environment deployed in a computer device such as a client, a host, an agent device, a server, etc. In addition, in one example, the computer program and any associated data, data files and data structures are distributed on a networked computer system, so that the computer program and any associated data, data files and data structures are stored, accessed and executed in a distributed manner by one or more processors or computers.
[0115] According to an exemplary embodiment of the present disclosure, a computer program product may also be provided, including computer instructions, which, when executed by at least one processor, execute the cross-view multi-target tracking method as described in the above exemplary embodiment.
[0116] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The description and examples are intended to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the appended claims.
[0117] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A cross-view multi-target tracking method, characterized in that: include: At each point in the tracking process, perform the following steps: Performing target detection processing on the current frames of the videos from different perspectives respectively, and obtaining target information of each target in each perspective; Using a single-view feature extraction network, extracting single-view features from the target information; Using a cross-view feature extraction network to extract cross-view features from the target information, wherein the cross-view feature extraction network is used to extract similar features for the same target at different viewpoints; Constructing an undirected graph, wherein the nodes in the undirected graph include a target node of each target of each perspective detected at the current moment, and a trajectory node of each target trajectory currently determined, wherein the target trajectory is a motion path of the same target in a time series; According to the attributes of each node, nodes corresponding to the same target in the undirected graph are divided into the same subgraph, wherein the attributes include cross-view features and single-view features of each view involved in the corresponding node; Based on the attributes of each node in each subgraph, a target trajectory corresponding to the subgraph and the attributes of the target trajectory are determined.
2. The cross-view multi-target tracking method according to claim 1, characterized in that: The undirected graph further includes an edge for connecting two nodes, wherein the nodes corresponding to the same target in the undirected graph are divided into the same subgraph according to the attribute of each node, including: For each edge in the undirected graph, determine an edge weight of the edge according to the attributes of two nodes connected by the edge, wherein the edge weight is used to represent the possibility that the corresponding two nodes correspond to the same target; Based on the edge weight of each edge in the undirected graph, the undirected graph is divided into a plurality of non-intersecting subgraphs by solving a minimum multicut problem of the undirected graph.
3. The cross-view multi-target tracking method according to claim 2, characterized in that: The attribute further includes location information, wherein determining the edge weight of the edge according to the attributes of the two nodes connected by the edge includes: Determine, according to the position information of two nodes connected by the edge, the position distance between the two nodes as the position weight of the edge; Determine, according to at least one of the single-view feature and the cross-view feature of two nodes connected by the edge, a feature similarity degree of the two nodes as a feature weight of the edge; An edge weight of the edge is determined according to the position weight and the feature weight of the edge.
4. The cross-view multi-target tracking method according to claim 3, characterized in that: The step of determining the edge weight of the edge according to the position weight and the feature weight of the edge includes: When the position distance between the two nodes is less than a distance threshold, determining a weighted sum of the position weight and the feature weight of the edge as the edge weight of the edge; When the position distance between the two nodes is greater than or equal to a distance threshold, a difference between the weighted sum value and a preset penalty value is determined as the edge weight of the edge.
5. The cross-view multi-target tracking method according to claim 3, characterized in that: The determining, based on at least one of the single-view feature and the cross-view feature of the two nodes connected by the edge, the feature similarity of the two nodes as the feature weight of the edge comprises: When two nodes connected by the edge are target nodes from different perspectives, determining the feature similarity of the two nodes according to the cross-perspective features of the two nodes as the feature weight of the edge; When the two nodes connected by the edge are respectively a target node and a trajectory node, the perspective to which the target node belongs is taken as a reference perspective; for the reference perspective, the feature similarity of the two nodes in the reference perspective is determined based on the single-perspective features of the two nodes; for each perspective other than the reference perspective, the feature similarity of the two nodes in each other perspective is determined based on the cross-perspective features of the two nodes; and the feature weight of the edge is determined based on the feature similarity of each perspective.
6. The cross-view multi-target tracking method according to any one of claims 1 to 5, characterized in that: The single-view feature extraction network is trained by the following steps: Obtain image frames of a video of the same viewing angle at two previous and next moments, record the image frame at the previous moment as a previous sample frame, and record the image frame at the next moment as a next sample frame, wherein the previous sample frame carries a true target label; Performing target detection processing on the prior sample frame to obtain target information of the target corresponding to the real target label as the prior sample target information; Using a single-view feature extraction network to be trained, extracting a single-view feature of a prior sample from the prior sample target information; Performing target detection processing on the subsequent sample frame to obtain target information of multiple candidate frames as subsequent sample target information of each candidate frame; Using the single-view feature extraction network to be trained, extracting subsequent sample single-view features from subsequent sample target information of each candidate box; Performing cross-correlation calculation on the single-view features of the previous sample and the single-view features of the subsequent sample of each candidate frame to obtain a temporal cue matrix; Determining a loss value according to the temporal cue matrix and the true target label; According to the loss value, the parameters of the single-view feature extraction network to be trained are adjusted to obtain the single-view feature extraction network.
7. A cross-view multi-target tracking device, characterized in that: include: At each moment in the tracking process, the following units are called: The detection unit is configured to perform target detection processing on current frames of videos of different viewing angles respectively, and obtain target information of each target in each viewing angle; A first extraction unit is configured to extract a single-view feature from the target information using a single-view feature extraction network; A second extraction unit is configured to extract cross-view features from the target information using a cross-view feature extraction network, wherein the cross-view feature extraction network is used to extract similar features for the same target at different views; A construction unit is configured to construct an undirected graph, wherein the nodes in the undirected graph include a target node of each target of each perspective detected at the current moment, and a trajectory node of each target trajectory currently determined, wherein the target trajectory is a motion path of the same target in a time series; A segmentation unit is configured to segment the nodes corresponding to the same target in the undirected graph into the same subgraph according to the attributes of each node, wherein the attributes include cross-view features and single-view features of each view involved in the corresponding node; The determination unit is configured to determine the target trajectory corresponding to the subgraph and the attributes of the target trajectory based on the attributes of each node in each subgraph.
8. An electronic device, characterized in that: include: at least one processor; at least one memory storing computer executable instructions, Wherein, when the computer executable instructions are executed by the at least one processor, the at least one processor is prompted to execute the cross-view multi-target tracking method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: When the instructions in the computer-readable storage medium are executed by at least one processor, the at least one processor is prompted to perform the cross-view multi-target tracking method as described in any one of claims 1 to 6.
10. A computer program product comprising computer instructions, characterized in that: When the computer instructions are executed by at least one processor, the at least one processor is prompted to perform the cross-view multi-target tracking method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Semi-online visual multi-target tracking method based on wavelet graph correlation model
CN108765459A
Online multi-target tracking algorithm fusing single-target tracking result
CN110363791A
Multi-target cross-mirror tracking method and device based on graph matching, equipment and medium
CN112131904A
Multi-view multi-target tracking method and device, computer equipment and storage medium
CN114067428A
Target tracking method, computer program product, storage medium and electronic equipment
CN114387304A
Cited By
Multi-view crowd tracking method, device and equipment with view-ground interaction and medium
CN121937488A