Appearance-motion collaborative modeling multi-target tracking method, system and program product
By constructing a joint appearance-motion matrix through appearance-motion co-modeling, the stability and accuracy problems caused by independent modeling of appearance and motion information in multi-target tracking are solved, and stable tracking in complex environments is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INST OF OPTICS & ELECTRONICS CHINESE ACAD OF SCI
- Filing Date
- 2026-02-05
- Publication Date
- 2026-05-26
AI Technical Summary
Existing multi-target tracking methods model appearance and motion information independently, failing to depict the spatiotemporal relationship between the two. This results in insufficient tracking stability and accuracy in complex environments, and is prone to problems such as identity switching or trajectory loss.
The appearance-motion co-modeling method is adopted. By acquiring video frame sequences, target location detection and appearance feature extraction are performed respectively. An appearance-motion joint matrix is constructed, which integrates appearance similarity and motion consistency to associate and match the current target with historical trajectories.
It effectively improves the performance of multi-target tracking in complex environments, solves the matching ambiguity caused by target offset between adjacent frames and similar appearance features, and achieves stable multi-target tracking.
Smart Images

Figure CN122089778A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision and multi-target tracking technology, and in particular to a multi-target tracking method, system and program product based on appearance-motion cooperative modeling. Background Technology
[0002] Multi-target tracking, a crucial research area in computer vision, plays a vital role in applications such as intelligent surveillance, drone patrols, autonomous driving assistance, and behavior analysis. Its core objective is to continuously detect and locate multiple moving targets in a video sequence while maintaining consistency in their identities across frames. With the continuous expansion of camera deployment scenarios, video data in real-world applications exhibits more complex dynamic characteristics, posing a greater challenge to the stability and accuracy of multi-target tracking. In most existing tracking frameworks, appearance and motion information are typically modeled independently, generating separate appearance cost matrices and motion cost matrices, which are then simply fused. However, this parallel rather than coupled modeling approach cannot characterize the inherent spatiotemporal relationship between the two. When a modality is disturbed (e.g., appearance failure or inaccurate motion prediction), the independent costs will conflict, leading to unstable matching decisions, identity switching, or trajectory loss. Furthermore, frequent occurrences of missed detections, short-term occlusion, and target disappearance-reproduction issues in real-world videos also disrupt trajectory continuity. It is evident that traditional methods struggle to reliably maintain trajectory status in the event of detection failures, thus necessitating a novel technical solution to address these challenges. Summary of the Invention
[0003] To address the aforementioned issues, this invention provides a multi-target tracking method, system, and program product based on appearance-motion collaborative modeling, which effectively improves multi-target tracking performance in complex environments.
[0004] The first aspect discloses a multi-target tracking method using appearance-motion co-modeling. The method includes: acquiring a video frame sequence containing multiple targets; performing target position detection and appearance feature extraction on each video frame in the video frame sequence to obtain the detection position and corresponding re-identification features of each target in each video frame; constructing a corresponding appearance-motion joint matrix for the detection position and corresponding re-identification features of each target in the current video frame, wherein the appearance-motion joint matrix is used to represent the association relationship between each target and historical trajectory, and the current video frame is any video frame in the video frame sequence; and performing association matching between each target in the current video frame and historical trajectory according to the corresponding appearance-motion joint matrix to output multi-target tracking results.
[0005] The second aspect discloses a multi-target tracking system based on appearance-motion co-modeling. The system includes: a video frame sequence acquisition module for acquiring a video frame sequence containing multiple targets; a target position detection and appearance feature extraction module for performing target position detection and appearance feature extraction on each video frame in the video frame sequence, obtaining the detection position and corresponding re-identification features of each target in each video frame; an appearance-motion joint matrix construction module for constructing a corresponding appearance-motion joint matrix for the detection position and corresponding re-identification features of each target in the current video frame, wherein the appearance-motion joint matrix represents the association between each target and historical trajectories, and the current video frame is any video frame in the video frame sequence; and a tracking result output module for performing association matching between each target in the current video frame and historical trajectories based on the corresponding appearance-motion joint matrix, outputting multi-target tracking results.
[0006] The third aspect discloses a computer device including a processor and a memory, the memory storing a computer program that, when executed, implements a multi-target tracking method for appearance-motion co-modeling as disclosed in the first aspect or any possible implementation thereof.
[0007] The fourth aspect discloses a computer program product that, when run on a computer, causes the computer to perform the appearance-motion co-modeling multi-target tracking method disclosed in the first aspect or any possible implementation of the first aspect.
[0008] As can be seen from the above technical solutions, the present invention can achieve the following beneficial effects:
[0009] This invention acquires a video frame sequence containing multiple targets, performs target position detection and appearance feature extraction on each video frame in the sequence, and obtains the detection position and corresponding re-identification features of each target in each video frame. Then, it constructs a corresponding appearance-motion joint matrix for the detection position and corresponding re-identification features of each target in the current video frame. The appearance-motion joint matrix integrates appearance similarity and motion consistency, which can well characterize the correlation between the current detected target and the historical trajectory. Finally, it performs association matching between each target in the current video frame and the historical trajectory based on the corresponding appearance-motion joint matrix, and outputs the multi-target tracking result. This solves the problems of large target offsets between adjacent frames, drastic changes in motion patterns, and matching ambiguities caused by similar appearance features of different targets in complex dynamic scenes, and effectively improves the multi-target tracking performance in complex environments. Attached Figure Description
[0010] Figure 1 The flowchart of a multi-target tracking method for appearance-motion cooperative modeling provided by the present invention is shown.
[0011] Figure 2 The present invention provides a flowchart of a method for constructing an appearance-motion joint matrix.
[0012] Figure 3 This invention provides a framework diagram for a multi-target tracking system based on appearance-motion collaborative modeling. Detailed Implementation
[0013] To make the objectives, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Several embodiments of the present invention are shown in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that the disclosure of the present invention will be thorough and complete.
[0014] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0015] It should be noted that when an element is referred to as being "fixed to" another element, it can be directly on the other element or there may be an intervening element. When an element is considered to be "connected to" another element, it can be directly connected to the other element or there may be an intervening element. The terms "vertical," "horizontal," "left," "right," "up," "down," and similar expressions used herein are for illustrative purposes only and are not intended to indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as limiting the invention.
[0016] In this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," "linking," "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal communication between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances. The term "and / or" as used herein includes any and all combinations of one or more of the related listed items.
[0017] In one embodiment, the present invention provides a multi-target tracking method using appearance-motion cooperative modeling, such as... Figure 1 As shown, the specific steps include:
[0018] S101. Obtain a video frame sequence containing multiple targets, wherein the video frame sequence includes video frames with a continuous input time sequence;
[0019] Specifically, the video frame sequence is a continuous sequence of video frames in a dynamic and complex scene, which provides temporal input for subsequent feature extraction and target association, and the video frame contains multiple targets that need to be tracked.
[0020] It should be noted that this invention acquires a continuous sequence of video frames from dynamic and complex scenes. These complex scenes include factors commonly seen in UAV perspectives, such as rapid changes in viewpoint, relative motion between the camera and ground objects, drastic changes in target scale, and long-distance displacement. These factors can lead to instability in target appearance features and increased position prediction errors, thus posing challenges to multi-target tracking.
[0021] S102. Target position detection and appearance feature extraction are performed on each video frame in the video frame sequence to obtain the detection position of each target and the corresponding re-identification feature in each video frame.
[0022] Specifically, a joint detection and feature extraction network is first used for target location detection and re-identification feature extraction, and the target detection location and corresponding re-identification (ReID) features are output simultaneously. The joint detection and feature extraction network includes three modules: a feature extraction module, a target detection location branch, and a ReID feature branch. The feature extraction module shares the feature encoding structure in a multi-task learning manner.
[0023] The target detection location branch predicts the target center position based on a shared feature map using a center point detection algorithm, and estimates the width and height of the target bounding box through a regression branch, thereby achieving accurate localization of multiple targets in each video frame. The shared feature map is obtained by the feature extraction module. Specifically, keypoint detection algorithms such as heatmap and CenterNet can be used; this invention does not impose specific limitations on the target location detection algorithm.
[0024] The ReID (Re-identification) feature branch extracts appearance embedding features for each target based on a shared feature map, forming a 128-dimensional feature vector. This feature vector describes the target's appearance attributes and supports cross-frame identity preservation. Specifically, the Fairmot algorithm can be used; however, this invention does not impose specific limitations on the ReID feature extraction algorithm.
[0025] It should be further explained that various loss functions can be introduced during the training phase of the network based on joint detection and feature extraction. These include Focal Loss for supervised target center detection, L1 Loss for bounding box size regression, and classification loss and triplet loss for appearance feature discriminativeness. This ensures that the network can achieve high-precision target localization and stable appearance feature representation in a single forward inference during the training process. After training, the network can simultaneously generate the spatial location and stable appearance embedding features of the target, laying the foundation for the subsequent construction of the appearance-motion joint matrix.
[0026] S103. Construct the corresponding appearance-motion joint matrix for the detection position of each target in the current video frame of the video frame sequence and the corresponding re-identification features;
[0027] Specifically, an appearance-motion joint matrix is constructed based on the detection location of each target in the current video frame and its corresponding re-identification features, fusing appearance similarity and motion consistency. This appearance-motion joint matrix represents the association between each target and its historical trajectory; that is, each element in the appearance-motion joint matrix quantifies the pairing reliability between detected targets and historical trajectory pairs. This matrix simultaneously reflects appearance similarity, spatial displacement constraints, and cross-frame temporal consistency, providing a more robust cost representation for data association than traditional single-modal methods.
[0028] In one embodiment, such as Figure 2 As shown, the process of constructing the appearance-motion joint matrix specifically includes:
[0029] S1031. Obtain the appearance guidance prediction position corresponding to each historical trajectory based on the re-identification features of the target to be detected in each historical trajectory and the re-identification feature map of the current video frame, wherein the current video frame is any video frame in the video frame sequence.
[0030] In one embodiment, the step of obtaining the appearance-guided prediction position corresponding to each historical trajectory based on the re-identification features of the target to be detected in each historical trajectory and the re-identification feature map of the current video frame specifically includes:
[0031] For each target's re-identification features in the historical trajectory, the similarity between the re-identification feature map of the current video frame and the re-identification feature map is calculated pixel by pixel to generate the first spatial response map;
[0032] The location of the maximum value in the first spatial response map is used as the appearance guidance prediction position of the historical trajectory in the current video frame.
[0033] Specifically, a first spatial response map is generated by calculating the cosine similarity between the re-identification features of the detected target in each historical trajectory and the re-identification feature map of the current video frame. The location with the maximum response is then extracted as the appearance-guided prediction position of that historical trajectory in the current video frame. It should be noted that each historical trajectory corresponds to one detected target, and each historical trajectory includes at least one trajectory point. For continuous video frame sequences, processing is performed frame-by-frame according to the input time sequence; that is, the same processing operation needs to be performed on each video frame in the video frame sequence. The processing procedure for the current video frame is described in detail below.
[0034] Firstly, utilize the re-identification features of the detected target from the j-th historical trajectory. In the current video frame feature map The first spatial response map is obtained by performing a cosine similarity search, and is determined by the following formula:
[0035] (1)
[0036] in, The sim function represents the location coordinates of the target being detected and is used to determine the similarity between different features.
[0037] Secondly, the location of the maximum response is extracted as the appearance-guided prediction position of the historical trajectory j in the current video frame.
[0038] (2)
[0039] in, The spatial response map representing the historical trajectory j. This represents the predicted position of historical trajectory j within the current video frame, guided by its appearance. M represents the number of historical trajectories, and j represents the index of the historical trajectory. Represents the maximum function. This indicates the coordinates of the predicted location guided by the appearance.
[0040] S1032. Based on the re-identification features of each target in the current video frame and the re-identification feature map of the previous video frame, obtain the predicted position of the target appearance corresponding to each target.
[0041] In one embodiment, the step of obtaining the predicted target appearance location corresponding to each target based on the re-identification features of each target in the current video frame and the re-identification feature map of the previous video frame specifically includes:
[0042] For each target's re-identification features in the current video frame, the similarity between the re-identification feature map and the previous video frame is calculated pixel by pixel to generate a second spatial response map.
[0043] The location of the maximum value in the second spatial response map is used as the target appearance prediction location for each target in the current video frame.
[0044] Specifically, the re-identification features of the i-th detected target are utilized. Re-identification feature map in the previous video frame A cosine similarity inverse search is performed to obtain the second spatial response map, which is determined by the following formula:
[0045] (3)
[0046] (4)
[0047] in, It is the second spatial response map corresponding to the detected target i. This is the predicted position of the detected target's appearance in the previous video frame, where N represents the number of detected targets. The coordinates represent the predicted location of the target's appearance.
[0048] S1033. Calculate the forward spatial distance between the appearance guidance prediction position corresponding to each historical trajectory and the target appearance prediction position corresponding to each target;
[0049] Specifically, the forward spatial distance between the appearance-guided predicted position corresponding to historical trajectory j and the target appearance predicted position corresponding to detected target i is calculated. This measures whether the appearance features of the historical trajectory can accurately point to a certain current detected target, and is expressed by the following formula:
[0050] (5)
[0051] in, Represented as forward spatial distance. This represents the predicted position of the historical trajectory j in the current video frame based on its appearance. The detection position of target i in the current video frame.
[0052] S1034. Calculate the reverse spatial distance between the predicted position of the target appearance and the center position of each historical trajectory for each target, where the center position of the historical trajectory is the position of the last trajectory point in the historical trajectory.
[0053] Specifically, the distance between the predicted position of the detected target i and the position of the last trajectory point of the historical trajectory j is calculated. This measures whether the appearance features of the current detected target can accurately trace back to the historical position of the historical trajectory, and is expressed by the following formula:
[0054] (6)
[0055] in, Represented as reverse spatial distance, Detect target i: Predict the target's appearance and location based on its appearance features predicted in the previous video frame. This represents the actual position of the historical trajectory j in the previous video frame. The smaller these two distances are, the stronger the spatial consistency.
[0056] S1035. The forward spatial distance and the reverse spatial distance are fused using a Gaussian kernel function to obtain the appearance-motion joint matrix corresponding to the current video frame.
[0057] The steps to obtain the appearance-motion joint matrix are determined by the following formula:
[0058] (7)
[0059] Where i represents the target index and j represents the historical trajectory index. Indicates forward spatial distance. Indicates the reverse spatial distance. This is the value of the Gaussian function.
[0060] The above formula can be used to obtain an i. The appearance-motion joint matrix of j, where each element represents the association cost between the detected target i and the historical trajectory j, i.e., the pairing reliability between the two.
[0061] S104. Based on the corresponding appearance-motion joint matrix, perform association matching between each target in the current video frame and the historical trajectory, and output the multi-target tracking result.
[0062] Specifically, a candidate matching pair set is obtained by solving the appearance-motion joint matrix using a global optimal matching algorithm. The candidate matching pair set includes trajectory indices and target indices paired by the minimum association cost. The global optimal matching algorithm can be the Hungarian algorithm, and this invention does not impose any specific limitations.
[0063] When the association cost in the candidate matching pair set is less than the preset validity threshold, the association is determined to be successful. The corresponding historical trajectory is updated using the state information of the detected target, and the updated trajectory state is output.
[0064] The preset validity threshold can be set based on practical experience. The final output is a unique trajectory number and continuous position sequence for each target throughout the entire sequence, forming a stable and complete multi-target tracking result.
[0065] In one embodiment, the present invention further includes:
[0066] If the association cost of a candidate matching pair is greater than or equal to the preset validity threshold, or if the detected target is not included in the candidate matching pair set, it is determined to be a non-match detection, and a new target trajectory is initialized based on the non-match detected target.
[0067] If a historical trajectory is not included in the candidate matching pair set, or the association cost of pairing it with a historical trajectory is greater than or equal to the preset validity threshold, it is determined to be an unmatched trajectory, and a loss count or termination operation is performed on the unmatched trajectory.
[0068] It should be further explained that the multi-target tracking results of the current video frame in the video frame sequence are output and recorded, specifically including the unique trajectory number and continuous position sequence of each target in the entire sequence, so as to be applied when calculating the corresponding target tracking method for the next video frame.
[0069] This invention acquires a video frame sequence containing multiple targets, performs target position detection and appearance feature extraction on each video frame in the sequence, and obtains the detection position and corresponding re-identification features of each target in each video frame. Then, it constructs a corresponding appearance-motion joint matrix for the detection position and corresponding re-identification features of each target in the current video frame. The appearance-motion joint matrix integrates appearance similarity and motion consistency, which can well characterize the correlation between the current detected target and the historical trajectory. Finally, it performs association matching between each target in the current video frame and the historical trajectory based on the corresponding appearance-motion joint matrix, and outputs the multi-target tracking result. This solves the problems of large target offsets between adjacent frames, drastic changes in motion patterns, and matching ambiguities caused by similar appearance features of different targets in complex dynamic scenes, and effectively improves the multi-target tracking performance in complex environments.
[0070] In one embodiment, the present invention discloses a multi-target tracking system based on appearance-motion cooperative modeling, such as... Figure 3 As shown, the system includes four modules: video frame sequence acquisition module, target position detection and appearance feature extraction module, appearance-motion joint matrix construction module, and tracking result output module.
[0071] Specifically, the video frame sequence acquisition module is used to acquire a video frame sequence containing multiple targets; wherein, the video frame sequence includes video frames with a continuous input time sequence;
[0072] The target location detection and appearance feature extraction module is used to perform target location detection and appearance feature extraction on each video frame in the video frame sequence, respectively, to obtain the detection location and corresponding re-identification features of each target in each video frame;
[0073] The appearance-motion joint matrix construction module is used to construct the corresponding appearance-motion joint matrix for the detection position and corresponding re-identification features of each target in the current video frame. The appearance-motion joint matrix is used to represent the association between each target and the historical trajectory. The current video frame is any video frame in the video frame sequence.
[0074] The tracking result output module is used to associate and match each target in the current video frame with the historical trajectory based on the corresponding appearance-motion joint matrix, and output multi-target tracking results.
[0075] This application also provides a computer device, which includes a processor and a memory. The memory stores at least one instruction or at least one program, which is loaded by the processor and executed to perform the appearance-motion co-modeling multi-target tracking method provided in the above-described method embodiments.
[0076] Furthermore, a computer device is provided for implementing the methods provided in the embodiments of this application. This device may participate in constituting or including the apparatus or system provided in the embodiments of this application. The computer device may include one or more processors (processors may include, but are not limited to, processing devices such as microprocessors (MCUs) or programmable logic devices (FPGAs), a memory for storing data, and a transmission device for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply, and / or a camera.
[0077] It should be noted that the aforementioned one or more processors and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits can be implemented wholly or partially as software, hardware, firmware, or any other combination. Furthermore, the data processing circuits can be a single, independent processing module, or wholly or partially integrated into any other element within a device (or mobile device). As involved in the embodiments of this application, the data processing circuit serves as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).
[0078] The memory can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the method described in the embodiments of this application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby realizing the above-mentioned data processing method. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to electronic devices via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0079] The transmission device is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the device's communication provider. In one example, the transmission device includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device may be a Radio Frequency (RF) module, used for wireless communication with the Internet.
[0080] The display can be, for example, a touchscreen liquid crystal display (LCD), which allows users to interact with the user interface of a computer device (or mobile device).
[0081] This application also provides a computer program product or computer program that includes computer instructions stored in a computer storage medium. The processor of a computer device reads the computer instructions from the computer storage medium and executes the computer instructions, causing the computer device to perform the appearance-motion co-modeling multi-target tracking method provided in the above-described method embodiments.
[0082] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, specific embodiments have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in a different order than that shown in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0083] It should be understood that the above description of the preferred embodiments is quite detailed, but it should not be considered as a limitation on the scope of protection of this invention. Those skilled in the art, under the guidance of this invention, can make substitutions or modifications without departing from the scope of protection of the claims of this invention, and all such substitutions or modifications fall within the scope of protection of this invention. The scope of protection of this invention should be determined by the appended claims.
Claims
1. A multi-target tracking method using appearance-motion cooperative modeling, characterized in that, The method includes: Obtain a video frame sequence containing multiple targets, wherein the video frame sequence includes video frames with a continuous input time sequence; Target location detection and appearance feature extraction are performed on each video frame in the video frame sequence to obtain the detection location and corresponding re-identification features of each target in each video frame; For each target in the current video frame in the video frame sequence, a corresponding appearance-motion joint matrix is constructed based on the detection position and corresponding re-identification features. The appearance-motion joint matrix is used to represent the association between each target and the historical trajectory. The current video frame is any video frame in the video frame sequence. Based on the corresponding appearance-motion joint matrix, the various targets in the current video frame are associated and matched with historical trajectories, and the multi-target tracking results are output.
2. The method according to claim 1, characterized in that, The construction of the appearance-motion joint matrix for the detection location and corresponding re-identification features of each target in the current video frame includes: The appearance-guided prediction position corresponding to each historical trajectory is obtained by comparing the re-identification features of the target to be detected in each historical trajectory with the re-identification feature map of the current video frame. The predicted position of the target appearance for each target is obtained by comparing the re-identification features of each target in the current video frame with the re-identification feature map of the previous video frame. Calculate the forward spatial distance between the appearance guidance prediction position corresponding to each historical trajectory and the target appearance prediction position corresponding to each target; Calculate the reverse spatial distance between the predicted position of the target's appearance and the center position of each historical trajectory for each target, where the center position of the historical trajectory is the position of the last trajectory point in the historical trajectory; The forward spatial distance and the backward spatial distance are fused using a Gaussian kernel function to obtain the appearance-motion joint matrix corresponding to the current video frame.
3. The method according to claim 2, characterized in that, The step of obtaining the appearance-guided prediction position corresponding to each historical trajectory based on the re-identification features of the target to be detected in each historical trajectory and the re-identification feature map of the current video frame includes: For each target's re-identification features in the historical trajectory, the similarity between the re-identification feature map of the current video frame and the re-identification feature map is calculated pixel by pixel to generate the first spatial response map; The location of the maximum value in the first spatial response map is used as the appearance guidance prediction position of the historical trajectory in the current video frame.
4. The method according to claim 2, characterized in that, The step of obtaining the predicted target appearance location for each target based on the re-identification features of each target in the current video frame and the re-identification feature map of the previous video frame includes: For each target's re-identification features in the current video frame, the similarity between the re-identification feature map and the previous video frame is calculated pixel by pixel to generate a second spatial response map. The location of the maximum value of the second spatial response map is used as the target appearance prediction location for each target in the current video frame.
5. The method according to claim 2, characterized in that, The step of fusing the forward spatial distance and the backward spatial distance using a Gaussian kernel function to obtain the appearance-motion joint matrix corresponding to the current video frame is determined by the following formula: (7); Where i represents the target index and j represents the historical trajectory index. Indicates forward spatial distance. Indicates the reverse spatial distance. This is the value of the Gaussian function.
6. The method according to claim 1, characterized in that, The step involves associating and matching each target in the current video frame with its historical trajectory based on the corresponding appearance-motion joint matrix, and outputting multi-target tracking results, including: A candidate matching pair set is obtained by solving the appearance-motion joint matrix using a global optimal matching algorithm, wherein the candidate matching pair set includes trajectory index and target index paired by minimum association cost; When the association cost in the candidate matching pair set is less than the preset validity threshold, the association is determined to be successful. The corresponding historical trajectory is updated using the state information of the detected target, and the updated trajectory state is output.
7. The method according to claim 6, characterized in that, The method further includes: If the association cost of a candidate matching pair is greater than or equal to the preset validity threshold, or if the detected target is not included in the candidate matching pair set, it is determined to be a non-match detection, and a new target trajectory is initialized based on the non-match detected target; If a historical trajectory is not included in the candidate matching pair set, or the association cost of pairing it with a historical trajectory is greater than or equal to the validity threshold, it is determined to be an unmatched trajectory, and a loss count or termination operation is performed on the unmatched trajectory.
8. A multi-target tracking system using appearance-motion cooperative modeling, characterized in that, The system includes: a video frame sequence acquisition module, used to acquire a video frame sequence containing multiple targets; The target location detection and appearance feature extraction module is used to perform target location detection and appearance feature extraction on each video frame in the video frame sequence, so as to obtain the detection location and corresponding re-identification features of each target in each video frame. The appearance-motion joint matrix construction module is used to construct a corresponding appearance-motion joint matrix for the detection position and corresponding re-identification features of each target in the current video frame. The appearance-motion joint matrix is used to represent the association between each target and the historical trajectory. The current video frame is any video frame in the video frame sequence. The tracking result output module is used to associate and match each target in the current video frame with the historical trajectory based on the corresponding appearance-motion joint matrix, and output the multi-target tracking result.
9. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing a computer program that is loaded and executed by the processor to implement the method as claimed in any one of claims 1 to 7.
10. A computer program product, characterized in that, The computer program product includes a computer program stored in a computer-readable storage medium, which a processor reads from and executes to implement the method as described in any one of claims 1 to 7.