Unsupervised multi-object tracking method and system based on spatio-temporal features, and electronic device
Patent Information
- Application Number
- CN202610909718.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-23
- Publication Date
- 2026-09-25
AI Technical Summary
[0003]该框架存在三大核心问题:一是检测与重识别优化目标相悖,共享特征无法同时兼顾定位精度与特征判别性;二是空间外观特征与时序运动信息相互割裂,在遮挡、目标密集场景下易出现 ID 切换、轨迹断裂;三是现有无监督训练手段难以同步约束双分支任务,造成特征失衡、误差累积,跟踪鲁棒性不足,难以满足车载实时性与座舱高可靠性要求
[0046]本发明通过检测与重识别双向促进由重识别特征生成空间注意力突出目标区域,由检测位置提供运动先验,两分支相互校准、协同优化。
Smart Images

Figure CN122820767A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent driving, and in particular to an unsupervised multi-target tracking method based on spatiotemporal features, an unsupervised multi-target tracking system based on spatiotemporal features, an electronic device, and a storage medium. Background Technology
[0002] Multi-target tracking is a core perception technology for autonomous driving and smart cockpits, and is widely used in safety scenarios such as vehicle surround view and in-vehicle target tracking. Currently, the mainstream approach adopts a joint tracking framework of detection-recognition, which relies on a shared backbone network to extract features, with two independent branches completing target detection and identity re-identification respectively, and then achieving cross-frame target association through feature similarity.
[0003] The framework has three core problems: First, the optimization goals of detection and re-identification are contradictory, and shared features cannot simultaneously take into account both positioning accuracy and feature discriminativeness; second, spatial appearance features and temporal motion information are disconnected, which can easily lead to ID switching and trajectory breakage in occluded or dense target scenarios; third, existing unsupervised training methods are difficult to simultaneously constrain the dual-branch tasks, resulting in feature imbalance, error accumulation, insufficient tracking robustness, and difficulty in meeting the real-time requirements of vehicle and high reliability requirements of the cockpit.
[0004] Existing improvement solutions in the industry mostly only add attention, motion prediction, or spatiotemporal memory modules, failing to achieve two-way collaboration between detection and re-identification, and failing to fundamentally unify spatiotemporal information or optimize unsupervised training mechanisms, thus still failing to completely solve the aforementioned pain points. Summary of the Invention
[0005] In view of this, the purpose of this invention is to provide an unsupervised multi-target tracking method based on spatiotemporal features, an unsupervised multi-target tracking system based on spatiotemporal features, an electronic device, and a storage medium, aiming to solve one of the technical problems of detection and re-identification tasks being independent of each other, target optimization being contradictory, ID switching and trajectory breakage easily occurring in occluded and dense scenes, and feature imbalance and error accumulation under unsupervised training.
[0006] This invention provides the following solution:
[0007] According to one aspect of the present invention, an unsupervised multi-target tracking method based on spatiotemporal features is provided, comprising the following steps:
[0008] The system acquires continuous video frames captured by the vehicle-mounted camera, inputs the continuous video frames into the shared encoding and decoding backbone network, extracts basic features, and then inputs them into the detection branch and the re-identification branch respectively.
[0009] The feature map output by the re-identification branch is subjected to channel compression and spatial weighting to generate a spatial attention feature map, which enhances the weight of the target region and suppresses the weight of the background interference region.
[0010] Two adjacent video frames are selected, and a feature correlation matrix is constructed using spatial attention feature maps. The location correlation matrix is then constructed by combining the target bounding box, intersection-union ratio, and spatial distance output by the detection branch.
[0011] The feature correlation matrix and the location correlation matrix are subjected to distribution consistency constraints to ensure that the correlation distribution of the two matrices tends to be consistent.
[0012] A dual-matrix consistency loss function is established, and the distribution deviation between the two matrices is eliminated by a distribution difference metric function. End-to-end unsupervised training is then performed on the entire network based on the continuous video frames.
[0013] Target tracking is performed based on the trained network. The inter-frame spatiotemporal displacement is calculated and matched with the historical trajectory database. The target identity is matched by comprehensively utilizing appearance features, spatial location and temporal motion information, and the target trajectory and corresponding identity identifier are output.
[0014] Furthermore, the process of extracting basic features and then inputting them into the detection branch and the re-identification branch includes:
[0015] Extract basic features to input the detection branch, generate detection branch features, and output the target bounding box based on the detection branch features. Perform motion prior on the detection position corresponding to the target bounding box.
[0016] Based on the characteristics of the detection branches, the detection location is obtained;
[0017] Extract basic features and input them into the re-identification branch to generate re-identification branch features;
[0018] Based on the re-identification branch features, channel compression and spatial weighting are performed to generate a spatial attention feature map. At the same time, the effective area of the spatial attention feature map is constrained by the motion prior corresponding to the detection position.
[0019] The detection branch and the re-identification branch are calibrated based on the effective region of the spatial attention feature map and the target bounding box.
[0020] Furthermore, channel compression and spatial weighting are applied to the feature map output by the re-identification branch to generate a spatial attention feature map, specifically including:
[0021] The re-identification branch features are compressed to complete the channel dimensionality reduction process. Then, spatial weights are assigned to each pixel of the dimensionality-reduced feature map. The weight of the target subject region is increased, and the weight of the background and interference regions is reduced to obtain a spatial attention feature map that enhances the target region and suppresses background interference.
[0022] Furthermore, the feature correlation matrix and the location correlation matrix are constructed, specifically including:
[0023] Selecting adjacent t-th and t-1-th video frames, and based on the spatial attention feature maps corresponding to the two frames, calculating the similarity of appearance features of different targets between the frames, and constructing a feature correlation matrix; the elements of the feature correlation matrix represent the degree of matching of appearance features between targets in adjacent frames;
[0024] Extract the target bounding boxes output by the detection branches of frame t and frame t-1, calculate the cross-union ratio and pixel spatial distance between each pair of bounding boxes, combine the cross-union ratio and spatial distance to characterize the degree of spatial matching, and construct the position correlation matrix;
[0025] The elements of the location correlation matrix represent the degree of spatial location matching between targets in adjacent frames.
[0026] Furthermore, distribution consistency constraints are applied to the feature correlation matrix and the location correlation matrix, including:
[0027] Aligning the data distribution of the feature correlation matrix and the location correlation matrix forces the correlation values of the two matrices to be consistent at the corresponding positions of the same target; this makes targets with high similarity in appearance features have a synchronously higher spatial location matching degree, and targets with adjacent spatial locations have a synchronously corresponding appearance feature matching degree.
[0028] Furthermore, the bimatrix consistency loss function is calculated using a distribution difference measure function, which includes KL divergence, JS divergence, or L2 loss.
[0029] By minimizing the distribution deviation between the feature correlation matrix and the location correlation matrix using the bimatrix consistency loss function, end-to-end joint training is performed on the shared encoder-decoder backbone network, detection branch, and re-identification branch under the condition of no target identity ID labeling.
[0030] Furthermore, target tracking is performed based on the trained network, specifically including:
[0031] The video images are input frame by frame and are sequentially processed by the backbone network, the detection branch, and the re-identification branch to generate spatial attention feature maps and target bounding boxes.
[0032] Calculate the inter-frame spatiotemporal displacement by combining the target bounding boxes of adjacent frames, and record the target information of the current frame into the historical trajectory database;
[0033] The system integrates target appearance features, spatial location information, and inter-frame temporal motion information to complete target identity matching and outputs the corresponding identity identifier and continuous motion trajectory for each target.
[0034] According to a second aspect of the present invention, an unsupervised multi-target tracking system based on spatiotemporal features is provided, comprising:
[0035] The module includes a feature extraction module, a spatial attention enhancement module, a matrix construction module, a distribution constraint module, a training optimization module, and a tracking and inference module.
[0036] The feature extraction module is used to acquire continuous video frames captured by the vehicle-mounted camera. The continuous video frames are input into the shared encoding and decoding backbone network, and the basic features are extracted and then input into the detection branch and the re-identification branch respectively.
[0037] The spatial attention enhancement module is used to perform channel compression and spatial weighting on the feature map output by the re-identification branch to generate a spatial attention feature map, which enhances the weight of the target region and suppresses the weight of the background interference region.
[0038] The matrix construction module is used to select two adjacent video frames, construct a feature correlation matrix using spatial attention feature maps, and construct a location correlation matrix by combining the target bounding box, intersection-over-union ratio, and spatial distance output by the detection branch.
[0039] The distribution constraint module is used to perform distribution consistency constraints on the feature correlation matrix and the location correlation matrix to ensure that the correlation distribution of the two matrices tends to be consistent.
[0040] The training optimization module is used to establish a dual-matrix consistency loss function, eliminate the distribution bias of the two matrices through a distribution difference metric function, and perform end-to-end unsupervised training on the entire network based on continuous video frames.
[0041] The tracking inference module is used to track targets based on the trained network, calculate inter-frame spatiotemporal displacement and match it with the historical trajectory database, and comprehensively utilize appearance features, spatial location and temporal motion information to achieve target identity matching, and output the target trajectory and corresponding identity identifier.
[0042] According to three aspects of the present invention, an electronic device is provided, comprising: a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus;
[0043] The memory stores a computer program, which, when executed by a processor, causes the processor to perform steps of an unsupervised multi-target tracking method based on spatiotemporal features.
[0044] According to four aspects of the present invention, a computer-readable storage medium is provided that stores a computer program executable by an electronic device, which, when run on the electronic device, causes the electronic device to perform the steps of an unsupervised multi-target tracking method based on spatiotemporal features.
[0045] Compared with the prior art, the present invention has the following advantages:
[0046] This invention promotes spatial attention to highlight the target region by generating re-identification features through detection and re-identification, and provides motion priors by detection location. The two branches are mutually calibrated and optimized collaboratively.
[0047] This invention constructs a feature correlation matrix and a location correlation matrix through temporal correlation, forces the two to have the same distribution, and binds the re-identified appearance with the detection location depth.
[0048] This invention constructs a more stable spatiotemporal balance loss function through unsupervised training, achieves end-to-end training without ID labeling, reduces data costs, and is adaptable to multiple vehicle scenarios.
[0049] This invention adapts to the high reliability requirements of vehicles / cabins, improving trajectory integrity and reducing ID switching rate, supporting occupant tracking and pedestrian warning, and enhancing driving and cabin safety. Attached Figure Description
[0050] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0051] Figure 1 This is a flowchart of an unsupervised multi-target tracking method based on spatiotemporal features provided by one or more embodiments of the present invention.
[0052] Figure 2 This is a structural diagram of an unsupervised multi-target tracking system based on spatiotemporal features provided by one or more embodiments of the present invention.
[0053] Figure 3 This is a schematic diagram related to unsupervised multi-target tracking based on spatiotemporal features provided by one or more embodiments of the present invention.
[0054] Figure 4 This is a block diagram of an electronic device structure for an unsupervised multi-target tracking method based on spatiotemporal features, provided by one or more embodiments of the present invention. Detailed Implementation
[0055] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0056] The terminology used in the embodiments of this invention is for the purpose of describing particular embodiments only and is not intended to limit the invention. The singular forms “a,” “the,” and “the” as used in the embodiments of this invention and the appended claims are also intended to include the plural forms, and “multiple” generally includes at least two unless the context clearly indicates otherwise.
[0057] It should be understood that the term "and / or" used in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0058] It should be understood that although the terms first, second, third, etc., may be used in the embodiments of the present invention, these descriptions should not be limited to these terms. These terms are only used to distinguish the descriptions. For example, first may also be referred to as second without departing from the scope of the embodiments of the present invention, and similarly, second may also be referred to as first.
[0059] Depending on the context, the words "if" or "if" as used here can be interpreted as "when," "when," "in response to determination," or "in response to detection." Similarly, depending on the context, the phrases "if determination" or "if detection (of the stated condition or event)" can be interpreted as "when determination," "in response to determination," "when detection (of the stated condition or event)," or "in response to detection (of the stated condition or event)."
[0060] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that an article or device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such an article or device. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the article or device that includes said element.
[0061] Figure 1 This is a flowchart of an unsupervised multi-target tracking method based on spatiotemporal features provided by one or more embodiments of the present invention.
[0062] like Figure 1 As shown, it includes the following steps:
[0063] Step S1: Acquire continuous video frames captured by the vehicle-mounted camera, input the continuous video frames into the shared encoding and decoding backbone network, extract basic features, and then input them into the detection branch and re-identification branch respectively;
[0064] Specifically, after extracting basic features, the inputs are given to the detection branch and the re-identification branch, including:
[0065] Extract basic features to input the detection branch, generate detection branch features, and output the target bounding box based on the detection branch features. Perform motion prior on the detection position corresponding to the target bounding box.
[0066] Based on the characteristics of the detection branches, the detection location is obtained;
[0067] Extract basic features and input them into the re-identification branch to generate re-identification branch features;
[0068] Based on the re-identification branch features, channel compression and spatial weighting are performed to generate a spatial attention feature map. At the same time, the effective area of the spatial attention feature map is constrained by the motion prior corresponding to the detection position.
[0069] The detection branch and the re-identification branch are calibrated based on the effective region of the spatial attention feature map and the target bounding box.
[0070] Step S2: Channel compression and spatial weighting are performed on the feature map output by the re-identification branch to generate a spatial attention feature map, which enhances the weight of the target region and suppresses the weight of the background interference region.
[0071] Specifically, the feature map output from the re-identification branch is subjected to channel compression and spatial weighting to generate a spatial attention feature map, including:
[0072] The re-identification branch features are compressed to complete the channel dimensionality reduction process. Then, spatial weights are assigned to each pixel of the dimensionality-reduced feature map. The weight of the target subject region is increased, and the weight of the background and interference regions is reduced to obtain a spatial attention feature map that enhances the target region and suppresses background interference.
[0073] Step S3: Select two adjacent video frames, construct a feature correlation matrix using the spatial attention feature map, and construct a position correlation matrix by combining the target bounding box, intersection-over-union ratio, and spatial distance output by the detection branch.
[0074] Specifically, constructing the feature correlation matrix and the location correlation matrix includes:
[0075] Selecting adjacent t-th and t-1-th video frames, and based on the spatial attention feature maps corresponding to the two frames, calculating the similarity of appearance features of different targets between the frames, and constructing a feature correlation matrix; the elements of the feature correlation matrix represent the degree of matching of appearance features between targets in adjacent frames;
[0076] Extract the target bounding boxes output by the detection branches of frame t and frame t-1, calculate the cross-union ratio and pixel spatial distance between each pair of bounding boxes, combine the cross-union ratio and spatial distance to characterize the degree of spatial matching, and construct the position correlation matrix;
[0077] The elements of the location correlation matrix represent the degree of spatial location matching between targets in adjacent frames.
[0078] Step S4: Perform distribution consistency constraint processing on the feature correlation matrix and the location correlation matrix to ensure that the correlation distribution of the two matrices tends to be consistent.
[0079] Specifically, distribution consistency constraints are applied to the feature correlation matrix and the location correlation matrix, including:
[0080] Aligning the data distribution of the feature correlation matrix and the location correlation matrix forces the correlation values of the two matrices to be consistent at the corresponding positions of the same target; this makes targets with high similarity in appearance features have a synchronously higher spatial location matching degree, and targets with adjacent spatial locations have a synchronously corresponding appearance feature matching degree.
[0081] Step S5: Establish a dual-matrix consistency loss function, eliminate the distribution bias of the two matrices through a distribution difference metric function, and conduct end-to-end unsupervised training on the entire network based on continuous video frames.
[0082] Specifically, the bimatrix consistency loss function is calculated using a distribution difference measure function, which includes KL divergence, JS divergence, or L2 loss.
[0083] By minimizing the distribution deviation between the feature correlation matrix and the location correlation matrix using the bimatrix consistency loss function, end-to-end joint training is performed on the shared encoder-decoder backbone network, detection branch, and re-identification branch under the condition of no target identity ID labeling.
[0084] Step S6: Perform target tracking based on the trained network, calculate inter-frame spatiotemporal displacement and match it with the historical trajectory database, comprehensively utilize appearance features, spatial location and temporal motion information to achieve target identity matching, and output the target trajectory and corresponding identity identifier.
[0085] Specifically, target tracking is performed based on the trained network, including:
[0086] The video images are input frame by frame and are sequentially processed by the backbone network, the detection branch, and the re-identification branch to generate spatial attention feature maps and target bounding boxes.
[0087] Calculate the inter-frame spatiotemporal displacement by combining the target bounding boxes of adjacent frames, and record the target information of the current frame into the historical trajectory database;
[0088] The system integrates target appearance features, spatial location information, and inter-frame temporal motion information to complete target identity matching and outputs the corresponding identity identifier and continuous motion trajectory for each target.
[0089] Specifically, by combining bi-branch feature extraction with motion prior constraints and spatial attention calibration mechanisms, the accuracy and anti-interference capability of target feature extraction are significantly improved. This method relies on a shared encoder-decoder backbone network to achieve shared extraction of basic features. The detection branch and the re-identification branch each perform their respective functions, accurately acquiring target location and appearance feature information respectively, avoiding the information loss problem of single-branch feature extraction. Simultaneously, utilizing the effective region of the spatial attention feature map with motion prior constraints output by the detection branch, combined with channel compression and spatial weighting operations, the feature weights of the target main body region are precisely strengthened, while the feature weights of background, clutter, occlusion, and other interfering regions are suppressed. This achieves adaptive optimization calibration of features, effectively solving the feature extraction failure problem caused by cluttered backgrounds, target occlusion, and weakened features of small targets in complex vehicle scenarios, thus improving the matching accuracy of target appearance and location features.
[0090] Secondly, this embodiment constructs a dual correlation matrix of features and location, combined with distribution consistency constraints, to completely solve the problem of disconnect between appearance feature matching and spatial location matching in traditional tracking methods. Traditional multi-target tracking methods often rely solely on appearance features or spatial location for target matching, which easily leads to problems such as mismatch of appearance-similar targets, mismatch of spatially adjacent targets, and trajectory jumps between frames. This method constructs a feature correlation matrix based on appearance features optimized by spatial attention and a location correlation matrix based on the intersection-union ratio of target bounding boxes and spatial distance, respectively, to represent the target matching relationship between frames from both appearance and spatiotemporal dimensions; at the same time, by aligning the data distribution of the two matrices through distribution consistency constraints, it forces a linkage matching mechanism in which targets with high appearance similarity correspond to high spatial matching degree, and targets with close spatial location correspond to high feature matching degree, effectively avoiding the limitations of single-dimensional matching and significantly improving the accuracy and stability of target association matching between adjacent frames.
[0091] Furthermore, by adopting an unsupervised end-to-end training model, this invention eliminates the reliance on manually labeled target ID data in traditional tracking networks, significantly reducing model training costs and deployment barriers. Existing mainstream multi-target tracking methods largely rely on large amounts of accurate target ID labeled data for supervised training, resulting in a large workload and high cost for data labeling, and labeling errors can easily lead to a decline in model generalization ability. This invention innovatively constructs a dual-matrix consistency loss function, minimizing the distribution deviation of the feature and location correlation matrix through distribution difference measures such as KL divergence, JS divergence, and L2 loss. Even without target ID labeling, end-to-end joint training of the backbone network, detection branch, and re-identification branch can be completed, significantly reducing data preprocessing and manual labeling costs, improving model training efficiency, and effectively avoiding model performance degradation caused by manual labeling bias.
[0092] By fusing multi-dimensional information on appearance, space, and temporal motion, this method significantly improves the robustness and continuity of multi-target tracking in complex vehicle scenarios. The trained model can accurately output target bounding boxes and optimized spatial attention feature maps frame by frame. By calculating inter-frame spatiotemporal displacement and linking with a historical trajectory database, it integrates target appearance features, spatial location information, and temporal motion information for multi-dimensional identity matching. This effectively adapts to complex conditions in vehicle scenarios, such as high-speed target movement, dense intersections, brief occlusion, and scale changes. It effectively reduces problems such as target trajectory breaks, identity jumps, missed tracking, and mistracking, and can stably output continuous and accurate target identity identifiers and motion trajectories, greatly improving the overall performance of multi-target tracking. It is applicable to various complex dynamic scenarios such as autonomous driving, vehicle perception, and intelligent monitoring, and has strong engineering applicability and generalization capabilities.
[0093] Figure 2 This is a structural diagram of an unsupervised multi-target tracking system based on spatiotemporal features provided by one or more embodiments of the present invention.
[0094] like Figure 2 As shown, it includes:
[0095] The module includes a feature extraction module, a spatial attention enhancement module, a matrix construction module, a distribution constraint module, a training optimization module, and a tracking and inference module.
[0096] The feature extraction module is used to acquire continuous video frames captured by the vehicle-mounted camera. The continuous video frames are input into the shared encoding and decoding backbone network, and the basic features are extracted and then input into the detection branch and the re-identification branch respectively.
[0097] The spatial attention enhancement module is used to perform channel compression and spatial weighting on the feature map output by the re-identification branch to generate a spatial attention feature map, which enhances the weight of the target region and suppresses the weight of the background interference region.
[0098] The matrix construction module is used to select two adjacent video frames, construct a feature correlation matrix using spatial attention feature maps, and construct a location correlation matrix by combining the target bounding box, intersection-over-union ratio, and spatial distance output by the detection branch.
[0099] The distribution constraint module is used to perform distribution consistency constraints on the feature correlation matrix and the location correlation matrix to ensure that the correlation distribution of the two matrices tends to be consistent.
[0100] The training optimization module is used to establish a dual-matrix consistency loss function, eliminate the distribution bias of the two matrices through a distribution difference metric function, and perform end-to-end unsupervised training on the entire network based on continuous video frames.
[0101] The tracking inference module is used to track targets based on the trained network, calculate inter-frame spatiotemporal displacement and match it with the historical trajectory database, and comprehensively utilize appearance features, spatial location and temporal motion information to achieve target identity matching, and output the target trajectory and corresponding identity identifier.
[0102] It is worth noting that although only some basic functional modules are disclosed in this embodiment, it does not mean that the composition of this system is limited to the above-mentioned basic functional modules. On the contrary, what this embodiment intends to express is that, based on the above-mentioned basic functional modules, those skilled in the art can arbitrarily add one or more functional modules in combination with existing technology to form an infinite number of embodiments or technical solutions. That is to say, this system is open rather than closed. The fact that this embodiment only discloses a few basic functional modules does not mean that the scope of protection of the claims of this invention is limited to the disclosed basic functional modules. At the same time, for the convenience of description, the above device is described separately according to its functions as various units and modules. Of course, in implementing this invention, the functions of each unit and module can be implemented in one or more software and / or hardware.
[0103] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0104] In one embodiment, a vehicle-mounted video multi-target stable tracking method based on spatiotemporal dual-matrix consistency constraints is applicable to continuous video frame sequences captured by vehicle-mounted forward-looking, surround-looking, and in-cabin cameras. It can achieve stable target detection, re-identification, and continuous trajectory tracking in complex vehicle-mounted scenarios, solving technical problems such as the disconnect between the detection and re-identification branches in traditional algorithms, feature contamination caused by background interference, target matching misalignment, and frequent trajectory ID jumps. This method relies on an encoder-decoder backbone network to achieve unified feature extraction, and combines it with a spatial attention enhancement module, a temporal correlation dual-matrix consistency calculation module, and a consistency loss constraint module to complete multi-branch collaborative optimization, ultimately outputting a continuous and stable target trajectory and a unique tracking ID.
[0105] Includes the following steps:
[0106] The input to this method is a sequence of continuous video frames captured in real time by an in-vehicle camera. After image preprocessing (size normalization, pixel standardization, and noise reduction filtering), the video frames are input into a shared Encoder-Decoder backbone network. The Encoder module is responsible for extracting shallow texture, edge, and deep semantic features step by step, while the Decoder module upsamples and restores the deep features, outputting a multi-scale fused basic feature map.
[0107] The basic feature maps output by the backbone network are fed into two functional branches in parallel: an object detection branch and an object re-identification branch. The detection branch extracts candidate bounding boxes, classifies objects, and regresses their locations based on the basic feature maps, outputting the bounding box coordinates, confidence scores, and category information of all objects within a single frame. The re-identification branch extracts object appearance features based on the basic feature maps, which are used for feature matching and identity association of objects between different frames.
[0108] To address issues such as independent optimization of two branches, feature interference, and inconsistent spatiotemporal matching, this system incorporates three core innovative modules: a spatial attention enhancement module, a temporal correlation-dual-matrix consistency calculation module, and a consistency loss constraint module. Through spatial attention-based feature purification, dual-matrix constraints to unify feature and position matching logic, and a dedicated loss function to force branch collaborative optimization, the system ultimately integrates multi-dimensional information during the inference phase to complete target ID matching and trajectory association, outputting a stable target tracking trajectory and a fixed ID for vehicle-mounted scenarios.
[0109] Traditional vehicle-mounted weight recognition features are easily affected by road debris, background buildings, vehicle interior decorations, and changes in lighting, resulting in feature contamination and inaccurate representation of target appearance features, leading to large matching errors. This module optimizes the original feature map output by the re-recognition branch. Through channel compression and adaptive spatial weight allocation, it strengthens the target's main features and suppresses background interference features to obtain a clean target attention feature map. The specific implementation steps are as follows.
[0110] Step A1, Feature Channel Compression. Receive the multi-channel original feature map output from the re-identification branch, and use a 1×1 convolutional layer to perform dimensionality reduction and compression on the feature channels. Eliminate redundant channel features, retain the core appearance feature information of the target, reduce subsequent computation, and avoid noise interference caused by redundant features.
[0111] Step A2, Spatial Weight Generation. Global pooling and convolutional activation are performed on the compressed feature map. Feature response weights are calculated pixel-by-pixel to generate a spatial weight matrix with the same size as the feature map. The magnitude of the weight matrix corresponds to the saliency of the target features in the region. Target entity regions have high feature saliency, and their weight values adaptively increase; invalid regions such as background, shadows, and interference have low feature saliency, and their weight values adaptively approach 0.
[0112] Step A3, Feature Weighting Enhancement. The original compressed feature map is multiplied and weighted pixel by pixel with the spatial weight matrix to complete feature reconstruction and generate the final spatial attention-enhanced feature map.
[0113] After implementation, this module can accurately locate the spatial position of the target subject in the video frame, thoroughly filter background noise and interference information in complex vehicle scenarios, and provide a clean target feature basis for subsequent inter-frame feature matching and position association calculation, thus solving the problems of re-identification feature pollution and target feature representation distortion from the source.
[0114] For consecutive video frames at time t and t-1, this module constructs a feature correlation matrix and a position correlation matrix based on the clean features enhanced by spatial attention. Through the dual-matrix consistency constraint mechanism, it realizes the spatiotemporal binding of the spatial position information of the detection branch and the appearance feature information of the re-identification branch, solving the problem of independent optimization of the two branches and the contradiction between feature matching and position matching in traditional algorithms. The specific implementation process is divided into three parts.
[0115] Select consecutive adjacent historical frames t-1 and the current frame t, and extract the target appearance features of the two frames after spatial attention enhancement. Assuming there are M valid targets in frame t-1 and N valid targets in frame t, traverse all target feature pairs between the two frames and calculate the appearance feature similarity between each pair of targets.
[0116] This embodiment uses cosine similarity to calculate feature matching degree. It measures the consistency of target appearance features by the angle between vector spaces, effectively avoiding interference from amplitude differences in the matching results. M×N sets of target feature similarity values are calculated one by one, and all values are arranged in an ordered manner to construct a feature correlation matrix M_feat of dimension M×N. The element in the i-th row and j-th column of the matrix represents the degree of matching of appearance features between the i-th target in frame t-1 and the j-th target in frame t. The closer the value is to 1, the higher the similarity of the appearance features of the two targets, and the greater the probability that they are the same target.
[0117] Based on the bounding box information of all targets in frame t-1 and frame t from the detection branch output, the spatial matching degree between targets is calculated by combining spatial positional relationships, and a positional correlation matrix M_loc is constructed. Similarly, based on M targets in frame t-1 and N targets in frame t, the intersection-over-union ratio (IOU) and Euclidean spatial distance are fused to comprehensively evaluate the spatial correlation between targets in the two frames.
[0118] First, the Intersection over Union (IOU) value of each pair of target bounding boxes is calculated to characterize the degree of overlap between the boxes. Then, the Euclidean distance between the center coordinates of the target boxes is calculated to characterize the spatial displacement difference between the targets. These two metrics are then normalized and weighted to obtain a unified spatial location matching degree value. All M×N sets of location matching degree values are arranged in an ordered manner to construct an M×N location correlation matrix M_loc. The closer the matrix element value is to 1, the higher the spatial correlation between the two targets.
[0119] Figure 4 This is a block diagram of an electronic device structure for an unsupervised multi-target tracking method based on spatiotemporal features, provided by one or more embodiments of the present invention.
[0120] like Figure 4 As shown, the present invention provides an electronic device, including: a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus;
[0121] The memory stores a computer program that, when executed by a processor, causes the processor to perform steps of an unsupervised multi-target tracking method based on spatiotemporal features.
[0122] The present invention also provides a computer-readable storage medium storing a computer program executable by an electronic device, which, when run on the electronic device, causes the electronic device to perform the steps of an unsupervised multi-target tracking method based on spatiotemporal features.
[0123] For the sake of simplicity, the method embodiments are described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.
[0124] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present invention.
[0125] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An unsupervised multi-target tracking method based on spatiotemporal features, characterized in that, Includes the following steps: The system acquires continuous video frames captured by the vehicle-mounted camera. Based on the continuous video frames, it inputs them into the backbone network to extract basic features. Then, it extracts the corresponding features through detection and re-identification. The corresponding features of the re-identified region are compressed and spatially weighted to generate a spatial attention feature map, which enhances the weight of the target region and suppresses the weight of the background interference region. Select two adjacent video frames, construct a feature correlation matrix using the spatial attention feature map, and construct a position correlation matrix by combining the detected corresponding features; Constraints are applied to the feature correlation matrix and the location correlation matrix; A dual-matrix consistency loss function is established, and the distribution deviation between the two matrices is eliminated by a distribution difference metric function. End-to-end unsupervised training is then performed on the backbone network using the continuous video frames. Target tracking is performed based on the trained network. The inter-frame spatiotemporal displacement is calculated and matched with the historical trajectory database. Target identity matching is achieved by using appearance features, spatial location, and temporal motion information, and the target trajectory and corresponding identity identifier are output.
2. The unsupervised multi-target tracking method based on spatiotemporal features according to claim 1, characterized in that, The feature extraction through detection and re-identification includes: Extract basic features to input the detection branch, generate detection branch features, and output target bounding boxes based on the detection branch features. Calculate the interaction ratio and spatial distance between adjacent target bounding boxes. The spatial location matching degree of the target is obtained by fusing the interaction ratio and spatial distance between adjacent target bounding boxes; A location correlation matrix is constructed based on the matching degree of all spatial locations, and the matrix elements of the location correlation matrix represent the degree of matching of the target spatial locations in adjacent frames. Extract basic features and input them into the re-identification branch to generate re-identification branch features; Based on the re-identification branch features, channel compression and spatial weighting are performed to generate a spatial attention feature map, while the effective area of the spatial attention feature map is constrained by the detection position correlation matrix. The detection branch and the re-identification branch are calibrated based on the effective region of the spatial attention feature map and the target bounding box.
3. The unsupervised multi-target tracking method based on spatiotemporal features according to claim 2, characterized in that, The process of performing channel compression and spatial weighting on the feature map output by the re-identification branch to generate a spatial attention feature map specifically includes: The re-identification branch features are compressed to complete the channel dimensionality reduction process. Then, spatial weights are assigned to each pixel of the dimensionality-reduced feature map. The weight of the target subject region is increased, and the weight of the background and interference regions is reduced to obtain a spatial attention feature map that enhances the target region and suppresses background interference.
4. The unsupervised multi-target tracking method based on spatiotemporal features according to claim 2, characterized in that, The construction of the feature correlation matrix and the location correlation matrix specifically includes: Selecting adjacent t-th and t-1-th video frames, and based on the spatial attention feature maps corresponding to the two frames, calculating the similarity of appearance features of different targets between the frames, and constructing a feature correlation matrix; the elements of the feature correlation matrix represent the degree of matching of appearance features between targets in adjacent frames; Extract the target bounding boxes output by the detection branches of frame t and frame t-1, calculate the cross-union ratio and pixel spatial distance between each pair of bounding boxes, and combine the cross-union ratio and pixel spatial distance to output the spatial matching degree and construct the position correlation matrix.
5. The unsupervised multi-target tracking method based on spatiotemporal features according to claim 1, characterized in that, The constraint processing of the feature correlation matrix and the location correlation matrix includes: Align the data distribution of the feature correlation matrix and the location correlation matrix to force the correlation values of the two matrices to be consistent at the corresponding positions of the same target; This results in targets with high similarity in appearance having a correspondingly high spatial matching degree, and targets that are spatially adjacent have a correspondingly high appearance matching degree.
6. The unsupervised multi-target tracking method based on spatiotemporal features according to claim 1, characterized in that, The bimatrix consistency loss function is calculated using a distribution difference metric function. The distribution difference measurement function includes: KL divergence, JS divergence, or L2 loss; Based on the bimatrix consistency loss function minimizing the distribution deviation between the feature correlation matrix and the location correlation matrix, the backbone network, detection branch, and re-identification branch are jointly trained end-to-end without target identity ID labeling.
7. The unsupervised multi-target tracking method based on spatiotemporal features according to claim 1, characterized in that, The target tracking based on the trained network includes: Frame by frame, continuous video frames are input and generated into spatial attention feature maps and target bounding boxes through the backbone network, detection branch, and re-identification branch. Calculate the inter-frame spatiotemporal displacement by combining the target bounding boxes of adjacent frames, and record the target information of the current frame into the historical trajectory database; The system integrates target appearance features, spatial location information, and inter-frame temporal motion information to complete target identity matching and outputs the corresponding identity identifier and continuous motion trajectory for each target.
8. An unsupervised multi-target tracking system based on spatiotemporal features, characterized in that, The system is used to perform the tracking method according to any one of claims 1-7, including: The module includes a feature extraction module, a spatial attention enhancement module, a matrix construction module, a distribution constraint module, a training optimization module, and a tracking and inference module. The feature extraction module is used to acquire continuous video frames captured by the vehicle-mounted camera, input the continuous video frames into the backbone network, extract basic features, and then input them into the detection branch and the re-identification branch respectively. The spatial attention enhancement module is used to perform channel compression and spatial weighting on the feature map output by the re-identification branch to generate a spatial attention feature map, enhance the weight of the target region and suppress the weight of the background interference region. The matrix construction module is used to select two adjacent video frames, construct a feature correlation matrix using the spatial attention feature map, and construct a position correlation matrix by combining the corresponding features extracted from the detection branch. The distribution constraint module is used to perform constraint processing on the feature correlation matrix and the location correlation matrix; The training optimization module is used to establish a dual-matrix consistency loss function, eliminate the distribution bias of the two matrices through a distribution difference measurement function, and perform end-to-end unsupervised training on the backbone network based on the continuous video frames. The tracking inference module is used to track targets based on the trained network, calculate inter-frame spatiotemporal displacement and match it with the historical trajectory database, use appearance features, spatial location and temporal motion information to achieve target identity matching, and output the target trajectory and corresponding identity identifier.
9. An electronic device, characterized in that, include: The processor, communication interface, memory, and communication bus are connected, with the processor, communication interface, and memory communicating with each other via the communication bus. The memory stores a computer program that, when executed by the processor, causes the processor to perform the steps of the unsupervised multi-target tracking method based on spatiotemporal features as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, It stores a computer program executable by an electronic device, which, when run on the electronic device, causes the electronic device to perform the steps of the unsupervised multi-target tracking method based on spatiotemporal features as described in any one of claims 1-7.