Cross-view fusion multi-person detection and tracking method for outdoor view variety
By performing time synchronization and geometric calibration on outdoor variety show videos, a unified world coordinate system is established. Combined with sparsely downsampled in-view feature maps and 3D candidate sets, multi-person detection and tracking across viewpoints are achieved, solving the problem of difficult identification and tracking of target characters in outdoor variety shows and improving production quality and real-time performance.
Patent Information
- Application Number
- CN202610059566.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-16
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2046-01-16
AI Technical Summary
Existing technical solutions cannot effectively meet the needs of real-time, accurate, and stable target detection, tracking, and identity association in complex shooting environments for outdoor variety shows. In particular, when multiple cameras are shooting at different locations and the camera angles are changing rapidly, it is difficult to identify and track the target person. Moreover, existing methods are either costly to implement or time-consuming to process.
By performing time synchronization and geometric calibration on multi-camera video frame sequences, a standard projection and back-projection operator under a unified world coordinate system is established. Combined with sparsely downsampled in-view feature maps and 3D candidate sets, a method of cross-view unified geometry, online temporal update and robust identity management is adopted to achieve cross-view multi-person detection and tracking.
Without adding extra sensors or complex reconstruction, it achieves accurate detection, stable tracking, and unified identity positioning in outdoor variety shows, improving the production quality and real-time performance of the program and meeting the requirements of live broadcasting and rapid editing.
Smart Images

Figure CN121545185A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of image processing and target tracking technology, and in particular to a cross-view fusion multi-person detection and tracking method for outdoor variety shows. Background Technology
[0002] The filming and production of outdoor variety shows presents numerous unique and pressing technical challenges that significantly impact the quality and efficiency of program production. These challenges are primarily manifested in the following aspects:
[0003] The filming locations for outdoor variety shows are highly complex and dynamic. Multi-camera cross-shooting is a common technique in outdoor variety shows, with different cameras recording the same scene from different angles to capture exciting moments from all angles. However, this method also brings many problems. Dense crowds and frequent obstruction are extremely common. Numerous actors, guests, and staff move frequently within the scene, blocking each other's view, resulting in significant differences in how a target appears in different camera shots. This poses a major challenge to target identification and tracking.
[0004] Meanwhile, rapid changes in lens and focal length are also a major characteristic of outdoor variety show filming. To create rich visual effects, cameramen quickly adjust the lens focal length and shooting angle according to shooting needs. This causes the size, shape, and position of the target person in the frame to change instantly, further increasing the difficulty of tracking. Moreover, outdoor scenes often have a lightly structured nature and lack stable feature points, unlike indoor scenes which have fixed building structures or iconic objects for reference. This makes it difficult to effectively apply traditional feature point matching techniques.
[0005] Existing technical solutions to address the aforementioned problems in outdoor variety show filming have significant limitations. Some existing methods rely on single-view 2D detection and appearance matching technology. This technology primarily involves 2D detection of targets within a single camera shot and matching and tracking them based on their appearance features. However, in practical applications, frequent occlusion due to dense crowds and significant changes in target appearance during camera transitions easily lead to target ID loss and drift. For example, when a target reappears after being obscured by other objects or people, the system may fail to accurately identify their original identity, resulting in tracking interruptions or incorrect associations.
[0006] Another set of existing solutions employs heavy-duty 3D reconstruction or offline registration methods. These methods attempt to improve target tracking accuracy by constructing complex 3D models or performing offline data registration. However, these methods are costly to implement, requiring substantial computing resources and specialized equipment, and the processing is time-consuming. In the context of live streaming or rapid editing for outdoor variety shows, this latency is unacceptable and fails to meet the requirements of real-time program production.
[0007] In conclusion, existing technical solutions cannot effectively meet the real-time, accurate, and stable target detection, tracking, and identity association requirements of outdoor variety shows in complex shooting environments.
[0008] Therefore, there is an urgent need for a cross-perspective, multi-person detection and tracking method for outdoor variety shows to solve the above problems. Summary of the Invention
[0009] To address the aforementioned technical problems in related technologies, this invention proposes a cross-view fusion multi-person detection and tracking method for outdoor variety shows. The aim is to achieve spatiotemporal structured output of "accurate detection + stable tracking + unified ID + visual positioning" in outdoor variety shows without adding extra sensors or complex reconstruction. This method can directly serve the director's scheduling and post-production, improving the overall production level and quality of outdoor variety shows.
[0010] This invention provides a cross-view fusion method for multi-person detection and tracking in outdoor variety shows, comprising the following steps:
[0011] S1. Time synchronization of the multi-camera video frame sequence to obtain a time-aligned multi-camera video frame sequence. Then, geometric calibration is performed on each camera position to obtain a calibration list. And establish a standard projection operator under a unified world coordinate system for each camera position. With standard return operator This yields a list of standard projection operators. List of standard return operators k is the gate index; K is the number of gates.
[0012] S2, Current video frame for camera position k Sparse downsampling yields in-view feature maps A single-shot detection head is then applied, and a standard return-to-base operator based on position k is used. A 3D candidate set is obtained, and based on the standard projection operator... The corresponding 2D keypoints are generated to construct a fused 3D candidate set. Then, the fused 3D candidate set is subjected to in-view nonmaximum suppression to obtain an in-view candidate set. and its corresponding single-frame trajectory state set ;in, m represents the number of candidates within the viewpoint after nonmaximum suppression; m is the candidate index.
[0013] S3, Viewpoint Feature Map Through standard projection operator Obtaining three-dimensional feature volumes in voxel space with a unified world coordinate system And aggregated into BEV aggregated features along the Z-axis. Later compressed into a 3D perception token Then based on a single 3D perception token Feature map within view Spatial augmentation attention is performed to obtain a refined in-view feature map. Finally, a sequence of detailed feature maps from each camera position was obtained. ;
[0014] S4. Set the trajectory from the previous time step. The query token set is obtained by combining the codes of each trajectory and its 2D predicted key points. This will refine the feature map within the viewpoint. Flatten the linear mapping to obtain image tokens and concatenate them to obtain multi-view image token Y, then use the 3D perception token. After projection, it is concatenated with the multi-view image token Y to obtain a cross-view key value sequence. Finally, the query token set will be... With cross-perspective key value sequences After obtaining the cross features through linear mapping, the current time step trajectory set is obtained through pooling, multilayer perceptron processing, and residual correction. And at each camera position, based on the current time step trajectory and the standard projection operator Real-time projection obtains the projection key points of each camera position. Finally, the set of projection key points for all camera positions is obtained. ;
[0015] S5. Calculate the current time step trajectory set. Ground plane coordinates of each trajectory and candidate set within the viewpoint Candidate ground plane coordinates within each viewpoint Then, based on the two, the OKS-3D paradigm is constructed to calculate the current time step trajectory set. With the set of projection key points Candidates and trajectory similarity And construct the cost matrix accordingly. Then, the Hungarian algorithm is used in the cost matrix. A one-to-one matching of trajectory index i and candidate index m is performed to obtain a unified matching result across different viewpoints. The trajectory set at the current time step is then updated based on the matching result. and set the trajectory at the current time step. As input for step S4 at the next time step.
[0016] Specifically, step S1 includes the following steps:
[0017] S11. Perform frame-level time synchronization on the multi-camera video frame sequence to obtain a time-aligned multi-camera video frame sequence. ;
[0018] S12. Perform geometric calibration on each machine position to obtain a calibration list. And establish standard projection operators for each machine position under the unified world system. This yields a list of standard projection operators. ;in For camera internal parameters; For camera external parameters; For rotation, For translation;
[0019] S13, Based on calibration list Establish standard return-to-base operators for each camera position under a unified world coordinate system. This yields a list of standard return operators. .
[0020] Specifically, step S2 includes the following steps:
[0021] S21. Extract the current video frame of camera position k using the in-view encoder. Sparse downsampling in-view feature map ;
[0022] S22, Feature map within viewpoint Apply a single-shot detection head to multi-scale grid locations based on the standard return-to-field operator. By directly regressing the SMPL triples and confidence scores of each candidate, a 3D candidate set for position k is obtained. ;in, For 3D candidate translation, 3D candidate joint angles 3D candidate body type; 3D candidate confidence scores;
[0023] S23. For each 3D candidate Based on standard projection operator Generate corresponding 2D key points Construct fused 3D candidates, and then perform in-view nonmaximum suppression on the fused 3D candidates to obtain the in-view candidate set. ;in, This represents the number of candidates within the viewpoint after nonmaximum suppression. These are candidate 2D keypoints after nonmaximum suppression; m is the candidate index. For candidate 3D translation within the viewpoint; Candidate joint angles within the field of view; Candidate body types within the field of view; The confidence score of candidates within the viewpoint;
[0024] S24. Based on the candidate set within the viewpoint Construct the single-frame trajectory state set at time step t Among them, single-frame trajectory state .
[0025] Specifically, step S3 includes the following steps:
[0026] S31. For each voxel center in the voxel space of the unified world coordinate system. Using the standard projection operator at position k Calculate its pixel coordinates Then based on pixel coordinates Feature map within view Bilinear sampling is performed, and the average value of the sampling results from all camera positions is calculated to obtain the three-dimensional feature volume. ;
[0027] S32. For three-dimensional feature bodies BEV aggregated features are obtained by applying a 1D convolution along the Z-axis. The BEV score graph is output from the BEV header. ;
[0028] S33, Aggregation features of BEV After feature extraction, it is mapped to a 3D perceptual token using a multilayer perceptron. ;
[0029] S34. Feature map within the viewpoint Linear embedding and flattening are performed to obtain the token sequence. Then, the 3D perception token With token sequence The refined feature map is obtained by concatenating the token dimensions and then refining the result. Finally, a sequence of detailed feature maps from each camera position was obtained. .
[0030] Specifically, step S4 includes the following steps:
[0031] S41. Set the trajectory of the previous time step. Each trajectory and its 2D prediction key points The tokens are encoded into M tokens and then combined to obtain the query token set. ;
[0032] ,
[0033] in, For the 3D translation of the i-th trajectory at time step t-1; Let be the joint angle at time step t-1 on the i-th trajectory; Let be the size of the trajectory at time step t-1; Let be the hidden state of the i-th trajectory at time step t-1; The standard projection operator for camera position k; Indicates by The set of 3D key points obtained from decoding; N represents the number of trajectories that were active at the previous time step t-1; i is the trajectory index; Indicates a trajectory encoder;
[0034] S42. Refine the feature map sequence within the viewpoint. The feature maps within each refined viewpoint are flattened and linearly mapped to image tokens. Then, image tokens for all camera positions. Multi-view image tokens are obtained by sequentially piecing them together. Then, the 3D perception token The projection is a priori token concatenated with the multi-view image token Y to obtain a cross-view key-value sequence. ; It is a linear transformation matrix;
[0035] S43, Set the query tokens With cross-perspective key value sequences Cross features are obtained by performing scaled dot product attention through linear mapping and residual stacking. Then, for the cross features Pooling is performed along each person's token dimension and the hidden state is updated. Finally, the pose increment is regressed through a multilayer perceptron and residual correction is applied to obtain the trajectory at the current time step, ultimately resulting in the set of trajectories at the current time step. ;
[0036] in, The 3D translation of the i-th trajectory at the current time step t; Let be the joint angle of the i-th trajectory at the current time step t; Let be the size of the i-th trajectory at the current time step t; Let represent the hidden state of the i-th trajectory at the current time step t.
[0037] S44. At each camera position, based on the current time step trajectory... With standard projection operator Real-time projection obtains the projection key points of each camera position. Finally, the set of projection key points for all camera positions is obtained. .
[0038] Specifically, step S5 includes the following steps:
[0039] S51. Calculate the current time step trajectory set using the BEV plane position function. 3D translation of trajectory i Ground plane coordinates and candidate set within the viewpoint 3D translation of candidates within each viewpoint Ground plane coordinates ;
[0040] S52. Measure the projection key points of each camera position according to the 2D-OKS paradigm. The similarity between candidate 2D keypoints is used, and then the candidate and trajectory similarity of each camera position is obtained through the OKS-3D paradigm based on the BEV proximity term and similarity. And the trajectory similarity of each camera position Aggregation yields cross-view similarity And construct the cost matrix ;
[0041] S53. Using the Hungarian algorithm in the cost matrix A one-to-one matching of trajectory index i and candidate index m is performed to obtain a unified matching result across viewpoints. Then, the cross-viewpoint similarity corresponding to the matching result is used to obtain the matching result. Similarity threshold with preset matching Perform gating management to update the current time step trajectory set. and set the trajectory at the current time step. As input for step S4 at the next time step.
[0042] Specifically, the 2D-OKS paradigm described in step S52 is shown in the following formula:
[0043] ,
[0044] in, For projection key points and candidate 2D key points The Euclidean distance between them; Visibility is indicated; s is the scale term; is the key point constant; exp() is an exponential function with the natural constant e as its base; Let J be the similarity between keypoints; where J is the number of SMPL keypoints; and j is the keypoint index.
[0045] Specifically, the BEV affinity term mentioned in step S52 As shown in the following formula:
[0046] ,
[0047] in, This indicates the scale of permissible spatial deviation. These represent the tracking location and the predicted location, respectively.
[0048] Specifically, in step S52, candidate positions and trajectory similarities for each camera position are obtained using the OKS-3D paradigm based on the BEV proximity term and similarity. And the trajectory similarity of each camera position Aggregation yields cross-view similarity And construct the cost matrix Specifically, the OKS-3D paradigm is shown in the following formula:
[0049] ,
[0050] in, Let k be the candidate trajectory similarity; It serves as a similarity adjustment factor;
[0051] Candidates for each camera position and trajectory similarity Aggregation yields optimal similarity Based on this, a correlation cost matrix is constructed. ; where max is the function for finding the maximum value.
[0052] Specifically, step S53 calculates the cross-view similarity based on the matching results. Similarity threshold with preset matching Perform gating management to update the current time step trajectory set. Specifically, it includes:
[0053] If the cross-view similarity of the matching pairs (i, m) in the matching results Greater than or equal to the matching similarity threshold If the trajectory index i is 0, it is considered a match, and the trajectory corresponding to the trajectory index i is taken as the surviving trajectory. The surviving trajectory is then subject to deletion judgment. If the cross-view similarity of the matching pair (i, m) in the matching result is 0, it is considered a match. Less than the matching similarity threshold If the match is not found, it is considered a mismatch, and a new match is created for the mismatched pair.
[0054] The deletion determination includes: for each surviving trajectory i, maintaining an exponential moving average of the matching quality to obtain a statistic. ;
[0055] ,
[0056] in, This represents the statistic of the i-th trajectory at the previous time step t-1; For smoothing coefficients;
[0057] Let be the best match similarity between the i-th trajectory and all candidates in the current frame;
[0058] like If the condition is met, then delete the trajectory and add it to the deletion list; where, The similarity-based deletion threshold;
[0059] The newly established discrimination includes: for a non-matching pair (i, m), if its corresponding If so, the newly created trajectory will be added to the new list.
[0060] This invention provides a cross-view fusion method for multi-person detection and tracking in outdoor variety shows, which is achieved collaboratively through three main lines: "unified geometry across views + online temporal update + robust identity management". First, 2-3 camera positions are synchronized and geometrically calibrated to unify multiple feeds into the same world coordinate system. Second, the 3D pose parameters (including translation and posture) of the characters are directly regressed within each viewpoint, resulting in clean 3D candidates for multiple people. Then, the features of each viewpoint are reprojected onto a unified voxel space according to the camera model and aggregated along the height axis to form a BEV (bird's-eye view) prior; this BEV prior is then compressed into a "3D perception token," and spatial enhancement attention is used to feed back the features of each camera position, allowing even occluded viewpoints to leverage geometric evidence from other viewpoints. Next, "tracking-style attention + lightweight memory units" are used to perform online iterative updates to the 3D position and posture of each target: querying the trajectory from the previous frame, with key / value pairs from the current multi-view pixels and the 3D perception token, thus stabilizing the trajectory even when temporarily invisible or during camera position changes. Finally, the "OKS-3D" gating system is used to jointly measure the consistency of 2D keypoints and the distance to the BEV plane, completing optimal matching, instantiation, and deletion across camera positions to maintain a unified ID and consistent trajectory over the long term. Through this process, the system achieves integrated output of "accurate detection + stable tracking + spatial positioning + unified identity" in outdoor variety show scenes without adding extra sensors or relying on heavy 3D reconstruction. This directly serves the director's scheduling and post-production, meeting the real-time and engineering simplicity requirements of live broadcasts and fast editing, thereby improving the overall production level and quality of outdoor variety shows. Attached Figure Description
[0061] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0062] Figure 1 This is a schematic diagram of a cross-view fusion multi-person detection and tracking method for outdoor variety shows, provided by an embodiment of the present invention. Detailed Implementation
[0063] The present invention will be explained in detail through the following embodiments. The purpose of this invention is to protect all technical improvements within its scope. In the description of this invention, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0064] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.
[0065] Example 1
[0066] refer to Figure 1 This embodiment provides a cross-view fusion multi-person detection and tracking method for outdoor variety shows, including the following steps:
[0067] S1. Time synchronization of the multi-camera video frame sequence to obtain a time-aligned multi-camera video frame sequence. Then, geometric calibration is performed on each camera position to obtain a calibration list. And establish a standard projection operator under a unified world coordinate system for each camera position. With standard return operator This yields a list of standard projection operators. List of standard return operators ;
[0068] S11. Perform frame-level time synchronization on the multi-camera video frame sequence to obtain a time-aligned multi-camera video frame sequence. ;
[0069] Multi-camera video frame sequences were obtained by capturing data from multiple camera positions. The original video frames from each camera position are recorded as follows: , representing the video frame t of camera position k; where K is the number of camera positions, and K is a positive integer greater than or equal to 2. In this embodiment... H is the video frame height; W is the video frame width; 3 represents the RGB color channels; k is the camera position index. ; For discrete frame indexing;
[0070] A global timeline is established, with time step t. t is also the discrete frame index of the video, and there is a one-to-one correspondence between the two. That is, in this embodiment, t is both the time step index and the video frame index.
[0071] First, time synchronization is performed to unify the timeline to the global timeline. Using the first camera as a reference, the frame-level constant offset of each camera is estimated. And use it to align to obtain aligned video frames. :
[0072] ,
[0073] Frame-level constant offset It is the first The position of the station relative to the reference position (take) The integer frame offset on the timeline is used to unify multi-camera sequences to the same global time step t. When hardware timecode or Genlock is available, it is directly calculated from the timecode. ;
[0074] in, The timecode for camera position K; The timecode for reference camera position (camera position 1); The time interval on the time axis;
[0075] When there is no hardware, take discrete search. To maximize cross-perspective consistency, we get:
[0076]
[0077] in, For the similarity function, a weighted score of "consistency of the number of people + key point OKS / IoU" or audio peak cross-correlation can be used, and the value corresponding to the peak value can be taken. This allows us to obtain the constant frame offset for each camera position; where... The upper bound of the maximum allowed frame-level search offset (in frames) is used to limit "the maximum possible difference in frames" during exhaustive alignment. This represents the set of detection results (such as number of people / detection boxes / key points, etc.) of the reference camera position (camera position 1) at global time step t. This represents the set of detection results corresponding to camera position k at time step (t+d) after the offset;
[0078] First apply the above formula Time alignment is obtained If necessary, drop / patch frames to match the global frame rate;
[0079] To avoid symbol inflation, time alignment will be performed. Assign to Therefore, a time-aligned multi-camera video frame sequence is obtained. Subsequent multi-camera video frame sequences All are time-aligned multi-camera video frame sequences;
[0080] S12. Perform geometric calibration on each machine position to obtain a calibration list. And establish standard projection operators for each machine position under the unified world system. This yields a list of standard projection operators. ;
[0081] Geometric calibration is performed on each machine position to obtain the calibration. ;in For camera internal parameters; For camera external parameters; For rotation, For translation;
[0082] It is worth noting The "K" in the bottom right corner has no practical meaning and is only used as an identifier; the "K" in the top right corner is the gate index.
[0083] The geometric calibration is performed using a standard multi-view checkerboard calibration board, ChArUco calibration, or AprilTag calibration.
[0084] Calibration objects such as checkerboard calibration boards, ChArUco (a mixture of checkerboard and ArUco markings), and AprilTag have known geometric structures and feature point distributions. By capturing images of the calibration objects from multiple different perspectives, and utilizing the correspondence between feature points in the images and the actual 3D coordinates of the calibration objects, combined with a camera model (such as a pinhole camera model), the intrinsic parameters (focal length, optical center, distortion coefficients, etc.) and extrinsic parameters (camera position and orientation in space) of the camera can be solved. These principles have been researched and verified over many years and are mature existing technologies, which will not be elaborated further here.
[0085] Using the outdoor ground as the world coordinate system (origin at the center of the scene), (Vertical axis), establish standard projection operators for each camera position under a unified world coordinate system. :
[0086] For any world point Its pixel coordinates (u,v) at camera position k satisfy:
[0087] ,
[0088] In the above formula For camera internal references, For rotation, For translation, the dimensions are respectively , , This geometric relationship will subsequently be used to uniformly map multi-view 2D features / keypoints to 3D / BEV space.
[0089] Finally, a calibration list of all camera positions is obtained. List of standard projection operators ;
[0090] S13, Based on calibration list Establish standard return-to-base operators for each camera position under a unified world coordinate system. This yields a list of standard return operators. ;
[0091] Based on the unified world system calibration results, standard return operators for each station are established. Given pixel (u,v) and depth The 3D points in its camera coordinate system are:
[0092] ,
[0093] in, For camera internal parameters;
[0094] This formula is used to generate 3D sampling positions or character translations from 2D features / keypoints at a given depth (related to subsequent pose and BEV construction). In multi-view fusion scenarios, only key shapes are retained as the minimum necessary operators. This allows for direct reuse in subsequent steps, and other optional details are not introduced to keep the implementation concise.
[0095] Finally, a list of standard repaint operators for all camera positions is obtained. ;
[0096] S2, Current video frame for camera position k Sparse downsampling yields in-view feature maps A single-shot detection head is then applied, and a standard return-to-base operator based on position k is used. A 3D candidate set is obtained, and based on the standard projection operator... The corresponding 2D keypoints are generated to construct a fused 3D candidate set. Then, the fused 3D candidate set is subjected to in-view nonmaximum suppression to obtain an in-view candidate set. and its corresponding single-frame trajectory state set ;in, This represents the number of candidates within the viewpoint after nonmaximum suppression. Candidate 2D keypoints;
[0097] S21. Extract the current video frame of camera position k using the in-view encoder. Sparse downsampling in-view feature map :
[0098] ,
[0099] Where C represents the number of channels in the feature map; This represents the height of the feature map after sparse downsampling; Represents the width of the feature map after sparse downsampling; in-view feature map Represents the current video frame 2D feature map from camera position k-view;
[0100] In-view feature map With video frames The pixels in the original image have a one-to-one correspondence; they are used for detection and regression in this step and will also be projected by S3 onto a unified 3D / BEV space.
[0101] In-view encoder This refers to a lightweight backbone (such as a small CNN / ViT) shared by each machine, used to... Mapped to viewpoint feature maps with one-to-one correspondence between pixels. (default It only performs downsampling and channel compression, maintaining grid alignment, which facilitates the subsequent projection of features onto a unified 3D / BEV space using the camera model, allowing them to be directly reused by the detection head and tracking module.
[0102] S22, Feature map within viewpoint Apply a single-shot detection head to multi-scale grid locations based on the standard return-to-field operator. By directly regressing the SMPL triples and confidence scores of each candidate, a 3D candidate set for position k is obtained. ;
[0103] To generate multi-person 3D candidates within a single viewpoint, the feature map within the viewpoint is processed. Apply a single-shot detection head D to multi-scale grid locations based on the standard back-projection operator. The 3D candidate set is obtained by directly regressing the SMPL triples and confidence scores of each candidate. :
[0104] ,
[0105] Among them, the single-shot detection head D is an anchor-box-independent per-grid prediction head. The detection head does not depend on the anchor box, but generates multiple sets of candidate poses and confidence scores at each grid position. For 3D candidate translation (camera coordinate system). 3D candidate joint angles (including global orientation). 3D candidate body type; 3D candidate confidence scores; SMPL (Skinned Multi-Person Linear Model) is a skinned, vertex-based 3D model of the human body that can accurately represent different shapes and poses of the human body.
[0106] SMPL triplet refers to Construct triples to represent the parameters of the SMPL model; where This represents the pose parameters of human joints, also known as joint angles; This represents the shape parameter of the human body, that is, body size; This represents the global translation parameters of the human body, also known as 3D translation;
[0107] Understandable, This is the symbol for the SMPL parameter; since there may be multiple candidates, Then it is used to represent the prediction parameters of the j-th 3D candidate; The prediction parameters used to represent the j-th 3D candidate at camera position k can be considered as follows: and yes Instantiation;
[0108] SMPL triplet refers to Construct triples to represent the parameters of the SMPL model;
[0109] To obtain geometrically consistent The detection head first regresses pixel offset and logarithmic depth Then, using the standard re-projection operator based on camera intrinsic parameters Reproject the candidate center (u,v) as a 3D translation:
[0110]
[0111] in, This is the reference pixel position for the grid. For internal reference focal length, It is a scale constant. The standard back-feed operator defined for S1; This represents the horizontal offset of a pixel. This represents the vertical offset of a pixel. The logarithmic depth is represented by (u,v); (u,v) are pixel coordinates, and z is the depth.
[0112] S23. For each 3D candidate Based on standard projection operator Generate corresponding 2D key points Construct fused 3D candidates, and then perform in-view nonmaximum suppression on the fused 3D candidates to obtain the in-view candidate set. ;
[0113] To support subsequent association and online updates, for each 3D candidate Based on standard projection operator Generating SMPL parameter 3D candidates for regression Generate corresponding 2D key points :
[0114] ,
[0115] in, Indicates by The set of 3D joints obtained by analysis (number) (projection after unit pixel) The standard projection operator in S1;
[0116] Then based on 3D candidates With 2D key points Constructing fused 3D candidates This leads to the fusion of 3D candidate sets. Subsequently, the fusion of 3D candidate sets was performed from a single perspective. Perform non-maximum suppression (NMS) to suppress duplicate candidates and obtain the candidate set within the viewpoint. ;
[0117] in, This represents the number of candidates within the viewpoint after nonmaximum suppression. These are candidate 2D keypoints after nonmaximum suppression; m is the candidate index. For candidate 3D translation within the viewpoint; Candidate joint angles within the field of view; Candidate body types within the field of view; The confidence score of candidates within the viewpoint;
[0118] NMS (Non-Maximum Suppression) is a commonly used method for removing redundant detection boxes in object detection tasks in computer vision, and it is the existing technology.
[0119] In this embodiment, the NMS used is OKS-NMS;
[0120] OKS (Object Keypoint Similarity) is a metric used to measure the similarity of keypoint detections in human pose estimation. OKS-NMS builds upon NMS by incorporating object keypoint information to remove redundant bounding boxes. It considers not only the overlap of bounding boxes but also the matching of keypoints, thus more accurately determining whether two bounding boxes correspond to the same target.
[0121] S24. Based on the candidate set within the viewpoint Construct the single-frame trajectory state set at time step t ;
[0122] video frames This represents the feature map within the viewpoint of camera position k in the t-th video frame. With video frames The original image pixels have a one-to-one correspondence, while the 3D candidate... Again, based on 3D candidates With 2D key points The fusion is constructed and obtained, and then the fusion 3D candidate set is processed in a single view. Performing nonmaximum suppression yields the candidate set within the viewpoint. Although some variables do not explicitly reflect t, they are actually all related to t. Therefore, based on the candidates within each perspective... Initialize its single-frame trajectory state at time step t :
[0123] ,
[0124] in, The hidden state / memory vector of this trajectory (used for subsequent inter-frame association and online updates, such as fusing appearance, motion, and pose consistency information) is set to zero during initialization. . It is the current frame confidence / score output by the detection head, used for reliability weighting of the candidate during initialization or update; the relationship between the two is: Provides observational credibility. Carrying cross-frame accumulated information, available As determined by the gating coefficient The update intensity.
[0125] Single-frame trajectory state This indicates the state of the m-th detection trajectory at time step t (frame t), which is used for subsequent S4 online updates and S5 cross-view fusion;
[0126] Finally, the candidate set within the viewpoint is obtained. The corresponding single-frame trajectory state set ;
[0127] This step first involves the in-view encoder extracting the current video frame from camera position k. Feature maps of sparse downsampling And then A single-shot detector head D is applied to directly regress the multi-person SMPL triplet and confidence level at multi-scale grid locations, and then the regression operator is used. The candidate center is projected from the pixel domain back to a 3D translation in the camera coordinate system. Subsequently, the 3D keypoint set is obtained by regressing the SMPL parameters and then processed by the projection operator. Generate corresponding 2D key points Within a single viewpoint, NMS filtering (OKS-NMS can be used) is performed on the candidates to obtain the candidate set within that viewpoint. Based on this, the construction of the "candidate-instantiation" set and the initialization of the single-frame trajectory state are completed, and finally each video frame... Candidate set corresponding to a viewpoint and single-frame trajectory state set .
[0128] S3, Viewpoint Feature Map Through standard projection operator Obtaining three-dimensional feature volumes in voxel space with a unified world coordinate system And aggregated into BEV aggregated features along the Z-axis. Later compressed into a 3D perception token Then based on a single 3D perception token Feature map within view Spatial augmentation attention is performed to obtain a refined in-view feature map. Finally, a sequence of detailed feature maps from each camera position was obtained. ;
[0129] S31, 2D→3D voxelization projection and filling: For each voxel center in the voxel space of the unified world coordinate system Using the standard projection operator at position k Calculate its pixel coordinates Then based on pixel coordinates Feature map within view Bilinear sampling is performed, and the average value of the sampling results from all camera positions is calculated to obtain the three-dimensional feature volume. ;
[0130] The voxel space of the unified world coordinate system refers to the 3D discrete mesh under the unified world coordinate system. Its scope comes from the shootable area of the outdoor location: first, the image boundaries of each camera position are determined using calibration parameters. Ground-level projection is obtained by taking the intersection / outer rectangle of the visible areas of multiple camera positions. , and then (e.g., 0–3 m, corresponding to human height, selected according to actual needs), its resolution is measured in steps. Uniform discretization yields: , , ;
[0131] Where X is the number of discrete points in the voxel space along the x-direction; Y is the number of discrete points in the voxel space along the y-direction; and Z is the number of discrete points in the voxel space along the z-direction.
[0132] In this embodiment, based on experience... (Approximately shoulder width / standing distance) This not only covers the program's movement area but also matches the granularity of BEV, facilitating subsequent integration and gating.
[0133] Then, for each voxel center in the 3D discrete mesh under the unified world coordinate system Using the standard projection operator at position k Calculate its projected pixel coordinates Then based on the projected pixel coordinates Feature map within view Bilinear sampling is performed, and the average value of the sampling results from all camera positions is calculated to obtain the three-dimensional feature volume. As shown in the following formula:
[0134] ,
[0135] Where C is the number of feature map channels; bilinear() is the bilinear sampling function;
[0136] This process aligns the discriminative features of the "visible viewpoint" in 3D space and provides an aggregateable geometric carrier for the "occluded viewpoint".
[0137] S32, Z-axis aggregation and BEV supervision: for 3D feature volumes BEV aggregated features are obtained by applying a 1D convolution along the Z-axis. The BEV score graph is output from the BEV header. ;
[0138] To form a supervised spatial prior on the horizontal plane (X,Y), the three-dimensional feature volume is... BEV aggregated features are obtained by applying 1D convolution along the Z-axis. :
[0139] ,
[0140] in, This indicates that a 1D convolution is applied along the Z-axis;
[0141] In another possible implementation, a three-dimensional feature body can also be used. Weighted summation along the Z-axis yields ;
[0142] The BEV aggregate feature is output from the BEV header. BEV score graph It is used to characterize the "probability of the existence of the target on the horizontal plane" and serves as a priori for subsequent spatial location.
[0143] A BEV (Bird's Eye View) head is a module or group in a deep learning model used to generate a bird's-eye view (BEV) representation. The BEV head is part of the model's output layer; it receives features extracted from the backbone network (such as a convolutional neural network) and transforms these features into a bird's-eye view. A bird's-eye view is a perspective of observing a scene from directly above, providing information about the scene's planar layout.
[0144] The BEV header is used after Z-axis aggregation to aggregate BEV features. Mapped to a supervised / interpretable BEV score graph The introduced lightweight prediction branch is an output head in the network structure, similar to the common detection / segmentation head.
[0145] The BEV head can be constructed using 1×1 / 3×3 convolutions plus non-linear activation (such as sigmoid), with a very small number of parameters; during training, supervision labels can be generated using the human body center / foot points projected onto the BEV mesh. Supervised learning is conducted to ensure that the source of the "BEV score graph output by the BEV head" is clear and the logic is closed.
[0146] S33, Perceptual Token Embedding: Aggregating Features of BEV After feature extraction, it is mapped to a 3D perceptual token using a multilayer perceptron. ;
[0147] BEV aggregation features After feature extraction (convolution / pooling), it is mapped to a single token through a multilayer perceptron (MLP), i.e., a 3D perceptual token. , That is, the feature map within each viewpoint. That is, it corresponds to a 3D perception token. ;
[0148] This token encapsulates the global geometric semantics of multiple perspectives in 3D / BEV, which is used for subsequent refinement of features of each perspective, and participates in cross-perspective / cross-temporal information interaction as a spatial prior.
[0149] A multilayer perceptron (MLP) is a feedforward artificial neural network model that maps multiple input datasets to a single output dataset.
[0150] S34. Spatial Enhancement Attention: Feature Maps Within the Viewpoint Linear embedding and flattening are performed to obtain the token sequence. Then, the 3D perception token With token sequence The refined feature map is obtained by concatenating the token dimensions and then refining the result. Finally, a sequence of detailed feature maps from each camera position was obtained. ;
[0151] Feature map within view Pass Get the token sequence :
[0152]
[0153] in is a linear flattening function, representing the result of linear embedding and flattening (implemented by convolution) of input features to obtain a token sequence;
[0154] Then the 3D perception token With token sequence The concatenated token sequence is obtained by inputting SEA after concatenating the token dimensions. :
[0155] ,
[0156] in, The spatial enhancement attention module performs self-attention / interactive attention on the spliced sequence, enabling information interaction with each 2D token. This injects the global geometric prior of 3D / BEV into the in-view token sequence, outputting the enhanced result. ;
[0157] Then concatenate the token sequence The image is fed into several Transformer attention blocks for refinement to obtain a refined in-view feature map. :
[0158] ,
[0159] in, The decoding mapping function refines the token (usually by removing the first token or only taking the last token). ( ) are linearly mapped from dimension D back to channel dimension C, resulting in a length of The characteristic sequence; For the rearrangement operator, the length is... The sequence according to Restored to a spatial grid, we obtain .
[0160] Finally, a sequence of detailed feature maps within each camera position was obtained. ;
[0161] This step first "lifts" the 2D features of each viewpoint to a unified 3D voxel space through geometric projection and aggregates them into a BEV along the Z-axis; then, the cross-view geometric prior of the BEV is compressed into a single 3D perception token; finally, spatial augmented attention is performed on the unrefined features of each viewpoint using this token to obtain viewpoint features constrained by 3D / BEV, thereby providing stable spatial priors and cross-view consistency constraints for subsequent online updates.
[0162] S4. Set the trajectory from the previous time step. The query token set is obtained by combining the codes of each trajectory and its 2D predicted key points. This will refine the feature map within the viewpoint. Flatten the linear mapping to obtain image tokens and concatenate them to obtain multi-view image token Y, then use the 3D perception token. After projection, it is concatenated with the multi-view image token Y to obtain a cross-view key value sequence. Finally, the query token set will be... With cross-perspective key value sequences Cross features are obtained through linear mapping, and then processed by pooling and multilayer perceptron with residual correction to obtain the trajectory at the current time step. At each camera position, the trajectory at the current time step is compared with the standard projection operator. Real-time projection obtains the projection key points of each camera position. Finally, the set of current time step trajectories for all camera positions is obtained. and projection key point set ;
[0163] S41, Encoded trajectory as query token: The trajectory set from the previous time step... Each trajectory and its 2D prediction key points The tokens are encoded into M tokens and then combined to obtain the query token set. :
[0164] ,
[0165] ,
[0166] in, Let i be the trajectory at time step t-1 of the i-th trajectory; For the 3D translation of the i-th trajectory at time step t-1; Let be the joint angle at time step t-1 on the i-th trajectory; Let be the size of the trajectory at time step t-1; Let be the hidden state of the i-th trajectory at time step t-1;
[0167] M represents the number of tokens per person. M is a hyperparameter that is set manually, indicating "how many query tokens are used to represent each trajectory / target". In engineering, it can be a fixed small integer (e.g., M=1, 4, 8). M=1 is equivalent to one query token per person. M>1 means that multiple tokens are used to carry different information about the person (posture / appearance / movement, etc.), which improves expressive power and matching robustness. D is the hidden dimension.
[0168] N represents the number of trajectories active at the previous time step t-1 (the size of the trajectory pool), which is dynamically determined by the matching results of the previous time step. The number of trajectories N also corresponds to the number of people detected. The number detected in the current frame is usually denoted as... (Number of candidates for each camera position), while the number of trajectories N will increase or decrease with the "new / termination / loss and reconnection" mechanism: candidates that cannot be matched can trigger the generation of new trajectories, and trajectories that have failed to match for a long time will be terminated.
[0169] A general symbol representing the trajectory state at the previous time step t-1; with subscript i. This represents a specific instance of the i-th trajectory at the previous time step t-1;
[0170] The standard projection operator for camera position k (determined by intrinsic and extrinsic parameters). This represents the set of 3D joints obtained by decoding the SMPL parameters; therefore It means "projecting 3D key points onto the 2D key points of camera position k"; it is related to S23. The same operator has the same meaning; the only difference is that S23 is used for generating 2D keypoints for "candidate j," while S4 is used here for generating 2D predicted keypoints for "trajectory i (previous time step or after prediction)" to match the current candidate. For consistent notation, we will write them as follows: (Candidate) and (Trajectory);
[0171] This represents a trajectory encoder, used to map the trajectory and its projected key points from the previous time step to... Each query token (dimension D) is implemented as a learnable linear / MLP (may contain a small number of convolutions), is fully differentiable, and is used to generate queries for attention.
[0172] Understandable, The current set of valid trajectories is the set of trajectory states output at the previous time step t-1. After cross-view association and update are completed at time t-1, the current set of valid trajectory states is obtained, which is the "tracked trajectory pool" formed by the previous step (the output of the previous round S4 or the first update after initializing S24).
[0173] It is the m-th single-frame candidate of camera position k at time step t (the observation / candidate state generated by the detector head); This refers to the confirmed trajectory status of the global trajectory layer at the previous time step t-1 (used to generate the query token); at time step t, it will... Prediction / projection is performed to obtain the prediction key points for each camera position. and with candidate set Perform a match; if a match is successful, the corresponding candidate will be... Used to update trajectory i Therefore, the subscript evolution rule is: the candidate (k,m,t) is mapped to trajectory index i through "cross-view / cross-target association", and the time is updated from t-1 to t.
[0174] S42. Refine the feature map sequence within the viewpoint. The feature maps within each refined viewpoint are flattened and linearly mapped to image tokens. Then, image tokens for all camera positions. Multi-view image tokens are obtained by sequentially piecing them together. Then, the 3D perception token The projection is a priori token concatenated with the multi-view image token Y to obtain a cross-view key-value sequence. ;
[0175] It can be understood that in this context, "token" means the same thing;
[0176] in, This indicates that the feature map within each refined viewpoint is... First flatten it Each pixel token is then linearly mapped to the hidden dimension. (Commonly obtained by 1×1 convolution + reshape or linear layer) is the image token; concat indicates that the sequence of feature maps within the viewpoint will be refined. Feature maps within all refined viewpoints The corresponding image tokens are concatenated to obtain the multi-view image token Y; , These are feature maps within a refined viewpoint. Height and width in spatial dimensions. Flattened out as... Each pixel token represents a single token, where each spatial location (grid point) in the feature map is considered a token, and the total number of tokens is [number missing]. ;
[0177] 3D perception token The projection is a priori token concatenated with the multi-view image token Y to obtain a cross-view key-value sequence. As shown in the following formula:
[0178]
[0179] in, It is a linear transformation matrix;
[0180] In the token dimension, combine the multi-view image token Y with the 3D perception token. Sequentially concatenated to form a unified key / value sequence for cross-attention, ensuring dimensional consistency; this allows each trajectory to "see" multi-view pixels and the fused 3D prior during evidence collection;
[0181] S43, Set the query tokens With cross-perspective key value sequences Cross features are obtained by performing scaled dot product attention through linear mapping and residual stacking. Then, for the cross features Pooling is performed along each person's token dimension and the hidden state is updated. Finally, the pose increment is regressed through a multilayer perceptron and residual correction is applied to obtain the trajectory at the current time step, ultimately resulting in the set of trajectories at the current time step. ;
[0182] Query token set With cross-perspective key value sequences The query Q, key K, and value V are obtained through linear mapping:
[0183]
[0184] in, To query the mapping matrix; The key mapping matrix; Value mapping matrix;
[0185] Then, scaled dot product attention is performed on the query Q, key K, and value V, and the residuals are stacked to obtain the cross features. :
[0186] ,
[0187] Then, the cross features are analyzed. Pooling and updating the hidden state along each person's token dimension:
[0188] ,
[0189] in, This represents the feature vector after pooling; This indicates that a pooling operation with each person's token dimension M is performed on the input; GRU stands for Gated Recurrent Unit. This represents the output of the GRU at time step t, i.e., the updated hidden state, which combines the hidden state from the previous time step t-1. The pooled feature vector of the current input The hidden state is used in recurrent neural networks to transmit and accumulate contextual information in the sequence.
[0190] GRU (Gated Recurrent Unit) is a variant of recurrent neural network (RNN) that uses a gating mechanism to control the flow of information, enabling it to better capture long-term dependencies in sequential data and alleviate the gradient vanishing problem in traditional RNNs to some extent.
[0191] Finally, the pose increment is recovered using a multilayer perceptron and the trajectory from the previous time step is analyzed. Perform residual correction to obtain the updated trajectory at the current time step. :
[0192] ,
[0193] in, This indicates that the multilayer perceptron (MLP) is specifically designed for regressing pose increments;
[0194] First, the GRU updates and retains the trajectory in a hidden state. by Input only output The residual increment is used to obtain the current time step trajectory set. ;
[0195] in, The 3D translation of the i-th trajectory at the current time step t; Let be the joint angle of the i-th trajectory at the current time step t; Let be the size of the i-th trajectory at the current time step t; Let represent the hidden state of the i-th trajectory at the current time step t.
[0196] S44. At each camera position, based on the current time step trajectory... With standard projection operator Real-time projection obtains the projection key points of each camera position. Finally, the set of projection key points for all camera positions is obtained. ;
[0197] To facilitate subsequent correlation and gating, the trajectory at position k is determined based on the current time step. With standard projection operator Real-time projection obtains the projection key points of camera position k. This leads to the set of projection key points for all camera positions. Where J is the number of SMPL joints;
[0198] S5, via BEV plane position function Calculate the current time step trajectory set 3D translation of each trajectory Ground plane coordinates and candidate set within the viewpoint 3D translation of candidates within each viewpoint Ground plane coordinates Then, based on the two, the OKS-3D paradigm is constructed to calculate the current time step trajectory set. With the set of projection key points Candidates and trajectory similarity And construct the cost matrix accordingly. Then, the Hungarian algorithm is used in the cost matrix. A one-to-one matching of trajectory index i and candidate index m is performed to obtain a unified matching result across different viewpoints. The trajectory set at the current time step is then updated based on the matching result. and set the trajectory at the current time step. As input for step S4 at the next time step.
[0199] S51. To introduce spatial priors, the BEV plane position function is used. Calculate the current time step trajectory set 3D translation of trajectory i Ground plane coordinates and candidate set within the viewpoint Candidates within each viewpoint 3D translation Ground plane coordinates ;
[0200] ,
[0201] in, This represents the 3D translation of trajectory i at time t. After BEV position function The two-dimensional position after projection onto the ground plane; This represents the m-th 3D candidate (or observation) for camera position k at time t. through The two-dimensional position mapped onto the BEV plane;
[0202] S52. Measure the projection key points of each camera position according to the 2D-OKS paradigm. The similarity between candidate 2D keypoints is used, and then the candidate and trajectory similarity of each camera position is obtained through the OKS-3D paradigm based on the BEV proximity term and similarity. And the trajectory similarity of each camera position Aggregation yields cross-view similarity And construct the cost matrix ;
[0203] First, the standard form of 2D-OKS is given. Used to measure two sets of 2D key points , Similarity between them (normalized to [0,1]):
[0204] ,
[0205] in, For projection key points and candidate 2D key points The Euclidean distance between them; Visibility is indicated; s is the scale term; is the key point constant; exp() is an exponential function with the natural constant e as its base; Similarity between key points;
[0206] 2D-OKS (Object Keypoint Similarity) is a core metric for evaluating the accuracy of 2D keypoint detection, and it is widely used, especially in human pose estimation tasks. By comprehensively considering the Euclidean distance between keypoints, visibility markers, and the target scale factor, it provides a similarity measurement method that is more in line with human perception.
[0207] Define BEV affinity terms :
[0208]
[0209] in, This indicates the permissible spatial deviation scale, which is related to the site scale / BEV unit and should be selected according to actual needs. These represent the tracking location and the predicted location, respectively.
[0210] 3D translation Ground plane coordinates and 3D translation Ground plane coordinates Bringing in BEV-friendly features ;
[0211] Based on this, the OKS-3D paradigm is defined (for the similarity between the k-th camera candidate and the trajectory):
[0212] ,
[0213] in, Let k be the candidate trajectory similarity; It serves as a similarity adjustment factor;
[0214] To consider multi-camera redundancy and enhance robustness, candidate cameras for each camera position are compared with trajectory similarity. Aggregation yields optimal similarity Based on this, a correlation cost matrix is constructed. .
[0215] `max` is a function that takes the maximum value and aggregates the similarity between candidates and trajectories of each camera position. It extracts the maximum value from the similarity measurements of multiple camera positions (K camera positions) as the optimal similarity to enhance robustness and provide a basis for the subsequent construction of the association cost matrix.
[0216] S53. Using the Hungarian algorithm in the cost matrix A one-to-one matching of trajectory index i and candidate index m is performed to obtain a unified matching result across viewpoints. Then, the cross-viewpoint similarity corresponding to the matching result is used to obtain the matching result. Similarity threshold with preset matching Perform gating management to update the current time step trajectory set. ;
[0217] The Hungarian algorithm is a combinatorial optimization algorithm that solves the task assignment problem in polynomial time and spurred the development of later primal-dual methods. In 1955, W.W. Kuhn constructed this solution using a theorem by the Hungarian mathematician König, hence the name Hungarian method.
[0218] In this embodiment, the Hungarian algorithm is used to make the global association problem of "trajectory-candidate" a one-to-one optimal allocation: among all candidate matches, the set of matches that minimizes the total cost (equivalent to maximizing the overall similarity) is selected, thereby avoiding identity cross-referencing caused by "local optimal preemption" in crowded multi-person scenarios.
[0219] In the cost matrix The two objects that are matched one-to-one according to the Hungarian algorithm are:
[0220] Line: Track index i has been updated, from ,
[0221] Column: Candidate index m, the candidate set from the perspective of each camera position. ;
[0222] The matching result is a set of associated pairs (a one-to-one mapping of trajectory to detection), which essentially outputs the candidate m to which each trajectory i is assigned (or the assignment fails), and from this, we obtain: matched pairs, the set of unmatched trajectories, and the set of unmatched detections (which are subsequently used for maintenance / deletion and instantiation, respectively).
[0223] Since the Hungarian algorithm only guarantees "one-to-one and globally optimal", but does not guarantee that the matching quality is high enough, this embodiment also needs to use threshold judgment to forcibly remove low similarity assignments to ensure that only credible associations are retained; the removed trajectories / detections will enter instantiation, deletion and other management branches.
[0224] Based on the cross-view similarity of the matching results Similarity threshold with preset matching Perform gating management to update the current time step trajectory set. Specifically, it includes:
[0225] If the cross-view similarity of the matching pairs (i, m) in the matching results Greater than or equal to the matching similarity threshold If the trajectory index i is 0, it is considered a match. The trajectory corresponding to the trajectory index i is taken as the surviving trajectory. The surviving trajectory is then subject to deletion and merging judgments. If the cross-view similarity of the matching pair (i, m) in the matching result is... Less than the matching similarity threshold If the match is not found, it is considered a mismatch, and a new match is created for the mismatched pair.
[0226] Deletion criterion: For each surviving trajectory i, the exponential moving average of the matching quality is used to obtain the statistic. ;
[0227] ,
[0228] in, This represents the statistic of the i-th trajectory at the previous time step t-1. For smoothing coefficients, ;
[0229] Let be the best match similarity between the i-th trajectory and all candidates in the current frame;
[0230] like If the condition is met, then delete the trajectory and add it to the deletion list; where, In this embodiment, the similarity deletion threshold is used. Take 0.15;
[0231] The matching similarity threshold The value is set to 0.2, but can be changed according to actual needs.
[0232] Also includes:
[0233] Merge discrimination (collapse repair): If a pair of trajectories has high similarity for multiple frames (e.g., >20 consecutive frames) If the alignment with the most recent detection is worse, then delete the one that is worse than the one detected recently, to avoid the two trajectories being associated with the same person.
[0234] New discrimination: For non-matching pairs (i, m), detect the corresponding candidates within the viewpoint. If its corresponding If so, the newly created trajectory will be added to the new list;
[0235] Specifically, for the current frame, the best similarity score is obtained when no existing trajectory matches. From the candidates, generate a new trajectory ID, and use the candidate's... Initialize the trajectory state, setting the hidden state to zero (h=0) (a refinement update module can be performed immediately afterwards). This new trajectory will be written to the trajectory management list and added to the trajectory set at the current time step. As the next moment Input;
[0236] After deletion, merging, and creation, the final gated set of surviving trajectories is obtained, and the set of surviving trajectories is used to update the current time step trajectory set. As input for the next time step, S4;
[0237] Update the trajectory management list based on the matching results (delete and create lists), and record the BEV plane positions of surviving trajectories. This data will be compiled to form a spatial reference that can be sustainably transmitted back; at the same time, it will retain... As a statistical basis for deletion or merging in the next frame;
[0238] The trajectory management list is a data structure table that is tracked and maintained online, used to record key information about all active trajectories: trajectory ID, latest status. Historical statistics The list includes recent match timestamps, lifecycle markers (birth / active / lost / death), etc. Each frame performs addition (birth), update (update), keep (keep), deletion (death), and optional merge operations on this list based on the matching results, thereby achieving continuous tracking and status feedback.
[0239] The output is:
[0240] (1) Unified matching results across perspectives (one-to-one mapping between detection and trajectory) and new / deleted lists;
[0241] (2) Set of survival trajectories after gating Its BEV planar location set ,in ;
[0242] (3) Statistics Used for deletion / merging determination in the next frame.
[0243] This embodiment provides a cross-view fusion multi-person detection and tracking method for outdoor variety shows, which is implemented collaboratively along three main lines: "cross-view unified geometry + online temporal update + robust identity management". First, 2-3 camera positions are synchronized in time and calibrated geometrically to unify multiple images into the same world coordinate system. Second, the 3D pose parameters (including translation and posture) of the characters are directly regressed within each viewpoint to obtain clean 3D candidates for multiple people. Then, the features of each viewpoint are reprojected onto a unified voxel space according to the camera model and aggregated along the height axis to form a BEV (bird's-eye view) prior; this BEV prior is then compressed into a "3D perception token", and spatial enhanced attention is used to feed back the features of each camera position, so that even occluded viewpoints can benefit from geometric evidence from other viewpoints. Next, the 3D position and posture of each target are iteratively updated online using "tracking attention + lightweight memory unit": querying the trajectory from the previous frame, with the key / value from the current multi-view pixels and the 3D perception token, so that the trajectory can still be stabilized when the target is temporarily invisible or when the camera position changes. Finally, the "OKS-3D" gating system is used to jointly measure the consistency of 2D keypoints and the distance to the BEV plane, completing optimal matching, instantiation, and deletion across camera positions to maintain a unified ID and consistent trajectory over the long term. Through this process, the system achieves integrated output of "accurate detection + stable tracking + spatial positioning + unified identity" in outdoor variety show scenes without adding extra sensors or relying on heavy 3D reconstruction. This directly serves the director's scheduling and post-production, meeting the real-time and engineering simplicity requirements of live broadcasts and fast editing, thereby improving the overall production level and quality of outdoor variety shows.
[0244] This specification is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this specification. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations and / or block diagrams. Figure 1 A process, multiple processes, and / or boxes Figure 1 Devices that specify the functions in one or more boxes.
[0245] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including an instruction device, which is implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0246] These computer program instructions can also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0247] The parts of this invention not described in detail are prior art. It will be apparent to those skilled in the art that this invention is not limited to the details of the above exemplary embodiments, and that the invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be regarded as exemplary and non-limiting in all respects, and are intended to encompass all changes falling within the meaning and scope of equivalents within this invention.
Claims
1. A cross-view fusion multi-person detection and tracking method for outdoor variety shows, characterized in that, Includes the following steps: S1. Time synchronization of the multi-camera video frame sequence to obtain a time-aligned multi-camera video frame sequence. Then, geometric calibration is performed on each camera position to obtain a calibration list. And establish a standard projection operator under a unified world coordinate system for each camera position. With standard return operator This yields a list of standard projection operators. List of standard return operators k is the gate index; K is the number of gates. S2, Current video frame for camera position k Sparse downsampling yields in-view feature maps A single-shot detection head is then applied, and a standard return-to-base operator based on position k is used. A 3D candidate set is obtained, and based on the standard projection operator... The corresponding 2D keypoints are generated to construct a fused 3D candidate set. Then, the fused 3D candidate set is subjected to in-view nonmaximum suppression to obtain an in-view candidate set. and its corresponding single-frame trajectory state set ;in, m represents the number of candidates within the viewpoint after nonmaximum suppression; m is the candidate index. S3, Viewpoint Feature Map Through standard projection operator Obtaining three-dimensional feature volumes in voxel space with a unified world coordinate system And aggregated into BEV aggregated features along the Z-axis. Later compressed into a 3D perception token Then based on a single 3D perception token Feature map within view Spatial augmentation attention is performed to obtain a refined in-view feature map. Finally, a sequence of detailed feature maps from each camera position was obtained. ; S4. Set the trajectory from the previous time step. The query token set is obtained by combining the codes of each trajectory and its 2D predicted key points. This will refine the feature map within the viewpoint. Flatten the linear mapping to obtain image tokens and concatenate them to obtain multi-view image token Y, then use the 3D perception token. After projection, it is concatenated with the multi-view image token Y to obtain a cross-view key value sequence. Finally, the query token set will be... With cross-perspective key value sequences After obtaining the cross features through linear mapping, the current time step trajectory set is obtained through pooling, multilayer perceptron processing, and residual correction. And at each camera position, based on the current time step trajectory and the standard projection operator Real-time projection obtains the projection key points of each camera position. Finally, the set of projection key points for all camera positions is obtained. ; S5. Calculate the current time step trajectory set. Ground plane coordinates of each trajectory and candidate set within the viewpoint Candidate ground plane coordinates within each viewpoint Then, based on the two, the OKS-3D paradigm is constructed to calculate the current time step trajectory set. With the set of projection key points Candidates and trajectory similarity And construct the cost matrix accordingly. Then, the Hungarian algorithm is used in the cost matrix. A one-to-one matching of trajectory index i and candidate index m is performed to obtain a unified matching result across different viewpoints. The trajectory set at the current time step is then updated based on the matching result. and set the trajectory at the current time step. As input for step S4 at the next time step.
2. The method according to claim 1, characterized in that, Step S1 specifically includes the following steps: S11. Perform frame-level time synchronization on the multi-camera video frame sequence to obtain a time-aligned multi-camera video frame sequence. ; S12. Perform geometric calibration on each machine position to obtain a calibration list. And establish standard projection operators for each machine position under the unified world system. This yields a list of standard projection operators. ;in For camera internal parameters; For camera external parameters; For rotation, For translation; S13, Based on calibration list Establish standard return-to-base operators for each camera position under a unified world coordinate system. This yields a list of standard return operators. .
3. The method according to claim 2, characterized in that, Step S2 specifically includes the following steps: S21. Extract the current video frame of camera position k using the in-view encoder. Sparse downsampling in-view feature map ; S22, Feature map within viewpoint Apply a single-shot detection head to multi-scale grid locations based on the standard return-to-field operator. By directly regressing the SMPL triples and confidence scores of each candidate, a 3D candidate set for position k is obtained. ;in, For 3D candidate translation, 3D candidate joint angles, 3D candidate body type; 3D candidate confidence scores; S23. For each 3D candidate Based on standard projection operator Generate corresponding 2D key points Construct fused 3D candidates, and then perform in-view nonmaximum suppression on the fused 3D candidates to obtain the in-view candidate set. ;in, This represents the number of candidates within the viewpoint after nonmaximum suppression. These are candidate 2D keypoints after nonmaximum suppression; m is the candidate index. For candidate 3D translation within the viewpoint; Candidate joint angles within the field of view; Candidate body types within the field of view; The confidence score of candidates within the viewpoint; S24. Based on the candidate set within the viewpoint Construct the single-frame trajectory state set at time step t Among them, single-frame trajectory state .
4. The method according to claim 3, characterized in that, Step S3 specifically includes the following steps: S31. Apply the standard projection operator of position k to each voxel center in the voxel space of the unified world coordinate system. Calculate its pixel coordinates Then based on pixel coordinates Feature map within view Bilinear sampling is performed, and the average value of the sampling results from all camera positions is calculated to obtain the three-dimensional feature volume. ; S32. For three-dimensional feature bodies BEV aggregated features are obtained by applying a 1D convolution along the Z-axis. The BEV score graph is output from the BEV header. ; S33, Aggregation features of BEV After feature extraction, it is mapped to a 3D perceptual token using a multilayer perceptron. ; S34. Feature map within the viewpoint Linear embedding and flattening are performed to obtain the token sequence. Then, the 3D perception token With token sequence The refined feature map is obtained by concatenating the token dimensions and then refining the result. Finally, a sequence of detailed feature maps from each camera position was obtained. .
5. The method according to claim 4, characterized in that, Step S4 specifically includes the following steps: S41. Set the trajectory of the previous time step. Each trajectory and its 2D prediction key points The tokens are encoded into M tokens and then combined to obtain the query token set. ; , in, For the 3D translation of the i-th trajectory at time step t-1; Let be the joint angle at time step t-1 on the i-th trajectory; Let be the size of the trajectory at time step t-1; Let be the hidden state of the i-th trajectory at time step t-1; The standard projection operator for camera position k; Indicates by The set of 3D key points obtained from decoding; N represents the number of trajectories that were active at the previous time step t-1; i is the trajectory index; Indicates a trajectory encoder; S42. Refine the feature map sequence within the viewpoint. The feature maps within each refined viewpoint are flattened and linearly mapped to image tokens. Then, image tokens for all camera positions. Multi-view image tokens are obtained by sequentially piecing them together. Then, the 3D perception token The projection is a priori token concatenated with the multi-view image token Y to obtain a cross-view key-value sequence. ; It is a linear transformation matrix; S43, Set the query tokens With cross-perspective key value sequences Cross features are obtained by performing scaled dot product attention through linear mapping and residual stacking. Then, for the cross features Pooling is performed along each person's token dimension and the hidden state is updated. Finally, the pose increment is regressed through a multilayer perceptron and residual correction is applied to obtain the trajectory at the current time step, ultimately resulting in the set of trajectories at the current time step. ; in, The 3D translation of the i-th trajectory at the current time step t; Let be the joint angle of the i-th trajectory at the current time step t; Let be the size of the i-th trajectory at the current time step t; Let represent the hidden state of the i-th trajectory at the current time step t. S44. At each camera position, based on the current time step trajectory... With standard projection operator Real-time projection obtains the projection key points of each camera position. Finally, the set of projection key points for all camera positions is obtained. .
6. The method according to claim 5, characterized in that, Step S5 specifically includes the following steps: S51. Calculate the current time step trajectory set using the BEV plane position function. 3D translation of trajectory i Ground plane coordinates and candidate set within the viewpoint 3D translation of candidates within each viewpoint Ground plane coordinates ; S52. Measure the projection key points of each camera position according to the 2D-OKS paradigm. The similarity between candidate 2D keypoints is used, and then the candidate and trajectory similarity of each camera position is obtained through the OKS-3D paradigm based on the BEV proximity term and similarity. And the trajectory similarity of each camera position Aggregation yields cross-view similarity And construct the cost matrix ; S53. Using the Hungarian algorithm in the cost matrix A one-to-one matching of trajectory index i and candidate index m is performed to obtain a unified matching result across viewpoints. Then, the cross-viewpoint similarity corresponding to the matching result is used to obtain the matching result. Similarity threshold with preset matching Perform gating management to update the current time step trajectory set. and set the trajectory at the current time step. As input for step S4 at the next time step.
7. The method according to claim 6, characterized in that, The 2D-OKS paradigm described in step S52 is shown in the following formula: , in, For projection key points and candidate 2D key points The Euclidean distance between them; Visibility is indicated; s is the scale term; is the key point constant; exp() is an exponential function with the natural constant e as its base; Let J be the similarity between keypoints; where J is the number of SMPL keypoints; and j is the keypoint index.
8. The method according to claim 7, characterized in that, The BEV affinity term mentioned in step S52 As shown in the following formula: , in, This indicates the scale of permissible spatial deviation. These represent the tracking location and the predicted location, respectively.
9. The method according to claim 8, characterized in that, In step S52, candidate positions and trajectory similarities for each camera position are obtained using the OKS-3D paradigm based on the BEV proximity term and similarity. And the trajectory similarity of each camera position Aggregation yields cross-view similarity And construct the cost matrix Specifically, the OKS-3D paradigm is shown in the following formula: , in, Let k be the candidate trajectory similarity; It serves as a similarity adjustment factor; Candidates for each camera position and trajectory similarity Aggregation yields optimal similarity Based on this, a correlation cost matrix is constructed. ; where max is the function for finding the maximum value.
10. The method according to claim 9, characterized in that, Step S53: Based on the cross-view similarity corresponding to the matching results Similarity threshold with preset matching Perform gating management to update the current time step trajectory set. Specifically, it includes: If the cross-view similarity of the matching pairs (i, m) in the matching results Greater than or equal to the matching similarity threshold If the trajectory index i is 0, it is considered a match, and the trajectory corresponding to the trajectory index i is taken as the surviving trajectory. The surviving trajectory is then subject to deletion judgment. If the cross-view similarity of the matching pair (i, m) in the matching result is 0, it is considered a match. Less than the matching similarity threshold If the match is not found, it is considered a mismatch, and a new match is created for the mismatched pair. The deletion determination includes: for each surviving trajectory i, maintaining an exponential moving average of the matching quality to obtain a statistic. ; , in, This represents the statistic of the i-th trajectory at the previous time step t-1; For smoothing coefficients; Let be the best match similarity between the i-th trajectory and all candidates in the current frame; like If the condition is met, then delete the trajectory and add it to the deletion list; where, The similarity-based deletion threshold; The newly established discrimination includes: for a non-matching pair (i, m), if its corresponding If so, the newly created trajectory will be added to the new list.
Citation Information
Patent Citations
Multi-target detection method, system and server
CN117218628A
Multi-modal target tracking method based on coupling-decoupling feature enhancement
CN120495347A
Multi-view construction personnel tracking method and system based on attention perception
CN121191074A
Multi-modal sensor-based detection and tracking of objects using bounding boxes
US20250381983A1
Cited By
Image data retrieval method and system based on monitoring big data
CN121765104A
Multi-view crowd tracking method, device and equipment with view-ground interaction and medium
CN121937488A