Cross-camera target alignment and deduplication counting method based on timing consistency

By adopting a cross-camera target alignment and deduplication counting method based on temporal consistency, this method utilizes temporal relationships and motion behavior characteristics to solve the problem of cross-camera target alignment and deduplication counting in industrial conveyor belt scenarios, achieving a high-precision, low-cost, and highly adaptable counting solution.

CN121962209BActive Publication Date: 2026-06-02HEFEI YIWEI QUANTUM TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HEFEI YIWEI QUANTUM TECH CO LTD
Filing Date
2026-03-27
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing technologies for cross-camera target alignment and deduplication in industrial conveyor belt scenarios suffer from problems such as over-reliance on ideal conditions, lack of adaptability, and high system coupling, making it difficult to achieve accurate cross-camera target alignment and deduplication in environments with strong dynamic characteristics.

Method used

A cross-camera target alignment and deduplication counting method based on temporal consistency is adopted. This method involves establishing a multi-camera temporal reference, single-camera target detection and trajectory generation, generating a candidate trajectory pair set, matching scoring, global matching and conflict resolution, and deduplication counting and uncertainty labeling. It utilizes temporal relationships and motion behavior features to perform target association and deduplication.

Benefits of technology

It achieves high-precision cross-camera target alignment and deduplication without the need for 3D reconstruction and spatial calibration. It is suitable for field-of-view separation scenarios, meets industrial real-time requirements, reduces deployment and maintenance costs, and provides a modular architecture for easy integration, improving the system's adaptability and interpretability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121962209B_ABST
    Figure CN121962209B_ABST
Patent Text Reader

Abstract

The present application relates to machine vision, in particular to a cross-camera target alignment and deduplication counting method based on timing consistency, obtaining a video stream of a camera, establishing a unified time reference system; independently executing target detection and tracking algorithms in each camera, and generating a local trajectory for each target; generating a candidate trajectory pair set according to the number of consecutive tracking frames of the target in a single camera and the time window of the trajectory appearing in different cameras; calculating a matching score for each candidate trajectory pair based on motion behavior consistency; globally optimizing and conflict resolving all candidate trajectory pairs, and outputting a cross-camera associated trajectory pair set; merging the corresponding cross-camera associated trajectory pairs into the same target by using an uncertainty quantification mechanism, directly performing deduplication counting, and simultaneously performing uncertainty marking; the present application can overcome the defect that accurate cross-camera target alignment and deduplication counting cannot be performed in an industrial conveyor belt scene with dynamic characteristics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to machine vision, and more specifically to a cross-camera target alignment and deduplication counting method based on temporal consistency. Background Technology

[0002] In the field of industrial automation visual inspection, especially in large-scale production lines, warehousing and logistics, and port loading and unloading scenarios, conveyor belt material counting is often a key link in production management and quality control. With the expansion of production scale and the increase in automation, a single camera can no longer cover the entire conveyor area, necessitating the deployment of multiple cameras working collaboratively. However, this multi-camera deployment method brings technical challenges to cross-camera target alignment and deduplication counting.

[0003] Existing cross-camera association and deduplication methods can be mainly divided into the following five categories, all of which have significant limitations:

[0004] 1) Methods based on 3D reconstruction and spatial calibration: Relying on precise camera calibration and stereo vision baseline, parameters are easily drifted due to environmental disturbances in industrial sites, and there is no effective baseline for long-distance deployment. Calibration and maintenance costs are high, and positioning accuracy drops sharply as the distance between cameras increases.

[0005] 2) Matching method based on spatial overlap area: requires that there is an overlapping transition zone in the field of view of the camera, but the overlap rate of the field of view in more than 70% of industrial scenes is extremely low or even non-overlapping, and the difference in viewing angle and lighting leads to poor matching reliability.

[0006] 3) Appearance feature re-identification (ReID) method: Due to the high homogeneity of industrial materials, changes in lighting, posture and occlusion, the appearance feature has low distinguishability and poor stability, and the matching accuracy is difficult to meet industrial needs.

[0007] 4) Deduplication method based on simple time window: It adopts a fixed time threshold, which does not adapt to the dynamic characteristics such as conveyor belt speed fluctuation and material slippage and stagnation. Time correlation fails, and the false alarm rate remains high.

[0008] 5) Comprehensive method of integrating multi-source information: Excellent laboratory results, but complex on-site deployment, poor real-time performance, high coupling degree that cannot be independently integrated, and high cost of upgrade and adaptation.

[0009] Therefore, it is evident that existing technologies have the following three fundamental limitations in solving the problem of cross-camera target alignment and deduplication counting in industrial conveyor belt scenarios:

[0010] 1) Over-reliance on ideal conditions: Most existing methods assume precise calibration, overlapping fields of view, or stable appearance features, which are often unmet in industrial settings;

[0011] 2) Lack of adaptability: It cannot adapt to dynamic characteristics such as changes in conveyor belt speed and uncertainties in material behavior;

[0012] 3) High system coupling: It is difficult to integrate into the existing system as an independent module, and the deployment and maintenance costs are high.

[0013] In numerous industrial field practices, cameras are often deployed only in segments along conveyor belts. Targets between different cameras only exhibit "temporal and motion continuity," without reliable spatial overlap or stable appearance features. Therefore, there is an urgent need for a cross-camera target alignment and deduplication method that does not rely on 3D reconstruction, spatial calibration, or field-of-view overlap, but solely on temporal relationships and motion behavior characteristics to address the practical pain points in industrial settings. Summary of the Invention

[0014] In view of the above-mentioned shortcomings of the existing technology, the present invention provides a cross-camera target alignment and deduplication counting method based on temporal consistency, which can effectively overcome the shortcomings of the existing technology in that it cannot perform accurate cross-camera target alignment and deduplication counting in industrial conveyor belt scenarios with dynamic characteristics.

[0015] To achieve the above objectives, the present invention provides the following technical solution:

[0016] A cross-camera target alignment and deduplication counting method based on temporal consistency includes the following steps:

[0017] S1. Multi-camera time reference establishment: Acquire video streams from at least two cameras and establish a unified time reference system;

[0018] S2. Single-camera target detection and trajectory generation: Target detection and tracking algorithms are executed independently within each camera, and a local trajectory is generated for each target;

[0019] S3. Generation of candidate trajectory pair set: Based on the number of frames the target is continuously tracked in a single camera and the time window in which the trajectory appears in different cameras, candidate trajectory pairs are filtered to generate a candidate trajectory pair set.

[0020] S4. Matching score: Calculate the matching score for each candidate trajectory pair based on the consistency of motion behavior;

[0021] S5. Global Matching and Conflict Resolution: Perform global optimization and conflict resolution on all candidate trajectory pairs, and output a set of cross-camera associated trajectory pairs;

[0022] S6. Deduplication and Uncertainty Labeling: Using an uncertainty quantification mechanism, corresponding cross-camera associated trajectory pairs are merged into the same target, and deduplication is performed directly, while uncertainty is labeled.

[0023] Preferably, in S1, the multi-camera time reference is established by acquiring video streams from at least two cameras and establishing a unified time reference system, including:

[0024] Acquire video streams from at least two cameras and establish a unified time reference system using NTP protocol, GPS timestamps, hardware synchronization signals, or software synchronization mechanisms;

[0025] Assign a globally unique frame number to each video frame and establish a mapping relationship between the frame number and physical time;

[0026] Among them, a clock drift compensation mechanism is introduced: a sliding window algorithm is used to monitor and compensate for clock drift between cameras in real time;

[0027] Introducing a time calibration offset mechanism: Allowing clock errors within ±100ms range, improving the system's adaptability in real industrial environments.

[0028] Preferably, in S2, single-camera target detection and trajectory generation involves: independently executing target detection and tracking algorithms within each camera and generating a local trajectory for each target, including:

[0029] S21. Execute the target detection and tracking algorithm independently within each camera and assign a locally unique ID to each target to maintain the temporal continuity of the target within a single camera;

[0030] S22. Generate a local trajectory for each target, including:

[0031] The sequence of target positions within the image plane;

[0032] Timestamps of the target entering / leaving the Region of Interest (ROI);

[0033] The inter-frame displacement and velocity change sequence of the target;

[0034] The sequence of changes in the confidence level of the target;

[0035] S23. A lightweight temporal filtering algorithm is used to smooth the trajectory and reduce noise interference in a single frame.

[0036] Preferably, in S3, the candidate trajectory pair set is generated by: filtering candidate trajectory pairs based on the number of consecutive frames the target is tracked within a single camera and the time windows in which the trajectory appears within different cameras, and generating a candidate trajectory pair set, including:

[0037] S31. For all trajectories within cameras A and B, cross-camera matching is only allowed when the target is continuously tracked for more than a preset frame rate within a single camera.

[0038] S32, Regarding the trajectory T within camera AA Based on the time t that the target leaves camera A A_exit Predict the entry time t of the target within camera B. A_enter And based on time constraints, candidate trajectories are selected from the trajectories within camera B:

[0039] ;

[0040] in, This is the time window threshold, with the initial value calculated based on the nominal speed of the conveyor belt and dynamically adjusted based on historical matching results.

[0041] S33. Generate a set of candidate trajectory pairs. ;

[0042] Among them, T B1 T B2 These are the first and second candidate trajectories within camera B, respectively.

[0043] Preferably, in S3, the candidate trajectory pair set is generated by filtering candidate trajectory pairs based on the number of consecutive frames the target is tracked within a single camera and the time windows in which the trajectory appears within different cameras. This process includes:

[0044] Appearance verification: Based on appearance consistency, candidate trajectory pairs with obvious mismatches are eliminated, specifically including:

[0045] 1) Extract low-dimensional appearance features of the target, including size variation trends, color statistics, and shape invariant moments;

[0046] 2) Calculate the feature distance d:

[0047] ;

[0048] Among them, f A f B These are the appearance feature vectors of the targets within cameras A and B, respectively. This indicates the calculation of the L2 norm;

[0049] 3) Convert the feature distance d into an appearance consistency verification score f. score :

[0050] ;

[0051] Among them, the appearance consistency verification score f score A larger value indicates a more similar appearance. d max The maximum allowable distance threshold is determined based on historical data statistics;

[0052] 4) Appearance consistency verification score fscore The candidate trajectory pairs are compared with a preset lenient score threshold, and those that do not match are removed based on the comparison results.

[0053] Preferably, in S4, the matching score is calculated based on the consistency of motion behavior for each candidate trajectory pair, including:

[0054] S41. Calculate the time continuity score f time :

[0055] ;

[0056] S42. Compare the trend of target's motion speed changes in different cameras:

[0057] S421, Calculate the standard deviation of the speed difference. :

[0058] ;

[0059] in, , are the velocity of the target in the i-th frame within cameras A and B, respectively, and n is the number of comparison frames;

[0060] S422. Calculate the motion speed consistency score f. speed :

[0061] ;

[0062] Where k is the scaling factor, used to adjust the standard deviation of the speed difference. Consistency score for movement speed f speed The extent of the impact;

[0063] S43. Compare the trend of the target's motion direction changes in different cameras:

[0064] S431. Extract the target's motion direction angle sequences dA and dB within cameras A and B:

[0065] ;

[0066] ;

[0067] in, , These are the motion direction angles of the target in the i-th frame within cameras A and B, respectively, which are the angles relative to the X-axis direction of the image coordinate system.

[0068] S432. Determine the angular difference sequence of the target's motion direction within cameras A and B. , :

[0069] ;

[0070] ;

[0071] in, Let be the difference in the motion direction angle of the target in the i-th and i+1-th frames within camera A. Let be the difference in the motion direction angle of the target in the i-th and i+1-th frames within camera B;

[0072] S433, Calculate the motion direction consistency score f direction :

[0073] ;

[0074] Where cos_sim is the similarity function;

[0075] S44. Compare the temporal characteristics of the target in different cameras, including acceleration change trends and pause patterns, and calculate the motion state consistency score f. behavior ;

[0076] S45. Calculate the matching score:

[0077] ;

[0078] in, , , , All are weighting coefficients, and .

[0079] Preferably, in S5, global matching and conflict resolution involve performing global optimization and conflict resolution on all candidate trajectory pairs, outputting a set of cross-camera associated trajectory pairs, including:

[0080] S51. The maximum weight matching algorithm is used to perform global optimization on all candidate trajectory pairs to obtain an initial global matching scheme.

[0081] S52. Utilize conflict resolution strategies to resolve conflicts in the initial global matching scheme, avoiding one-to-many and many-to-one matching patterns:

[0082] Set a minimum score threshold; candidate trajectory pairs with scores below the minimum score threshold are directly rejected.

[0083] Introducing a time consistency constraint: the trajectories of the same target in multiple cameras do not overlap in time;

[0084] If multiple candidate trajectory pairs share the same target, the candidate trajectory pair with the highest matching score is retained;

[0085] S53. Introduce a backtracking correction mechanism: When a new trajectory is added, a decision is made on whether to accept the new match based on the confidence level. At the same time, when the addition of a new trajectory causes the original match to become invalid, the historical match can be readjusted.

[0086] S54, Output a set of cross-camera associated trajectory pairs.

[0087] Preferably, in S6, deduplication and uncertainty labeling: The corresponding cross-camera associated trajectory pairs are merged into the same target using an uncertainty quantization mechanism, and deduplication is performed directly. Simultaneously, uncertainty labeling is performed, including:

[0088] S61. Introducing an uncertainty quantification mechanism:

[0089] S611. Calculate the matching confidence level C. match :

[0090] ;

[0091] Where max(Score) and second_max(Score) are the highest and second highest matching scores, respectively;

[0092] S612. Introduce the time window coverage factor:

[0093] ;

[0094] Wherein, min_end and max_end are the smaller and larger values ​​of the end time of the two segments of the trajectory in the cross-camera associated trajectory pair, respectively; and min_start and max_start are the smaller and larger values ​​of the start time of the two segments of the trajectory in the cross-camera associated trajectory pair, respectively.

[0095] S613, Calculate the final uncertainty:

[0096] ;

[0097] Among them, A history Historical matching accuracy measures the system's success rate in matching similar historical scenarios. , , All are weighting coefficients, and ;

[0098] S62. Based on the final uncertainty, perform hierarchical processing on all cross-camera associated trajectory pairs:

[0099] For highly deterministic cross-camera associated trajectory pairs with uncertainty < 0.3, they are merged into the same target and deduplication is performed directly.

[0100] For deterministic cross-camera associated trajectory pairs with a probability of 0.3 ≤ uncertainty < 0.7, they are marked as "requires verification" and provided to the upper-level system for further processing.

[0101] For cross-camera associated trajectory pairs with low uncertainty (uncertainty ≥ 0.7), the count is rejected directly, the anomaly is recorded, and manual review is requested.

[0102] S63. Generate metadata for the cross-camera associated trajectory pairs that "require verification". The metadata includes:

[0103] The chain of evidence that matches the decision;

[0104] Detailed scoring for each dimension;

[0105] Visual summary of candidate trajectories;

[0106] S64. Output the deduplication count results, as well as the cross-camera associated trajectories and their metadata that "need verification", for use by the upper-layer system.

[0107] Compared with existing technologies, the cross-camera target alignment and deduplication counting method based on temporal consistency provided by this invention has the following beneficial effects:

[0108] 1) It does not rely on 3D reconstruction and spatial calibration, significantly reducing the deployment threshold.

[0109] This invention completely eliminates the reliance on camera intrinsic and extrinsic parameter calibration, stereo vision reconstruction, and world coordinate system mapping. In industrial settings, vibration, temperature and humidity changes, and installation limitations often lead to calibration parameter drift, causing 3D reconstruction-based methods to fail. This invention utilizes only the temporal relationship and motion behavior characteristics of the target appearing in different cameras, enabling stable operation in harsh environments where calibration fails or is impossible, significantly reducing deployment and maintenance costs. Actual industrial tests show that even without precise calibration, this invention can still maintain high-precision counting performance, while traditional 3D reconstruction-based methods show a significant performance decline under the same conditions.

[0110] 2) Applicable to scenarios with separated fields of view, solving the "blind spot" problem commonly found in industrial sites.

[0111] In most industrial conveyor belt deployments, the overlap rate of the fields of view of adjacent cameras is generally low, or even non-overlapping. Existing methods based on spatial overlap regions fail completely in these scenarios. This invention is specifically designed for industrial environments where the fields of view are completely separated. Through a temporal gating mechanism and motion behavior consistency analysis, it can achieve accurate target association across cameras under conditions where there is no overlap in the fields of view, filling a key technological gap in the field of industrial visual counting. Practice has proven that even when the distance between cameras is large and there is no overlap in the fields of view, this invention can still maintain excellent association performance.

[0112] 3) High computational efficiency, meeting industrial real-time requirements.

[0113] The core algorithm of this invention has a complexity of O(n·m), where n and m are the number of trajectories in the two cameras, respectively. It adopts a lightweight temporal filtering and fast matching strategy, without involving complex 3D reconstruction and high-dimensional feature comparison. On standard industrial edge computing devices, a single cross-camera association calculation can be completed in a very short time, supporting the real-time processing requirements of high frame rate. Compared with appearance feature re-identification methods based on deep learning, this invention has significantly improved computational efficiency and is more suitable for deployment in resource-constrained industrial sites.

[0114] 4) Highly adaptable, able to cope with complex changes in dynamic industrial environments.

[0115] In industrial conveyor belt scenarios, the movement of materials is complex and ever-changing: conveyor belt speed fluctuations, irregular material shapes, temporary obstructions, and drastic changes in lighting. This invention, through an adaptive time window, uncertainty quantification, and hierarchical processing mechanism, can dynamically adapt to these changes. When the conveyor belt speed changes significantly, this invention automatically adjusts the time window parameters, and the counting accuracy remains stable. In contrast, the performance of methods based on fixed time thresholds fluctuates significantly. This adaptive capability significantly improves the robustness in real industrial environments.

[0116] 5) Modular architecture design, which can be used as the underlying algorithm module of any counting system.

[0117] This invention adopts a loosely coupled architecture design, requiring only standardized trajectory input (position sequence, timestamp, displacement and velocity change sequence, etc.), without relying on specific sensing technologies or system architectures. This design enables the invention to be seamlessly integrated into existing counting systems without requiring large-scale modifications to existing systems. Actual deployment cases show that integrating this invention into counting systems from different vendors significantly reduces integration time, while traditional 3D reconstruction-based methods typically require a longer cycle. This "plug-and-play" characteristic greatly reduces the technical upgrade threshold and enterprise transformation costs.

[0118] 6) Quantifying and interpreting uncertainty to improve credibility.

[0119] This invention introduces a matching confidence quantification and uncertainty labeling mechanism, which can clearly distinguish between "definite counts" and "suspicious counts", providing a basis for decision-making for upper-level systems. When it is impossible to determine whether they are the same target, uncertainty labels and associated metadata are generated to avoid the accumulation of errors caused by blind counting. This transparent design greatly improves interpretability and credibility. In practical applications, it significantly reduces the workload of manual review and increases users' trust in the output results.

[0120] 7) Clear technological boundaries and positioning

[0121] This invention focuses on the underlying algorithm level in terms of technical positioning, solving the core problem of cross-camera target association. Its "pure temporal sequence + motion behavior" technical approach is in stark contrast to existing methods that rely on spatial information, establishing a unique and irreplaceable technical advantage. This clear technical boundary provides a clear value proposition for industrial applications.

[0122] In summary, this invention successfully solves the core challenges in the field of industrial visual counting, demonstrating significant advantages in terms of technological advancement, ease of deployment, and environmental adaptability, and providing key technical support for the upgrading of industrial automation and intelligence. Attached Figure Description

[0123] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.

[0124] Figure 1 This is a schematic diagram of the process of the present invention;

[0125] Figure 2 This is a flowchart of the deduplication counting and uncertainty marking process in this invention. Detailed Implementation

[0126] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0127] The core of this invention lies in the fact that it does not rely on 3D reconstruction, spatial calibration, or field-of-view overlap, but only on the temporal relationship and motion behavior characteristics of the target appearing in different cameras to achieve accurate alignment and deduplication of targets across cameras. This invention can be used as the underlying algorithm module of any counting system and is suitable for industrial field environments with separated camera fields of view and limited installation conditions.

[0128] The following describes the specific process of the cross-camera target alignment and deduplication counting method based on temporal consistency provided by this invention, using a concrete example (e.g.) Figure 1 (as shown) and technical effects.

[0129] S1. Multi-camera time reference establishment: Acquire video streams from at least two cameras and establish a unified time reference system, including:

[0130] Acquire video streams from at least two cameras and establish a unified time reference system using NTP protocol, GPS timestamps, hardware synchronization signals, or software synchronization mechanisms;

[0131] Assign a globally unique frame number to each video frame and establish a mapping relationship between the frame number and physical time;

[0132] Among them, a clock drift compensation mechanism is introduced: a sliding window algorithm is used to monitor and compensate for clock drift between cameras in real time;

[0133] Introducing a time calibration offset mechanism: Allowing clock errors within ±100ms range, improving the system's adaptability in real industrial environments.

[0134] S2. Single-camera target detection and trajectory generation: Target detection and tracking algorithms are executed independently within each camera, and a local trajectory is generated for each target, including:

[0135] S21. Execute the target detection and tracking algorithm independently within each camera and assign a locally unique ID to each target to maintain the temporal continuity of the target within a single camera;

[0136] S22. Generate a local trajectory for each target, including:

[0137] The sequence of the target's position (normalized coordinates) within the image plane;

[0138] Timestamps of the target entering / leaving the Region of Interest (ROI);

[0139] The inter-frame displacement and velocity change sequence of the target;

[0140] The sequence of changes in the confidence level of the target;

[0141] S23. Lightweight time-series filtering algorithms (such as first-order low-pass filtering) are used to smooth the trajectory and reduce noise interference in a single frame.

[0142] S3. Generation of Candidate Trajectory Pair Set: Based on the number of consecutive frames the target is tracked within a single camera and the time windows in which the trajectory appears in different cameras, candidate trajectory pairs are filtered to generate a candidate trajectory pair set, including:

[0143] S31. For all trajectories within cameras A and B, cross-camera matching is only allowed if the target is continuously tracked for more than a preset frame threshold (e.g., 5 frames) within a single camera.

[0144] S32, Regarding the trajectory T within camera A A Based on the time t that the target leaves camera A A_exit Predict the entry time t of the target within camera B. A_enter And based on time constraints, candidate trajectories are selected from the trajectories within camera B:

[0145] ;

[0146] in, This is the time window threshold, with the initial value calculated based on the nominal speed of the conveyor belt and dynamically adjusted based on historical matching results.

[0147] S33. Generate a set of candidate trajectory pairs. ;

[0148] Among them, T B1 T B2 These are the first and second candidate trajectories within camera B, respectively.

[0149] Optionally, in the technical solution of this application, the generation of the candidate trajectory pair set in S3 involves: filtering candidate trajectory pairs based on the number of consecutive frames the target is tracked within a single camera and the time windows in which the trajectory appears within different cameras, and generating the candidate trajectory pair set, which includes:

[0150] Appearance verification: Based on appearance consistency, candidate trajectory pairs with obvious mismatches are eliminated (this step does not affect the operation of the core algorithm and can be turned off in resource-constrained environments), specifically including:

[0151] 1) Extract low-dimensional appearance features of the target, including size change trends (width / height ratio changes over time), color statistics (dominant color distribution in HSV space), and shape invariant moments (low-order moments such as Hu moments).

[0152] 2) Calculate the feature distance d:

[0153] ;

[0154] Among them, f A f B These are the appearance feature vectors of the targets within cameras A and B, respectively. This indicates the calculation of the L2 norm;

[0155] 3) Convert the feature distance d into an appearance consistency verification score f. score :

[0156] ;

[0157] Among them, the appearance consistency verification score f score A larger value indicates a more similar appearance. d max The maximum allowable distance threshold is determined based on historical data statistics;

[0158] 4) Appearance consistency verification score f score The candidate trajectory pairs are compared with a preset lenient score threshold (to avoid false rejection under common industrial conditions such as changes in lighting and posture differences), and candidates that do not match are rejected based on the comparison results.

[0159] S4. Matching Score: Calculate a matching score for each candidate trajectory pair based on the consistency of motion behavior, including:

[0160] S41. Calculate the time continuity score f time :

[0161] ;

[0162] S42. Compare the trend of target's motion speed changes in different cameras:

[0163] S421, Calculate the standard deviation of the speed difference. :

[0164] ;

[0165] in, , are the velocity of the target in the i-th frame within cameras A and B, respectively, and n is the number of comparison frames;

[0166] S422. Calculate the motion speed consistency score f. speed :

[0167] ;

[0168] Where k is the scaling factor, used to adjust the standard deviation of the speed difference. Consistency score for movement speed f speed The extent of the impact;

[0169] S43. Compare the trend of the target's motion direction changes in different cameras:

[0170] S431. Extract the target's motion direction angle sequences dA and dB within cameras A and B:

[0171] ;

[0172] ;

[0173] in, , These are the motion direction angles of the target in the i-th frame within cameras A and B, respectively, which are the angles relative to the X-axis direction of the image coordinate system.

[0174] S432. Determine the angular difference sequence of the target's motion direction within cameras A and B. , :

[0175] ;

[0176] ;

[0177] in, Let be the difference in the motion direction angle of the target in the i-th and i+1-th frames within camera A. Let be the difference in the motion direction angle of the target in the i-th and i+1-th frames within camera B;

[0178] S433, Calculate the motion direction consistency score f direction :

[0179] ;

[0180] Where cos_sim is the similarity function;

[0181] S44. Compare the temporal characteristics of the target in different cameras, including acceleration change trends and pause patterns, and calculate the motion state consistency score f. behavior ;

[0182] S45. Calculate the matching score:

[0183] ;

[0184] in, , , , All are weighting coefficients, and .

[0185] S5. Global Matching and Conflict Resolution: Perform global optimization and conflict resolution on all candidate trajectory pairs, outputting a set of cross-camera associated trajectory pairs, including:

[0186] S51. Use the maximum weight matching algorithm (such as the Hungarian algorithm) to perform global optimization on all candidate trajectory pairs to obtain the initial global matching scheme;

[0187] S52. Utilize conflict resolution strategies to resolve conflicts in the initial global matching scheme, avoiding one-to-many and many-to-one matching patterns:

[0188] Set a minimum score threshold; candidate trajectory pairs with scores below the minimum score threshold are directly rejected.

[0189] Introducing a time consistency constraint: the trajectories of the same target in multiple cameras do not overlap in time;

[0190] If multiple candidate trajectory pairs share the same target, the candidate trajectory pair with the highest matching score is retained;

[0191] S53. Introduce a backtracking correction mechanism: When a new trajectory is added, a decision is made on whether to accept the new match based on the confidence level. At the same time, when the addition of a new trajectory causes the original match to become invalid, the historical match can be readjusted.

[0192] S54, Output a set of cross-camera associated trajectory pairs.

[0193] S6. Deduplication and Uncertainty Labeling: Using an uncertainty quantization mechanism, corresponding cross-camera associated trajectory pairs are merged into a single target, and deduplication is performed directly. Simultaneously, uncertainty labeling is applied. Figure 2 As shown, it includes:

[0194] S61. Introducing an uncertainty quantification mechanism:

[0195] S611. Calculate the matching confidence level C. match :

[0196] ;

[0197] Where max(Score) and second_max(Score) are the highest and second highest matching scores, respectively;

[0198] S612. Introduce the time window coverage factor:

[0199] ;

[0200] Wherein, min_end and max_end are the smaller and larger values ​​of the end time of the two segments of the trajectory in the cross-camera associated trajectory pair, respectively; and min_start and max_start are the smaller and larger values ​​of the start time of the two segments of the trajectory in the cross-camera associated trajectory pair, respectively.

[0201] S613, Calculate the final uncertainty:

[0202] ;

[0203] Among them, A history Historical matching accuracy measures the system's success rate in matching similar historical scenarios. , , All are weighting coefficients, and ;

[0204] S62. Based on the final uncertainty, perform hierarchical processing on all cross-camera associated trajectory pairs:

[0205] For highly deterministic cross-camera associated trajectory pairs with uncertainty < 0.3, they are merged into the same target and deduplication is performed directly.

[0206] For deterministic cross-camera associated trajectory pairs with a probability of 0.3 ≤ uncertainty < 0.7, they are marked as "requires verification" and provided to the upper-level system for further processing.

[0207] For cross-camera associated trajectory pairs with low uncertainty (uncertainty ≥ 0.7), the count is rejected directly, the anomaly is recorded, and manual review is requested.

[0208] S63. Generate metadata for the cross-camera associated trajectory pairs that "require verification". The metadata includes:

[0209] The chain of evidence that matches the decision;

[0210] Detailed scoring for each dimension;

[0211] Visual summary of candidate trajectories;

[0212] S64. Output the deduplication count results, as well as the cross-camera associated trajectories and their metadata that "need verification", for use by the upper-layer system.

[0213] The core innovations of this invention are as follows:

[0214] 1) Purely time-series driven cross-camera target association mechanism

[0215] This invention proposes a temporal gating mechanism for the first time, transforming the cross-camera target association problem into a temporal consistency verification problem. Traditional methods rely on spatial coordinate mapping or appearance feature matching, while this invention accurately captures the temporal relationship of the target's appearance / departure in different cameras and constructs an adaptive time window, fundamentally solving the problem of cross-camera target association in areas without spatial overlap. This mechanism does not rely on any camera parameter calibration, greatly reducing the threshold for industrial deployment.

[0216] 2) Motor Behavior Consistency Assessment Model

[0217] This invention innovatively proposes the concept of "motion behavior fingerprint," which extracts the motion behavior features of a target from different perspectives (movement speed change trend, movement direction change trend, acceleration change trend, pause patterns, etc.), rather than simple position or speed values, to construct a highly robust cross-camera matching and scoring system. This model can effectively distinguish between real, identical targets and coincidentally similar, non-identical targets, significantly improving matching accuracy while avoiding excessive reliance on appearance features. It performs particularly well in scenarios with homogeneous materials.

[0218] 3) Uncertainty Quantification and Hierarchical Processing Mechanism

[0219] This invention introduces an uncertainty quantification mechanism, which dynamically evaluates the reliability of the matching results by calculating the difference between the highest and second-highest matching scores, and implements a hierarchical processing strategy accordingly. This mechanism enables the algorithm to maintain high accuracy in complex environments, while providing interpretable decision-making basis for upper-level systems, avoiding information loss caused by traditional binary decision-making (match / non-match).

[0220] 4) Lightweight and pluggable algorithm architecture

[0221] This invention designs a modular and loosely coupled algorithm architecture, which can be seamlessly integrated into any counting system as an independent algorithm module. It only requires standardized trajectory input (position sequence, timestamp, displacement and velocity change sequence, etc.) to output deduplication counting results and uncertainty markers. It does not depend on specific sensing technologies or system architectures. This design greatly improves the versatility and deployability of the algorithm, making it an "infrastructure-level" technology in the field of industrial vision counting.

[0222] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A cross-camera target alignment and deduplication counting method based on temporal consistency, characterized in that: Includes the following steps: S1. Multi-camera time reference establishment: Acquire video streams from at least two cameras and establish a unified time reference system; S2. Single-camera target detection and trajectory generation: Target detection and tracking algorithms are executed independently within each camera, and a local trajectory is generated for each target; S3. Generation of candidate trajectory pair set: Based on the number of frames the target is continuously tracked in a single camera and the time window in which the trajectory appears in different cameras, candidate trajectory pairs are filtered to generate a candidate trajectory pair set. S4. Matching score: Calculate the matching score for each candidate trajectory pair based on the consistency of motion behavior; S5. Global Matching and Conflict Resolution: Perform global optimization and conflict resolution on all candidate trajectory pairs, and output a set of cross-camera associated trajectory pairs; S6. Deduplication and Uncertainty Labeling: Using an uncertainty quantification mechanism, corresponding cross-camera associated trajectory pairs are merged into the same target, and deduplication is performed directly, while uncertainty is labeled.

2. The cross-camera target alignment and deduplication counting method based on temporal consistency according to claim 1, characterized in that: Multi-camera time reference establishment in S1: Acquire video streams from at least two cameras and establish a unified time reference system, including: Acquire video streams from at least two cameras and establish a unified time reference system using NTP protocol, GPS timestamps, hardware synchronization signals, or software synchronization mechanisms; Assign a globally unique frame number to each video frame and establish a mapping relationship between the frame number and physical time; Among them, a clock drift compensation mechanism is introduced: a sliding window algorithm is used to monitor and compensate for clock drift between cameras in real time; Introducing a time calibration offset mechanism: Allowing clock errors within ±100ms range, improving the system's adaptability in real industrial environments.

3. The cross-camera target alignment and deduplication counting method based on temporal consistency according to claim 1, characterized in that: S2 Single-camera target detection and trajectory generation: Target detection and tracking algorithms are executed independently within each camera, and a local trajectory is generated for each target, including: S21. Execute the target detection and tracking algorithm independently within each camera and assign a locally unique ID to each target to maintain the temporal continuity of the target within a single camera; S22. Generate a local trajectory for each target, including: The sequence of target positions within the image plane; Timestamps of the target entering / leaving the Region of Interest (ROI); The inter-frame displacement and velocity change sequence of the target; The sequence of changes in the confidence level of the target; S23. A lightweight temporal filtering algorithm is used to smooth the trajectory and reduce noise interference in a single frame.

4. The cross-camera target alignment and deduplication counting method based on temporal consistency according to claim 1, characterized in that: Candidate trajectory pair set generation in S3: Based on the number of consecutive frames the target is tracked within a single camera and the time windows in which the trajectory appears in different cameras, candidate trajectory pairs are filtered to generate a candidate trajectory pair set, including: S31. For all trajectories within cameras A and B, cross-camera matching is only allowed when the target is continuously tracked for more than a preset frame rate within a single camera. S32, Regarding the trajectory T within camera A A Based on the target's departure time t within camera A A_exit Predict the entry time t of the target within camera B. A_enter And based on time constraints, candidate trajectories are selected from the trajectories within camera B: ; in, This is the time window threshold, with the initial value calculated based on the nominal speed of the conveyor belt and dynamically adjusted based on historical matching results. S33. Generate a set of candidate trajectory pairs. ; Among them, T B1 T B2 These are the first and second candidate trajectories within camera B, respectively.

5. The cross-camera target alignment and deduplication counting method based on temporal consistency according to claim 1, characterized in that: Candidate trajectory pair set generation in S3: Based on the number of consecutive frames the target is tracked within a single camera and the time windows in which the trajectory appears in different cameras, candidate trajectory pairs are filtered to generate a candidate trajectory pair set, which includes: Appearance verification: Based on appearance consistency, candidate trajectory pairs with obvious mismatches are eliminated, specifically including: 1) Extract low-dimensional appearance features of the target, including size variation trends, color statistics, and shape invariant moments; 2) Calculate the feature distance d: ; Among them, f A f B These are the appearance feature vectors of the targets within cameras A and B, respectively. This indicates the calculation of the L2 norm; 3) Convert the feature distance d into an appearance consistency verification score f. score : ; Among them, the appearance consistency verification score f score A larger value indicates a more similar appearance. d max The maximum allowable distance threshold is determined based on historical data statistics; 4) Appearance consistency verification score f score The candidate trajectory pairs are compared with a preset lenient score threshold, and those that do not match are removed based on the comparison results.

6. The cross-camera target alignment and deduplication counting method based on temporal consistency according to claim 1, characterized in that: Matching score in S4: A matching score is calculated for each candidate trajectory pair based on the consistency of motion behavior, including: S41. Calculate the time continuity score f time : ; S42. Compare the trend of target's motion speed changes in different cameras: S421, Calculate the standard deviation of the speed difference. : ; in, , are the velocity of the target in the i-th frame within cameras A and B, respectively, and n is the number of comparison frames; S422. Calculate the motion speed consistency score f. speed : ; Where k is the scaling factor, used to adjust the standard deviation of the speed difference. Consistency score for movement speed f speed The extent of the impact; S43. Compare the trend of the target's motion direction changes in different cameras: S431. Extract the target's motion direction angle sequences dA and dB within cameras A and B: ; ; in, , These are the motion direction angles of the target in the i-th frame within cameras A and B, respectively, which are the angles relative to the X-axis direction of the image coordinate system. S432. Determine the angular difference sequence of the target's motion direction within cameras A and B. , : ; ; in, Let be the difference in the motion direction angle of the target in the i-th and i+1-th frames within camera A. Let be the difference in the motion direction angle of the target in the i-th and i+1-th frames within camera B; S433, Calculate the motion direction consistency score f direction : ; Where cos_sim is the similarity function; S44. Compare the temporal characteristics of the target in different cameras, including acceleration change trends and pause patterns, and calculate the motion state consistency score f. behavior ; S45. Calculate the matching score: ; in, , , , All are weighting coefficients, and .

7. The cross-camera target alignment and deduplication counting method based on temporal consistency according to claim 1, characterized in that: Global matching and conflict resolution in S5: Global optimization and conflict resolution are performed on all candidate trajectory pairs, outputting a set of cross-camera associated trajectory pairs, including: S51. The maximum weight matching algorithm is used to perform global optimization on all candidate trajectory pairs to obtain an initial global matching scheme. S52. Utilize conflict resolution strategies to resolve conflicts in the initial global matching scheme, avoiding one-to-many and many-to-one matching patterns: Set a minimum score threshold; candidate trajectory pairs with scores below the minimum score threshold are directly rejected. Introducing a time consistency constraint: the trajectories of the same target in multiple cameras do not overlap in time; If multiple candidate trajectory pairs share the same target, the candidate trajectory pair with the highest matching score is retained; S53. Introduce a backtracking correction mechanism: When a new trajectory is added, a decision is made on whether to accept the new match based on the confidence level. At the same time, when the addition of a new trajectory causes the original match to become invalid, the historical match can be readjusted. S54, Output a set of cross-camera associated trajectory pairs.

8. The cross-camera target alignment and deduplication counting method based on temporal consistency according to claim 1, characterized in that: In S6, deduplication and uncertainty labeling are performed: An uncertainty quantization mechanism is used to merge corresponding cross-camera associated trajectory pairs into a single target, directly performing deduplication and uncertainty labeling, including: S61. Introducing an uncertainty quantification mechanism: S611. Calculate the matching confidence level C. match : ; Where max(Score) and second_max(Score) are the highest and second highest matching scores, respectively; S612. Introduce the time window coverage factor: ; Wherein, min_end and max_end are the smaller and larger values ​​of the end time of the two segments of the trajectory in the cross-camera associated trajectory pair, respectively; and min_start and max_start are the smaller and larger values ​​of the start time of the two segments of the trajectory in the cross-camera associated trajectory pair, respectively. S613, Calculate the final uncertainty: ; Among them, A history Historical matching accuracy measures the system's success rate in matching similar historical scenarios. , , All are weighting coefficients, and ; S62. Based on the final uncertainty, perform hierarchical processing on all cross-camera associated trajectory pairs: For highly deterministic cross-camera associated trajectory pairs with uncertainty < 0.3, they are merged into the same target and deduplication is performed directly. For deterministic cross-camera associated trajectory pairs with a probability of 0.3 ≤ uncertainty < 0.7, they are marked as "requires verification" and provided to the upper-level system for further processing. For cross-camera associated trajectory pairs with low uncertainty (uncertainty ≥ 0.7), the count is rejected directly, the anomaly is recorded, and manual review is requested. S63. Generate metadata for the "verification required" cross-camera associated trajectory pairs. The metadata includes: The chain of evidence that matches the decision; Detailed scoring for each dimension; Visual summary of candidate trajectories; S64. Output the deduplication count results, as well as the cross-camera associated trajectories and their metadata that "need verification", for use by the upper-layer system.

Citation Information

Patent Citations

  • Unsupervised video sequence pedestrian re-identification method based on joint space-time sampling

    CN111914730A

  • Target tracking method and device and storage medium

    CN117474947A