A passenger flow statistics and false counting prevention method based on multi-view feature fusion motion tracking

By employing a passenger flow statistics method based on multi-view feature fusion and motion trajectory tracking, the problems of occlusion, environmental interference, and duplicate counting in complex scenarios are solved, achieving high-precision passenger flow monitoring.

CN122493529APending Publication Date: 2026-07-31SHANGHAI JUNYU DIGITAL TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI JUNYU DIGITAL TECH
Filing Date
2026-06-22
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing passenger flow statistics technologies suffer from poor adaptability to obstruction in complex scenarios, weak environmental interference resistance, large statistical errors in dense scenes, and frequent occurrences of duplicate or missed counts, making it difficult to meet the needs of high-precision passenger flow monitoring.

Method used

A multi-view feature fusion motion tracking method is adopted. Through multi-view passenger flow video stream data preprocessing, adaptive correction of device offset and view deviation, and multiple coupled feature constraints, target matching pairs are obtained and pedestrian movement trajectories are recorded. Combined with real-time passenger flow density, the passage boundary and dwell time judgment threshold are dynamically updated to achieve the prevention of miscounting statistics.

Benefits of technology

It improves the accuracy and stability of passenger flow statistics, reduces statistical errors in complex scenarios, and enhances the ability to resist interference and the reliability of results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493529A_ABST
    Figure CN122493529A_ABST
Patent Text Reader

Abstract

This invention relates to a passenger flow statistics and error prevention method based on multi-view feature fusion motion tracking. The method includes: acquiring multi-view passenger flow video stream data and obtaining a standardized video frame sequence through data preprocessing; extracting pedestrian geometric features to obtain a multi-view passenger flow geometric feature set; obtaining a multi-view alignment feature sequence by adaptively correcting device offset and viewpoint deviation; stripping redundant features such as background pseudo-contours, obstacle residues, and lighting noise to obtain fused feature data; fusing feature similarity, spatial distance, and temporal interval constraints to obtain target matching pairs; generating smooth motion trajectories by recording pedestrian dynamic motion information and splicing pedestrian motion trajectory segments; acquiring passenger flow data with independent identification; dynamically updating passage boundaries and dwell time determination thresholds; and obtaining error prevention statistics results through a multi-dimensional discrimination mechanism. This achieves passenger flow statistics and error prevention based on multi-view feature fusion motion tracking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent passenger flow monitoring technology, specifically involving a passenger flow statistics and error prevention method based on multi-view feature fusion motion tracking. Background Technology

[0002] The current passenger flow monitoring still has the following areas for improvement: Passenger flow statistics are a core foundational technology for intelligent management of public places such as smart shopping malls, transportation hubs, scenic spots, and office buildings. Accurate passenger flow data is crucial for business operation analysis, crowd control and management, site resource allocation, and safety risk early warning. Currently, mainstream passenger flow statistics methods in the industry mainly include infrared sensor counting, access control card counting, monocular video intelligent counting, and traditional single-feature target tracking counting. These methods can achieve basic personnel counting, but in complex real-world applications, they still have many technical shortcomings, with insufficient overall statistical accuracy, anti-interference capabilities, and scenario adaptability, requiring urgent improvement.

[0003] The specific shortcomings of existing traditional passenger flow statistics techniques are as follows: First, single-view acquisition suffers from limited field of view and severe occlusion issues, resulting in an extremely high rate of missed detections. Currently, most video-based passenger flow statistics solutions employ a single-camera, single-view acquisition mode. The fixed field of view of the camera has limited coverage, making it unsuitable for passenger flow monitoring needs across large spans and wide scenarios. Furthermore, during peak hours in public places, crowds are dense, with pedestrians walking side-by-side and overlapping, easily leading to limb overlap and mutual occlusion. Combined with interference from obstacles such as luggage, pillars, and walls, single-view devices cannot acquire complete pedestrian feature information, resulting in missed detections and undercounting of target pedestrians. According to industry statistics, the error of traditional single-view statistics in dense scenarios can exceed 35%, severely impacting data validity.

[0004] Secondly, traditional algorithms suffer from limited feature extraction and weak resistance to environmental interference, making them prone to miscalculations. Existing visual pedestrian flow tracking and statistical algorithms mostly extract only single, shallow visual features such as pedestrian outlines and colors, failing to integrate multi-dimensional deep feature information. In real-world applications, environmental interference is frequent, including sudden changes in lighting, flickering lights, reflections from glass curtain walls, and shifting ground shadows. Furthermore, non-pedestrian targets such as human-shaped posters, human projections, and moving objects are easily misidentified by the algorithm as real pedestrians, generating a large amount of invalid miscalculation data. In addition, traditional algorithms cannot distinguish between customers, on-site employees, and temporary visitors, further contributing to inflated pedestrian flow data and poor data accuracy.

[0005] Third, the lack of a multi-view feature fusion mechanism results in extremely poor adaptability to dense crowds. Existing technologies all rely on independent calculations and statistics from a single camera, failing to utilize image information from multiple angles and directions for feature fusion. Cameras from different perspectives can capture feature information from different postures and body parts of pedestrians, effectively compensating for the lack of features from a single perspective. However, traditional solutions completely abandon the complementary advantages of multi-view data. In high-density pedestrian traffic scenarios, they cannot accurately distinguish overlapping and clustered individual pedestrians, often resulting in problems such as multiple people being counted together and a single person being identified multiple times, leading to a significant decrease in statistical accuracy.

[0006] Fourth, the lack of long-term motion trajectory tracking and targeted anti-false counting logic leads to significant duplicate counting issues. Traditional passenger flow statistics are mostly based on single-frame instantaneous images for identification and counting, without real-time tracking, correlation, and verification of continuous pedestrian movement trajectories. This makes it impossible to identify behaviors such as pedestrians lingering in place, making short-distance turns, or repeatedly entering and exiting the monitoring area, easily resulting in the same pedestrian being counted repeatedly. Furthermore, existing technologies lack a robust false counting filtering mechanism, failing to effectively filter stationary individuals, rapidly passing non-target individuals, and false or distracting images, making accurate deduplication and error correction difficult and unable to meet the application requirements of refined passenger flow statistics.

[0007] In summary, current passenger flow statistics technologies generally suffer from poor adaptability to occlusion, weak environmental resistance to interference, large statistical errors in densely populated areas, frequent double counting and undercounting, and insufficient generalization ability, making them unsuitable for the high-precision passenger flow monitoring needs of various complex public places. Therefore, there is an urgent need to design a passenger flow statistics and miscount prevention method based on multi-view feature fusion and motion tracking. This method, through complementary fusion of multi-view features combined with long-term motion trajectory tracking, can address the various pain points of traditional technologies, such as miscounting and undercounting, and improve the accuracy and stability of passenger flow statistics. Summary of the Invention

[0008] To address the aforementioned problems in the existing technology, this invention provides a method for passenger flow statistics and error prevention based on multi-view feature fusion motion tracking; The objective of this invention can be achieved through the following technical solutions: S1: Acquire multi-view passenger flow video stream data, obtain standardized video frame sequences through data preprocessing; segment densely connected and slightly moving pedestrian foreground targets, combine sub-pixel contour detection and invariant geometric moment operation to extract the geometric features of pedestrians, and obtain a multi-view passenger flow geometric feature set; S2: Based on the multi-view passenger flow geometric feature set, the device offset and view deviation are adaptively corrected to obtain the multi-view alignment feature sequence; by stripping noise point features, the fused feature data is obtained. S3: Based on the fused feature data, target matching pairs are obtained through multiple coupled feature constraints; pedestrian dynamic motion information is recorded, pedestrian motion trajectory segments are spliced ​​to generate smooth motion trajectories; passenger flow data is obtained by distinguishing pedestrian trajectories and binding unique identifiers; S4: Based on the passenger flow data and real-time passenger flow density, dynamically update the passage boundary and dwell time determination threshold; obtain the error prevention statistical results through a multi-dimensional discrimination mechanism.

[0009] As a preferred technical solution of the present invention, the specific process of acquiring multi-view passenger flow video stream data includes: The left, right, and front multi-directional data acquisition terminals of the preset passenger flow monitoring area are pre-processed for corresponding equipment installation and calibration, pixel parameter calibration and hardware timing. The PTP time synchronization protocol is used to perform clock synchronization calibration on the acquisition terminal to unify the sampling timing reference of the equipment; Collect continuous real-time video frames, synchronously bind the timestamp of the video frame to the device space identifier, and obtain multi-view passenger flow video stream data.

[0010] Specifically, the data preprocessing process includes: Based on the multi-view passenger flow video stream data, perspective deviation detection is performed to locate areas of image distortion and pixel offset. Adaptive local Gamma layering correction is used to perform layered light and shadow equalization processing on the bright, backlit and shadow areas of the image to obtain a standardized video frame sequence.

[0011] Specifically, the process of segmenting densely clustered and slightly moving pedestrian foreground targets includes: Perform pixel difference calculation on three consecutive frames to extract the dynamic moving pixel region of the image and filter the dynamic target region of pedestrians. Coupled with a dynamic background update algorithm, the static background model of the scene is iteratively updated in real time to adapt to background shifts caused by changes in light and shadow; By dynamically adjusting the foreground segmentation threshold, the foreground and background are distinguished between densely clustered and slightly moving pedestrians, and the foreground target frame image of densely clustered and slightly moving pedestrians is segmented.

[0012] Specifically, the process of extracting the geometric features of pedestrians includes: Extract continuous contour lines of pedestrian targets, and calculate the basic geometric parameters of pedestrian contour curvature and target size range by using contour curvature and bounding rectangle scale; Calculate the centroid coordinates, body moment of inertia, and proportional moment parameters of the pedestrian target, and extract body proportion features and edge gradient features; Integrate multi-dimensional geometric features for feature classification, aggregation, and encapsulation to obtain a multi-view passenger flow geometric feature set.

[0013] Specifically, the process of adaptively correcting device offset and viewing angle deviation includes: Collect real-time installation pose, lens shake, and angle pitch offset parameters of the corresponding equipment, and construct a real-time equipment deviation parameter matrix; The dynamic epipolar constraint model is invoked to iteratively revise the cross-view epipolar matching equation and correct the feature matching bias. The mapping relationship of multi-view feature space is dynamically calibrated through a real-time iterative update mechanism of homography matrix.

[0014] Specifically, the process of obtaining the fused feature data includes: The gradient change difference of surrounding pixels is statistically analyzed for each feature point, and a hierarchical adaptive neighborhood threshold purification strategy is adopted to peel off the false contours of the background, the residual contours of static obstacles, and the noise features of light and shadow distortion in layers. Distinguish the differences in pedestrian contour curvature and body proportion gradient to obtain fused feature data.

[0015] Specifically, the process of obtaining the target matching pair includes: Extract three core parameters: single-frame pedestrian feature vector, pixel spatial coordinates, and frame timestamp. Construct a triple-constraint matching model based on feature similarity, spatial Euclidean distance, and temporal interval, and simultaneously set multi-level matching confidence thresholds. By performing cross-view pedestrian feature matching operations, associated targets with multiple reset information thresholds are obtained.

[0016] Specifically, the process of splicing and generating pedestrian movement trajectory segments includes: Analyze the pixel displacement, instantaneous motion rate and travel angle of the pedestrian target to establish a single-frame motion parameter ledger; Perform inter-frame motion parameter association matching according to the temporal sequence, and connect the same pedestrian motion node in adjacent frames; Fill in the short-term motion gaps between frames and stitch them together to form fragmented pedestrian trajectory segments.

[0017] Specifically, the process of distinguishing pedestrian trajectories includes: Extract the core features of the corresponding trajectory. The core features include at least: starting coordinates, motion trend, velocity range, and time span. Trajectories with similar motion patterns are grouped and classified. Through a motion feature uniqueness verification mechanism, the independence of trajectories in scenarios of trajectory overlap, intersection, and adhesion is verified, thus distinguishing the independent trajectories of pedestrians.

[0018] Specifically, the process of dynamically updating the passage boundary and dwell time determination threshold includes: Real-time statistical monitoring of pedestrian density, average pedestrian speed, and regional congestion status within the area; Using pedestrian density and passage speed as the core correction factors, the pixel threshold of the area passage boundary is iteratively optimized. By combining regional congestion status with pedestrian loitering behavior characteristics, the threshold for determining pedestrian loitering time is dynamically adjusted to distinguish between normal passage and loitering behavior. Real-time fixed and iteratively updated boundary thresholds and dwell thresholds adapt to changes in corresponding passenger flow scenarios.

[0019] Specifically, the process of obtaining the error prevention statistics includes: The system determines pedestrian behavior across the entire area, identifying behaviors such as entering, exiting, crossing, turning back, and loitering, and generates statistical data on passenger flow across the entire area. By using a multi-dimensional trajectory discrimination model, abnormal targets such as densely clustered targets, occluded frame breaks, in-situ reversals, and interference from resident personnel are identified, resulting in recounting, omissions, and false counts. Based on the unique trajectory ID of pedestrians, full lifecycle correction is performed. At the same time, passenger flow data is iteratively calibrated in real time and archived and solidified at regular intervals to correct statistical deviations in dynamic scenarios and obtain intelligent passenger flow statistics results to prevent errors.

[0020] The beneficial effects of this invention are as follows: By building a full-domain, multi-view video acquisition system and adopting the PTP precision time synchronization protocol to achieve millisecond-level clock synchronization, the asynchronous sampling deviation of multiple devices is eliminated. At the same time, relying on the technical design of adaptive correction of device offset and viewpoint deviation and layered stripping of redundant features, the background pseudo-contour, obstacle residue and light and shadow noise interference are eliminated. The accuracy foundation is solidified from the data acquisition and feature purification stage, making the extraction and fusion of multi-view passenger flow features more stable and accurate.

[0021] The system innovatively integrates feature similarity, spatial distance, and temporal interval constraints to achieve target matching. It generates smooth pedestrian trajectories by recording dynamic motion information and splicing trajectory fragments, and then binds a globally unique identifier to each independent trajectory. This efficiently solves the problem of pedestrian tracking in complex scenarios such as dense clustering, occlusion, frame breaks, and trajectory intersections, enabling accurate differentiation and full-cycle tracking of pedestrian targets and reducing statistical errors caused by target confusion.

[0022] Based on real-time pedestrian density and traffic status, the passage boundaries and dwell time thresholds are dynamically updated. Combined with a multi-dimensional discrimination mechanism, error prevention statistics are completed. It can effectively identify abnormal issues such as double counting, omissions, and false counting, and distinguish between normal passage, staying in place, and turning back. This allows the passenger flow statistics results to adapt to complex scenarios with different levels of congestion and different lighting conditions, greatly improving the anti-interference ability and reliability of passenger flow statistics. Attached Figure Description

[0023] To facilitate understanding by those skilled in the art, the present invention will be further described below with reference to the accompanying drawings.

[0024] Figure 1 This is a flowchart illustrating a passenger flow statistics and miscounting prevention method based on multi-view feature fusion motion tracking according to the present invention. Figure 2 This is a structural block diagram of target matching and trajectory generation in this invention. Detailed Implementation

[0025] To further illustrate the technical means and effects of the present invention in achieving the intended purpose, the following detailed description of the specific implementation methods, structures, features and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided.

[0026] Please see Figure 1-2 A method for passenger flow statistics and error prevention based on multi-view feature fusion and motion tracking includes: S1: Acquire multi-view passenger flow video stream data, obtain standardized video frame sequences through data preprocessing; segment densely connected and slightly moving pedestrian foreground targets, combine sub-pixel contour detection and invariant geometric moment operation to extract the geometric features of pedestrians, and obtain a multi-view passenger flow geometric feature set; S2: Based on the multi-view passenger flow geometric feature set, the device offset and view deviation are adaptively corrected to obtain the multi-view alignment feature sequence; by stripping noise point features, the fused feature data is obtained. S3: Based on the fused feature data, target matching pairs are obtained through multiple coupled feature constraints; pedestrian dynamic motion information is recorded, pedestrian motion trajectory segments are spliced ​​to generate smooth motion trajectories; passenger flow data is obtained by distinguishing pedestrian trajectories and binding unique identifiers; S4: Based on the passenger flow data and real-time passenger flow density, dynamically update the passage boundary and dwell time determination threshold; obtain the error prevention statistical results through a multi-dimensional discrimination mechanism.

[0027] As a preferred technical solution of the present invention, the specific process of acquiring multi-view passenger flow video stream data includes: The left, right, and front multi-directional data acquisition terminals of the preset passenger flow monitoring area are pre-processed for corresponding equipment installation and calibration, pixel parameter calibration and hardware timing. The PTP time synchronization protocol is used to perform clock synchronization calibration on the acquisition terminal to unify the sampling timing reference of the equipment; Collect continuous real-time video frames, synchronously bind the timestamp of the video frame to the device space identifier, and obtain multi-view passenger flow video stream data.

[0028] In this embodiment, the accurate acquisition of multi-view passenger flow video stream data is mainly used to solve the problems of large blind spots, asynchronous timing of multi-device acquisition, and inconsistent image parameters in traditional single-view acquisition, which lead to subsequent pedestrian matching errors and statistical benchmark failures. For example, in core passenger flow monitoring areas such as mall entrances and exits, passages, and turnstiles, high-definition acquisition terminals are deployed and installed in the left, right, and front directions to cover the entire passage area without monitoring blind spots. After the equipment is installed, unified equipment calibration, pixel parameter calibration, and hardware timing preprocessing are carried out to correct inherent pixel deviations of the lenses and installation tilt errors, and to unify the basic parameters of image resolution and frame rate of all terminals, thus building a standardized multi-view acquisition hardware system. On this basis, the traditional independent timing mode of devices is abandoned, and a high-precision PTP timing protocol is used to perform millisecond-level clock synchronization calibration on all acquisition terminals, unifying the sampling timing benchmark of all devices, completely eliminating the core problems of multi-device frame sampling misalignment and timing asynchrony, and ensuring the consistency of the time dimension of multi-view images. After timing calibration is completed, all terminals are controlled to synchronously acquire continuous real-time video frames within the area. During the acquisition process, a high-precision timestamp and a unique spatial identifier for each frame are precisely bound, clearly distinguishing the acquired video data from different perspectives and at different times. Through an integrated acquisition mechanism of hardware calibration, timing synchronization, and spatiotemporal tag binding, the integrity, temporal consistency, and spatial correspondence of passenger flow video data are guaranteed from the source.

[0029] Specifically, the data preprocessing process includes: Based on the multi-view passenger flow video stream data, perspective deviation detection is performed to locate areas of image distortion and pixel offset. Adaptive local Gamma layering correction is used to perform layered light and shadow equalization processing on the bright, backlit and shadow areas of the image to obtain a standardized video frame sequence.

[0030] In this embodiment, based on the collected multi-view passenger flow video stream data, lens perspective deviation and radial distortion are detected frame by frame. A pixel comparison algorithm is used to accurately locate distortion areas, pixel offset areas, and distortion areas at the edges of the image, pinpointing the location of all image defects. Basic distortion correction is performed on the detected image distortions and pixel deviations to restore the true spatial form of the image. For complex lighting scenes with alternating outdoor strong light, indoor backlight, and corridor shadows, an adaptive local Gamma layered correction technique is adopted. This technique overcomes the limitations of traditional global correction, performing layered and zoned light and shadow equalization processing on bright areas, backlit dark areas, and shadow-occluded areas. It dynamically adjusts the contrast and brightness of different areas, effectively weakening the interference of extreme lighting on pedestrian targets and maximizing the preservation of pedestrian outline details. After light and shadow equalization processing, lightweight filtering is applied to residual minor noise and pixel artifacts, optimizing image noise reduction while preserving pedestrian edge features. The size, resolution, and aspect ratio of all video frames are unified, data normalization is completed, and a standardized video frame sequence with uniform specifications, clear image quality, no distortion, and balanced lighting is output. Through a multi-level preprocessing workflow including distortion correction, layered light and shadow equalization, and standardization, the video quality in complex scenes is significantly improved, ensuring the accuracy of subsequent pedestrian segmentation and feature extraction.

[0031] Specifically, the process of segmenting densely clustered and slightly moving pedestrian foreground targets includes: Perform pixel difference calculation on three consecutive frames to extract the dynamic moving pixel region of the image and filter the dynamic target region of pedestrians. Coupled with a dynamic background update algorithm, the static background model of the scene is iteratively updated in real time to adapt to background shifts caused by changes in light and shadow; By dynamically adjusting the foreground segmentation threshold, the foreground and background are distinguished between densely clustered and slightly moving pedestrians, and the foreground target frame image of densely clustered and slightly moving pedestrians is segmented.

[0032] In this embodiment, a preprocessed standardized video frame sequence is retrieved, and three consecutive valid frames are selected for pixel difference calculation. By analyzing the changes in pixel grayscale differences between frames, dynamic pixel regions with displacement changes are accurately extracted from the image, initially screening dynamic target regions of pedestrians and eliminating static background regions with no changes in the image. To adapt to background shift issues caused by gradual changes in scene lighting and minor device vibrations, a dynamic background update algorithm is coupled, iteratively updating the static background model of the scene in real time at a fixed frame rate. This dynamically adapts to background parameter shifts caused by changes in ambient lighting and minor device movements, avoiding foreground missegmentation caused by background solidification. Supported by the dynamic background model, an adaptive threshold adjustment mechanism is adopted to dynamically adjust the foreground segmentation threshold according to the real-time density of pedestrian flow. For scenarios prone to missed segmentation, such as densely packed crowds, pedestrians moving slightly in place, and slow-moving pedestrians, pixel-level accurate distinction between foreground and background is achieved. Scattered interference pixels, minor pseudo-contours, and invalid noise are removed, and false foreground targets are filtered out, resulting in complete foreground target frame images of densely packed pedestrians and slightly moving pedestrians.

[0033] Specifically, the process of extracting the geometric features of pedestrians includes: Extract continuous contour lines of pedestrian targets, and calculate the basic geometric parameters of pedestrian contour curvature and target size range by using contour curvature and bounding rectangle scale; Calculate the centroid coordinates, body moment of inertia, and proportional moment parameters of the pedestrian target, and extract body proportion features and edge gradient features; Integrate multi-dimensional geometric features for feature classification, aggregation, and encapsulation to obtain a multi-view passenger flow geometric feature set.

[0034] In this embodiment, based on the segmented, clean pedestrian foreground target frame image, a contour detection method is used to extract the continuous edge contour lines of individual pedestrian targets, completely preserving the details of the pedestrian's body contour and avoiding contour breaks and missing parts. Based on the extracted complete pedestrian contour lines, contour curvature calculation and circumscribed rectangle scale calculation are performed to statistically analyze the basic geometric parameters such as the contour curvature, target length and width, and body size range of each pedestrian, constructing a basic contour feature system. On this basis, an invariant geometric moment calculation mechanism is introduced to calculate pixel moment values ​​for the pedestrian target region, solving for the core parameters of pedestrian centroid coordinates, body inertia moment, and proportional moment. Relying on the advantage of geometric moment invariance, highly stable body proportion features and edge gradient features that are not affected by shooting angle, image scaling, or slight deformation are extracted. The contour features, scale features, body features, and gradient features of all pedestrians from a single viewpoint are integrated to complete the classification, collection, regularization, and structured encapsulation of multi-dimensional geometric features, unify the feature data format and storage specifications, and obtain a multi-view pedestrian flow geometric feature set that is dimensionally complete, highly stable, and has good anti-interference properties.

[0035] Specifically, the process of adaptively correcting device offset and viewing angle deviation includes: Collect real-time installation pose, lens shake, and angle pitch offset parameters of the corresponding equipment, and construct a real-time equipment deviation parameter matrix; The dynamic epipolar constraint model is invoked to iteratively revise the cross-view epipolar matching equation and correct the feature matching bias. The mapping relationship of multi-view feature space is dynamically calibrated through a real-time iterative update mechanism of homography matrix.

[0036] In this embodiment, based on the extracted multi-view passenger flow geometric feature set, the installation pose deviation, lens micro-shake amplitude, and viewing angle pitch and yaw offset parameters of each acquisition terminal are collected in real time frame by frame. All equipment deviation data are integrated to construct a real-time dynamic equipment deviation parameter matrix, recording the equipment state error for each frame. Based on the constructed deviation parameter matrix, a dynamic epipolar constraint model is invoked, abandoning the rigid mode of traditional fixed epipolar matching. The cross-view epipolar matching equation is iteratively corrected according to the real-time equipment deviation, dynamically correcting the epipolar correspondence. After completing the epipolar constraint correction, a real-time iterative update mechanism for the homography matrix is ​​initiated. The multi-view feature space mapping relationship is dynamically calibrated according to real-time viewing angle changes, achieving unified transformation of feature spaces for different shooting angles and installation positions, ensuring that multi-view features are in the same spatial coordinate system. Through a multi-level dynamic correction mechanism, pixel-level registration and temporal alignment of multi-view geometric features are achieved, eliminating cross-view spatial misalignment and temporal disconnection problems.

[0037] Specifically, the process of obtaining the fused feature data includes: The gradient change difference of surrounding pixels is statistically analyzed for each feature point, and a hierarchical adaptive neighborhood threshold purification strategy is adopted to peel off the false contours of the background, the residual contours of static obstacles, and the noise features of light and shadow distortion in layers. Distinguish the differences in pedestrian contour curvature and body proportion gradient to obtain fused feature data.

[0038] In this embodiment, based on the multi-view feature sequence completed by spatiotemporal alignment, the gradient change difference of surrounding pixels is statistically analyzed for each feature point. The effective feature points of pedestrians are distinguished from background interference feature points based on pixel gradient differences, completing the initial feature identification. On this basis, a hierarchical adaptive neighborhood threshold purification strategy is adopted to perform refined stripping operations on interference features at different levels. Various redundant interference information, such as false outlines of the background, residual outlines of static obstacles, and light and shadow distortion noise, are removed layer by layer, filtering out invalid impurity features and retaining the core effective features of pedestrians. After redundant feature stripping, to address the problem of insufficient homogeneity and differentiation in the features of densely clustered pedestrians, the difference in contour curvature and body proportion gradient between different pedestrians is adaptively amplified to enhance the feature uniqueness of individual pedestrians, increase the feature distinguishability of densely clustered pedestrians, and avoid confusion and overlap of features between adjacent pedestrians. The purified and enhanced feature data is then regularized, removing extreme abnormal features and supplementing locally missing features, completing feature structure optimization, and obtaining low-redundancy, high-discrimination, and highly differentiated multi-view fused feature data.

[0039] Specifically, the process of obtaining the target matching pair includes: Extract three core parameters: single-frame pedestrian feature vector, pixel spatial coordinates, and frame timestamp. Construct a triple-constraint matching model based on feature similarity, spatial Euclidean distance, and temporal interval, and simultaneously set multi-level matching confidence thresholds. By performing cross-view pedestrian feature matching operations, associated targets with multiple reset information thresholds are obtained.

[0040] In this embodiment, based on the purified high-purity fused feature data, three core matching parameters are extracted frame by frame: single-frame pedestrian feature vectors, pixel spatial coordinates, and frame timestamps. These three parameters are integrated to construct a multi-dimensional matching basic parameter set, covering the three major matching dimensions of features, space, and time. Based on this basic parameter set, a triple-constraint matching model is built, fusing feature similarity, spatial Euclidean distance, and temporal interval. This breaks the limitations of traditional single-dimensional matching. Simultaneously, multi-level refined matching confidence thresholds are set to distinguish between high, medium, and low matching confidence levels, filtering out low-quality matching results. After the model is built, cross-frame and cross-viewpoint pedestrian feature matching operations are performed on continuous video frames. Simultaneously, feature similarity, spatial distance, and temporal interval rationality are combined for comprehensive verification, selecting highly correlated target combinations that meet the multi-level confidence threshold constraints. After matching, the initial matching results are verified a second time to eliminate interfering targets such as low-confidence matches, cross-pedestrian mismatches, and temporally disordered matches, retaining valid matching combinations to obtain highly accurate, unambiguous, and temporally compliant inter-frame target matching pairs.

[0041] Specifically, the process of splicing and generating pedestrian movement trajectory segments includes: Analyze the pixel displacement, instantaneous motion rate and travel angle of the pedestrian target to establish a single-frame motion parameter ledger; Perform inter-frame motion parameter association matching according to the temporal sequence, and connect the same pedestrian motion node in adjacent frames; Fill in the short-term motion gaps between frames and stitch them together to form fragmented pedestrian trajectory segments.

[0042] In this embodiment, based on the acquired high-confidence target matching pairs, the dynamic motion parameters such as pixel displacement, instantaneous motion rate, and travel angle of each matched pedestrian target are analyzed frame by frame. This comprehensively records the motion state information of the pedestrian in each frame, establishing a standardized single-frame motion parameter ledger to achieve traceable and associative motion data. After completing the single-frame parameter recording, inter-frame motion parameter association matching is performed according to the video's temporal sequence, connecting the motion nodes of the same pedestrian in adjacent frames one by one to construct a continuous trajectory node chain. Addressing the issue of missing inter-frame motion data caused by short-term pedestrian occlusion or slight image stuttering, the system reasonably fills in the short-term motion gaps between frames based on the changing patterns of motion parameters between frames, repairing trajectory breakpoint defects and initially splicing together fragmented pedestrian trajectory segments with higher continuity. All fragmented trajectories are then regularized, eliminating invalid trajectory segments that are too short, lack effective displacement, or exhibit static pseudo-motion, retaining valid trajectory segments with genuine pedestrian movement characteristics, thus completing the initial trajectory screening and optimization.

[0043] Specifically, the process of distinguishing pedestrian trajectories includes: Extract the core features of the corresponding trajectory. The core features include at least: starting coordinates, motion trend, velocity range, and time span. Trajectories with similar motion patterns are grouped and classified. Through a motion feature uniqueness verification mechanism, the independence of trajectories in scenarios of trajectory overlap, intersection, and adhesion is verified, thus distinguishing the independent trajectories of pedestrians.

[0044] In this embodiment, based on the fragmented pedestrian motion trajectory segments assembled from the network, core feature parameters of each independent trajectory are extracted in batches, including: trajectory start coordinates, overall motion trend, instantaneous speed range, and time span duration, constructing a complete trajectory feature dataset to provide data support for trajectory differentiation and classification. Based on the trajectory feature dataset, a temporal clustering classification mechanism is adopted to aggregate and integrate trajectories with similar motion patterns and similar spatiotemporal travel intervals, while trajectories with significant differences in motion trends, speeds, and start and end positions are split and differentiated, initially completing the classification and aggregation of different pedestrian trajectories. Addressing the complex situations of trajectory overlap, intersection, and adhesion that frequently occur in passenger flow scenarios, a motion feature uniqueness verification mechanism is activated. The independence of each trajectory is verified through trajectory detail parameters and motion change patterns, accurately distinguishing different pedestrian trajectories that overlap or adhere, and preventing misjudgments due to trajectory mixing. After completing the independent differentiation of trajectories, a globally unique dynamic ID is assigned to each verified independent pedestrian trajectory, achieving precise binding of a single pedestrian and a single identifier, and obtaining full-domain passenger flow trajectory data with an independent identity identifier.

[0045] Specifically, the process of dynamically updating the passage boundary and dwell time determination threshold includes: Real-time statistical monitoring of pedestrian density, average pedestrian speed, and regional congestion status within the area; Using pedestrian density and passage speed as the core correction factors, the pixel threshold of the area passage boundary is iteratively optimized. By combining regional congestion status with pedestrian loitering behavior characteristics, the threshold for determining pedestrian loitering time is dynamically adjusted to distinguish between normal passage and loitering behavior. Real-time fixed and iteratively updated boundary thresholds and dwell thresholds adapt to changes in corresponding passenger flow scenarios.

[0046] In this embodiment, based on the uniquely identified full-area passenger flow trajectory data, three core scenario parameters—passenger flow distribution density, average pedestrian speed, and regional congestion status—are monitored and statistically analyzed in real time within the area. A real-time scenario dynamic parameter set is dynamically constructed to reflect the current passenger flow status. Based on the scenario parameter set, an adaptive threshold iterative calculation model is established, setting passenger flow density and pedestrian speed as core correction factors. The model iteratively optimizes the boundary pixel thresholds of the monitored area in real time, dynamically adjusting the passage judgment range to adapt to different passenger flow densities. Simultaneously, combining regional congestion status and pedestrian trajectory dwelling behavior characteristics, the effective dwell time judgment threshold for pedestrians is dynamically adjusted to distinguish between different behaviors such as fast passage, slow crossing, lingering in place, and brief pauses, avoiding misjudgments caused by fixed thresholds. After each iteration optimization, the updated boundary thresholds and dwell time thresholds are fixed in real time, continuously adapting to different scenario changes such as peak-hour dense passenger flow, off-peak sparse passenger flow, and congestion.

[0047] Specifically, the process of obtaining the error prevention statistics includes: The system determines pedestrian behavior across the entire area, identifying behaviors such as entering, exiting, crossing, turning back, and loitering, and generates statistical data on passenger flow across the entire area. By using a multi-dimensional trajectory discrimination model, abnormal targets such as densely clustered targets, occluded frame breaks, in-situ reversals, and interference from resident personnel are identified, resulting in recounting, omissions, and false counts. Based on the unique trajectory ID of pedestrians, full lifecycle correction is performed. At the same time, passenger flow data is iteratively calibrated in real time and archived and solidified at regular intervals to correct statistical deviations in dynamic scenarios and obtain intelligent passenger flow statistics results to prevent errors.

[0048] In this embodiment, based on dynamically updated passage boundaries and dwell time thresholds, comprehensive passage behavior determination is performed on all pedestrian trajectories with unique IDs across the entire area. Four core behaviors are identified: pedestrians entering and exiting areas, walking in straight lines, turning back in place, and lingering at fixed points. Preliminary overall passenger flow statistics are generated. To eliminate statistical errors in complex scenarios, a multi-dimensional trajectory discrimination model is built, incorporating trajectory path, movement speed, dwell time, and behavioral trends. The initial statistical results undergo secondary verification, filtering out various abnormal statistical targets such as misjudgments of densely packed crowds, missed counts due to occlusion and frame breaks, recounting of pedestrians turning back in place, and false counts of staff stationed on-site. This forms a passenger flow dataset to be calibrated. Based on the unique pedestrian trajectory IDs, full lifecycle correction processing is performed on the trajectories, systematically removing duplicate turning trajectories, supplementing occluded and broken track targets, filtering out invalid lingering trajectories, and eliminating interfering trajectories, comprehensively correcting various statistical biases. Equipped with a real-time iterative calibration and timed archiving mechanism at the edge, it dynamically corrects the deviation of statistical parameters under dynamic scenarios, and timed solidifies and archives passenger flow data to obtain intelligent anti-miscount passenger flow statistics results with high accuracy, zero redundancy, and ultra-low miscount rate, realizing the closed-loop optimization of the entire process of multi-view tracking passenger flow statistics.

[0049] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.

Claims

1. A method for passenger flow statistics and error prevention based on multi-view feature fusion and motion tracking, characterized in that, include: S1: Acquire multi-view passenger flow video stream data and obtain a standardized video frame sequence through data preprocessing; By segmenting densely clustered and slightly moving pedestrian foreground targets, and combining sub-pixel contour detection and invariant geometric moment calculation, the geometric features of pedestrians are extracted to obtain a multi-view passenger flow geometric feature set; S2: Based on the multi-view passenger flow geometric feature set, the device offset and view deviation are adaptively corrected to obtain the multi-view alignment feature sequence; By stripping noise point features, fused feature data is obtained; S3: Based on the fused feature data, target matching pairs are obtained through multiple coupled feature constraints; pedestrian dynamic motion information is recorded, pedestrian motion trajectory segments are spliced ​​to generate smooth motion trajectories; passenger flow data is obtained by distinguishing pedestrian trajectories and binding unique identifiers; S4: Based on the passenger flow data and real-time passenger flow density, dynamically update the passage boundary and dwell time determination threshold; By using a multi-dimensional discrimination mechanism, we can obtain statistical results to prevent errors.

2. The method according to claim 1, characterized in that, The specific process of acquiring multi-view passenger flow video stream data includes: The left, right, and front multi-directional data acquisition terminals of the preset passenger flow monitoring area are pre-processed for corresponding equipment installation and calibration, pixel parameter calibration and hardware timing. The PTP time synchronization protocol is used to perform clock synchronization calibration on the acquisition terminal to unify the sampling timing reference of the equipment; Collect continuous real-time video frames, synchronously bind the timestamp of the video frame to the device space identifier, and obtain multi-view passenger flow video stream data.

3. The method according to claim 1, characterized in that, The specific process of data preprocessing includes: Based on the multi-view passenger flow video stream data, perspective deviation detection is performed to locate areas of image distortion and pixel offset. Adaptive local Gamma layering correction is used to perform layered light and shadow equalization processing on the bright, backlit and shadow areas of the image to obtain a standardized video frame sequence.

4. The method according to claim 1, characterized in that, The specific process of segmenting densely clustered and slightly moving pedestrian foreground targets includes: Perform pixel difference calculation on three consecutive frames to extract the dynamic moving pixel region of the image and filter the dynamic target region of pedestrians. Coupled with a dynamic background update algorithm, the static background model of the scene is iteratively updated in real time to adapt to background shifts caused by changes in light and shadow; By dynamically adjusting the foreground segmentation threshold, the foreground and background are distinguished between densely clustered and slightly moving pedestrians, and the foreground target frame image of densely clustered and slightly moving pedestrians is segmented.

5. The method according to claim 1, characterized in that, The specific process for extracting the geometric features of pedestrians includes: Extract continuous contour lines of pedestrian targets, and calculate the basic geometric parameters of pedestrian contour curvature and target size range by using contour curvature and bounding rectangle scale; Calculate the centroid coordinates, body moment of inertia, and proportional moment parameters of the pedestrian target, and extract body proportion features and edge gradient features; Integrate multi-dimensional geometric features for feature classification, aggregation, and encapsulation to obtain a multi-view passenger flow geometric feature set.

6. The method according to claim 1, characterized in that, The specific process of adaptively correcting device offset and viewing angle deviation includes: Collect real-time installation pose, lens shake, and angle pitch offset parameters of the corresponding equipment, and construct a real-time equipment deviation parameter matrix; The dynamic epipolar constraint model is invoked to iteratively revise the cross-view epipolar matching equation and correct the feature matching bias. The mapping relationship of multi-view feature space is dynamically calibrated through a real-time iterative update mechanism of homography matrix.

7. The method according to claim 1, characterized in that, The specific process for obtaining the fused feature data includes: The gradient change difference of surrounding pixels is statistically analyzed for each feature point, and a hierarchical adaptive neighborhood threshold purification strategy is adopted to peel off the false contours of the background, the residual contours of static obstacles, and the noise features of light and shadow distortion in layers. Distinguish the differences in pedestrian contour curvature and body proportion gradient to obtain fused feature data.

8. The method according to claim 1, characterized in that, The specific process for obtaining the target matching pair includes: Extract three core parameters: single-frame pedestrian feature vector, pixel spatial coordinates, and frame timestamp. Construct a triple-constraint matching model based on feature similarity, spatial Euclidean distance, and temporal interval, and simultaneously set multi-level matching confidence thresholds. By performing cross-view pedestrian feature matching operations, associated targets with multiple reset information thresholds are obtained.

9. The method according to claim 1, characterized in that, The specific process of splicing together to generate pedestrian movement trajectory segments includes: Analyze the pixel displacement, instantaneous motion rate and travel angle of the pedestrian target to establish a single-frame motion parameter ledger; Perform inter-frame motion parameter association matching according to the temporal sequence, and connect the same pedestrian motion node in adjacent frames; Fill in the short-term motion gaps between frames and stitch them together to form fragmented pedestrian trajectory segments.

10. The method according to claim 1, characterized in that, The specific process of distinguishing pedestrian trajectories includes: Extract the core features of the corresponding trajectory. The core features include at least: starting coordinates, motion trend, velocity range, and time span. Trajectories with similar motion patterns are grouped and classified. Through a motion feature uniqueness verification mechanism, the independence of trajectories in scenarios of trajectory overlap, intersection, and adhesion is verified, thus distinguishing the independent trajectories of pedestrians.

11. The method according to claim 1, characterized in that, The specific process for dynamically updating the passage boundary and dwell time determination threshold includes: Real-time statistical monitoring of pedestrian density, average pedestrian speed, and regional congestion status within the area; Using pedestrian density and passage speed as the core correction factors, the pixel threshold of the area passage boundary is iteratively optimized. By combining regional congestion status with pedestrian loitering behavior characteristics, the threshold for determining pedestrian loitering time is dynamically adjusted to distinguish between normal passage and loitering behavior. Real-time fixed and iteratively updated boundary thresholds and dwell thresholds adapt to changes in corresponding passenger flow scenarios.

12. The method according to claim 1, characterized in that, The specific process for obtaining the error prevention statistics includes: The system determines pedestrian behavior across the entire area, identifying behaviors such as entering, exiting, crossing, turning back, and loitering, and generates statistical data on passenger flow across the entire area. By using a multi-dimensional trajectory discrimination model, abnormal targets such as densely clustered targets, occluded frame breaks, in-situ reversals, and interference from resident personnel are identified, resulting in recounting, omissions, and false counts. Based on the unique trajectory ID of pedestrians, full lifecycle correction is performed. At the same time, passenger flow data is iteratively calibrated in real time and archived and solidified at regular intervals to correct statistical deviations in dynamic scenarios and obtain intelligent passenger flow statistics results to prevent errors.