Intelligent bus shelter passenger behavior analysis system
Through multimodal data acquisition and processing technology, combined with high-definition cameras, millimeter-wave radars and temperature and humidity sensors, the passenger behavior analysis problem of bus shed monitoring system in an open environment is solved, accurate identification and real-time management of passenger behavior is achieved, and the safety and operational efficiency of bus sheds are improved.
Patent Information
- Application Number
- CN202510583339.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-08-01
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing bus shed monitoring system lacks intelligent analysis capabilities and cannot detect passenger abnormal behaviors in a timely manner. In an open environment, multimodal interference and high model complexity cannot meet the real-time reasoning requirements at the edge and lacks adaptability to changes in the dynamic environment.
High-definition camera, millimeter-wave radar and temperature and humidity sensors are used to collect multimodal data, and space-time synchronization and fusion are performed through Retinex illumination equalization and data preprocessing, combined with Kalman filtering and attention mechanism, and behavioral analysis is performed using the improved YOLOv8n target detection network and Softmax multi-label classifier to generate density level signals and trigger hierarchical warnings.
It realizes accurate analysis of passenger behavior in complex environments, improves the robustness and accuracy of the system, provides rich data support, and ensures the safety and management efficiency of bus shelters.
Smart Images

Figure CN120412097A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of artificial intelligence (AI) and computer vision technology, and more specifically, to an intelligent bus shelter passenger behavior analysis system. Background Art
[0002] In the modern urban public transportation system, as an important place for passengers to wait for transportation, bus shelters face many challenges. With the acceleration of urbanization and the continuous growth of public transportation demand, the passenger flow in bus shelters is increasing day by day, and passenger behavior has become more complex and diverse.
[0003] However, most of the existing bus shelter monitoring systems only have basic video monitoring functions and lack the ability to intelligently analyze passenger behavior. This results in the inability to detect abnormal passenger behaviors, such as crowding, staying, running, etc., in special situations such as peak hours or holidays in a timely manner, which may lead to potential safety hazards and affect the normal operation of public transportation.
[0004] In addition, although the existing subway platform behavior recognition technology (Patent CN118015701A) can solve the problem of passenger behavior recognition in high-density passenger flow scenarios, it faces the following challenges in the open bus shelter scenario:
[0005] 1. Multi-modal interference: Interference factors in the open environment such as moving vehicles and billboard reflections are not effectively suppressed;
[0006] 2. High model complexity: It cannot meet the real-time inference requirements of the edge side;
[0007] 3. Limitations of static models: Lack of adaptability to dynamic environmental changes such as weather and time periods.
[0008] Therefore, an intelligent bus shelter passenger behavior analysis system is provided. Summary of the Invention
[0009] To solve the above technical problems, this application is proposed.
[0010] Specifically, according to one aspect of this application, an intelligent bus shelter passenger behavior analysis system is provided, which includes:
[0011] A multi-modal data acquisition module for collecting panoramic video streams of the bus shelter, radar point cloud data, and environmental parameters through a high-definition camera, a millimeter-wave radar, and a temperature and humidity sensor respectively, where the environmental parameters include temperature, humidity, and light intensity;
[0012] A data preprocessing module for performing Retinex illumination equalization and resolution compression on the video stream, coordinate transformation and dynamic filtering on the radar point cloud data, and dynamic calibration on the environmental parameters;
[0013] A spatio-temporal synchronization and fusion module, which is used to align the spatio-temporal information of video frames and radar point clouds, and fuse video spatial features, radar temporal features, and environmental embedding features based on an attention mechanism to obtain fused features;
[0014] A target detection module, which is used to capture the positions of passengers in the video frames after spatio-temporal synchronization through an improved YOLOv8n target detection network to generate detection boxes;
[0015] A cross-modal tracking module, which is used to bind the detection boxes and radar point clouds through the Hungarian algorithm, and combine environmental parameters to generate detection boxes with unique identity IDs;
[0016] A behavior classification module, which is used to perform feature recognition on the detection boxes with unique identity IDs and fused features using a multi-label classifier based on Softmax to obtain the final behavior labels;
[0017] A density analysis module, which is used to calculate the regional density based on the detection boxes with unique identity IDs and environmental parameters to obtain density level signals;
[0018] An analysis result output module, which is used to display abnormal behavior annotations, density heatmaps, and environmental parameters through a real-time visualization interface, and trigger hierarchical warnings based on the density level and abnormal behavior types.
[0019] Preferably, the data preprocessing module includes: performing illumination equalization on each frame image in the panoramic video stream of the bus stop through the Retinex algorithm, and the specific formula is: L enhanced = log(I) - log(F * I), where I represents the image and F represents the spatial frequency filter; downsampling the 4K video to 1080p; converting the polar coordinates (ρ, θ, v) of the radar point cloud to Cartesian coordinates (x, y, z), and using the Savitzky-Golay filter to remove velocity outliers; performing temporal smoothing on temperature, humidity, and light intensity to eliminate sensor noise.
[0020] Preferably, the spatio-temporal synchronization and fusion module includes: performing spatio-temporal alignment of video frames in the waiting area video stream and radar point clouds through Kalman filtering; fusing the spatial features of video frames, the temporal features of radar point clouds, and the embedding features of environmental parameters based on the attention mechanism, and the specific formula is expressed as:
[0021] X fusion = αCNN video + βTransformer radar + γMLP env
[0022] where, CNN videoRepresents the spatial features of video frames, Transformer radar Represents the temporal modeling features of radar point clouds, MLP env Represents the embedded features of environmental parameters, and α, β, γ represent modal weight coefficients.
[0023] Preferably, the improved YOLOv8n object detection network includes: replacing the backbone network with MobileNetV3-Small; introducing the Atrous Spatial Pyramid Pooling module ASPP to capture multi-scale context information through atrous convolutions with different sampling rates; modeling spatio-temporal feature correlations through the multi-head self-attention mechanism Transformer; and retaining the Fast Spatial Pyramid Pooling module SPPF of YOLOv8.
[0024] Compared with the prior art, an intelligent bus stop passenger behavior analysis system provided by the present application has the following technical effects:
[0025] 1. By fusing multi-modal data from high-definition cameras, millimeter-wave radars, and temperature and humidity sensors, more comprehensive passenger behavior and environmental information can be captured, especially improving the robustness and accuracy of the system in complex environments, providing rich data support for bus stop management and optimization;
[0026] 2. Through Kalman filtering for spatio-temporal alignment and multi-modal feature fusion based on the attention mechanism, the consistency of different modal data in space and time and the dynamic highlighting of important features are achieved, improving the quality of feature representation, better capturing the key information of passenger behavior, and providing accurate and comprehensive data support for behavior analysis;
[0027] 3. Through the improved YOLOv8n object detection network (replacing the backbone network, introducing ASPP and Transformer), multi-scale context information capture and spatio-temporal feature correlation modeling are achieved, improving the accuracy and efficiency of object detection, and providing reliable detection results for tracking and behavior analysis;
[0028] 4. By adopting a multi-label classifier based on Softmax and a multi-task loss function, the multi-label classification problem is effectively solved, improving the accuracy and robustness of behavior recognition, supporting the simultaneous recognition and marking of multiple behaviors, and providing accurate behavior analysis results for bus stop safety management and operation optimization;
[0029] 5. By introducing weather impact factors and time period weights for regional density calculation, the accurate reflection of regional density under different environments and time periods is achieved, providing a scientific basis for bus stop management and scheduling, and helping to optimize the operation efficiency of bus stops and the passenger experience;
[0030] 6. Display abnormal behavior annotations, density heat maps, and environmental parameters through a real-time visualization interface, trigger hierarchical early warnings based on density levels and abnormal behavior types, achieve an intuitive display of the real-time situation inside the bus shelter and timely detection of abnormal situations, adopt corresponding response strategies through the hierarchical early warning mechanism, improve the safety and management efficiency of the bus shelter, and provide data support for long-term management and optimization. Description of the Drawings
[0031] The embodiments of the present application will be described in more detail by combining the accompanying drawings. The above and other objects, features, and advantages of the present application will become more obvious. The accompanying drawings are used to provide a further understanding of the embodiments of the present application, and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application, and do not constitute a limitation to the present application. In the accompanying drawings, the same reference numerals generally represent the same components or steps.
[0032] Figure 1 The system block diagram according to an embodiment of the present application is illustrated. Detailed Embodiments
[0033] Next, the embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. It should be understood that the present application is not limited by the example embodiments described herein.
[0034] Embodiment:
[0035] Figure 1 The system block diagram according to an embodiment of the present application is illustrated. As Figure 1As shown in the figure, the intelligent waiting shelter passenger behavior analysis system 100 according to the embodiments of the present application includes: a multi-modal data acquisition module 101, which is used to collect the panoramic video stream of the waiting shelter, radar point cloud data, and environmental parameters through a high-definition camera, millimeter-wave radar, and temperature and humidity sensor respectively. The environmental parameters include temperature, humidity, and light intensity; a data preprocessing module 102, which is used to perform Retinex illumination equalization and resolution compression on the video stream, perform coordinate transformation and dynamic filtering on the radar point cloud data, and perform dynamic calibration on the environmental parameters; a spatio-temporal synchronization and fusion module 103, which is used to align the spatio-temporal information of the video frame and the radar point cloud, and fuse the video spatial feature, radar temporal feature, and environmental embedding feature based on the attention mechanism to obtain a fusion feature; a target detection module 104, which is used to capture the passenger position in the video frame after spatio-temporal synchronization through an improved YOLOv8n target detection network to generate a detection box; a cross-modal tracking module 105, which is used to bind the detection box and the radar point cloud through the Hungarian algorithm, and combine the environmental parameters to generate a detection box with a unique identity ID; a behavior classification module 106, which is used to perform feature recognition on the detection box with a unique identity ID and the fusion feature using a multi-label classifier based on Softmax to obtain the final behavior label; a density analysis module 107, which is used to calculate the regional density based on the detection box with a unique identity ID and the environmental parameters to obtain a density level signal; an analysis result output module 108, which is used to display abnormal behavior annotations, density heat maps, and environmental parameters through a real-time visualization interface, and trigger hierarchical warnings according to the density level and abnormal behavior type.
[0036] In the embodiments of the present application, the multi-modal data acquisition module 101 is used to collect the panoramic video stream of the waiting shelter, radar point cloud data, and environmental parameters through a high-definition camera, millimeter-wave radar, and temperature and humidity sensor respectively. The environmental parameters include temperature, humidity, and light intensity. It should be understood that the video stream of the high-definition camera can provide intuitive visual information for directly observing the behavior and distribution of passengers; the radar point cloud data of the millimeter-wave radar can supplement the deficiency of visual data. Especially in complex environments (such as low light and bad weather), it can accurately detect and track moving targets and provide information such as the speed, distance, and angle of objects; while the environmental parameters collected by the temperature and humidity sensor reflect the actual environmental conditions in the waiting shelter. These environmental factors may affect the behavior and comfort of passengers. For example, in a high-temperature or high-humidity environment, the behavior pattern of passengers may be different from that in a normal environment. By integrating these multi-modal data, the system can analyze passenger behavior more accurately, discover abnormal situations in time, and provide data support for optimizing the environment and management of the waiting shelter, thereby improving the safety and efficiency of the public transportation system. Therefore, in the embodiments of the present application, the multi-modal sensors are deployed first.
[0037] Specifically, the deployment of the multi-modal sensor array is as follows:
[0038] High-definition camera: Equipped with a 4K ultra-high-definition camera, with starlight-level night vision function, a frame rate of not less than 30fps, and a resolution reaching or exceeding 3840×2160 pixels. Through the high-definition camera, the panoramic video stream of the bus shelter is collected to ensure clear capture of the dynamic situations inside the bus shelter under various lighting conditions.
[0039] Millimeter-wave radar: Install a millimeter-wave radar with a detection range of not less than 150 meters and an angular resolution of not more than 1°. In this way, through the millimeter-wave radar, radar point cloud data is obtained to accurately detect the positions, speeds, and movement directions of passengers and objects inside the bus shelter, and maintain high-precision detection capabilities even in complex environments.
[0040] Temperature and humidity sensor: Deploy temperature and humidity sensors with accuracies of ±2%RH (relative humidity) and ±0.3℃ respectively. The temperature and humidity sensors can real-time monitor environmental parameters such as temperature, humidity, and light intensity inside the bus shelter, providing comprehensive environmental information for the system to better analyze passenger behavior and optimize the waiting environment.
[0041] In the embodiment of this application, the data preprocessing module 102 is used to perform Retinex illumination equalization and resolution compression on the video stream, perform coordinate transformation and dynamic filtering on the radar point cloud data, and perform dynamic calibration on the environmental parameters. It should be understood that since the collected original video stream may have problems such as uneven illumination and high resolution resulting in large amounts of data, Retinex illumination equalization and resolution compression are performed on the video stream to improve image quality and processing efficiency. At the same time, the coordinate system of the radar point cloud data may be inconsistent with the video stream, and there is noise and dynamic interference, so coordinate transformation and dynamic filtering are required to ensure the accuracy and consistency of the data. In addition, the environmental parameter sensors may be affected by environmental changes, resulting in measurement data deviations, so dynamic calibration is required to ensure the accuracy and reliability of the data. Through these preprocessing steps, the data quality can be significantly improved, providing more accurate and reliable basic data for subsequent feature analysis and behavior recognition, thereby ensuring the performance and accuracy of the entire system.
[0042] Specifically, the data preprocessing module 102 includes:
[0043] 1), Perform illumination equalization on each frame image in the panoramic video stream of the bus shelter through the Retinex algorithm. The specific formula is: L enhanced = log(I) - log(F*I), where I represents the image and F represents the spatial frequency filter. Then, downsample the 4K video to 1080p (1920×1080 pixels).
[0044] 2) Convert the polar coordinates (ρ, θ, v) of the radar point cloud to Cartesian coordinates (x, y, z), and use the Savitzky-Golay filter to remove velocity outliers (such as velocity mutations or data outside the reasonable range).
[0045] 3) Perform temporal smoothing (such as moving average filtering) on temperature, humidity, and light intensity to eliminate sensor noise.
[0046] In the embodiment of the present application, the spatio-temporal synchronization and fusion module 103 is used to align the spatio-temporal information of the video frame and the radar point cloud, and fuse the video spatial features, radar temporal features, and environmental embedding features based on the attention mechanism to obtain fused features. Considering that the video stream and radar point cloud data come from different sensors, their coordinate systems and timestamps may not match, and spatio-temporal alignment is required to ensure the consistency of the data in time and space. At the same time, data of different modalities have different features and advantages. For example, the video stream has rich visual information, the radar point cloud data is excellent in detection distance and angle, and the environmental parameters provide background information. Through the fusion based on the attention mechanism, important features can be dynamically highlighted, noise and irrelevant information can be suppressed, thereby improving the quality of feature representation. This fusion method can better capture the key information of passenger behavior and provide more accurate and comprehensive data support for subsequent behavior analysis.
[0047] Specifically, the spatio-temporal synchronization and fusion module 103 includes:
[0048] 1. Timestamp alignment: Let the timestamp of the video frame be t v , and the timestamp of the radar point cloud be t r . Through Kalman filtering, the time difference Δt = t v - t r between the video frame and the radar point cloud can be estimated, and the timestamps are adjusted accordingly to make them consistent in time.
[0049] 2. Spatial position alignment: Let the passenger position in the video frame be (x v , y v ), and the passenger position in the radar point cloud be (x r , y r ). Through Kalman filtering, the spatial transformation relationship between the video frame and the radar point cloud, such as the translation vector (dx, dy) and the rotation angle θ, can be estimated. According to these transformation relationships, the passenger position in the radar point cloud can be converted into the coordinate system of the video frame to achieve spatial position alignment.
[0050] 3. Based on the attention mechanism, fuse the spatial features of video frames, the temporal features of radar point clouds, and the embedded features of environmental parameters. The specific implementation steps are as follows:
[0051] 3.1 Extract the spatial features of video frames through a convolutional neural network (CNN), extract the temporal features of radar point clouds through a Transformer, and extract the embedded features of environmental parameters through a multi-layer perceptron (MLP);
[0052] 3.2 Calculate the weight of each feature through the attention mechanism. Among them, let the weight of the spatial features of video frames be α, the weight of the temporal features of radar point clouds be β, and the weight of the embedded features of environmental parameters be γ;
[0053] 3.3 According to the calculated weights, the features of different modalities can be weighted and fused. The specific formula is expressed as:
[0054] X fusion = αCNN video + βTransformer radar + γMLP env
[0055] where CNN video represents the spatial features of video frames, Transformer radar represents the temporal modeling features of radar point clouds, MLP env represents the embedded features of environmental parameters, and α, β, and γ represent the modality weight coefficients.
[0056] In the embodiment of the present application, the target detection module 104 is used to capture the positions of passengers in the video frames after spatio-temporal synchronization through an improved YOLOv8n target detection network to generate detection frames. It should be understood that in order to accurately identify and locate the passengers in the waiting pavilion, the positions of passengers in the video frames after spatio-temporal synchronization are captured through an improved YOLOv8n target detection network and detection frames are generated, so as to monitor the distribution and dynamic changes of passengers in real time, timely detect abnormal behaviors (such as crowding, staying, running, etc.), and thus ensure the safety of passengers, optimize the management efficiency of the waiting pavilion, and improve the overall operation effect of the public transportation system.
[0057] Specifically, the improved YOLOv8n object detection network is as follows: the backbone network is replaced by MobileNetV3-Small; the Atrous Spatial Pyramid Pooling module ASPP is introduced, and atrous convolutions with different sampling rates (such as 6×6, 12×12, 18×18) are used to capture multi-scale context information; the spatio-temporal feature correlation is modeled through the multi-head self-attention mechanism Transformer, where the number of heads is set to 8 and the hidden layer dimension is 512; meanwhile, the Fast Spatial Pyramid Pooling module SPPF of YOLOv8 is retained.
[0058] In the embodiment of the present application, the cross-modal tracking module 105 is used to bind the detection box and the radar point cloud through the Hungarian algorithm, and generate a detection box with a unique identity ID in combination with environmental parameters. It should be understood that in view of the problems of inconsistent and inaccurate matching between the passenger positions (i.e., detection boxes) detected in the video frames and the passenger positions identified in the radar point cloud in the multi-modal data sources, the Hungarian algorithm is introduced in the embodiment of the present application to achieve the best matching and association between the detection box and the radar point cloud. Through the application of this algorithm, not only can the matching problem during multi-modal data fusion be effectively processed, but also the tracking accuracy and the stability of the system can be improved. In addition, by taking into account the environmental parameter information, the tracking accuracy can be further improved, a detection box with a unique identifier can be generated, providing more reliable and accurate data support for subsequent passenger behavior analysis, thereby enhancing the perception and understanding ability of the intelligent waiting pavilion system in dealing with complex scenarios.
[0059] Specifically, the specific implementation steps of the cross-modal tracking module 105 are as follows:
[0060] Extract the motion features and environmental features in the radar point cloud;
[0061] Bind the detection box and the radar point cloud through the Hungarian algorithm to minimize the following cost matrix: C ij =||CNN(b i )-Transformer(p j )|| 2 +α||v i -v j || 2 +β||MLP env (e i )-MLP env (e j )|| 2
[0062] where, CNN(b i ) represents the spatial features of the video detection box, Transformer(p j ) represents the temporal features of the radar point cloud, v i 、vj represents the target speed, e i 、e j represents the environmental parameter embedding vector, and α, β represent weight coefficients;
[0063] According to the minimized cost matching result, a unique identity ID is assigned to each detection box.
[0064] In the embodiment of the present application, the behavior classification module 106 is used to perform feature recognition on the detection box with a unique identity ID and the fusion feature by using a multi-label classifier based on Softmax to obtain the final behavior label. Considering that in the multi-modal data source, the behavior of passengers may involve multiple labels, and there may be correlations between different labels. Therefore, in the embodiment of the present application, a multi-label classifier based on Softmax is used to perform feature recognition on the detection box with a unique identity ID and the fusion feature. In this way, the multi-label classification problem can be effectively solved, and the accuracy and robustness of behavior recognition are improved.
[0065] Specifically, the behavior classification module 106 includes:
[0066] 1. Output an 8-class behavior probability vector P:
[0067]
[0068] 2. Determine the final behavior category: It is determined by setting a probability threshold (for example, P>0.5), and multi-label classification is supported, such as simultaneously recognizing and labeling as "staying" and "crowded".
[0069] 3. Define a multi-task loss function: Loss = L cls +λL reg +γL triplet where L cls represents the cross-entropy classification loss, L reg represents the bounding box regression loss, and L triplet represents the cross-modal triplet loss.
[0070] In the embodiment of the present application, the density analysis module 107 is used to calculate the regional density based on the detection box with a unique identity ID and the environmental parameters to obtain a density level signal. Specifically, the specific implementation steps of the density analysis module 107 are as follows:
[0071] 1. Perform density calculation:
[0072] where N t represents the number of passengers detected at the current moment, A represents the area of the monitoring region, and f weather represents the weather influence factor (such as f weatherWhen it is 1.5, it indicates rainy days, f weather When it is 1, it indicates sunny days), f time Indicates the time period weight (such as f time When it is 1.2, it indicates the peak period, f time When it is 1, it indicates the off-peak).
[0073] 2. Density classification: According to D t Divide the density levels (for example: D t < 0.5 person / m2 is low density, 0.5 ≤ D t < 1.0 is medium density, D t ≥ 1.0 is high density), and obtain the density level signal.
[0074] In the embodiment of the present application, the analysis result output module 108 is used to display abnormal behavior annotations, density heat maps and environmental parameters through a real-time visualization interface, and trigger hierarchical warnings according to the density level and the type of abnormal behavior.
[0075] Specifically, the analysis result output module 108 includes:
[0076] Real-time visualization interface: Real-time annotate abnormal behaviors in the video stream based on behavior tags (such as using a red frame + flashing prompt for "running"); Generate a heat map of the waiting shed area based on the density level signal (low / medium / high), and intuitively display the degree of crowd density in different areas through color coding (green indicates low density, yellow indicates medium density, and red indicates high density); Set up an environmental parameter panel to display the current temperature, humidity, light intensity and weather impact factors in real time.
[0077] Hierarchical warning prompt interface: Set different response strategies according to the density level and the type of abnormal behavior. In the low-density area, only log records are made and no alarms are triggered; In the medium-density area, notifications (such as text messages or APP messages) are pushed to the administrator to prompt attention to the flow trend; In the high-density area, audible and visual alarms are triggered, the broadcast system is linked to guide evacuation, and the traffic dispatching center is notified to dispatch additional vehicles. For abnormal behaviors, such as single-person anomalies (such as running), local audible and visual reminders are made and recorded in the background, while group anomalies (such as crowding, staying) are synchronously pushed to the public security department.
[0078] In terms of data storage and report generation: Structurally store data such as the number of passengers, behavior tags, density levels, and environmental parameters according to timestamps, supporting SQL or NoSQL databases. At the same time, automatically generate daily or weekly reports, count the peak period, the frequency of abnormal events, and the impact of the environment on behaviors, providing data support for the long-term management and optimization of the waiting shed.
[0079] In summary, the intelligent bus stop passenger behavior analysis system according to the embodiments of the present application is elucidated. It collects the bus stop video stream, radar point cloud and environmental parameters through high-definition cameras, millimeter-wave radars and temperature and humidity sensors and preprocesses them; aligns the spatio-temporal information of video frames and radar point clouds, and fuses features based on the attention mechanism; uses the improved YOLOv8n network to detect the passenger positions and generate detection frames; binds the detection frames with the radar point cloud through the Hungarian algorithm, combines the environmental parameters to generate detection frames with unique identity IDs, and uses the Softmax multi-label classifier to identify behavior labels; calculates the regional density to obtain density level signals; displays abnormal behaviors, density heat maps and environmental parameters through a real-time visualization interface, and triggers hierarchical early warnings according to the density level and abnormal behaviors. The present application realizes the accurate analysis and real-time monitoring of the behaviors of passengers at the bus stop, and improves the safety and management efficiency.
[0080] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms.
[0081] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit of the technical solutions of the present invention.
Claims
1. An intelligent waiting shed passenger behavior analysis system, characterized in that Including: A multi-modal data acquisition module, which is used to collect panoramic video streams of the waiting shelter, radar point cloud data, and environmental parameters through a high-definition camera, a millimeter-wave radar, and a temperature and humidity sensor respectively. The environmental parameters include temperature, humidity, and light intensity; A data preprocessing module, which is used to perform Retinex illumination equalization and resolution compression on the video stream, perform coordinate transformation and dynamic filtering on the radar point cloud data, and perform dynamic calibration on the environmental parameters; A spatio-temporal synchronization and fusion module, which is used to align the spatio-temporal information of video frames and radar point clouds, and fuse video spatial features, radar temporal features, and environmental embedding features based on an attention mechanism to obtain fused features; A target detection module, which is used to capture the positions of passengers in the video frames after spatio-temporal synchronization through an improved YOLOv8n target detection network to generate detection boxes; A cross-modal tracking module, which is used to bind the detection boxes and radar point clouds through the Hungarian algorithm, and combine environmental parameters to generate detection boxes with unique identity IDs; A behavior classification module, which is used to perform feature recognition on the detection boxes with unique identity IDs and fused features using a multi-label classifier based on Softmax to obtain the final behavior labels; A density analysis module, which is used to calculate the regional density based on the detection boxes with unique identity IDs and environmental parameters to obtain a density level signal; An analysis result output module, which is used to display abnormal behavior annotations, density heat maps, and environmental parameters through a real-time visualization interface, and trigger hierarchical warnings according to the density level and abnormal behavior types.
2. The intelligent bus shelter passenger behavior analysis system according to claim 1, wherein The data preprocessing module includes: Perform illumination equalization on each frame image in the panoramic video stream of the bus stop pavilion through the Retinex algorithm. The specific formula is: L enhanced = log(I) - log(F * I), where I represents the image and F represents the spatial frequency filter; Downsampling a 4K video to 1080p; Converting the polar coordinates (ρ, θ, v) of the radar point cloud to Cartesian coordinates (x, y, z), and using a Savitzky-Golay filter to remove velocity outliers; Performing temporal smoothing on temperature, humidity, and light intensity to eliminate sensor noise.
3. The intelligent waiting shelter passenger behavior analysis system according to claim 2, wherein, The spatio-temporal synchronization and fusion module includes: Performing spatio-temporal alignment of video frames in the waiting area video stream and radar point clouds through Kalman filtering; Fusing the spatial features of video frames, the temporal features of radar point clouds, and the embedding features of environmental parameters based on an attention mechanism. The specific formula is expressed as: X fusion = αCNN video + βTransformer radar + γMLP env Among them, CNN video represents the spatial features of video frames, and Transformer radar represents the temporal modeling features of radar point clouds, and MLP env represents the embedding features of environmental parameters, and α, β, and γ represent modal weight coefficients.
4. The intelligent waiting shelter passenger behavior analysis system according to claim 3, wherein, The improved YOLOv8n target detection network includes: Replacing the backbone network with MobileNetV3-Small; introducing an Atrous Spatial Pyramid Pooling module (ASPP), capturing multi-scale context information through dilated convolutions with different sampling rates; modeling spatio-temporal feature correlations through a multi-head self-attention mechanism Transformer; retaining the Fast Spatial Pyramid Pooling module (SPPF) of YOLOv8.
5. The intelligent waiting shelter passenger behavior analysis system according to claim 4, wherein, The cross-modal tracking module includes: Extracting motion features and environmental features from the radar point cloud; Binding the detection boxes and radar point clouds through the Hungarian algorithm to minimize the following cost matrix: C ij = ||CNN(b i ) - Transformer(p j )|| 2 + α||v i - v j || 2 + β||MLP env (e i ) - MLP env (e j )|| 2 Among them, CNN(b i ) represents the spatial features of the video detection box, and Transformer(p j ) represents the temporal features of the radar point cloud. v i , v j represent the target speed, and e i , e j represent the environmental parameter embedding vectors. α and β represent the weight coefficients; According to the minimum cost matching result, assigning a unique identity ID to each detection box.
6. The intelligent bus stop passenger behavior analysis system according to claim 5, wherein The behavior classification module further includes: Define the multi-task loss function: Loss = L cls + λL reg + γL triplet , where L cls represents the cross-entropy classification loss, L reg represents the bounding box regression loss, and L triplet represents the cross-modal triplet loss.
7. The intelligent bus shelter passenger behavior analysis system according to claim 6, characterized in that, The density analysis module includes: Among them, N t represents the number of passengers detected at the current moment, A represents the area of the monitored area, and f weather represents the weather influence factor, and f time represents the time period weight.
8. The intelligent bus shelter passenger behavior analysis system according to claim 7, wherein The analysis result output module includes: Real-time visualization interface: Real-time annotation of abnormal behaviors, generation of heat maps to show the degree of crowd density, and display of environmental parameters; Hierarchical early warning prompt interface: Set different response strategies according to density levels and types of abnormal behaviors; In terms of data storage and report generation: Structured storage of data and automatic generation of daily or weekly reports.
Citation Information
Cited By
Line inspection channel crowd behavior identification method and system based on multi-modal data fusion
CN120853224A