Public area video monitoring system and method
By combining intelligent data acquisition and recognition, feature fusion, heterogeneous trajectory prediction, and dynamic risk field, the prediction error problem caused by differences in motion characteristics and environmental changes in existing monitoring systems has been solved, achieving high-precision public safety early warning.
Patent Information
- Application Number
- CN202511102221.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-11-11
AI Technical Summary
Existing public video surveillance systems fail to effectively distinguish the different motion characteristics of pedestrians, vehicles, and suddenly moving objects. They rely solely on coordinate overlap to assess risk without considering dynamic environmental changes, resulting in significant prediction errors and poor practicality.
The system employs an intelligent acquisition and recognition module to accurately distinguish pedestrians, vehicles, and suddenly moving objects. A feature fusion module integrates video data with meteorological and traffic flow information to construct a multi-dimensional environmental state vector. A heterogeneous trajectory prediction module uses Social-LSTM, TransMotion, and random field probability models for differentiated prediction. A dynamic risk field module calculates the trajectory intersection probability and combines it with dynamic environmental factors. A graded response module implements progressive early warnings, and a collaborative warning module executes multi-level alarm strategies.
It significantly improves the accuracy and practicality of public safety early warning, solves the error problem of risk prediction caused by differences in motion characteristics and changes in environmental factors, and achieves high-precision risk assessment and progressive early warning.
Smart Images

Figure CN120935329A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent security technology, and in particular to a video surveillance system and method for public areas. Background Technology
[0002] With the continuous development of modern cities and the improvement of various infrastructures, surveillance equipment has become widely used in public places. While these surveillance devices can regulate people's behavior, they cannot prevent impending dangers.
[0003] Patent document CN117636240A discloses a public video surveillance system based on big data, comprising a video acquisition module, a scene analysis module, a first prediction module, a second prediction module, a dangerous behavior analysis module, and a dangerous behavior alarm module. The video acquisition module acquires video within its coverage area and identifies features within the video based on big data, extracting people and objects from the video. The scene analysis module analyzes the movement status of people and objects in the video within its coverage area and predicts their movement trajectories using the first and second prediction modules. Finally, the dangerous behavior analysis module and the dangerous behavior alarm module issue alarms for potentially dangerous behaviors. This technical solution predicts the movement trajectories of people and objects within the monitored area and, combined with the real-time scene within the monitored area, accurately and comprehensively analyzes and judges whether personnel behavior poses a danger.
[0004] However, existing public video surveillance systems fail to differentiate the motion characteristics of pedestrians, vehicles, and suddenly moving objects in actual use. They also rely solely on coordinate overlap to determine risk without considering the impact of dynamic environmental changes on risk probability, resulting in large prediction errors and poor practicality. Summary of the Invention
[0005] The purpose of this invention is to provide a public area video surveillance system and method, which aims to solve the technical problems of existing public video surveillance systems in practical applications, which fail to distinguish the differences in motion characteristics of pedestrians, vehicles, and suddenly moving objects, and rely solely on coordinate overlap to judge risks without considering the impact of dynamic environmental changes on risk probability, resulting in large prediction errors and poor practicality.
[0006] To achieve the above objectives, the present invention provides a public area video surveillance system, including an intelligent acquisition and recognition module, a feature fusion module, a heterogeneous trajectory prediction module, a dynamic risk field module, a hierarchical response module, and a collaborative warning module. The intelligent acquisition and recognition module is used to acquire video data of the monitored area in real time and identify pedestrians, vehicles, and suddenly moving objects. The extracted video data is transmitted to the feature fusion module via a wireless communication module. The feature fusion module fuses the video data of pedestrians, vehicles, and suddenly moving objects extracted by the intelligent acquisition and recognition module with meteorological data and traffic flow data to construct an environmental state vector.
[0007] The heterogeneous trajectory prediction module uses the Social-LSTM model, TransMotion model and random field probability model respectively to predict the motion trajectories of pedestrians, vehicles and sudden moving objects based on the environmental state vector constructed by the feature fusion module.
[0008] The dynamic risk field module is connected to the heterogeneous trajectory prediction module. Based on the prediction results of the heterogeneous trajectory prediction module on the motion trajectories of pedestrians, vehicles and suddenly moving objects, the probability of trajectory intersection is calculated, and the spatiotemporal risk intensity is quantified by combining environmental dynamic factors and object hazard levels.
[0009] The graded response module obtains the corresponding risk level based on the spatiotemporal risk intensity output by the dynamic risk field module, and triggers an edge-end pre-warning or cloud-based linkage response.
[0010] The collaborative warning module executes a multi-level alarm strategy, including audible and visual alerts, traffic light control, and security notifications, based on the risk level instructions output by the hierarchical response module.
[0011] The intelligent acquisition and recognition module includes a data acquisition submodule, a pedestrian data extraction submodule, a vehicle data extraction submodule, and a sudden moving object extraction submodule. The data acquisition submodule acquires video information of the current monitoring area in real time. The pedestrian data extraction submodule detects pedestrians based on the real-time acquired video information of the current monitoring area using an improved YOLOv5-Pedestrian network and outputs the center coordinates x of the pedestrian bounding box. p ,y p Speed v p and orientation angle θ p ;
[0012] The vehicle data extraction submodule uses the YOLOv7-Vehicle network to detect vehicles based on real-time video information collected from the current monitoring area, and outputs the center coordinates x of the vehicle bounding box. v ,y v Speed v v and heading angle
[0013] The sudden moving object extraction submodule detects sudden moving objects based on real-time video information of the current monitoring area using an inter-frame difference and optical flow fusion algorithm, and outputs the center coordinates x of the bounding box of the sudden moving object. u ,y u Motion vector Δx u ,Δy u and area change rate A u .
[0014] The specific formula for constructing the environment state vector by the feature fusion module is as follows:
[0015] E t =[F v ,F p ,F u [,W,T];
[0016] Among them, E t This represents the environment state vector;
[0017] F v ={x v ,y v ,v v , Vehicle type c v};
[0018] F p ={x p ,y p ,v p ,θ p pedestrian density p p};
[0019] F u ={x u ,y u ,Δx u ,Δy u A u Sudden object category c u};
[0020] W = {wind speed, visibility, precipitation intensity};
[0021] T = {Traffic flow, average vehicle speed, lane occupancy rate}.
[0022] In the heterogeneous trajectory prediction module, the Social-LSTM model introduces an attention weight pooling layer to optimize the prediction of pedestrian group interaction trajectories and output the predicted trajectory point sequence of each pedestrian in the next few frames.
[0023] The TransMotion model combines prior knowledge of traffic rules to output a probability distribution of vehicle trajectories with confidence.
[0024] The random field probability model uses a spatiotemporal conditional random field model to model the probability distribution of multiple candidate trajectories of a sudden object, and outputs a set of trajectories with confidence scores to reflect the uncertainty of the motion of the sudden object.
[0025] The specific operation steps of the dynamic risk field module are as follows:
[0026] Receive all trajectory prediction results from the heterogeneous trajectory prediction module;
[0027] Calculate the probability of intersection of any two trajectories within the same future spatiotemporal window;
[0028] Based on the intersection probability, dynamic environmental factors are further introduced, including the impact of real-time weather conditions on braking distance and visibility;
[0029] Introduce an object hazard level coefficient to differentiate the risk weights of pedestrians, ordinary vehicles, hazardous materials transport vehicles, and unknown sudden objects.
[0030] Taking into account the above factors, a risk intensity map that varies with time and space is generated.
[0031] The specific operation steps of the hierarchical response module are as follows:
[0032] Two risk thresholds are preset, dividing the risk intensity into three levels: low, medium, and high.
[0033] The risk value at each spatiotemporal coordinate point of the risk intensity map generated by the dynamic risk field module is read in real time and compared with a preset threshold to determine the corresponding risk level strategy at that spatiotemporal coordinate.
[0034] Specifically, the low-risk strategy of the graded response module is to provide text prompts to on-site personnel via LED screens at the edge.
[0035] The specific strategy for the medium-risk level of the graded response module is as follows: trigger an audible and visual alarm at the edge and automatically extend the green light duration of the downstream traffic lights;
[0036] The specific strategy for high-risk levels in the graded response module is as follows: immediately upload alarm information to the cloud and coordinate with security personnel to handle the situation.
[0037] Specifically, the collaborative warning module operates by synchronously executing three levels of alarm actions based on the instructions of the hierarchical response module:
[0038] Audible and visual alerts: On-site warnings are provided via ≥90dB directional speakers and red and blue flashing lights;
[0039] Traffic signal control: Using the WebRTC protocol to remotely adjust the traffic signal phases at intersections to avoid conflicts between pedestrians and vehicles;
[0040] Security notice: Pushing an alarm package in JSON format containing risk coordinates, levels, object types, and predicted trajectory video clips to security terminals through the MQTT protocol.
[0041] The present invention also provides a method for video surveillance in public areas, which is applied to the public area video surveillance system as described above, and includes the following steps:
[0042] Using the intelligent acquisition and recognition module to collect and recognize pedestrians, vehicles, and sudden moving objects in the monitored area in real time;
[0043] Using the feature fusion module to fuse the recognized target data with meteorological and traffic flow data to construct an environmental state vector;
[0044] The heterogeneous trajectory prediction module predicts the movement trajectories of pedestrians, vehicles, and sudden moving objects respectively based on the environmental state vector using a differential model;
[0045] The dynamic risk field module quantifies and classifies the spatio-temporal risk intensity based on the trajectory prediction results and environmental dynamic factors;
[0046] The classification response module triggers edge-end warnings or cloud linkage responses according to the risk levels;
[0047] The collaborative warning module executes audible and visual reminders, traffic signal control, or security notices matching the risk levels.
[0048] A public area video surveillance system and method of the present invention include an intelligent acquisition and recognition module, a feature fusion module, a heterogeneous trajectory prediction module, a dynamic risk field module, a hierarchical response module, and a collaborative warning module. The intelligent acquisition and recognition module accurately distinguishes pedestrians, vehicles, and sudden moving objects, laying a foundation for subsequent differential processing. The feature fusion module integrates video data with real-time meteorological and traffic flow information to construct a multi-dimensional environmental state vector, breaking through the limitations of single coordinate matching. The heterogeneous trajectory prediction module uses Social-LSTM, TransMotion, and random field probability models to process the motion characteristic differences of three types of targets respectively, significantly improving the prediction accuracy. The dynamic risk field module calculates the trajectory intersection probability and weighted integrates environmental dynamic factors (such as wind speed and humidity changes) and object danger levels to achieve dynamic quantification of spatio-temporal risks. The hierarchical response module and the collaborative warning module cooperate to implement a progressive warning strategy according to the risk intensity, forming a closed-loop response from the edge-end acoustic and optical reminder to the cloud linkage control. This system not only solves the prediction error problem caused by motion characteristic differences but also overcomes the risk misjudgment defect affected by environmental factor changes, making the accuracy and practicality of public safety warning higher. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0050] Figure 1 It is a schematic block diagram of the public area video surveillance system according to the first embodiment of the present invention.
[0051] Figure 2 It is a flowchart of the specific operation steps of the dynamic risk field module in the first embodiment of the present invention.
[0052] Figure 3 It is a schematic block diagram of the public area video surveillance system according to the second embodiment of the present invention.
[0053] Figure 4 It is a flowchart of the steps of the public area video surveillance method provided by the present invention.
[0054] 101-Intelligent data acquisition and recognition module, 102-Feature fusion module, 103-Heterogeneous trajectory prediction module, 104-Dynamic risk field module, 105-Graded response module, 106-Collaborative warning module, 107-Data acquisition sub-module, 108-Pedestrian data extraction sub-module, 109-Vehicle data extraction sub-module, 110-Sudden moving object extraction sub-module, 201-Event review module, 202-Adaptive strategy optimization module. Detailed Implementation
[0055] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.
[0056] First embodiment:
[0057] Please see Figure 1 and Figure 2 ,in Figure 1 This is a schematic diagram of the public area video surveillance system according to the first embodiment. Figure 2 This is a flowchart of the specific operation steps of the dynamic risk field module in the first embodiment.
[0058] This invention provides a public area video surveillance system, comprising an intelligent acquisition and recognition module 101, a feature fusion module 102, a heterogeneous trajectory prediction module 103, a dynamic risk field module 104, a hierarchical response module 105, and a collaborative warning module 106. The intelligent acquisition and recognition module 101 includes a data acquisition submodule 107, a pedestrian data extraction submodule 108, a vehicle data extraction submodule 109, and a sudden moving object extraction submodule 110. This solution addresses the problems of existing public video surveillance systems in practical applications, such as failing to distinguish the motion characteristics of pedestrians, vehicles, and suddenly moving objects, and relying solely on coordinate overlap to determine risk without considering the impact of dynamic environmental changes on risk probability, leading to large prediction errors and poor practicality. It is understood that the aforementioned solution can be used in the architecture of public area video surveillance systems.
[0059] In this specific embodiment, the intelligent acquisition and recognition module 101 is used to acquire video data of the monitored area in real time and identify pedestrians, vehicles, and sudden moving objects. The extracted video data is transmitted to the feature fusion module 102 via a wireless communication module. The feature fusion module 102 fuses the video data of pedestrians, vehicles, and sudden moving objects extracted by the intelligent acquisition and recognition module 101 with meteorological data and traffic flow data to construct an environmental state vector. The heterogeneous trajectory prediction module 103 uses a Social-LSTM model, a TransMotion model, and a random field probability model respectively to predict the movement of pedestrians, vehicles, and sudden moving objects based on the environmental state vector constructed by the feature fusion module 102. The motion trajectory of moving objects is predicted; the dynamic risk field module 104 is connected to the heterogeneous trajectory prediction module 103, and calculates the trajectory intersection probability based on the prediction results of the heterogeneous trajectory prediction module 103 for the motion trajectories of pedestrians, vehicles and suddenly moving objects, and quantifies the spatiotemporal risk intensity by combining environmental dynamic factors and object hazard level; the graded response module 105 obtains the corresponding risk level based on the spatiotemporal risk intensity output by the dynamic risk field module 104, and triggers edge pre-warning or cloud linkage response; the collaborative warning module 106 executes a multi-level alarm strategy of sound and light reminders, traffic light control and security notification based on the risk level instructions output by the graded response module 105.
[0060] In this embodiment, the intelligent acquisition and recognition module 101 accurately distinguishes pedestrians, vehicles, and suddenly moving objects, laying the foundation for subsequent differentiated processing. The feature fusion module 102 integrates video data with real-time weather and traffic flow information to construct a multi-dimensional environmental state vector, breaking through the limitations of single coordinate matching. The heterogeneous trajectory prediction module 103 uses Social-LSTM, TransMotion, and random field probability models to process the differences in motion characteristics of the three types of targets, significantly improving prediction accuracy. The dynamic risk field module 104 calculates the trajectory intersection probability and weights and integrates dynamic environmental factors (such as wind speed and humidity changes) and object hazard levels to achieve dynamic quantification of spatiotemporal risks. The graded response module 105 and the collaborative warning module 106 work together to implement a progressive early warning strategy based on risk intensity, forming a closed-loop response from edge-end audio-visual alerts to cloud-based linkage control. This system not only solves the prediction error problem caused by differences in motion characteristics but also overcomes the risk misjudgment defect caused by changes in environmental factors, making public safety early warnings more accurate and practical.
[0061] The intelligent acquisition and recognition module 101 includes a data acquisition submodule 107, a pedestrian data extraction submodule 108, a vehicle data extraction submodule 109, and a sudden moving object extraction submodule 110. The data acquisition submodule 107 acquires video information of the current monitoring area in real time. The pedestrian data extraction submodule 108 detects pedestrians based on the real-time acquired video information of the current monitoring area using an improved YOLOv5-Pedestrian network and outputs the center coordinates x of the pedestrian bounding box. p ,y p Speed v p and orientation angle θ p ;
[0062] The vehicle data extraction submodule 109 detects vehicles using the YOLOv7-Vehicle network based on real-time video information collected from the current monitoring area, and outputs the center coordinates x of the vehicle bounding box. v ,y v Speed v v and heading angle
[0063] The sudden moving object extraction submodule 110 detects sudden moving objects based on real-time acquired video information of the current monitoring area, using an inter-frame difference and optical flow fusion algorithm, and outputs the center coordinates x of the bounding box of the sudden moving object. u ,y u Motion vector Δx u ,Δy u and area change rate A u .
[0064] Secondly, the specific construction formula for the environment state vector constructed by the feature fusion module 102 is as follows:
[0065] E t =[F v ,F p ,F u [,W,T];
[0066] Among them, E t This represents the environment state vector;
[0067]
[0068] F p ={x p ,y p ,v p ,θ p pedestrian density p p};
[0069] F u ={x u ,yu ,Δx u ,Δy u A u Sudden object category c u};
[0070] W = {wind speed, visibility, precipitation intensity};
[0071] T = {Traffic flow, average vehicle speed, lane occupancy rate}.
[0072] In this embodiment, the intelligent acquisition and recognition module 101 employs an improved YOLOv5-Pedestrian network, a YOLOv7-Vehicle network, and an inter-frame difference and optical flow fusion algorithm to accurately extract feature parameters of pedestrians (position, speed, orientation), vehicles (position, speed, heading), and suddenly moving objects (position, displacement, area change); the feature fusion module 102 constructs a multi-dimensional environmental state vector E. t =[F v ,F p ,F u [,W,T] deeply integrates the dynamic characteristics of three types of targets with real-time meteorological data (wind speed, visibility, etc.) and traffic flow data (flow rate, vehicle speed, etc.). This not only achieves differentiated characterization of the motion characteristics of pedestrians, vehicles, and sudden objects, but also effectively quantifies the impact of changes in external conditions on risk probability through dynamic weighting of environmental parameters. This fundamentally solves the prediction error problem caused by the simplification of target characteristics and the static nature of environmental factors in traditional systems.
[0073] Meanwhile, in the heterogeneous trajectory prediction module 103, the Social-LSTM model introduces an attention weight pooling layer to optimize the prediction of pedestrian group interaction trajectories and output the predicted trajectory point sequence of each pedestrian in the next few frames.
[0074] The TransMotion model combines prior knowledge of traffic rules to output a probability distribution of vehicle trajectories with confidence.
[0075] The random field probability model uses a spatiotemporal conditional random field model to model the probability distribution of multiple candidate trajectories of a sudden object, and outputs a set of trajectories with confidence scores to reflect the uncertainty of the motion of the sudden object.
[0076] In this implementation, for pedestrian groups, the improved Social-LSTM model effectively captures the social interaction features among pedestrians through attention weight pooling layers, outputting high-precision individual trajectory sequences. For vehicle motion, the TransMotion model integrates prior knowledge of traffic rules to generate probabilistic trajectory predictions that meet traffic constraints. For objects with sudden movement, a spatiotemporal conditional random field model is used to construct the probability distribution of multiple candidate trajectories, fully characterizing their motion uncertainty. This differentiated prediction method not only solves the limitation of traditional systems using a single prediction model for all targets, but also provides more reliable trajectory data support for subsequent risk assessment by outputting probabilistic prediction results with confidence, significantly improving the system's prediction accuracy for the motion trends of various targets in complex scenarios.
[0077] In addition, the specific operating steps of the dynamic risk field module 104 are as follows:
[0078] S101: Receive all trajectory prediction results from the heterogeneous trajectory prediction module 103;
[0079] S102: Calculate the probability of intersection of any two trajectories within the same future spatiotemporal window;
[0080] S103: Based on the intersection probability, further introduce dynamic environmental factors, including the impact of real-time weather conditions on braking distance and visibility;
[0081] S104: Introduce an object hazard level coefficient to distinguish the risk weights of pedestrians, ordinary vehicles, dangerous goods transport vehicles, and unknown sudden objects.
[0082] S105: Taking into account the above factors, generate a risk intensity map that varies with time and space.
[0083] In this embodiment, the dynamic risk field module 104 first integrates all trajectory data output by the heterogeneous trajectory prediction module 103, and identifies potential collision risks through spatiotemporal intersection probability calculation. Secondly, it introduces real-time meteorological parameters (such as the impact of rainfall on braking distance and the effect of strong winds on object stability) to dynamically correct the risk value. Simultaneously, it combines the risk coefficients differentiated by target type (such as the weight of hazardous material vehicles being greater than that of ordinary pedestrians) to finally generate a spatiotemporally continuous risk intensity heatmap. This three-dimensional assessment model, which integrates trajectory intersection probability, dynamic environmental influences, and target hazard levels, overcomes the shortcomings of traditional systems that rely solely on static coordinate matching. It enables risk warnings to possess both physical accuracy and environmental adaptability, significantly improving the reliability and practicality of public safety monitoring.
[0084] Furthermore, the specific operating steps of the hierarchical response module 105 are as follows:
[0085] Two risk thresholds are preset, dividing the risk intensity into three levels: low, medium, and high.
[0086] The risk value at each spatiotemporal coordinate point of the risk intensity map generated by the dynamic risk field module 104 is read in real time and compared with a preset threshold to determine the corresponding risk level strategy at that spatiotemporal coordinate.
[0087] In this embodiment, the low-risk strategy of the graded response module 105 is specifically as follows: providing text prompts to on-site personnel via LED screens at the edge.
[0088] The specific strategy for the medium-risk level of the graded response module 105 is as follows: trigger an audible and visual alarm at the edge and automatically extend the green light duration of the downstream traffic lights.
[0089] The specific strategy for high-risk levels in the graded response module 105 is as follows: immediately upload the alarm information to the cloud and coordinate with security personnel to handle the situation.
[0090] Furthermore, the specific operation of the collaborative warning module 106 is as follows: based on the instructions of the hierarchical response module 105, it synchronously executes three-level alarm actions:
[0091] Audible and visual alerts: On-site warnings are provided via ≥90dB directional speakers and red and blue flashing lights;
[0092] Traffic light control: Remotely adjust the phase of traffic lights at intersections using the WebRTC protocol to avoid conflicts between pedestrians and vehicles;
[0093] Security notification: Pushes JSON format alarm packets containing risk coordinates, level, object type, and predicted trajectory video clips to security terminals via the MQTT protocol.
[0094] When using a public area video surveillance system according to this embodiment, the intelligent acquisition and recognition module 101 accurately distinguishes pedestrians, vehicles, and suddenly moving objects, laying the foundation for subsequent differentiated processing. The feature fusion module 102 integrates video data with real-time weather and traffic flow information to construct a multi-dimensional environmental state vector, breaking through the limitations of single coordinate matching. The heterogeneous trajectory prediction module 103 uses Social-LSTM, TransMotion, and random field probability models to process the differences in motion characteristics of the three types of targets, significantly improving prediction accuracy. The dynamic risk field module 104 calculates the trajectory intersection probability and weights and integrates dynamic environmental factors (such as wind speed and humidity changes) and object danger levels to achieve dynamic quantification of spatiotemporal risks. The graded response module 105 and the collaborative warning module 106 work together to implement a progressive early warning strategy based on the risk intensity, forming a closed-loop response from edge-end audio-visual reminders to cloud-based linkage control. This system not only solves the prediction error problem caused by differences in motion characteristics but also overcomes the risk misjudgment defect caused by changes in environmental factors, making public safety early warnings more accurate and practical.
[0095] Second embodiment:
[0096] Based on the first embodiment, please refer to Figure 3 , Figure 3 This is a schematic diagram of the public area video surveillance system according to the second embodiment.
[0097] The present invention provides a public area video surveillance system, which also includes an event review module 201 and an adaptive strategy optimization module 202.
[0098] In this specific implementation, the event review module 201 is communicatively connected to the hierarchical response module 105. The event review module 201 is used to subscribe to all alarm events triggered by the hierarchical response module 105 in real time, automatically capture multiple original videos, environmental state vectors and risk intensity maps within a preset time period from before the alarm to after the alarm, form an immutable hash fingerprint, and encapsulate the captured multimodal data into an event dossier. At the same time, it provides a visual backtracking interface, supports trajectory restoration, responsibility determination and strategy optimization suggestions output with risk hotspots as the entry point.
[0099] The input of the adaptive strategy optimization module 202 is connected to the event review module 201, and the output of the adaptive strategy optimization module 202 is connected to the heterogeneous trajectory prediction module 103, the dynamic risk field module 104, the hierarchical response module 105, and the collaborative warning module 106. The specific operation steps of the adaptive strategy optimization module 202 are as follows:
[0100] A digital twin simulation environment is established in the cloud, and historical event files of the event replay module 201 are periodically imported.
[0101] A new parameter combination is generated by jointly optimizing the heterogeneous trajectory prediction model, dynamic risk field threshold and collaborative alarm strategy through reinforcement learning algorithm.
[0102] By adopting a canary release mechanism, the optimized strategy is distributed as an edge hot patch, and continuously iterated through A / B testing in real scenarios to achieve continuous self-evolution of system performance.
[0103] When using a public area video surveillance system according to this embodiment, the event review module 201 uses blockchain technology to solidify the entire process data of the alarm event (including original video, environmental status, and risk map), forming a traceable and tamper-proof complete evidence chain, providing multi-dimensional visual backtracking support for post-event analysis; the adaptive strategy optimization module 202 establishes a virtual simulation environment based on digital twin technology, and jointly optimizes the prediction model, risk assessment threshold, and alarm strategy through reinforcement learning, and uses a gray-scale release mechanism to achieve dynamic hot updates of system parameters. Through the innovative architecture of full-cycle data storage-virtual simulation optimization-gradual deployment verification, the system solves the problem of traditional systems lacking post-event analysis capabilities, and further enhances the system's dynamic evolution characteristics through continuous adaptive optimization, significantly improving long-term adaptability and early warning accuracy in complex scenarios.
[0104] Please see Figure 4 The present invention also provides a method for video surveillance in public areas, applied to the public area video surveillance system described above, comprising the following steps:
[0105] S1: The intelligent acquisition and recognition module 101 is used to collect and identify pedestrians, vehicles and suddenly moving objects in the monitoring area in real time;
[0106] S2: The feature fusion module 102 is used to fuse the target identification data with meteorological and traffic flow data to construct an environmental state vector;
[0107] S3: The heterogeneous trajectory prediction module 103 predicts the motion trajectories of pedestrians, vehicles and sudden moving objects respectively based on the environmental state vector and using a differentiated model.
[0108] S4: The dynamic risk field module 104 quantifies and classifies the spatiotemporal risk intensity based on trajectory prediction results and dynamic environmental factors;
[0109] S5: The graded response module 105 triggers an edge warning or cloud-based linkage response based on the risk level;
[0110] S6: The collaborative warning module 106 executes audible and visual alerts, traffic light adjustments, or security notifications that match the risk level.
[0111] In this embodiment, the intelligent acquisition and recognition module 101 accurately distinguishes pedestrians, vehicles, and suddenly moving objects, laying the foundation for subsequent differentiated processing. The feature fusion module 102 integrates video data with real-time weather and traffic flow information to construct a multi-dimensional environmental state vector, breaking through the limitations of single coordinate matching. The heterogeneous trajectory prediction module 103 uses Social-LSTM, TransMotion, and random field probability models to process the differences in motion characteristics of the three types of targets, significantly improving prediction accuracy. The dynamic risk field module 104 calculates the trajectory intersection probability and weights and integrates dynamic environmental factors (such as wind speed and humidity changes) and object hazard levels to achieve dynamic quantification of spatiotemporal risks. The graded response module 105 and the collaborative warning module 106 work together to implement a progressive early warning strategy based on risk intensity, forming a closed-loop response from edge-end audio-visual alerts to cloud-based linkage control. This system not only solves the prediction error problem caused by differences in motion characteristics but also overcomes the risk misjudgment defect caused by changes in environmental factors, making public safety early warnings more accurate and practical.
[0112] The above description discloses only one preferred embodiment of the present invention, and should not be construed as limiting the scope of the present invention. Those skilled in the art will understand that all or part of the processes of the above embodiments can be implemented, and equivalent changes made in accordance with the claims of the present invention are still within the scope of the invention.
Claims
1. A public area video surveillance system, characterized in that, It includes an intelligent acquisition and recognition module, a feature fusion module, a heterogeneous trajectory prediction module, a dynamic risk field module, a hierarchical response module, and a collaborative warning module. The intelligent acquisition and recognition module is used to collect video data of the monitored area in real time and identify pedestrians, vehicles, and sudden moving objects. It transmits the extracted video data to the feature fusion module through a wireless communication module. The feature fusion module fuses the video data of pedestrians, vehicles, and sudden moving objects extracted by the intelligent acquisition and recognition module with meteorological data and traffic flow data to construct an environmental state vector. The heterogeneous trajectory prediction module uses the Social-LSTM model, TransMotion model and random field probability model respectively to predict the motion trajectories of pedestrians, vehicles and sudden moving objects based on the environmental state vector constructed by the feature fusion module. The dynamic risk field module is connected to the heterogeneous trajectory prediction module. Based on the prediction results of the heterogeneous trajectory prediction module on the motion trajectories of pedestrians, vehicles and suddenly moving objects, the probability of trajectory intersection is calculated, and the spatiotemporal risk intensity is quantified by combining environmental dynamic factors and object hazard levels. The graded response module obtains the corresponding risk level based on the spatiotemporal risk intensity output by the dynamic risk field module, and triggers an edge-end pre-warning or cloud-based linkage response. The collaborative warning module executes a multi-level alarm strategy, including audible and visual alerts, traffic light control, and security notifications, based on the risk level instructions output by the hierarchical response module.
2. The public area video surveillance system as described in claim 1, characterized in that, The intelligent data acquisition and recognition module includes a data acquisition submodule, a pedestrian data extraction submodule, a vehicle data extraction submodule, and a sudden moving object extraction submodule. The data acquisition submodule acquires video information of the current monitoring area in real time. The pedestrian data extraction submodule detects pedestrians based on the real-time acquired video information of the current monitoring area using an improved YOLOv5-Pedestrian network and outputs the center coordinates x of the pedestrian bounding box. p ,y p Speed v p and orientation angle θ p ; The vehicle data extraction submodule uses the YOLOv7-Vehicle network to detect vehicles based on real-time video information collected from the current monitoring area, and outputs the center coordinates x of the vehicle bounding box. v ,y v Speed v v and heading angle The sudden moving object extraction submodule detects sudden moving objects based on real-time video information of the current monitoring area using an inter-frame difference and optical flow fusion algorithm, and outputs the center coordinates x of the bounding box of the sudden moving object. u ,y u Motion vector Δx u ,Δy u and area change rate A u .
3. The public area video surveillance system as described in claim 2, characterized in that, The specific formula for constructing the environment state vector by the feature fusion module is as follows: E t =[F v ,F p ,F u ,W,T]; Among them, E t This represents the environment state vector; F p ={x p ,y p ,v p ,θ p pedestrian density p p }; F u ={x u ,y u ,Δx u ,Δy u A u Sudden object category c u }; W = {wind speed, visibility, precipitation intensity}; T = {Traffic flow, average vehicle speed, lane occupancy rate}.
4. The public area video surveillance system as described in claim 3, characterized in that, In the heterogeneous trajectory prediction module, the Social-LSTM model introduces an attention weight pooling layer to optimize the prediction of pedestrian group interaction trajectories and outputs the predicted trajectory point sequence for each pedestrian in the next few frames. The TransMotion model combines prior knowledge of traffic rules to output a probability distribution of vehicle trajectories with confidence. The random field probability model uses a spatiotemporal conditional random field model to model the probability distribution of multiple candidate trajectories of a sudden object, and outputs a set of trajectories with confidence scores to reflect the uncertainty of the motion of the sudden object.
5. The public area video surveillance system as described in claim 4, characterized in that, The specific operation steps of the dynamic risk field module are as follows: Receive all trajectory prediction results from the heterogeneous trajectory prediction module; Calculate the probability of intersection of any two trajectories within the same future spatiotemporal window; Based on the intersection probability, dynamic environmental factors are further introduced, including the impact of real-time weather conditions on braking distance and visibility; Introduce the object danger level coefficient to distinguish the risk weights of pedestrians, ordinary vehicles, dangerous goods transport vehicles, and unknown sudden objects; Based on the above factors, generate a risk intensity map that changes with time and space.
6. The public area video surveillance system according to claim 5, wherein The specific operation steps of the hierarchical response module are as follows: Preset two levels of risk thresholds, and divide the risk intensity into three levels: low, medium, and high; Read the risk value of the risk intensity map generated by the dynamic risk field module at each space-time coordinate point in real time, and compare it with the preset threshold to determine the corresponding risk level strategy at this space-time coordinate.
7. The public area video surveillance system according to claim 6, wherein The strategy for the low risk level of the hierarchical response module is specifically: use the LED screen at the edge to prompt the on-site personnel with text; The strategy for the medium risk level of the hierarchical response module is specifically: trigger the audible and visual alarm at the edge, and automatically extend the green light duration of the downstream signal lights; The strategy for the high risk level of the hierarchical response module is specifically: immediately upload the alarm information to the cloud and link the security personnel to go to the scene for disposal.
8. The public area video surveillance system according to claim 7, wherein The specific operation content of the collaborative warning module is to synchronously execute three-level alarm actions according to the instructions of the hierarchical response module: Audible and visual reminder: Use a directional speaker with ≥90dB and a red-blue flashing light for on-site warning; Signal light regulation: Use the WebRTC protocol to remotely adjust the signal light phase at the intersection to avoid conflicts between people and vehicles; Security notification: Push a JSON-format alarm package containing risk coordinates, levels, object types, and predicted trajectory video clips to the security terminal through the MQTT protocol.
9. A method for video surveillance of public areas, applied to the public area video surveillance system as described in claim 1, characterized in that, Include the following steps: Use the intelligent acquisition and recognition module to collect and recognize pedestrians, vehicles, and sudden moving objects in the monitoring area in real time; Use the feature fusion module to fuse the recognized target data with meteorological and traffic flow data to construct an environmental state vector; The heterogeneous trajectory prediction module predicts the movement trajectories of pedestrians, vehicles, and sudden moving objects respectively based on the environmental state vector using a differential model; The dynamic risk field module quantifies and classifies the space-time risk intensity based on the trajectory prediction results and environmental dynamic factors; The hierarchical response module triggers an early warning at the edge end or a cloud linkage response according to the risk level; The collaborative warning module executes audible and visual reminders, signal light regulation, or security notifications that match the risk level.
Citation Information
Patent Citations
Public video monitoring system based on big data
CN117636240A
Cited By
Video monitoring information integrated management method for multi-scene security and protection
CN121214352A
Intelligent cruise ship passenger safety early warning system and method based on fusion of large model and visual identification
CN121600595A