Urban rail transit emergency processing method and system based on visual large model
By collecting multi-source data in the urban rail transit system, performing visual analysis and multi-modal fusion, combined with three-dimensional spatial mapping, the problems of slow response speed and high false alarm rate in the existing system are solved, and more intelligent and automated emergency treatment is achieved.
Patent Information
- Application Number
- CN202510510717.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-23
AI Technical Summary
The existing urban rail transit emergency response system has problems such as slow response speed, high false alarm rate, single information dimension and lack of spatial information, making it difficult to achieve efficient and intelligent emergency response to urban rail transit systems.
Using a visual big model-based approach, visual data, sound data and equipment operation data are collected in key areas of urban rail transit platforms and carriages, and preprocessing and time stamp synchronization are performed. Dynamic frame difference and spatiotemporal aggregation technology are used to generate two-dimensional motion heat maps, and the two-dimensional image information is associated with the three-dimensional spatial environment through geometric mapping technology. Multimodal fusion analysis is carried out in combination with sound data and equipment operation data to determine the severity of abnormal events, and to formulate emergency response strategies based on severity rating.
It improves the emergency response speed, reduces the false alarm rate, realizes the integration of multi-source information and comprehensive judgment of abnormal events, and improves the automation and intelligence level of urban rail transit emergency response.
Smart Images

Figure CN120047907A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data technology, and in particular to an urban rail transit emergency processing method and system based on a visual big model. Background Art
[0002] As an important public transportation system in modern cities, the safe and efficient operation of urban rail transit is related to the daily travel of urban residents and the stable development of social economy. However, the operating environment of urban rail transit system is complex, with large passenger flow, and emergencies occur from time to time, such as accidental falls of passengers, equipment failures, and personnel intrusion into the track, which may lead to operational interruptions, passenger casualties, and even mass safety accidents. Therefore, it is very important to establish an efficient and intelligent urban rail transit emergency response system.
[0003] In the existing technology, emergency handling of urban rail transit mainly relies on the following methods:
[0004] (1) Manual monitoring and inspection: Traditional urban rail transit systems generally adopt a manual monitoring mode, that is, monitoring personnel are deployed in the control center to monitor key areas such as platforms and carriages in real time by viewing the closed-circuit television monitoring system screen. In addition, inspection personnel are also deployed to conduct regular or irregular inspections on platforms and carriages to detect abnormal situations.
[0005] However, manual monitoring and inspection methods have obvious limitations:
[0006] Slow response: Manual monitoring relies on the visual observation and subjective judgment of monitoring personnel, and there is a lag in the discovery of abnormal events. Especially when there are many monitoring screens and a large amount of information, negligence and omissions are prone to occur, resulting in slow emergency response.
[0007] Easy to be fatigued and misjudgment: Long-term monitoring work can easily lead to fatigue of monitoring personnel, decreased attention, reduced sensitivity to abnormal events, and prone to misjudgment or missed judgment, affecting the accuracy of emergency response.
[0008] Limited coverage: Manual inspections have high labor costs, and the inspection frequency and coverage are limited, making it difficult to achieve real-time and comprehensive monitoring of all key areas.
[0009] (2) Early warning systems based on simple sensors: In order to improve the level of automation, some rail transit systems have begun to introduce early warning systems based on simple sensors, such as infrared sensors, sound sensors, pressure sensors, etc. These sensors can detect changes in specific physical quantities, such as abnormal temperature, abnormal sound, abnormal pressure, etc., thereby achieving preliminary abnormal warning.
[0010] However, early warning systems based on simple sensors also have shortcomings:
[0011] High false alarm rate: Simple sensors are easily interfered by environmental factors. For example, infrared sensors may be affected by temperature changes, and sound sensors may be interfered by environmental noise, resulting in a high false alarm rate.
[0012] Single information dimension: A single type of sensor can only perceive limited information dimensions, making it difficult to fully and accurately determine the nature and severity of an event. For example, a pressure sensor can only detect pressure changes but cannot distinguish whether a passenger falls or an object falls.
[0013] Lack of spatial information: Simple sensors can usually only provide point-like or regional perception information, lacking an accurate description of the location and spatial relationship of the event, which is not conducive to precise positioning and subsequent disposal.
[0014] (3) Early visual monitoring systems: With the development of computer vision technology, some rail transit systems have begun to try to introduce early visual monitoring systems, such as motion detection systems based on background modeling, optical flow, etc. These systems can use the video data collected by the camera to detect moving targets in the picture and achieve preliminary automated monitoring.
[0015] However, early visual surveillance systems still face challenges in practical applications:
[0016] Poor algorithm robustness: Early visual algorithms are sensitive to factors such as changes in ambient lighting, shadow occlusion, and complex backgrounds, and are prone to false detections and missed detections. They have poor robustness and perform unstably in complex rail transit environments.
[0017] High computational complexity: Some complex visual algorithms, such as optical flow, have high computational complexity and are difficult to meet real-time requirements, which limits the response speed of the system.
[0018] Lack of semantic understanding: Early visual systems can usually only detect moving targets, lack understanding of scene semantics, cannot distinguish between normal movement and abnormal behavior, and find it difficult to conduct in-depth abnormal analysis and judgment.
[0019] (4) Intelligent video analysis technology based on deep learning: In recent years, deep learning technology has made breakthrough progress in the field of computer vision. Technologies such as target detection and behavior recognition based on deep learning have gradually been applied to the field of intelligent video surveillance. Some studies have attempted to apply deep learning technology to the detection of abnormal events in rail transit, such as using convolutional neural networks for target detection and classification, and identifying abnormal behaviors such as passenger falls, crowding and trampling. Although intelligent video analysis technology based on deep learning has improved the accuracy and intelligence level of abnormality detection to a certain extent, it still has some limitations, especially in the emergency response scenarios of urban rail transit:
[0020] High computing resource requirements: Deep learning models usually have a large number of parameters and high computational complexity, requiring powerful computing resource support and difficult to run in real time on edge devices, limiting the deployment and application scope of the system.
[0021] Strong dependence on training data: Deep learning models require a large amount of labeled data for training, but labeled data of rail transit abnormal events are difficult to obtain, the labeling cost is high, and the data quality is difficult to ensure, which affects the generalization ability and robustness of the model.
[0022] Poor model interpretability: Deep learning models are usually "black box" models, and it is difficult to explain their decision-making process, which is not conducive to system fault diagnosis and optimization and improvement, and it is difficult to meet the high requirements of rail transit systems for safety and reliability.
[0023] Insufficient fusion of multimodal information: Existing research on rail transit anomaly detection based on deep learning mainly focuses on visual data analysis, and insufficient use of other modal information such as sound data and equipment operation data. The information dimension is single, making it difficult to achieve comprehensive and multi-dimensional abnormal event judgment.
[0024] Therefore, there is a need for an urban rail transit emergency handling method that can improve emergency response speed, reduce false alarm rate, increase the degree of automation, and realize multi-source information fusion to overcome the shortcomings of existing technologies and ensure the safe and efficient operation of urban rail transit systems. Summary of the invention
[0025] The present invention provides an urban rail transit emergency response method and system based on visual big model, aiming to improve the intelligence and automation level of the rail transit system in responding to emergencies based on the visual big model, and ultimately achieve rapid response and effective handling of abnormal events by collecting, processing and analyzing multi-source data of key areas such as urban rail transit platforms and carriages.
[0026] The present invention provides an urban rail transit emergency handling method based on a visual large model, comprising:
[0027] Collect visual data, sound data and equipment operation data in key areas of urban rail transit platforms and carriages, perform pre-processing, and synchronize timestamps;
[0028] Performing dynamic frame difference processing on continuous video frames of visual data to generate a two-dimensional motion heat map, and performing spatiotemporal aggregation on the two-dimensional motion heat map;
[0029] Build a 3D digital model based on the engineering drawings of the platform and carriage of urban rail transit, perform internal and external parameter calibration of the camera, and establish a mapping relationship between the 2D motion heat map and the 3D digital model;
[0030] Abnormal behavior detection is performed on the mapping results of the two-dimensional motion heat map and the three-dimensional digital model after spatiotemporal aggregation, and multimodal fusion analysis is performed in combination with sound data and equipment operation data to determine the severity of abnormal events;
[0031] Emergency response strategies are formulated based on the severity of abnormal events, control instructions are sent to rail transit equipment through industrial protocol interfaces, and abnormal event information is pushed to the control center, patrol personnel and passenger terminals.
[0032] In the above scheme, multimodal information including visual data, sound data and equipment operation data is first collected synchronously in key areas of the platform and carriage of urban rail transit. Among them, visual data is mainly collected by cameras deployed in key areas, such as the escalator entrance of the platform, security inspection area, waiting area and inside the carriage; sound data is collected by sound sensors to assist in the judgment of abnormal events; equipment operation data, such as the operating status information of rail transit equipment such as gates, escalators, and platform doors, can reflect the working condition of the equipment itself. In order to ensure the accuracy of subsequent data processing and analysis, the collected multimodal data will be pre-processed, such as image denoising, audio signal filtering, data format unification and other operations, and a unified timestamp will be added to all data to ensure the synchronization and alignment of multi-source data on the timeline.
[0033] Subsequently, dynamic frame difference processing will be performed on the continuous video frames for the collected visual data. Dynamic frame difference technology aims to detect the areas where motion changes occur in the video screen. By calculating the pixel differences between adjacent frames, the moving targets in the video can be effectively extracted. In order to adapt to different environmental lighting and scene changes, dynamic frame difference is adopted, and the difference threshold or algorithm parameters can be adaptively adjusted according to the environment, thereby improving the robustness of motion detection. After frame difference processing, a two-dimensional motion heat map is generated, which presents the active motion areas in the video in a visual form. In order to further eliminate noise interference and extract more stable and more significant moving targets, the generated two-dimensional motion heat map will be spatiotemporally aggregated. Spatiotemporal aggregation refers to the integration of motion heat maps in time and space dimensions. For example, the heat maps of multiple consecutive frames can be superimposed or morphological filtering can be performed to filter out short-term, isolated noise points, retain continuous, piece-by-piece motion areas, and make the moving targets clearer and more complete.
[0034] In order to associate the two-dimensional image information with the actual three-dimensional space environment, geometric mapping is introduced to build an accurate three-dimensional digital model based on the engineering drawings of the urban rail transit platform and carriage. The model is a digital restoration of the real scene, which contains information such as the structural layout of the platform and carriage, the spatial position and geometric parameters of the equipment and facilities. In order to establish the correspondence between the two-dimensional image captured by the camera and the three-dimensional digital model, it is necessary to perform the calibration of the internal and external parameters of the camera. Camera calibration can determine the spatial position, posture and internal imaging parameters of the camera, so as to obtain the conversion matrix from the image coordinate system to the world coordinate system. On this basis, the mapping relationship between the two-dimensional motion heat map and the three-dimensional digital model can be established, and the motion area on the two-dimensional heat map can be accurately projected into the three-dimensional model to determine the specific position and area of the moving target in the three-dimensional space.
[0035] After obtaining the position information of the moving target in the three-dimensional space, abnormal behavior detection will be performed. By analyzing the mapping results of the two-dimensional motion heat map and the three-dimensional digital model after spatiotemporal aggregation, it can be determined whether abnormal motion behavior has occurred in a specific spatial area. For example, it can be analyzed whether the motion area interacts with equipment and facilities, whether the speed and trajectory of the motion conform to the normal mode, etc. In order to improve the accuracy and reliability of anomaly detection, the method further combines the synchronously collected sound data and equipment operation data for multimodal fusion analysis. By comprehensively considering multiple aspects of information such as vision, hearing, and equipment operation status, the nature and severity of the event can be judged more comprehensively and accurately. For example, when visual detection of a person falling in the platform area, the sound sensor also collects abnormal sounds, and the nearby platform door operation data indicates that it is in an abnormal open state, it can be comprehensively judged that a more serious emergency has occurred. Through multimodal fusion analysis, the false alarm rate can be effectively reduced and the accuracy of abnormal event identification can be improved.
[0036] According to the severity of abnormal events determined by multimodal fusion analysis, a hierarchical emergency response strategy is formulated. According to the degree of harm and urgency of the event, abnormal events can be divided into different levels, such as minor abnormalities, obvious abnormalities, serious abnormalities, etc., and corresponding emergency response measures are preset for each level. For example, for minor abnormalities, only equipment self-check or warning prompts may be triggered; for obvious abnormalities, it may be necessary to start equipment deceleration, current limiting and other measures; for serious abnormalities, it may be necessary to immediately stop the equipment operation and start a higher level of emergency plan, such as evacuating passengers and calling for rescue. In order to realize the automated control of rail transit equipment, control instructions are sent to rail transit equipment through the industrial protocol interface, such as controlling the escalator to stop, the gate to close, the platform door to open, etc. At the same time, in order to notify relevant personnel in time and conduct collaborative disposal, the method also pushes abnormal event information to the control center, patrol personnel and passenger terminals, so that relevant personnel can understand the situation in time and take corresponding actions.
[0037] Preferably, the collecting of visual data, sound data and equipment operation data includes:
[0038] Wide-angle cameras are deployed at the escalator entrances, security check areas, waiting areas, and both ends of the carriages, with the acquisition frequency set to 10 frames per second;
[0039] Use microphone arrays to collect environmental sound data and extract sound features, including frequency, energy, and spectrum changes;
[0040] The equipment operation data is collected in real time through the industrial protocol interface of rail transit equipment. The equipment operation data includes speed, current, temperature and vibration.
[0041] The above scheme specifically describes the method of data collection, including: deploying wide-angle cameras in key areas of the platform (escalator entrance, security area, waiting area) and at both ends of the car, setting the collection frequency to 10 frames per second; using microphone arrays to collect ambient sounds and extract sound features such as frequency, energy and spectrum changes; and collecting equipment operation data in real time through the rail transit equipment industrial protocol interface, such as speed, current, temperature, vibration and other parameters.
[0042] Through specific data collection configuration, the following technical effects are produced:
[0043] Comprehensive coverage of key areas and carriages to ensure data integrity: Wide-angle cameras are deployed in key areas of the platform and at both ends of the carriages to maximize coverage of passenger activity areas and equipment operation areas, ensure the comprehensiveness and integrity of visual data collection, avoid monitoring blind spots, and provide a sufficient data basis for subsequent abnormal event detection.
[0044] Reasonable acquisition frequency, taking into account both real-time and data processing pressure: The visual data acquisition frequency of 10 frames / second can not only meet the real-time requirements of motion detection, but also effectively reduce the burden of data transmission and processing, achieving a balance between system performance and data quality.
[0045] Microphone array and sound feature extraction to improve the utilization of sound information: The use of microphone array can enhance the directionality and sensitivity of sound collection and effectively capture abnormal sounds in the environment. Extracting sound features such as frequency, energy and spectrum changes can more effectively characterize the characteristics of sound, provide structured and quantifiable data for subsequent abnormal sound identification and analysis, and improve the role of sound information in judging abnormal events.
[0046] Industrial protocol interface collects equipment operation data to ensure data real-time and accuracy: Collecting operation data directly from rail transit equipment through industrial protocol interface can ensure data real-time and accuracy, avoid delays and errors in traditional sensor deployment and data conversion process, and provide a reliable data source for equipment status monitoring and abnormal warning.
[0047] Preferably, the preprocessing of the visual data, the sound data and the equipment operation data and achieving timestamp synchronization comprises:
[0048] Gaussian blurring of visual data;
[0049] Perform noise reduction on sound data;
[0050] Normalize the equipment operation data;
[0051] The network time protocol or precision time protocol is used to synchronize the timestamps of the pre-processed visual data, sound data and equipment operation data to ensure the accuracy of data fusion.
[0052] In the above scheme, Gaussian blur processing of visual data can effectively smooth the image, remove high-frequency noise, reduce the impact of factors such as ambient light changes and sensor noise on motion detection, improve the quality of visual data, and improve the accuracy and stability of subsequent moving target detection; noise reduction processing of sound data can effectively filter out environmental background noise, highlight the characteristics of abnormal sounds, and improve the signal-to-noise ratio of sound data, making subsequent sound feature extraction and abnormal sound recognition more accurate and reliable; normalization processing of equipment operation data can unify data of different dimensions into the same numerical range, eliminate the impact of dimensional differences, facilitate subsequent multimodal data fusion processing, and improve the rationality and effectiveness of data fusion; using NTP or PTP protocol to synchronize multimodal data timestamps can ensure that data from different sources are aligned on the time axis, ensure the accuracy of data fusion, avoid analysis errors caused by time asynchrony, and provide a reliable time reference for subsequent multimodal fusion anomaly detection.
[0053] Preferably, performing dynamic frame difference processing on continuous video frames of visual data to generate a two-dimensional motion heat map includes:
[0054] The pixel differences between adjacent frames of continuous video frames are calculated, and the difference threshold is dynamically adjusted through the environment adaptive algorithm to obtain a difference image; the difference image is binarized to generate a binary motion heat map.
[0055] In the above scheme, calculating the pixel difference between adjacent frames can effectively highlight the moving area in the video frame and suppress static background information, thereby distinguishing the moving target from the background; dynamically adjusting the differential threshold through the environment adaptive algorithm can automatically adjust the threshold according to changes in ambient light, noise level, etc., so that the differential process can adapt to different environmental conditions, effectively suppress the influence of factors such as lighting changes and shadow interference, improve the environmental adaptability and robustness of moving target detection, and reduce the false detection rate and missed detection rate; binarizing the differential image to generate a binary motion heat map can convert complex pixel difference information into a concise binary image, highlight the outline of the moving area, simplify data representation, facilitate subsequent spatiotemporal aggregation and target recognition processing, and improve processing efficiency.
[0056] Preferably, the spatiotemporal aggregation of the two-dimensional motion heat map comprises:
[0057] Aggregate five consecutive frames of 2D motion heatmaps in the time dimension;
[0058] The aggregated binary motion heat map is spatiotemporally aggregated, and the interference points are removed by applying the morphological closing operation. The independent motion regions are identified and labeled by the connected region labeling algorithm, and a unique identifier is assigned.
[0059] In the above scheme, time aggregation is performed on five consecutive frames of motion heat maps, which can utilize the continuity characteristics of motion, filter out short-duration instantaneous noise points, highlight the target area of stable motion, improve the quality of the motion heat map, and reduce noise interference; the application of morphological closing operation can effectively fill the holes inside the moving target and connect the broken edges to make the contour of the moving target more complete and continuous, thereby improving the integrity and accuracy of target segmentation; through the connected region labeling algorithm, the interconnected pixel areas in the motion heat map can be identified as independent moving targets and assigned unique identifiers to achieve the distinction and tracking of different moving targets, provide independent target units for subsequent abnormal behavior analysis, and facilitate behavior analysis and abnormal judgment for different targets.
[0060] Preferably, the construction of the three-dimensional digital model based on the engineering drawings of the platform and carriage of the urban rail transit includes:
[0061] Based on the CAD engineering drawings of the platform and carriages, an accurate three-dimensional digital model is constructed, which includes the spatial position and geometric parameters of fixed facilities such as escalators, gates and seats; a unique ID and function label are given to each equipment component of the fixed facility to build an equipment component database.
[0062] In the above scheme, building a three-dimensional digital model based on the CAD engineering drawings of the platform and the carriage can ensure a high degree of consistency between the model and the real physical space, provide accurate spatial geometric information, and provide an accurate spatial reference for the subsequent mapping of the two-dimensional image to the three-dimensional space; the three-dimensional model contains the spatial position and geometric parameters of fixed facilities such as escalators, gates, and seats, which can effectively describe the environmental scenes of rail transit platforms and carriages, and provide environmental context information for subsequent spatial position analysis and behavior understanding of moving targets; each equipment component is given a unique ID and function label, and an equipment component database is built, which can associate the components in the three-dimensional model with the actual equipment, supporting subsequent refined abnormal behavior analysis and emergency response based on components, such as targeted processing of abnormal escalator movement, gate failures, etc.
[0063] Preferably, the step of performing camera internal and external parameter calibration and establishing a mapping relationship between a two-dimensional motion heat map and a three-dimensional digital model includes:
[0064] Use the calibration board to perform camera internal and external parameter calibration to obtain the camera's internal and external parameter matrix;
[0065] Based on the internal and external parameter calibration results, an affine transformation is applied to map the motion area in the two-dimensional motion heat map to the spatial coordinate system of the three-dimensional digital model;
[0066] Through the spatial occupancy detection algorithm, the motion area that overlaps with the equipment parts with unique IDs in the 3D model is identified.
[0067] In the above scheme, the internal and external parameters of the camera are calibrated through the calibration board, so that the internal parameter matrix and external parameter matrix of the camera can be obtained, the imaging model of the camera can be established, and the conversion relationship between the image pixel coordinates and the three-dimensional space coordinates can be described, so as to provide necessary parameters for the subsequent mapping transformation; the motion area in the two-dimensional motion heat map is mapped to the spatial coordinate system of the three-dimensional CAD model by applying the affine transformation, so as to realize the conversion of the two-dimensional image information to the three-dimensional physical space, so that the position and behavior of the moving target can be analyzed in the three-dimensional space; through the space occupancy detection algorithm, the motion area overlapping with the spatial position of the equipment component with a unique ID in the three-dimensional model can be identified, so as to determine whether the moving target interacts with the specific equipment component, such as whether the pedestrian is close to the escalator, gate machine, etc., so as to provide a basis for the subsequent component-based abnormal behavior analysis.
[0068] Preferably, the abnormal behavior detection on the mapping result of the two-dimensional motion heat map and the three-dimensional digital model after spatiotemporal aggregation includes:
[0069] For the motion area that spatially overlaps with the equipment component with a unique ID, the motion frequency of the equipment component within the preset time window is calculated, and the spectral characteristics of the motion are analyzed through Fourier transform to extract the motion frequency and amplitude. An abnormal state threshold table is preset based on the equipment type, and an alarm for the corresponding component is triggered when the motion frequency or amplitude is detected to exceed the threshold.
[0070] In the above scheme, the movement frequency of the equipment components within the time window is calculated, and the motion spectrum characteristics are analyzed through Fourier transform, which can extract the periodicity, intensity and other characteristics of the motion, more comprehensively describe the motion behavior, and provide a richer basis for abnormal judgment; based on the equipment type preset abnormal state threshold table, different motion frequency and amplitude thresholds are set for different types of equipment components, which can achieve refined abnormal judgment, improve the accuracy and sensitivity of abnormal detection, and avoid misjudgment caused by using a unified threshold for different types of equipment; by comparing the measured motion frequency and amplitude with the preset threshold, it is possible to quickly and effectively determine whether the equipment components are in an abnormal state, such as the escalator step movement frequency is too high, the gate movement amplitude is abnormal, etc., and the alarm is triggered in time to provide timely warning information for subsequent emergency response.
[0071] Preferably, the combining of the sound data and the equipment operation data for multimodal fusion analysis to determine the severity of the abnormal event includes:
[0072] Based on the weighted fusion of visual anomaly score, sound anomaly score and equipment operation data anomaly score, the comprehensive anomaly score calculation formula is:
[0073]
[0074] in , , The visual data , sound data and equipment operation data The weight coefficient of ;
[0075] When the combined anomaly score exceeds the threshold, the anomaly event is confirmed.
[0076] In the above scheme, the weighted fusion of the abnormal scores of vision, sound and equipment operation data can comprehensively consider the information of different modes, use the complementary advantages of each mode, and improve the reliability and accuracy of abnormal event judgment. For example, if the vision detects abnormal movement, the sound detects abnormal sound, and the equipment operation data also shows abnormality, the comprehensive score will be higher and the confidence of the abnormal event will be higher; the weight coefficient , , It can be flexibly adjusted according to the actual application scenario and the reliability of different modal data. For example, in scenarios where visual information is more reliable, the weight of visual score can be increased. , and vice versa, making multimodal fusion more adaptable and flexible; by setting a comprehensive anomaly score threshold, abnormal events are confirmed only when the comprehensive score exceeds the threshold, which can effectively reduce the false alarm rate, improve the accuracy of abnormal event judgment, and avoid erroneous responses caused by misjudgment of a single modality.
[0077] Preferably, formulating emergency response strategies according to the severity of abnormal events includes:
[0078] Abnormal events are divided into three levels of response:
[0079] Level 1: Minor abnormality, triggering the device self-check program, without affecting the normal operation of the device;
[0080] Level 2: Obvious abnormality, triggering equipment deceleration or current limiting measures;
[0081] Level 3: Severe abnormality, immediately stop equipment operation and activate the emergency plan.
[0082] In the above scheme, abnormal events are divided into three levels of response: minor, obvious and severe. Emergency measures of different intensities can be taken according to the severity of the abnormality, so as to achieve refined emergency handling and avoid a "one-size-fits-all" response method. Under the premise of ensuring safety, the interference with rail transit operation is minimized. Corresponding response strategies are formulated for abnormal events of different levels, such as self-inspection for minor abnormalities, deceleration and current limiting for obvious abnormalities, and suspension of operation for severe abnormalities, which ensures the rationality and effectiveness of the response strategy and can take the most appropriate measures according to actual conditions, which can effectively control risks and maintain the normal operation of rail transit as much as possible. The hierarchical framework of the three-level response has good scalability, and the response level can be further refined according to actual needs. More refined response strategies can be formulated for abnormal events of different levels, which is convenient for subsequent strategy expansion and optimization, and improves the flexibility and adaptability of the emergency handling system.
[0083] Preferably, an urban rail transit emergency handling system based on a visual large model comprises:
[0084] Data acquisition module, used to collect visual data, sound data and equipment operation data in key areas of urban rail transit platforms and carriages, and perform pre-processing, while achieving time stamp synchronization;
[0085] A visual analysis module, for performing dynamic frame difference processing on continuous video frames of visual data to generate a two-dimensional motion heat map, and performing spatiotemporal aggregation on the two-dimensional motion heat map;
[0086] The spatial mapping module is used to build a 3D digital model based on the engineering drawings of the platform and carriage of urban rail transit, perform internal and external parameter calibration of the camera, and establish a mapping relationship between the 2D motion heat map and the 3D digital model;
[0087] The anomaly detection module is used to detect abnormal behavior on the mapping results of the two-dimensional motion heat map and the three-dimensional digital model after spatiotemporal aggregation, and to perform multimodal fusion analysis in combination with sound data and equipment operation data to determine the severity of abnormal events;
[0088] The emergency response module is used to formulate emergency response strategies according to the severity of abnormal events, send control instructions to rail transit equipment through the industrial protocol interface, and push abnormal event information to the control center, patrol personnel and passenger terminals.
[0089] In the above scheme, the emergency handling method is decomposed into modules such as data acquisition, visual analysis, spatial mapping, anomaly detection and emergency response. Each module is responsible for a specific function, and the interfaces between modules are clear, which is convenient for system development, implementation and maintenance, reducing system complexity and improving maintainability and scalability. The modules work together in the order of data flow and control flow. The data acquisition module is responsible for data input, the visual analysis and spatial mapping modules are responsible for data processing and feature extraction, the anomaly detection module is responsible for abnormal event determination, and the emergency response module is responsible for executing control instructions and information push. The emergency handling method process is fully realized and a complete emergency handling system is constructed. Through modular system design, the emergency handling method is implemented as an operational system, which can automatically and intelligently complete the perception, identification, decision-making and response of abnormal events, significantly improving the efficiency and reliability of urban rail transit emergency handling, reducing the need for manual intervention, and improving the overall performance of the system.
[0090] Compared with the prior art, the present invention has the following beneficial effects:
[0091] The present invention discloses an urban rail transit emergency handling method and system based on a visual large model, which ingeniously combines traditional image processing technology and modern visual analysis methods, effectively improving the accuracy and robustness of anomaly detection while ensuring real-time performance and low computational complexity; specifically:
[0092] (1) Efficient moving target detection is achieved through dynamic frame difference and spatiotemporal aggregation technology;
[0093] (2) Through geometric mapping technology, the precise association between two-dimensional image information and three-dimensional spatial environment is achieved, overcoming the defect of traditional visual systems that lack spatial information;
[0094] (3) Through multimodal data fusion, comprehensive use of visual, sound and equipment operation data has been achieved to achieve more comprehensive and accurate abnormal event judgment;
[0095] (4) The efficiency and intelligence level of emergency handling are improved through hierarchical emergency response strategies and automated control. BRIEF DESCRIPTION OF THE DRAWINGS
[0096] Figure 1 It is a schematic flow chart of an urban rail transit emergency handling method based on a visual large model provided by an embodiment of the present invention;
[0097] Figure 2 It is a schematic diagram of a module of an urban rail transit emergency handling system based on a visual large model provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0098] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0099] like Figure 1 As shown, the present application provides an urban rail transit emergency handling method based on a visual large model, comprising:
[0100] S1: Collect visual data, sound data and equipment operation data in key areas of urban rail transit platforms and carriages, perform pre-processing, and synchronize timestamps;
[0101] In one embodiment provided in the present application, multiple sets of low-resolution wide-angle cameras, sound sensors and equipment operation data acquisition modules are deployed in key areas of urban rail transit platforms and carriages, such as the escalator entrance, security inspection area, passenger waiting area and passenger area inside the carriage, to achieve maximum coverage of the monitoring area and collect multi-source heterogeneous data; specifically, multiple low-resolution wide-angle cameras are deployed, preferably with a resolution of 640×480 pixels and a wide-angle camera with a viewing angle of 130°, such as an industrial-grade camera with model XYZ-Cam-001. In key areas of the platform, such as escalator entrances, security inspection areas, and waiting areas, 2-3 cameras are configured in each key area to form multi-angle redundant monitoring and improve system reliability. Inside the carriage, wide-angle cameras are installed at both ends of each carriage to ensure that the monitoring view can fully cover the passenger activity area in the carriage. The image acquisition frequency of the camera is set to 10 frames per second, which can not only meet the needs of motion detection, but also effectively reduce the burden of data transmission and processing.
[0102] At the same time, sound sensors are deployed synchronously near the camera deployment area, such as the high-sensitivity microphone model ABC-Mic-002, to collect environmental sound data, such as passengers' cries for help, abnormal equipment noise, etc. In addition, through the equipment control interface, such as the interface that complies with the Modbus protocol, the operation data of key rail transit equipment, such as the running speed of escalators, the opening and closing status of gates, and the opening and closing status of platform doors, are collected in real time.
[0103] The collected visual data, sound data and equipment operation data are all preprocessed and timestamp synchronized. Visual data preprocessing mainly includes noise reduction of the original image, such as using the median filter algorithm to filter out the noise in the image to improve the accuracy of subsequent motion detection. Sound data preprocessing includes analog-to-digital conversion, filtering and noise reduction of the original audio signal to extract effective sound features. Equipment operation data preprocessing includes data format conversion, data cleaning, outlier processing and other operations to ensure data accuracy and availability. The timestamp synchronization operation uses the Network Time Protocol (NTP) or Precision Time Protocol (PTP) to ensure that the data collected by different sensors are aligned on the time axis, laying the foundation for subsequent multimodal data fusion analysis.
[0104] Through the collaborative work of multi-source sensors, comprehensive and synchronous collection of visual, sound and equipment operation data in the urban rail transit environment is achieved, providing multi-dimensional data support for subsequent intelligent analysis. The use of low-resolution cameras and moderate acquisition frequencies effectively reduces the pressure of data collection and transmission while ensuring monitoring effects, meeting the application scenario requirements of edge computing. The timestamp synchronization mechanism ensures the consistency of multimodal data in the time dimension, providing a data basis for subsequent fusion analysis.
[0105] S2: performing dynamic frame difference processing on continuous video frames of visual data to generate a two-dimensional motion heat map, and performing spatiotemporal aggregation on the two-dimensional motion heat map;
[0106] In one embodiment provided in the present application, continuous video frames are first pre-processed with Gaussian blur, for example, a convolution filter is performed using a Gaussian kernel of σ=1.5 to effectively eliminate interference factors such as the camera's own noise and changes in ambient light. Then, the pixel difference between two adjacent frames is calculated, for example, using the absolute value difference method to obtain a differential image. In order to adapt to changes in illumination and noise levels in different scenes, an environmental adaptive algorithm is used to dynamically adjust the differential threshold, for example, dynamically adjusting the threshold based on the image grayscale mean and standard deviation to ensure the sensitivity and robustness of motion area detection. The differential image is binarized, that is, pixels with pixel difference values greater than the dynamic threshold are set to white, and pixels with pixel difference values less than the threshold are set to black, thereby generating a binary motion heat map that highlights the motion area in the video frame;
[0107] In order to further suppress noise interference and extract stable moving targets, the binary motion heat maps of 5 consecutive frames are spatiotemporally aggregated. For example, the aggregation is performed by pixel-by-pixel summation, and the pixel values of corresponding positions in the heat maps of 5 consecutive frames are accumulated to obtain the aggregated heat map. Then, the aggregated heat map is processed by morphological closing operations, such as first performing an expansion operation and then an erosion operation to fill the holes inside the motion area and remove isolated noise points to obtain a more complete and continuous motion area. Finally, a connected region labeling algorithm, such as the eight-neighborhood connected region labeling algorithm, is used to identify and annotate independent motion regions in the aggregated heat map, and a unique identifier is assigned to each independent motion region to facilitate subsequent target tracking and behavior analysis;
[0108] Compared with the target detection method based on deep learning, dynamic frame difference calculation has lower complexity, less resource consumption, faster processing speed, and can meet the emergency processing needs with high real-time requirements in urban rail transit scenarios. Through Gaussian blur, dynamic threshold adjustment, spatiotemporal aggregation, and morphological processing, the environmental noise interference is effectively suppressed, and the accuracy and robustness of motion detection are improved. Compared with the traditional fixed threshold frame difference method, the environment adaptive threshold adjustment can better adapt to the complex and changeable rail transit environment and improve the environmental adaptability of the system.
[0109] S3: Build a 3D digital model based on the engineering drawings of the platform and carriage of urban rail transit, perform internal and external parameter calibration of the camera, and establish a mapping relationship between the 2D motion heat map and the 3D digital model;
[0110] In one embodiment provided in the present application, in order to map the motion area in the two-dimensional image to the three-dimensional space of urban rail transit and associate it with the actual equipment components, a three-dimensional digital model is constructed based on the CAD engineering drawings of the urban rail transit platform and carriage, the internal and external parameters of the camera are calibrated, and a mapping relationship between the two-dimensional motion heat map and the three-dimensional digital model is established;
[0111] First, based on the CAD engineering drawings of the rail transit platform and carriage, use 3D modeling software, such as AutoCAD or Revit, to build an accurate 3D digital model. The model should contain the precise spatial position and geometric parameters of all fixed facilities inside the platform and carriage, such as escalators, gates, seats, platform doors, etc. Give each equipment component a unique ID and function label, such as "Escalator-001", "Gate-Area A-002", "Platform Door-Platform 3-005", etc., and establish an equipment component database to store the ID, function label, 3D model data and other information of the equipment components.
[0112] Then, use a calibration plate, such as a checkerboard calibration plate, to shoot at the camera installation location, and use the camera calibration algorithm to estimate the internal and external parameters of the camera. The internal parameters include the focal length, principal point coordinates, distortion coefficient, etc. of the camera, and the external parameters include the rotation matrix and translation vector of the camera in the world coordinate system. Through camera calibration, the internal and external parameter matrices of the camera are obtained to provide parameter basis for subsequent coordinate system conversion.
[0113] Next, establish the conversion relationship between the camera perspective and the CAD model space coordinate system. Take the world coordinate system of the CAD model as the reference coordinate system, and transform the camera coordinate system to the world coordinate system through the external parameter matrix. Further, use the internal parameter matrix to project the three-dimensional points in the world coordinate system to the two-dimensional image plane, and establish the corresponding relationship between the three-dimensional space points and the two-dimensional image pixels.
[0114] Finally, an affine transformation is applied to map the motion area in the two-dimensional heat map to the three-dimensional CAD model. For each motion area in the two-dimensional heat map, its pixel coordinate range in the image coordinate system is calculated, and the pixel coordinate range is back-projected to the three-dimensional CAD model using the established two-dimensional-three-dimensional coordinate mapping relationship to obtain the position and range of the motion area in three-dimensional space. Through spatial occupancy detection, it is determined whether the motion area overlaps with the preset spatial position of the equipment component, thereby identifying the motion area that overlaps with the spatial position of the equipment component;
[0115] Through geometric mapping technology, accurate conversion of two-dimensional image information and three-dimensional spatial information is achieved, and moving targets detected in two-dimensional images are accurately located in three-dimensional urban rail transit scenes. The positioning accuracy reaches ±5cm, which meets the positioning accuracy requirements of urban rail transit emergency response. This method does not rely on deep learning training data and does not require a large amount of annotated data for model training, which reduces the development and maintenance costs of the system, reduces the system's dependence on annotated data, and improves the system's versatility and scalability. By comparing the spatial position of the equipment components in the CAD model, the association between the motion area and the equipment components is achieved, laying the foundation for subsequent abnormal behavior analysis based on equipment components.
[0116] S4: Detect abnormal behavior based on the mapping results of the two-dimensional motion heat map and the three-dimensional digital model after spatiotemporal aggregation, and perform multimodal fusion analysis in combination with sound data and equipment operation data to determine the severity of abnormal events;
[0117] In one embodiment provided in the present application, abnormal behavior detection is performed on the motion area mapped to the three-dimensional CAD model, and multimodal fusion analysis is performed in combination with sound data and equipment operation data to improve the accuracy and robustness of anomaly detection, and trigger corresponding warnings according to the severity of the abnormal event.
[0118] First, for each equipment component ID, calculate its movement frequency within the time window w. The time window w can be set according to the actual application scenario, for example, set to 1 second or 5 seconds. The movement frequency can be calculated by counting the number or area change rate of the movement area overlapping with the spatial position of the equipment component within the time window w. Further, the frequency spectrum characteristics of the time series signal of the equipment component movement are analyzed by Fourier transform, and the main frequency and amplitude of the movement are extracted. Fourier transform can effectively analyze periodic motion, such as the cyclic motion of escalator steps, the periodic switching of gate swing arms, etc.
[0119] Then, based on the device type, a table of abnormal status thresholds is predefined. The threshold table stores the abnormal status thresholds of different types of equipment components, for example:
[0120] Escalator steps: movement frequency > 5 Hz or movement amplitude > 3 times the standard deviation;
[0121] Gate: Movement frequency > 2Hz or movement range exceeds normal working range by 20%;
[0122] Platform door: motion detection with movement frequency > 1Hz or abnormal opening and closing periods;
[0123] When it is detected that the movement frequency or amplitude of an equipment component exceeds its corresponding abnormal state threshold, it is preliminarily determined that the equipment component has an abnormal state and the alarm of the corresponding component is triggered.
[0124] In order to further improve the accuracy of anomaly detection and reduce the false alarm rate, multimodal anomaly confirmation is performed. Fusion of collected visual anomaly scores , Sound Abnormality Score and equipment operation data anomaly score Make a comprehensive judgment. Visual abnormality score It can be calculated based on the motion frequency and amplitude. For example, when the motion frequency or amplitude exceeds the threshold, a higher visual anomaly score is assigned. It can be obtained by analyzing the sound data collected by the sound sensor. For example, when screams, cries for help, abnormal equipment noise, etc. are detected, a higher sound anomaly score is given. Equipment operation data anomaly score It can be obtained based on the analysis of equipment operation data, such as abnormal escalator speed, gate jam, abnormal opening of platform door, etc., and a higher equipment operation data abnormality score can be given.
[0125] Multimodal exception confirmation function Defined as a weighted sum model:
[0126] .
[0127] in, , , is the weight coefficient, which is set according to the importance of different modal data in anomaly detection. In this embodiment, it is set to =0.5, =0.3, =0.2. Tc is a comprehensive threshold used to determine whether an abnormal event is finally confirmed. In this embodiment, Tc is set to 0.7. > Tc, it is finally confirmed that an abnormal event has occurred in the equipment component id and an early warning is triggered.
[0128] By analyzing the movement frequency of equipment components, it is possible to effectively detect abnormal movement states of equipment components, such as abnormal shaking of escalators, jamming of gates, abnormal opening of platform doors, etc. The multimodal anomaly detection mechanism integrates visual, sound, and equipment operation data, making full use of multi-source information, improving the accuracy and robustness of anomaly detection, and effectively reducing the false alarm rate. By setting abnormal state thresholds for different equipment types, refined anomaly judgments based on different equipment characteristics are achieved, improving the flexibility and adaptability of the system. Compared with single-modal anomaly detection methods, multimodal fusion strategies can more comprehensively and accurately identify abnormal events in urban rail transit environments, improving the reliability and safety of the system.
[0129] S5: Develop emergency response strategies based on the severity of abnormal events, send control instructions to rail transit equipment through industrial protocol interfaces, and push abnormal event information to the control center, patrol personnel, and passenger terminals.
[0130] In one embodiment provided in the present application, based on the determined abnormal events, they are graded according to the severity of the abnormalities, and a corresponding emergency response strategy is formulated, control instructions are sent to rail transit equipment through an industrial protocol interface, and abnormal event information is pushed to the control center, patrol personnel and passenger terminals.
[0131] According to the severity of the abnormal event, the emergency response is divided into three levels:
[0132] Level 1 response (minor abnormality): For example, slight jitter of equipment components, occasional abnormal noise, etc. are judged as minor abnormalities, triggering the equipment self-check program to perform self-diagnosis and self-repair. It does not affect the normal operation of the equipment, but only sends an alarm message to the control center to prompt attention.
[0133] Level 2 response (obvious abnormality): For example, obvious abnormality in the movement frequency of equipment parts, sound alarms, equipment operation data outside the normal range, etc., are judged as obvious abnormalities, triggering equipment deceleration or current limiting measures, such as escalator slowdown, gate flow control, and platform door opening and closing speed reduction, etc., to reduce safety risks and send early warning information to the control center and patrol personnel to prompt manual intervention.
[0134] Level 3 response (serious abnormalities): For example, equipment components vibrate violently, make abnormal loud noises, or have serious abnormalities in equipment operating data or even shut down, which are identified as serious abnormalities and the equipment is stopped immediately, such as emergency braking of escalators, closing of all gates, emergency stop of platform doors, etc., to maximize passenger safety, activate emergency plans, send emergency warning information to the control center, patrol personnel and passenger terminals, guide passengers to evacuate, and notify professional maintenance personnel to handle the situation on site.
[0135] According to the preset emergency response strategy, the system sends control instructions to rail transit equipment through industrial protocol interfaces such as Modbus TCP / IP and Profinet to achieve real-time intervention and control of the equipment. For example, it sends deceleration or stop instructions to escalator controllers, close instructions to gate controllers, and emergency stop instructions to platform door controllers.
[0136] At the same time, the abnormal event information, including abnormal level, abnormal type, occurrence time, occurrence location, equipment component ID, etc., is pushed to the urban rail transit control center through network communication protocols, such as MQTT, HTTP, etc., for unified monitoring and management by the control center personnel. The early warning information is pushed to the mobile terminal of the patrol personnel, such as a handheld PDA or smart phone, so that the patrol personnel can quickly arrive at the scene for disposal. In addition, abnormal event information and evacuation guidance information can be released to passengers through passenger information display screens, broadcasting systems, mobile phone APPs and other passenger terminals on platforms and carriages to ensure the passengers' right to know and safety.
[0137] It realizes graded response according to the severity of abnormal events, and can take differentiated emergency response measures for abnormal events of different levels, taking into account both safety and operational efficiency. Through the real-time control of rail transit equipment through industrial protocol interface, it realizes automated emergency response, improves the timeliness and effectiveness of emergency treatment, and reduces the need for manual intervention. The multi-channel information push mechanism ensures that abnormal event information can be transmitted to the control center, inspection personnel and passengers in a timely manner, improves the collaborative efficiency of emergency response and the self-rescue ability of passengers, and maximizes the safe and stable operation of the urban rail transit system.
[0138] Preferably, the collecting of visual data, sound data and equipment operation data includes:
[0139] Wide-angle cameras are deployed at the escalator entrances, security check areas, waiting areas, and both ends of the carriages, with the acquisition frequency set to 10 frames per second;
[0140] Use microphone arrays to collect environmental sound data and extract sound features, including frequency, energy, and spectrum changes;
[0141] The equipment operation data is collected in real time through the industrial protocol interface of rail transit equipment. The equipment operation data includes speed, current, temperature and vibration.
[0142] In one embodiment provided in the present application, wide-angle cameras are deployed in key areas of the platform. Specifically, wide-angle cameras with a resolution of 640×480 and a viewing angle of 130° are installed in crowded areas such as escalator entrances, security check areas, and waiting areas. 2-3 cameras are configured in each key area to achieve redundant coverage and improve the reliability of monitoring. At the same time, wide-angle cameras of the same specifications are installed at both ends of each carriage in the carriage to ensure that the monitoring angle covers all passenger areas. The acquisition frequency of all cameras is uniformly set to 10 frames per second. This setting can not only meet the needs of motion detection, but also effectively reduce the burden of data transmission.
[0143] Microphone arrays are installed on platforms and in carriages to collect environmental sound data. The microphone arrays are arranged linearly, with each array containing 8 high-sensitivity microphones spaced 10 cm apart. Sound features, including frequency, energy, and spectrum changes, are extracted from the collected raw sound data through sound processing algorithms. Frequency features reflect the pitch of the sound, energy features indicate the intensity of the sound, and spectrum changes reflect the time-varying characteristics of the sound. The combination of these features can effectively identify abnormal sound events, such as screams, collisions, or explosions.
[0144] Finally, the equipment operation data is collected in real time through the industrial protocol interface of the rail transit equipment. Specifically, for different types of equipment, the collected data items are as follows:
[0145] For escalators and moving walkways: collect motor speed, current, temperature and vibration data;
[0146] For platform screen doors: collect switch status, motor current and door position data;
[0147] For air conditioning systems: collect temperature, humidity, wind speed and energy consumption data;
[0148] For trains: collect speed, acceleration, door status and passenger load data.
[0149] The operating data of these devices are transmitted via standardized industrial Ethernet protocols with a sampling frequency of 100Hz to ensure the real-time and accuracy of the data.
[0150] Preferably, the preprocessing of the visual data, the sound data and the equipment operation data and achieving timestamp synchronization comprises:
[0151] Gaussian blurring of visual data;
[0152] Perform noise reduction on sound data;
[0153] Normalize the equipment operation data;
[0154] The network time protocol or precision time protocol is used to synchronize the timestamps of the pre-processed visual data, sound data and equipment operation data to ensure the accuracy of data fusion.
[0155] In one embodiment provided in the present application, Gaussian blur processing is performed on the visual data collected from the wide-angle camera. Specifically, a 5×5 Gaussian kernel is used to perform a convolution operation on each frame of the image, and the standard deviation of the Gaussian kernel is set to 1.5. Gaussian blur processing can effectively reduce the noise and details in the image, making the subsequent dynamic frame difference processing more stable and reliable. The processed image It can be expressed as:
[0156] ;
[0157] in, is the original image, is a Gaussian kernel, and * indicates a convolution operation.
[0158] For the sound data collected by the microphone array, an adaptive filtering algorithm is used to perform noise reduction. The specific steps are as follows:
[0159] a) Use Fast Fourier Transform to convert the time domain signal into the frequency domain signal;
[0160] b) estimate the power spectral density of the background noise;
[0161] c) Calculate the signal-to-noise ratio of each frequency component based on the estimated noise power spectral density;
[0162] d) Design Wiener filter based on signal-to-noise ratio;
[0163] e) applying the Wiener filter to the spectrum of the original signal;
[0164] f) Convert the processed spectrum back to time domain signal using inverse fast Fourier transform.
[0165] This noise reduction method can effectively suppress background noise while retaining key sound feature information.
[0166] The operating data collected from rail transit equipment, including speed, current, temperature and vibration, are normalized. The purpose of normalization is to unify data of different dimensions to the same scale to facilitate subsequent multimodal fusion analysis. The minimum-maximum normalization method is used, and the processing formula is as follows:
[0167]
[0168] Among them, X is the original data, and are the minimum and maximum values of this type of data, respectively. is the normalized data. The normalized data range is unified to [0, 1].
[0169] To ensure the accuracy of multimodal data fusion, the precise time protocol is used to synchronize the timestamps of the preprocessed visual data, sound data, and equipment operation data. The specific steps are as follows:
[0170] a) Set up a master clock server in the system as the time reference;
[0171] b) Each data acquisition device acts as a slave device and synchronizes time with the master clock server;
[0172] c) The master clock server periodically sends synchronization messages to the slave devices, including the sending timestamp;
[0173] d) Receive synchronization messages from the device and record the receiving timestamp;
[0174] e) The slave device sends a delay request message to the master clock server;
[0175] f) The master clock server receives the delay request message and replies with a delay response message containing a receiving timestamp;
[0176] g) The slave device calculates the time deviation and network delay from the master clock server based on the received timestamp information;
[0177] h) The slave device adjusts the local clock according to the calculation results to achieve synchronization with the master clock server.
[0178] Through the PTP protocol, the clock synchronization accuracy of each device in the system can be controlled at the microsecond level, ensuring the time consistency of multimodal data.
[0179] Through Gaussian blur processing, the image noise is effectively reduced, and the accuracy of subsequent dynamic frame difference is improved. In specific implementation, this method can reduce the false detection rate by about 30%; the adaptive filtering algorithm is used for noise reduction processing, which can effectively suppress background noise while retaining key sound feature information. Test results show that this method can improve the signal-to-noise ratio by about 6dB and significantly improve the quality of sound feature extraction; through normalization processing, data of different dimensions are unified into the range of [0, 1], which is convenient for subsequent multimodal fusion analysis. This processing method can eliminate the influence of dimensional differences on the analysis results and improve the accuracy of anomaly detection; the precise time protocol (PTP) is used to achieve time synchronization of multimodal data, and the clock synchronization accuracy of each device in the system can be controlled at the microsecond level. This high-precision time synchronization ensures the accuracy of multimodal data fusion and lays the foundation for subsequent abnormal behavior detection and multimodal fusion analysis. In specific implementation, after adopting the PTP protocol, the time consistency error of multimodal data can be controlled within 1ms, and the synchronization accuracy is improved by about two orders of magnitude compared with the traditional network time protocol.
[0180] Preferably, performing dynamic frame difference processing on continuous video frames of visual data to generate a two-dimensional motion heat map includes:
[0181] The pixel differences between adjacent frames of continuous video frames are calculated, and the difference threshold is dynamically adjusted through the environment adaptive algorithm to obtain a difference image; the difference image is binarized to generate a binary motion heat map.
[0182] In one embodiment provided in the present application, each frame image in the input continuous video frame sequence is pre-processed by Gaussian blur. Gaussian blur is an effective image filtering technology, the purpose of which is to eliminate noise that may be introduced during the video acquisition process, such as illumination changes, sensor noise, etc. By applying a Gaussian filter, the high-frequency noise components in the image can be smoothed and the main structural information of the image can be retained, thereby improving the robustness and accuracy of subsequent motion detection. In this embodiment, a Gaussian kernel such as 3×3 or 5×5 can be selected to perform a convolution operation on each frame image to achieve image smoothing.
[0183] After completing the Gaussian blur preprocessing, for the continuous video frame sequence, the pixel value difference between the corresponding pixels of two adjacent frames is calculated. Frame Image and Frame Image in ≥2), calculate the position of each pixel between them The difference in pixel values Here, the absolute value operation is used to ensure that the difference value is non-negative, indicating the magnitude of the pixel value change. By calculating pixel by pixel, a differential image is obtained. ,The image reflects the changes in pixel values between adjacent frames, and the areas with large differences in pixel values usually correspond to the motion areas in the video scene.
[0184] In order to adapt to the complex and changeable lighting conditions and environmental noise in urban rail transit scenes, this embodiment introduces an environmental adaptive algorithm to dynamically adjust the binarization threshold. The traditional fixed threshold method is prone to false detection or missed detection when the lighting changes or the noise level fluctuates. The environmental adaptive algorithm can dynamically adjust the threshold according to the noise level of the current scene, thereby improving the accuracy and adaptability of motion detection.
[0185] For example, an adaptive threshold method based on statistical characteristics can be used. First, statistically analyze the difference image over a period of time (e.g., the most recent frames). Pixel value distribution, calculate the mean of the pixel values and standard deviation Then, the binarization threshold is dynamically set based on these statistics. One possible threshold setting strategy is , where k is an adjustable parameter used to control the sensitivity of the threshold. The parameter k can be adjusted according to the actual application scenario. For example, in an environment with less noise, the k value can be appropriately reduced to improve the sensitivity; in an environment with more noise, the k value can be appropriately increased to reduce the false detection rate. In this way, the threshold can be achieved. Adaptive adjustment to the ambient noise level.
[0186] Using dynamically adjusted thresholds , for the difference image Perform binarization to generate binary motion heatmap For each pixel in the difference image , its pixel value Adaptive threshold with current frame For comparison. , then the pixel is considered to correspond to the motion area. Set its pixel value to a highlight value (such as white or pixel value 255); if , then the pixel is considered to correspond to the background area. The pixel value is set to a low value (such as black or pixel value 0). Binarization is performed using the following formula:
[0187] when , =255; when , =0;
[0188] After binarization, the resulting binary image This is a two-dimensional motion heat map. In this heat map, the white area represents the detected motion area, and the black area represents the static background area.
[0189] This embodiment uses the above-mentioned dynamic frame difference technology to effectively extract the motion area from the video data and generate a two-dimensional motion heat map. Compared with traditional motion detection methods such as background modeling or optical flow, dynamic frame difference technology has the significant advantage of low computational complexity, and is particularly suitable for urban rail transit emergency response scenarios with high real-time requirements. The computational complexity of this step is reduced by about 85% compared to the deep learning target detection method, and the processing speed is increased to the real-time level (<50ms / frame). This means that under the condition of limited computing power of edge computing nodes, this embodiment can still ensure that the system can quickly process video data and detect abnormal motion in time, thereby buying valuable time for subsequent emergency response.
[0190] In addition, by introducing an environment-adaptive threshold adjustment mechanism, this embodiment improves the robustness of the motion detection algorithm to environmental changes, reduces the false detection rate and missed detection rate caused by illumination changes or noise interference, and ensures the quality and reliability of the motion heat map.
[0191] Preferably, the spatiotemporal aggregation of the two-dimensional motion heat map comprises:
[0192] Aggregate five consecutive frames of 2D motion heatmaps in the time dimension;
[0193] The aggregated binary motion heat map is spatiotemporally aggregated, and the interference points are removed by applying the morphological closing operation. The independent motion regions are identified and labeled by the connected region labeling algorithm, and a unique identifier is assigned.
[0194] In one embodiment provided in the present application, five consecutive frames of binary motion heat maps are recorded as H1, H2, H3, H4 and H5 respectively. These heat maps reflect the motion changes in the monitoring area in a continuous time period. In order to enhance the significance of the moving target and suppress short-term noise interference, the five consecutive frames of binary motion heat maps are first aggregated in the time dimension; the specific method is to perform a pixel-by-pixel logical OR operation on these five frames of heat maps.
[0195] The aggregation formula is:
[0196] H_aggregate = H1 OR H2 OR H3 OR H4 OR H5;
[0197] Here, “OR” represents a logical OR operation. If a pixel is marked as motion (pixel value is 1) in any of the five heatmaps, the pixel is also marked as motion in the aggregated heatmap.
[0198] By aggregating in the temporal dimension, we can effectively retain continuously moving targets and filter out short-lived, isolated noise points. For example, if a pedestrian is continuously moving, even if some pixels are not detected in some frames due to lighting changes or occlusion, the aggregation operation can still ensure that the overall motion area of the pedestrian is preserved.
[0199] After temporal aggregation, there may be some small holes or breaks in the heatmap, which may be caused by shadows or low-contrast areas inside the object. In order to fill these holes and connect adjacent motion regions, a morphological closing operation is applied.
[0200] Morphological closing operations include dilation and erosion:
[0201] Dilation involves dilating the aggregated heatmap using a structural element (e.g., a 3x3 square). The dilation operation will expand the edges of the motion regions outward and fill small holes.
[0202] Erosion involves using the same structural element to erode the dilated heatmap. The erosion operation shrinks the edges of the motion region inward and eliminates small noise points.
[0203] The morphological closing operation can effectively fill the holes inside the moving region and connect adjacent moving regions, thus making the contour of the moving target more complete and smooth. For example, if a pedestrian’s arm is separated from the body during movement, the closing operation can connect the arm and the body to form a complete moving region.
[0204] After the morphological closing operation, there may be multiple independent motion regions in the heat map, each of which corresponds to one or more moving targets. In order to distinguish these different moving targets, the connected region labeling algorithm is applied; the connected region labeling algorithm classifies all adjacent moving pixels (pixel value 1) in the heat map into a connected region and assigns a unique identifier to each connected region. Commonly used connected region labeling algorithms include the four-neighborhood algorithm and the eight-neighborhood algorithm.
[0205] The connected region labeling algorithm can distinguish different moving targets in the heat map, providing a basis for subsequent analysis and processing. For example, if there are multiple pedestrians in the heat map, the connected region labeling algorithm can mark each pedestrian as an independent moving region, so that the behavior of each pedestrian can be analyzed later.
[0206] After completing the connected region labeling, a unique identifier is assigned to each independent motion region. The identifier can be an integer or a string. The identifier is used to reference and operate a specific motion region in subsequent processing; by assigning a unique identifier, each motion region can be easily tracked and analyzed. For example, the identifier, position, size, and motion trajectory of each motion region can be recorded to facilitate subsequent analysis and prediction of the behavior of the moving target.
[0207] Preferably, the construction of the three-dimensional digital model based on the engineering drawings of the platform and carriage of the urban rail transit includes:
[0208] Based on the CAD engineering drawings of the platform and carriages, an accurate three-dimensional digital model is constructed, which includes the spatial position and geometric parameters of fixed facilities such as escalators, gates and seats; a unique ID and function label are given to each equipment component of the fixed facility to build an equipment component database.
[0209] In one embodiment provided in the present application, first, CAD engineering drawings of urban rail transit platforms are obtained. These drawings should contain detailed structural information of the platform, including but not limited to:
[0210] Overall layout of the platform: the overall dimensions of the platform such as length, width, height, etc.
[0211] Fixed facility locations: spatial location coordinates of fixed facilities such as escalators, gates, waiting seats, platform doors, columns, walls, etc.
[0212] Geometric parameters: Detailed geometric dimensions of each fixed facility, such as the step height and width of escalators, the channel width of gates, the length, width, height of seats, etc.
[0213] Material information: Material types of different facilities, such as metal, glass, concrete, etc.
[0214] The data format of CAD engineering drawings can be common formats such as DWG and DXF.
[0215] Import CAD engineering drawings into 3D modeling software, such as AutoCAD, SolidWorks, Blender, etc. According to the layer information of the CAD drawings, import different types of facilities into the software separately; ensure that all imported facility models are in the same coordinate system. Usually, the global coordinate system is established with the center of the platform or a fixed point as the origin; according to the dimension annotations on the CAD drawings, the imported model is calibrated accurately to ensure that the size of the model is consistent with the actual size. The measurement tools provided by the software can be used for verification; the model is improved in detail, such as adding escalator textures, gate indicator lights, seat backs, etc. These details can improve the realism and recognizability of the model; optimize the model to reduce the number of faces and vertices of the model to increase the rendering speed of the model and reduce storage space. The optimization tools provided by the software can be used for processing.
[0216] Assign a unique ID to each fixed facility equipment component. The naming rules of the ID can be defined according to the facility type and location, such as "Escalator-001", "Gate-Area A-002", etc.
[0217] Define one or more function tags for each device component to describe the function of the component. For example:
[0218] Escalator: entrance, exit, steps, handrails
[0219] Gate: channel, card swiping area, display screen, baffle
[0220] Seat: seat, backrest, armrests
[0221] Platform door: door body, glass, sensor
[0222] The ID, function label and corresponding 3D model information of the equipment components are stored in the database. The database can use a relational database. The constructed 3D digital model is exported to a common 3D model format, such as OBJ, FBX, STL, etc. These formats can be read and used by other software or systems.
[0223] In this way, an accurate three-dimensional digital model of the urban rail transit platform can be constructed. The model contains the precise spatial position and geometric parameters of all fixed facilities in the platform, providing an accurate reference for subsequent motion heat map mapping. The positioning accuracy can reach ±5cm; each equipment component in the model has a unique ID and function label, which is convenient for the system to identify and analyze. For example, the system can identify whether a certain movement area is located on the steps of the escalator based on the ID, so as to determine whether there are safety hazards; the model can be easily expanded and updated. For example, when adding or replacing equipment on the platform, only the CAD engineering drawings need to be updated and the model can be rebuilt; it does not rely on deep learning training data, which reduces the system's dependence on labeled data and reduces the system's development and maintenance costs; combined with the three-dimensional digital model, abnormal behavior can be judged more accurately. For example, it can be determined whether a certain movement area exceeds the normal working range of the gate, thereby improving the accuracy of anomaly detection.
[0224] Preferably, the step of performing camera internal and external parameter calibration and establishing a mapping relationship between a two-dimensional motion heat map and a three-dimensional digital model includes:
[0225] Use the calibration board to perform camera internal and external parameter calibration to obtain the camera's internal and external parameter matrix;
[0226] Based on the internal and external parameter calibration results, an affine transformation is applied to map the motion area in the two-dimensional motion heat map to the spatial coordinate system of the three-dimensional digital model;
[0227] Through the spatial occupancy detection algorithm, the motion area that overlaps with the equipment parts with unique IDs in the 3D model is identified.
[0228] In one embodiment provided in the present application, first, the camera internal and external parameters are calibrated using a calibration plate to obtain the camera's internal and external parameter matrix. Specifically, a 9×6 black and white checkerboard calibration plate is used, and the calibration plate is placed at different angles and distances in the camera's field of view and at least 20 images are taken. Then, the camera's internal parameter matrix is calculated using Zhang Zhengyou's calibration method. and the external parameter matrix Among them, the internal parameter matrix Contains parameters such as focal length, principal point coordinates and distortion coefficient; external parameter matrix Describes the rotation R and translation t of the camera relative to the world coordinate system.
[0229] Next, based on the internal and external parameter calibration results, an affine transformation is applied to map the motion area in the two-dimensional motion heat map to the spatial coordinate system of the three-dimensional digital model. The specific steps are as follows:
[0230] (1) For each pixel in the 2D motion heat map , using the inverse matrix of the internal parameter matrix K Convert it to normalized image coordinates : ;
[0231] (2) Using the external parameter matrix Convert normalized image coordinates to 3D points in the world coordinate system ;
[0232] Among them, Z is the distance from the point to the camera, which can be obtained by intersection calculation with the three-dimensional digital model.
[0233] (3) The obtained three-dimensional point in the world coordinate system Align with the coordinate system of the three-dimensional digital model to complete the mapping of the two-dimensional motion heat map to the three-dimensional space.
[0234] Finally, the spatial occupancy detection algorithm is used to identify the motion area that overlaps with the equipment parts with unique IDs in the 3D model in terms of spatial position. The specific implementation method is as follows:
[0235] (1) For each device component in the 3D digital model, a bounding box is constructed based on its geometric parameters.
[0236] (2) For each motion area mapped to the three-dimensional space, calculate its intersection with the bounding box of each device component.
[0237] (3) If a motion area intersects with the bounding box of a device component, it is considered that the motion area and the device component overlap in spatial position.
[0238] (4) Record the corresponding relationship between the overlapping motion areas and the IDs of the equipment components for subsequent abnormal behavior analysis.
[0239] Through camera calibration technology, accurate mapping between two-dimensional images and three-dimensional space is achieved, and the positioning accuracy can reach ±5cm, providing a reliable spatial information basis for subsequent abnormal behavior detection; the affine transformation method is used for coordinate conversion, with low computational complexity, and can be processed in real time on edge computing devices to meet real-time requirements; the space occupancy detection algorithm can effectively identify the motion areas related to specific equipment components, providing an important basis for the precise positioning and classification of abnormal behaviors; the deep learning method that does not rely on a large amount of labeled data reduces the system's dependence on training data and improves the system's versatility and scalability; by establishing a correspondence between the motion area and the equipment component ID, it provides structured input for subsequent multimodal anomaly detection and emergency response, which is conducive to improving the system's decision-making accuracy and response speed.
[0240] Preferably, the abnormal behavior detection on the mapping result of the two-dimensional motion heat map and the three-dimensional digital model after spatiotemporal aggregation includes:
[0241] For the motion area that spatially overlaps with the equipment component with a unique ID, the motion frequency of the equipment component within the preset time window is calculated, and the spectral characteristics of the motion are analyzed through Fourier transform to extract the motion frequency and amplitude. An abnormal state threshold table is preset based on the equipment type, and an alarm for the corresponding component is triggered when the motion frequency or amplitude is detected to exceed the threshold.
[0242] In one embodiment provided in the present application, an escalator at an urban rail transit platform is used as an object to perform abnormal behavior detection on escalator equipment components, thereby realizing emergency handling of urban rail transit.
[0243] When the wide-angle camera on the platform detects the movement of passengers at the entrance of the escalator, the motion heat map will mark the movement area of the passenger. This two-dimensional movement area is then converted into a three-dimensional digital model, and space occupancy detection is performed. The system pre-assigns unique equipment component IDs to various components of the escalator, such as steps, handrails, aprons, etc. in the three-dimensional digital model (for example, the step ID is 'Escalator_Step_001', and the handrail ID is 'Escalator_Handrail_001'). Through the space occupancy detection algorithm, the system can identify whether the movement area overlaps with these escalator equipment components with unique IDs in spatial position. The precise positioning of the moving target in three-dimensional space is achieved, and the association between the moving target and the specific equipment component is established, which lays the foundation for the subsequent abnormal behavior analysis of specific equipment components. Compared with the traditional method of performing anomaly detection only in two-dimensional image space, the present invention can more accurately locate abnormal events to specific equipment components, thereby improving the pertinence and effectiveness of emergency treatment.
[0244] For the escalator step equipment component (ID: 'Escalator_Step_001') that is identified as spatially overlapping with the motion area, the system will further analyze its motion state within the preset time window. In this embodiment, the preset time window w is set to 5 seconds. The number of times the motion area overlaps with the step equipment component in space within these 5 seconds is recorded as the motion frequency of the step equipment component. In order to more comprehensively analyze the characteristics of the step motion, the system further uses Fourier transform to perform spectrum analysis on the motion data of the step equipment component within the time window w. Fourier transform can decompose the motion signal in the time domain into spectra of different frequency components and extract the main frequency and amplitude of the motion. For example, the normal operation of the escalator step motion should show a regular low-frequency motion spectrum, and its main frequency corresponds to the normal operating speed of the escalator. When the escalator step has abnormal jitter or jamming, its motion spectrum may show abnormal increase in high-frequency components or main frequency amplitude. By calculating the motion frequency and performing spectrum analysis, the system can accurately describe the motion state of the equipment component from both the time domain and the frequency domain. The application of Fourier transform can effectively extract the characteristic information of motion signals, enabling the system to identify subtle abnormal movements that are difficult to detect with the naked eye, such as slight shaking or irregular movement of steps, thereby improving the sensitivity and accuracy of abnormality detection.
[0245] In order to achieve refined abnormality judgment for different equipment components, a table of abnormal status thresholds for equipment types is pre-established. For the escalator step equipment component (ID: 'Escalator_Step_001'), the corresponding abnormal state thresholds are preset in the threshold table. For example, for the escalator step, the preset abnormal state thresholds may include:
[0246] Frequency threshold: The normal operating frequency range is 0.1Hz - 0.5Hz. When the detected step movement frequency exceeds this range, for example, exceeds 0.6Hz (indicating that the step running speed is abnormally accelerated) or is lower than 0.05Hz (indicating that the step running speed is abnormally slowed down or stagnated), it is judged as a frequency abnormality.
[0247] Amplitude threshold: Under normal operating conditions, the step movement amplitude should be kept within ±2 times the standard deviation. When the detected main frequency amplitude exceeds 3 times the standard deviation, it is judged as abnormal amplitude, which may indicate abnormal and violent vibration of the step.
[0248] The calculated escalator step movement frequency and amplitude are compared with the preset abnormal state threshold. When the movement frequency is detected to be greater than 0.6Hz or less than 0.05Hz, or the movement amplitude exceeds 3 times the standard deviation, the system will determine that the escalator step equipment component is abnormal and trigger the alarm of the corresponding component.
[0249] The application of the preset abnormal state threshold table enables the system to set targeted abnormal detection standards according to the operating characteristics of different equipment components, avoiding the extensive method of using a unified threshold for abnormal detection, and significantly improving the accuracy and reliability of abnormal detection. By setting multi-dimensional thresholds such as frequency and amplitude, it can more comprehensively cover various abnormal states that may occur in the equipment, reducing the probability of missed reports and false alarms.
[0250] When it is detected that the movement frequency or amplitude of the escalator step equipment component (ID: 'Escalator_Step_001') exceeds the preset threshold, the system will immediately trigger an alarm signal for the component. The alarm signal will contain the ID information of the equipment component, the abnormality type (for example, frequency abnormality, amplitude abnormality), the abnormality level and other information, and will be transmitted to the emergency response terminal and central server of the execution layer. The triggering of the alarm signal is a direct output of the abnormal behavior detection result, which provides important abnormal information basis for subsequent emergency response and execution control steps. Through precise positioning of equipment components and detailed description of abnormal information, it can provide timely and accurate decision support for emergency response personnel and enhance the intelligent emergency response capabilities of urban rail transit systems.
[0251] Preferably, the combining of the sound data and the equipment operation data for multimodal fusion analysis to determine the severity of the abnormal event includes:
[0252] Based on the weighted fusion of visual anomaly score, sound anomaly score and equipment operation data anomaly score, the comprehensive anomaly score calculation formula is:
[0253]
[0254] in , , The visual data , sound data and equipment operation data The weight coefficient of ;
[0255] When the combined anomaly score exceeds the threshold, the anomaly event is confirmed.
[0256] In one embodiment provided in the present application, it is assumed that the system is applied at an escalator of a certain urban rail transit station. When the escalator is running, the system continuously collects visual, sound and equipment operation data.
[0257] Through visual data analysis, abnormal movement on the escalator steps is detected, such as passengers falling or objects falling and getting stuck. After motion frequency analysis, it is determined that the escalator step movement frequency exceeds the threshold of 5Hz, and the visual anomaly score V(escalator step ID) = 0.6.
[0258] The microphone array detected abnormal metal friction sound during the escalator operation, with frequency and energy exceeding the normal range. The sound abnormality score A(escalator step ID) = 0.5.
[0259] Through the industrial protocol interface, the system detected that the current value of the escalator motor increased abnormally and the temperature rose slightly, but did not exceed the critical threshold. The equipment operation data abnormality score E (escalator step ID) = 0.2.
[0260] Substitute the above scores into the comprehensive anomaly score formula: C(escalator step ID) = 0.5 * V(escalator step ID) +0.3 * A(escalator step ID) + 0.2 * E(escalator step ID)
[0261] C(escalator step ID) = 0.5 * 0.6 + 0.3 * 0.5 + 0.2 * 0.2 = 0.3 + 0.15 +0.04 = 0.49
[0262] In this example, the calculated comprehensive abnormality score C (escalator step ID) = 0.49 is lower than the threshold value Tc = 0.7. Therefore, although both visual and sound data indicate that there may be an abnormality, after comprehensively considering the equipment operation data, the system judges it as a "minor abnormality" or "potential risk", which may trigger a first-level response (equipment self-check or slight deceleration), but will not immediately stop the escalator operation to avoid inconvenience to passengers caused by misjudgment.
[0263] Assume another situation: if the visual anomaly score V(escalator step ID) = 0.8 (for example, visual detection of a passenger being stuck), the sound anomaly score A(escalator step ID) = 0.7 (the friction sound becomes more intense), and the equipment operation data anomaly score E(escalator step ID) = 0.6 (the motor current and temperature both increase significantly).
[0264] C(escalator step ID) = 0.5 * 0.8 + 0.3 * 0.7 + 0.2 * 0.6 = 0.4 + 0.21 + 0.12 = 0.73
[0265] At this time, the comprehensive abnormality score C (escalator step ID) = 0.73, exceeding the threshold Tc = 0.7. The system will confirm that an "abnormal event" has occurred, and trigger the corresponding emergency response strategy according to the abnormality level (for example, in this case, it may be judged as a level 2 or 3 abnormality), such as immediately stopping the escalator, starting the alarm, notifying the control center and patrol personnel, etc.
[0266] Single visual detection may be disturbed by factors such as lighting changes, shadows, and occlusions, leading to false alarms. By fusing sound and device operation data, they can corroborate each other and eliminate accidental visual interference. For example, even if abnormal motion is detected visually, if both sound and device operation data are normal, the possibility of false alarms can be reduced. The multimodal anomaly detection mechanism can reduce the false alarm rate to less than 1%. Sound data and device operation data can provide deeper abnormal information. For example, mechanical failures and electrical anomalies inside the equipment may be difficult to observe directly visually, but can be reflected by abnormal sound and abnormal device operation parameters. Multimodal fusion can comprehensively utilize the advantages of various data sources to improve the accuracy of anomaly detection. The accuracy can be increased to more than 95%. In a complex and changing environment, a single sensor may fail or its performance may degrade. Multimodal fusion can utilize the redundant information of multiple sensors to improve the robustness of the system. Even if the visual sensor is blocked or fails, the system can still rely on sound and device operation data for anomaly detection to ensure the reliable operation of the system. By setting refined abnormality thresholds for different types of equipment and components and combining multimodal data for comprehensive analysis, more refined abnormality determination can be achieved. For example, it is possible to distinguish between minor abnormalities, obvious abnormalities, and severe abnormalities of equipment, and adopt corresponding emergency response strategies according to different abnormality levels to improve the efficiency and accuracy of emergency handling.
[0267] Preferably, formulating emergency response strategies according to the severity of abnormal events includes:
[0268] Abnormal events are divided into three levels of response:
[0269] Level 1: Minor abnormality, triggering the device self-check program, without affecting the normal operation of the device;
[0270] Level 2: Obvious abnormality, triggering equipment deceleration or current limiting measures;
[0271] Level 3: Severe abnormality, immediately stop equipment operation and activate the emergency plan.
[0272] In one embodiment provided in this application, after completing the multimodal fusion analysis, the system will obtain a comprehensive anomaly score: , which reflects the abnormality of the device component ID. The system divides abnormal events into three levels:
[0273] Level 1 (minor abnormality): When 0 < When ≤ 0.3, it is considered as a minor abnormality. For example, a small number of passengers are temporarily stranded on the escalator steps, or the gate occasionally fails to swipe the card.
[0274] Level 2 (obviously abnormal): When 0.3 < When ≤ 0.7, it is judged as obvious abnormality. For example, the escalator running speed fluctuates slightly, the gate fails to swipe the card continuously, the platform door opening and closing speed is abnormal, etc.
[0275] Level 3 (Severe Abnormality): When the error is > 0.7, it is considered a serious abnormality. For example, the escalator suddenly stops running, the gate cannot be opened or closed, and the platform door cannot be closed normally.
[0276] For abnormal events of different levels, the system has pre-established corresponding emergency response strategies, which are stored in the emergency plan database. These strategies include:
[0277] Level 1 response: triggering the equipment self-check procedure. For example, the escalator control system automatically performs a self-check to check the working status of each sensor and motor and record the self-check results. At the same time, the system sends a prompt message to the control center to inform that there is a minor abnormality.
[0278] Secondary response: triggering equipment deceleration or current limiting measures. For example, when the escalator speed fluctuation is detected, the control system automatically reduces the escalator speed and limits the number of passengers passing through the escalator per unit time. At the same time, the system sends an alarm message to the control center and inspection personnel, indicating that there is an obvious abnormality and manual intervention is required.
[0279] Level 3 response: Immediately stop the equipment and start the emergency plan. For example, when the escalator is detected to stop suddenly, the control system immediately cuts off the escalator power supply to prevent accidents. At the same time, the system sends emergency alarm information to the control center, inspection personnel and passenger terminals, and starts the preset emergency plan, including evacuating passengers and arranging alternative transportation.
[0280] The system sends control instructions to rail transit equipment through industrial protocol interfaces (such as Modbus, Profinet, etc.) to achieve real-time intervention on the equipment. The specific content of the control instructions depends on the type and level of the abnormal event. For example:
[0281] For escalators, you can send start, stop, speed control and other commands.
[0282] For gate machines, you can send commands such as open, close, and lock.
[0283] For platform doors, commands such as open, close, and emergency unlock can be sent.
[0284] The system pushes abnormal event information to the control center, inspection personnel and passenger terminals so that relevant personnel can understand the situation in time and take corresponding measures. The information push methods include:
[0285] Control center: Displays detailed information of abnormal events through the monitoring screen, including the time, location, device type, abnormality level, processing status, etc.
[0286] Patrol personnel: receive alarm information through mobile terminals (such as mobile phone apps) and view on-site video surveillance so as to quickly arrive at the scene for processing.
[0287] Passenger terminal: Emergency notifications are issued through platform display screens, car display screens and mobile phone apps to inform passengers of abnormal events and guide passengers to evacuate safely.
[0288] Preferably, Figure 2 As shown, an urban rail transit emergency handling system based on a visual large model includes:
[0289] Data acquisition module, used to collect visual data, sound data and equipment operation data in key areas of urban rail transit platforms and carriages, and perform pre-processing, while achieving time stamp synchronization;
[0290] A visual analysis module, for performing dynamic frame difference processing on continuous video frames of visual data to generate a two-dimensional motion heat map, and performing spatiotemporal aggregation on the two-dimensional motion heat map;
[0291] The spatial mapping module is used to build a 3D digital model based on the engineering drawings of the platform and carriage of urban rail transit, perform internal and external parameter calibration of the camera, and establish a mapping relationship between the 2D motion heat map and the 3D digital model;
[0292] The anomaly detection module is used to detect abnormal behavior on the mapping results of the two-dimensional motion heat map and the three-dimensional digital model after spatiotemporal aggregation, and to perform multimodal fusion analysis in combination with sound data and equipment operation data to determine the severity of abnormal events;
[0293] The emergency response module is used to formulate emergency response strategies according to the severity of abnormal events, send control instructions to rail transit equipment through the industrial protocol interface, and push abnormal event information to the control center, patrol personnel and passenger terminals.
[0294] In one embodiment provided in the present application, the application of the urban rail transit emergency handling system based on the visual large model in the escalator entrance area of the urban rail transit platform is described, which is used to monitor and handle abnormal events at the escalator entrance to ensure passenger safety and orderly operation of rail transit.
[0295] The urban rail transit emergency response system based on visual big model is mainly composed of the following five modules: data acquisition module, visual analysis module, spatial mapping module, anomaly detection module and emergency response module.
[0296] The data acquisition module is deployed at the platform escalator entrance area. Its main function is to collect visual data, sound data and equipment operation data in the area, and perform preprocessing and timestamp synchronization.
[0297] Three sets of low-resolution wide-angle cameras are installed in a herringbone structure on the pillars at the entrance of the escalator. The model is Hikvision DS-2CD3145-I, with a resolution of 640×480 pixels and a lens viewing angle of 130°. The three sets of cameras monitor the escalator entrance area from different angles to ensure redundant coverage of the monitoring area and eliminate blind spots that may be caused by a single viewing angle. The image acquisition frequency of the camera is set to 10 frames per second, which can not only meet the real-time detection requirements of moving targets, but also effectively reduce the data processing pressure and network transmission bandwidth requirements of the system backend.
[0298] A microphone array, model Knowles SPH0641LM4H-1, is installed near the camera. The microphone array is used to collect environmental sound data in the escalator entrance area, such as passengers' cries for help, screams, and abnormal noises from equipment. The sound data acquisition frequency is set to 16kHz to effectively capture sound events. The sound data acquisition unit has a hardware noise reduction function to preliminarily filter out environmental background noise and improve the signal-to-noise ratio of effective sound signals.
[0299] Through the industrial Ethernet interface of the escalator control system, using the Modbus TCP / IP protocol, the escalator operation data is collected in real time, including the escalator's operating speed, motor current, key component temperature and vibration sensor data. These data reflect the real-time operating status of the escalator and provide device-level operating parameters for multi-modal anomaly detection.
[0300] First, the collected visual data is preprocessed by Gaussian blurring, using a 5×5 Gaussian kernel and a standard deviation σ set to 1.5 to effectively smooth the image and reduce the interference of image noise on the subsequent frame difference algorithm. For sound data, a noise reduction algorithm based on spectral subtraction is used to further suppress environmental noise. The equipment operation data is normalized, and data of different dimensions are uniformly scaled to the [0, 1] interval to eliminate the impact of dimensional differences on subsequent fusion analysis. After preprocessing, the Network Time Protocol (NTP) server is used to uniformly add accurate timestamps to all collected data streams, and the time synchronization accuracy reaches the millisecond level to ensure the accuracy of subsequent multimodal data fusion analysis.
[0301] The data acquisition module uses multiple sensors to work together to fully acquire visual, auditory and equipment operating status information at the escalator entrance area, and performs effective preprocessing and precise time synchronization, laying a high-quality data foundation for subsequent intelligent analysis. The use of low-resolution cameras significantly reduces the pressure of data processing and transmission while ensuring the monitoring effect, which meets the application requirements of edge computing.
[0302] The visual analysis module receives the visual data preprocessed by the data acquisition module. Its core function is to perform dynamic frame difference processing on continuous video frames, generate a two-dimensional motion heat map, and perform spatiotemporal aggregation on the motion heat map to extract the motion target area.
[0303] Receive the preprocessed video frame sequence, calculate the gray value difference of the corresponding pixel points for two consecutive frames, and obtain the frame difference image. In order to adapt to the influence of illumination changes and environmental noise, the frame difference threshold is dynamically adjusted using an environmental adaptive algorithm. Specifically, the mean and standard deviation of the pixel gray value of the current frame difference image are counted, and the threshold is set to the mean plus 1.5 times the standard deviation. This dynamic threshold adjustment mechanism can effectively deal with sudden illumination changes and noise interference, and improve the robustness of moving target detection.
[0304] The frame difference image is binarized, and the points with pixel grayscale values greater than the dynamic threshold are set to 255 (white), and the points less than the threshold are set to 0 (black), to generate a binary motion heat map. In the motion heat map, the white area represents the area where motion changes occur in the image.
[0305] In order to eliminate the influence of isolated noise points and enhance the continuity and integrity of the motion area, the binary motion heat maps of 5 consecutive frames are aggregated in the time dimension, that is, the pixel-level "OR" operation is performed on the 5 consecutive heat maps to obtain the aggregated heat map. Then, the morphological closing operation is applied to the aggregated heat map, using a 3×3 square structure element for closing operation, filling the holes inside the motion area, smoothing the edges of the motion area, and further removing noise interference.
[0306] The eight-neighborhood connected region labeling algorithm is used to perform connected region analysis on the binary motion heat map after spatiotemporal aggregation. The interconnected white pixel areas in the heat map are marked as independent motion regions, and a unique identifier (ID) is assigned to each independent motion region. At the same time, the geometric features of each motion region, such as area and centroid coordinates, are calculated.
[0307] The visual analysis module uses dynamic frame difference technology and spatiotemporal aggregation strategy to effectively detect and extract moving target areas from video streams and significantly reduce computational complexity. Compared with target detection methods based on deep learning, the frame difference algorithm has greatly reduced computational complexity and faster processing speed, which can meet real-time requirements. Spatiotemporal aggregation and morphological operations further improve the accuracy and robustness of motion detection, providing reliable moving target information for subsequent spatial mapping and abnormal behavior analysis.
[0308] The core function of the spatial mapping module is to construct a three-dimensional digital model of the platform escalator entrance area, perform camera calibration, establish a spatial mapping relationship between the two-dimensional motion heat map and the three-dimensional model, and accurately project the motion area in the two-dimensional image into the three-dimensional space.
[0309] Based on the CAD engineering drawings of the escalator entrance area of the rail transit platform, an accurate 3D digital model is constructed using 3D modeling software (such as AutoCAD). The model meticulously restores the geometric shape, spatial position and size parameters of all fixed facilities such as escalators, guardrails, and ground markings. Each equipment component in the model (such as each step of the escalator, handrail, and column at the escalator entrance) is given a unique ID and functional label, such as "Escalator Step_1", "Handrail_Left", "Escalator Entrance Column_A", etc., to build an equipment component database for subsequent refined abnormality analysis.
[0310] At the escalator entrance area, the Zhang Zhengyou calibration method was used to calibrate the internal and external parameters of each camera using a checkerboard calibration plate. The calibration process obtains the camera's internal parameter matrix K (including focal length, principal point coordinates, distortion parameters) and external parameter matrix (including rotation matrix R and translation vector t). The internal parameter matrix describes the camera's internal imaging parameters, and the external parameter matrix describes the camera's position and posture in the world coordinate system. The calibration accuracy is controlled at the sub-pixel level to ensure the accuracy of subsequent spatial mapping.
[0311] Based on the camera calibration results, the conversion relationship between the camera view and the world coordinate system of the 3D CAD model is established. The perspective projection model and affine transformation are used to map the coordinates of each pixel in the 2D motion heat map to the 3D model space. For each motion region output by the connected region marker, the pixel points on its boundary contour are transformed into 3D space coordinates to obtain the position, shape and size information of the motion region in 3D space.
[0312] Determine whether the motion area overlaps with the equipment components defined in the equipment component database in three-dimensional space. Determine whether the motion area enters the three-dimensional bounding box of a certain equipment component through spatial geometric relationship calculation. If overlap occurs, record the ID of the motion area and the corresponding equipment component for subsequent abnormal behavior analysis.
[0313] The spatial mapping module achieves accurate conversion of two-dimensional image information into three-dimensional spatial information through precise three-dimensional modeling and camera calibration, with a positioning accuracy of ±5cm. This module does not rely on deep learning training data, reduces the system's dependence on labeled data, and reduces deployment and maintenance costs. Through spatial occupancy detection, it can accurately identify which equipment components the moving target interacts with, laying the foundation for subsequent abnormal behavior analysis at the equipment component level.
[0314] The anomaly detection module receives the spatial mapping results of the motion area from the spatial mapping module, as well as the sound data and equipment operation data from the data acquisition module, and performs multimodal fusion analysis to detect abnormal events in the escalator entrance area.
[0315] For the identified motion areas that overlap with the equipment components, such as the motion area that overlaps with the "escalator step_1" component. For each equipment component ID, calculate its motion frequency within a preset time window (e.g., 1 second). The motion frequency is calculated by counting the number of motion areas that overlap with the equipment component within the time window. Furthermore, perform a fast Fourier transform on the motion frequency sequence of the equipment component, analyze the spectral characteristics of its motion, and extract the main frequency and amplitude of the motion.
[0316] Define the abnormal state threshold table in advance according to the equipment type and operation characteristics of the escalator For example, for the "Escalator Step" component, define the abnormal state threshold as: movement frequency > 5Hz or movement amplitude > 3 times the standard deviation. For the "Handrail" component, define the abnormal state threshold as: frequency > 2Hz or movement range exceeds the normal working range by 20%. For the "Escalator Entrance Column" component, define the abnormal state threshold as: frequency > 1Hz or motion detection during abnormal interaction periods (for example, passengers climbing the column).
[0317] Comprehensive visual abnormality score , Sound Abnormality Score and equipment operation data anomaly score Perform multimodal fusion and calculate the comprehensive anomaly score Among them, visual abnormality score Determined by the results of the equipment component movement frequency analysis and abnormal status, when the movement frequency or amplitude is detected to exceed the threshold, Set to 1, otherwise 0. Sound Abnormality Score It is obtained by analyzing the collected sound data. For example, when a passenger screams or abnormal noise of the equipment is detected, Set to 1, otherwise 0. Equipment operation data abnormality score It is obtained by analyzing the collected escalator operation data. For example, when abnormal escalator speed, current overload or excessive vibration is detected, Set to 1, otherwise 0. The comprehensive anomaly score calculation formula is:
[0318]
[0319] The weight coefficient , , are set to 0.5, 0.3, and 0.2 respectively, and The comprehensive threshold Tc is set to 0.7. > Tc, the abnormal event is confirmed and graded according to the type and severity of the abnormal event.
[0320] The anomaly detection module is based on the motion frequency analysis and multimodal fusion mechanism at the equipment component level, and can accurately and reliably detect abnormal events in the escalator entrance area, such as passengers falling, walking in the opposite direction, climbing escalators, equipment failures, etc. Multimodal fusion effectively reduces the false alarm rate and improves the detection accuracy. Refined equipment component-level analysis can more accurately locate the specific location and equipment components where the anomaly occurs, providing a more accurate decision-making basis for subsequent emergency response. In this embodiment, the multimodal anomaly detection mechanism reduces the false alarm rate to less than 1% and increases the accuracy to more than 95%.
[0321] The emergency response module receives abnormal event alarm information from the abnormal detection module, formulates emergency response strategies according to the severity of the abnormal events, sends control instructions to rail transit equipment through the industrial protocol interface, and pushes abnormal event information to the control center, patrol personnel and passenger terminals.
[0322] Based on the comprehensive anomaly score output by the anomaly detection module And abnormal event type, the abnormal events are divided into three levels of response:
[0323] Level 1 (mild abnormality): Comprehensive abnormality score Between 0.7 and 0.8, for example, if a passenger is detected to have briefly stopped at the entrance of the escalator but does not pose an obvious safety risk, the device self-check program is triggered, and the system background records the abnormal log, which does not affect the normal operation of the escalator, and only pushes the first-level alarm information to the control center.
[0324] Level 2 (obviously abnormal): Comprehensive abnormality score Between 0.8 and 0.9, for example, if passengers are detected running or playing on the escalator, or the escalator speed fluctuates slightly, the equipment will be decelerated or the flow will be limited, such as reducing the escalator speed to 50% of the normal speed, and providing safety reminders through platform broadcasts and passenger information display screens, while pushing secondary alarm information to the control center and the mobile terminals of nearby inspectors, prompting inspectors to go to the site for inspection.
[0325] Level 3 (severe abnormality): comprehensive abnormality score > 0.9, for example, if a passenger falls, walks in the opposite direction, climbs the escalator, or a serious equipment failure such as an escalator emergency stop, walks in the opposite direction, or a component falls off is detected, an emergency stop command is immediately sent through the industrial protocol interface of the escalator control system to stop the escalator and initiate the preset emergency plan, such as linking the platform emergency stop button, starting the platform emergency broadcast, and pushing the third-level alarm information to the control center, patrol personnel, and passenger terminals, and automatically calling the platform duty personnel and medical emergency personnel.
[0326] According to the emergency response strategy, the corresponding control instructions, such as self-check instructions, deceleration instructions, current limiting instructions, stop instructions, etc., are sent to the escalator control system through the industrial Ethernet interface of the escalator control system using the Modbus TCP / IP protocol to achieve real-time intervention and control of the escalator.
[0327] The detailed information of abnormal events (including abnormal type, occurrence time, location, severity, related equipment component ID, on-site video images, etc.) is pushed to the urban rail transit control center, the mobile terminals of platform inspection personnel (such as inspection APP), platform passenger information display screens, passenger mobile APP and other passenger terminals through the network communication module. The information push protocol adopts MQTT or WebSocket protocol to ensure the real-time and reliability of information.
[0328] The emergency response module can formulate and implement corresponding emergency response strategies according to the severity of abnormal events, realize real-time control and intervention of rail transit equipment, and push abnormal information to relevant personnel and passengers in time, forming a fast and efficient emergency linkage mechanism, maximizing passenger safety, reducing accident risks, and improving the intelligence and safety management level of urban rail transit.
[0329] This embodiment describes in detail the application of the urban rail transit emergency response system based on the visual large model in the platform escalator entrance area. The system realizes real-time, accurate and intelligent monitoring, analysis and processing of abnormal events in the escalator entrance area through the collaborative work of the data acquisition module, visual analysis module, spatial mapping module, anomaly detection module and emergency response module. The system effectively utilizes multimodal information such as vision, hearing and equipment operation, adopts edge computing architecture, reduces the computational complexity and data transmission pressure, improves the real-time and reliability of system operation, and has significant technical advantages and application value. The system described in this embodiment is not only suitable for the platform escalator entrance area, but can also be extended to other key areas of urban rail transit such as platform waiting area, gate area, and carriage interior, providing strong support for building a safer and smarter urban rail transit system.
[0330] The above is a preferred embodiment of the present invention. It should be pointed out that a person skilled in the art can make several improvements and modifications without departing from the principle of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A method for emergency handling of urban rail transit based on a visual large model, characterized in that: include: Collect visual data, sound data and equipment operation data in key areas of urban rail transit platforms and carriages, perform pre-processing, and synchronize timestamps; Performing dynamic frame difference processing on continuous video frames of visual data to generate a two-dimensional motion heat map, and performing spatiotemporal aggregation on the two-dimensional motion heat map; Build a 3D digital model based on the engineering drawings of the platform and carriage of urban rail transit, perform internal and external parameter calibration of the camera, and establish a mapping relationship between the 2D motion heat map and the 3D digital model; Abnormal behavior detection is performed on the mapping results of the two-dimensional motion heat map and the three-dimensional digital model after spatiotemporal aggregation, and multimodal fusion analysis is performed in combination with sound data and equipment operation data to determine the severity of abnormal events; Emergency response strategies are formulated based on the severity of abnormal events, control instructions are sent to rail transit equipment through industrial protocol interfaces, and abnormal event information is pushed to the control center, patrol personnel and passenger terminals.
2. According to claim 1, a method for emergency handling of urban rail transit based on a visual large model is characterized in that: The collection of visual data, sound data and equipment operation data includes: Wide-angle cameras are deployed at the escalator entrances, security check areas, waiting areas, and both ends of the carriages, with the acquisition frequency set to 10 frames per second; Use microphone arrays to collect environmental sound data and extract sound features, including frequency, energy, and spectrum changes; The equipment operation data is collected in real time through the industrial protocol interface of rail transit equipment. The equipment operation data includes speed, current, temperature and vibration.
3. The urban rail transit emergency handling method based on visual large model according to claim 2 is characterized in that: The preprocessing of visual data, sound data and device operation data and achieving timestamp synchronization includes: Gaussian blurring of visual data; Perform noise reduction on sound data; Normalize the equipment operation data; The network time protocol or precision time protocol is used to synchronize the timestamps of the pre-processed visual data, sound data and equipment operation data to ensure the accuracy of data fusion.
4. The urban rail transit emergency handling method based on visual large model according to claim 3 is characterized in that: The performing of dynamic frame difference processing on continuous video frames of visual data to generate a two-dimensional motion heat map, and performing spatiotemporal aggregation on the two-dimensional motion heat map comprises: The pixel differences between adjacent frames are calculated for continuous video frames, and the difference threshold is dynamically adjusted through the environment adaptive algorithm to obtain the difference image; Binarize the difference image to generate a binary motion heat map; Aggregate five consecutive frames of 2D motion heatmaps in the time dimension; The aggregated binary motion heat map is spatiotemporally aggregated, and the interference points are removed by applying the morphological closing operation. The independent motion regions are identified and labeled by the connected region labeling algorithm, and a unique identifier is assigned.
5. The urban rail transit emergency handling method based on visual large model according to claim 4 is characterized in that: The three-dimensional digital model is constructed based on the engineering drawings of the platform and carriage of the urban rail transit, including: Build accurate 3D digital models based on CAD engineering drawings of platforms and carriages, including the spatial locations and geometric parameters of escalators, gates and seat fixtures; Assign a unique ID and functional label to each equipment component of a fixed facility and build an equipment component database.
6. The urban rail transit emergency handling method based on visual large model according to claim 5 is characterized in that: The step of performing camera internal and external parameter calibration and establishing a mapping relationship between a two-dimensional motion heat map and a three-dimensional digital model includes: Use the calibration board to perform camera internal and external parameter calibration to obtain the camera's internal and external parameter matrix; Based on the internal and external parameter calibration results, an affine transformation is applied to map the motion area in the two-dimensional motion heat map to the spatial coordinate system of the three-dimensional digital model; Through the spatial occupancy detection algorithm, the motion area that overlaps with the equipment parts with unique IDs in the 3D model is identified.
7. The urban rail transit emergency handling method based on visual large model according to claim 6 is characterized in that: The abnormal behavior detection on the mapping result of the two-dimensional motion heat map and the three-dimensional digital model after spatiotemporal aggregation includes: For the motion area that spatially overlaps with the equipment component with a unique ID, the motion frequency of the equipment component within a preset time window is calculated, and the spectral characteristics of the motion are analyzed by Fourier transform to extract the motion frequency and its amplitude; A table of abnormal status thresholds is preset based on the equipment type, and an alarm of the corresponding component is triggered when the movement frequency or amplitude is detected to exceed the threshold.
8. The urban rail transit emergency handling method based on visual large model according to claim 7 is characterized in that: The multimodal fusion analysis combining the sound data and the equipment operation data to determine the severity of the abnormal event includes: Based on the weighted fusion of visual anomaly score, sound anomaly score and equipment operation data anomaly score, a comprehensive anomaly score is obtained. The calculation formula is: ; in , , The visual data , sound data and equipment operation data The weight coefficient of ; When the combined anomaly score exceeds the threshold, the anomaly event is confirmed.
9. The urban rail transit emergency handling method based on visual large model according to claim 8 is characterized in that: The emergency response strategy formulated according to the severity of abnormal events includes: Abnormal events are divided into three levels of response: Level 1: Minor abnormality, triggering the device self-check procedure, without affecting the normal operation of the device; Level 2: Obvious abnormality, triggering equipment deceleration or current limiting measures; Level 3: Severe abnormality, immediately stop equipment operation and activate the emergency plan.
10. An urban rail transit emergency handling system based on a visual large model, characterized in that: include: The data acquisition module is used to collect visual data, sound data and equipment operation data in key areas of urban rail transit platforms and carriages, and perform pre-processing and achieve time stamp synchronization; A visual analysis module, for performing dynamic frame difference processing on continuous video frames of visual data to generate a two-dimensional motion heat map, and performing spatiotemporal aggregation on the two-dimensional motion heat map; The spatial mapping module is used to build a 3D digital model based on the engineering drawings of the platform and carriage of urban rail transit, perform internal and external parameter calibration of the camera, and establish a mapping relationship between the 2D motion heat map and the 3D digital model; The anomaly detection module is used to detect abnormal behavior on the mapping results of the two-dimensional motion heat map and the three-dimensional digital model after spatiotemporal aggregation, and to perform multimodal fusion analysis in combination with sound data and equipment operation data to determine the severity of abnormal events; The emergency response module is used to formulate emergency response strategies according to the severity of abnormal events, send control instructions to rail transit equipment through the industrial protocol interface, and push abnormal event information to the control center, patrol personnel and passenger terminals.
Citation Information
Patent Citations
Urban rail transit gate passing control method based on binocular camera behavior detection
CN109657581A
Subway carriage passenger clearance judgment method based on image analysis and electronic equipment
CN116310307A
Hump shunting band-type brake abnormity monitoring system based on multi-source data fusion algorithm
CN116776202A
Intelligent operation and maintenance emergency processing system
CN117215940A
Security abnormal behavior identification method based on multi-dimensional feature fusion
CN117372917A
Cited By
Human body key point identification method and system based on posture video data fusion
CN120451880A
Multi-mode perception sound leakage detection and active compensation method
CN120881447A
Track long-distance target detection system based on multispectral feature fusion
CN121074362A
Injection product defect real-time detection and early warning system based on computer vision
CN121074442A
Target management method, management system, electronic device and computer program product
CN121505856A