An urban rail transit emergency handling method and system based on a large vision model

Through visual large models combined with multi-source data processing technology, the response speed and multi-modal information fusion problems of urban rail transit emergency processing systems are solved, and efficient and intelligent abnormal event detection and emergency response are achieved.

CN120047907BActive Publication Date: 2025-07-29PCI TECH GRP CO LTD +2

Patent Information

Application Number
CN202510510717.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-29
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

The existing urban rail transit emergency response system has shortcomings in response speed, false alarm rate, multimodal information fusion and computing resource requirements, making it difficult to achieve efficient and intelligent emergency handling.

Method used

The visual big model is adopted, combining visual data, sound data and device operation data, and through dynamic frame difference, spatiotemporal aggregation, geometric mapping and multimodal fusion, rapid detection and accurate judgment of abnormal events are achieved, and a hierarchical emergency response strategy is formulated.

Benefits of technology

It improves the accuracy and robustness of abnormal detection, reduces the false alarm rate, and realizes efficient and intelligent emergency response to urban rail transit systems, meeting the needs of real-time and low computing complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047907B_ABST
    Figure CN120047907B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of big data technology, and discloses an urban rail transit emergency handling method and system based on a visual large model, including: collecting visual data, sound data, and equipment operation data at the platforms and carriages of urban rail transit, performing preprocessing and timestamp synchronization; performing dynamic frame difference processing on consecutive video frames of the visual data to generate a two-dimensional motion heat map, and performing spatio-temporal aggregation; constructing a three-dimensional digital model based on the engineering drawings of the platforms and carriages, and establishing a mapping relationship between the two-dimensional motion heat map and the three-dimensional digital model; performing abnormal behavior detection on the mapping results, and performing multi-modal fusion analysis in combination with the sound data and equipment operation data to determine the severity of the abnormal event; formulating an emergency response strategy according to the severity level of the abnormal event, sending control instructions to the rail transit equipment through an industrial protocol interface, and pushing the abnormal event information to the control center, inspection personnel, and passenger terminals.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of big data technology, and in particular to an urban rail transit emergency handling method and system based on a visual large model. Background Art

[0002] As an important public transportation system in modern cities, the safe and efficient operation of urban rail transit is related to the daily travel of urban residents and the stable development of social economy. However, the operating environment of the urban rail transit system is complex, with a large passenger flow, and emergencies occur from time to time, such as passengers accidentally falling, equipment failures, and people intruding into the track. These events may lead to service interruptions, passenger casualties, and even trigger group safety accidents. Therefore, it is crucial to establish a set of efficient and intelligent urban rail transit emergency handling systems.

[0003] In the prior art, the emergency handling of urban rail transit mainly relies on the following methods:

[0004] (1) Manual monitoring and inspection: Traditional urban rail transit systems generally adopt a manual monitoring mode, that is, monitoring personnel are set up in the control center to monitor key areas such as platforms and carriages in real time by viewing the closed-circuit television monitoring system images. In addition, inspection personnel are also equipped to conduct regular or irregular inspections on platforms and in carriages to detect abnormal situations.

[0005] However, the method of manual monitoring and inspection has obvious limitations:

[0006] Slow response speed: Manual monitoring relies on the visual observation and subjective judgment of monitoring personnel, and there is a lag in the discovery of abnormal events. Especially in the case of a large number of monitoring images and a large amount of information, it is easy to be negligent and omitted, resulting in a slow emergency response speed.

[0007] Prone to fatigue and misjudgment: Long-term monitoring work is likely to cause fatigue of monitoring personnel, a decrease in attention, a reduction in sensitivity to abnormal events, and is prone to misjudgment or omission of judgment, affecting the accuracy of emergency handling.

[0008] Limited coverage: The labor cost of manual inspection is high, the inspection frequency and coverage are limited, and it is difficult to achieve real-time and comprehensive monitoring of all key areas.

[0009] (2) Early warning system based on simple sensors: In order to improve the automation level, some rail transit systems have begun to introduce early warning systems based on simple sensors, such as infrared sensors, sound sensors, pressure sensors, etc. These sensors can detect changes in specific physical quantities, such as abnormal temperature, abnormal sound, abnormal pressure, etc., so as to achieve preliminary abnormal early warning.

[0010] However, the early warning system based on simple sensors also has deficiencies:

[0011] High false alarm rate: Simple sensors are vulnerable to environmental factors. For example, infrared sensors may be affected by temperature changes, and sound sensors may be interfered by environmental noise, resulting in a relatively high false alarm rate.

[0012] Single information dimension: A single type of sensor can only perceive a limited number of information dimensions and it is difficult to comprehensively and accurately judge the nature and severity of events. For example, a pressure sensor can only detect pressure changes but cannot distinguish whether it is a passenger fall or an item drop.

[0013] Lack of spatial information: Simple sensors usually can only provide point-like or regional perception information, lacking an accurate description of the location and spatial relationship where an event occurs, which is not conducive to precise positioning and subsequent handling.

[0014] (3) Early visual surveillance systems: With the development of computer vision technology, some rail transit systems have begun to attempt to introduce early visual surveillance systems, such as motion detection systems based on background modeling, optical flow method, etc. These systems can use the video data collected by cameras to detect moving targets in the images and achieve preliminary automated surveillance.

[0015] However, early visual surveillance systems still face challenges in practical applications:

[0016] Poor algorithm robustness: Early visual algorithms are sensitive to factors such as environmental light changes, shadow occlusion, and complex backgrounds, and are prone to false detections and missed detections. Their robustness is poor and they perform unstably in complex rail transit environments.

[0017] High computational complexity: Some complex visual algorithms, such as the optical flow method, have a relatively high computational complexity and are difficult to meet the real-time requirements, which limits the response speed of the system.

[0018] Lack of semantic understanding: Early visual systems usually can only detect moving targets and lack an understanding of scene semantics. They cannot distinguish normal movements from abnormal behaviors and are difficult to perform in-depth abnormal analysis and judgment.

[0019] (4) Deep learning-based intelligent video analysis technology: In recent years, deep learning technology has made breakthrough progress in the field of computer vision. Deep learning-based object detection, behavior recognition and other technologies have gradually been applied to the field of intelligent video surveillance. Some studies have attempted to apply deep learning technology to the detection of abnormal events in rail transit, such as using convolutional neural networks for object detection and classification to identify abnormal behaviors such as passenger falls and crowded stampedes. Although deep learning-based intelligent video analysis technology has improved the accuracy and intelligence level of abnormal detection to a certain extent, there are still some limitations, especially in the urban rail transit emergency handling scenarios:

[0020] High computational resource requirements: Deep learning models usually have a large number of parameters and high computational complexity, requiring powerful computational resources for support. It is difficult to run in real time on edge devices, which limits the deployment and application scope of the system.

[0021] Strong dependence on training data: Deep learning models require a large amount of labeled data for training. However, it is difficult to obtain labeled data for rail transit abnormal events, the labeling cost is high, and the data quality is difficult to guarantee, which affects the generalization ability and robustness of the model.

[0022] Poor model interpretability: Deep learning models are usually "black box" models, making it difficult to explain their decision-making processes. This is not conducive to system fault diagnosis and optimization improvement, and it is also difficult to meet the high requirements of rail transit systems for safety and reliability.

[0023] Insufficient multi-modal information fusion: Existing research on rail transit anomaly detection based on deep learning mainly focuses on visual data analysis, with insufficient utilization of other modal information such as sound data and equipment operation data. The information dimension is single, making it difficult to achieve comprehensive and multi-dimensional anomaly event judgment.

[0024] Therefore, there is a need for an urban rail transit emergency handling method that can improve the emergency response speed, reduce the false alarm rate, increase the degree of automation, and achieve multi-source information fusion to overcome the deficiencies of existing technologies and ensure the safe and efficient operation of urban rail transit systems. Summary of the Invention

[0025] The present invention provides an urban rail transit emergency handling method and system based on a vision large model, aiming to improve the intelligence and automation level of rail transit systems in responding to emergencies. By collecting, processing, and analyzing multi-source data in key areas such as urban rail transit platforms and carriages, rapid response and effective disposal of abnormal events are ultimately achieved.

[0026] The present invention provides an urban rail transit emergency handling method based on a vision large model, including:

[0027] Collect visual data, sound data, and equipment operation data in key areas of urban rail transit platforms and carriages, and perform preprocessing while achieving timestamp synchronization;

[0028] Perform dynamic frame difference processing on consecutive video frames of visual data to generate a two-dimensional motion heat map, and perform spatio-temporal aggregation on the two-dimensional motion heat map;

[0029] Construct a three-dimensional digital model based on the engineering drawings of urban rail transit platforms and carriages, perform camera internal and external parameter calibration, and establish a mapping relationship between the two-dimensional motion heat map and the three-dimensional digital model;

[0030] Abnormal behavior detection is performed on the mapping results of the spatiotemporally aggregated 2D motion heat map and the 3D digital model. Multimodal fusion analysis is then performed in combination with sound data and equipment operation data to determine the severity of the abnormal event.

[0031] Emergency response strategies are formulated based on the severity of abnormal events, control instructions are sent to rail transit equipment through industrial protocol interfaces, and abnormal event information is pushed to the control center, inspection personnel, and passenger terminals.

[0032] In the above solution, multimodal information, including visual data, audio data, and equipment operation data, is first collected synchronously in key areas of urban rail transit platforms and carriages. Visual data is primarily collected by cameras deployed in key areas, such as escalator entrances, security check areas, waiting areas, and the interior of carriages. Audio data is collected by sound sensors to assist in the identification of abnormal events. Equipment operation data, such as the operating status of rail transit equipment such as gates, escalators, and platform doors, can reflect the operating status of the equipment itself. To ensure the accuracy of subsequent data processing and analysis, the collected multimodal data is preprocessed, such as image denoising, audio signal filtering, and data format unification. A unified timestamp is also added to all data to ensure synchronous alignment of multi-source data on the timeline.

[0033] Subsequently, dynamic frame differencing is performed on the collected visual data across consecutive video frames. This technology aims to detect areas of motion within the video. By calculating pixel differences between adjacent frames, it can effectively extract moving targets within the video. To adapt to varying ambient lighting and scene changes, dynamic frame differencing is employed. The difference threshold or algorithm parameters can be adaptively adjusted based on the environment, thereby improving the robustness of motion detection. After frame differencing, a two-dimensional motion heatmap is generated, visualizing areas of active motion within the video. To further eliminate noise interference and extract more stable and prominent moving targets, the generated two-dimensional motion heatmap is subjected to spatiotemporal aggregation. Spatiotemporal aggregation involves integrating the motion heatmap across time and space. For example, multiple consecutive frames of heatmaps can be superimposed or subjected to morphological filtering to filter out transient, isolated noise points while retaining continuous, continuous motion areas, making the moving targets clearer and more complete.

[0034] To associate two-dimensional image information with the actual three-dimensional space environment, geometric mapping is introduced. Based on the engineering drawings of urban rail transit platforms and carriages, an accurate three-dimensional digital model is constructed. This model is a digital restoration of the real scene, containing information such as the structural layout of the platform and carriage, the spatial positions and geometric parameters of equipment and facilities, etc. To establish the correspondence between the two-dimensional images captured by the camera and the three-dimensional digital model, camera internal and external parameter calibration needs to be performed. Camera calibration can determine the spatial position, attitude, and internal imaging parameters of the camera, thereby obtaining the transformation matrix from the image coordinate system to the world coordinate system. On this basis, the mapping relationship between the two-dimensional motion heat map and the three-dimensional digital model can be established, accurately projecting the motion area on the two-dimensional heat map into the three-dimensional model to determine the specific position and area of the moving object in the three-dimensional space.

[0035] After obtaining the position information of the moving object in the three-dimensional space, abnormal behavior detection will be carried out. By analyzing the mapping results of the two-dimensional motion heat map after spatio-temporal aggregation and the three-dimensional digital model, it can be judged whether abnormal motion behavior has occurred in a specific spatial area. For example, it can be analyzed whether the motion area interacts with equipment and facilities, and whether the speed and trajectory of the motion conform to the normal mode, etc. To improve the accuracy and reliability of abnormal detection, the method further combines the simultaneously collected sound data and equipment operation data for multi-modal fusion analysis. By comprehensively considering information from multiple aspects such as vision, hearing, and equipment operation status, the nature and severity of the event can be judged more comprehensively and accurately. For example, when vision detects a person falling in the platform area, at the same time the sound sensor also collects abnormal sounds, and the operation data of the nearby platform door indicates that it is in an abnormal opening state, it can be comprehensively judged that a relatively serious emergency has occurred. Through multi-modal fusion analysis, the false alarm rate can be effectively reduced and the accuracy of abnormal event recognition can be improved.

[0036] Based on the severity of the abnormal event determined by multimodal fusion analysis, a hierarchical emergency response strategy is formulated. According to the harm degree and urgency of the event, abnormal events can be classified into different levels, such as minor anomalies, obvious anomalies, severe anomalies, etc., and corresponding emergency response measures are preset for each level. For example, for minor anomalies, only device self-check or warning prompts may be triggered; for obvious anomalies, measures such as starting device deceleration and current limiting may be required; for severe anomalies, the device operation may need to be stopped immediately, and a higher-level emergency plan may be activated, such as evacuating passengers and calling for rescue. To achieve automated control of rail transit equipment, control instructions are sent to rail transit equipment through industrial protocol interfaces, such as controlling escalators to stop, turnstiles to close, and platform doors to open. At the same time, to notify relevant personnel in a timely manner and carry out collaborative disposal, the method also pushes abnormal event information to the control center, inspection personnel, and passenger terminals, so that relevant personnel can understand the situation in a timely manner and take corresponding actions.

[0037] Preferably, the acquisition of visual data, sound data, and equipment operation data includes:

[0038] Deploy wide-angle cameras at the escalator entrances, security check areas, waiting areas on the platform, and both ends of the carriages, and set the acquisition frequency to 10 frames per second;

[0039] Use a microphone array to collect ambient sound data and extract sound features, where the sound features include frequency, energy, and spectral changes;

[0040] Real-time collect equipment operation data through the industrial protocol interface of rail transit equipment, and the equipment operation data includes speed, current, temperature, and vibration.

[0041] In the above solution, the data acquisition method is specifically described, including: deploying wide-angle cameras at key areas on the platform (escalator entrances, security check areas, waiting areas) and both ends of the carriages, setting the acquisition frequency to 10 frames per second; using a microphone array to collect ambient sound and extract sound features such as frequency, energy, and spectral changes; real-time collect equipment operation data through the industrial protocol interface of rail transit equipment, such as parameters such as speed, current, temperature, and vibration.

[0042] Through specific data acquisition configurations, the following technical effects are produced:

[0043] Full coverage of key areas and carriages to ensure data integrity: Deploying wide-angle cameras at key areas on the platform and both ends of the carriages can maximize the coverage of the passenger activity area and equipment operation area, ensure the comprehensiveness and integrity of visual data acquisition, avoid monitoring blind spots, and provide a sufficient data basis for subsequent abnormal event detection.

[0044] Reasonable acquisition frequency, balancing real-time performance and data processing pressure: The visual data acquisition frequency of 10 frames per second can not only meet the real-time requirements of motion detection but also effectively reduce the burden of data transmission and processing, achieving a balance between system performance and data quality.

[0045] Microphone array and sound feature extraction to improve the utilization rate of sound information: Using a microphone array can enhance the directivity and sensitivity of sound collection, effectively capturing abnormal sounds in the environment. Extracting sound features such as frequency, energy, and spectral changes can more effectively characterize the characteristics of sounds, providing structured and quantifiable data for subsequent abnormal sound recognition and analysis, and improving the role of sound information in abnormal event judgment.

[0046] Industrial protocol interface to collect device operation data, ensuring data real-time performance and accuracy: Directly collecting operation data from rail transit devices through the industrial protocol interface can ensure data real-time performance and accuracy, avoiding delays and errors in the traditional sensor deployment and data conversion processes, and providing a reliable data source for device status monitoring and abnormal warning.

[0047] Preferably, the preprocessing of visual data, sound data, and device operation data, and the simultaneous realization of timestamp synchronization include:

[0048] Perform Gaussian blur processing on visual data;

[0049] Perform noise reduction processing on sound data;

[0050] Perform normalization processing on device operation data;

[0051] Use the Network Time Protocol or Precision Time Protocol to synchronize timestamps for the preprocessed visual data, sound data, and device operation data to ensure the accuracy of data fusion.

[0052] In the above solution, performing Gaussian blur processing on visual data can effectively smooth the image, remove high-frequency noise, reduce the influence of factors such as environmental light changes and sensor noise on motion detection, improve the quality of visual data, and enhance the accuracy and stability of subsequent moving object detection; performing noise reduction processing on sound data can effectively filter out environmental background noise, highlight the characteristics of abnormal sounds, improve the signal-to-noise ratio of sound data, and make subsequent sound feature extraction and abnormal sound recognition more accurate and reliable; performing normalization processing on device operation data can unify data with different dimensions into the same numerical range, eliminate the influence brought by dimensional differences, facilitate subsequent multi-modal data fusion processing, and improve the rationality and effectiveness of data fusion; using the NTP or PTP protocol to synchronize timestamps for multi-modal data can ensure that data from different sources are aligned on the time axis, ensure the accuracy of data fusion, avoid analysis errors caused by time asynchronization, and provide a reliable time reference for subsequent multi-modal fusion anomaly detection.

[0053] Preferably, performing dynamic frame difference processing on consecutive video frames of the visual data to generate a two-dimensional motion heat map includes:

[0054] Calculating the pixel differences between adjacent frames of consecutive video frames, and dynamically adjusting the difference threshold through an environment adaptive algorithm to obtain a difference image; performing binary processing on the difference image to generate a binary motion heat map.

[0055] In the above solution, calculating the pixel differences between adjacent frames can effectively highlight the motion regions in the video frames and suppress static background information, thereby distinguishing moving objects from the background; dynamically adjusting the difference threshold through an environment adaptive algorithm can automatically adjust the threshold according to changes in environmental light, noise level, etc., enabling the difference process to adapt to different environmental conditions, effectively suppressing the influence of factors such as light changes and shadow interference, improving the environmental adaptability and robustness of moving object detection, and reducing the false detection rate and missed detection rate; performing binary processing on the difference image to generate a binary motion heat map can convert complex pixel difference information into a simple binary image, highlight the contours of the motion regions, simplify data representation, facilitate subsequent spatio-temporal aggregation and object recognition processing, and improve processing efficiency.

[0056] Preferably, performing spatio-temporal aggregation on the two-dimensional motion heat map includes:

[0057] Performing aggregation in the time dimension on the two-dimensional motion heat maps of five consecutive frames;

[0058] Performing spatio-temporal aggregation on the aggregated binary motion heat map, applying morphological closing operation to remove interference points; identifying and labeling independent motion regions through a connected component labeling algorithm and assigning unique identifiers.

[0059] In the above solution, time aggregation is performed on five consecutive frames of motion heatmaps, which can utilize the continuity characteristics of motion, filter out instantaneous noise points with short durations, highlight the target areas of stable motion, improve the quality of the motion heatmaps, and reduce noise interference; applying morphological closing operation can effectively fill the holes inside the moving targets, connect the broken edges, make the contours of the moving targets more complete and continuous, and improve the integrity and accuracy of target segmentation; through the connected component labeling algorithm, the pixel regions that are connected to each other in the motion heatmap can be identified as independent moving targets and assigned unique identifiers, realizing the distinction and tracking of different moving targets, providing independent target units for subsequent abnormal behavior analysis, and facilitating behavior analysis and abnormal judgment for different targets.

[0060] Preferably, the construction of the three-dimensional digital model based on the engineering drawings of the platform and carriage of urban rail transit includes:

[0061] Constructing an accurate three-dimensional digital model based on the CAD engineering drawings of the platform and carriage, where the three-dimensional digital model includes the spatial positions and geometric parameters of fixed facilities such as escalators, turnstiles, and seats; assigning unique IDs and function labels to the equipment components of each fixed facility to construct an equipment component database.

[0062] In the above solution, constructing a three-dimensional digital model based on the CAD engineering drawings of the platform and carriage can ensure a high degree of consistency between the model and the real physical space, provide accurate spatial geometric information, and provide an accurate spatial reference for the subsequent mapping from two-dimensional images to three-dimensional space; the three-dimensional model includes the spatial positions and geometric parameters of fixed facilities such as escalators, turnstiles, and seats, which can effectively describe the environmental scenes of the rail transit platform and carriage, and provide environmental context information for subsequent analysis of the spatial positions of moving targets and behavior understanding; assigning unique IDs and function labels to each equipment component and constructing an equipment component database can associate the components in the three-dimensional model with the actual equipment, support subsequent refined abnormal behavior analysis and emergency response based on components, such as targeted handling of abnormal movements of escalators, turnstile failures, etc.

[0063] Preferably, the execution of the internal and external parameter calibration of the camera and the establishment of the mapping relationship between the two-dimensional motion heatmap and the three-dimensional digital model include:

[0064] Using a calibration board to perform the internal and external parameter calibration of the camera to obtain the internal parameter matrix and the external parameter matrix of the camera;

[0065] Based on the results of the internal and external parameter calibration, applying an affine transformation to map the motion region in the two-dimensional motion heatmap to the spatial coordinate system of the three-dimensional digital model;

[0066] Through a spatial occupancy detection algorithm, identifying the motion regions that overlap with the equipment components with unique IDs in the three-dimensional model in terms of spatial position.

[0067] In the above solution, the internal and external parameters of the camera are calibrated through a calibration board, and the internal parameter matrix and external parameter matrix of the camera can be obtained, an imaging model of the camera is established, the conversion relationship between the image pixel coordinates and the three-dimensional space coordinates is described, and necessary parameters are provided for subsequent mapping transformation; an affine transformation is applied to map the moving area in the two-dimensional motion heat map to the space coordinate system of the three-dimensional CAD model, realizing the conversion of two-dimensional image information to the three-dimensional physical space, enabling the analysis of the position and behavior of moving targets in the three-dimensional space; through a space occupancy detection algorithm, the moving areas that overlap with the space positions of device components with unique IDs in the three-dimensional model are identified, and it can be determined whether the moving target interacts with specific device components, such as whether a pedestrian approaches an escalator, a turnstile, etc., providing a basis for subsequent component-based abnormal behavior analysis.

[0068] Preferably, the abnormal behavior detection of the mapping result of the two-dimensional motion heat map after spatio-temporal aggregation and the three-dimensional digital model includes:

[0069] For the moving area that spatially overlaps with the device component with a unique ID, calculate the movement frequency of the device component within a preset time window, and analyze the spectral characteristics of the movement through Fourier transform, extract the movement frequency and its amplitude; based on a preset abnormal state threshold table for device types, when the detected movement frequency or amplitude exceeds the threshold, trigger an alarm for the corresponding component.

[0070] In the above solution, calculating the movement frequency of the device component within the time window and analyzing the movement spectral characteristics through Fourier transform can extract characteristics such as the periodicity and intensity of the movement, more comprehensively describe the movement behavior, and provide a richer basis for abnormal judgment; based on a preset abnormal state threshold table for device types, setting different movement frequency and amplitude thresholds for different types of device components can achieve refined abnormal determination, improve the accuracy and sensitivity of abnormal detection, and avoid misjudgment caused by using a unified threshold for different types of devices; by comparing the measured movement frequency and amplitude with the preset threshold, it can quickly and effectively determine whether the device component is in an abnormal state, such as the movement frequency of the escalator steps being too high, the movement amplitude of the turnstile being abnormal, etc., and trigger an alarm in a timely manner, providing timely warning information for subsequent emergency response.

[0071] Preferably, the multi-modal fusion analysis by combining sound data and device operation data to determine the severity of abnormal events includes:

[0072] Perform weighted fusion based on visual abnormal score, sound abnormal score, and device operation data abnormal score. The comprehensive abnormal score calculation formula is:

[0073]

[0074] Among them 、 、 are the weight coefficients of visual data ,audio data and device operation data respectively, and ;

[0075] When the comprehensive anomaly score exceeds the threshold, an anomaly event is confirmed.

[0076] In the above solution, the anomaly scores of visual, audio, and device operation data are weighted and fused, which can comprehensively consider information of different modalities, utilize the complementary advantages of each modality, and improve the reliability and accuracy of anomaly event judgment. For example, if visual detection finds abnormal movement, audio detection detects abnormal sound, and device operation data also shows anomalies, the comprehensive score will be higher and the confidence level of the anomaly event will be higher; the weight coefficients 、 、 can be flexibly adjusted according to the actual application scenario and the reliability of different modality data. For example, in a scenario where visual information is more reliable, the weight of the visual score can be increased , and vice versa, so that the multi-modal fusion is more adaptable and flexible; by setting the comprehensive anomaly score threshold, an anomaly event is only confirmed when the comprehensive score exceeds the threshold, which can effectively reduce the false alarm rate, improve the accuracy of anomaly event judgment, and avoid incorrect responses caused by misjudgment of a single modality.

[0077] Preferably, the formulation of the emergency response strategy according to the severity level of the anomaly event includes:

[0078] The anomaly event is divided into three levels of response:

[0079] Level 1: Minor anomaly, triggering the device self-check program, without affecting the normal operation of the device;

[0080] Level 2: Obvious anomaly, triggering device deceleration or current limiting measures;

[0081] Level 3: Severe anomaly, immediately stopping the device operation and starting the emergency plan.

[0082] In the above solution, the abnormal events are divided into three levels of response: minor, obvious, and severe. Different intensities of emergency measures can be taken according to the severity of the abnormality, enabling refined emergency handling, avoiding the "one-size-fits-all" response method, and minimizing the interference to the operation of rail transit to the greatest extent while ensuring safety; corresponding response strategies are formulated for different levels of abnormal events. For example, self-checking is performed for minor abnormalities, speed reduction and flow restriction are carried out for obvious abnormalities, and operation is stopped for severe abnormalities, ensuring the rationality and effectiveness of the response strategy, being able to take the most appropriate measures according to the actual situation, effectively controlling risks, and maintaining the normal operation of rail transit as much as possible; the hierarchical framework of the three-level response has good scalability, and the response levels can be further refined according to actual needs, and more refined response strategies can be formulated for different levels of abnormal events, facilitating subsequent strategy expansion and optimization, and enhancing the flexibility and adaptability of the emergency handling system.

[0083] Preferably, an urban rail transit emergency handling system based on a large vision model includes:

[0084] A data acquisition module, which is used to collect visual data, sound data, and equipment operation data in key areas of the platforms and carriages of urban rail transit, perform preprocessing, and achieve timestamp synchronization at the same time;

[0085] A visual analysis module, which is used to perform dynamic frame difference processing on consecutive video frames of visual data to generate a two-dimensional motion heat map, and perform spatio-temporal aggregation on the two-dimensional motion heat map;

[0086] A space mapping module, which is used to construct a three-dimensional digital model based on the engineering drawings of the platforms and carriages of urban rail transit, perform calibration of the internal and external parameters of the camera, and establish a mapping relationship between the two-dimensional motion heat map and the three-dimensional digital model;

[0087] An abnormal detection module, which is used to detect abnormal behaviors in the mapping results of the spatio-temporally aggregated two-dimensional motion heat map and the three-dimensional digital model, and perform multi-modal fusion analysis in combination with sound data and equipment operation data to determine the severity of abnormal events;

[0088] An emergency response module, which is used to formulate emergency response strategies according to the severity level of abnormal events, send control instructions to rail transit equipment through an industrial protocol interface, and push abnormal event information to the control center, inspection personnel, and passenger terminals.

[0089] In the above scheme, the emergency handling method is decomposed into modules such as data acquisition, visual analysis, spatial mapping, anomaly detection and emergency response. Each module is responsible for a specific function, and the interfaces between modules are clear, which facilitates the development, implementation and maintenance of the system, reduces the complexity of the system, and improves maintainability and scalability. The modules work together in the order of data flow and control flow. The data acquisition module is responsible for data input, the visual analysis and spatial mapping modules are responsible for data processing and feature extraction, the anomaly detection module is responsible for abnormal event judgment, and the emergency response module is responsible for executing control instructions and information push. The emergency handling method process is fully realized and a complete emergency handling system is constructed. Through modular system design, the emergency handling method is implemented as an operational system, which can automatically and intelligently complete the perception, identification, decision-making and response of abnormal events, significantly improving the efficiency and reliability of urban rail transit emergency handling, reducing the need for manual intervention, and improving the overall performance of the system.

[0090] Compared with the prior art, the present invention has the following beneficial effects:

[0091] The present invention discloses an urban rail transit emergency response method and system based on a large-scale visual model. This method cleverly combines traditional image processing technology with modern visual analysis methods to effectively improve the accuracy and robustness of anomaly detection while ensuring real-time performance and low computational complexity. Specifically:

[0092] (1) Efficient moving target detection is achieved through dynamic frame difference and spatiotemporal aggregation technology;

[0093] (2) Through geometric mapping technology, the precise association between two-dimensional image information and three-dimensional spatial environment is achieved, overcoming the defect of traditional visual systems that lack spatial information;

[0094] (3) Through multimodal data fusion, comprehensive utilization of visual, sound and equipment operation data, a more comprehensive and accurate abnormal event judgment is achieved;

[0095] (4) Through hierarchical emergency response strategies and automated control, the efficiency and intelligence level of emergency handling are improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0096] Figure 1 This is a flow chart of an urban rail transit emergency response method based on a large visual model provided by an embodiment of the present invention;

[0097] Figure 2 This is a schematic diagram of a module of an urban rail transit emergency response system based on a visual large model provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0098] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0099] As Figure 1 shown, this application provides an urban rail transit emergency handling method based on a large vision model, including:

[0100] S1: Collect visual data, sound data, and equipment operation data in key areas of the platform and carriage of urban rail transit, perform preprocessing, and simultaneously achieve timestamp synchronization;

[0101] In an embodiment provided by this application, in key areas of the urban rail transit platform and carriage, such as the escalator entrance, security check area, passenger waiting area on the platform, and passenger area inside the carriage, multiple groups of low-resolution wide-angle cameras, sound sensors, and equipment operation data collection modules are deployed to achieve maximum coverage of the monitoring area and collect multi-source heterogeneous data; specifically, multiple low-resolution wide-angle cameras are deployed, preferably wide-angle cameras with a resolution of 640×480 pixels and a viewing angle of 130°, such as industrial cameras of model XYZ-Cam-001. In key areas of the platform, such as the escalator entrance, security check area, and waiting area, 2-3 cameras are configured in each key area to form multi-angle redundant monitoring and improve the reliability of the system. Inside the carriage, wide-angle cameras are installed at both ends of each carriage to ensure that the monitoring view can fully cover the passenger activity area inside the carriage. The image acquisition frequency of the camera is set to 10 frames per second, which can not only meet the requirements of motion detection but also effectively reduce the burden of data transmission and processing.

[0102] At the same time, sound sensors, such as high-sensitivity microphones of model ABC-Mic-002, are synchronously deployed near the camera deployment area to collect environmental sound data, such as passenger calls for help and abnormal equipment noises. In addition, through an equipment control interface, such as an interface compliant with the Modbus protocol, the operation data of key rail transit equipment is collected in real time, such as the running speed of the escalator, the opening and closing status of the turnstile, and the opening and closing status of the platform door.

[0103] The collected visual data, sound data, and device operation data are all preprocessed and timestamp synchronized. The preprocessing of visual data mainly includes denoising the original images. For example, the median filter algorithm is used to filter out the noise in the images to improve the accuracy of subsequent motion detection. The preprocessing of sound data includes operations such as analog-to-digital conversion and filter denoising of the original audio signals to extract effective sound features. The preprocessing of device operation data includes operations such as data format conversion, data cleaning, and outlier processing to ensure the accuracy and availability of the data. The timestamp synchronization operation uses the Network Time Protocol (NTP) or the Precision Time Protocol (PTP) to ensure that the data collected by different sensors are aligned on the time axis, laying a foundation for subsequent multi-modal data fusion analysis.

[0104] Through the collaborative work of multi-source sensors, the comprehensive and synchronous collection of visual, sound, and device operation data in the urban rail transit environment is realized, providing multi-dimensional data support for subsequent intelligent analysis. By using low-resolution cameras and a moderate acquisition frequency, the data acquisition and transmission pressure are effectively reduced on the premise of ensuring the monitoring effect, meeting the application scenario requirements of edge computing. The timestamp synchronization mechanism ensures the consistency of multi-modal data in the time dimension, providing a data basis for subsequent fusion analysis.

[0105] S2: Perform dynamic frame difference processing on consecutive video frames of visual data to generate a two-dimensional motion heat map, and perform spatio-temporal aggregation on the two-dimensional motion heat map;

[0106] In an embodiment provided by the present application, first, the consecutive video frames are preprocessed by Gaussian blur. For example, convolution filtering is performed using a Gaussian kernel with σ = 1.5 to effectively eliminate interference factors such as the noise of the camera itself and environmental light changes. Then, the pixel differences between adjacent two frames of images are calculated. For example, the absolute difference method is used to obtain the difference image. To adapt to the changes in light and noise levels in different scenarios, an environment-adaptive algorithm is used to dynamically adjust the difference threshold. For example, the threshold is dynamically adjusted based on the average value and standard deviation of the image grayscale to ensure the sensitivity and robustness of motion area detection. The difference image is binarized, that is, the pixel points with pixel difference values greater than the dynamic threshold are set to white, and the pixel points less than the threshold are set to black, thereby generating a binary motion heat map that highlights the motion areas in the video frames;

[0107] To further suppress noise interference and extract stable moving objects, spatio-temporal aggregation is performed on the binary motion heatmaps of 5 consecutive frames. For example, aggregation is carried out in a pixel-by-pixel summation manner, where the pixel values at corresponding positions in the 5 consecutive frame heatmaps are accumulated to obtain the aggregated heatmap. Then, morphological closing operation is applied to the aggregated heatmap. For example, dilation operation is first performed, followed by erosion operation, to fill the holes inside the moving region and remove isolated noise points, resulting in a more complete and continuous moving region. Finally, a connected component labeling algorithm, such as the eight-neighbor connected component labeling algorithm, is used to identify and label the independent moving regions in the aggregated heatmap, and a unique identifier is assigned to each independent moving region to facilitate subsequent object tracking and behavior analysis;

[0108] Compared with the object detection method based on deep learning, the dynamic frame difference has lower computational complexity, less resource consumption, and faster processing speed, and can meet the high real-time requirements of emergency handling in the urban rail transit scenario. Through means such as Gaussian blur, dynamic threshold adjustment, spatio-temporal aggregation, and morphological processing, environmental noise interference is effectively suppressed, and the accuracy and robustness of motion detection are improved. Compared with the traditional fixed threshold frame difference method, the environment adaptive threshold adjustment can better adapt to the complex and changeable rail transit environment and improve the environmental adaptability of the system.

[0109] S3: Construct a 3D digital model based on the engineering drawings of the platform and carriage of urban rail transit, perform calibration of the internal and external parameters of the camera, and establish a mapping relationship between the 2D motion heatmap and the 3D digital model;

[0110] In an embodiment provided by the present application, in order to map the moving region in the 2D image to the 3D space of urban rail transit and associate it with actual equipment components, a 3D digital model is constructed based on the CAD engineering drawings of the platform and carriage of urban rail transit, the internal and external parameters of the camera are calibrated, and a mapping relationship between the 2D motion heatmap and the 3D digital model is established;

[0111] First, based on the CAD engineering drawings of the rail transit platform and carriage, a precise 3D digital model is constructed using 3D modeling software, such as AutoCAD or Revit. The model should include the precise spatial positions and geometric parameters of all fixed facilities inside the platform and carriage, such as escalators, turnstiles, seats, platform doors, etc. A unique ID and functional label are assigned to each equipment component, such as "escalator - 001", "turnstile - Area A - 002", "platform door - Platform 3 - 005", etc., and an equipment component database is established to store information such as the ID, functional label, and 3D model data of the equipment components.

[0112] Then, using a calibration board, such as a checkerboard calibration board, take pictures at the camera installation position and use the camera calibration algorithm to estimate the internal and external parameters of the camera. The internal parameters include the focal length of the camera, the coordinates of the principal point, the distortion coefficients, etc., and the external parameters include the rotation matrix and the translation vector of the camera in the world coordinate system. Through camera calibration, the internal parameter matrix and the external parameter matrix of the camera are obtained, providing a parameter basis for subsequent coordinate system conversion.

[0113] Next, establish the conversion relationship between the camera view and the CAD model space coordinate system. Take the world coordinate system of the CAD model as the reference coordinate system and convert the camera coordinate system to the world coordinate system through the external parameter matrix. Further, use the internal parameter matrix to project the three-dimensional points in the world coordinate system onto the two-dimensional image plane to establish the correspondence between the three-dimensional space points and the two-dimensional image pixel points.

[0114] Finally, apply an affine transformation to map the moving regions in the two-dimensional heat map to the three-dimensional CAD model. For each moving region in the two-dimensional heat map, calculate the pixel coordinate range in the image coordinate system, and use the established two-dimensional to three-dimensional coordinate mapping relationship to back-project this pixel coordinate range into the three-dimensional CAD model to obtain the position and range of the moving region in three-dimensional space. Through spatial occupancy detection, determine whether the moving region overlaps with the pre-set spatial position of the equipment components, thereby identifying the moving regions that overlap with the spatial position of the equipment components;

[0115] Through geometric mapping technology, the accurate conversion between two-dimensional image information and three-dimensional space information is realized, and the moving targets detected in the two-dimensional image are accurately located in the three-dimensional urban rail transit scene. The positioning accuracy reaches ±5 cm, meeting the requirements for positioning accuracy in urban rail transit emergency handling. This method does not rely on deep learning training data, does not require a large amount of labeled data for model training, reduces the development and maintenance costs of the system, reduces the dependence of the system on labeled data, and improves the versatility and scalability of the system. By comparing with the spatial positions of the equipment components in the CAD model, the association between the moving regions and the equipment components is realized, laying a foundation for subsequent abnormal behavior analysis based on the equipment components.

[0116] S4: Detect abnormal behaviors in the mapping results of the two-dimensional motion heat map after spatio-temporal aggregation and the three-dimensional digital model, and perform multi-modal fusion analysis by combining sound data and equipment operation data to determine the severity of the abnormal event;

[0117] In an embodiment provided by the present application, detect abnormal behaviors in the moving regions mapped onto the three-dimensional CAD model, and perform multi-modal fusion analysis by combining sound data and equipment operation data to improve the accuracy and robustness of abnormal detection, and trigger corresponding warnings according to the severity of the abnormal event.

[0118] First, for each device component ID, calculate its movement frequency within the time window w. The time window w can be set according to the actual application scenario, for example, set to 1 second or 5 seconds. The calculation method of the movement frequency can be to count the number or the rate of change of the area of the movement region that overlaps with the spatial position of the device component within the time window w. Further, through Fourier transform, analyze the spectral characteristics of the time series signal of the device component's movement, and extract the main frequency and its amplitude of the movement. Fourier transform can effectively analyze periodic movements, such as the cyclic movement of escalator steps, the periodic opening and closing of the turnstile swing arm, etc.

[0119] Then, based on the device type, pre-define an abnormal state threshold table . This threshold table stores the abnormal state thresholds for different types of device components, for example:

[0120] Escalator steps: movement frequency > 5Hz or movement amplitude > 3 times the standard deviation;

[0121] Turnstile: movement frequency > 2Hz or movement range exceeds the normal working range by 20%;

[0122] Platform door: movement frequency > 1Hz or movement detection during abnormal opening and closing periods;

[0123] When it is detected that the movement frequency or amplitude of the device component exceeds its corresponding abnormal state threshold, it is preliminarily determined that the device component is in an abnormal state, and an alarm for the corresponding component is triggered.

[0124] To further improve the accuracy of abnormal detection and reduce the false alarm rate, multi-modal abnormal confirmation is performed. Integrate the collected visual abnormal score , sound abnormal score and device operation data abnormal score for comprehensive judgment. The visual abnormal score can be calculated based on the movement frequency and amplitude. For example, when the movement frequency or amplitude exceeds the threshold, a higher visual abnormal score is given. The sound abnormal score can be obtained by analyzing the sound data collected by the sound sensor. For example, when screams, calls for help, abnormal device noises, etc. are detected, a higher sound abnormal score is given. The device operation data abnormal score can be obtained by analyzing the device operation data. For example, abnormal escalator speed, turnstile jamming, abnormal opening of the platform door, etc., a higher device operation data abnormal score is given.

[0125] The multi-modal abnormal confirmation function is defined as a weighted summation model:

[0126] .

[0127] in, 、 、 is the weight coefficient, which is set according to the importance of different modal data in anomaly detection. In this embodiment, =0.5, =0.3, =0.2. Tc is a comprehensive threshold used to determine whether an abnormal event is finally confirmed. In this embodiment, Tc is set to 0.7. > Tc, it is finally confirmed that an abnormal event has occurred in the equipment component ID and an early warning is triggered.

[0128] By analyzing the motion frequency of equipment components, it is possible to effectively detect abnormal motion states of equipment components, such as abnormal escalator vibration, gate jamming, and abnormal platform door opening. The multimodal anomaly detection mechanism integrates visual, acoustic, and equipment operation data, fully utilizing multi-source information to improve the accuracy and robustness of anomaly detection and effectively reduce the false alarm rate. By setting abnormal state thresholds for different equipment types, refined anomaly determination based on different equipment characteristics is achieved, enhancing the flexibility and adaptability of the system. Compared with single-modality anomaly detection methods, multimodal fusion strategies can more comprehensively and accurately identify abnormal events in urban rail transit environments, improving system reliability and safety.

[0129] S5: Develop emergency response strategies based on the severity of abnormal events, send control instructions to rail transit equipment through industrial protocol interfaces, and push abnormal event information to the control center, inspection personnel, and passenger terminals.

[0130] In one embodiment provided in the present application, based on the determined abnormal events, they are graded according to the severity of the abnormalities, and corresponding emergency response strategies are formulated. Control instructions are sent to rail transit equipment through the industrial protocol interface, and abnormal event information is pushed to the control center, patrol personnel and passenger terminals.

[0131] According to the severity of the abnormal event, the emergency response is divided into three levels:

[0132] Level 1 response (minor abnormality): For example, slight vibration of equipment components, occasional abnormal noise, etc. are judged as minor abnormalities, triggering the equipment self-test program to perform self-diagnosis and self-repair. It does not affect the normal operation of the equipment, but only sends an alarm information to the control center to prompt attention.

[0133] Secondary response (obvious anomaly): For example, if the movement frequency of equipment components is significantly abnormal, there is a sound alarm, or the equipment operation data exceeds the normal range, it is determined as an obvious anomaly, triggering equipment deceleration or current limiting measures, such as escalators running at a reduced speed, turnstiles restricting the flow of people, platform doors reducing the switching speed, etc., to reduce safety risks and send early warning information to the control center and inspection personnel to prompt for manual intervention.

[0134] Tertiary response (severe anomaly): For example, if the equipment components experience severe vibrations, emit abnormal loud noises, the equipment operation data is severely abnormal or even the equipment stops operating, it is determined as a severe anomaly, and the equipment operation is immediately stopped, such as escalators making an emergency brake, all turnstiles closing, platform doors making an emergency stop, etc., to ensure the safety of passengers to the greatest extent, and start the emergency plan, sending emergency early warning information to the control center, inspection personnel, and passenger terminals to guide the evacuation of passengers and notify professional maintenance personnel to conduct on-site disposal.

[0135] According to the preset emergency response strategy, the system sends control instructions to rail transit equipment through industrial protocol interfaces, such as Modbus TCP / IP, Profinet, etc., to achieve real-time intervention and control of the equipment. For example, sending deceleration or stop instructions to the escalator controller, sending closing instructions to the turnstile controller, sending emergency stop instructions to the platform door controller, etc.

[0136] At the same time, the abnormal event information, including the anomaly level, anomaly type, occurrence time, occurrence location, equipment component ID, etc., is pushed to the urban rail transit control center through network communication protocols, such as MQTT, HTTP, etc., for the control center personnel to conduct unified monitoring and management. The early warning information is pushed to the mobile terminals of inspection personnel, such as handheld PDAs or smartphones, to facilitate the inspection personnel to quickly reach the scene for disposal. In addition, through passenger terminals such as passenger information display screens, broadcast systems, and mobile phone APPs in the platform and carriage, abnormal event information and evacuation guidance information can be released to passengers to ensure the right to know and safety of passengers.

[0137] It realizes hierarchical response according to the severity of abnormal events, can take differentiated emergency treatment measures for different levels of abnormal events, taking into account both safety and operation efficiency. Through the industrial protocol interface, real-time control of rail transit equipment is achieved, realizing automated emergency response, improving the timeliness and effectiveness of emergency handling, and reducing the need for manual intervention. The multi-channel information push mechanism ensures that abnormal event information can be transmitted to the control center, inspection personnel, and passengers in a timely manner, improving the collaborative efficiency of emergency response and the self-rescue ability of passengers, and ensuring the safe and stable operation of the urban rail transit system to the greatest extent.

[0138] Preferably, the collection of visual data, sound data, and equipment operation data includes:

[0139] Deploy wide-angle cameras at the escalator entrances, security check areas, waiting areas on the platform, and at both ends of the carriages, with the acquisition frequency set at 10 frames per second;

[0140] Use a microphone array to collect ambient sound data and extract sound features, where the sound features include frequency, energy, and spectral changes;

[0141] Collect device operation data in real time through the industrial protocol interface of rail transit equipment, where the device operation data includes speed, current, temperature, and vibration.

[0142] In an embodiment provided by this application, wide-angle cameras are deployed in key areas on the platform. Specifically, in crowded areas such as escalator entrances, security check areas, and waiting areas, wide-angle cameras with a resolution of 640×480 and a viewing angle of 130° are installed respectively. 2 - 3 cameras are configured in each key area to achieve redundant coverage and improve the reliability of monitoring. At the same time, wide-angle cameras of the same specification are also installed at both ends of each carriage in the carriages to ensure that the monitoring view covers all passenger areas. The acquisition frequency of all cameras is uniformly set at 10 frames per second. This setting can not only meet the requirements of motion detection but also effectively reduce the burden of data transmission.

[0143] Install microphone arrays on the platform and in the carriages to collect ambient sound data. The microphone array adopts a linear arrangement, and each array contains 8 high-sensitivity microphones with a spacing of 10 cm. Through sound processing algorithms, sound features including frequency, energy, and spectral changes are extracted from the collected original sound data. The frequency feature reflects the pitch of the sound, the energy feature represents the intensity of the sound, and the spectral change can reflect the time-varying characteristics of the sound. The combination of these features can effectively identify abnormal sound events such as screams, collisions, or explosions.

[0144] Finally, collect device operation data in real time through the industrial protocol interface of rail transit equipment. Specifically, for different types of equipment, the collected data items are as follows:

[0145] For escalators and moving walkways: collect motor speed, current, temperature, and vibration data;

[0146] For platform screen doors: collect switch status, motor current, and door position data;

[0147] For air conditioning systems: collect temperature, humidity, wind speed, and energy consumption data;

[0148] For trains: collect speed, acceleration, door status, and passenger load data.

[0149] These device operation data are transmitted through a standardized industrial Ethernet protocol with a sampling frequency of 100Hz to ensure the real-time and accuracy of the data.

[0150] Preferably, the preprocessing of the visual data, sound data, and device operation data while achieving timestamp synchronization includes:

[0151] Performing Gaussian blur processing on the visual data;

[0152] Performing noise reduction processing on the sound data;

[0153] Performing normalization processing on the device operation data;

[0154] Using the Network Time Protocol or the Precision Time Protocol to perform timestamp synchronization on the preprocessed visual data, sound data, and device operation data to ensure the accuracy of data fusion.

[0155] In an embodiment provided by the present application, Gaussian blur processing is performed on the visual data collected from the wide-angle camera. Specifically, a 5×5 Gaussian kernel is used to perform a convolution operation on each frame of the image, and the standard deviation of the Gaussian kernel is set to 1.5. Gaussian blur processing can effectively reduce the noise and details in the image, making the subsequent dynamic frame difference processing more stable and reliable. The processed image can be expressed as:

[0156] ;

[0157] where is the original image, is the Gaussian kernel, and * represents the convolution operation.

[0158] For the sound data collected through the microphone array, an adaptive filtering algorithm is used for noise reduction processing. The specific steps are as follows:

[0159] a) Using the fast Fourier transform to convert the time-domain signal into a frequency-domain signal;

[0160] b) Estimating the power spectral density of the background noise;

[0161] c) Calculating the signal-to-noise ratio of each frequency component according to the estimated noise power spectral density;

[0162] d) Designing a Wiener filter based on the signal-to-noise ratio;

[0163] e) Applying the Wiener filter to the spectrum of the original signal;

[0164] f) Using the inverse fast Fourier transform to convert the processed spectrum back into a time-domain signal.

[0165] This noise reduction method can effectively suppress the background noise while retaining the key sound feature information.

[0166] For the operation data collected from rail transit equipment, including speed, current, temperature, vibration, etc., normalization processing is carried out. The purpose of normalization processing is to unify data with different dimensions to the same scale, facilitating subsequent multi-modal fusion analysis. Specifically, the min-max normalization method is adopted, and the processing formula is as follows:

[0167]

[0168] Among them, X is the original data, and are the minimum and maximum values of this type of data respectively, is the normalized data. The range of the normalized data is uniformly [0, 1].

[0169] To ensure the accuracy of multi-modal data fusion, the Precision Time Protocol is used to synchronize the timestamps of the preprocessed visual data, sound data, and equipment operation data. The specific steps are as follows:

[0170] a) Set a master clock server in the system as the time reference;

[0171] b) Each data acquisition device acts as a slave device and synchronizes time with the master clock server;

[0172] c) The master clock server periodically sends synchronization messages to the slave devices, including the sending timestamp;

[0173] d) The slave device receives the synchronization message and records the receiving timestamp;

[0174] e) The slave device sends a delay request message to the master clock server;

[0175] f) The master clock server receives the delay request message and replies with a delay response message containing the receiving timestamp;

[0176] g) The slave device calculates the time deviation and network delay from the master clock server according to the received timestamp information;

[0177] h) The slave device adjusts its local clock according to the calculation result to achieve synchronization with the master clock server.

[0178] Through the PTP protocol, the clock synchronization accuracy of each device in the system can be controlled at the microsecond level, ensuring the time consistency of multi-modal data.

[0179] Through Gaussian blur processing, the image noise is effectively reduced, and the accuracy of subsequent dynamic frame difference is improved. In specific implementation, this method can reduce the false detection rate by about 30%. By using an adaptive filtering algorithm for noise reduction, background noise can be effectively suppressed while key sound feature information is retained. The test results show that this method can increase the signal-to-noise ratio by about 6 dB, significantly improving the quality of sound feature extraction. Through normalization processing, data with different dimensions are unified into the range of [0, 1], which is convenient for subsequent multi-modal fusion analysis. This processing method can eliminate the influence of dimension differences on the analysis results and improve the accuracy of anomaly detection. By using the Precision Time Protocol (PTP) to achieve time synchronization of multi-modal data, the clock synchronization accuracy of each device in the system can be controlled at the microsecond level. This high-precision time synchronization ensures the accuracy of multi-modal data fusion and lays a foundation for subsequent anomaly behavior detection and multi-modal fusion analysis. In specific implementation, after adopting the PTP protocol, the time consistency error of multi-modal data can be controlled within 1 ms, and the synchronization accuracy is improved by about two orders of magnitude compared with the traditional Network Time Protocol.

[0180] Preferably, performing dynamic frame difference processing on consecutive video frames of visual data to generate a two-dimensional motion heat map includes:

[0181] Calculating the pixel difference between adjacent frames of consecutive video frames, and dynamically adjusting the difference threshold through an environment adaptive algorithm to obtain a difference image; binarizing the difference image to generate a binary motion heat map.

[0182] In an embodiment provided by the present application, each frame image in the input consecutive video frame sequence is preprocessed by Gaussian blur. Gaussian blur is an effective image filtering technology, the purpose of which is to eliminate the noise that may be introduced during video acquisition, such as light changes, sensor noise, etc. By applying a Gaussian filter, the high-frequency noise components in the image can be smoothed, and the main structural information of the image can be retained, thereby improving the robustness and accuracy of subsequent motion detection. In this embodiment, a Gaussian kernel of, for example, 3×3 or 5×5 can be selected to perform a convolution operation on each frame image to achieve image smoothing.

[0183] After completing the Gaussian blur preprocessing, for the consecutive video frame sequence, calculate the pixel value difference between corresponding pixel points of adjacent two frame images. Specifically, for the frame image and the frame image where ≥2), calculate the pixel value difference at each pixel position Here, absolute value operation is adopted to ensure that the difference value is non - negative, representing the magnitude of pixel value change. Through pixel - by - pixel calculation, a difference image is obtained. , which reflects the change of pixel values between adjacent frames. Areas with large pixel value differences usually correspond to moving areas in the video scene.

[0184] To adapt to the complex and changeable lighting conditions and environmental noise in the urban rail transit scene, this embodiment introduces an environment - adaptive algorithm to dynamically adjust the binarization threshold. Traditional fixed - threshold methods are prone to false detection or missed detection when the lighting changes or the noise level fluctuates. The environment - adaptive algorithm can dynamically adjust the threshold according to the current scene's noise level, improving the accuracy and adaptability of motion detection.

[0185] For example, an adaptive threshold method based on statistical characteristics can be adopted. First, within a certain period of time (such as the recent several frames), the pixel value distribution of the difference image is statistically analyzed, and the mean value and standard deviation of the pixel values are calculated. Then, based on these statistics, the binarization threshold is dynamically set. A possible threshold - setting strategy is , where k is an adjustable parameter used to control the sensitivity of the threshold. The parameter k can be adjusted according to the actual application scenario. For example, in an environment with less noise, the value of k can be appropriately reduced to increase the sensitivity; in an environment with more noise, the value of k can be appropriately increased to reduce the false detection rate. In this way, the threshold can be adaptively adjusted according to the environmental noise level.

[0186] Using the dynamically adjusted threshold , the difference image is binarized to generate a binary motion heat map . For each pixel point in the difference image, its pixel value is compared with the adaptive threshold of the current frame. If , it is considered that this pixel point corresponds to a moving area, and its pixel value in the motion heat map is set to a highlight value (such as white or pixel value 255); if , it is considered that this pixel point corresponds to a background area, and its pixel value in the motion heat map is set to a low value (such as black or pixel value 0). The binarization is performed through the following formula:

[0187] When , = 255; when , = 0;

[0188] After binarization, the obtained binary image is the two-dimensional motion heat map. In this heat map, the white area represents the detected motion area, and the black area represents the static background area.

[0189] In this embodiment, through the above dynamic frame difference technology, the motion area is effectively extracted from the video data, and a two-dimensional motion heat map is generated. Compared with traditional motion detection methods such as background modeling or optical flow method, the dynamic frame difference technology has the significant advantage of low computational complexity, and is especially suitable for the urban rail transit emergency handling scenario with high real-time requirements. The computational complexity of this step is reduced by about 85% compared with the deep learning object detection method, and the processing speed is increased to the real-time level (<50 ms / frame). This means that under the condition of limited computing power of the edge computing node, this embodiment can still ensure that the system quickly processes the video data, timely detects abnormal motion, and gains valuable time for subsequent emergency response.

[0190] In addition, by introducing an environment adaptive threshold adjustment mechanism, this embodiment improves the robustness of the motion detection algorithm to environmental changes, reduces the false detection rate and missed detection rate caused by light changes or noise interference, and ensures the quality and reliability of the motion heat map.

[0191] Preferably, the spatio-temporal aggregation of the two-dimensional motion heat map includes:

[0192] Performing temporal aggregation on the two-dimensional motion heat maps of five consecutive frames;

[0193] Performing spatio-temporal aggregation on the aggregated binary motion heat map, applying morphological closing operation to remove interference points; identifying and labeling independent motion regions through the connected component labeling algorithm, and assigning unique identifiers.

[0194] In an embodiment provided by the present application, for the binary motion heat maps of five consecutive frames, they are respectively denoted as H1, H2, H3, H4, and H5. These heat maps reflect the motion changes in the monitoring area during a continuous time period. In order to enhance the saliency of the motion target and suppress transient noise interference, first perform temporal aggregation on the binary motion heat maps of five consecutive frames; the specific method is to perform a pixel-by-pixel logical OR operation on these five heat maps.

[0195] The aggregation formula is:

[0196] H_aggregation = H1 OR H2 OR H3 OR H4 OR H5;

[0197] Among them, "OR" represents a logical OR operation. If a certain pixel is marked as moving (pixel value is 1) in any one of the five-frame heatmaps, then this pixel in the aggregated heatmap is also marked as moving.

[0198] Through aggregation in the time dimension, it is possible to effectively retain the targets with continuous movement and filter out short-term and isolated noise points. For example, if a pedestrian moves continuously, even if some pixels are not detected due to changes in lighting or occlusion in some frames, the aggregation operation can still ensure that the overall moving area of the pedestrian is retained.

[0199] After aggregation in the time dimension, there may be some small holes or breaks in the heatmap. These holes may be caused by shadows or low-contrast regions inside the target. To fill these holes and connect adjacent moving regions, morphological closing operations are applied.

[0200] Morphological closing operations include dilation and erosion:

[0201] Dilation includes using a structuring element (e.g., a 3x3 square) to perform a dilation operation on the aggregated heatmap. The dilation operation will expand the edges of the moving region outward to fill small holes.

[0202] Erosion includes using the same structuring element to perform an erosion operation on the dilated heatmap. The erosion operation will contract the edges of the moving region inward to eliminate small noise points.

[0203] Morphological closing operations can effectively fill the holes inside the moving region and connect adjacent moving regions, thus making the contour of the moving target more complete and smooth. For example, if a pedestrian's arm is separated from the body during movement, the closing operation can connect the arm and the body to form a complete moving region.

[0204] After morphological closing operations, there may be multiple independent moving regions in the heatmap, and each region corresponds to one or more moving targets. To distinguish these different moving targets, a connected component labeling algorithm is applied; the connected component labeling algorithm will group all adjacent moving pixels (pixel value is 1) in the heatmap into a connected component and assign a unique identifier to each connected component. Commonly used connected component labeling algorithms include the four-neighborhood algorithm and the eight-neighborhood algorithm.

[0205] The connected component labeling algorithm can distinguish different moving targets in the heatmap and provide a basis for subsequent analysis and processing. For example, if there are multiple pedestrians in the heatmap, the connected component labeling algorithm can label each pedestrian as an independent moving region for subsequent analysis of the behavior of each pedestrian.

[0206] After completing the labeling of connected regions, a unique identifier is assigned to each independent motion region. This identifier can be an integer or a string. The identifier is used to reference and operate on specific motion regions in subsequent processing; by assigning unique identifiers, it is convenient to track and analyze each motion region. For example, the identifier, location, size, and motion trajectory of each motion region can be recorded for subsequent analysis and prediction of the behavior of moving objects.

[0207] Preferably, the construction of the three-dimensional digital model based on the engineering drawings of the platform and carriage of urban rail transit includes:

[0208] Construct an accurate three-dimensional digital model based on the CAD engineering drawings of the platform and carriage. The three-dimensional digital model includes the spatial positions and geometric parameters of fixed facilities such as escalators, ticket gates, and seats; assign unique IDs and function labels to the equipment components of each fixed facility to construct an equipment component database.

[0209] In an embodiment provided by the present application, first, obtain the CAD engineering drawings of the urban rail transit platform. These drawings should contain detailed structural information of the platform, including but not limited to:

[0210] Overall layout of the platform: Overall dimensions such as the length, width, and height of the platform.

[0211] Positions of fixed facilities: Spatial position coordinates of fixed facilities such as escalators, ticket gates, waiting seats, platform doors, columns, and walls.

[0212] Geometric parameters: Detailed geometric dimensions of each fixed facility, such as the step height and width of the escalator, the passage width of the ticket gate, and the length, width, and height of the seat.

[0213] Material information: Material types of different facilities, such as metal, glass, concrete, etc.

[0214] The data format of the CAD engineering drawings can be common formats such as DWG and DXF.

[0215] Import the CAD engineering drawings into 3D modeling software, such as AutoCAD, SolidWorks, Blender, etc. According to the layer information of the CAD drawings, import different types of facilities into the software separately; ensure that all imported facility models are in the same coordinate system. Usually, establish a global coordinate system with the center position of the platform or a certain fixed point as the origin; according to the dimension markings on the CAD drawings, calibrate the accuracy of the imported models to ensure that the dimensions of the models are consistent with the actual dimensions. The provided measurement tools in the software can be used for verification; improve the details of the models, such as adding the texture of the escalator, the indicator lights of the turnstile, the backrest of the seat, etc. These details can improve the realism and recognizability of the models; optimize the models to reduce the number of faces and vertices of the models, so as to improve the rendering speed of the models and reduce the storage space. The provided optimization tools in the software can be used for processing.

[0216] Assign a unique ID to each equipment component of the fixed facilities. The naming rules of the ID can be defined according to the facility type and location, such as "Escalator-001", "Turnstile-A Area-002", etc.;

[0217] Define one or more functional labels for each equipment component to describe the function of the component. For example:

[0218] Escalator: Entrance, Exit, Steps, Handrail

[0219] Turnstile: Passage, Card Swiping Area, Display Screen, Baffle

[0220] Seat: Seat, Backrest, Handrail

[0221] Platform Door: Door Body, Glass, Sensor

[0222] Store the ID, functional labels of the equipment components and the corresponding 3D model information in the database. A relational database can be used for the database. Export the constructed 3D digital model into a common 3D model format, such as OBJ, FBX, STL, etc. These formats can be read and used by other software or systems.

[0223] Thus, an accurate three-dimensional digital model of the urban rail transit platform can be constructed. The model includes the accurate spatial positions and geometric parameters of all fixed facilities within the platform, providing an accurate reference for subsequent motion heat map mapping. The positioning accuracy can reach ±5 cm; each device component in the model has a unique ID and functional label, facilitating system identification and analysis. For example, the system can identify whether a certain motion area is on the steps of the escalator based on the ID, thereby determining whether there are potential safety hazards; the model can be easily extended and updated. For example, when new equipment is added or replaced at the platform, only the CAD engineering drawings need to be updated, and then the model can be reconstructed; it does not rely on deep learning training data, reducing the system's dependence on labeled data and lowering the system's development and maintenance costs; combined with the three-dimensional digital model, abnormal behaviors can be judged more accurately. For example, it can be determined whether a certain motion area exceeds the normal working range of the turnstile, thereby improving the accuracy of abnormal detection.

[0224] Preferably, the execution of the internal and external parameter calibration of the camera and the establishment of the mapping relationship between the two-dimensional motion heat map and the three-dimensional digital model include:

[0225] Using a calibration board to perform the internal and external parameter calibration of the camera, and obtaining the internal parameter matrix and external parameter matrix of the camera;

[0226] Based on the results of the internal and external parameter calibration, applying an affine transformation to map the motion area in the two-dimensional motion heat map to the spatial coordinate system of the three-dimensional digital model;

[0227] Through a spatial occupancy detection algorithm, identifying the motion areas that overlap with the device components with unique IDs in the three-dimensional model in terms of spatial position.

[0228] In an embodiment provided by the present application, first, use a calibration board to perform the internal and external parameter calibration of the camera, and obtain the internal parameter matrix and external parameter matrix of the camera. Specifically, use a 9×6 black and white checkerboard calibration board, place the calibration board at different angles and distances within the camera's field of view and take at least 20 images. Then, calculate the internal parameter matrix of the camera through the Zhang Zhengyou calibration method and the external parameter matrix . Among them, the internal parameter matrix contains parameters such as focal length, principal point coordinates, and distortion coefficients; the external parameter matrix describes the rotation R and translation t of the camera relative to the world coordinate system.

[0229] Next, based on the results of the internal and external parameter calibration, apply an affine transformation to map the motion area in the two-dimensional motion heat map to the spatial coordinate system of the three-dimensional digital model. The specific steps are as follows:

[0230] (1) For each pixel point in the two-dimensional motion heat map , using the inverse matrix of the intrinsic parameter matrix K Convert it to normalized image coordinates : ;

[0231] (2) Using the extrinsic parameter matrix Convert the normalized image coordinates to three-dimensional points in the world coordinate system ;

[0232] Among them, Z is the distance from this point to the camera, which can be obtained through intersection calculation with the three-dimensional digital model.

[0233] (3) Align the obtained three-dimensional points in the world coordinate system with the coordinate system of the three-dimensional digital model to complete the mapping from the two-dimensional motion heat map to the three-dimensional space.

[0234] Finally, through the space occupancy detection algorithm, identify the motion areas that overlap with the device components with unique IDs in the three-dimensional model in terms of spatial position. The specific implementation method is as follows:

[0235] (1) For each device component in the three-dimensional digital model, construct a bounding box according to its geometric parameters.

[0236] (2) For each motion area mapped to the three-dimensional space, calculate its intersection with the bounding boxes of each device component.

[0237] (3) If there is an intersection between a certain motion area and the bounding box of a device component, it is considered that the motion area overlaps with the device component in terms of spatial position.

[0238] (4) Record the corresponding relationship between the overlapping motion areas and the device component IDs for subsequent abnormal behavior analysis.

[0239] Through camera calibration technology, an accurate mapping between the two-dimensional image and the three-dimensional space is achieved, and the positioning accuracy can reach ±5 cm, providing a reliable spatial information basis for subsequent abnormal behavior detection; the affine transformation method is used for coordinate conversion, with relatively low computational complexity, which can be processed in real time on edge computing devices to meet the real-time requirements; the space occupancy detection algorithm can effectively identify the motion areas related to specific device components, providing an important basis for the accurate positioning and classification of abnormal behaviors; it does not rely on the deep learning method of a large amount of labeled data, reducing the system's dependence on training data and improving the system's versatility and scalability; by establishing the corresponding relationship between the motion areas and the device component IDs, it provides a structured input for subsequent multi-modal abnormal detection and emergency response, which is conducive to improving the decision-making accuracy and response speed of the system.

[0240] Preferably, the abnormal behavior detection of the mapping result of the two-dimensional motion heat map after spatio-temporal aggregation and the three-dimensional digital model includes:

[0241] For the motion area that spatially overlaps with the device component with a unique ID, calculate the motion frequency of the device component within a preset time window, analyze the spectral characteristics of the motion through Fourier transform, and extract the motion frequency and its amplitude; based on the preset abnormal state threshold table of the device type, trigger the alarm of the corresponding component when it is detected that the motion frequency or amplitude exceeds the threshold.

[0242] In an embodiment provided by the present application, taking the escalator on the urban rail transit platform as an object, the abnormal behavior of the escalator device components is detected to realize the emergency treatment of urban rail transit.

[0243] When the wide-angle camera on the platform monitors that there are passengers moving at the escalator entrance, the motion heat map will mark the motion area of the passengers. Subsequently, this two-dimensional motion area is converted into the three-dimensional digital model for spatial occupancy detection. The system has pre-assigned unique device component IDs to each component of the escalator in the three-dimensional digital model, such as the step ID is 'Escalator_Step_001', and the handrail ID is 'Escalator_Handrail_001'. Through the spatial occupancy detection algorithm, the system can identify whether the motion area overlaps with these escalator device components with unique IDs in terms of spatial position. It realizes the precise positioning of the moving target in the three-dimensional space and establishes the association between the moving target and the specific device component, laying a foundation for the subsequent abnormal behavior analysis of specific device components. Compared with the traditional method of only performing abnormal detection in the two-dimensional image space, the present invention can more accurately locate the abnormal event to a specific device component, thereby improving the pertinence and effectiveness of emergency treatment.

[0244] For the escalator step equipment component (ID: 'Escalator_Step_001') identified as having a spatial overlap with the movement area, the system will further analyze its movement state within a preset time window. In this embodiment, the preset time window w is set to 5 seconds. Record the number of times the movement area has a spatial overlap with this step equipment component within these 5 seconds, and use this as the movement frequency of this step equipment component. To more comprehensively analyze the characteristics of the step movement, the system further uses Fourier transform to perform spectral analysis on the movement data of this step equipment component within the time window w. Fourier transform can decompose the movement signal in the time domain into spectra of different frequency components and extract the main frequency and its amplitude of the movement. For example, the movement of a normally operating escalator step should exhibit a regular low-frequency movement spectrum, and its main frequency corresponds to the normal operating speed of the escalator. When the escalator step experiences abnormal jitter or jamming, etc., its movement spectrum may show high-frequency components or the phenomenon that the amplitude of the main frequency increases abnormally. By calculating the movement frequency and performing spectral analysis, the system can accurately describe the movement state of the equipment component from two dimensions of time domain and frequency domain. The application of Fourier transform can effectively extract the characteristic information of the movement signal, enabling the system to identify subtle abnormal movements that are difficult to detect by the naked eye, such as slight jitter or irregular movement of the step, improving the sensitivity and accuracy of anomaly detection.

[0245] To achieve refined anomaly determination for different equipment components, an equipment type anomaly state threshold table is established in advance. For the escalator step equipment component (ID: 'Escalator_Step_001'), corresponding anomaly state thresholds are preset in this threshold table. For example, for escalator steps, the preset anomaly state thresholds can include:

[0246] Frequency threshold: The normal operating frequency range is 0.1 Hz - 0.5 Hz. When the detected step movement frequency exceeds this range, for example, exceeds 0.6 Hz (indicating that the step running speed is abnormally fast) or is lower than 0.05 Hz (indicating that the step running speed is abnormally slow or stagnant), it is determined as a frequency anomaly.

[0247] Amplitude threshold: Under normal operating conditions, the step movement amplitude should be maintained within the range of ±2 times the standard deviation. When the detected amplitude of the main frequency exceeds 3 times the standard deviation, it is determined as an amplitude anomaly, which may indicate that the step has abnormal severe vibration.

[0248] Compare the calculated escalator step movement frequency and amplitude with the preset anomaly state thresholds. When it is detected that the movement frequency exceeds 0.6 Hz or is lower than 0.05 Hz, or the movement amplitude exceeds 3 times the standard deviation, the system will determine that the escalator step equipment component is abnormal and trigger an alarm for the corresponding component.

[0249] The application of the preset abnormal state threshold table enables the system to set targeted abnormal detection criteria according to the operating characteristics of different equipment components, avoiding the rough method of using a unified threshold for abnormal detection, and significantly improving the accuracy and reliability of abnormal detection. By setting multi-dimensional thresholds such as frequency and amplitude, various abnormal states that may occur in the equipment can be more comprehensively covered, reducing the probability of missed reports and false alarms.

[0250] When the movement frequency or amplitude of the escalator step equipment component (ID: 'Escalator_Step_001') exceeds the preset threshold, the system will immediately trigger an alarm signal for this component. The alarm signal will include the ID information of the equipment component, the type of abnormality (e.g., frequency abnormality, amplitude abnormality), the level of abnormality, etc., and be transmitted to the emergency response terminal and the central server at the execution layer. The triggering of the alarm signal is the direct output of the abnormal behavior detection result, providing an important basis for abnormal information for subsequent emergency response and execution control steps. Through precise equipment component positioning and detailed abnormal information description, it can provide timely and accurate decision-making support for emergency handlers, enhancing the intelligent emergency handling ability of the urban rail transit system.

[0251] Preferably, the multi-modal fusion analysis combining sound data and equipment operation data to determine the severity of abnormal events includes:

[0252] Performing weighted fusion based on visual abnormality scores, sound abnormality scores, and equipment operation data abnormality scores. The comprehensive abnormality score calculation formula is:

[0253]

[0254] where , , are the weight coefficients of visual data , sound data and equipment operation data respectively, and ;

[0255] When the comprehensive abnormality score exceeds the threshold, the abnormal event is confirmed.

[0256] In an embodiment provided by the present application, it is assumed to be applied at the escalator of a certain urban rail transit platform. When the escalator is running, the system continuously collects visual, sound, and equipment operation data.

[0257] Through visual data analysis, abnormal movements are detected on the escalator steps, such as passengers falling or items dropping and getting stuck. After motion frequency analysis, it is determined that the motion frequency of the escalator steps exceeds the threshold of 5 Hz, and the visual abnormality score V (escalator step ID) = 0.6.

[0258] The microphone array captures abnormal metallic friction sounds during the operation of the escalator, with the frequency and energy exceeding the normal range. The abnormal sound score A (escalator step ID) = 0.5.

[0259] Through the industrial protocol interface, the system captures that the current value of the escalator motor has increased abnormally, and the temperature has risen slightly but has not exceeded the severe threshold. The abnormal score of equipment operation data E (escalator step ID) = 0.2.

[0260] Substitute the above scores into the comprehensive abnormal score formula: C (escalator step ID) = 0.5 * V (escalator step ID) + 0.3 * A (escalator step ID) + 0.2 * E (escalator step ID)

[0261] C (escalator step ID) = 0.5 * 0.6 + 0.3 * 0.5 + 0.2 * 0.2 = 0.3 + 0.15 + 0.04 = 0.49

[0262] In this example, the calculated comprehensive abnormal score C (escalator step ID) = 0.49, which is lower than the threshold Tc = 0.7. Therefore, although both visual and sound data indicate possible abnormalities, after comprehensively considering the equipment operation data, the system determines it as "slight abnormality" or "potential risk", which may trigger a first-level response (equipment self-check or slight deceleration), but will not immediately stop the escalator operation to avoid inconvenience to passengers caused by misjudgment.

[0263] Assume another situation: If the visual abnormal score V (escalator step ID) = 0.8 (for example, visual detection shows that a passenger is stuck), the abnormal sound score A (escalator step ID) = 0.7 (the friction sound is more intense), and the abnormal score of equipment operation data E (escalator step ID) = 0.6 (both the motor current and temperature have increased significantly).

[0264] C (escalator step ID) = 0.5 * 0.8 + 0.3 * 0.7 + 0.2 * 0.6 = 0.4 + 0.21 + 0.12 = 0.73

[0265] At this time, the comprehensive abnormal score C (escalator step ID) = 0.73, which exceeds the threshold Tc = 0.7. The system will confirm that an "abnormal event" has occurred and trigger corresponding emergency response strategies according to the abnormal level (for example, it may be judged as a second-level or third-level abnormality in this example), such as immediately stopping the escalator operation, activating the alarm, notifying the control center and inspection personnel, etc.

[0266] Single visual detection may be interfered by factors such as illumination changes, shadows, occlusions, etc., resulting in false alarms. By fusing sound and device operation data, they can corroborate each other and eliminate accidental visual interferences. For example, even if abnormal movement is detected visually, if both the sound and device operation data are normal, the possibility of false alarms can be reduced. The multi-modal anomaly detection mechanism can reduce the false alarm rate to below 1%. Sound data and device operation data can provide deeper anomaly information. For example, mechanical failures, electrical anomalies, etc. inside the device may be difficult to directly observe visually but can be reflected by abnormal sounds and abnormal device operation parameters. Multi-modal fusion can comprehensively utilize the advantages of various data sources to improve the accuracy of anomaly detection. The accuracy can be increased to over 95%. In a complex and changeable environment, a single sensor may fail or its performance may decline. Multi-modal fusion can utilize the redundant information of multiple sensors to improve the robustness of the system. Even if the visual sensor is occluded or fails, the system can still rely on sound and device operation data for anomaly detection to ensure the reliable operation of the system. By setting refined anomaly thresholds for different types of devices and components and conducting comprehensive analysis in combination with multi-modal data, more refined anomaly determination can be achieved. For example, it is possible to distinguish between minor anomalies, obvious anomalies, and severe anomalies of the device and adopt corresponding emergency response strategies according to different anomaly levels to improve the efficiency and accuracy of emergency handling.

[0267] Preferably, the formulating of emergency response strategies according to the severity level of anomaly events includes:

[0268] Dividing anomaly events into three levels of response:

[0269] Level 1: Minor anomaly, triggering the device self-check program, without affecting the normal operation of the device;

[0270] Level 2: Obvious anomaly, triggering device deceleration or current limiting measures;

[0271] Level 3: Severe anomaly, immediately stopping the device operation and starting the emergency plan.

[0272] In an embodiment provided by the present application, after completing the multi-modal fusion analysis, the system will obtain a comprehensive anomaly score , which reflects the anomaly degree of the device component id. According to 's size, the system divides anomaly events into three levels:

[0273] Level 1 (Minor anomaly): When 0 < ≤ 0.3, it is determined as a minor anomaly. For example, there are a small number of passengers staying on the escalator steps for a short time, and the card swiping fails occasionally at the turnstile, etc.

[0274] Level 2 (Obvious anomaly): When 0.3 < When it is ≤ 0.7, it is determined as an obvious anomaly. For example, the running speed of the escalator fluctuates slightly, the turnstile fails to swipe the card continuously, the opening and closing speed of the platform screen door is abnormal, etc.

[0275] Level 3 (severe anomaly): When > 0.7, it is determined as a severe anomaly. For example, the escalator suddenly stops running, the turnstile cannot be opened or closed, the platform screen door cannot be closed normally, etc.

[0276] For different levels of anomaly events, the system has pre-developed corresponding emergency response strategies and stored them in the emergency plan database. These strategies include:

[0277] Level 1 response: Trigger the device self-check program. For example, the escalator control system automatically performs a self-check, checks the working status of each sensor and motor, and records the self-check results. At the same time, the system sends a prompt message to the control center, informing of a minor anomaly.

[0278] Level 2 response: Trigger the device deceleration or flow-limiting measures. For example, when it is detected that the running speed of the escalator fluctuates, the control system automatically reduces the running speed of the escalator and limits the number of passengers passing through the escalator per unit time. At the same time, the system sends an alarm message to the control center and the inspection personnel, indicating an obvious anomaly and requiring manual intervention for investigation.

[0279] Level 3 response: Immediately stop the device operation and start the emergency plan. For example, when it is detected that the escalator suddenly stops running, the control system immediately cuts off the power supply of the escalator to prevent accidents. At the same time, the system sends an emergency alarm message to the control center, the inspection personnel and the passenger terminal, and starts the preset emergency plan, including evacuating passengers and arranging alternative means of transportation.

[0280] The system sends control instructions to the rail transit equipment through industrial protocol interfaces (such as Modbus, Profinet, etc.) to achieve real-time intervention in the equipment. The specific content of the control instructions depends on the type and level of the anomaly event. For example:

[0281] For the escalator, instructions such as start, stop, and speed adjustment can be sent.

[0282] For the turnstile, instructions such as open, close, and lock can be sent.

[0283] For the platform screen door, instructions such as open, close, and emergency unlock can be sent.

[0284] The system pushes the anomaly event information to the control center, the inspection personnel and the passenger terminal so that the relevant personnel can understand the situation in time and take corresponding measures. The ways of information push include:

[0285] Control Center: Displays the detailed information of abnormal events through a monitoring large screen, including the occurrence time, location, equipment type, abnormal level, handling status, etc.

[0286] Inspection Personnel: Receive alarm information through a mobile terminal (such as a mobile phone App) and view on-site video surveillance to quickly rush to the scene for handling.

[0287] Passenger Terminal: Issues emergency notifications through platform display screens, carriage display screens, and mobile phone Apps to inform passengers of the occurrence of abnormal events and guide passengers to evacuate safely.

[0288] Preferably, as Figure 2 shown, an urban rail transit emergency handling system based on a large vision model includes:

[0289] Data Acquisition Module: Used to collect visual data, sound data, and equipment operation data in key areas of the platform and carriage of urban rail transit, perform preprocessing, and achieve timestamp synchronization at the same time;

[0290] Visual Analysis Module: Used to perform dynamic frame difference processing on consecutive video frames of visual data, generate a two-dimensional motion heat map, and perform spatio-temporal aggregation on the two-dimensional motion heat map;

[0291] Spatial Mapping Module: Used to build a three-dimensional digital model based on the engineering drawings of the platform and carriage of urban rail transit, perform camera internal and external parameter calibration, and establish a mapping relationship between the two-dimensional motion heat map and the three-dimensional digital model;

[0292] Abnormal Detection Module: Used to detect abnormal behaviors in the mapping results of the spatio-temporally aggregated two-dimensional motion heat map and the three-dimensional digital model, and perform multi-modal fusion analysis in combination with sound data and equipment operation data to determine the severity of abnormal events;

[0293] Emergency Response Module: Used to formulate emergency response strategies according to the severity level of abnormal events, send control instructions to rail transit equipment through an industrial protocol interface, and push abnormal event information to the control center, inspection personnel, and passenger terminals.

[0294] In an embodiment provided by the present application, the application of the urban rail transit emergency handling system based on the large vision model in the escalator entrance area of the urban rail transit platform is described, which is used to monitor and handle abnormal events at the escalator entrance to ensure the safety of passengers and the orderly operation of rail transit.

[0295] The urban rail transit emergency handling system based on the large vision model mainly consists of the following five modules: data acquisition module, visual analysis module, spatial mapping module, abnormal detection module, and emergency response module.

[0296] The data acquisition module is deployed in the area at the entrance of the platform escalator. Its main function is to collect visual data, sound data, and equipment operation data in this area, and perform preprocessing and timestamp synchronization.

[0297] On the columns at the escalator entrance, three groups of low-resolution wide-angle cameras are installed in a triangular structure. The model is Hikvision DS-2CD3145-I, with a resolution set to 640×480 pixels and a lens viewing angle of 130°. The three groups of cameras monitor the escalator entrance area from different angles for all-round monitoring, ensuring redundant coverage of the monitoring area and eliminating blind spots that may be caused by a single viewing angle. The image acquisition frequency of the cameras is set to 10 frames per second. This frequency can not only meet the real-time detection requirements of moving targets but also effectively reduce the data processing pressure on the system backend and the network transmission bandwidth requirements.

[0298] A microphone array, model Knowles SPH0641LM4H-1, is integrally installed near the cameras. The microphone array is used to collect environmental sound data in the escalator entrance area, such as passengers' calls for help, screams, and abnormal equipment noises. The sound data acquisition frequency is set to 16 kHz to meet the effective capture of sound events. The sound data acquisition unit has a hardware noise reduction function to preliminarily filter out environmental background noise and improve the signal-to-noise ratio of effective sound signals.

[0299] Through the industrial Ethernet interface of the escalator control system and using the Modbus TCP / IP protocol, the operation data of the escalator is collected in real time, including the running speed of the escalator, the current of the motor, the temperature of key components, and vibration sensor data. These data reflect the real-time operation status of the escalator and provide operation parameters at the equipment level for multimodal anomaly detection.

[0300] First, Gaussian blur preprocessing is performed on the collected visual data. A 5×5 Gaussian kernel is used, and the standard deviation σ is set to 1.5 to effectively smooth the image and reduce the interference of image noise on the subsequent frame difference algorithm. For sound data, a noise reduction algorithm based on spectral subtraction is used to further suppress environmental noise. The equipment operation data is normalized to uniformly scale data with different dimensions to the [0, 1] interval, eliminating the influence of dimension differences on subsequent fusion analysis. After preprocessing, the Network Time Protocol (NTP) server is used to uniformly add accurate timestamps to all collected data streams, with a time synchronization accuracy reaching the millisecond level to ensure the accuracy of subsequent multimodal data fusion analysis.

[0301] The data acquisition module works in collaboration with multiple sensors to comprehensively obtain visual, auditory, and equipment operating status information in the escalator entrance area, and performs effective preprocessing and precise time synchronization, laying a high-quality data foundation for subsequent intelligent analysis. The adoption of low-resolution cameras significantly reduces the pressure of data processing and transmission while ensuring the monitoring effect, meeting the application requirements of edge computing.

[0302] The visual analysis module receives the preprocessed visual data from the data acquisition module. Its core function is to perform dynamic frame difference processing on consecutive video frames to generate a two-dimensional motion heat map, and conduct spatio-temporal aggregation on the motion heat map to extract the moving target area.

[0303] Receive the preprocessed video frame sequence. For two consecutive frames of images, calculate the gray value difference of corresponding pixel points to obtain the frame difference image. To adapt to the influence of light changes and environmental noise, an environment adaptive algorithm is used to dynamically adjust the frame difference threshold. Specifically, the mean and standard deviation of the pixel gray values of the current frame difference image are statistically calculated, and the threshold is set to the mean plus 1.5 times the standard deviation. This dynamic threshold adjustment mechanism can effectively cope with sudden light changes and noise interference, improving the robustness of moving target detection.

[0304] Perform binary processing on the frame difference image. Points with pixel gray values greater than the dynamic threshold are set to 255 (white), and points less than the threshold are set to 0 (black) to generate a binary motion heat map. In the motion heat map, the white area represents the area where motion changes occur in the image.

[0305] To eliminate the influence of isolated noise points and enhance the continuity and integrity of the moving area, perform temporal aggregation on the binary motion heat maps of 5 consecutive frames, that is, perform pixel-level "OR" operation on the 5 consecutive heat maps to obtain the aggregated heat map. Then, apply morphological closing operation to the aggregated heat map, using a 3×3 square structuring element for closing operation to fill the holes inside the moving area, smooth the edges of the moving area, and further remove noise interference.

[0306] Adopt the eight-neighborhood connected region labeling algorithm to perform connected region analysis on the spatio-temporally aggregated binary motion heat map. Label the connected white pixel regions in the heat map as independent moving regions, assign a unique identifier (ID) to each independent moving region, and calculate geometric features such as the area and centroid coordinates of each moving region.

[0307] The visual analysis module adopts dynamic frame difference technology and spatio-temporal aggregation strategy, effectively detecting and extracting moving target regions from the video stream, and significantly reducing the computational complexity. Compared with the object detection method based on deep learning, the computational amount of the frame difference algorithm is greatly reduced, the processing speed is faster, and it can meet the real-time requirements. Spatio-temporal aggregation and morphological operations further improve the accuracy and robustness of motion detection, providing reliable moving target information for subsequent spatial mapping and abnormal behavior analysis.

[0308] The core function of the spatial mapping module is to construct a three-dimensional digital model of the escalator entrance area of the platform, perform camera calibration, establish the spatial mapping relationship between the two-dimensional motion heat map and the three-dimensional model, and accurately project the motion area in the two-dimensional image into the three-dimensional space.

[0309] Based on the CAD engineering drawings of the escalator entrance area of the rail transit platform, use three-dimensional modeling software (such as AutoCAD) to construct an accurate three-dimensional digital model. The geometric shapes, spatial positions, and dimensional parameters of all fixed facilities such as escalators, guardrails, and ground marking lines are refined and restored in the model. Assign a unique ID and function label to each device component in the model (such as each step of the escalator, handrail belt, column at the escalator entrance, etc.), for example, "escalator step_1", "handrail belt_left", "escalator entrance column_A", etc., to construct a device component database for subsequent refined anomaly analysis.

[0310] On-site at the escalator entrance area, use the Zhang Zhengyou calibration method to calibrate the internal and external parameters of each camera using a checkerboard calibration board. The calibration process obtains the internal parameter matrix K of the camera (including focal length, principal point coordinates, distortion parameters) and the external parameter matrix (including rotation matrix R and translation vector t). The internal parameter matrix describes the internal imaging parameters of the camera, and the external parameter matrix describes the position and attitude of the camera in the world coordinate system. The calibration accuracy is controlled at the sub-pixel level to ensure the accuracy of subsequent spatial mapping.

[0311] Based on the camera calibration results, establish the conversion relationship between the camera view and the world coordinate system of the three-dimensional CAD model. Using the perspective projection model and affine transformation, map the pixel coordinates in the two-dimensional motion heat map to the three-dimensional model space. For each motion area marked and output in the connected area, convert the pixel points on its boundary contour to three-dimensional space coordinates to obtain the position, shape, and size information of the motion area in the three-dimensional space.

[0312] Judge whether the motion area overlaps with the device components defined in the device component database in the three-dimensional space. Through spatial geometric relationship calculation, judge whether the motion area enters the three-dimensional bounding box of a certain device component. If an overlap occurs, record the ID of the motion area and the corresponding device component for subsequent abnormal behavior analysis.

[0313] The spatial mapping module realizes the accurate conversion of two-dimensional image information into three-dimensional spatial information through precise three-dimensional modeling and camera calibration, with a positioning accuracy of up to ±5 cm. This module does not rely on deep learning training data, reducing the system's dependence on labeled data and lowering the deployment and maintenance costs. Through spatial occupancy detection, it can accurately identify which device components the moving targets interact with, laying a foundation for subsequent device component-level abnormal behavior analysis.

[0314] The anomaly detection module receives the spatial mapping results of the moving areas from the spatial mapping module, as well as the sound data and device operation data from the data acquisition module, and performs multi-modal fusion analysis to detect abnormal events in the escalator entrance area.

[0315] For the identified moving areas that have spatial overlap with device components, such as the moving area that overlaps with the "escalator step_1" component. For each device component ID, calculate its movement frequency within a preset time window (e.g., 1 second). The calculation method of the movement frequency is: count the number of moving areas that have spatial overlap with this device component within the time window. Further, perform a fast Fourier transform on the movement frequency sequence of the device component, analyze its movement spectral characteristics, and extract the main frequency and its amplitude of the movement.

[0316] Define an abnormal state threshold table in advance according to the device type and operation characteristics of the escalator . For example, for the "escalator step" component, define the abnormal state threshold as: movement frequency > 5 Hz or movement amplitude > 3 times the standard deviation. For the "handrail belt" component, define the abnormal state threshold as: frequency > 2 Hz or the movement range exceeds the normal working range by 20%. For the "escalator entrance column" component, define the abnormal state threshold as: frequency > 1 Hz or movement detection during abnormal interaction periods (e.g., passengers climbing the column).

[0317] Comprehensive visual anomaly score 、Sound anomaly score And device operation data anomaly score Perform multi-modal fusion and calculate the comprehensive anomaly score . Among them, the visual anomaly score Is determined by the results of device component movement frequency analysis and abnormal states. When it is detected that the movement frequency or amplitude exceeds the threshold, Set it to 1, otherwise 0. The sound anomaly score Is obtained from the analysis of the collected sound data. For example, when the screams of passengers or abnormal noises of the device are detected, Set it to 1, otherwise 0. The device operation data anomaly score Obtained from the analysis of the collected escalator operation data. For example, when the escalator speed is detected to be abnormal, the current is overloaded, or the vibration exceeds the standard, Set it to 1, otherwise set it to 0. The comprehensive anomaly score calculation formula is:

[0318]

[0319] Where the weight coefficients 、 、 Are respectively set to 0.5, 0.3, 0.2, and . The comprehensive threshold Tc is set to 0.7. When the comprehensive anomaly score > Tc, it is confirmed that an abnormal event has occurred, and the abnormal event is classified according to the type and severity of the abnormal event.

[0320] The anomaly detection module, based on the motion frequency analysis and multi-modal fusion mechanism at the device component level, can accurately and reliably detect abnormal events in the escalator entrance area, such as passengers falling, going against the flow, climbing the escalator, equipment failures, etc. The multi-modal fusion effectively reduces the false alarm rate and improves the detection accuracy. The refined device component-level analysis can more accurately locate the specific location and device components where the anomaly occurs, providing a more accurate decision-making basis for subsequent emergency responses. In this embodiment, the multi-modal anomaly detection mechanism reduces the false alarm rate to less than 1% and increases the accuracy to more than 95%.

[0321] The emergency response module receives the abnormal event alarm information from the anomaly detection module, formulates an emergency response strategy according to the severity level of the abnormal event, sends control instructions to the rail transit equipment through the industrial protocol interface, and pushes the abnormal event information to the control center, inspection personnel, and passenger terminals.

[0322] According to the comprehensive anomaly score Output by the anomaly detection module and the type of abnormal event, the abnormal event is divided into three-level responses:

[0323] Level 1 (minor anomaly): The comprehensive anomaly score Is between 0.7 and 0.8. For example, it is detected that a passenger stays briefly at the escalator entrance, but no obvious safety risk is caused. Trigger the device self-check program, the system background records the anomaly log, which does not affect the normal operation of the escalator, and only pushes the level 1 alarm information to the control center.

[0324] Level 2 (obvious anomaly): The comprehensive anomaly score Between 0.8 and 0.9. For example, it is detected that passengers are running or playing on the escalator, or there are slight fluctuations in the running speed of the escalator. Trigger device deceleration or flow-limiting measures, such as reducing the running speed of the escalator to 50% of the normal speed, and giving safety prompts through platform announcements and passenger information displays. At the same time, push secondary alarm information to the control center and the mobile terminals of nearby patrol personnel to prompt the patrol personnel to go to the scene for inspection.

[0325] Level 3 (serious anomaly): Comprehensive anomaly score > 0.9. For example, it is detected that passengers have fallen, are going against the flow, or climbing the escalator, or serious equipment failures such as the escalator suddenly stopping, running in reverse, or parts falling off. Immediately send an emergency stop command through the industrial protocol interface of the escalator control system to stop the escalator operation, and start the preset emergency plan, such as linking the platform emergency stop button, starting the platform emergency broadcast, pushing level 3 alarm information to the control center, patrol personnel, and passenger terminals, and automatically calling the platform duty personnel and medical emergency personnel at the same time.

[0326] According to the emergency response strategy, through the industrial Ethernet interface of the escalator control system, using the Modbus TCP / IP protocol, send corresponding control commands to the escalator control system, such as self-check commands, deceleration commands, flow-limiting commands, stop commands, etc., to achieve real-time intervention and control of the escalator.

[0327] Push the detailed information of the abnormal event (including the type of anomaly, occurrence time, location, severity, relevant equipment part ID, on-site video images, etc.) to the urban rail transit control center, the mobile terminals of platform patrol personnel (such as the patrol APP), the platform passenger information display, and passenger terminals such as the passenger mobile APP through the network communication module. The information push protocol uses the MQTT or WebSocket protocol to ensure the real-time and reliability of the information.

[0328] The emergency response module can formulate and execute corresponding emergency response strategies at different levels according to the severity of the abnormal event, achieve real-time control and intervention of rail transit equipment, and promptly push the abnormal information to relevant personnel and passengers, forming a fast and efficient emergency linkage mechanism, maximizing the safety of passengers, reducing accident risks, and improving the intelligent and safety management level of urban rail transit.

[0329] This embodiment details the application of the urban rail transit emergency handling system based on the vision large model in the escalator entrance area of the platform. Through the collaborative work of the data acquisition module, vision analysis module, spatial mapping module, anomaly detection module, and emergency response module, this system realizes real-time, accurate, and intelligent monitoring, analysis, and handling of abnormal events in the escalator entrance area. This system effectively utilizes multi-modal information such as vision, hearing, and equipment operation, adopts an edge computing architecture, reduces the computational complexity and data transmission pressure, improves the real-time performance and reliability of the system operation, and has significant technical advantages and application value. The system described in this embodiment is not only applicable to the escalator entrance area of the platform but can also be extended to other key areas of urban rail transit such as the waiting area of the platform, gate area, and inside the carriage, providing strong support for building a safer and more intelligent urban rail transit system.

[0330] The above is the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.

Claims

1. An urban rail transit emergency handling method based on a large vision model, characterized in that, Including: Collecting visual data, sound data, and equipment operation data in key areas of the platform and carriage of urban rail transit, preprocessing them, and achieving timestamp synchronization simultaneously; Performing dynamic frame difference processing on consecutive video frames of visual data to generate a two-dimensional motion heatmap, and performing spatio-temporal aggregation on the two-dimensional motion heatmap, including: Calculating the pixel differences between adjacent frames for consecutive video frames to obtain a difference image, statistically analyzing the pixel value distribution of the difference image, and dynamically adjusting the threshold based on the calculated pixel value mean, pixel value standard deviation, and the noise level of the current scene; Performing binaryzation processing on the difference image according to the dynamically adjusted threshold to generate a two-dimensional motion heatmap, where the white area in the two-dimensional motion heatmap represents the detected motion area, and the black area represents the static background area; performing a pixel-by-pixel logical OR operation on the two-dimensional motion heatmaps of five consecutive frames to achieve aggregation in the time dimension; Performing spatio-temporal aggregation on the aggregated binary motion heatmap and applying morphological closing operation to remove interference points; Identifying and labeling independent motion areas through a connected component labeling algorithm and assigning a unique identifier to each motion area; Constructing a three-dimensional digital model based on the engineering drawings of the platform and carriage of urban rail transit, performing camera internal and external parameter calibration, and establishing a mapping relationship between the two-dimensional motion heatmap and the three-dimensional digital model, including: Constructing an accurate three-dimensional digital model based on the CAD engineering drawings of the platform and carriage, where the three-dimensional digital model includes the spatial positions and geometric parameters of escalators, turnstiles, and seat fixtures; Assigning a unique ID and a functional label to each equipment component of the fixed facilities and constructing an equipment component database; Performing camera internal and external parameter calibration using a calibration board to obtain the internal parameter matrix and external parameter matrix of the camera; Based on the internal and external parameter calibration results, applying an affine transformation to map the motion areas in the two-dimensional motion heatmap to the spatial coordinate system of the three-dimensional digital model; Identifying the motion areas that overlap with the equipment components with unique IDs in the three-dimensional model through a spatial occupancy detection algorithm; Performing abnormal behavior detection on the mapping results of the spatio-temporally aggregated two-dimensional motion heatmap and the three-dimensional digital model, and performing multi-modal fusion analysis in combination with sound data and equipment operation data to determine the severity of the abnormal event; The abnormal behavior detection includes: calculating the motion frequency of the equipment component within a preset time window for the motion areas that overlap with the equipment component with a unique ID, where the motion frequency is the number or area change rate of the motion areas that overlap with the spatial position of the equipment component, and analyzing the spectral characteristics of the motion through Fourier transform to extract the main motion frequency and its amplitude; based on a preset abnormal state threshold table for the equipment type, triggering an alarm for the corresponding component when the detected main motion frequency or amplitude exceeds the threshold; Formulating an emergency response strategy according to the severity level of the abnormal event, sending control instructions to the rail transit equipment through an industrial protocol interface, and pushing the abnormal event information to the control center, inspection personnel, and passenger terminals.

2. The urban rail transit emergency handling method based on a large vision model according to claim 1, characterized in that The collection of visual data, sound data, and equipment operation data includes: Deploy wide - angle cameras at the escalator entrances, security check areas, waiting areas of the platform and both ends of the carriages, and set the acquisition frequency to 10 frames per second; Use a microphone array to collect environmental sound data and extract sound features, where the sound features include frequency, energy, and spectral changes; Real - time collect equipment operation data through the industrial protocol interface of rail transit equipment, and the equipment operation data includes speed, current, temperature, and vibration.

3. The urban rail transit emergency handling method based on a large vision model according to claim 2, characterized in that, Pre - process the visual data, sound data, and equipment operation data, and at the same time achieve timestamp synchronization, including: Perform Gaussian blur processing on the visual data; Perform noise reduction processing on the sound data; Perform normalization processing on the equipment operation data; Use the Network Time Protocol or Precision Time Protocol to synchronize timestamps for the pre - processed visual data, sound data, and equipment operation data to ensure the accuracy of data fusion.

4. A method for urban rail transit emergency handling based on a large vision model according to claim 3, characterized in that, Perform multi - modal fusion analysis by combining sound data and equipment operation data to determine the severity of abnormal events, including: Weighted fusion is performed based on the visual anomaly score, the sound anomaly score, and the device operation data anomaly score to obtain a comprehensive anomaly score The calculation formula is as follows: ; Among them , , are the visual data , the sound data and the device operation data respectively, and ; When the comprehensive anomaly score exceeds the threshold, confirm the abnormal event.

5. The urban rail transit emergency handling method based on a large vision model according to claim 4, wherein, Formulate emergency response strategies according to the severity level of abnormal events, including: Classify abnormal events into three - level responses: Level 1: Slight anomaly, trigger the equipment self - inspection program, which does not affect the normal operation of the equipment; Level 2: Obvious anomaly, trigger equipment deceleration or current - limiting measures; Level 3: Severe anomaly, immediately stop the equipment operation and start the emergency plan.

6. An urban rail transit emergency handling system based on a large vision model, characterized in that, Including: A data acquisition module, which is used to collect visual data, sound data, and equipment operation data in key areas of the platform and carriages of urban rail transit, perform pre - processing, and at the same time achieve timestamp synchronization; A visual analysis module, which is used to perform dynamic frame difference processing on consecutive video frames of visual data to generate a two - dimensional motion heat map, and perform spatio - temporal aggregation on the two - dimensional motion heat map, including: Calculate the pixel difference between adjacent frames of consecutive video frames to obtain a difference image, statistically analyze the pixel value distribution of the difference image, and dynamically adjust the threshold through the calculated pixel value mean, pixel value standard deviation, and the noise level of the current scene; Perform binary processing on the difference image according to the dynamically adjusted threshold to generate a two - dimensional motion heat map, where the white area in the two - dimensional motion heat map represents the detected motion area, and the black area represents the static background area; perform a pixel - by - pixel logical OR operation on the two - dimensional motion heat maps of five consecutive frames to achieve aggregation in the time dimension; Perform spatio - temporal aggregation on the aggregated binary motion heat map, and apply morphological closing operations to remove interference points; Identify and label independent motion areas through the connected - component labeling algorithm, and assign a unique identifier to each motion area; A space mapping module, which is used to construct a three - dimensional digital model based on the engineering drawings of the platform and carriages of urban rail transit, perform camera internal and external parameter calibration, and establish a mapping relationship between the two - dimensional motion heat map and the three - dimensional digital model, including: Construct an accurate three - dimensional digital model based on the CAD engineering drawings of the platform and carriages, and the three - dimensional digital model includes the spatial positions and geometric parameters of escalators, turnstiles, and seat fixtures; Assign a unique ID and function label to each equipment component of the fixed facilities, and construct an equipment component database; Execute the internal and external parameter calibration of the camera using a calibration board to obtain the internal parameter matrix and external parameter matrix of the camera; Based on the results of the internal and external parameter calibration, apply an affine transformation to map the motion area in the two-dimensional motion heat map to the spatial coordinate system of the three-dimensional digital model; Through a spatial occupancy detection algorithm, identify the motion areas that overlap with the device components with unique IDs in the three-dimensional model in terms of spatial position; Anomaly detection module, which is used to detect abnormal behaviors in the mapping results of the two-dimensional motion heat map after spatio-temporal aggregation and the three-dimensional digital model, and perform multi-modal fusion analysis in combination with sound data and device operation data to determine the severity of abnormal events; The abnormal behavior detection includes: for the motion areas that overlap with the device components with unique IDs in terms of space, calculate the motion frequency of the device components within a preset time window, where the motion frequency is the number or area change rate of the motion areas with overlapping spatial positions of the device components, and analyze the spectral characteristics of the motion through Fourier transform to extract the main motion frequency and its amplitude; based on a preset anomaly status threshold table for device types, trigger an alarm for the corresponding component when it is detected that the main motion frequency or amplitude exceeds the threshold; Emergency response module, which is used to formulate emergency response strategies according to the severity level of abnormal events, send control instructions to rail transit equipment through an industrial protocol interface, and push the abnormal event information to the control center, inspection personnel, and passenger terminals.

Citation Information

Patent Citations

  • Hump shunting band-type brake abnormity monitoring system based on multi-source data fusion algorithm

    CN116776202A

  • Intelligent operation and maintenance emergency processing system

    CN117215940A

Cited By

  • Rail transit station vertical domain intelligent agent inspection system and method

    CN122244779A