Railway passenger transport operation safety risk hidden danger troubleshooting method and system

By synchronously collecting and multimodally analyzing railway passenger operation data, the blind spot problem of hidden danger identification in single-modal monitoring in existing technologies has been solved, and real-time and accurate inspection of railway passenger operation safety risks has been achieved, thereby improving emergency response efficiency.

CN120808794APending Publication Date: 2025-10-17INST OF COMPUTING TECH CHINA ACAD OF RAILWAY SCI +2
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510739465.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing railway passenger operation safety risk detection technology relies on single-modal monitoring, which makes it difficult to fully and accurately grasp the operation site conditions. It lacks the dynamic alignment and causal reasoning capabilities of multi-modal data, resulting in blind spots in hazard identification and information delays, affecting emergency response efficiency.

Method used

It simultaneously collects operational voices from multiple communication channels and broadcast channels, operational status data of station guidance screens, and data from various environmental sensors. Through voiceprint recognition, speech-to-text conversion, and multimodal risk knowledge graphs, it performs logical reasoning to achieve the fusion and synchronous investigation of multi-source data.

Benefits of technology

It has achieved real-time and accurate detection of safety risks in railway passenger operations, improved the level of safety management, increased the error detection rate of displayed content to 98%, the response time for high-risk events to ≤1 second, and the accuracy rate of multimodal conflict handling to ≥95%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808794A_ABST
    Figure CN120808794A_ABST
Patent Text Reader

Abstract

The invention provides a railway passenger transport operation safety risk hidden danger troubleshooting method and system, and the method comprises the steps: synchronously collecting various operation data, including communication and broadcast voice, station guide screen state, fault log and operation environment sensor data; and determining a homework voice main body by using the voiceprint recognition model, and converting the homework voice main body into a homework text after noise Meanwhile, key entity information is extracted from all kinds of data and is aligned with the homework text. And performing semantic extraction and logical reasoning comparison on the aligned text and information in combination with a multi-modal risk knowledge graph, outputting a real-time troubleshooting result according to a preset risk assessment mechanism, and feeding back the real-time troubleshooting result to an operation text main body, so as to realize troubleshooting of the railway passenger operation safety risk hidden danger. Real-time and accurate troubleshooting of the railway passenger transport operation safety risk hidden danger is realized, and the safety management level of the railway passenger transport operation is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of railway management, and in particular to a railway passenger operation safety risk hidden danger checking method and system. BACKGROUND

[0002] Currently, there are many deficiencies in the field of railway passenger operation safety risk hidden danger checking. Traditional checking methods mostly rely on single modal monitoring, such as pure voice monitoring or pure video observation. In this mode, there is a lack of effective association between personnel operation and display device anomalies, and it is difficult to fully and accurately grasp the real situation of the operation site. For example, through voice monitoring alone, it is not possible to timely detect whether the display content of the station large screen is consistent with the actual running state of the train; and relying solely on video monitoring, it is also difficult to accurately capture key information in voice instructions.

[0003] For monitoring the display content of the station large screen, the existing technology mostly relies on manual inspection, which not only consumes a lot of manpower and time, but also is prone to risk omissions due to human negligence. The periodicity and uncertainty of manual inspection make it difficult to timely discover and correct some transient display errors or anomalies, which may cause safety hazards such as passenger misdirection.

[0004] In addition, the existing technology lacks deep reasoning ability for the causal relationship between display content and voice instructions in the hidden danger determination process. When there is inconsistency or conflict in information, it is difficult to quickly and accurately determine the root cause of the problem, and it is not possible to provide a strong basis for subsequent risk response. For example, when the train arrival broadcast and the platform number displayed on the large screen are inconsistent, the existing system cannot automatically analyze whether the broadcast is incorrect or the large screen is faulty, and manual intervention is required for checking, which affects the efficiency of emergency handling. SUMMARY

[0005] In view of this, the embodiments of the present application provide a railway passenger operation safety risk hidden danger checking method and system to eliminate or improve one or more defects in the prior art, and solve the problem that the prior art cannot systematically check safety risk hidden dangers for railway passenger operation safety management due to the single form.

[0006] One aspect of the present application provides a railway passenger operation safety risk hidden danger checking method, which comprises the following steps:

[0007] synchronously collecting operation voice of a plurality of communication channels and broadcast channels, station guide screen operation state data, fault logs, and a plurality of preset operation environment sensor data; the operation environment sensor data is image or text data collected for a plurality of railway passenger management operation elements;

[0008] The operation voice of multiple communication channels and broadcast channels in the process of railway passenger operation is obtained, and the preset voiceprint recognition model is used to obtain the corresponding voiceprint features and find the subject of the operation voice in the preset operation personnel voiceprint library.

[0009] After the operation voice is denoised and input into a preset voice text conversion model, the corresponding operation text is obtained. The station guide screen operation state data, the fault log and the operation environment sensor data are extracted and corrected to obtain key entity information about the railway passenger management operation elements, and the operation text is aligned according to the production time.

[0010] The operation text and the key entity information of the same railway passenger management operation element in the same period are aligned, and the semantic extraction is performed based on the preset multi-modal risk knowledge graph. The logic inference comparison is performed based on the preset risk assessment mechanism, and the feedback and response are performed. The real-time railway passenger operation safety risk hidden danger investigation result is output, and is fed back to the subject of the operation text. The multi-modal risk knowledge graph is set according to the preset railway passenger operation rules.

[0011] In some embodiments, the operation voice of multiple communication channels and broadcast channels is synchronously collected, including:

[0012] The intercom operation voice is obtained from the multiple communication channels according to a first set sampling frequency, the broadcast operation voice is obtained from the multiple broadcast channels according to a second set sampling frequency, and the voice endpoint detection is divided into multiple effective instruction segments and the sampling time point is marked.

[0013] In some embodiments, the station guide screen operation state data is obtained by the OCR text extraction after the image sensor collects the image data of the station guide screen based on the preset image sensor; and the fault log is recorded in the JSON format.

[0014] In some embodiments, the operation environment sensor data is derived from a station environment monitoring sensor, a train arrival and running state sensing sensor, a safety warning sensor and a device health management sensor.

[0015] The station environment monitoring sensor includes a temperature and humidity sensor, an air quality sensor, a noise sensor, an illumination sensor and a passenger flow monitoring sensor deployed at multiple positions in the railway station.

[0016] The train arrival and running state sensing sensor includes a track deformation sensor, a train positioning sensor and a train speed sensor deployed at multiple positions on the track.

[0017] The safety warning sensor includes a fire sensor and a security sensor.

[0018] The device health management type sensor includes a circuit monitoring sensor, a mechanical vibration sensor, and a water leakage detection sensor.

[0019] In some embodiments, based on the preset voiceprint recognition model, the corresponding voiceprint feature is obtained, and the subject of the operation voice is searched in the preset operation personnel voiceprint library, comprising: using a pre-trained ResCNN-GRU voiceprint recognition model to extract the voiceprint feature, and comparing the similarity with the voiceprint template stored in the preset operation personnel voiceprint library, and outputting the object with the highest similarity as the subject;

[0020] The creation step of the preset operation personnel voiceprint library comprises:

[0021] Based on the preset link, the multi-segment voice of each railway passenger transport management operation person is collected, and the corresponding voiceprint feature is extracted based on the pre-trained ResCNN-GRU voiceprint recognition model, and the voiceprint templates of each railway passenger transport management operation person are obtained by mean or clustering;

[0022] The main storage database is constructed in the form of PostgreSQL combined with pgvector extension, which is used to store the voiceprint template, identity and collection time of each railway passenger transport management operation person; Redis cache is introduced for real-time comparison request.

[0023] In some embodiments, the noise reduction processing of the operation voice comprises: using a pre-trained voice enhancement model to perform noise reduction on the operation voice;

[0024] The pre-training step of the voice enhancement model comprises:

[0025] Obtain a training sample set, the training sample set contains multiple samples, each sample contains a clear voice and a noisy voice after adding noise to the clear voice;

[0026] Perform short-time Fourier transform on the noisy voice in each sample to extract noisy amplitude spectrum and noisy phase spectrum as input features; obtain an initial neural network, including an encoder and a decoder; the input features corresponding to each sample are obtained by the encoder based on the convolutional neural network to obtain time-frequency domain features; the decoder includes a parallel amplitude spectrum decoder and a phase spectrum decoder, the time-frequency domain features are output through the amplitude spectrum decoder to obtain the reconstructed amplitude spectrum, and the time-frequency domain features are output through the phase spectrum decoder to obtain the reconstructed phase spectrum; combine the reconstructed amplitude spectrum and the reconstructed phase spectrum and perform inverse short-time Fourier transform to obtain the reconstructed voice signal;

[0027] The initial neural network is trained by using the training sample set, a loss is constructed by combining the deviation of the noisy amplitude spectrum and the reconstructed amplitude spectrum, the deviation of the noisy phase spectrum and the reconstructed phase spectrum, and the deviation of the clear voice and the reconstructed voice signal, the initial neural network is updated in parameters to obtain the voice enhancement model.

[0028] In some embodiments, the preset voice text conversion model adopts a Conformer model.

[0029] The method further comprises correcting the station guide screen operation state data, the fault log and the operation environment sensor data by using a Seq2Seq text correction model.

[0030] In some embodiments, the method further comprises aligning the key entity information and the operation text based on dynamic time warping.

[0031] In another aspect, the present application also provides a railway passenger operation safety risk hidden danger checking system, comprising:

[0032] A railway passenger operation communication subsystem is composed of a plurality of mobile communication devices and performs voice communication based on a preset communication link.

[0033] A railway passenger operation broadcasting subsystem is composed of a plurality of broadcasting devices.

[0034] A plurality of station guide screens are used to display guide operation information.

[0035] A sensor subsystem comprises station environment monitoring sensors, train arrival and running state sensing sensors, safety warning sensors and equipment health management sensors.

[0036] A full risk hidden danger checking processor is used to execute the steps of the above method.

[0037] In another aspect, the present application also provides a computer readable storage medium having a computer program / instruction stored thereon, which is executed by a processor to implement the steps of the above method.

[0038] The present application has at least the following advantages:

[0039] The application provides the railway passenger operation safety risk hidden danger checking method and system, through synchronous collection of operation voice of multiple communication channels and broadcast channels, station guide screen operation state data, fault log and various operation environment sensor data, all-round and multi-dimensional data acquisition of railway passenger operation is realized. Voiceprint recognition is carried out on the operation voice to find the subject in the preset operation personnel voiceprint library. The operation voice is denoised by a voice enhancement model, and then the operation text is obtained by using a preset voice text conversion model. The station guide screen operation state data, fault log and operation environment sensor data are extracted and corrected, key entity information is obtained, and the operation text is aligned according to the production time, the fusion and synchronization of the multi-source data are realized, and the basis for subsequent logical reasoning is provided. Based on the preset multi-modal risk knowledge graph, the operation text and the key entity information after semantic extraction are compared by logical reasoning, feedback and response are carried out according to the preset risk evaluation mechanism, the railway passenger operation safety risk hidden danger checking result is output in real time, and is fed back to the subject of the corresponding operation text. The real-time and accurate checking of the railway passenger operation safety risk hidden danger is realized, and the safety management level of the railway passenger operation is effectively improved.

[0040] Additional advantages, objects, and features of the application will be set forth in part by the description that follows, and will become apparent to those skilled in the art upon examination of the following figures and detailed description thereof or can be learned by practice of the application. The objects and other advantages of the application can be realized and attained by the structure particularly pointed out in the description and claims hereof as well as the appended drawings.

[0041] It will be understood by those within the art that the objects and advantages of the application can be met by the present application, and that the present application can be used to achieve the above and other objects and advantages. BRIEF DESCRIPTION OF DRAWINGS

[0042] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments of the application and together with the description, serve to explain the principles of the application. In the drawings:

[0043] Figure 1 The flowchart of the railway passenger operation safety risk hidden danger checking method of an embodiment of the application.

[0044] Figure 2 The technical framework diagram of the railway passenger operation safety risk hidden danger checking method of an embodiment of the application.

[0045] Figure 3 The logic diagram of the railway passenger operation safety risk hidden danger checking method of an embodiment of the application. DETAILED DESCRIPTION

[0046] In order to make the objects, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the embodiments and drawings. Herein, the illustrative embodiments of the present application and their descriptions are used to explain the present application but not as a limitation of the present application.

[0047] It should also be noted that, in order to avoid obscuring the present application with unnecessary details, only the structures and / or processing steps closely related to the solutions according to the present application are shown in the drawings, and other details not closely related to the present application are omitted.

[0048] It should be emphasized that the term "comprises / comprising" as used herein is used to indicate the presence of a feature, element, step or component but does not preclude the presence or addition of one or more other features, elements, steps or components.

[0049] The current railway passenger operation safety risk hidden danger investigation technology mainly relies on a single mode (such as pure voice or video monitoring), which is difficult to associate personnel operation and equipment state, resulting in blind area of hidden danger identification. The traditional method lacks automatic verification of large screen display content, and needs to rely on manual inspection, which is easy to cause information delay or error omission due to negligence; at the same time, the hidden danger judgment lacks dynamic alignment and causal reasoning ability of multi-modal data (such as voice instruction, broadcast content, large screen information), and it is difficult to effectively detect multi-source information conflict (such as inconsistent train number and platform). In addition, the existing technology is rigid in risk grading and response mechanism, and cannot dynamically adjust the priority according to the severity of information delay, loss or conflict, and the management of emergency broadcast instructions and regular intercom content lacks intelligent integration, which may cause key information to be disturbed by low priority conversation, affecting the timeliness of emergency disposal.

[0050] In view of this, the present application provides a railway passenger operation safety risk hidden danger investigation method, as shown in Figure 1 and Figure 2 The method comprises the following steps S101-S104:

[0051] Step S101: synchronously collecting operation voice of a plurality of communication channels and broadcast channels, station guide screen operation state data, fault log and a plurality of preset operation environment sensor data; the operation environment sensor data is image or text data collected for a plurality of railway passenger management operation elements.

[0052] Step S102: for the operation voice of the plurality of communication channels and broadcast channels in the railway passenger operation process, obtaining corresponding voiceprint features based on a preset voiceprint recognition model and finding the subject of the operation voice in a preset operation personnel voiceprint library.

[0053] Step S103: input the corresponding job text into the preset voice text conversion model after noise reduction processing of the job voice; extract and correct the text from the station guide screen job state data, fault log, and job environment sensor data to obtain key entity information about the railway passenger management job elements, and align the job text with the production time.

[0054] Step S104: perform semantic extraction on the job text and key entity information of the same railway passenger management job element in the same period after alignment, perform logical inference comparison based on the preset multi-modal risk knowledge graph, and perform feedback and response according to the preset risk assessment mechanism, output real-time railway passenger operation safety risk hidden danger investigation results, and feedback to the subject of the corresponding job text; the multi-modal risk knowledge graph is set according to the preset railway passenger operation rules.

[0055] In step S101, the job voice based on the communication channel can be generated on a walkie-talkie, a mobile phone, or a special device, and the broadcast channel is a broadcast line for passengers or management personnel in the railway station. In some embodiments, the walkie-talkie job voice is obtained from multiple communication channels according to a first set sampling frequency, which can be 16 kHz sampling frequency and H.265 encoding format. The broadcast job voice is obtained from multiple broadcast channels according to a second set sampling frequency, which can be greater than or equal to 16 kHz sampling frequency. Further, for the job voice, the voice endpoint detection technology (VAD) can be used to divide it into multiple effective instruction segments and mark the sampling time point.

[0056] The station guide screen job state data can include the number, time, and stop platform of the railway arrival or departure vehicle displayed through the station guide screen, the passenger flow, temperature and humidity in the station, and the emergency information. The fault log can be a series of problem records generated during the railway passenger operation process based on the set system, including the equipment inside and outside the station, the train operation scheduling state, and the personnel management and guidance process in the station. In some embodiments, the station guide screen job state data is obtained by the OCR text extraction method based on the image data collected by the preset image sensor on the station guide screen; the fault log is recorded in JSON format.

[0057] The job environment sensor data is a series of data collected by sensors in the passenger management process of a railway station. In some embodiments, the job environment sensor data is derived from station environment monitoring sensors, train arrival and running state perception sensors, safety warning sensors, and equipment health management sensors. The station environment monitoring sensors include temperature and humidity sensors, air quality sensors, noise sensors, light sensors, and passenger flow monitoring sensors deployed at multiple locations in the railway station. The train arrival and running state perception sensors include track deformation sensors, train positioning sensors, and train speed sensors deployed at multiple locations on the track. The safety warning sensors include fire sensors and security sensors. The equipment health management sensors include circuit monitoring sensors, mechanical vibration sensors, and water leakage detection sensors.

[0058] In step S102, the corresponding voiceprint features are obtained based on the preset voiceprint recognition model, and the subject of the operation voice is searched in the preset operation personnel voiceprint library, including: using a pre-trained ResCNN-GRU voiceprint recognition model to extract voiceprint features, and comparing the similarity with the voiceprint templates stored in the preset operation personnel voiceprint library, and outputting the object with the highest similarity as the subject.

[0059] The creation steps of the preset operation personnel voiceprint library include steps S201-S202:

[0060] Step S201: Based on the preset link, collect multiple voice segments of each railway passenger management operation person in charge, and extract the corresponding voiceprint features based on the pre-trained ResCNN-GRU voiceprint recognition model, and obtain the voiceprint templates of each railway passenger management operation person in charge through mean value or clustering.

[0061] Step S202: Use the form of PostgreSQL combined with pgvector extension to build a main storage database for storing the voiceprint templates, identity and collection time of each railway passenger management operation person in charge; introduce Redis cache for real-time comparison request.

[0062] In step S201, the ResCNN-GRU voiceprint recognition model is a deep learning model combining residual convolutional neural network (ResCNN) and gated recurrent unit (GRU) for voiceprint recognition task. The ResCNN part is composed of multiple residual blocks, each containing two branches. One branch undergoes several convolution operations, and the last convolution does not perform activation. The other branch does not perform any convolution operation. The outputs of the two branches are summed and then output through the ReLU activation function. This structure helps to alleviate the gradient vanishing problem in deep networks, enabling the network to more effectively learn local features in the speech signal. The GRU part is an optimized version of the LSTM structure, with fewer parameters than LSTM. It can handle long-time dependent sequence data and is suitable for capturing the timing patterns in speech, better modeling the dynamic changes of speech signals.

[0063] The speech signal is converted into a time-frequency graph (such as a mel spectrum graph) or a waveform as the input of the model. The ResCNN part extracts local features from the input speech features, and the GRU part models the sequence of these local features to capture the temporal dynamic information of the speech. Through the combination of ResCNN and GRU, the model can learn both local spatial features and global temporal features contained in the speech signal, thereby generating a voiceprint embedding vector that can represent the speaker's identity. In voiceprint verification, the similarity is calculated by comparing the embedding vector of the new input speech with the voiceprint embedding vector stored in the database, and it is determined whether they match. Cosine similarity is usually used for comparison because the length has been normalized, saving the calculation of the numerator of the cosine.

[0064] In step S202, a main storage database based on PostgreSQL is constructed, and the pgvector extension is used to store the vectorized features of the voiceprint templates, supporting efficient feature similarity retrieval (such as Euclidean distance calculation), while associating the identity identifier with the collection time to realize version management of voiceprint data; by introducing a Redis cache layer, the frequently accessed voiceprint templates are preloaded into memory, taking advantage of its low latency characteristics to speed up the processing of real-time comparison requests, and through cache eviction strategies (such as LRU) to dynamically maintain hot data, ensuring that the system can still respond quickly to identity verification needs in high concurrency scenarios, while the asynchronous synchronization mechanism between the database and the cache ensures data consistency.

[0065] In step S103, the noise reduction processing of the task speech includes: using a pre-trained speech enhancement model to perform noise reduction on the task speech. The pre-training steps of the speech enhancement model include steps S301-S303:

[0066] Step S301: obtaining a training sample set, where the training sample set includes multiple samples, each sample including a clear speech and a noisy speech obtained by adding noise to the clear speech.

[0067] Step S302: Perform a short-time Fourier transform on the noisy speech in each sample, and extract the noisy amplitude spectrum and the noisy phase spectrum as input features; obtain an initial neural network, including an encoder and a decoder; the input features corresponding to each sample are used to obtain time-frequency domain features through an encoder based on a convolutional neural network; the decoder includes a parallel amplitude spectrum decoder and a phase spectrum decoder, and the time-frequency domain features are outputted as a reconstructed amplitude spectrum through the amplitude spectrum decoder, and as a reconstructed phase spectrum through the phase spectrum decoder; the reconstructed amplitude spectrum and the reconstructed phase spectrum are combined and inverse short-time Fourier transform is performed to obtain a reconstructed speech signal.

[0068] Step S303: The initial neural network is trained using the training sample set. The loss is constructed by combining the deviation between the noisy amplitude spectrum and the reconstructed amplitude spectrum, the deviation between the noisy phase spectrum and the reconstructed phase spectrum, and the deviation between the clear speech and the reconstructed speech signal. The parameters of the initial neural network are updated to obtain a speech enhancement model.

[0069] In some embodiments, the preset speech-to-text conversion model uses the Conformer model. The Conformer model is a neural network model that combines a convolutional neural network (CNN) with the Transformer architecture and is specifically designed for speech recognition tasks. By introducing convolutional layers into the Transformer architecture, it can simultaneously capture both local and global dependencies in speech signals, resulting in excellent performance in speech recognition tasks.

[0070] The method further includes: using a Seq2Seq text error correction model to correct the station guide screen operation status data, fault logs and operation environment sensor data.

[0071] The Seq2Seq (Sequence-to-Sequence) text error correction model is a deep learning model based on an encoder-decoder architecture. It maps an input sequence of erroneous text to a corrected sequence. It is commonly used to address spelling errors, grammatical errors, or optical character recognition (OCR) errors. Its core approach is to train on a large number of pairs of erroneous and correct text, enabling the model to learn the mapping from error patterns to standard representations. For example, this model uses a Transformer architecture to capture long-range dependencies and incorporates an attention mechanism to dynamically align input and output.

[0072] In the railway passenger scenario, for the operation state data of the station guide screen (such as the display text extracted by OCR), the fault log (such as the error field in JSON), and the sensor data (such as the textualized environmental parameters), the specific implementation of correction through the Seq2Seq model is as follows: First, a training set containing typical error samples (such as the "G5678 times → G56B8 times" misrecognized by the guide screen OCR and the abnormal unit "30℃ → 300℃" in the sensor data) is constructed, and the corresponding correct text is labeled; then, the model is trained to learn the error patterns and correction rules (such as the confusion of numbers and letters, the absence of unit symbols); in actual application, the originally extracted text is input into the model to generate the corrected text (such as "3 station platform → 3 station platform"), and through the post-processing module and the structured field (train number, time), logical verification (such as whether the platform number is within a reasonable range) is performed, and finally standardized data aligned with the operation time is output to ensure the accuracy of subsequent risk reasoning. At the same time, the decoding process can be constrained in combination with the railway joint control language specification library to limit the model output to conform to industry terminology and avoid semantic deviation.

[0073] In some embodiments, the method further comprises: aligning the key entity information and the operation text based on dynamic time warping (DTW).

[0074] Using the dynamic time warping (DTW) algorithm to align the voice events (such as intercom, broadcast) and the large screen state changes needs to follow the following steps:

[0075] 1. Data preprocessing

[0076] Voice event data processing: Convert voice events into time series data, recording the start time, end time, and corresponding feature values (such as volume, frequency, etc.) of each event. For example, you can use a speech processing library (such as Librosa) to extract key features of the voice signal to form a time series, where each time point corresponds to a feature vector.

[0077] Large screen state change data processing: Similarly, organize the data of large screen state changes into time series, recording the time points and state values of each state change. This may involve extracting relevant information from log files or monitoring data and converting it into a format compatible with voice event data.

[0078] 2. Feature extraction and selection

[0079] Select appropriate features: According to the specific alignment target, select parameters that can effectively reflect the characteristics of voice events and large screen state changes. For voice events, it may include duration, frequency change, energy, etc.; for large screen state changes, it may include the time interval of state switching, the amplitude of state value change, etc.

[0080] Feature Normalization: To eliminate the influence of different feature dimensions and magnitudes, the extracted features are normalized to be in the same numerical range, for example, mapping all feature values to the interval [0, 1].

[0081] 3. DTW Algorithm Implementation and Application

[0082] Distance Matrix Construction: Calculate the distance between each corresponding point in the speech event time series and the large screen state change time series, usually using Euclidean distance or other appropriate distance metrics, to form a two-dimensional distance matrix.

[0083] Finding the Optimal Path: Use dynamic programming techniques to find the path in the distance matrix that minimizes the cumulative distance, i.e., the optimal alignment path. This path represents how to match the points in one time series with the points in another time series to achieve the best time alignment.

[0084] In step S104, the semantic extraction of the job text and key entity information can use natural language processing technology (NLP) for semantic extraction. The present application constructs a multi-modal risk knowledge graph, logically compares data collected from various channels, monitors abnormalities and faults, and conducts risk assessment. The comparison logic between various projects in the multi-modal risk knowledge graph should be set in accordance with the railway passenger transport management specifications, for example, the train arrival time announced in the broadcast should be consistent with the content displayed on the station guide screen.

[0085] For risk assessment, risk classification and response mechanisms can be used. For example, a delay in the display or voice information on the station guide screen is defined as a low risk, and is automatically recorded and notified to the on-duty officer. Any two of the intercom, broadcast, and large screen information conflict, defined as medium risk, triggers manual review. Complete lack of multi-source information or serious conflict is defined as high risk, which forces the switching of the standby screen and broadcasts an alert.

[0086] On the other hand, the present application also provides a railway passenger operation safety risk hidden danger investigation system, comprising: a railway passenger operation communication subsystem, a railway passenger operation broadcast subsystem, a plurality of station guide screens, a sensor subsystem and a full risk hidden danger investigation processor.

[0087] The railway passenger operation communication subsystem is composed of a plurality of mobile communication devices, and voice communication is based on a pre-set communication link. The railway passenger operation broadcast subsystem is composed of a plurality of broadcast devices. The station guide screen is used to display guide operation information. The sensor subsystem includes station environment monitoring sensors, train arrival and running state sensing sensors, safety warning sensors and device health management sensors. The full risk hidden danger investigation processor is used to execute the steps of the method described in steps S101-S104.

[0088] In another aspect, the present application also provides a computer readable storage medium, which stores computer programs / instructions, and the computer programs / instructions are executed by a processor to implement the steps of the above method.

[0089] The present application will be described below in combination with a specific embodiment:

[0090] The embodiment provides a railway passenger operation safety risk hidden danger checking method, referring to Figure 3 , comprising the following steps:

[0091] 1. Multi-modal data fusion

[0092] Voice stream: intercom audio (16 kHz sampling, H.265 encoding), broadcast system audio stream (real-time parsing, sampling rate ≥ 16 kHz), and effective instruction segments are segmented through voice activity detection (VAD).

[0093] Large screen data: display content screenshot (OCR text extraction), fault log (JSON format).

[0094] Environmental data: temperature and humidity, fire sensor (MQTT protocol transmission).

[0095] 2. Core algorithm

[0096] Voiceprint-voice joint modeling: speaker features are extracted through ResCNN-GRU, broadcast system and artificial instruction sources are distinguished synchronously, noise reduction is performed through an encoder-decoder voice enhancement model (amplitude mask and phase spectrum reconstruction), and high-precision conversion from voice to text is realized in combination with a Conformer model.

[0097] Cross-modal alignment and error correction: DTW algorithm is used to align voice events (intercom, broadcast) and large screen state changes, a Seq2Seq text error correction model is used, recognized text is corrected in combination with a railway joint control language specification library, and key entities such as train number and operation time are extracted.

[0098] Dynamic knowledge graph reasoning: rules in “Railway Technical Management Regulations” are embedded into a differentiable reasoning engine, logical consistency of voice instructions, broadcast content and large screen display is calculated, and an example rule is: if the broadcast instruction is “stop operation notice”, then the large screen is forced to display the stop operation train number and trigger a red warning identifier.

[0099] Multi-modal priority management: the weight of broadcast and intercom voice is dynamically allocated through a Transformer multi-source voice feature fusion module, and emergency broadcast instructions (such as “emergency evacuation”) automatically cover low-priority intercom content.

[0100] 3. Risk classification and response mechanism

[0101] Level 1 (low risk): display or voice information delay (<30 seconds), automatically record and notify the on-duty staff.

[0102] Level 2 (medium risk): any two of intercom, broadcast, and large screen information conflict, trigger manual review.

[0103] Level 3 (high risk): complete lack of multi-source information or serious conflict, forced switching to backup screen and broadcast warning.

[0104] 4. Implementation effect

[0105] Display content error detection rate increased to 98% (compared to manual inspection), high-risk event response time ≤1 second; broadcast instruction priority management effectively avoids key information being disturbed by low-priority conversations; multi-modal conflict (such as broadcast and screen information contradiction) processing accuracy ≥95%.

[0106] 5. Implementation case

[0107] Scenario: Station large screen information error risk inspection.

[0108] Intercom voice instruction: staff issued "display G5678 train stops at platform 3".

[0109] Broadcast voice instruction: broadcast system broadcasted "G5678 train stops at platform 3".

[0110] Large screen data: screen capture OCR detects display content as "G5678 stops at platform 2".

[0111] Model action: match "train number-platform" mapping rules in the knowledge base; detect inconsistent platform numbers (3 ≠ 2); trigger Level 2 warning and perform manual review as required.

[0112] Corresponding to the above method, the application also provides a device / system, which comprises a computer device including a processor and a memory, the memory storing computer instructions, and the processor being configured to execute the computer instructions stored in the memory, so that the device / system implements the steps of the above method.

[0113] The embodiment of the present application also provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the steps of the aforementioned edge computing server deployment method. The computer readable storage medium can be a tangible storage medium, such as a random access memory (RAM), a memory, a read only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a floppy disk, a hard disk, a removable storage disk, a CD-ROM, or any other form of storage medium known in the technical field.

[0114] To sum up, the present application provides the railway passenger operation safety risk hidden danger checking method and system, through synchronously collecting operation voice of multiple communication channels and broadcast channels, station guide screen operation state data, fault logs and various operation environment sensor data, all-round and multi-dimensional data acquisition of railway passenger operation is realized. Voiceprint recognition is performed on the operation voice to find the subject in the preset operation personnel voiceprint library. The operation voice is denoised by using the voice enhancement model, and then the operation text is obtained by using the preset voice text conversion model. The station guide screen operation state data, fault logs and operation environment sensor data are extracted and corrected to obtain key entity information, and are aligned with the operation text according to the generation time, the fusion and synchronization of the multi-source data are realized, and the basis for subsequent logical reasoning is provided. Based on the preset multi-modal risk knowledge graph, the operation text and the key entity information after semantic extraction are compared by logical reasoning, and feedback and response are performed according to the preset risk evaluation mechanism, the railway passenger operation safety risk hidden danger checking result is output in real time, and is fed back to the subject of the corresponding operation text. Real-time and accurate checking of the railway passenger operation safety risk hidden danger is realized, and the safety management level of the railway passenger operation is effectively improved.

[0115] Those of ordinary skill in the art will appreciate that the various illustrative components, systems and methods described in connection with the embodiments disclosed herein can be implemented as hardware, software, or hardware and software in combination. The particular implementation as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans appreciate the fact that the hardware and software depicted in the figures can be implemented in an order different than that of the figures. Skilled artisans also appreciate that the depicted examples are only exemplary and can be implemented in various other ways than those explicitly described herein. When implemented in hardware, the hardware can be implemented in, for example, electronic circuitry, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a programmable logic device (PLD), a controller, a microcontroller, a microprocessor, or any other suitable hardware component. When implemented in software, the software can be executed by a processing unit, such as a microprocessor, software modules, or code can be stored on a machine-readable medium, such as a floppy disk, a hard disk, a CD ROM, a RAM, a ROM, a DVD, a Blu-ray Disc, a magnetic tape, a flash memory, a microprocessor, a microcontroller, a microprocessor, or any other suitable hardware component.

[0116] It is to be expressly understood that the invention is not limited to the specific configurations and process described above and illustrated in the accompanying drawings. For the sake of clarity, detailed descriptions of known methods are omitted. In the above-described embodiments, several specific steps are described and illustrated as examples. However, the method processes of the present invention are not limited to the specific steps described and illustrated, and various changes, modifications and additions can be made thereto by one of ordinary skill in the art without departing from the spirit of the present invention, and the order of the steps can be changed.

[0117] In the present invention, features described and / or illustrated with respect to one embodiment can be used in the same or a similar way in one or more other embodiments, and / or in combination with or instead of features of other embodiments.

[0118] The above description is merely illustrative of the application, and is not intended to limit the scope of the application. Various modifications and changes can be made by one of ordinary skill in the art without departing from the spirit and scope of the application. Any modification, equivalent replacement, improvement, and the like made within the spirit and principle of the application should be included in the scope of the application.

Claims

1. A method for troubleshooting safety risks in railway passenger transport operations, characterized in that: The method comprises the following steps: Synchronously collects operation voices from multiple communication channels and broadcast channels, operation status data of station guidance screens, fault logs, and multiple preset operation environment sensor data; the operation environment sensor data is image or text data collected from multiple railway passenger management operation elements; For the operation voices on multiple communication channels and broadcast channels during railway passenger operations, the corresponding voiceprint features are obtained based on a preset voiceprint recognition model, and the speaker of the operation voice is searched in a preset operator voiceprint library; The operation voice is subjected to noise reduction processing and then input into a preset speech-to-text conversion model to obtain a corresponding operation text; text extraction and correction are performed on the operation status data of the station guide screen, the fault log, and the operation environment sensor data to obtain key entity information about the railway passenger management operation elements, and the key entity information is aligned with the operation text according to the generation time; Semantic extraction is performed on the operation text and the key entity information after the same railway passenger management operation elements in the railway passenger operation are aligned in the same period, and logical reasoning and comparison are performed based on the preset multimodal risk knowledge graph, and feedback and response are performed according to the preset risk assessment mechanism, and real-time railway passenger operation safety risk hazard inspection results are output and fed back to the corresponding issuer of the operation text; the multimodal risk knowledge graph is set according to the preset railway passenger operation rules.

2. The method for troubleshooting safety risks in railway passenger transport operations according to claim 1, characterized in that: Synchronously collects operational voice from multiple communication and broadcast channels, including: The walkie-talkie operation voice is obtained from multiple communication channels according to the first set sampling frequency, and the broadcast operation voice is obtained from multiple broadcast channels according to the second set sampling frequency, and is divided into multiple valid instruction segments through voice endpoint detection and the sampling time points are marked.

3. The method for troubleshooting safety risks in railway passenger transport operations according to claim 1, characterized in that: The station guide screen operation status data is obtained by OCR text extraction after a preset image sensor collects image data of the station guide screen; the fault log is recorded in JSON format.

4. The method for troubleshooting safety risks in railway passenger transport operations according to claim 1, characterized in that: The operating environment sensor data is derived from station environment monitoring sensors, train arrival and operation status perception sensors, safety warning sensors, and equipment health management sensors; The station environment monitoring sensors include temperature and humidity sensors, air quality sensors, noise sensors, light sensors and passenger flow monitoring sensors deployed at multiple locations in railway stations; The train arrival and running status sensing sensors include track deformation sensors, train positioning sensors and train speed sensors deployed at multiple locations on the track; The safety warning sensors include fire protection sensors and security sensors; The equipment health management sensors include: circuit monitoring sensors, mechanical vibration sensors and water leakage detection sensors.

5. The method for troubleshooting safety risks in railway passenger transport operations according to claim 1, characterized in that: Acquiring corresponding voiceprint features based on a preset voiceprint recognition model and searching for the speaker of the work speech in a preset operator voiceprint library, including: extracting the voiceprint features using a pre-trained ResCNN-GRU voiceprint recognition model, performing a similarity comparison with the voiceprint templates stored in the preset operator voiceprint library, and outputting the object with the highest similarity as the speaker; The steps of creating the preset operator voiceprint database include: Based on the preset link, multiple voice segments of the person in charge of each railway passenger transport management operation are collected, and the corresponding voiceprint features are extracted based on the pre-trained ResCNN-GRU voiceprint recognition model. The voiceprint templates of each person in charge of railway passenger transport management operations are obtained through averaging or clustering; The main storage database is constructed using PostgreSQL combined with the pgvector extension to store the voiceprint template, identity identifier and collection time of each person in charge of railway passenger management operations; Redis is introduced to cache real-time comparison requests.

6. The method for troubleshooting safety risks in railway passenger transport operations according to claim 1, characterized in that: The noise reduction processing of the operation speech includes: using a pre-trained speech enhancement model to reduce the noise of the operation speech; The pre-training step of the speech enhancement model includes: Acquire a training sample set, wherein the training sample set includes a plurality of samples, each sample including a clear speech and a noisy speech obtained by adding noise to the clear speech; Performing a short-time Fourier transform on the noisy speech in each sample, extracting a noisy amplitude spectrum and a noisy phase spectrum as input features; obtaining an initial neural network, including an encoder and a decoder; obtaining time-frequency domain features from the input features corresponding to each sample through the encoder based on a convolutional neural network; the decoder including a parallel amplitude spectrum decoder and a phase spectrum decoder, the time-frequency domain features outputting a reconstructed amplitude spectrum through the amplitude spectrum decoder, and outputting a reconstructed phase spectrum through the phase spectrum decoder; combining the reconstructed amplitude spectrum and the reconstructed phase spectrum and performing an inverse short-time Fourier transform to obtain a reconstructed speech signal; The initial neural network is trained using the training sample set, and a loss is constructed by combining the deviation of the noisy amplitude spectrum and the reconstructed amplitude spectrum, the deviation of the noisy phase spectrum and the reconstructed phase spectrum, and the deviation of the clear speech and the reconstructed speech signal. The parameters of the initial neural network are updated to obtain the speech enhancement model.

7. The method for troubleshooting safety risks in railway passenger transport operations according to claim 1, characterized in that: The preset speech-to-text conversion model adopts the Conformer model; The method further includes: using a Seq2Seq text error correction model to correct the station guide screen operation status data, the fault log and the operation environment sensor data.

8. The method for troubleshooting safety risks in railway passenger transport operations according to claim 1, characterized in that: The method further includes: aligning the key entity information and the job text based on dynamic time warping.

9. A railway passenger operation safety risk hidden danger investigation system, characterized by: include: The railway passenger operation communication subsystem consists of multiple mobile communication devices and conducts voice communication based on preset communication links; Railway passenger operation broadcasting subsystem, consisting of multiple broadcasting devices; Multiple station guidance screens for displaying guidance operation information; Sensor subsystem, including station environment monitoring sensors, train arrival and operation status perception sensors, safety warning sensors, and equipment health management sensors; A full-risk hidden danger investigation processor is used to execute the steps of the method as claimed in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Truck key fault intelligent identification method, device and equipment and medium

    CN121686113A

  • Method and system for call priority control for column intercoms

    CN122372524A