A rail transit equipment anomaly detection method, system, device and medium
Patent Information
- Application Number
- CN202611077270.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-20
- Publication Date
- 2026-09-15
AI Technical Summary
[0010]针对现有技术中对轨道交通设备的声学监测无法同时实现多异常类型的精确分类以及异常事件的精确里程定位的缺陷,本申请提出了一种轨道交通设备异常检测方法、系统、设备和介质,该方法能够在复杂背景噪声下,精确识别列车/站台门开关异常、列车牵引启动异常、轮轨摩擦高频啸叫、接触件摩擦异响、轴承故障等多种异常事件,并自动将异常事件绑定至具体的列车运行轨道里程点,以便于进行及时维护
[0051] This application proposes a method for detecting anomalies in rail transit equipment. By combining audio log-Mel spectrum analysis with a two-level cascaded classifier, it can accurately identify various abnormal events in rail transit equipment, such as abnormal opening and closing of train/platform doors, abnormal train traction and starting, high-frequency whistling due to wheel-rail friction, abnormal friction noise of contact parts, and bearing failure. It automatically binds the abnormal events to specific track mileage points for timely maintenance. At the same time, it can also provide accurate and actionable maintenance suggestions directly to relevant departments (command center, rolling stock maintenance, station maintenance, etc.) and personnel through a maintenance suggestion expert knowledge base. Compared with existing technologies, this method has advantages such as high accuracy, good real-time performance, rapid response, and low cost.
Smart Images

Figure CN122761901A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical fields of anomaly recognition and speech recognition, and specifically relates to a method, system, equipment and medium for anomaly detection of rail transit equipment. Background Technology
[0002] The operational safety of rail transit systems is highly dependent on the health status of key equipment such as the train running gear, traction drive system, track structure, and door system. When equipment experiences early failures, its physical response often manifests as abnormal vibrations and acoustic emissions, such as high-frequency whistling caused by abnormal wheel-rail friction, periodic impacts caused by bearing raceway spalling, and modulation sidebands caused by gear meshing errors.
[0003] These acoustic signals are non-contact, perceptible, and contain fault information. If voice capture, voice recognition and classification can be performed accurately and the location can be determined, then all-weather anomaly monitoring and predictive maintenance can be achieved without intruding on the equipment structure.
[0004] Currently, the industry mainly relies on the following technologies for monitoring the operational status and identifying anomalies of rail transit equipment, but these technologies have significant shortcomings:
[0005] (1) Manual inspection and subjective listening: Maintenance personnel listen closely to the sound after the vehicle returns to the depot, or the train driver / passenger makes a subjective judgment by perceiving abnormal noises by human ear during train operation. This technology relies entirely on the personal experience of maintenance personnel or drivers, without objective quantitative data support, and its accuracy and reliability are easily affected by subjectivity; at the same time, some audio frequency bands are high and difficult to be captured by human ear; it cannot cover unmanned driving lines; and it is very easy to miss detection.
[0006] (2) Contact vibration monitoring: Accelerometers are installed in parts such as bogie bearings and gearbox housings to acquire vibration time-domain signals, calculate fault characteristic frequencies, and determine whether the equipment is faulty. This technology requires the vibration sensors and cables to be installed in the vehicle equipment structure, which involves a large amount of engineering work and high costs for modifying existing equipment; in addition, the transmission path of vibration signals in solid components is complex, and high-frequency attenuation is severe, making it difficult to capture certain specific abnormal sounds.
[0007] (3) Periodic inspection based on track inspection vehicles: Using specialized equipment such as track inspection vehicles and rail flaw detection vehicles, abnormalities such as corrugation, scratches, and unevenness are detected on the entire track at fixed intervals. This technology has a single detection target, and the detection results are mostly geometric dimensional parameters; the detection vehicles are specialized equipment with high costs, complex scheduling, and are not easy to maintain; and it cannot achieve all-weather real-time monitoring with the vehicle in operation.
[0008] (4) Audio analysis-based technology: Audio pickup devices are deployed along the track to acquire audio data when vehicles pass by, extracting short-time energy features, Mel-frequency cepstral coefficients, etc., and using models such as support vector machines and convolutional neural networks to achieve binary classification of "normal / abnormal". This technology has relatively coarse classification capabilities, and can only output fuzzy abnormal alarms, unable to distinguish fault types in a fine-grained manner (such as abnormal noises when opening and closing doors, gear wear, wheel-rail friction whistling, etc.); or it only focuses on the abnormal faults of the vehicle itself, without involving non-vehicle equipment such as tracks and platform doors, resulting in the final alarm information lacking operational guidance value.
[0009] In addition, existing technologies have largely failed to solve the problem of geospatial location of abnormal acoustic events, making it impossible to accurately link abnormal events to specific track mileage points or platform sections. This results in the inability to perform maintenance operations such as grinding, applying lubricant, and replacing parts in a timely manner based on alarm information. Summary of the Invention
[0010] To address the shortcomings of existing technologies in acoustic monitoring of rail transit equipment, which cannot simultaneously achieve accurate classification of multiple anomaly types and precise mileage location of abnormal events, this application proposes a method, system, equipment, and medium for anomaly detection in rail transit equipment. This method can accurately identify multiple abnormal events under complex background noise, such as abnormal opening and closing of train / platform doors, abnormal train traction and starting, high-frequency whistling due to wheel-rail friction, abnormal friction noise of contact parts, and bearing failure. It also automatically binds the abnormal events to specific train track mileage points to facilitate timely maintenance.
[0011] This application is achieved through the following technical solution:
[0012] A method for detecting anomalies in rail transit equipment, comprising:
[0013] Simultaneously acquire signals from multiple sources, including audio signals, train speed, beacon mileage, and timestamps;
[0014] The acquired audio signal is divided into audio segments of a preset duration, and the start and end timestamps of each audio segment are recorded.
[0015] Generate a log-Mel spectrum corresponding to each audio segment;
[0016] The log-Mel spectrum of each audio segment is input into the first-level classifier to extract features and identify the specific type of acoustic event. Different types of acoustic events and their corresponding features are input into the corresponding second-level classifier for normal or abnormal classification.
[0017] Based on the train speed, beacon mileage, and timestamp, the audio segments identified as abnormal acoustic events are precisely located.
[0018] Based on the combined abnormal event type, abnormal location results, and log-Mel spectrogram of the abnormal audio segment, an alarm report is generated and the alarm information is synchronized to relevant departments.
[0019] In some implementations, the anomaly detection method further includes:
[0020] The abnormal event type is input into the maintenance solution expert knowledge base, which automatically associates and generates corresponding maintenance suggestions, which are then synchronized to relevant departments.
[0021] In some implementations, generating the log-Melb spectrogram corresponding to each audio segment includes:
[0022] The audio segment is pre-emphasized;
[0023] The pre-emphasized audio segment is divided into short time frames;
[0024] Perform an N-point discrete Fourier transform on each short-time frame to obtain a complex spectrum, and calculate the power spectrum of the short-time frame by taking the complex spectrum of the first N / 2+1 frequency points.
[0025] Construct a Mel filter bank;
[0026] The power spectrum of each short-time frame is passed through the Mel filter bank to obtain the Mel spectrum energy, and the logarithm of the Mel spectrum energy is taken to obtain the logarithmic Mel spectrum vector of each short-time frame.
[0027] The log-Mel spectrum vectors of all short-time frames are concatenated column by column to form a two-dimensional matrix, which is the log-Mel spectrum.
[0028] In some implementations, inputting the log-Mel spectrum of each audio segment into a first-level classifier for feature extraction and identification of the specific type of acoustic event includes:
[0029] The log-Mel spectrum of each audio segment is input into the feature extraction layer to extract a high-dimensional feature map;
[0030] The high-dimensional feature map is then subjected to global average pooling layer to obtain a feature vector of fixed length.
[0031] The feature vectors are then used for preliminary multi-event classification through a SoftMax layer to obtain acoustic event categories.
[0032] In some implementations, the step of inputting different types of acoustic events and their corresponding features into a corresponding second-level classifier for normal or abnormal classification includes:
[0033] The acoustic event categories and feature vectors after preliminary classification are sent to the model manager;
[0034] The model manager sends the corresponding feature vectors to the classifier of the corresponding acoustic event according to the acoustic event category for normal or abnormal classification.
[0035] In some implementations, the precise location of audio segments identified as anomalous acoustic events based on the train speed, beacon mileage, and timestamp includes:
[0036] Extract the beacon reference signal that is closest to the start timestamp of the audio segment of the abnormal acoustic event, obtain the end timestamp of the audio segment, obtain the absolute mileage corresponding to the beacon and the time when the train passes the beacon;
[0037] The time when the abnormal acoustic event occurred was calculated based on the start and end timestamps of the audio segment.
[0038] The train speed is integrated between the time the train passes the beacon and the time the abnormal acoustic event occurs to obtain the distance offset of the event relative to the beacon.
[0039] Based on the train's direction of travel, the mileage offset of the event to the beacon is increased or decreased on the absolute mileage of the beacon to obtain the absolute mileage of the abnormal acoustic event.
[0040] Secondly, this application proposes a rail transit equipment anomaly detection system, comprising:
[0041] The synchronous acquisition module is used to synchronously acquire signals from multiple sources, including audio signals, train speed, beacon mileage, and timestamps.
[0042] An audio preprocessing module is used to divide the acquired audio signal into audio segments of a preset duration and record the start and end timestamps of each audio segment.
[0043] A spectrogram generation module is used to generate a log-Mel spectrogram corresponding to each audio segment;
[0044] The second-level cascaded classification module is used to input the log-Mel spectrogram corresponding to each audio segment into the first-level classifier to extract features and identify the specific type of acoustic event. Different types of acoustic events and their corresponding features are input into the corresponding second-level classifier for normal or abnormal classification.
[0045] The multi-source fusion positioning module is used to accurately locate audio segments identified as abnormal acoustic events based on the train's operating speed, beacon mileage, and timestamp.
[0046] In addition, an anomaly handling module, including an alarm unit, is used to generate alarm information and synchronize the alarm information to relevant departments by comprehensively considering the anomaly event type, anomaly location result, and log-Mel spectrum of the anomaly audio segment.
[0047] In some embodiments, the exception handling module further includes:
[0048] The built-in maintenance solution expert knowledge base is used to automatically associate abnormal event types with corresponding maintenance suggestions and push the maintenance suggestions to relevant departments simultaneously.
[0049] Thirdly, this application proposes an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the rail transit equipment anomaly detection method described in any of the above embodiments.
[0050] Fourthly, this application proposes a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the rail transit equipment anomaly detection method described in any of the above embodiments.
[0051] This application proposes a method for detecting anomalies in rail transit equipment. By combining audio log-Mel spectrum analysis with a two-level cascaded classifier, it can accurately identify various abnormal events in rail transit equipment, such as abnormal opening and closing of train / platform doors, abnormal train traction and starting, high-frequency whistling due to wheel-rail friction, abnormal friction noise of contact parts, and bearing failure. It automatically binds the abnormal events to specific track mileage points for timely maintenance. At the same time, it can also provide accurate and actionable maintenance suggestions directly to relevant departments (command center, rolling stock maintenance, station maintenance, etc.) and personnel through a maintenance suggestion expert knowledge base. Compared with existing technologies, this method has advantages such as high accuracy, good real-time performance, rapid response, and low cost.
[0052] Accordingly, the rail transit equipment anomaly detection system, electronic equipment, and computer-readable storage medium proposed in this application also possess the same technical effects as described above. Attached Figure Description
[0053] The accompanying drawings, which are included to provide a further understanding of the embodiments of this application and form part of this application, do not constitute a limitation on the embodiments of this application. In the drawings:
[0054] Figure 1 This is a schematic diagram of the anomaly detection method proposed in the embodiments of this application;
[0055] Figure 2 Two log-Melbourne spectrograms of the train during normal operation at different sampling times were generated for embodiments of this application;
[0056] Figure 3 Two log-Mel spectrograms of the train doors opening at different sampling times were generated for embodiments of this application;
[0057] Figure 4 Two log-Mel spectrograms of the train at different sampling times were generated for the embodiments of this application.
[0058] Figure 5 Two log-Mel spectra of wheel-rail friction noise at different sampling times were generated for embodiments of this application;
[0059] Figure 6 This is a schematic diagram of the ResNet-50 network structure proposed in an embodiment of this application;
[0060] Figure 7 This is a schematic diagram of the two-stage cascaded audio classification process based on ResNet-50 and SVM classifiers proposed in the embodiments of this application;
[0061] Figure 8 This is a schematic diagram of the anomaly detection system architecture proposed in an embodiment of this application;
[0062] Figure 9 This is a schematic diagram of the electronic device proposed in the embodiments of this application;
[0063] Figure 10 This is a schematic diagram of a computer-readable storage medium proposed in an embodiment of this application;
[0064] Figure reference numerals and corresponding component names:
[0065] 200 - Anomaly detection system; 201 - Synchronization acquisition module; 202 - Audio preprocessing module; 203 - Spectrum diagram generation module; 204 - Second-level cascade classification module; 205 - Multi-source fusion localization module; 206 - Anomaly handling module; 300 - Electronic device; 310 - Memory; 320 - Processor; 311 - Computer program A; 400 - Computer-readable storage medium; 411 - Computer program B. Detailed Implementation
[0066] In the following, the terms “comprising” or “may include” as used in the various embodiments of this application indicate the presence of a function, operation, or element of the invention and do not limit the addition of one or more functions, operations, or elements. Furthermore, as used in the various embodiments of this application, the terms “comprising,” “having,” and their cognates are intended only to indicate a specific feature, number, step, operation, element, component, or combination of the foregoing and should not be construed as primarily excluding the presence of one or more other features, numbers, steps, operations, elements, components, or combinations of the foregoing, or adding one or more combinations of the foregoing.
[0067] In various embodiments of this application, the expression "or" or "at least one of A and / or B" includes any combination or all combinations of the words listed simultaneously. For example, the expression "A or B" or "at least one of A and / or B" may include A, may include B, or may include both A and B.
[0068] The terms used in the various embodiments of this application (such as "first," "second," etc.) may modify various constituent elements in the various embodiments, but do not limit the corresponding constituent elements. For example, the above terms do not limit the order and / or importance of the elements. The above terms are only used for the purpose of distinguishing one element from other elements. For example, a first user device and a second user device refer to different user devices, although both are user devices. For example, without departing from the scope of the various embodiments of this application, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element.
[0069] It should be noted that if a description is made of "connecting" one component to another, then the first component can be directly connected to the second component, and a third component can be "connected" between the first and second components. Conversely, when a component is "directly connected" to another component, it can be understood that there is no third component between the first and second components.
[0070] The terminology used in the various embodiments of this application is for the purpose of describing particular embodiments only and is not intended to limit the various embodiments of this application. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which the various embodiments of this application pertain. The terms (such as those defined in a generally used dictionary) are to be interpreted as having the same meaning as in the context of the relevant technical field and are not to be interpreted as having an idealized or overly formal meaning, unless clearly defined in the various embodiments of this application.
[0071] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the embodiments and accompanying drawings. The illustrative embodiments and descriptions of this application are only for explaining this application and are not intended to limit this application.
[0072] like Figure 1 As shown in the figure, this application proposes a method for detecting anomalies in rail transit equipment, including the following steps:
[0073] Step 1, Multi-source signal synchronous acquisition: Multiple industrial-grade broadband microphones are installed near each train door on the outside of the train body to continuously record audio as the train runs. At the same time, the real-time running speed of the train, the ID of the trackside beacon, the absolute mileage value, and the time when the train passes the beacon are synchronously acquired through the train control system, acceleration sensor, wheel speed sensor, trackside beacon, and other units.
[0074] Step 2, Audio preprocessing: Divide the acquired audio signal into short audio segments of a preset duration (e.g., 3-5 seconds) and record the start and end timestamps of each audio segment.
[0075] Step 3, Log-Mel spectrogram generation: For each audio segment, generate a log-Mel spectrogram with high fault characterization capability.
[0076] Furthermore, in step 3, the log-Melbourne spectrum generation process includes:
[0077] Audio pre-emphasis: Pre-emphasizes each audio segment to boost high-frequency components, specifically as follows:
[0078] ;
[0079] in, Indicates the first The amplitude value of the nth sampling point in an audio segment;
[0080] Indicates the first after pre-emphasis The output value of the nth sample point of an audio segment; This represents the pre-emphasis coefficient, which is usually taken as... Preferably, ; Indicates the first The amplitude value of the (n-1)th sample point in the audio segment.
[0081] Frame segmentation and windowing: The k-th audio segment after pre-emphasis... The data is divided into short frames, each with a length of N = 1024 samples, and the frame shift is H = 512 samples. A Hamming window is added to reduce spectral leakage.
[0082] ;
[0083] ;
[0084] in, Indicates the first The nth sample point (after windowing) of the mth frame in an audio segment; Represents the Hamming window function; This represents the value of the mH+nth sampling point in the pre-emphasized audio segment, where N is the frame length and H is the frame shift.
[0085] Short-time Fourier Transform: Calculate the N-point Discrete Fourier Transform for each short-time frame to obtain the complex spectrum. :
[0086] ;
[0087] Due to symmetry, the power spectrum is calculated using the first N / 2 + 1 frequency points. :
[0088] ;
[0089] Mel filter bank: Defines a set of F triangular filters, with the frequency axis mapped from the linear frequency domain to the Mel domain:
[0090] ;
[0091] in, This indicates the corresponding linear frequency. The Mel frequency, measured in Mel (Mel); This represents the linear frequency, and the unit is Hertz (Hz).
[0092] The filters are uniformly distributed in the Mel domain, exhibiting a dense distribution at low frequencies and a sparse distribution at high frequencies in the linear frequency domain. The transfer function of each filter for:
[0093] ;
[0094] in, Indicates the first The FFT index corresponding to the center frequency of the filter; Indicates the first The FFT index corresponding to the center frequency of each filter; Indicates the first The FFT index corresponding to the center frequency of the filter.
[0095] The power spectrum of each short frame is passed through a Mel filter bank to obtain the Mel spectral energy:
[0096] ;
[0097] in, This indicates that the power spectrum of each frame will be displayed. Through the first The Mel spectrum energy is obtained from a Mel filter. Preferably, F is 64 or 128.
[0098] Logarithmic compression: To simulate the human ear's perception of audio loudness and compress the dynamic range of strong and weak signals, the Mel spectrum energy is logarithmically calculated.
[0099] ;
[0100] in, It is a very small constant, used to prevent the overall calculation result from being 0.
[0101] Generating the log-Mel spectrum: Concatenate the log-Mel spectral vectors of all short-time frames column-wise to form a two-dimensional matrix, which is the log-Mel spectrum. :
[0102] ;
[0103] ;
[0104] in, ( ) indicates the first A short time frame; This refers to the number of short-term frames.
[0105] The log-Melb spectrum of the train equipment generated through the above process is as follows: Figure 2-5 As shown, where, Figure 2 The log-Melbourne spectrograms of the train during normal operation at different sampling times are shown. Figure 3 The log-Mel spectrum of the train doors at different sampling times is shown (the obvious voiceprint features generated by the train door opening and unlocking and the beeping prompt tone). Figure 4 The log-Mel spectrum of the train at different sampling times during departure from the station is shown (the obvious acoustic signature generated by the speed-up of the train traction converter and motor). Figure 5 The log-Mel spectrum of wheel-rail friction whistling at different sampling times is shown (the acoustic characteristics generated by the strong angular contact friction between the train wheel flange and the track).
[0106] Step 4, Two-stage cascaded audio classification using deep feature extraction and classification + backend binary classification: The first stage uses a deep feature extractor and classifier to map the log-Mel spectrum to a high-dimensional space for preliminary classification, focusing on accurately identifying the specific type of acoustic event under various noise interferences, such as classifying the extracted audio into specific acoustic events like "door opening / closing sound," "wheel-rail contact sound," and "bearing rotation sound." The second stage uses a binary classification model to classify acoustic events into "normal" or "abnormal" states, focusing on analyzing subtle anomalies within specific acoustic events. The two stages are linked by a model manager, responsible for distributing the preliminary classification results from the first stage to the second stage for precise classification. This decoupled two-stage cascaded classification architecture allows each model to achieve optimal performance on its respective task, avoiding problems such as insufficient accuracy, robustness, and performance degradation that may occur when using a single model.
[0107] Preferably, in step 4, ResNet-50 is used to construct a deep feature extractor and classifier, an SVM classifier is used as a binary classification model, and an SVM model manager is used. The specific implementation process is as follows:
[0108] Multi-event Feature Extraction and Preliminary Classification Based on ResNet-50: The log-Melogram is input into a ResNet-50-based deep convolutional neural network (i.e., the feature extraction layer). ResNet-50 addresses the feature degradation problem in deep neural networks through residual connections, effectively extracting multi-level time-frequency texture features from the log-Melogram. High-dimensional feature maps extracted by the ResNet-50 feature extraction layer (Conv1-Conv5). for:
[0109] ;
[0110] After the high-dimensional feature map is processed by a global average pooling layer, a fixed-length feature vector is obtained. The feature vector is then passed through a SoftMax layer for preliminary multi-event classification to obtain the acoustic event categories. The specific network structure of ResNet-50 is as follows: Figure 6 As shown. The training process includes:
[0111] Collect sufficient raw audio clips and background noise from the industrial-grade broadband microphones on the outside of the train body, classify and store them, and assign labels to each type of audio.
[0112] Following the methods described in steps 2 and 3, each type of raw audio data is converted into a log-Mel spectrogram as the original dataset.
[0113] The original dataset is divided into training, testing, and validation sets according to a preset ratio (e.g., 8:1:1) to train the ResNet-50 network. The training is iterated for 500-1000 epochs, and the model with the highest classification accuracy on the testing set is selected as the target classification model.
[0114] Acoustic event normal / abnormal classification based on SVM classifier: For each acoustic event after preliminary classification, the feature vector output by the global average pooling layer of ResNet-50 is taken. categorizing acoustic events and eigenvectors The data is fed into the SVM model manager, which then inputs the corresponding feature vectors according to the acoustic event category into the SVM classifier for normal / abnormal classification. The RBF kernel is used as the kernel function for the SVM classifier, and the decision function of the SVM classifier is as follows:
[0115] ;
[0116] in, Indicates the classification result; Represents a symbolic function; Indicates the number of support vectors; Let represent the Lagrange multiplier of the i-th support vector, obtained through training optimization; This represents the true class label of the i-th support vector; Represents the RBF kernel function; This represents the feature vector of the i-th support vector; This represents the feature vector obtained after the current input audio segment has been processed by the ResNet-50 network; This indicates the bias term.
[0117] The training process is as follows:
[0118] In the initial model training phase, the SVM classifier only needs feature vector samples when the device is running normally to complete the training, which is the One-Class SVM mode. Using the Single-Class SVM classifier, the model only needs to learn the boundary of "normal" sound. Any sample that deviates from this boundary will be marked as "abnormal". This can alleviate the problem that fault samples are very scarce and difficult to obtain in the rail transit scenario.
[0119] After accumulating enough fault samples, a more accurate SVM binary classification model, namely Binary SVM, can be learned using normal and abnormal samples, which can further improve the recognition accuracy of abnormal acoustic events.
[0120] The detailed process of two-stage cascaded audio classification based on ResNet-50 and SVM classifiers is as follows: Figure 7 As shown.
[0121] Step 5, Abnormal mileage localization based on train speed and trackside beacon fusion: Accurately locate audio segments identified as abnormal acoustic events.
[0122] Furthermore, in step 5, the abnormal mileage localization process includes:
[0123] Baseline Mileage Acquisition: Extracting the start timestamps of audio segments related to anomalous acoustic events Find the nearest trackside beacon reference signal and obtain the end timestamp of the audio segment. Let the known absolute mileage corresponding to the beacon be . For example, "Metro Line XX - [XX Station - XX Station] Downward Section K3+100m", and record the time when the train passes this beacon. .
[0124] Mileage integral calculation: Let the train's speed within the section be... The time when the abnormal acoustic event occurred was:
[0125] ;
[0126] The event's mileage offset relative to the beacon for:
[0127] ;
[0128] In actual calculations, the train speed is a discrete value that varies with time. Let's assume the speed from the beacon time... At the moment of the anomalous acoustic event A total of n+1 discrete velocity values were obtained. The corresponding time is And the acquisition frequency is fixed. That is, then for The calculation can be simplified to:
[0129] ;
[0130] Finally, based on the train's direction of travel (up or down), the absolute mileage of the abnormal acoustic event is obtained. :
[0131] ;
[0132] For example, in this embodiment of the application, the abnormality of "high-frequency whistling of wheel-rail contact" is identified by audio classification. Then, the absolute mileage of the abnormality is calculated by the above-mentioned abnormal mileage location to be "K3+200m in the downlink section of Metro Line XX-[XX Station-XX Station]". This location result will be written into the alarm information and uploaded to the maintenance scheme expert knowledge base built into the system.
[0133] Step 6, Automatic generation of abnormal alarms and maintenance suggestions: Based on the abnormal event type, the absolute mileage of the abnormal event, and the log-Mel spectrum of the abnormal audio segment, generate alarm information and synchronize the alarm information with relevant departments. If necessary, an audible and visual alarm mode can be used.
[0134] Furthermore, the anomaly detection method proposed in this application embodiment also includes:
[0135] The abnormal event type is input into the maintenance solution expert knowledge base, which automatically associates and generates corresponding maintenance suggestions, and then pushes these suggestions to relevant departments. This maintenance solution expert knowledge base pre-stores several different types of abnormal events and their corresponding maintenance suggestions, and can also be updated online during actual application.
[0136] Based on the same technical concept described above, this application also proposes an anomaly detection system for rail transit equipment, such as... Figure 8 As shown, the anomaly detection system 200 includes:
[0137] The synchronization acquisition module 201 is used to synchronously acquire multi-source signals, including audio signals, train speed, beacon mileage, and timestamps. The specific signal acquisition method is as described in step 1 above, and will not be repeated here.
[0138] The audio preprocessing module 202 is used to segment the acquired audio signal into audio segments of preset duration and record the start and end timestamps of each audio segment. The specific segmentation method is as described in step 2 above, and will not be repeated here.
[0139] The spectrogram generation module 203 is used to generate the log-Mel spectrogram corresponding to each audio segment. The specific method for generating the log-Mel spectrogram is as described in step 3 above, and will not be repeated here.
[0140] The second-level cascaded classification module 204 is used to input the log-Melogram of each audio segment into the first-level classifier for feature extraction and identification of the specific type of acoustic event. Different types of acoustic events and their corresponding features are then input into the corresponding second-level classifier for normal or abnormal classification. The specific second-level cascaded classification method is as described in step 4 above and will not be repeated here.
[0141] The multi-source fusion positioning module 205 is used to accurately locate audio segments identified as abnormal acoustic events based on train speed, beacon mileage, and timestamps. The specific anomaly localization method is as described in step 5 above and will not be repeated here.
[0142] Additionally, the anomaly handling module 206 includes an alarm unit. This alarm unit is used to generate alarm information and synchronize the alarm information with relevant departments by comprehensively considering the anomaly event type, anomaly location results, and the log-Mel spectrum of the anomaly audio segment. The specific anomaly handling method is as described in step 6 above, and will not be repeated here.
[0143] Furthermore, the anomaly handling module 206 in this embodiment also includes a built-in maintenance solution expert knowledge base, which can automatically associate the type of abnormal event with the corresponding maintenance suggestion: for example, "High-frequency wheel-rail contact whistling occurs at K3+200m in the downlink section of Metro Line XX - [XX Station - XX Station]" is associated with "It is recommended to dispatch maintenance personnel to apply wheel-rail friction modifier or arrange for a rail grinding vehicle to perform fixed-point operation"; "There is an abnormality in the opening and closing of the platform door on the platform of Metro Line XX - [XX Station]" is associated with "It is recommended to dispatch personnel to the platform to further inspect the platform door". This maintenance suggestion will be pushed and synchronized to the ground monitoring center, vehicle command center and other departments in real time, further improving the maintenance efficiency of rail transit equipment.
[0144] Based on the same technical concept described above, this application also proposes an electronic device, such as... Figure 9 As shown, the electronic device 300 includes: a memory 310, a processor 320, and a computer program A311 stored in the memory 310 and executable on the processor 320. When the processor 320 executes the computer program A311, it performs the following steps:
[0145] Simultaneously acquire signals from multiple sources, including audio signals, train speed, beacon mileage, and timestamps;
[0146] The acquired audio signal is divided into audio segments of preset duration, and the start and end timestamps of each audio segment are recorded.
[0147] Generate a log-Mel spectrogram for each audio segment;
[0148] The log-Mel spectrogram of each audio segment is input into the first-level classifier to extract features and identify the specific type of acoustic event. Different types of acoustic events and their corresponding features are input into the corresponding second-level classifier for normal or abnormal classification.
[0149] Based on train speed, beacon mileage, and timestamps, audio segments identified as anomalous acoustic events are precisely located.
[0150] Based on the combined abnormal event type, abnormal location results, and log-Mel spectrogram of the abnormal audio segment, alarm information is generated and synchronized with relevant departments.
[0151] Optionally, when the processor 320 executes the computer program A311, it can implement any of the embodiments in the corresponding examples of the above-described anomaly detection method.
[0152] It should be noted that the electronic device proposed in this application embodiment is a device used to implement the above-mentioned anomaly detection method. Therefore, based on the above-mentioned anomaly detection method proposed in this application embodiment, those skilled in the art can understand the specific implementation method and various variations of the electronic device in this application embodiment. Therefore, the specific implementation method of the above-mentioned anomaly detection method will not be described in detail here. Any electronic device used by those skilled in the art to implement the above-mentioned anomaly detection method is within the scope of protection of this application.
[0153] Based on the same technical concept described above, embodiments of this application also propose a computer-readable storage medium, such as... Figure 10 As shown, the computer-readable storage medium 400 stores a computer program B411, which, when executed by a processor, performs the following steps:
[0154] Simultaneously acquire signals from multiple sources, including audio signals, train speed, beacon mileage, and timestamps;
[0155] The acquired audio signal is divided into audio segments of preset duration, and the start and end timestamps of each audio segment are recorded.
[0156] Generate a log-Mel spectrogram for each audio segment;
[0157] The log-Mel spectrogram of each audio segment is input into the first-level classifier to extract features and identify the specific type of acoustic event. Different types of acoustic events and their corresponding features are input into the corresponding second-level classifier for normal or abnormal classification.
[0158] Based on train speed, beacon mileage, and timestamps, audio segments identified as anomalous acoustic events are precisely located.
[0159] Based on the combined abnormal event type, abnormal location results, and log-Mel spectrogram of the abnormal audio segment, alarm information is generated and synchronized with relevant departments.
[0160] Optionally, when the computer program B411 is executed by the processor, it can implement any of the embodiments corresponding to the above-described anomaly detection method.
[0161] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0162] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0163] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0164] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0165] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0166] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of this application. It should be understood that the above description is only a specific embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A rail transit equipment anomaly detection method, characterized in that, include: Simultaneously acquire signals from multiple sources, including audio signals, train speed, beacon mileage, and timestamps; The acquired audio signal is divided into audio segments of a preset duration, and the start and end timestamps of each audio segment are recorded. Generate a log-Mel spectrum corresponding to each audio segment; The log-Mel spectrum of each audio segment is input into the first-level classifier to extract features and identify the specific type of acoustic event. Different types of acoustic events and their corresponding features are input into the corresponding second-level classifier for normal or abnormal classification. Based on the train speed, beacon mileage, and timestamp, the audio segments identified as abnormal acoustic events are precisely located. Based on the combined abnormal event type, abnormal location results, and log-Mel spectrogram of the abnormal audio segment, an alarm report is generated and the alarm information is synchronized to relevant departments.
2. The rail transit equipment anomaly detection method of claim 1, wherein, Also includes: The abnormal event type is input into the maintenance solution expert knowledge base, which automatically associates and generates corresponding maintenance suggestions, which are then synchronized to relevant departments. 3.The rail transit equipment anomaly detection method of claim 1 or 2, characterized in that, The generation of the log-Mel spectrum corresponding to each audio segment includes: The audio segment is pre-emphasized; The pre-emphasized audio segment is divided into short time frames; Perform an N-point discrete Fourier transform on each short-time frame to obtain a complex spectrum, and calculate the power spectrum of the short-time frame by taking the complex spectrum of the first N / 2+1 frequency points. Construct a Mel filter bank; The power spectrum of each short-time frame is passed through the Mel filter bank to obtain the Mel spectrum energy, and the logarithm of the Mel spectrum energy is taken to obtain the logarithmic Mel spectrum vector of each short-time frame. The log-Mel spectrum vectors of all short-time frames are concatenated column by column to form a two-dimensional matrix, which is the log-Mel spectrum.
4. The rail transit equipment anomaly detection method of claim 1 or 2, wherein, The step of inputting the log-Mel spectrum of each audio segment into a first-level classifier for feature extraction and identification of the specific type of acoustic event includes: The log-Mel spectrum of each audio segment is input into the feature extraction layer to extract a high-dimensional feature map; The high-dimensional feature map is then subjected to global average pooling layer to obtain a feature vector of fixed length. The feature vectors are then used for preliminary multi-event classification through a SoftMax layer to obtain acoustic event categories.
5. The method for detecting anomalies in rail transit equipment according to claim 4, characterized in that, The process of inputting different types of acoustic events and their corresponding features into the corresponding second-level classifier for normal or abnormal classification includes: The acoustic event categories and feature vectors after preliminary classification are sent to the model manager; The model manager sends the corresponding feature vectors to the classifier of the corresponding acoustic event according to the acoustic event category for normal or abnormal classification.
6. The rail transit equipment anomaly detection method of claim 1 or 2, wherein, The method of accurately locating audio segments identified as anomalous acoustic events based on the train's operating speed, beacon mileage, and timestamp includes: Extract the beacon reference signal that is closest to the start timestamp of the audio segment of the abnormal acoustic event, obtain the end timestamp of the audio segment, obtain the absolute mileage corresponding to the beacon and the time when the train passes the beacon; The time when the abnormal acoustic event occurred was calculated based on the start and end timestamps of the audio segment. The train speed is integrated between the time the train passes the beacon and the time the abnormal acoustic event occurs to obtain the distance offset of the event relative to the beacon. Based on the train's direction of travel, the mileage offset of the event to the beacon is increased or decreased on the absolute mileage of the beacon to obtain the absolute mileage of the abnormal acoustic event.
7. A rail transit equipment abnormality detection system characterized by comprising: include: The synchronous acquisition module is used to synchronously acquire signals from multiple sources, including audio signals, train speed, beacon mileage, and timestamps. An audio preprocessing module is used to divide the acquired audio signal into audio segments of a preset duration and record the start and end timestamps of each audio segment. A spectrogram generation module is used to generate a log-Mel spectrogram corresponding to each audio segment; The second-level cascaded classification module is used to input the log-Mel spectrogram corresponding to each audio segment into the first-level classifier to extract features and identify the specific type of acoustic event. Different types of acoustic events and their corresponding features are input into the corresponding second-level classifier for normal or abnormal classification. The multi-source fusion positioning module is used to accurately locate audio segments that are determined to be abnormal acoustic events based on the train's operating speed, beacon mileage, and timestamp. In addition, an anomaly handling module, including an alarm unit, is used to generate alarm information and synchronize the alarm information to relevant departments by comprehensively considering the anomaly event type, anomaly location result, and log-Mel spectrogram of the anomaly audio segment.
8. The rail transit equipment anomaly detection system of claim 7, wherein, The exception handling module further includes: The built-in maintenance solution expert knowledge base is used to automatically associate abnormal event types with corresponding maintenance suggestions and push the maintenance suggestions to relevant departments simultaneously. 9.An electronic device comprising a memory and a processor, the memory storing a computer program, wherein, When the processor executes the computer program, it implements the rail transit equipment anomaly detection method according to any one of claims 1-6.
10. A computer-readable storage medium having stored thereon a computer program, characterized in that, When the computer program is executed by the processor, it implements the rail transit equipment anomaly detection method according to any one of claims 1-6.