Urban rail transit engineering structure health identification method based on train running audio

By deploying audio acquisition equipment in urban rail transit projects and combining edge computing and multi-model fusion architecture, non-intrusive, all-time, and high-precision structural health identification is achieved, solving the problems of strong invasiveness, high cost, incomplete coverage, and poor real-time performance of existing monitoring methods, and providing real-time early warning capabilities.

CN122135738APending Publication Date: 2026-06-02BEIJING URBAN CONSTR EXPLORATION & SURVEYING DESIGN RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING URBAN CONSTR EXPLORATION & SURVEYING DESIGN RES INST
Filing Date
2026-02-27
Publication Date
2026-06-02

Smart Images

  • Figure CN122135738A_ABST
    Figure CN122135738A_ABST
Patent Text Reader

Abstract

This invention discloses a method for structural health identification in urban rail transit engineering based on train operation audio. The method involves deploying non-invasive audio acquisition devices at fixed locations along the track and at key parts of the train. These devices collect multi-source audio signals generated during train operation. Standardized audio feature vectors are input into a pre-trained structural health identification model. Combined with train operation parameters and environmental parameters, acoustic pattern matching and multi-dimensional correlation analysis are used to output the structural health status identification result. This invention provides all-weather coverage, unaffected by nighttime, rain, snow, or other environmental conditions, effectively eliminating coverage blind spots of traditional monitoring methods. Real-time noise reduction, signal filtering, and feature extraction are performed through an edge computing module, combined with a lightweight data transmission mechanism, significantly improving monitoring real-time performance and achieving low end-to-end response latency, meeting the need for timely early warning of structural health anomalies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of technology, and in particular to a method for structural health identification of urban rail transit engineering based on train travel audio. Background Technology

[0002] The tracks, bridges, tunnels, ballast, fasteners and other components of urban rail transit engineering structures are subjected to train loads and environmental erosion over a long period of time, which can easily lead to defects such as track deformation, rail damage, bridge cracks, tunnel lining spalling, and ballast loosening. If these defects are not identified and dealt with in a timely manner, they may cause safety accidents such as train derailment and structural collapse, which seriously threaten operational safety and passenger safety.

[0003] Existing methods for monitoring the structural health of urban rail transit projects mainly fall into the following two categories:

[0004] 1. Contact Sensor Monitoring Method: This method involves installing contact devices such as strain gauges, displacement gauges, and accelerometers at key structural locations to collect physical parameters such as structural stress, displacement, and vibration, thereby analyzing the structural health status. This method has significant drawbacks: installation requires modification of the existing structure, resulting in complex and costly construction; sensors are susceptible to environmental interference (such as humidity and vibration), leading to failure and making maintenance difficult; and the limited number of monitoring points makes it difficult to achieve full coverage of the entire line and structure, resulting in monitoring blind spots.

[0005] 2. Visual Monitoring Method: This method utilizes machine vision technology to identify surface defects by deploying equipment such as cameras and lidar. However, this method is significantly affected by environmental conditions; for example, its accuracy drops drastically at night, in rainy or snowy weather, or in dark tunnel environments. It is also susceptible to obstruction leading to missed detections; the data transmission volume is enormous, requiring significant bandwidth and resulting in poor real-time performance; and it carries high privacy and data security risks, failing to comply with urban rail transit operation data standards.

[0006] Therefore, there is an urgent need to develop a non-invasive, all-time coverage, low-cost, high-precision, and real-time structural health identification method to address the pain points of existing monitoring technologies and provide technical support for the safe operation and maintenance of urban rail transit engineering structures. Summary of the Invention

[0007] To address the aforementioned technical problems, this invention provides a method for structural health identification in urban rail transit engineering based on train operation audio. The technical solution adopted is as follows:

[0008] A method for structural health identification of urban rail transit engineering based on train operation audio includes the following steps:

[0009] Step 1: Deploy non-invasive audio acquisition devices at fixed locations along the track and at key parts of the train. The non-invasive audio acquisition devices collect multi-source audio signals generated during the train's operation.

[0010] Step 2: The edge computing module built into the audio acquisition device performs real-time noise reduction, invalid signal filtering, and feature extraction on the original audio signal to generate a standardized audio feature vector.

[0011] Step 3: Input the standardized audio feature vector into the pre-trained structural health recognition model, combine it with train operation parameters and environmental parameters, and output the structural health status recognition result through acoustic pattern matching and multidimensional correlation analysis.

[0012] Step 4: Based on the identification results, construct a health status assessment system, output the risk level, and generate a structural health report.

[0013] Optionally, step 5 is also included: after receiving the early warning information, the backend operation and maintenance platform triggers the corresponding level of operation and maintenance response mechanism and links the train dispatching system and the operation and maintenance management system to achieve closed-loop handling.

[0014] Optionally, the non-invasive audio acquisition device is a plug-and-play IoT device, with fixed deployment locations along the track including overhead contact line supports, tunnel sidewalls, and bridge crash barriers; deployment locations at key train components include below the train bogies and at the bottom of the carriages; the acquired multi-source audio signals include wheel-rail contact audio, structural vibration audio, component friction audio, and environmental background audio.

[0015] Optionally, in step 2, the real-time noise reduction adopts an adaptive noise reduction algorithm, constructs a dedicated noise feature library based on typical urban rail noise, and decomposes, thresholds and reconstructs the original audio signal through wavelet transform to achieve targeted filtering of interference noise and retain effective audio signals related to structural health.

[0016] Optionally, in step 2, invalid signal filtering adopts a dual threshold screening mechanism based on signal strength and frequency, setting a strength threshold and an effective frequency range to eliminate irrelevant signals with insufficient strength or exceeding the effective frequency range, thereby reducing data processing redundancy.

[0017] Optionally, in step 2, feature extraction includes basic acoustic features and engineering structure-related features; basic acoustic features include Mel frequency cepstral coefficients and audio physical features, and audio physical features include differential features, audio spectrum flatness, and short-time energy.

[0018] The engineering structure-related features include wheel-rail impact frequency, structural resonance frequency, and audio pulse duration. The basic acoustic features and engineering structure-related features are fused to form a standardized audio feature vector.

[0019] Optionally, in step 3, the structural health identification model is a multi-model fusion architecture. The structural health identification model includes a convolutional neural network sub-model for extracting audio spatial features, a long short-term memory network sub-model for learning the temporal variation pattern of audio, and an attention mechanism for enhancing the sensitivity features of diseases. The preliminary identification result of health status is output through the collaborative operation of each sub-model.

[0020] Optionally, in step 3, the multidimensional correlation analysis uses a probabilistic model to integrate train operation parameters and environmental parameters, construct the correlation between audio features, operating status, and environmental conditions, and correct the preliminary identification results.

[0021] The acoustic pattern matching is based on a structural health state audio feature library, which includes a normal state feature subset, a potential risk feature subset, and an abnormal state feature subset.

[0022] The normal state feature subset collects audio signals of healthy structures under different vehicle models, different speeds from 30km / h to 80km / h, and different load conditions such as unloaded and fully loaded. After processing in step 2, the feature vector is constructed, and each feature vector is labeled as [F,S0].

[0023] The potential risk feature subset is constructed by collecting audio signals of early minor structural anomalies in the early stage of the structure, and after processing in step 2, each feature vector is labeled as [F,S1].

[0024] The abnormal state feature subset targets typical defects, including track deformation, rail damage, bridge cracks, tunnel lining spalling, track bed loosening, and fastener failure. Audio signals were collected through laboratory simulation and field measurement, and constructed after processing in step 2. Each feature vector is labeled as [F,S2].

[0025] The feature library supports incremental learning and updating, and the update formula is as follows: ,in For the existing feature library, Features of newly collected valid samples This is the updated feature library.

[0026] Optionally, in step 4, the health status assessment system uses the confidence level of the identification results and the degree of disease damage as core indicators, and the risk levels are classified as follows:

[0027] Level 1 emergency warning refers to the structural health status identification result being an abnormal fault S2 with a confidence level greater than or equal to 95%, corresponding to an immediate safety hazard. Immediate safety hazards include rail fractures, severe cracks in bridges, and large-scale spalling of tunnel lining.

[0028] Level 2 potential early warning refers to the structural health status identification result being potential risk S1, with a confidence level greater than or equal to 90%, corresponding to early anomalies. Early anomalies include track bed loosening, minor track deformation, and fastener failure.

[0029] Level 3 normal means that the structural health status identification result is normal S0, and the corresponding structure has no abnormal signals; the structural health report includes identification time, collection location, disease type, risk level, confidence level and feature matching details.

[0030] Optionally, in step 5, the operation and maintenance response mechanism and closed-loop handling process are as follows:

[0031] Level 1 Emergency Warning: The backend operation and maintenance platform immediately links with the train dispatching system to send speed limit or stop instructions to trains traveling on the affected section of the track. At the same time, it pushes an emergency response work order to the operation and maintenance emergency center. After the operation and maintenance personnel complete the response, they will send the results back to the platform to form a closed loop.

[0032] Level 2 potential warning: The backend operation and maintenance platform generates a targeted investigation work order and pushes it to the operation and maintenance management system. After the operation and maintenance personnel complete the investigation and rectification, they send the results back to the platform, and the platform updates the structural health status ledger.

[0033] In summary, the present invention has at least one of the following beneficial technical effects:

[0034] This invention provides a method for structural health identification of urban rail transit projects based on train operation audio. It enables non-intrusive monitoring without requiring modification of existing track structures, is easy to deploy and has low cost, and is adaptable to complex outdoor and onboard environments of urban rail transit, avoiding interference with existing operations.

[0035] It has all-day coverage capability, is not limited by environmental conditions such as nighttime, rain, or snow, and effectively eliminates the coverage blind spots of traditional monitoring methods.

[0036] The edge computing module performs noise reduction, signal filtering, and feature extraction in real time. Combined with a lightweight data transmission mechanism, it significantly improves the real-time performance of monitoring, with low end-to-end response latency, meeting the need for timely early warning of structural health anomalies.

[0037] By integrating basic acoustic features with engineering structural features, employing a multi-model fusion architecture and multi-dimensional correlation analysis, and combined with comprehensive feature library support, the accuracy of structural health status identification is significantly improved, effectively distinguishing different types and degrees of structural defects.

[0038] Establish a tiered early warning system and a closed-loop operation and maintenance response mechanism to trigger corresponding handling procedures for different risk levels, ensuring timely intervention for structural anomalies, improving operation and maintenance efficiency, and reducing safety risks.

[0039] The feature library supports incremental learning and updates, continuously incorporating new and effective samples to enhance the ability to identify new types of defects and adapt to the operational changes and defect evolution characteristics of urban rail transit engineering structures. Attached Figure Description

[0040] Figure 1 This is a flowchart illustrating the structural health identification method for urban rail transit engineering based on train driving audio, as per the present invention.

[0041] Figure 2 This is a schematic diagram of the first page of the structural health report of a specific embodiment;

[0042] Figure 3 This is a schematic diagram on the second page of the structural health report of a specific embodiment;

[0043] Figure 4 This is a schematic diagram on the third page of the structural health report of a specific embodiment;

[0044] Figure 5 This is a schematic diagram on page four of the structural health report of a specific embodiment. Detailed Implementation

[0045] The present invention will be further described in detail below with reference to the accompanying drawings.

[0046] This invention discloses a method for structural health identification of urban rail transit engineering based on train travel audio.

[0047] Reference Figures 1-5 Example 1, a method for structural health identification of urban rail transit engineering based on train operation audio, includes the following steps:

[0048] Step 1: Deploy non-invasive audio acquisition devices at fixed locations along the track and at key parts of the train. The non-invasive audio acquisition devices collect multi-source audio signals generated during the train's operation.

[0049] Step 2: The edge computing module built into the audio acquisition device performs real-time noise reduction, invalid signal filtering, and feature extraction on the original audio signal to generate a standardized audio feature vector.

[0050] Step 3: Input the standardized audio feature vector into the pre-trained structural health recognition model, combine it with train operation parameters and environmental parameters, and output the structural health status recognition result through acoustic pattern matching and multidimensional correlation analysis.

[0051] Step 4: Based on the identification results, construct a health status assessment system, output the risk level, and generate a structural health report.

[0052] Example 2 also includes step 5, whereby after receiving the early warning information, the backend operation and maintenance platform triggers the corresponding level of operation and maintenance response mechanism and links the train dispatching system and the operation and maintenance management system to achieve closed-loop processing.

[0053] By adopting the above technical solutions, the physical state of the urban rail transit engineering structure directly determines the generation characteristics of audio signals during the interaction between the structure and the train. Healthy structures exhibit stable acoustic laws in wheel-rail contact and vibration response, while abnormalities such as deformation, cracks, and loosening lead to changes in wheel-rail contact patterns, shifts in structural vibration frequencies, and variations in frictional audio intensity. By deploying audio acquisition equipment along the track and at key train locations in a dual-dimensional manner, both macroscopic audio signals from the entire line and microscopic audio signals from key local areas can be comprehensively captured, ensuring coverage of all critical structural regions and providing a comprehensive and complementary data source for condition identification.

[0054] The raw audio signal contains interference from environmental noise, irrelevant human sounds, and other factors that can mask valid signals related to the structural state. The edge computing module processes the raw signal locally in real time. On one hand, it uses an adaptive noise reduction algorithm to selectively filter typical interference in urban rail scenarios, retaining valid signals directly related to the structural state. On the other hand, it uses dual threshold filtering to eliminate invalid signals with insufficient intensity or mismatched frequencies, reducing data redundancy. Simultaneously, it extracts basic acoustic features and engineering structure-related features, transforming the raw audio signal into a standardized feature vector that accurately reflects the structural state. This ensures the effectiveness of the features while avoiding the latency and bandwidth consumption caused by the original data transmission.

[0055] The pre-trained multi-model fusion architecture possesses the ability to deeply mine feature information. The convolutional neural network sub-model can extract spatial features from feature vectors and capture the spectral differences corresponding to different defects. The long short-term memory network sub-model can learn the temporal variation patterns of audio features and adapt to the dynamic characteristics of signals during train operation. The attention mechanism can strengthen the weight of defect-sensitive features and improve the targeting of identification. Acoustic pattern matching initially locks down the structural health status by comparing the current feature vector with a feature library of known structural states. Train operating parameters affect the magnitude and frequency of wheel-rail forces, while environmental parameters change the mechanical properties and vibration response of the structure. Multidimensional correlation analysis that integrates these two types of parameters can correct the interference of the environment and operating status on the identification results, further improving the identification accuracy.

[0056] The confidence level of the identification results directly reflects the reliability of the judgment, while the severity of the damage determines the urgency level of the safety risk. Building an assessment system around these two elements allows for the scientific classification of risk levels and clarifies the priority of structural safety conditions. The health report, by integrating key information such as identification time, collection location, and risk level, transforms abstract feature analysis results into intuitive and usable safety data, providing a clear basis for subsequent handling and ensuring accurate control over structural safety conditions.

[0057] Different risk levels correspond to varying degrees of urgency in structural safety hazards, requiring differentiated handling procedures. Level 1 emergency warnings correspond to immediate safety risks; linking with the train dispatching system can quickly control the spread of risks and prevent accidents. Level 2 potential warnings correspond to early anomalies; pushing out investigation work orders enables early detection and rectification of hazards. By transmitting operation and maintenance results back to the platform, both the timeliness of handling is ensured, and practical data support is provided for the optimization of feature libraries and models, continuously improving the adaptability and reliability of the monitoring system.

[0058] Example 3: The non-invasive audio acquisition device is a plug-and-play IoT device. Fixed deployment locations along the track include contact wire supports, tunnel sidewalls, and bridge crash barriers. Deployment locations at key train components include below the train bogies and at the bottom of the carriages. The acquired multi-source audio signals include wheel-rail contact audio, structural vibration audio, component friction audio, and environmental background audio.

[0059] By adopting the above technical solutions, existing urban rail transit lines are in high demand, and it is necessary to avoid interference with existing structures and operations caused by the deployment of monitoring equipment. The plug-and-play nature of the equipment requires no modification to existing structures such as tracks and trains, and installation and commissioning can be completed quickly, significantly reducing deployment difficulty and cost. At the same time, IoT devices are lightweight, low-power, and easy to network, adapting to the dual needs of outdoor distributed deployment along the track and mobile deployment on trains, ensuring long-term stable operation of the equipment and providing hardware support for 24 / 7 uninterrupted monitoring.

[0060] The deployment of equipment along the track and at key train components forms a complementary monitoring system. The overhead contact line supports, tunnel sidewalls, and bridge crash barriers along the track are all located in stable and unobstructed areas, allowing for macroscopic capture of overall audio signals from the main structures of the track, tunnels, and bridges, achieving full line coverage. The area beneath the train bogies is the core region of wheel-rail contact, and the bottom of the carriage is adjacent to the track bed, fasteners, and other components. Deploying equipment in these locations allows for close-range capture of microscopic audio signals generated by wheel-rail interaction and component friction, accurately identifying subtle changes in the local structure. This dual-dimensional deployment ensures comprehensive monitoring while enhancing the sensitivity of local defect identification, avoiding blind spots inherent in single-deployment models.

[0061] Wheel-rail contact audio is directly related to track smoothness and rail integrity; structural vibration audio reflects the stiffness and stability of main structures such as bridges and tunnels; and component friction audio corresponds to the tightness of auxiliary structures such as fasteners and track bed. All three types of signals are directly related to structural health and are core data sources for identifying defects. Environmental background audio contains typical interference components of urban rail scenarios; collecting this type of signal provides a reference for subsequent noise reduction processing. By comparing and analyzing it with the core signals, accurate filtering of interference signals can be achieved, ensuring the purity of effective signals. The coordinated acquisition of these four types of signals ensures comprehensive capture of acoustic information related to structural health and provides necessary interference references for signal preprocessing, laying a data foundation for subsequent accurate identification.

[0062] In Example 4, step 2, the real-time noise reduction adopts an adaptive noise reduction algorithm. Based on the typical noise of urban rail transit, a dedicated noise feature library is constructed. The original audio signal is decomposed, thresholded, and reconstructed through wavelet transform to achieve targeted filtering of interference noise and retain effective audio signals related to structural health.

[0063] In Example 5, step 2, invalid signal filtering adopts a dual threshold screening mechanism based on signal strength and frequency. A strength threshold and an effective frequency range are set to eliminate irrelevant signals with insufficient strength or exceeding the effective frequency range, thereby reducing data processing redundancy.

[0064] In Example 6, step 2, feature extraction includes basic acoustic features and engineering structure correlation features; basic acoustic features include Mel frequency cepstral coefficients and audio physical features, and audio physical features include differential features, audio spectrum flatness, and short-time energy.

[0065] The engineering structure-related features include wheel-rail impact frequency, structural resonance frequency, and audio pulse duration. The basic acoustic features and engineering structure-related features are fused to form a standardized audio feature vector.

[0066] By adopting the above technical solution, specific typical interference noise exists in urban rail transit scenarios. This type of noise has inherent differences in frequency and temporal characteristics compared to effective signals related to structural health. A dedicated noise feature library constructed based on typical urban rail transit noise can accurately characterize the acoustic profiles of various interferences, providing a basis for targeted filtering. The adaptive noise reduction algorithm performs multi-scale decomposition of the original audio signal through wavelet transform, separating the signal into components in different frequency ranges, thus separating the effective signal from the noise components at different scales. Subsequently, threshold processing is used to selectively suppress components belonging to the noise feature library, retaining effective components related to structural vibration and wheel-rail contact. Finally, signal reconstruction restores the high-purity effective audio signal. This design avoids the generalized filtering defects of traditional noise reduction algorithms, ensuring that key acoustic information reflecting the structural state is not lost while eliminating interference.

[0067] Audio signals related to structural health (such as wheel-rail contact signals and structural vibration signals) exhibit relatively fixed distribution patterns in intensity and frequency range: signals with excessively low intensity are mostly environmental clutter or electronic noise from equipment, failing to reflect the true state of the structure; signals exceeding a specific frequency range are unrelated to physical processes such as wheel-rail interaction and structural vibration, constituting irrelevant interference. A dual-threshold screening mechanism based on signal intensity and frequency eliminates weak clutter by setting reasonable intensity thresholds and excludes irrelevant signals exceeding the defined effective frequency range. This dual screening precisely distinguishes between effective and invalid signals from two core dimensions, significantly reducing redundancy in subsequent data processing, improving feature extraction efficiency, and simultaneously preventing interference from invalid signals to effective features, ensuring the accuracy of feature extraction.

[0068] The physical properties of audio signals are deeply correlated with the structural health status, and different types of features reflect this correlation from different dimensions. Among basic acoustic features, Mel-frequency cepstral coefficients can accurately characterize the spectral envelope of audio, adapting to the laws of human auditory perception and effectively distinguishing audio spectral differences under different structural states. Differential features, audio spectral flatness, and short-time energy supplement acoustic information from the perspectives of dynamic changes, spectral distribution uniformity, and signal strength, comprehensively capturing the physical properties of audio. Engineering structural correlation features directly relate to structural mechanical response: wheel-rail impact frequency is directly related to track smoothness and rail integrity; structural resonance frequency reflects the stiffness and stability of the main structure such as bridges and tunnels; and audio pulse duration corresponds to the intensity and duration of component friction and impact. These three types of features establish a direct mapping relationship between audio signals and structural states. Fusing these two types of features to form a standardized audio feature vector not only preserves the basic physical information of the audio signal but also strengthens its correlation with structural health, giving the feature vector both universality and specificity, providing comprehensive and high-value input data for subsequent AI models to accurately identify structural health status.

[0069] In Example 7, in step 3, the structural health identification model is a multi-model fusion architecture. The structural health identification model includes a convolutional neural network sub-model for extracting audio spatial features, a long short-term memory network sub-model for learning the temporal variation law of audio, and an attention mechanism for enhancing the sensitivity features of diseases. The preliminary identification result of health status is output through the collaborative operation of each sub-model.

[0070] In Example 8, step 3, the multidimensional correlation analysis uses a probabilistic model to integrate train operation parameters and environmental parameters, construct the correlation between audio features, operating status, and environmental conditions, and correct the preliminary identification results;

[0071] The acoustic pattern matching is based on a structural health state audio feature library, which includes a normal state feature subset, a potential risk feature subset, and an abnormal state feature subset.

[0072] The normal state feature subset collects audio signals of healthy structures under different vehicle models, different speeds from 30km / h to 80km / h, and different load conditions such as unloaded and fully loaded. After processing in step 2, the feature vector is constructed, and each feature vector is labeled as [F,S0].

[0073] The potential risk feature subset is constructed by collecting audio signals of early minor structural anomalies in the early stage of the structure, and after processing in step 2, each feature vector is labeled as [F,S1].

[0074] The abnormal state feature subset targets typical defects, including track deformation, rail damage, bridge cracks, tunnel lining spalling, track bed loosening, and fastener failure. Audio signals were collected through laboratory simulation and field measurement, and constructed after processing in step 2. Each feature vector is labeled as [F,S2].

[0075] The feature library supports incremental learning and updating, and the update formula is as follows: ,in For the existing feature library, Features of newly collected valid samples This is the updated feature library.

[0076] By adopting the above technical solutions, the audio features related to structural health encompass both spatial distribution characteristics and temporal dynamic patterns. Furthermore, different features exhibit varying sensitivities to defects, making it difficult for a single model to comprehensively capture this complex information. The convolutional neural network sub-model possesses powerful local feature extraction capabilities, capable of mining spatial differences such as spectral distribution and intensity variations from standardized audio feature vectors, accurately distinguishing the acoustic profiles corresponding to different defects. The long short-term memory network sub-model excels at processing temporal data, tracking the continuous changes in audio features during train operation, adapting to the dynamic characteristics of wheel-rail interaction and structural vibration over time, and avoiding misjudgments of structural status based on features at a single moment. The attention mechanism, through dynamic weight allocation, strengthens features highly sensitive to structural defects (such as impact features corresponding to rail damage and resonance features corresponding to cracked structures) while weakening interference from irrelevant features, improving the targeting of identification. The collaborative computation of these three components covers both spatial and temporal feature dimensions while highlighting the weight of key information, achieving complementary advantages and ensuring the comprehensiveness and reliability of the preliminary health status identification results.

[0077] There is a fixed mapping relationship between structural health status and audio features. The core of building the feature library is to establish a clear correspondence between status and features. The normal state feature subset covers different vehicle models, speeds, and load conditions because these operating factors affect wheel-rail forces and structural vibration intensity. Comprehensive collection can avoid misjudgments caused by differences in operating conditions and provide a unified reference standard for health status. The potential risk feature subset focuses on minor early structural anomalies, capturing subtle feature changes in the initial stage of defects, providing data support for early detection of hidden dangers. The abnormal state feature subset targets various typical defects. Through laboratory simulation, the type and degree of defects can be accurately controlled. Combined with on-site measurements, the authenticity and practicality of the samples are ensured, achieving accurate matching of specific defects. The feature library supports incremental learning and updates because new defects or uncovered operating scenarios may appear in actual operation. By incorporating new effective samples, the coverage of the feature library can be continuously enriched, allowing the pattern matching capability to continuously improve with actual application and adapt to the evolution characteristics of structural defects and changes in operating scenarios.

[0078] Train operating parameters and environmental parameters indirectly affect the correspondence between audio features and structural states: travel speed and load determine wheel-rail contact pressure and impact frequency, thus altering the intensity and frequency distribution of audio signals; environmental factors such as temperature, humidity, and rainfall affect the mechanical properties of the structure (e.g., stiffness, tightness), causing fluctuations in audio features under the same health conditions. Probabilistic models can quantify these relationships. By integrating the probability distributions of audio features, operating parameters, and environmental conditions, the posterior probability of different health states can be calculated, correcting initial identification biases caused by differences in operation or environment. For example, the same audio features may correspond to different structural states under high-speed heavy-load and low-speed empty-load conditions. Combining parameter correlation analysis can eliminate interference from such scenarios, making the final identification result more closely match the actual health condition of the structure and significantly improving identification accuracy.

[0079] In Example 9, step 4, the health status assessment system uses the confidence level of the identification results and the degree of disease damage as core indicators, and the risk levels are classified as follows:

[0080] Level 1 emergency warning refers to the structural health status identification result being an abnormal fault S2 with a confidence level greater than or equal to 95%, corresponding to an immediate safety hazard. Immediate safety hazards include rail fractures, severe cracks in bridges, and large-scale spalling of tunnel lining.

[0081] Level 2 potential early warning refers to the structural health status identification result being potential risk S1, with a confidence level greater than or equal to 90%, corresponding to early anomalies. Early anomalies include track bed loosening, minor track deformation, and fastener failure.

[0082] Level 3 normal means that the structural health status identification result is normal S0, and the corresponding structure has no abnormal signals; the structural health report includes identification time, collection location, disease type, risk level, confidence level and feature matching details.

[0083] In Example 10, step 5, the operation and maintenance response mechanism and closed-loop handling process are as follows:

[0084] Level 1 Emergency Warning: The backend operation and maintenance platform immediately links with the train dispatching system to send speed limit or stop instructions to trains traveling on the affected section of the track. At the same time, it pushes an emergency response work order to the operation and maintenance emergency center. After the operation and maintenance personnel complete the response, they will send the results back to the platform to form a closed loop.

[0085] Level 2 potential warning: The backend operation and maintenance platform generates a targeted investigation work order and pushes it to the operation and maintenance management system. After the operation and maintenance personnel complete the investigation and rectification, they send the results back to the platform, and the platform updates the structural health status ledger.

[0086] By adopting the above technical solutions, the audio features related to structural health encompass both spatial distribution characteristics and temporal dynamic patterns. Furthermore, different features exhibit varying sensitivities to defects, making it difficult for a single model to comprehensively capture this complex information. The convolutional neural network sub-model possesses powerful local feature extraction capabilities, capable of mining spatial differences such as spectral distribution and intensity variations from standardized audio feature vectors, accurately distinguishing the acoustic profiles corresponding to different defects. The long short-term memory network sub-model excels at processing temporal data, tracking the continuous changes in audio features during train operation, adapting to the dynamic characteristics of wheel-rail interaction and structural vibration over time, and avoiding misjudgments of structural status based on features at a single moment. The attention mechanism, through dynamic weight allocation, strengthens features highly sensitive to structural defects, such as impact features corresponding to rail damage and resonance features corresponding to cracked structures, while weakening interference from irrelevant features, thus improving the targeting of identification. The collaborative computation of these three components covers both spatial and temporal feature dimensions, while highlighting the weight of key information, achieving complementary advantages and ensuring the comprehensiveness and reliability of the preliminary health status identification results.

[0087] There is a fixed mapping relationship between structural health status and audio features. The core of building the feature library is to establish a clear correspondence between status and features. The normal state feature subset covers different vehicle models, speeds, and load conditions because these operating factors affect wheel-rail forces and structural vibration intensity. Comprehensive collection can avoid misjudgments caused by differences in operating conditions and provide a unified reference standard for health status. The potential risk feature subset focuses on minor early structural anomalies, capturing subtle feature changes in the initial stage of defects, providing data support for early detection of hidden dangers. The abnormal state feature subset targets various typical defects. Through laboratory simulation, the type and degree of defects can be accurately controlled. Combined with on-site measurements, the authenticity and practicality of the samples are ensured, achieving accurate matching of specific defects. The feature library supports incremental learning and updates because new defects or uncovered operating scenarios may appear in actual operation. By incorporating new effective samples, the coverage of the feature library can be continuously enriched, allowing the pattern matching capability to continuously improve with actual application and adapt to the evolution characteristics of structural defects and changes in operating scenarios.

[0088] Train operating parameters and environmental parameters indirectly affect the correspondence between audio features and structural states: travel speed and load determine wheel-rail contact pressure and impact frequency, thus altering the intensity and frequency distribution of audio signals; environmental factors such as temperature, humidity, and rainfall affect the mechanical properties of structures, such as stiffness and tightness, leading to fluctuations in audio features under the same health conditions. Probabilistic models can quantify these relationships. By integrating the probability distributions of audio features, operating parameters, and environmental conditions, the posterior probability of different health states can be calculated, correcting initial identification biases caused by differences in operation or environment. For example, the same audio features may correspond to different structural states under high-speed heavy-load and low-speed empty-load conditions. Combining parameter correlation analysis can eliminate interference from such scenarios, making the final identification result more closely match the actual health condition of the structure and significantly improving identification accuracy.

[0089] The following describes the implementation principle of the present invention using specific embodiments:

[0090] Taking a certain city's Metro Line 1 as an application scenario, the line is 35 kilometers long, including 20 kilometers of underground tunnels, 8 kilometers of surface track, and 7 kilometers of elevated bridges. It operates with both Type A and Type B trains, with a daily operating speed of 30 to 80 kilometers per hour and an average daily operating time of 18 hours. It requires 24 / 7 health monitoring of core engineering structures such as tracks, bridges, tunnels, track bed, and fasteners. The specific implementation is as follows:

[0091] Step 1: Audio signal acquisition

[0092] Non-invasive audio acquisition devices are deployed at 50-meter intervals along the track, on the catenary supports, tunnel sidewalls, and bridge crash barriers. One device is installed under the bogie and at the bottom of each train car. The devices continuously capture wheel-rail contact audio, structural vibration audio, component friction audio, and environmental background audio generated during train operation. The acquisition process is uninterrupted, ensuring full coverage of the line's operating and shutdown periods.

[0093] Step 2: Signal preprocessing

[0094] After the built-in edge computing module of the audio acquisition device is activated, it first performs real-time noise reduction on the original audio signal, then filters out invalid signals, and finally extracts features. The noise reduction process targets and filters typical interferences common in urban rail scenarios, such as traffic noise, crowd noise, and wind and rain noise; the filtering stage removes irrelevant signals with insufficient intensity or exceeding a specific frequency range; the extraction stage simultaneously acquires basic acoustic features and engineering structure-related features, and finally integrates them to form a standardized audio feature vector. The entire preprocessing process is completed locally on the device, without storing the original audio data.

[0095] Step 3: Health Status Identification

[0096] Standardized audio feature vectors are input into a pre-trained structural health recognition model, along with operational parameters such as the train's current speed and load, as well as environmental parameters such as ambient temperature, humidity, and rainfall. The model compares the current feature vectors with a pre-set feature library through acoustic pattern matching, and then corrects for the influence of operational and environmental parameters through multi-dimensional correlation analysis. Finally, it outputs the structural health status recognition results, which are categorized into three types: normal, potential risk, and abnormal fault.

[0097] Step 4: Health Assessment and Report Generation

[0098] Based on the confidence level of the identification results and the severity of the corresponding diseases, a health status assessment system is constructed to classify risk levels. Simultaneously, a structural health report is generated, clearly recording the identification time, collection location, disease type, risk level, confidence level, and feature matching details, providing complete data support for subsequent treatment.

[0099] Step 5: After receiving the early warning information through an encrypted transmission channel, the backend operation and maintenance platform triggers the corresponding operation and maintenance response mechanism according to the risk level. The platform achieves data interoperability with the train dispatching system and operation and maintenance management system of the subway line to ensure that early warning information can be quickly transferred and the handling results can be promptly fed back, forming a closed-loop management process of "identification-early warning-handling-feedback".

[0100] A plug-and-play IoT audio acquisition device was selected. This device has a sampling rate of 48kHz, a frequency response range of 20Hz to 20kHz, and is waterproof and dustproof, which can withstand the corrosion of the complex outdoor environment of urban rail transit. It also has anti-electromagnetic interference capabilities and can adapt to the electromagnetic environment on the train.

[0101] The equipment along the track is strictly installed in stable and unobstructed positions such as the crossbeams of the contact wire support, the middle of the tunnel sidewall, and the inside of the bridge anti-collision guardrail, to ensure that the acquisition view is not obstructed by debris around the track or the train body; the train-mounted equipment is installed under the bogie near the wheel-rail contact point and in the area corresponding to the track bed at the bottom of the carriage through special fixed brackets to capture the audio signal of the core part at close range.

[0102] During the data acquisition process, the equipment focuses on capturing the impact sound generated when the wheel and rail come into contact, the resonance sound generated by structural vibration, and the friction sound generated by the friction between the fastener and the rail. At the same time, it records background audio such as traffic noise, crowd noise, and wind and rain sounds in the environment, providing complete interference reference data for subsequent noise reduction processing.

[0103] A typical noise feature database for urban rail transit was constructed, containing acoustic feature data for 12 common types of interference, including traffic noise, crowd noise, wind and rain noise, and noise from the train traction power supply system. An adaptive noise reduction algorithm was adopted, using a db4 wavelet basis to perform a 5-level multi-scale decomposition of the original audio signal, separating the signal into components in different frequency ranges.

[0104] The median of the high-frequency decomposition coefficients of the first layer is calculated and then divided by 0.6745 to obtain the noise standard deviation. Based on this standard deviation and the signal length, the adaptive threshold of the decomposition coefficients of each layer is calculated. The components belonging to the noise feature library are suppressed by the soft thresholding method, and the effective components related to structural health are retained. Finally, the high-purity effective audio signal is reconstructed by wavelet inverse transform.

[0105] The signal strength threshold is set to 0.05Vpp, and the effective frequency range is 20Hz to 20kHz. The edge computing module performs frame-by-frame detection on the preprocessed audio signal, removing weak noise with peak intensity below 0.05Vpp and filtering irrelevant signals with frequencies exceeding the 20Hz to 20kHz range, retaining only the effective signals that meet both the strength and frequency requirements, thus reducing the amount of data required for subsequent feature extraction and model calculation.

[0106] Basic acoustic feature extraction:

[0107] The effective audio signal is divided into frames with a frame length of 25 milliseconds and a frame shift of 10 milliseconds. Each frame is subjected to FFT transformation after being windowed by a Hanning window. After being processed by a 24-dimensional Mel filter bank, logarithmic operation is performed, and then DCT transformation is performed to extract the first 13-dimensional Mel frequency cepstral coefficients. Based on these 13-dimensional coefficients, the first-order difference and the second-order difference are calculated to obtain 26-dimensional difference features. At the same time, the audio spectrum flatness and short-time energy of each frame are calculated to complete the basic acoustic feature extraction.

[0108] Engineering structure association feature extraction:

[0109] The signal power spectrum was calculated using the Welch method with a window length of 512 and an overlap rate of 50%, and the peak frequency in the 50-500Hz band was extracted as the wheel-rail impact frequency. Characteristic frequencies in the 10-200Hz band were extracted as structural resonance frequencies through modal analysis. The duration of the pulse signal generated by wheel-rail contact from the initial moment to the point of decay to 10% of the peak value was recorded as the audio pulse duration. The basic acoustic features and engineering structural features were integrated to form a standardized audio feature vector.

[0110] The structural health recognition model adopts a multi-model fusion architecture, in which the CNN sub-model contains two convolutional layers and two max pooling layers. The first convolutional layer sets 32 filters of size 3×1, uses the ReLU activation function, and has a stride of 1. The second convolutional layer sets 64 filters of size 3×1, also using the ReLU activation function, and has a stride of 1. Both max pooling layers use pooling kernels of size 2×1 and a stride of 2.

[0111] The LSTM sub-model has 128 hidden layer units, a dropout rate of 0.3, and does not return sequence output. The attention mechanism calculates the scores of each dimension of features using a 32×128 dimensional weight matrix and a 32-dimensional bias vector. After normalization, the feature weights are obtained, and then the temporal feature vectors output by the LSTM are weighted and summed to obtain the enhanced feature vector.

[0112] The enhanced feature vector is input into the fully connected layer. The first fully connected layer is 64-dimensional and uses the ReLU activation function. The second fully connected layer is 3-dimensional and uses the Softmax activation function. The output is a preliminary identification result of three health states: normal, potential risk, and abnormal failure.

[0113] Feature library construction and updates:

[0114] The structural health status audio feature library includes a normal state feature subset, a potential risk feature subset, and an abnormal state feature subset. The normal state feature subset is constructed by collecting audio signals from healthy structures of Type A and Type B vehicles at different speeds (30-80 km / h) and under different load conditions (empty and fully loaded). After preprocessing in step 2, it contains 100,000 feature vectors, each labeled as a normal state identifier. The potential risk feature subset collects audio signals of early-stage minor structural anomalies, and after preprocessing, it contains 50,000 feature vectors, each labeled as a potential risk identifier. The abnormal state feature subset targets typical defects such as track deformation, rail damage, bridge cracks, tunnel lining spalling, track bed loosening, and fastener failure. Audio signals are collected through laboratory simulations of different defect degrees and on-site measurements of known defect scenarios. After preprocessing, it contains 80,000 feature vectors, each labeled as an abnormal fault identifier.

[0115] The feature library is updated incrementally on a monthly basis, selecting new and valid sample features with a signal-to-noise ratio of not less than 20dB and a duration of not less than 3 seconds, and incorporating them into the original feature library. The coverage of the updated feature library continues to expand.

[0116] Multidimensional correlation analysis employs a Bayesian network probabilistic model, integrating operational parameters such as train speed and load with environmental parameters such as ambient temperature, humidity, and rainfall to construct the correlation between audio features, operational status, and environmental conditions. The model is trained using historical data to obtain the conditional probabilities and prior probabilities of each parameter. After inputting the current audio feature vector and real-time parameters, the posterior probabilities of different health states are calculated, correcting the initial identification results and improving recognition accuracy.

[0117] The health status assessment system uses the confidence level of the identification results and the degree of damage as core indicators. When the identification result is an abnormal fault and the confidence level reaches 95% or above, it is judged as a Level 1 emergency warning, corresponding to immediate safety hazards such as rail breakage, severe bridge cracks, and large-scale spalling of tunnel lining; when the identification result is a potential risk and the confidence level reaches 90% or above, it is judged as a Level 2 potential warning, corresponding to early anomalies such as track bed loosening, slight track deformation, and fastener failure; when the identification result is normal, or the confidence level is lower than the corresponding threshold mentioned above, it is judged as a Level 3 normal, corresponding to no abnormal signals in the structure.

[0118] The structural health report is automatically generated once per hour. The report clearly records the identification time, the specific mileage marker of the collection location, the suspected disease type, the risk level, the confidence level of the identification result, and the matching details of the features with the feature library. The report is stored in a structured format on the backend platform, allowing maintenance personnel to access and view it at any time.

[0119] Level 1 Emergency Warning and Response:

[0120] Upon receiving a Level 1 emergency alert, the backend operations and maintenance platform, within 30 seconds, activates the train dispatching system to send speed limit instructions to trains traveling within a 3-kilometer radius of the affected section, reducing their speed to below 20 kilometers per hour. If the hazard is severe, a train stop order is sent. Simultaneously, an emergency response work order is pushed to the operations and maintenance emergency center, specifying the suspected location of the hazard, the risk level, the priority of response, and emergency response recommendations. After operations and maintenance personnel arrive at the scene and complete the response, they use a mobile operations and maintenance app to transmit the response results, on-site photos, and structural retest data back to the backend platform. Once the platform confirms that the hazard has been eliminated, it lifts the alert and records the entire response process, forming a closed loop.

[0121] Level 2 potential early warning response:

[0122] Upon receiving a Level 2 potential alert, the backend operations and maintenance platform automatically generates a targeted investigation work order. This work order includes the mileage marker of the suspected fault location, key areas to investigate, suggested investigation methods, and a completion deadline of 72 hours. After the work order is pushed to the operations and maintenance management system, it is assigned to the corresponding regional operations and maintenance team. Once the personnel complete the investigation and rectification, they send the investigation results and rectification records back to the backend platform. The platform then updates the structural health status ledger, tracks the rectification effectiveness, and ensures that potential hazards are eliminated promptly.

[0123] The above are all preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Therefore, all equivalent changes made in accordance with the structure, shape and principle of the present invention should be covered within the scope of protection of the present invention.

Claims

1. A method for structural health identification of urban rail transit engineering based on train operation audio, characterized in that, Includes the following steps: Step 1: Deploy non-invasive audio acquisition devices at fixed locations along the track and at key parts of the train. The non-invasive audio acquisition devices collect multi-source audio signals generated during the train's operation. Step 2: The edge computing module built into the audio acquisition device performs real-time noise reduction, invalid signal filtering, and feature extraction on the original audio signal to generate a standardized audio feature vector. Step 3: Input the standardized audio feature vector into the pre-trained structural health recognition model, combine it with train operation parameters and environmental parameters, and output the structural health status recognition result through acoustic pattern matching and multidimensional correlation analysis. Step 4: Based on the identification results, construct a health status assessment system, output the risk level, and generate a structural health report.

2. The method for structural health identification of urban rail transit engineering based on train operation audio according to claim 1, characterized in that, It also includes step 5, where the backend operation and maintenance platform receives the early warning information, triggers the corresponding level of operation and maintenance response mechanism, and links the train dispatching system and the operation and maintenance management system to achieve closed-loop processing.

3. The method for structural health identification of urban rail transit engineering based on train operation audio according to claim 2, characterized in that, The non-invasive audio acquisition device is a plug-and-play IoT device. Fixed deployment locations along the track include contact wire supports, tunnel sidewalls, and bridge crash barriers. Deployment locations at key train components include under the train bogies and at the bottom of the carriages. The acquired multi-source audio signals include wheel-rail contact audio, structural vibration audio, component friction audio, and environmental background audio.

4. The method for structural health identification of urban rail transit engineering based on train operation audio according to claim 3, characterized in that, In step 2, the real-time noise reduction adopts an adaptive noise reduction algorithm. Based on the typical noise of urban rail transit, a dedicated noise feature library is constructed. The original audio signal is decomposed, thresholded and reconstructed through wavelet transform to achieve targeted filtering of interference noise and retain effective audio signals related to structural health.

5. The method for structural health identification of urban rail transit engineering based on train operation audio according to claim 4, characterized in that, In step 2, invalid signal filtering adopts a dual threshold screening mechanism based on signal strength and frequency. It sets a strength threshold and an effective frequency range to eliminate irrelevant signals with insufficient strength or exceeding the effective frequency range, thereby reducing data processing redundancy.

6. The method for structural health identification of urban rail transit engineering based on train operation audio according to claim 5, characterized in that, In step 2, feature extraction includes basic acoustic features and engineering structure correlation features; basic acoustic features include Mel frequency cepstral coefficients and audio physical features, and audio physical features include differential features, audio spectrum flatness, and short-time energy. The engineering structure-related features include wheel-rail impact frequency, structural resonance frequency, and audio pulse duration. The basic acoustic features and engineering structure-related features are fused to form a standardized audio feature vector.

7. The method for structural health identification of urban rail transit engineering based on train operation audio according to claim 6, characterized in that, In step 3, the structural health identification model is a multi-model fusion architecture. The structural health identification model includes a convolutional neural network sub-model for extracting audio spatial features, a long short-term memory network sub-model for learning the temporal variation pattern of audio, and an attention mechanism for enhancing the sensitivity features of diseases. The preliminary identification results of health status are output through the collaborative operation of each sub-model.

8. The method for structural health identification of urban rail transit engineering based on train operation audio according to claim 7, characterized in that, In step 3, the multidimensional correlation analysis uses a probabilistic model to integrate train operation parameters and environmental parameters, construct the correlation between audio features, operating status, and environmental conditions, and correct the preliminary identification results. The acoustic pattern matching is based on a structural health state audio feature library, which includes a normal state feature subset, a potential risk feature subset, and an abnormal state feature subset. The normal state feature subset collects audio signals of healthy structures under different vehicle models, different speeds from 30km / h to 80km / h, and different load conditions such as unloaded and fully loaded. After processing in step 2, the feature vector is constructed, and each feature vector is labeled as [F,S0]. The potential risk feature subset is constructed by collecting audio signals of early minor structural anomalies in the early stage of the structure, and after processing in step 2, each feature vector is labeled as [F,S1]. The abnormal state feature subset targets typical defects, including track deformation, rail damage, bridge cracks, tunnel lining spalling, track bed loosening, and fastener failure. Audio signals were collected through laboratory simulation and field measurement, and constructed after processing in step 2. Each feature vector is labeled as [F,S2]. The feature library supports incremental learning and updating, and the update formula is as follows: ,in For the existing feature library, Features of newly collected valid samples This is the updated feature library.

9. The method for structural health identification of urban rail transit engineering based on train operation audio according to claim 8, characterized in that, In step 4, the health status assessment system uses the confidence level of the identification results and the degree of disease damage as core indicators, and the risk levels are classified as follows: Level 1 emergency warning refers to the structural health status identification result being an abnormal fault S2 with a confidence level greater than or equal to 95%, corresponding to an immediate safety hazard. Immediate safety hazards include rail fractures, severe cracks in bridges, and large-scale spalling of tunnel lining. Level 2 potential early warning refers to the structural health status identification result being potential risk S1, with a confidence level greater than or equal to 90%, corresponding to early anomalies. Early anomalies include track bed loosening, minor track deformation, and fastener failure. Level 3 normal means that the structural health status identification result is normal S0, and the corresponding structure has no abnormal signals; the structural health report includes identification time, collection location, disease type, risk level, confidence level and feature matching details.

10. The method for structural health identification of urban rail transit engineering based on train operation audio according to claim 9, characterized in that, In step 5, the operation and maintenance response mechanism and closed-loop handling process are as follows: Level 1 Emergency Warning: The backend operation and maintenance platform immediately links with the train dispatching system to send speed limit or stop instructions to trains traveling on the affected section of the track. At the same time, it pushes an emergency response work order to the operation and maintenance emergency center. After the operation and maintenance personnel complete the response, they will send the results back to the platform to form a closed loop. Level 2 potential warning: The backend operation and maintenance platform generates a targeted investigation work order and pushes it to the operation and maintenance management system. After the operation and maintenance personnel complete the investigation and rectification, they send the results back to the platform, and the platform updates the structural health status ledger.