Multi-section rail fault collaborative diagnosis method and system based on federated learning
By adopting a multi-segment rail fault collaborative diagnosis method based on federated learning, the problems of data silos, insufficient cross-domain adaptation, and weak anti-interference ability in multi-segment collaborative diagnosis are solved. This method achieves improved data privacy and security, diagnostic accuracy and real-time performance, enhanced adaptability, and meets the safety operation and maintenance needs of large-scale railway networks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHENGZHOU RAILWAY VOCATIONAL & TECH COLLEGE
- Filing Date
- 2026-02-24
- Publication Date
- 2026-05-12
AI Technical Summary
Existing railway fault diagnosis technologies suffer from problems such as data silos, insufficient cross-domain adaptability, difficulty in balancing security and real-time performance, and weak anti-interference capabilities in multi-segment collaborative diagnosis scenarios, making it difficult to meet the safety operation and maintenance needs of large-scale railway networks.
A collaborative fault diagnosis method for multiple railway segments based on federated learning is adopted. Heterogeneous monitoring data is collected and encrypted at the segment client, preprocessed and labeled to construct a local fault feature dataset. The weighted aggregation strategy and robust encrypted transmission of the central server are used to optimize the global fault diagnosis module and update the local model. Combined with multi-scale feature extraction and cross-segment feature fusion, the diagnostic accuracy and anti-interference ability are improved.
It has improved the data privacy and security protection capabilities of multi-segment data, increased the accuracy of cross-domain fault diagnosis, enhanced the model generalization ability, and improved the anti-interference and fault tolerance performance in the collaborative training process. It has balanced the real-time diagnosis and the adaptability for large-scale promotion, and solved the problems of data silos, insufficient cross-domain adaptation and weak anti-interference capabilities.
Smart Images

Figure CN122020484A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of railway track safety monitoring technology, and in particular to a collaborative diagnosis method and system for multi-segment railway track faults based on federated learning. Background Technology
[0002] As the railway industry transitions from a large-scale "design and construction" phase to a phase emphasizing both design, construction, operation, and maintenance, railway mileage and train speeds continue to increase, along with the ever-growing volume of heavy-haul freight. Under the repeated rolling of train wheels, alternating temperature changes, and complex geological environments, railway tracks are prone to various faults such as unevenness and minor deformations. If these faults are not diagnosed promptly and accurately, these small hidden dangers can gradually accumulate and evolve into major safety accidents, directly threatening train operation safety and the lives and property of passengers. Industry forecasts indicate that by 2025, the scale of railway maintenance equipment will reach 30.102 billion yuan, with a compound annual growth rate of 11.64%. As a core support for the railway operation and maintenance system, the demand for track fault diagnosis technology—its accuracy, real-time performance, and collaborative capabilities—is becoming increasingly urgent.
[0003] To address the challenge of railway fault diagnosis, existing technologies have gradually evolved from traditional manual inspection and periodic track inspection vehicle checks to data-driven, AI-assisted online diagnosis. For example, existing technologies have proposed machine vision-based track condition monitoring methods, achieving preliminary identification of track faults through visual data acquisition and analysis. Simultaneously, diagnostic methods based on data-driven approaches and expert knowledge fusion, as well as online diagnostic methods utilizing knowledge transfer learning, are being gradually applied. By constructing a mapping relationship between fault features and diagnostic results, and combining techniques such as reinforcement learning and average impact analysis, the sensitivity and real-time performance of small-sample fault diagnosis are optimized. Some methods can already achieve millimeter-level fault feature capture and graded early warning, improving diagnostic efficiency and accuracy to a certain extent. Furthermore, the deployment of the world's first intelligent track diagnostic analysis system has broken through the bottleneck of fragmented data through a dynamic programming flexible matching algorithm, achieving orderly association and rapid screening of data from multiple time periods, and promoting the upgrade of track maintenance towards a "real-time monitoring - intelligent early warning - precise handling" model.
[0004] However, existing track fault diagnosis technologies still have many bottlenecks in multi-segment collaborative diagnosis scenarios, making it difficult to meet the safety operation and maintenance needs of large-scale railway networks: First, the problem of data silos is prominent, and the foundation for collaborative diagnosis is weak. Railway track data is usually stored in different railway bureaus, engineering sections, and operation and maintenance units. It covers sensitive information such as track vibration signals, deformation data, and environmental parameters. Due to data privacy protection regulations, industry security standards, and departmental management boundaries, this data is prohibited from being transferred or centrally aggregated across units, forming isolated data silos. Existing diagnostic methods are mostly based on training models on a single track segment or a single data source, which cannot integrate the complementary value of heterogeneous data from multiple track segments. This results in limited model generalization ability and difficulty in adapting to the needs of track fault diagnosis under different regions and operating conditions (such as climate differences, geological conditions, and traffic volume levels).
[0005] Secondly, cross-domain adaptability is insufficient, and diagnostic accuracy is limited by differences in data distribution. While existing transfer learning methods can reduce model bias between the source domain (such as the track inspection vehicle environment) and the target domain (such as the actual operating environment) to some extent, for multi-segment scenarios, the rail materials, wear levels, operating years, and monitoring equipment types vary from segment to segment, resulting in significant heterogeneity in fault data distribution. A single transfer model cannot simultaneously adapt to the data characteristics of multiple segments, easily leading to fluctuations in diagnostic accuracy, missed diagnoses, and misdiagnoses. Furthermore, some segments suffer from a scarcity of high-level fault samples. Existing small-sample reinforcement learning methods can only optimize for a single segment and cannot leverage similar fault data from other segments to enhance the model's diagnostic capabilities.
[0006] Third, data transmission and security risks coexist, making it difficult to balance real-time performance and security. Traditional centralized diagnostic models require transmitting monitoring data from each road segment to a central server for training and analysis. This not only faces bandwidth pressure from massive data transmission, reducing real-time diagnostic performance, but also carries the risk of leakage and tampering during data transmission. Although some technologies use dedicated encrypted channels and physical isolation to ensure data security, these methods are costly to deploy and cannot fundamentally resolve the conflict between privacy protection and collaborative utilization arising from data centralization, making it difficult to scale up for application in multi-segment collaborative diagnostic scenarios.
[0007] Fourth, it lacks anti-interference capabilities and has limited adaptability to complex collaborative scenarios. During multi-segment collaborative diagnosis, monitoring data from different segments may have issues such as noise interference and inconsistent features. Existing diagnostic methods are weak in handling such heterogeneous interference and lack fault-tolerant mechanisms for multi-node collaborative training. They are easily affected by abnormal data from individual segments (such as Byzantine node attacks), leading to performance degradation of the global diagnostic model and an inability to stably support multi-segment joint fault diagnosis decisions.
[0008] Federated learning, as a distributed machine learning paradigm, enables multiple participants to jointly train a global model without sharing raw data, balancing data privacy protection with collaborative value mining. This provides a new technical approach to addressing the core pain points of collaborative diagnosis of track faults across multiple railway sections. However, the application of existing federated learning technology in railway track diagnosis is still in its early stages. It lacks customized designs to address the heterogeneity of track fault data, the differences in operating conditions across multiple sections, and the real-time requirements of diagnosis, making it difficult to directly adapt to collaborative diagnosis scenarios for track faults across multiple sections.
[0009] Therefore, how to provide a collaborative diagnosis method and system for multi-segment railway track faults based on federated learning, so as to improve the data privacy and security protection capabilities of multi-segment tracks, the accuracy of cross-domain fault diagnosis and the generalization ability of models, while strengthening the anti-interference and fault tolerance performance in the collaborative training process, and taking into account the real-time diagnosis and the adaptability for large-scale promotion, so as to break through the core bottlenecks of data silos, insufficient cross-domain adaptability, contradiction between security and real-time performance and weak anti-interference capabilities, has become an urgent technical problem to be solved. Summary of the Invention
[0010] The technical problem to be solved by this invention is to provide a collaborative diagnosis method and system for multi-segment railway track faults based on federated learning, which improves the data privacy and security protection capabilities of multi-segment tracks, the accuracy of cross-domain fault diagnosis and the generalization ability of the model, while strengthening the anti-interference and fault tolerance performance in the collaborative training process, and taking into account both the real-time diagnostic performance and the adaptability for large-scale promotion, so as to overcome the core bottlenecks of data silos, insufficient cross-domain adaptability, contradiction between security and real-time performance and weak anti-interference capabilities.
[0011] In a first aspect, the present invention provides a collaborative diagnosis method for multi-segment railway track faults based on federated learning, comprising the following steps: Step S1: Deploy the client on different sections of the railway track, collect heterogeneous monitoring data of the railway track in the section, including at least track vibration signals, deformation data and environmental parameters, and store them locally with encryption. Step S2: The clients of each road segment preprocess and label the heterogeneous monitoring data to construct a local fault feature dataset; Step S3: The central server initializes a global fault diagnosis module and a global output module, and encrypts and sends the initial global diagnosis parameters of the global fault diagnosis module and the initial global output parameters of the global output module to the clients of each road segment. Step S4: Each road segment client initializes a local fault diagnosis module based on the initial global diagnostic parameters, initializes a personalized output module based on the initial global output parameters, and constructs a local fault diagnosis model based on the local fault diagnosis module and the personalized output module. Step S5: Each road segment client trains the local fault diagnosis model locally based on the fault feature dataset; after training, the quality evaluation index of the training data in this round is calculated, and the local update parameters of the local fault diagnosis module and the quality evaluation index are encrypted and uploaded to the central server. Step S6: The central server adopts a weighted aggregation strategy with Byzantine robustness, calculates the aggregation weight based on the quality assessment indicators uploaded by each road segment client, aggregates the local update parameters based on the aggregation weights, updates the global fault diagnosis module to obtain optimized global diagnosis parameters, encrypts and sends the optimized global diagnosis parameters to each road segment client, and repeats the local training and parameter aggregation steps until the global fault diagnosis module reaches the preset convergence condition or training round, and encrypts and sends the latest optimized global diagnosis parameters to the road segment client to update the local fault diagnosis module. Step S7: Each section client uses the latest local fault diagnosis model to identify and classify faults in the local real-time heterogeneous monitoring data, outputs real-time fault diagnosis results, and realizes collaborative diagnosis of track faults across multiple sections.
[0012] Furthermore, step S1 specifically includes: The client-side components deployed on different sections of the railway track collect heterogeneous monitoring data of the track section via sensor arrays, including at least track vibration signals, deformation data, and environmental parameters. The track vibration signals include at least the vertical vibration acceleration of the rail, the lateral vibration acceleration of the rail, the sleeper vibration amplitude, and the vibration frequency. The deformation data includes at least the vertical displacement of the rail, the lateral displacement of the rail, the gauge deviation, and the rail curvature. The environmental parameters include at least the ambient temperature, ambient humidity, precipitation, wind speed, and dust concentration. The road segment client will collect the heterogeneous monitoring data and store it in a local database encrypted with the AES-256 encryption algorithm.
[0013] Furthermore, the database has a role-based hierarchical authorization mechanism, which only assigns corresponding operation permissions to authorized operation and maintenance personnel, and combines multi-factor authentication to prevent unauthorized access. At the same time, it records all data access, data modification and data export operation logs, and retains the operation logs for at least 90 days. When the road segment client updates each of the databases, it automatically calculates the data fingerprint of the database based on the HASH-256 algorithm and stores the data fingerprint on the blockchain. Before updating the database, it performs an integrity check based on the data fingerprint stored on the blockchain in the previous time, and periodically performs cloud-encrypted backup of the data stored in the database. Based on the storage space of the road segment client and the preset clearing rules, it continuously clears the invalid data in the database.
[0014] Furthermore, step S2 specifically includes: Each road segment client performs preprocessing on the heterogeneous monitoring data, including outlier handling, noise filtering, data alignment and synchronization, normalization, and feature engineering, to obtain standardized data; Based on railway industry standards, the standardized data are labeled with fault types and fault levels to obtain corresponding labeling information; By employing SMOTE oversampling technology or random undersampling technology, the proportion of fault samples to normal samples in each standardized dataset is balanced based on the annotation information, thereby constructing a local fault feature dataset. The outlier handling specifically involves: using statistical methods based on interquartile range or the 3σ criterion to identify and remove abnormal data points in the heterogeneous monitoring data caused by momentary sensor failure or external interference; and for continuous data in the heterogeneous monitoring data, using linear interpolation or the mean of the effective values before and after the data to fill in the gaps. The noise filtering specifically involves: using a Butterworth bandpass filter or wavelet threshold denoising method to filter out high-frequency noise and power frequency interference from the track vibration signal in the heterogeneous monitoring data, while retaining the effective frequency band signal related to the characteristics of rail faults; using a moving average filter to smooth the deformation data; and using a sliding window value filter method for the environmental parameters, while introducing physical constraint-based adaptive filtering and multi-sensor data verification. By using a preset instantaneous rate of change threshold, physically impossible abnormal jumps are identified and processed, and cross-validation is performed using the correlation between parameters. The data alignment and synchronization specifically involves: based on a unified first timestamp, performing time alignment on the heterogeneous monitoring data from sensors with different sampling frequencies, and unifying the data to the same time series through an interpolation resampling method to ensure the temporal correlation between the data; The normalization process specifically involves performing Z-score standardization or Min-Max normalization on the heterogeneous monitoring data to eliminate the influence of dimensions. The feature engineering specifically involves extracting time-domain features, frequency-domain features, and time-frequency-domain features from the heterogeneous monitoring data.
[0015] Furthermore, step S3 specifically includes: The central server initializes a global fault diagnosis module and a global output module; The global fault diagnosis module is constructed based on a heterogeneous data adaptation unit, a multi-scale feature extraction unit, a cross-road segment feature fusion unit, and a feature calibration unit; the global output module is constructed based on a feature enhancement unit, a multi-task prediction unit, and a probability calibration unit. Obtain the initial global diagnostic parameters of the global fault diagnosis module and the initial global output parameters of the global output module; Obtain the current second timestamp, calculate the hash value of the initial global diagnostic parameters, initial global output parameters, and second timestamp using the SM3 algorithm, encrypt the initial global diagnostic parameters, initial global output parameters, second timestamp, and hash value into an encrypted parameter packet using a preset SM4 raw key, retrieve the public key of each road segment client through the auxiliary server, encrypt the SM4 raw key with each public key to obtain the corresponding encryption key, and send the encrypted parameter packet and encryption key to the corresponding road segment client through an SSL encrypted transmission link.
[0016] Furthermore, the heterogeneous data adaptation unit is constructed based on vibration signal branches, deformation data branches, environmental parameter branches, and fusion subunits; The vibration signal branch is used to capture the time-frequency local features of non-stationary track vibration signals by using wavelet basis functions through wavelet convolution layers. The deformation data branch is used to capture the spatial correlation features of rail deformation from the deformation data through a spatial convolutional layer using a 3×3 convolutional kernel; The environmental parameter branch is used to capture the temporal variation trend characteristics of environmental parameters through a temporal convolutional layer using a 1×k convolutional kernel; The fusion subunit is used to dynamically calculate the attention weights of the time-frequency local features, spatial correlation features, and temporal change trend features through a multi-head attention weighting layer, and then weight and fuse them into a unified embedded feature. The multi-scale feature extraction unit is constructed based on multi-scale dilated convolutional blocks, channel attention subunits, and spatial attention subunits. The multi-scale dilated convolutional block is used to capture local and global features from the unified embedding features through four parallel dilated convolutional layers; The channel attention subunit is used to perform channel-dimensional "squeeze-excitement" on the output of the multi-scale dilated convolutional block through the squeeze-excitement module, thereby enhancing fault-sensitive channels and suppressing redundant channels. The spatial attention subunit is used to introduce position information into the output of the channel attention subunit through the coordinate attention module, focus on the spatial location-related features of the fault occurrence, improve the spatial discriminativeness of the features, and output multi-scale fusion features; The cross-segment feature fusion unit is constructed based on a graph construction subunit, a graph convolutional layer, and a knowledge distillation subunit. The graph construction subunit is used to extract a normalized adjacency matrix from the multi-scale fusion features through a dynamic federated graph generator; The graph convolutional layer is used to aggregate the features of adjacent or similar road segments in the normalized adjacency matrix through weighted graph convolution, and output a common fault feature matrix. The knowledge distillation subunit is used to take the strong road segment features with abundant fault samples and high feature quality in the common fault feature matrix as teacher signals and distill them into the weak road segment features with scarce fault samples, so as to improve the feature expression ability of the weak road segment and output cross-road segment common features. The feature calibration unit is constructed based on a feature confidence evaluation subunit, an adaptive filtering subunit, and a feature smoothing subunit. The feature credibility assessment subunit is used to calculate the credibility score matrix of the cross-segment common features of each railway segment using a Mahalanobis distance calculator, based on each of the multi-scale fusion features and the global feature distribution statistics; the global feature distribution statistics are the mean vector μ and covariance matrix Σ of the multi-scale fusion feature statistics of all railway segments. The adaptive filtering subunit is used to perform weighted suppression on the cross-segment common features through a soft threshold filtering layer, based on the confidence score matrix, to obtain a filtered common feature matrix. The feature smoothing subunit is used to perform Gaussian smoothing on the filtered common feature matrix through a Gaussian kernel smoothing layer to reduce feature fluctuations, improve the stability of common features, and output calibration common fault features. The feature enhancement unit is constructed based on residual connection blocks and task attention subunits; The residual connection block is used to enhance the common fault features of calibration through the ResNet bottleneck structure to obtain residual enhanced common fault features; The task attention subunit is used to dynamically adjust the feature enhancement weights of the residual-enhanced common fault features based on the task differences between fault type prediction and fault level prediction through the dynamic attention layer, thereby improving the adaptability of features to tasks and obtaining task-enhanced common features. The multi-task prediction unit is constructed based on a shared feature layer, a fault type prediction branch, a fault level prediction branch, and a task interaction subunit. The shared feature layer is used to extract common dual-task shared basic features for fault type prediction task and fault level prediction task from the common features of task enhancement through two fully connected layers. The fault type prediction branch is used to identify the fault type probability distribution defined by railway industry standards from the dual-task shared basic features through a 3-layer MLP and Softmax activation. The fault level prediction branch is used to identify the fault level probability distribution from the shared basic features of the dual tasks through a 3-layer MLP and Sigmoid activation. The task interaction subunit is used to calculate the interaction loss value through the interaction loss function, and to backpropagate and update the network parameters of the shared feature layer, the fault type prediction branch and the fault level prediction branch based on the interaction loss value. The probability calibration unit is constructed based on the distribution adaptation subunit and the probability correction subunit; The distribution adaptation subunit is used to perform temperature scaling on the fault type probability distribution and the fault level probability distribution. The probability correction subunit is used to perform nonlinear correction on the temperature-scaled fault type probability distribution and fault level probability distribution through the Beta calibration layer, so as to output fault diagnosis results carrying fault type and fault level.
[0017] Furthermore, step S4 specifically includes: Each client section receives the encrypted parameter packet and encryption key sent by the central server, calls the locally stored private key, decrypts the received encryption key to obtain the SM4 original key, and then decrypts the received encrypted parameter packet using the SM4 original key to obtain the initial global diagnostic parameters, initial global output parameters, second timestamp, and hash value. After performing integrity verification on the initial global diagnostic parameters, initial global output parameters, and second timestamp based on the hash value, a timeliness verification is performed based on the second timestamp. If the verification passes, the parameter reception is completed; if the verification fails, a retransmission request is sent to the central server. The central server, in conjunction with the access logs of the auxiliary server, investigates transmission anomalies and retransmits the data. Each road segment client initializes a local fault diagnosis module with the same network architecture as the global fault diagnosis module based on the initial global diagnostic parameters, and initializes a personalized output module with the same architecture as the global output module based on the initial global output parameters. A local fault diagnosis model is then constructed based on the local fault diagnosis module and the personalized output module.
[0018] Furthermore, step S5 specifically includes: Each road segment client uses a 30-day base time window and initially divides the fault feature dataset into a training set and a validation set in a 7:3 ratio to ensure that the training set covers samples of all fault types. Calculate the fault distribution entropy of the validation set. If the fault distribution entropy is less than a preset entropy threshold, dynamically adjust the ratio of the training set to the validation set within the basic time window. The training and validation sets are subjected to fault scenario-driven online data augmentation, specifically as follows: For the sample of track vibration signal, preset amplitude scaling, time stretching and Gaussian noise superposition are applied based on physical simulation rules to simulate the changes of fault signal under different train loads and speeds. For the deformation data samples, combined with the historical distribution of local environmental parameters, deformation-derived samples under different environmental couplings are generated through linear transformation. For samples of environmental parameters, virtual samples of the synergistic effects of multiple environmental factors are generated by permutation and combination to make up for the scarcity of local composite scene samples. The local fault diagnosis model is trained locally using the online data-augmented training set, and the trained local fault diagnosis model is validated using the online data-augmented validation set. After every 5 rounds of training, the base time window is updated (slid forward 1 day), and the training set and validation set are re-divided until the preset number of training rounds is completed. After training is completed, the quality assessment index of the training data for this round is calculated. The local update parameters of the local fault diagnosis module and the quality assessment index are encrypted using a homomorphic encryption algorithm and then uploaded to the central server through an SSL encrypted transmission link.
[0019] Furthermore, in step S6, the weighted aggregation strategy with Byzantine robustness specifically involves: the central server calculating the aggregation weight of each local update parameter based on the quality assessment indicators uploaded by the clients of each road segment, using a trimmed average algorithm to remove extreme outlier parameters from the local update parameters, and then performing a weighted summation of the effective parameters in the local update parameters based on the aggregation weight.
[0020] Secondly, this invention provides a multi-segment railway track fault collaborative diagnosis system based on federated learning, comprising the following modules: The heterogeneous monitoring data acquisition module is used to deploy on the client side of different railway sections to collect heterogeneous monitoring data of the railway section, including at least track vibration signals, deformation data and environmental parameters, and store them locally with encryption. The fault feature dataset construction module is used by clients of each road segment to preprocess and label the heterogeneous monitoring data in order to construct a local fault feature dataset. The server model initialization module is used by the central server to initialize a global fault diagnosis module and a global output module, and to encrypt and send the initial global diagnosis parameters of the global fault diagnosis module and the initial global output parameters of the global output module to the clients of each road segment. The local fault diagnosis model initialization module is used by each road segment client to initialize a local fault diagnosis module based on the initial global diagnosis parameters, initialize a personalized output module based on the initial global output parameters, and construct a local fault diagnosis model based on the local fault diagnosis module and the personalized output module. The local training module is used by each road segment client to train the local fault diagnosis model locally based on the fault feature dataset. After training is completed, the quality evaluation index of the training data in this round is calculated, and the local update parameters of the local fault diagnosis module and the quality evaluation index are encrypted and uploaded to the central server. The parameter aggregation module is used by the central server to adopt a weighted aggregation strategy with Byzantine robustness, calculate the aggregation weight based on the quality assessment indicators uploaded by each road segment client, aggregate each locally updated parameter based on the aggregation weight, update the global fault diagnosis module to obtain optimized global diagnosis parameters, encrypt and send the optimized global diagnosis parameters to each road segment client, repeat the local training and parameter aggregation steps until the global fault diagnosis module reaches the preset convergence condition or training round, and encrypt and send the latest optimized global diagnosis parameters to the road segment client to update the local fault diagnosis module; The rail fault diagnosis module is used by clients of each section to identify and classify faults in local real-time heterogeneous monitoring data using the latest local fault diagnosis model, and output real-time fault diagnosis results to achieve collaborative diagnosis of rail faults across multiple sections.
[0021] Furthermore, the heterogeneous monitoring data acquisition module is specifically used for: The client-side components deployed on different sections of the railway track collect heterogeneous monitoring data of the track section via sensor arrays, including at least track vibration signals, deformation data, and environmental parameters. The track vibration signals include at least the vertical vibration acceleration of the rail, the lateral vibration acceleration of the rail, the sleeper vibration amplitude, and the vibration frequency. The deformation data includes at least the vertical displacement of the rail, the lateral displacement of the rail, the gauge deviation, and the rail curvature. The environmental parameters include at least the ambient temperature, ambient humidity, precipitation, wind speed, and dust concentration. The road segment client will collect the heterogeneous monitoring data and store it in a local database encrypted with the AES-256 encryption algorithm.
[0022] Furthermore, the database has a role-based hierarchical authorization mechanism, which only assigns corresponding operation permissions to authorized operation and maintenance personnel, and combines multi-factor authentication to prevent unauthorized access. At the same time, it records all data access, data modification and data export operation logs, and retains the operation logs for at least 90 days. When the road segment client updates each of the databases, it automatically calculates the data fingerprint of the database based on the HASH-256 algorithm and stores the data fingerprint on the blockchain. Before updating the database, it performs an integrity check based on the data fingerprint stored on the blockchain in the previous time, and periodically performs cloud-encrypted backup of the data stored in the database. Based on the storage space of the road segment client and the preset clearing rules, it continuously clears the invalid data in the database.
[0023] Furthermore, the fault feature dataset construction module is specifically used for: Each road segment client performs preprocessing on the heterogeneous monitoring data, including outlier handling, noise filtering, data alignment and synchronization, normalization, and feature engineering, to obtain standardized data; Based on railway industry standards, the standardized data are labeled with fault types and fault levels to obtain corresponding labeling information; By employing SMOTE oversampling technology or random undersampling technology, the proportion of fault samples to normal samples in each standardized dataset is balanced based on the annotation information, thereby constructing a local fault feature dataset. The outlier handling specifically involves: using statistical methods based on interquartile range or the 3σ criterion to identify and remove abnormal data points in the heterogeneous monitoring data caused by momentary sensor failure or external interference; and for continuous data in the heterogeneous monitoring data, using linear interpolation or the mean of the effective values before and after the data to fill in the gaps. The noise filtering specifically involves: using a Butterworth bandpass filter or wavelet threshold denoising method to filter out high-frequency noise and power frequency interference from the track vibration signal in the heterogeneous monitoring data, while retaining the effective frequency band signal related to the characteristics of rail faults; using a moving average filter to smooth the deformation data; and using a sliding window value filter method for the environmental parameters, while introducing physical constraint-based adaptive filtering and multi-sensor data verification. By using a preset instantaneous rate of change threshold, physically impossible abnormal jumps are identified and processed, and cross-validation is performed using the correlation between parameters. The data alignment and synchronization specifically involves: based on a unified first timestamp, performing time alignment on the heterogeneous monitoring data from sensors with different sampling frequencies, and unifying the data to the same time series through an interpolation resampling method to ensure the temporal correlation between the data; The normalization process specifically involves performing Z-score standardization or Min-Max normalization on the heterogeneous monitoring data to eliminate the influence of dimensions. The feature engineering specifically involves extracting time-domain features, frequency-domain features, and time-frequency-domain features from the heterogeneous monitoring data.
[0024] Furthermore, the server model initialization module is specifically used for: The central server initializes a global fault diagnosis module and a global output module; The global fault diagnosis module is constructed based on a heterogeneous data adaptation unit, a multi-scale feature extraction unit, a cross-road segment feature fusion unit, and a feature calibration unit; the global output module is constructed based on a feature enhancement unit, a multi-task prediction unit, and a probability calibration unit. Obtain the initial global diagnostic parameters of the global fault diagnosis module and the initial global output parameters of the global output module; Obtain the current second timestamp, calculate the hash value of the initial global diagnostic parameters, initial global output parameters, and second timestamp using the SM3 algorithm, encrypt the initial global diagnostic parameters, initial global output parameters, second timestamp, and hash value into an encrypted parameter packet using a preset SM4 raw key, retrieve the public key of each road segment client through the auxiliary server, encrypt the SM4 raw key with each public key to obtain the corresponding encryption key, and send the encrypted parameter packet and encryption key to the corresponding road segment client through an SSL encrypted transmission link.
[0025] Furthermore, the heterogeneous data adaptation unit is constructed based on vibration signal branches, deformation data branches, environmental parameter branches, and fusion subunits; The vibration signal branch is used to capture the time-frequency local features of non-stationary track vibration signals by using wavelet basis functions through wavelet convolution layers. The deformation data branch is used to capture the spatial correlation features of rail deformation from the deformation data through a spatial convolutional layer using a 3×3 convolutional kernel; The environmental parameter branch is used to capture the temporal variation trend characteristics of environmental parameters through a temporal convolutional layer using a 1×k convolutional kernel; The fusion subunit is used to dynamically calculate the attention weights of the time-frequency local features, spatial correlation features, and temporal change trend features through a multi-head attention weighting layer, and then weight and fuse them into a unified embedded feature. The multi-scale feature extraction unit is constructed based on multi-scale dilated convolutional blocks, channel attention subunits, and spatial attention subunits. The multi-scale dilated convolutional block is used to capture local and global features from the unified embedding features through four parallel dilated convolutional layers; The channel attention subunit is used to perform channel-dimensional "squeeze-excitement" on the output of the multi-scale dilated convolutional block through the squeeze-excitement module, thereby enhancing fault-sensitive channels and suppressing redundant channels. The spatial attention subunit is used to introduce position information into the output of the channel attention subunit through the coordinate attention module, focus on the spatial location-related features of the fault occurrence, improve the spatial discriminativeness of the features, and output multi-scale fusion features; The cross-segment feature fusion unit is constructed based on a graph construction subunit, a graph convolutional layer, and a knowledge distillation subunit. The graph construction subunit is used to extract a normalized adjacency matrix from the multi-scale fusion features through a dynamic federated graph generator; The graph convolutional layer is used to aggregate the features of adjacent or similar road segments in the normalized adjacency matrix through weighted graph convolution, and output a common fault feature matrix. The knowledge distillation subunit is used to take the strong road segment features with abundant fault samples and high feature quality in the common fault feature matrix as teacher signals and distill them into the weak road segment features with scarce fault samples, so as to improve the feature expression ability of the weak road segment and output cross-road segment common features. The feature calibration unit is constructed based on a feature confidence evaluation subunit, an adaptive filtering subunit, and a feature smoothing subunit. The feature credibility assessment subunit is used to calculate the credibility score matrix of the cross-segment common features of each railway segment using a Mahalanobis distance calculator, based on each of the multi-scale fusion features and the global feature distribution statistics; the global feature distribution statistics are the mean vector μ and covariance matrix Σ of the multi-scale fusion feature statistics of all railway segments. The adaptive filtering subunit is used to perform weighted suppression on the cross-segment common features through a soft threshold filtering layer, based on the confidence score matrix, to obtain a filtered common feature matrix. The feature smoothing subunit is used to perform Gaussian smoothing on the filtered common feature matrix through a Gaussian kernel smoothing layer to reduce feature fluctuations, improve the stability of common features, and output calibration common fault features. The feature enhancement unit is constructed based on residual connection blocks and task attention subunits; The residual connection block is used to enhance the common fault features of calibration through the ResNet bottleneck structure to obtain residual enhanced common fault features; The task attention subunit is used to dynamically adjust the feature enhancement weights of the residual-enhanced common fault features based on the task differences between fault type prediction and fault level prediction through the dynamic attention layer, thereby improving the adaptability of features to tasks and obtaining task-enhanced common features. The multi-task prediction unit is constructed based on a shared feature layer, a fault type prediction branch, a fault level prediction branch, and a task interaction subunit. The shared feature layer is used to extract common dual-task shared basic features for fault type prediction task and fault level prediction task from the common features of task enhancement through two fully connected layers. The fault type prediction branch is used to identify the fault type probability distribution defined by railway industry standards from the dual-task shared basic features through a 3-layer MLP and Softmax activation. The fault level prediction branch is used to identify the fault level probability distribution from the shared basic features of the dual tasks through a 3-layer MLP and Sigmoid activation. The task interaction subunit is used to calculate the interaction loss value through the interaction loss function, and to backpropagate and update the network parameters of the shared feature layer, the fault type prediction branch and the fault level prediction branch based on the interaction loss value. The probability calibration unit is constructed based on the distribution adaptation subunit and the probability correction subunit; The distribution adaptation subunit is used to perform temperature scaling on the fault type probability distribution and the fault level probability distribution. The probability correction subunit is used to perform nonlinear correction on the temperature-scaled fault type probability distribution and fault level probability distribution through the Beta calibration layer, so as to output fault diagnosis results carrying fault type and fault level.
[0026] Furthermore, the local fault diagnosis model initialization module is specifically used for: Each client section receives the encrypted parameter packet and encryption key sent by the central server, calls the locally stored private key, decrypts the received encryption key to obtain the SM4 original key, and then decrypts the received encrypted parameter packet using the SM4 original key to obtain the initial global diagnostic parameters, initial global output parameters, second timestamp, and hash value. After performing integrity verification on the initial global diagnostic parameters, initial global output parameters, and second timestamp based on the hash value, a timeliness verification is performed based on the second timestamp. If the verification passes, the parameter reception is completed; if the verification fails, a retransmission request is sent to the central server. The central server, in conjunction with the access logs of the auxiliary server, investigates transmission anomalies and retransmits the data. Each road segment client initializes a local fault diagnosis module with the same network architecture as the global fault diagnosis module based on the initial global diagnostic parameters, and initializes a personalized output module with the same architecture as the global output module based on the initial global output parameters. A local fault diagnosis model is then constructed based on the local fault diagnosis module and the personalized output module.
[0027] Furthermore, the local training module is specifically used for: Each road segment client uses a 30-day base time window and initially divides the fault feature dataset into a training set and a validation set in a 7:3 ratio to ensure that the training set covers samples of all fault types. Calculate the fault distribution entropy of the validation set. If the fault distribution entropy is less than a preset entropy threshold, dynamically adjust the ratio of the training set to the validation set within the basic time window. The training and validation sets are subjected to fault scenario-driven online data augmentation, specifically as follows: For the sample of track vibration signal, preset amplitude scaling, time stretching and Gaussian noise superposition are applied based on physical simulation rules to simulate the changes of fault signal under different train loads and speeds. For the deformation data samples, combined with the historical distribution of local environmental parameters, deformation-derived samples under different environmental couplings are generated through linear transformation. For samples of environmental parameters, virtual samples of the synergistic effects of multiple environmental factors are generated by permutation and combination to make up for the scarcity of local composite scene samples. The local fault diagnosis model is trained locally using the online data-augmented training set, and the trained local fault diagnosis model is validated using the online data-augmented validation set. After every 5 rounds of training, the base time window is updated (slid forward 1 day), and the training set and validation set are re-divided until the preset number of training rounds is completed. After training is completed, the quality assessment index of the training data for this round is calculated. The local update parameters of the local fault diagnosis module and the quality assessment index are encrypted using a homomorphic encryption algorithm and then uploaded to the central server through an SSL encrypted transmission link.
[0028] Furthermore, in the parameter aggregation module, the weighted aggregation strategy with Byzantine robustness specifically involves: the central server calculating the aggregation weight of each local update parameter based on the quality assessment indicators uploaded by the clients of each road segment, using a trimmed average algorithm to remove extreme abnormal parameters from the local update parameters, and then performing a weighted summation of the effective parameters in the local update parameters based on the aggregation weight.
[0029] The advantages of this invention are: 1. By deploying client machines on different railway sections, heterogeneous monitoring data of the railway sections is collected and stored locally with encryption. Each client machine preprocesses and labels the heterogeneous monitoring data to construct a local fault feature dataset. The central server initializes the global fault diagnosis module and the global output module, and encrypts and distributes the initial global diagnostic parameters of the global fault diagnosis module and the initial global output parameters of the global output module to each client machine. Each client machine initializes its local fault diagnosis module based on the initial global diagnostic parameters, initializes its personalized output module based on the initial global output parameters, and initializes its local fault diagnosis module and personalized output module based on the local fault diagnosis module and personalized output module. The personalized output module constructs a local fault diagnosis model and trains it locally based on a fault feature dataset. After training, it calculates the quality assessment index of the training data for this round, and encrypts and uploads the locally updated parameters of the local fault diagnosis module and the quality assessment index to the central server. The central server adopts a weighted aggregation strategy with Byzantine robustness, calculates the aggregation weight based on each quality assessment index, and aggregates the parameters of each locally updated parameter based on the aggregation weight to update the global fault diagnosis module and obtain optimized global diagnosis parameters. The optimized global diagnosis parameters are then encrypted and distributed to the clients of each road segment, and the local training and parameter aggregation are repeated. The process involves several steps, continuing until the global fault diagnosis module reaches the preset convergence condition or training rounds. The latest optimized global diagnostic parameters are then encrypted and sent to the road segment clients to update their local fault diagnosis modules. Finally, each road segment client uses the latest local fault diagnosis model to identify and classify faults in its local real-time heterogeneous monitoring data, outputting real-time fault diagnosis results. This demonstrates how a federated learning-based collaborative diagnosis framework, with encrypted storage and training of data locally on each road segment, only exchanges encrypted model parameters rather than the original data, fundamentally solving the data silo problem and ensuring privacy and security. Furthermore, the "global fault diagnosis module" shares common knowledge and... The "personalized output module" is designed to adapt to local characteristics, effectively overcoming the differences in data distribution across multiple road segments and significantly improving cross-domain diagnostic accuracy and model generalization ability. Furthermore, the introduction of a Byzantine robust aggregation strategy based on quality assessment metrics empowers the system to identify and suppress abnormal or low-quality parameter updates, thereby enhancing the anti-interference and fault-tolerant performance of collaborative training. Simultaneously, massive data transmission is transformed into lightweight parameter interaction to alleviate bandwidth pressure, and the final real-time diagnostic task is deployed locally on the road segment client, successfully balancing the real-time nature of diagnosis with the system's scalability and adaptability, forming a comprehensive solution that overcomes multiple technical bottlenecks.
[0030] 2. By constructing a collaborative diagnostic framework based on federated learning, and under the premise of local encrypted storage and processing of data in each section, global knowledge sharing is achieved through parameter encryption transmission and aggregation, effectively solving the data silo dilemma. Privacy and security protection are strengthened by relying on national cryptographic algorithms and blockchain evidence storage. Through a customized model architecture integrating heterogeneous data adaptation, knowledge distillation, and personalized output modules, the cross-domain adaptation accuracy and generalization ability of the model to different road conditions are significantly improved. Furthermore, by adopting a Byzantine robust aggregation strategy and a multi-level quality assessment mechanism, interference from low-quality data and malicious nodes is effectively resisted, enhancing the anti-interference and fault tolerance of collaborative training. Finally, through real-time edge-side diagnosis and lightweight communication design, efficient real-time diagnosis is achieved while ensuring security, providing a reliable technical path for the large-scale promotion of large-scale railway networks.
[0031] 3. By adopting a federated learning framework, clients on each track section can process sensitive data locally without uploading raw monitoring data to the central server, thus effectively avoiding the risk of data leakage. At the same time, it integrates multi-layer encryption measures (such as AES-256 encrypted database, SM4 encrypted parameter transmission), blockchain-based data fingerprint storage, and role-based hierarchical authorization mechanisms to ensure the integrity and confidentiality of data during storage, transmission, and access, reducing compliance risks. It is particularly suitable for sensitive applications involving critical infrastructure, such as railway fault diagnosis.
[0032] 4. By integrating heterogeneous monitoring data (such as track vibration signals, deformation data, and environmental parameters) and combining multi-scale feature extraction, cross-segment feature fusion, and knowledge distillation techniques, more comprehensive fault characteristics can be captured, reducing false alarms and missed alarms. In addition, by adopting a weighted aggregation strategy with Byzantine robustness, the central server can remove parameters from abnormal clients, avoiding malicious or faulty nodes from affecting the global model, thereby improving the system's stability and diagnostic accuracy in real-world complex environments.
[0033] 5. The modular design, including a global fault diagnosis module, a personalized output module, and a local model, allows new railway segment clients to be easily added to the system without reconstructing the overall architecture. The distributed nature of federated learning allows each railway segment to be trained in a personalized manner based on local data characteristics, while knowledge sharing is achieved through parameter aggregation, adapting to environmental differences (such as climate and load conditions) of different railway segments. This design supports the smooth upgrading of the system as the railway network expands, reducing maintenance costs.
[0034] 6. Supports real-time fault identification and classification. It can quickly analyze real-time heterogeneous monitoring data through local fault diagnosis models, reducing cloud processing latency. Preprocessing steps (such as data alignment and normalization) and online data augmentation techniques (such as sample generation based on physical simulation) optimize training efficiency and ensure that the model can respond to sudden faults in a timely manner. Periodic sliding time windows and dynamic adjustment of training set ratios further enhance the model's ability to adapt to changes and achieve efficient collaborative diagnosis.
[0035] 7. By deploying client-side systems on various railway tracks, comprehensive multi-dimensional heterogeneous monitoring data, including vibration, deformation, and environmental data, are collected. This provides high-value feature inputs for the federated learning model, significantly improving the accuracy and robustness of fault diagnosis. Based on this, a multi-layered data security and privacy protection system is constructed through AES-256 encrypted storage, role-based hierarchical authorization, and operation log auditing, effectively safeguarding data sovereignty and compliance. Furthermore, blockchain technology is used for trusted data fingerprint storage and integrity verification, ensuring data tamper-proofing and traceability. Simultaneously, cloud-based encrypted backup and intelligent data lifecycle management strategies enhance the system's disaster recovery capabilities and local storage efficiency. The entire design achieves secure, reliable, and efficient multi-party collaborative diagnosis while ensuring data remains within the local system, demonstrating innovation, practicality, and high engineering feasibility.
[0036] 8. Through a systematic, refined, and engineering-practice-oriented data preprocessing and construction process, the feasibility, accuracy, and reliability of collaborative diagnosis of multi-segment railway track faults under the federated learning framework are fundamentally improved. By integrating statistical methods with physical constraints for outlier handling and adaptive noise filtering, the high quality and physical rationality of heterogeneous monitoring data are ensured. Based on unified standard labeling and targeted sample balancing strategies, the problems of data class imbalance and labeling consistency are solved. Simultaneously, temporal alignment, normalization, and multi-dimensional feature engineering provide standardized and information-rich input for subsequent federated learning models. These measures work together to significantly enhance the quality and representativeness of local datasets, thereby effectively supporting the efficient training and optimization of the federated global model, ultimately achieving more accurate and robust intelligent diagnosis of faults in complex railway environments.
[0037] 9. Through a heterogeneous data adaptation unit, vibration signal branches, deformation data branches, and environmental parameter branches were specifically designed. Wavelet convolutional layers were used to capture the time-frequency local features of non-stationary vibration signals, spatial convolutional layers to extract the spatial correlation features of deformation data, and temporal convolutional layers to capture the temporal variation trend features of environmental parameters. These features were then dynamically fused into a unified embedded feature through a multi-head attention weighting layer. This method effectively integrates multi-source heterogeneous data in railway fault diagnosis, overcomes the limitations of traditional methods that rely on a single data source, significantly improves the richness of fault features and the accuracy of diagnosis, and thus enhances the model's adaptability to complex working conditions.
[0038] 10. By using a cross-segment feature fusion unit, a normalized adjacency matrix is generated based on a graph construction sub-unit. Features of similar road segments are aggregated using graph convolutional layers, and then strong road segment features are distilled into weak road segment features using a knowledge distillation sub-unit. This design achieves effective transfer of fault knowledge between multiple road segments, especially for weak road segments with scarce fault samples, improving their feature representation ability and diagnostic performance. This solves the data imbalance problem in practical applications, enhances the model's generalization ability and robustness, and enables the diagnostic system to maintain high accuracy on different road segments, reducing the risk of overfitting.
[0039] 11. The feature calibration unit calculates the Mahalanobis distance confidence score through the feature confidence evaluation subunit, sets a dynamic soft threshold to suppress noise through the adaptive filtering subunit, and performs Gaussian smoothing through the feature smoothing subunit to reduce fluctuations. This series of operations ensures that the output calibration common fault features have high confidence and stability, and reduces diagnostic errors caused by data noise or abnormal fluctuations. This improves the reliability of the model in real-time monitoring, provides more robust feature inputs for subsequent prediction tasks, and thus enhances the stability and practicality of the overall diagnostic system.
[0040] 12. A multi-task prediction unit is adopted, which extracts shared basic features between the two tasks through a shared feature layer. The fault type prediction branch and the fault level prediction branch are processed in parallel, and the network parameters are optimized by combining task interaction sub-units. This design realizes synchronous prediction of fault type and level, avoids redundant calculations, and improves diagnostic efficiency. At the same time, the probability calibration unit performs nonlinear correction on the output probability through temperature scaling and Beta calibration layer, ensuring that the probability distribution of the diagnostic results is more accurate and reliable. This optimizes the utilization of computing resources and enhances the credibility of the model output in practical applications, meeting the railway industry's demand for efficient and accurate diagnosis.
[0041] 13. By using a 30-day base time window and initially dividing the training and validation sets into a fixed ratio, while ensuring that the training set covers samples of all fault types, the model can learn comprehensive fault modes. More importantly, by calculating the fault distribution entropy of the validation set and dynamically adjusting the ratio of the training and validation sets when the entropy value is lower than a preset threshold, this method can automatically identify imbalances or anomalies in data distribution, thereby optimizing data partitioning, preventing overfitting or underfitting, and ensuring that the model maintains high diagnostic accuracy and robustness in complex and ever-changing rail fault scenarios. This dynamic adjustment mechanism enhances the local model adaptability of each section's client in federated learning, providing a reliable foundation for subsequent collaborative diagnosis.
[0042] 14. Targeted online data augmentation strategies were designed for common data types in railway fault diagnosis. For track vibration signals, amplitude scaling, time stretching, and Gaussian noise superposition were applied based on physical simulation rules to simulate fault signal changes under different train loads and speeds. This effectively expanded the training samples and made the model closer to actual operating conditions. For deformation data, linear transformation was performed by combining the historical distribution of local environmental parameters to generate derived samples under different environmental couplings, enhancing the model's adaptability to environmental factors. For environmental parameters, virtual samples were generated through permutation and combination to compensate for the scarcity of local composite scenario samples, thereby improving data diversity and the model's generalization performance. These augmentation methods directly target fault scenarios, improving the recognition accuracy of the local fault diagnosis model in complex realities.
[0043] 15. During training, the base time window is updated (sliding forward by 1 day) after every 5 rounds of training, and the training set and validation set are re-divided until the preset number of training rounds is completed. This sliding window mechanism enables the model to continuously use the latest data for training, adapting to the trend of changes in rail fault data over time (such as seasonal effects or equipment aging), thereby maintaining the timeliness and accuracy of the diagnostic model. It avoids the model obsolescence problem that may be caused by static datasets, supports the dynamic optimization of the federated learning system in long-term operation, and improves the real-time response capability of collaborative diagnosis of rail faults. Attached Figure Description
[0044] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0045] Figure 1 This is a flowchart of a multi-segment railway track fault collaborative diagnosis method based on federated learning, according to the present invention.
[0046] Figure 2 This is a schematic diagram of the structure of a multi-segment railway track fault collaborative diagnosis system based on federated learning according to the present invention.
[0047] Figure 3This is a schematic diagram of the global fault diagnosis module of the present invention.
[0048] Figure 4 This is a schematic diagram of the global output module of the present invention. Detailed Implementation
[0049] The overall approach of the technical solution in this application is as follows: By constructing a collaborative diagnostic framework based on federated learning, and under the premise that data is stored and trained locally in encrypted form in each road segment, only encrypted model parameters are interacted with instead of the original data, fundamentally solving the data silo problem and ensuring privacy and security; through the design of a "global fault diagnosis module" sharing common knowledge and a "personalized output module" adapting to local characteristics, the differences in data distribution across multiple road segments are effectively overcome, significantly improving cross-domain diagnostic accuracy and model generalization ability; furthermore, a Byzantine robust aggregation strategy based on quality evaluation indicators is introduced, giving the system the ability to identify and suppress abnormal or low-quality parameter updates, thereby strengthening collaboration. The system enhances the anti-interference and fault tolerance of training; simultaneously, it transforms massive data transmission into lightweight parameter interaction to alleviate bandwidth pressure, and deploys the final real-time diagnostic task locally on the road segment client. This successfully balances the real-time nature of diagnosis with the system's scalability and adaptability, forming a comprehensive solution that overcomes multiple technical bottlenecks. Ultimately, it effectively improves the data privacy and security protection capabilities of multiple road segments, the accuracy of cross-domain fault diagnosis, and the model's generalization ability. At the same time, it strengthens the anti-interference and fault tolerance of collaborative training, balancing the real-time nature of diagnosis with scalability and adaptability, thereby overcoming the core bottlenecks of data silos, insufficient cross-domain adaptability, the contradiction between security and real-time performance, and weak anti-interference capabilities.
[0050] The hardware architecture of this invention consists of a road segment client cluster, a central server, an auxiliary server, and a blockchain node.
[0051] Section Client: One edge computing gateway is deployed every 5-10 kilometers of track as the section client, equipped with an Intel Core i7-12700H processor, 32GB DDR5 memory, and a 2TB NVMe encrypted solid-state drive, enabling sensor data access, local model training, and real-time diagnostics. Sensor arrays are configured according to the principle of "intensified deployment in critical sections and conventional deployment in ordinary sections." Rail vibration sensors (selected as PCB 356A16 accelerometers) are installed 10cm from the sleeper on the rail web, with one set every 200 meters (including vertical / lateral dual-axis). Deformation sensors (selected as Keyence GT2-H12K laser displacement meters) are installed at the sleeper ends to monitor gauge and rail displacement. Environmental sensors (selected as Sensirion SHT3x temperature and humidity sensors and RainWise sensors) are also included. The MK-III rain gauge is deployed at signal towers along the route, with one set every 1 kilometer. All sensors communicate with the road segment clients through LoRa gateways or 5G industrial modules, with a data transmission rate of no less than 1Mbps and a latency of ≤50ms.
[0052] Central Server: Deployed in the regional operations and maintenance center, it adopts a dual-machine hot standby architecture, configured with an Intel Xeon Platinum 8470C processor, 512GB DDR5 memory and 20TB enterprise-level storage, running the Ubuntu 22.04 LTS operating system and the TensorFlow 2.10 deep learning framework, responsible for global model initialization, parameter aggregation and system scheduling; Auxiliary Server is deployed in the same data center as the central server, mainly responsible for public key management, access log storage and transmission anomaly troubleshooting, with a log storage capacity of no less than 10TB, supporting log backtracking for more than 90 days.
[0053] Blockchain nodes: Adopting a consortium blockchain architecture, the nodes are jointly maintained by various railway bureaus, engineering sections and third-party supervision units. Hyperledger Fabric 2.5 is selected, the block generation cycle is 10 minutes, and the data fingerprint storage transaction confirmation time is ≤3 seconds to ensure that data tampering is traceable.
[0054] Please refer to Figures 1 to 4 As shown, a preferred embodiment of the multi-segment railway track fault collaborative diagnosis method based on federated learning of the present invention includes the following steps: Step S1: Deploy the client on different sections of the railway track, collect heterogeneous monitoring data of the railway track in the section, including at least track vibration signals, deformation data and environmental parameters, and store them locally with encryption. Step S2: The clients of each road segment preprocess and label the heterogeneous monitoring data to construct a local fault feature dataset; Step S3: The central server initializes a global fault diagnosis module and a global output module, and encrypts and sends the initial global diagnosis parameters of the global fault diagnosis module and the initial global output parameters of the global output module to the clients of each road segment. Step S4: Each road segment client initializes a local fault diagnosis module based on the initial global diagnostic parameters, initializes a personalized output module based on the initial global output parameters, and constructs a local fault diagnosis model based on the local fault diagnosis module and the personalized output module. Step S5: Each road segment client trains the local fault diagnosis model locally based on the fault feature dataset; after training, the quality evaluation index of the training data in this round is calculated, and the local update parameters of the local fault diagnosis module and the quality evaluation index are encrypted and uploaded to the central server. Step S6: The central server adopts a weighted aggregation strategy with Byzantine robustness, calculates the aggregation weight based on the quality assessment indicators uploaded by each road segment client, aggregates the local update parameters based on the aggregation weights, updates the global fault diagnosis module to obtain optimized global diagnosis parameters, encrypts and sends the optimized global diagnosis parameters to each road segment client, and repeats the local training and parameter aggregation steps until the global fault diagnosis module reaches the preset convergence condition or training round, and encrypts and sends the latest optimized global diagnosis parameters to the road segment client to update the local fault diagnosis module. Step S7: Each section client uses the latest local fault diagnosis model to identify and classify faults in the local real-time heterogeneous monitoring data, outputs real-time fault diagnosis results, and realizes collaborative diagnosis of track faults across multiple sections.
[0055] Step S1 specifically involves: The client-side components deployed on different sections of the railway track collect heterogeneous monitoring data of the track section via sensor arrays, including at least track vibration signals, deformation data, and environmental parameters. The track vibration signals include at least the vertical vibration acceleration of the rail, the lateral vibration acceleration of the rail, the sleeper vibration amplitude, and the vibration frequency. The deformation data includes at least the vertical displacement of the rail, the lateral displacement of the rail, the gauge deviation, and the rail curvature. The environmental parameters include at least the ambient temperature, ambient humidity, precipitation, wind speed, and dust concentration. The sampling frequency of the track vibration signal is set to 1024Hz (to meet the high-frequency capture requirement of rail crack impact signal), the sampling frequency of deformation data is 10Hz (to adapt to the slow change characteristics of rail deformation), and the sampling frequency of environmental parameters is 1Hz (to balance data volume and timeliness). During the acquisition process, the sensor’s built-in self-calibration module performs zero-point calibration once per hour to avoid data deviation caused by drift.
[0056] The road segment client will collect the heterogeneous monitoring data and store it in a local database encrypted with the AES-256 encryption algorithm.
[0057] The database has a role-based hierarchical authorization mechanism, which only assigns corresponding operation permissions to authorized operation and maintenance personnel, and combines multi-factor authentication to prevent unauthorized access. At the same time, it records all data access, data modification and data export operation logs and retains the operation logs for at least 90 days. In practice, the database uses MySQL 8.0 Enterprise Edition with Transparent Data Encryption (TDE) enabled. AES-256 keys are stored and managed through a hardware security module (HSM, selected as Gemalto SafeNet Luna 7000), and the keys are automatically rotated every 90 days. The role-based authorization mechanism is divided into three levels: "administrator, operation and maintenance operator, and auditor". Administrators have full permissions (including key management), operation and maintenance operators can only read data and export diagnostic results, and auditors can only view operation logs and have no data modification permissions.
[0058] When the road segment client updates each of the databases, it automatically calculates the data fingerprint of the database based on the HASH-256 algorithm and stores the data fingerprint on the blockchain. Before updating the database, it performs an integrity check based on the data fingerprint stored on the blockchain in the previous time, and periodically performs cloud-encrypted backup of the data stored in the database. Based on the storage space of the road segment client and the preset clearing rules, it continuously clears the invalid data in the database.
[0059] In practice, data fingerprint calculation is performed on a daily incremental data basis in the database. After generating a unique fingerprint using the HASH-256 algorithm, a blockchain smart contract is invoked to complete the notarization. The notarized information includes the data fingerprint, road segment number, and update timestamp. During integrity verification, if the current data fingerprint is inconsistent with the notarized fingerprint on the blockchain, local data freezing and maintenance alarms are immediately triggered. Cloud-based encrypted backup uses AWS S3 compatible storage, with a backup frequency of one incremental backup per day and one full backup per week. Backup data is stored using AES-256 encryption. Failed data is defined as "historical monitoring data that is more than one year old and has no fault association" and "normal sample data that has completed three cloud backups and has no access records for six consecutive months." The clearing process is performed according to the "old first, new last" principle to ensure that the local storage space utilization rate is kept below 70%.
[0060] Step S2 specifically involves: Each road segment client performs preprocessing on the heterogeneous monitoring data, including outlier handling, noise filtering, data alignment and synchronization, normalization, and feature engineering, to obtain standardized data; Based on railway industry standards, the standardized data are labeled with fault types and fault levels to obtain corresponding labeling information; The fault type labeling is based on "TB / T 2340-2012 Classification and Code of Railway Rail Damage", covering 8 core faults including cracks, wear, loose bolts, gauge deviation, rail bending, and unevenness. The fault level labeling is based on "Railway Line Maintenance Rules", divided into 4 levels: "Minor (does not affect train operation safety, requires regular observation), Moderate (affects train operation stability, requires handling within 15 days), Severe (threatens train operation safety, requires handling within 72 hours), and Critical (immediate interruption of train operation)". The labeling process is cross-reviewed by 2 senior maintenance engineers to ensure that the labeling accuracy rate is ≥99%.
[0061] By employing SMOTE oversampling technology or random undersampling technology, the proportion of fault samples to normal samples in each standardized dataset is balanced based on the annotation information, thereby constructing a local fault feature dataset. The outlier handling specifically involves: using statistical methods based on interquartile range or the 3σ criterion to identify and remove abnormal data points in the heterogeneous monitoring data caused by momentary sensor failure or external interference; and for continuous data in the heterogeneous monitoring data, using linear interpolation or the mean of the effective values before and after the data to fill in the gaps. When the sample size of heterogeneous monitoring data is ≥10000, the 3σ criterion is adopted (to adapt to normally distributed data); when the sample size is <10000, the interquartile range (IQR) method is adopted (with stronger resistance to extreme value interference); linear interpolation is preferred for continuous data filling (suitable for scenarios with stable data trends). If the data fluctuates greatly (such as vibration signals during train passage), the mean of the five effective values before and after is used for filling to ensure that the filled data conforms to the actual change trend.
[0062] The noise filtering specifically involves: using a Butterworth bandpass filter or wavelet threshold denoising method to filter out high-frequency noise and power frequency interference from the track vibration signal in the heterogeneous monitoring data, while retaining the effective frequency band signal related to the characteristics of rail faults; using a moving average filter to smooth the deformation data; and using a sliding window value filter method for the environmental parameters, while introducing physical constraint-based adaptive filtering and multi-sensor data verification. By using a preset instantaneous rate of change threshold, physically impossible abnormal jumps are identified and processed, and cross-validation is performed using the correlation between parameters. The Butterworth bandpass filter cutoff frequency is set to 10-500Hz (based on industry experimental data, rail fault characteristic signals are mainly distributed in this frequency band, which can filter out 50Hz power frequency interference and high-frequency noise >500Hz); the wavelet threshold denoising method uses the db4 wavelet basis (taking into account both time-frequency localization characteristics and computational efficiency), with 4 decomposition layers, and the threshold is adaptively calculated using an improved Birge-Massart strategy; the moving average filter window size is 5 (to adapt to the slow variation characteristics of deformation data); the moving average filter window size for environmental parameters is 3, and the instantaneous change rate threshold is set according to industry standards (e.g., instantaneous temperature change ≤5℃ / minute, instantaneous humidity change ≤10% RH / minute), and the cross-validation between parameters is based on the physical correlation model of "temperature-humidity-precipitation" (e.g., precipitation in high temperature and high humidity environments is positively correlated with the lateral displacement of the rail).
[0063] The data alignment and synchronization specifically involves: based on a unified first timestamp, performing time alignment on the heterogeneous monitoring data from sensors with different sampling frequencies, and unifying the data to the same time series through an interpolation resampling method to ensure the temporal correlation between the data; The normalization process specifically involves performing Z-score standardization or Min-Max normalization on the heterogeneous monitoring data to eliminate the influence of dimensions. The first timestamp uses UTC time format, accurate to the millisecond level; interpolation and resampling use linear interpolation (balancing accuracy and efficiency), and the unified time series sampling interval is 100ms (matching the sampling frequency characteristics of vibration signals); during normalization, vibration signals and deformation data are standardized using Z-score (adapting to scenarios where the data is normally distributed), and environmental parameters are normalized using Min-Max (mapping values to the [0,1] interval to adapt to non-normally distributed data) to eliminate the impact of dimensional differences on model training.
[0064] The feature engineering specifically involves: extracting time-domain features (extracting time-domain statistical features including mean, variance, peak value, kurtosis, and waveform factor from the heterogeneous monitoring data), frequency-domain features (performing fast Fourier transform on the track vibration signal to extract frequency-domain features such as amplitude, centroid, and mean square frequency of the main frequency components), and time-frequency-domain features (using wavelet transform or empirical mode decomposition to extract joint time-frequency features from the non-stationary vibration signal to capture transient information of the fault).
[0065] The time-domain feature extraction includes six statistical measures: mean, variance, peak value, kurtosis, waveform factor, and impulse factor. The frequency-domain features are extracted by converting the vibration signal to the frequency domain using Fast Fourier Transform (FFT), extracting four features: amplitude, frequency centroid, mean square frequency, and band energy of the top 20 main frequency components. The time-frequency domain features are further extracted using wavelet packet decomposition (three decomposition levels), extracting three features: energy entropy, singular value entropy, and approximate entropy for each sub-band. Each sample ultimately generates a 6+4+3=13 dimensional feature vector. During sample balancing, if the proportion of faulty samples is <10% (severe imbalance), SMOTE oversampling technology is used (to generate synthetic faulty samples and avoid information loss). If the proportion of faulty samples is 10%-30% (mild imbalance), random undersampling technology is used (to remove some normal samples and reduce computational load). After balancing, the ratio of faulty samples to normal samples is controlled between 1:3 and 1:5 (balancing model generalization ability and training efficiency).
[0066] Step S3 specifically involves: The central server initializes a global fault diagnosis module and a global output module; In specific implementation, the initialization parameters of the global fault diagnosis module and the global output module include: In the heterogeneous data adaptation unit of the global fault diagnosis module, the number of convolutional kernels in the wavelet convolutional layer is 64, and the kernel size is 3×3; the number of convolutional kernels in the spatial convolutional layer is 32, and the kernel size is 3×3; the number of convolutional kernels in the temporal convolutional layer is 16, and the kernel size is 1×5 (temporal window size 5); the number of output channels of the four parallel dilated convolutional layers in the multi-scale dilated convolutional block is 64; the number of output channels of the graph convolutional layer in the cross-segment feature fusion unit is 128; the window size of the Gaussian kernel smoothing layer in the feature calibration unit is 3; the residual connection block of the global output module adopts the bottleneck structure of ResNet-18, and the number of neurons in the fully connected layer of the shared feature layer is 512 and 256, respectively; the number of neurons in the MLP hidden layer of the fault type prediction branch and the fault level prediction branch is 128, 64, and 32, respectively, and all trainable parameters are initialized using Xavier (to avoid gradient vanishing or exploding).
[0067] The global fault diagnosis module is constructed based on a heterogeneous data adaptation unit, a multi-scale feature extraction unit, a cross-road segment feature fusion unit, and a feature calibration unit; the global output module is constructed based on a feature enhancement unit, a multi-task prediction unit, and a probability calibration unit. The initial global diagnostic parameters of the global fault diagnosis module and the initial global output parameters of the global output module are obtained. The initial global diagnostic parameters are a set that covers the weights and configuration parameters of all trainable components in the global fault diagnosis module, ensuring that the road segment client can initialize a local fault diagnosis module with basic feature extraction and fusion capabilities. The initial global output parameters are a set that covers the weights and calibration parameters of all trainable components in the global output module, ensuring that the road segment client can initialize a personalized output module for final fault identification and classification. Obtain the current second timestamp, calculate the hash value of the initial global diagnostic parameters, initial global output parameters, and second timestamp using the SM3 algorithm, encrypt the initial global diagnostic parameters, initial global output parameters, second timestamp, and hash value into an encrypted parameter packet using a preset SM4 raw key, retrieve the public key of each road segment client through the auxiliary server, encrypt the SM4 raw key with each public key to obtain the corresponding encryption key, and send the encrypted parameter packet and encryption key to the corresponding road segment client through an SSL encrypted transmission link.
[0068] The heterogeneous data adaptation unit is constructed based on vibration signal branch, deformation data branch, environmental parameter branch, and fusion subunit; The vibration signal branch is used to capture the time-frequency local features of non-stationary track vibration signals (such as the instantaneous high-frequency components of crack impact) through a wavelet convolution layer using wavelet basis functions (such as db4); replacing the ordinary convolution of traditional CNNs and adapting to the non-stationarity of vibration signals. The deformation data branch is used to capture the spatial correlation features of rail deformation (such as the local spread features of gauge deviation) from the deformation data through a spatial convolution layer with a 3×3 convolution kernel; and combines zero-filling to maintain the feature map size to adapt to the spatial continuity of the deformation data. The environmental parameter branch is used to capture the temporal variation trend characteristics of environmental parameters (such as the cumulative effect of a sudden drop in temperature on rail deformation) through a temporal convolution layer (TCL) using a 1×k convolution kernel (k is the temporal window size), thus adapting to the temporal dependence of environmental parameters. The fusion subunit is used to dynamically calculate the attention weights of the time-frequency local features, spatial correlation features, and temporal change trend features through a multi-head attention fusion layer (e.g., the vibration signal weight is higher than the environmental parameter in a fault scenario), and then weight and fuse them into a unified embedded feature. The multi-scale feature extraction unit is constructed based on multi-scale dilated convolutional blocks, channel attention subunits, and spatial attention subunits. The multi-scale dilated convolutional block is used to capture local and global features from the unified embedding features through four parallel dilated convolutional layers. The dilation rates of the four parallel dilated convolutional layers are 1, 2, 4, and 8, respectively. The small dilation rate (1, 2) captures local features, while the large dilation rate (4, 8) expands the receptive field to capture global features, avoiding the loss of details caused by pooling. The channel attention subunit is used to perform channel-dimensional "squeeze-excite" on the output of the multi-scale dilated convolutional block through the squeeze-excite (SE) module, thereby enhancing fault-sensitive channels (such as high-frequency channels of vibration signals) and suppressing redundant channels. The spatial attention subunit is used to introduce position information into the output of the channel attention subunit through the coordinate attention (CA) module, focus on the spatial location-related features of the fault occurrence (such as the local area of abnormal sleeper vibration amplitude), improve the spatial discriminativeness of the features, and output multi-scale fused features; The cross-segment feature fusion unit is constructed based on a graph construction subunit, a graph convolutional layer, and a knowledge distillation subunit. The graph construction subunit is used to extract a normalized adjacency matrix from the multi-scale fusion features through a dynamic federated graph generator; that is, using each road segment as a node and the similarity between road segments (a weighted sum of similarity in track type, years of operation, and distribution of environmental parameters) as edge weights to construct a normalized adjacency matrix. Each round of training updates the edge weights based on the latest data distribution to avoid insufficient adaptation of the static graph. The graph convolutional layer is used to aggregate features of adjacent or similar road segments in the normalized adjacency matrix through weighted graph convolution (e.g., if road segment A and road segment B have the same track type, their features are aggregated), and outputs a common fault feature matrix; the formula is: ;in, Represents the normalized adjacency matrix; Indicates the convolution weights; This indicates that the graph convolutional layer has passed through the first... The common fault feature matrix output after round operation; The knowledge distillation subunit is used to take the strong road segment features with abundant fault samples and high feature quality in the common fault feature matrix as teacher signals and distill them into the weak road segment features with scarce fault samples, so as to improve the feature expression ability of the weak road segment and output cross-road segment common features. The feature calibration unit is constructed based on a feature confidence evaluation subunit, an adaptive filtering subunit, and a feature smoothing subunit. The feature credibility assessment subunit is used to calculate the credibility score matrix of the common features of the railway tracks across the railway segments using a Mahalanobis distance calculator, based on the multi-scale fusion features and the global feature distribution statistics. Each element corresponds to the feature credibility of a certain railway segment and a certain batch of samples (the value range is [0,1], and the higher the score, the more closely the feature fits the global distribution and the less abnormal deviation). The global feature distribution statistics are the mean vector μ and covariance matrix Σ of the multi-scale fusion feature statistics of all railway tracks. The adaptive filtering subunit is used to set a dynamic soft threshold based on the confidence score matrix through the soft threshold filtering layer, and to perform weight suppression (rather than direct discard) on the common features across road segments, so as to avoid the loss of effective features due to misjudgment, and obtain a filtered common feature matrix (features of low confidence road segments are dynamically suppressed, and features of high confidence road segments are retained and strengthened). The feature smoothing subunit is used to perform Gaussian smoothing on the filtered common feature matrix through a Gaussian kernel smoothing layer to reduce feature fluctuations, improve the stability of common features, and output calibration common fault features. The feature enhancement unit is constructed based on residual connection blocks and task attention subunits; The residual connection block is used to enhance the common fault features of calibration through the ResNet bottleneck structure (1×1 convolution dimensionality reduction → 3×3 convolution feature extraction → 1×1 convolution dimensionality increase) to obtain residual enhanced common fault features; it solves the gradient vanishing problem in deep networks and preserves the original common feature information; The task attention subunit is used to dynamically adjust the feature enhancement weights of the residual enhanced common fault features based on the task differences between fault type prediction (emphasizing high-frequency features) and fault level prediction (emphasizing low-frequency cumulative features) through the dynamic attention layer, thereby improving the adaptability of features to tasks and obtaining task-enhanced common features. The multi-task prediction unit is constructed based on a shared feature layer, a fault type prediction branch, a fault level prediction branch, and a task interaction subunit. The shared feature layer is used to extract common dual-task shared basic features for fault type prediction and fault level prediction tasks from the common features of the task enhancement through two fully connected layers (with 512 and 256 hidden units respectively); it shields task-specific interference, reduces parameter redundancy, and retains the core common characteristics of rail faults (such as the frequency domain peak characteristics of vibration signals and the spatial deviation characteristics of deformation data). The fault type prediction branch is used to identify the probability distribution of fault types (such as cracks, wear, loose bolts, gauge deviation, etc., default 8 types) from the shared basic features of the dual tasks through 3-layer MLP and Softmax activation. The fault level prediction branch is used to identify the fault level probability distribution (minor, moderate, severe, critical, 4 ordered levels) from the shared basic features of the dual tasks through a 3-layer MLP and Sigmoid activation, and to adapt the ordinal characteristics of the level. The task interaction subunit is used to calculate the interaction loss value through the interaction loss function, and to backpropagate and update the network parameters of the shared feature layer, the fault type prediction branch and the fault level prediction branch based on the interaction loss value, so as to achieve synchronous improvement of the accuracy of the two tasks. The formula for the interaction loss function is: ; in, Indicates the interaction loss value; This represents the cosine similarity of the features from both tasks; α represents the type loss weight, defaulting to 0.5; β represents the feature similarity weight, defaulting to 0.1. Indicates the type prediction loss (calculated using cross-entropy loss); This represents the rank prediction loss (calculated using ordered logit loss, which aligns with the rank ordinal relationship). This represents the intermediate features of the type branch (output of the hidden layer in the middle of the branch); This represents the intermediate features of the hierarchical branches (output of the hidden layer in the middle of the branch). The probability calibration unit is constructed based on the distribution adaptation subunit and the probability correction subunit; The distribution adaptation subunit is used to perform temperature scaling on the fault type probability distribution and the fault level probability distribution. The probability correction subunit is used to perform nonlinear correction on the temperature-scaled fault type probability distribution and fault level probability distribution through the Beta calibration layer, thereby improving the prediction reliability of small sample fault types (such as rare rail bending faults) and outputting fault diagnosis results carrying fault type and fault level.
[0069] The model architecture of this invention has the following advantages: Customized embedding design for heterogeneous data: Based on the physical characteristics of the three types of heterogeneous data (vibration / deformation / environment) from railway track monitoring, branch embedding methods using wavelet convolution, spatial convolution, and temporal convolution are employed respectively. This overcomes the limitation of traditional federated learning where "a single embedding method adapts to all data," improving the targeting of feature extraction. For example, the non-stationarity of vibration signals is accurately captured through wavelet convolution, and the spatial correlation of deformation data is preserved through spatial convolution, closely aligning with the actual data characteristics of railway fault diagnosis.
[0070] Cross-segment fusion using federated dynamic graph convolution: A dynamic federated graph is constructed based on segment similarity. Common features of similar segments are aggregated through graph convolution, while knowledge distillation is introduced to improve the feature quality of weakly data segments, overcoming the bottleneck of traditional federated learning which only uses parameter weighting and ignores semantic relationships between segments. For example, fault features of adjacent segments or segments with the same track type can mutually enhance each other, adapting to scenarios with uneven data distribution across multiple segments (such as scarce fault samples in some segments).
[0071] Full-link Byzantine robustness design: Not only does it adopt a weighted strategy in the parameter aggregation stage of the central server, but it also adds a Byzantine robustness calibration unit (feature credibility assessment based on Mahalanobis distance + soft threshold filtering) at the feature level to suppress the interference of abnormal nodes from the source, solve the industry pain point of "abnormal or malicious attacks on some road segments" in multi-segment collaborative diagnosis, and improve the robustness of the model.
[0072] Task-oriented multi-task prediction and federated calibration: Adopting a "hard sharing-soft separation" multi-task architecture, it simultaneously predicts fault types and levels, and improves collaboration by combining task interaction losses; Dynamic temperature scaling is introduced in the probabilistic calibration stage to adapt to the heterogeneity of data distribution in federated scenarios, solve the bias problem caused by the "globally unified parameters" of traditional calibration methods, and meet the actual needs of the railway industry for "accurate classification and graded early warning".
[0073] Step S4 specifically involves: Each client section receives the encrypted parameter packet and encryption key sent by the central server, calls the locally stored private key, decrypts the received encryption key to obtain the SM4 original key, and then decrypts the received encrypted parameter packet using the SM4 original key to obtain the initial global diagnostic parameters, initial global output parameters, second timestamp, and hash value. After performing integrity verification on the initial global diagnostic parameters, initial global output parameters, and second timestamp based on the hash value, a timeliness verification is performed based on the second timestamp. If the verification passes, the parameter reception is completed; if the verification fails, a retransmission request is sent to the central server. The central server, in conjunction with the access logs of the auxiliary server, investigates transmission anomalies and retransmits the data. Each road segment client initializes a local fault diagnosis module with the same network architecture as the global fault diagnosis module based on the initial global diagnostic parameters, and initializes a personalized output module with the same architecture as the global output module based on the initial global output parameters. A local fault diagnosis model is then constructed based on the local fault diagnosis module and the personalized output module.
[0074] Step S5 specifically involves: Each road segment client uses a 30-day base time window and initially divides the fault feature dataset into a training set and a validation set in a 7:3 ratio to ensure that the training set covers samples of all fault types. Calculate the fault distribution entropy of the validation set. If the fault distribution entropy is less than a preset entropy threshold, dynamically adjust the ratio of the training set to the validation set within the basic time window, prioritizing ensuring that the validation set contains at least 20 valid samples for each type of fault. The formula for calculating the fault distribution entropy is: ;in, The fault distribution entropy is represented by n; n represents the total number of fault types. This represents the proportion of samples of the i-th type of fault; the entropy threshold is 1.2, and a fault distribution entropy less than 1.2 indicates that the fault distribution is extremely unbalanced; The training and validation sets are subjected to fault scenario-driven online data augmentation, specifically as follows: For the sample of track vibration signal, preset amplitude scaling (scaling factor 0.8-1.2), time stretching (stretching ratio 0.9-1.1), and Gaussian noise superposition (noise intensity ≤ 5% of the original signal amplitude) are applied based on physical simulation rules to simulate the changes in fault signal under different train loads and speeds. For the deformation data samples, combined with the historical distribution of local environmental parameters (such as temperature and humidity), deformation derivative samples under different environmental couplings are generated through linear transformation (such as reasonable derivative samples where the lateral displacement of the rail increases by 0.1 mm for every 10°C increase in temperature). For samples of environmental parameters, virtual samples of the synergistic effects of multiple environmental factors (such as composite scene samples of "high temperature + high humidity + moderate wind") are generated by permutation and combination to make up for the lack of local composite scene samples. The local fault diagnosis model is trained locally using the online data-augmented training set, and the trained local fault diagnosis model is validated using the online data-augmented validation set. After every 5 rounds of training, the base time window is updated (slid forward by 1 day), and the training set and validation set are re-divided to avoid the decline in model generalization ability due to data timeliness, and to adapt to the cumulative characteristics of railway faults over time, until the preset number of training times is completed. After training is completed, the quality assessment index of the training data in this round is calculated. The local update parameters of the local fault diagnosis module and the quality assessment index are encrypted using a homomorphic encryption algorithm and then uploaded to the central server through an SSL encrypted transmission link. The quality assessment index includes the number of samples, inter-class balance, and stability relative to the distribution of local historical data.
[0075] The local fault diagnosis model is trained using the Adam optimizer with an initial learning rate of 0.001, which decays by 50% every 10 rounds. The training batch size is set to 32 (to adapt to the hardware computing power of the road segment clients), and the preset number of training rounds is 50. After every 5 rounds of training, the base time window is shifted forward by 1 day (removing the earliest day's data and adding the latest day's data), and the training and validation sets are re-divided to ensure that the model adapts to the temporal changes of the data. In the quality evaluation metrics, the number of samples is counted as the number of effective samples (after removing outliers), the inter-class balance is measured using the Gini coefficient (the closer the Gini coefficient is to 0, the more balanced the distribution), and the stability is measured using the KL divergence between the current training set and the historical training sets (when the KL divergence is <0.1, the distribution is considered stable). The homomorphic encryption algorithm used is the Paillier algorithm (supporting additive homomorphism and adapting to parameter aggregation operations). The encrypted parameters are uploaded via an SSL encrypted transmission link, with the upload bandwidth controlled within 1Mbps to avoid consuming too many communication resources.
[0076] In step S6, the weighted aggregation strategy with Byzantine robustness specifically involves: the central server calculating the aggregation weight of each local update parameter based on the quality assessment indicators uploaded by the clients of each road segment, using a trimmed average algorithm to remove extreme abnormal parameters from the local update parameters, and then performing a weighted summation of the effective parameters in the local update parameters based on the aggregation weight.
[0077] The formula for calculating the aggregate weight is: ; in, This represents the aggregate weight of the client in the i-th road segment; This represents the number of valid samples from the client in the i-th road segment; This represents the total number of valid client samples across all road segments; This represents the Gini coefficient for inter-class balance of the i-th client; Let KL divergence represent the distribution stability of the i-th client. These are all weighting coefficients, with default values of 0.5, 0.3, and 0.2.
[0078] A preferred embodiment of the multi-segment railway fault collaborative diagnosis system based on federated learning according to the present invention includes the following modules: The heterogeneous monitoring data acquisition module is used to deploy on the client side of different railway sections to collect heterogeneous monitoring data of the railway section, including at least track vibration signals, deformation data and environmental parameters, and store them locally with encryption. The fault feature dataset construction module is used by clients of each road segment to preprocess and label the heterogeneous monitoring data in order to construct a local fault feature dataset. The server model initialization module is used by the central server to initialize a global fault diagnosis module and a global output module, and to encrypt and send the initial global diagnosis parameters of the global fault diagnosis module and the initial global output parameters of the global output module to the clients of each road segment. The local fault diagnosis model initialization module is used by each road segment client to initialize a local fault diagnosis module based on the initial global diagnosis parameters, initialize a personalized output module based on the initial global output parameters, and construct a local fault diagnosis model based on the local fault diagnosis module and the personalized output module. The local training module is used by each road segment client to train the local fault diagnosis model locally based on the fault feature dataset. After training is completed, the quality evaluation index of the training data in this round is calculated, and the local update parameters of the local fault diagnosis module and the quality evaluation index are encrypted and uploaded to the central server. The parameter aggregation module is used by the central server to adopt a weighted aggregation strategy with Byzantine robustness, calculate the aggregation weight based on the quality assessment indicators uploaded by each road segment client, aggregate each locally updated parameter based on the aggregation weight, update the global fault diagnosis module to obtain optimized global diagnosis parameters, encrypt and send the optimized global diagnosis parameters to each road segment client, repeat the local training and parameter aggregation steps until the global fault diagnosis module reaches the preset convergence condition or training round, and encrypt and send the latest optimized global diagnosis parameters to the road segment client to update the local fault diagnosis module; The rail fault diagnosis module is used by clients of each section to identify and classify faults in local real-time heterogeneous monitoring data using the latest local fault diagnosis model, and output real-time fault diagnosis results to achieve collaborative diagnosis of rail faults across multiple sections.
[0079] The heterogeneous monitoring data acquisition module is specifically used for: The client-side components deployed on different sections of the railway track collect heterogeneous monitoring data of the track section via sensor arrays, including at least track vibration signals, deformation data, and environmental parameters. The track vibration signals include at least the vertical vibration acceleration of the rail, the lateral vibration acceleration of the rail, the sleeper vibration amplitude, and the vibration frequency. The deformation data includes at least the vertical displacement of the rail, the lateral displacement of the rail, the gauge deviation, and the rail curvature. The environmental parameters include at least the ambient temperature, ambient humidity, precipitation, wind speed, and dust concentration. The sampling frequency of the track vibration signal is set to 1024Hz (to meet the high-frequency capture requirement of rail crack impact signal), the sampling frequency of deformation data is 10Hz (to adapt to the slow change characteristics of rail deformation), and the sampling frequency of environmental parameters is 1Hz (to balance data volume and timeliness). During the acquisition process, the sensor’s built-in self-calibration module performs zero-point calibration once per hour to avoid data deviation caused by drift.
[0080] The road segment client will collect the heterogeneous monitoring data and store it in a local database encrypted with the AES-256 encryption algorithm.
[0081] The database has a role-based hierarchical authorization mechanism, which only assigns corresponding operation permissions to authorized operation and maintenance personnel, and combines multi-factor authentication to prevent unauthorized access. At the same time, it records all data access, data modification and data export operation logs and retains the operation logs for at least 90 days. In practice, the database uses MySQL 8.0 Enterprise Edition with Transparent Data Encryption (TDE) enabled. AES-256 keys are stored and managed through a hardware security module (HSM, selected as Gemalto SafeNet Luna 7000), and the keys are automatically rotated every 90 days. The role-based authorization mechanism is divided into three levels: "administrator, operation and maintenance operator, and auditor". Administrators have full permissions (including key management), operation and maintenance operators can only read data and export diagnostic results, and auditors can only view operation logs and have no data modification permissions.
[0082] When the road segment client updates each of the databases, it automatically calculates the data fingerprint of the database based on the HASH-256 algorithm and stores the data fingerprint on the blockchain. Before updating the database, it performs an integrity check based on the data fingerprint stored on the blockchain in the previous time, and periodically performs cloud-encrypted backup of the data stored in the database. Based on the storage space of the road segment client and the preset clearing rules, it continuously clears the invalid data in the database.
[0083] In practice, data fingerprint calculation is performed on a daily incremental data basis in the database. After generating a unique fingerprint using the HASH-256 algorithm, a blockchain smart contract is invoked to complete the notarization. The notarized information includes the data fingerprint, road segment number, and update timestamp. During integrity verification, if the current data fingerprint is inconsistent with the notarized fingerprint on the blockchain, local data freezing and maintenance alarms are immediately triggered. Cloud-based encrypted backup uses AWS S3 compatible storage, with a backup frequency of one incremental backup per day and one full backup per week. Backup data is stored using AES-256 encryption. Failed data is defined as "historical monitoring data that is more than one year old and has no fault association" and "normal sample data that has completed three cloud backups and has no access records for six consecutive months." The clearing process is performed according to the "old first, new last" principle to ensure that the local storage space utilization rate is kept below 70%.
[0084] The fault feature dataset construction module is specifically used for: Each road segment client performs preprocessing on the heterogeneous monitoring data, including outlier handling, noise filtering, data alignment and synchronization, normalization, and feature engineering, to obtain standardized data; Based on railway industry standards, the standardized data are labeled with fault types and fault levels to obtain corresponding labeling information; The fault type labeling is based on "TB / T 2340-2012 Classification and Code of Railway Rail Damage", covering 8 core faults including cracks, wear, loose bolts, gauge deviation, rail bending, and unevenness. The fault level labeling is based on "Railway Line Maintenance Rules", divided into 4 levels: "Minor (does not affect train operation safety, requires regular observation), Moderate (affects train operation stability, requires handling within 15 days), Severe (threatens train operation safety, requires handling within 72 hours), and Critical (immediate interruption of train operation)". The labeling process is cross-reviewed by 2 senior maintenance engineers to ensure that the labeling accuracy rate is ≥99%.
[0085] By employing SMOTE oversampling technology or random undersampling technology, the proportion of fault samples to normal samples in each standardized dataset is balanced based on the annotation information, thereby constructing a local fault feature dataset. The outlier handling specifically involves: using statistical methods based on interquartile range or the 3σ criterion to identify and remove abnormal data points in the heterogeneous monitoring data caused by momentary sensor failure or external interference; and for continuous data in the heterogeneous monitoring data, using linear interpolation or the mean of the effective values before and after the data to fill in the gaps. When the sample size of heterogeneous monitoring data is ≥10000, the 3σ criterion is adopted (to adapt to normally distributed data); when the sample size is <10000, the interquartile range (IQR) method is adopted (with stronger resistance to extreme value interference); linear interpolation is preferred for continuous data filling (suitable for scenarios with stable data trends). If the data fluctuates greatly (such as vibration signals during train passage), the mean of the five effective values before and after is used for filling to ensure that the filled data conforms to the actual change trend.
[0086] The noise filtering specifically involves: using a Butterworth bandpass filter or wavelet threshold denoising method to filter out high-frequency noise and power frequency interference from the track vibration signal in the heterogeneous monitoring data, while retaining the effective frequency band signal related to the characteristics of rail faults; using a moving average filter to smooth the deformation data; and using a sliding window value filter method for the environmental parameters, while introducing physical constraint-based adaptive filtering and multi-sensor data verification. By using a preset instantaneous rate of change threshold, physically impossible abnormal jumps are identified and processed, and cross-validation is performed using the correlation between parameters. The Butterworth bandpass filter cutoff frequency is set to 10-500Hz (based on industry experimental data, rail fault characteristic signals are mainly distributed in this frequency band, which can filter out 50Hz power frequency interference and high-frequency noise >500Hz); the wavelet threshold denoising method uses the db4 wavelet basis (taking into account both time-frequency localization characteristics and computational efficiency), with 4 decomposition layers, and the threshold is adaptively calculated using an improved Birge-Massart strategy; the moving average filter window size is 5 (to adapt to the slow variation characteristics of deformation data); the moving average filter window size for environmental parameters is 3, and the instantaneous change rate threshold is set according to industry standards (e.g., instantaneous temperature change ≤5℃ / minute, instantaneous humidity change ≤10% RH / minute), and the cross-validation between parameters is based on the physical correlation model of "temperature-humidity-precipitation" (e.g., precipitation in high temperature and high humidity environments is positively correlated with the lateral displacement of the rail).
[0087] The data alignment and synchronization specifically involves: based on a unified first timestamp, performing time alignment on the heterogeneous monitoring data from sensors with different sampling frequencies, and unifying the data to the same time series through an interpolation resampling method to ensure the temporal correlation between the data; The normalization process specifically involves performing Z-score standardization or Min-Max normalization on the heterogeneous monitoring data to eliminate the influence of dimensions. The first timestamp uses UTC time format, accurate to the millisecond level; interpolation and resampling use linear interpolation (balancing accuracy and efficiency), and the unified time series sampling interval is 100ms (matching the sampling frequency characteristics of vibration signals); during normalization, vibration signals and deformation data are standardized using Z-score (adapting to scenarios where the data is normally distributed), and environmental parameters are normalized using Min-Max (mapping values to the [0,1] interval to adapt to non-normally distributed data) to eliminate the impact of dimensional differences on model training.
[0088] The feature engineering specifically involves: extracting time-domain features (extracting time-domain statistical features including mean, variance, peak value, kurtosis, and waveform factor from the heterogeneous monitoring data), frequency-domain features (performing fast Fourier transform on the track vibration signal to extract frequency-domain features such as amplitude, centroid, and mean square frequency of the main frequency components), and time-frequency-domain features (using wavelet transform or empirical mode decomposition to extract joint time-frequency features from the non-stationary vibration signal to capture transient information of the fault).
[0089] The time-domain feature extraction includes six statistical measures: mean, variance, peak value, kurtosis, waveform factor, and impulse factor. The frequency-domain features are extracted by converting the vibration signal to the frequency domain using Fast Fourier Transform (FFT), extracting four features: amplitude, frequency centroid, mean square frequency, and band energy of the top 20 main frequency components. The time-frequency domain features are further extracted using wavelet packet decomposition (three decomposition levels), extracting three features: energy entropy, singular value entropy, and approximate entropy for each sub-band. Each sample ultimately generates a 6+4+3=13 dimensional feature vector. During sample balancing, if the proportion of faulty samples is <10% (severe imbalance), SMOTE oversampling technology is used (to generate synthetic faulty samples and avoid information loss). If the proportion of faulty samples is 10%-30% (mild imbalance), random undersampling technology is used (to remove some normal samples and reduce computational load). After balancing, the ratio of faulty samples to normal samples is controlled between 1:3 and 1:5 (balancing model generalization ability and training efficiency).
[0090] The server model initialization module is specifically used for: The central server initializes a global fault diagnosis module and a global output module; In specific implementation, the initialization parameters of the global fault diagnosis module and the global output module include: In the heterogeneous data adaptation unit of the global fault diagnosis module, the number of convolutional kernels in the wavelet convolutional layer is 64, and the kernel size is 3×3; the number of convolutional kernels in the spatial convolutional layer is 32, and the kernel size is 3×3; the number of convolutional kernels in the temporal convolutional layer is 16, and the kernel size is 1×5 (temporal window size 5); the number of output channels of the four parallel dilated convolutional layers in the multi-scale dilated convolutional block is 64; the number of output channels of the graph convolutional layer in the cross-segment feature fusion unit is 128; the window size of the Gaussian kernel smoothing layer in the feature calibration unit is 3; the residual connection block of the global output module adopts the bottleneck structure of ResNet-18, and the number of neurons in the fully connected layer of the shared feature layer is 512 and 256, respectively; the number of neurons in the MLP hidden layer of the fault type prediction branch and the fault level prediction branch is 128, 64, and 32, respectively, and all trainable parameters are initialized using Xavier (to avoid gradient vanishing or exploding).
[0091] The global fault diagnosis module is constructed based on a heterogeneous data adaptation unit, a multi-scale feature extraction unit, a cross-road segment feature fusion unit, and a feature calibration unit; the global output module is constructed based on a feature enhancement unit, a multi-task prediction unit, and a probability calibration unit. The initial global diagnostic parameters of the global fault diagnosis module and the initial global output parameters of the global output module are obtained. The initial global diagnostic parameters are a set that covers the weights and configuration parameters of all trainable components in the global fault diagnosis module, ensuring that the road segment client can initialize a local fault diagnosis module with basic feature extraction and fusion capabilities. The initial global output parameters are a set that covers the weights and calibration parameters of all trainable components in the global output module, ensuring that the road segment client can initialize a personalized output module for final fault identification and classification. Obtain the current second timestamp, calculate the hash value of the initial global diagnostic parameters, initial global output parameters, and second timestamp using the SM3 algorithm, encrypt the initial global diagnostic parameters, initial global output parameters, second timestamp, and hash value into an encrypted parameter packet using a preset SM4 raw key, retrieve the public key of each road segment client through the auxiliary server, encrypt the SM4 raw key with each public key to obtain the corresponding encryption key, and send the encrypted parameter packet and encryption key to the corresponding road segment client through an SSL encrypted transmission link.
[0092] The heterogeneous data adaptation unit is constructed based on vibration signal branch, deformation data branch, environmental parameter branch, and fusion subunit; The vibration signal branch is used to capture the time-frequency local features of non-stationary track vibration signals (such as the instantaneous high-frequency components of crack impact) through a wavelet convolution layer using wavelet basis functions (such as db4); replacing the ordinary convolution of traditional CNNs and adapting to the non-stationarity of vibration signals. The deformation data branch is used to capture the spatial correlation features of rail deformation (such as the local spread features of gauge deviation) from the deformation data through a spatial convolution layer with a 3×3 convolution kernel; and combines zero-filling to maintain the feature map size to adapt to the spatial continuity of the deformation data. The environmental parameter branch is used to capture the temporal variation trend characteristics of environmental parameters (such as the cumulative effect of a sudden drop in temperature on rail deformation) through a temporal convolution layer (TCL) using a 1×k convolution kernel (k is the temporal window size), thus adapting to the temporal dependence of environmental parameters. The fusion subunit is used to dynamically calculate the attention weights of the time-frequency local features, spatial correlation features, and temporal change trend features through a multi-head attention fusion layer (e.g., the vibration signal weight is higher than the environmental parameter in a fault scenario), and then weight and fuse them into a unified embedded feature. The multi-scale feature extraction unit is constructed based on multi-scale dilated convolutional blocks, channel attention subunits, and spatial attention subunits. The multi-scale dilated convolutional block is used to capture local and global features from the unified embedding features through four parallel dilated convolutional layers. The dilation rates of the four parallel dilated convolutional layers are 1, 2, 4, and 8, respectively. The small dilation rate (1, 2) captures local features, while the large dilation rate (4, 8) expands the receptive field to capture global features, avoiding the loss of details caused by pooling. The channel attention subunit is used to perform channel-dimensional "squeeze-excite" on the output of the multi-scale dilated convolutional block through the squeeze-excite (SE) module, thereby enhancing fault-sensitive channels (such as high-frequency channels of vibration signals) and suppressing redundant channels. The spatial attention subunit is used to introduce position information into the output of the channel attention subunit through the coordinate attention (CA) module, focus on the spatial location-related features of the fault occurrence (such as the local area of abnormal sleeper vibration amplitude), improve the spatial discriminativeness of the features, and output multi-scale fused features; The cross-segment feature fusion unit is constructed based on a graph construction subunit, a graph convolutional layer, and a knowledge distillation subunit. The graph construction subunit is used to extract a normalized adjacency matrix from the multi-scale fusion features through a dynamic federated graph generator; that is, using each road segment as a node and the similarity between road segments (a weighted sum of similarity in track type, years of operation, and distribution of environmental parameters) as edge weights to construct a normalized adjacency matrix. Each round of training updates the edge weights based on the latest data distribution to avoid insufficient adaptation of the static graph. The graph convolutional layer is used to aggregate features of adjacent or similar road segments in the normalized adjacency matrix through weighted graph convolution (e.g., if road segment A and road segment B have the same track type, their features are aggregated), and outputs a common fault feature matrix; the formula is: ;in, Represents the normalized adjacency matrix; Indicates the convolution weights; This indicates that the graph convolutional layer has passed through the first... The common fault feature matrix output after round operation; The knowledge distillation subunit is used to take the strong road segment features with abundant fault samples and high feature quality in the common fault feature matrix as teacher signals and distill them into the weak road segment features with scarce fault samples, so as to improve the feature expression ability of the weak road segment and output cross-road segment common features. The feature calibration unit is constructed based on a feature confidence evaluation subunit, an adaptive filtering subunit, and a feature smoothing subunit. The feature credibility assessment subunit is used to calculate the credibility score matrix of the common features of the railway tracks across the railway segments using a Mahalanobis distance calculator, based on the multi-scale fusion features and the global feature distribution statistics. Each element corresponds to the feature credibility of a certain railway segment and a certain batch of samples (the value range is [0,1], and the higher the score, the more closely the feature fits the global distribution and the less abnormal deviation). The global feature distribution statistics are the mean vector μ and covariance matrix Σ of the multi-scale fusion feature statistics of all railway tracks. The adaptive filtering subunit is used to set a dynamic soft threshold based on the confidence score matrix through the soft threshold filtering layer, and to perform weight suppression (rather than direct discard) on the common features across road segments, so as to avoid the loss of effective features due to misjudgment, and obtain a filtered common feature matrix (features of low confidence road segments are dynamically suppressed, and features of high confidence road segments are retained and strengthened). The feature smoothing subunit is used to perform Gaussian smoothing on the filtered common feature matrix through a Gaussian kernel smoothing layer to reduce feature fluctuations, improve the stability of common features, and output calibration common fault features. The feature enhancement unit is constructed based on residual connection blocks and task attention subunits; The residual connection block is used to enhance the common fault features of calibration through the ResNet bottleneck structure (1×1 convolution dimensionality reduction → 3×3 convolution feature extraction → 1×1 convolution dimensionality increase) to obtain residual enhanced common fault features; it solves the gradient vanishing problem in deep networks and preserves the original common feature information; The task attention subunit is used to dynamically adjust the feature enhancement weights of the residual enhanced common fault features based on the task differences between fault type prediction (emphasizing high-frequency features) and fault level prediction (emphasizing low-frequency cumulative features) through the dynamic attention layer, thereby improving the adaptability of features to tasks and obtaining task-enhanced common features. The multi-task prediction unit is constructed based on a shared feature layer, a fault type prediction branch, a fault level prediction branch, and a task interaction subunit. The shared feature layer is used to extract common dual-task shared basic features for fault type prediction and fault level prediction tasks from the common features of the task enhancement through two fully connected layers (with 512 and 256 hidden units respectively); it shields task-specific interference, reduces parameter redundancy, and retains the core common characteristics of rail faults (such as the frequency domain peak characteristics of vibration signals and the spatial deviation characteristics of deformation data). The fault type prediction branch is used to identify the probability distribution of fault types (such as cracks, wear, loose bolts, gauge deviation, etc., default 8 types) from the shared basic features of the dual tasks through 3-layer MLP and Softmax activation. The fault level prediction branch is used to identify the fault level probability distribution (minor, moderate, severe, critical, 4 ordered levels) from the shared basic features of the dual tasks through a 3-layer MLP and Sigmoid activation, and to adapt the ordinal characteristics of the level. The task interaction subunit is used to calculate the interaction loss value through the interaction loss function, and to backpropagate and update the network parameters of the shared feature layer, the fault type prediction branch and the fault level prediction branch based on the interaction loss value, so as to achieve synchronous improvement of the accuracy of the two tasks. The formula for the interaction loss function is: ; in, Indicates the interaction loss value; This represents the cosine similarity of the features from both tasks; α represents the type loss weight, with a default of 0.5; β represents the feature similarity weight, with a default of 0.1. Indicates the type prediction loss (calculated using cross-entropy loss); This represents the rank prediction loss (calculated using ordered logit loss, which aligns with the rank ordinal relationship). This represents the intermediate features of the type branch (output of the hidden layer in the middle of the branch); This represents the intermediate features of the hierarchical branches (output of the hidden layer in the middle of the branch). The probability calibration unit is constructed based on the distribution adaptation subunit and the probability correction subunit; The distribution adaptation subunit is used to perform temperature scaling on the fault type probability distribution and the fault level probability distribution. The probability correction subunit is used to perform nonlinear correction on the temperature-scaled fault type probability distribution and fault level probability distribution through the Beta calibration layer, thereby improving the prediction reliability of small sample fault types (such as rare rail bending faults) and outputting fault diagnosis results carrying fault type and fault level.
[0093] The model architecture of this invention has the following advantages: Customized embedding design for heterogeneous data: Based on the physical characteristics of the three types of heterogeneous data (vibration / deformation / environment) from railway track monitoring, branch embedding methods using wavelet convolution, spatial convolution, and temporal convolution are employed respectively. This overcomes the limitation of traditional federated learning where "a single embedding method adapts to all data," improving the targeting of feature extraction. For example, the non-stationarity of vibration signals is accurately captured through wavelet convolution, and the spatial correlation of deformation data is preserved through spatial convolution, closely aligning with the actual data characteristics of railway fault diagnosis.
[0094] Cross-segment fusion using federated dynamic graph convolution: A dynamic federated graph is constructed based on segment similarity. Common features of similar segments are aggregated through graph convolution, while knowledge distillation is introduced to improve the feature quality of weakly data segments, overcoming the bottleneck of traditional federated learning which only uses parameter weighting and ignores semantic relationships between segments. For example, fault features of adjacent segments or segments with the same track type can mutually enhance each other, adapting to scenarios with uneven data distribution across multiple segments (such as scarce fault samples in some segments).
[0095] Full-link Byzantine robustness design: Not only does it adopt a weighted strategy in the parameter aggregation stage of the central server, but it also adds a Byzantine robustness calibration unit (feature credibility assessment based on Mahalanobis distance + soft threshold filtering) at the feature level to suppress the interference of abnormal nodes from the source, solve the industry pain point of "abnormal or malicious attacks on some road segments" in multi-segment collaborative diagnosis, and improve the robustness of the model.
[0096] Task-oriented multi-task prediction and federated calibration: Adopting a "hard sharing-soft separation" multi-task architecture, it simultaneously predicts fault types and levels, and improves collaboration by combining task interaction losses; Dynamic temperature scaling is introduced in the probabilistic calibration stage to adapt to the heterogeneity of data distribution in federated scenarios, solve the bias problem caused by the "globally unified parameters" of traditional calibration methods, and meet the actual needs of the railway industry for "accurate classification and graded early warning".
[0097] The local fault diagnosis model initialization module is specifically used for: Each client section receives the encrypted parameter packet and encryption key sent by the central server, calls the locally stored private key, decrypts the received encryption key to obtain the SM4 original key, and then decrypts the received encrypted parameter packet using the SM4 original key to obtain the initial global diagnostic parameters, initial global output parameters, second timestamp, and hash value. After performing integrity verification on the initial global diagnostic parameters, initial global output parameters, and second timestamp based on the hash value, a timeliness verification is performed based on the second timestamp. If the verification passes, the parameter reception is completed; if the verification fails, a retransmission request is sent to the central server. The central server, in conjunction with the access logs of the auxiliary server, investigates transmission anomalies and retransmits the data. Each road segment client initializes a local fault diagnosis module with the same network architecture as the global fault diagnosis module based on the initial global diagnostic parameters, and initializes a personalized output module with the same architecture as the global output module based on the initial global output parameters. A local fault diagnosis model is then constructed based on the local fault diagnosis module and the personalized output module.
[0098] The local training module is specifically used for: Each road segment client uses a 30-day base time window and initially divides the fault feature dataset into a training set and a validation set in a 7:3 ratio to ensure that the training set covers samples of all fault types. Calculate the fault distribution entropy of the validation set. If the fault distribution entropy is less than a preset entropy threshold, dynamically adjust the ratio of the training set to the validation set within the basic time window, prioritizing ensuring that the validation set contains at least 20 valid samples for each type of fault. The formula for calculating the fault distribution entropy is: ;in, The fault distribution entropy is represented by n; n represents the total number of fault types. This represents the proportion of samples of the i-th type of fault; the entropy threshold is 1.2, and a fault distribution entropy less than 1.2 indicates that the fault distribution is extremely unbalanced; The training and validation sets are subjected to fault scenario-driven online data augmentation, specifically as follows: For the sample of track vibration signal, preset amplitude scaling (scaling factor 0.8-1.2), time stretching (stretching ratio 0.9-1.1), and Gaussian noise superposition (noise intensity ≤ 5% of the original signal amplitude) are applied based on physical simulation rules to simulate the changes in fault signal under different train loads and speeds. For the deformation data samples, combined with the historical distribution of local environmental parameters (such as temperature and humidity), deformation derivative samples under different environmental couplings are generated through linear transformation (such as reasonable derivative samples where the lateral displacement of the rail increases by 0.1 mm for every 10°C increase in temperature). For samples of environmental parameters, virtual samples of the synergistic effects of multiple environmental factors (such as composite scene samples of "high temperature + high humidity + moderate wind") are generated by permutation and combination to make up for the lack of local composite scene samples. The local fault diagnosis model is trained locally using the online data-augmented training set, and the trained local fault diagnosis model is validated using the online data-augmented validation set. After every 5 rounds of training, the base time window is updated (slid forward by 1 day), and the training set and validation set are re-divided to avoid the decline in model generalization ability due to data timeliness, and to adapt to the cumulative characteristics of railway faults over time, until the preset number of training times is completed. After training is completed, the quality assessment index of the training data in this round is calculated. The local update parameters of the local fault diagnosis module and the quality assessment index are encrypted using a homomorphic encryption algorithm and then uploaded to the central server through an SSL encrypted transmission link. The quality assessment index includes the number of samples, inter-class balance, and stability relative to the distribution of local historical data.
[0099] The local fault diagnosis model is trained using the Adam optimizer with an initial learning rate of 0.001, which decays by 50% every 10 rounds. The training batch size is set to 32 (to adapt to the hardware computing power of the road segment clients), and the preset number of training rounds is 50. After every 5 rounds of training, the base time window is shifted forward by 1 day (removing the earliest day's data and adding the latest day's data), and the training and validation sets are re-divided to ensure that the model adapts to the temporal changes of the data. In the quality evaluation metrics, the number of samples is counted as the number of effective samples (after removing outliers), the inter-class balance is measured using the Gini coefficient (the closer the Gini coefficient is to 0, the more balanced the distribution), and the stability is measured using the KL divergence between the current training set and the historical training sets (when the KL divergence is <0.1, the distribution is considered stable). The homomorphic encryption algorithm used is the Paillier algorithm (supporting additive homomorphism and adapting to parameter aggregation operations). The encrypted parameters are uploaded via an SSL encrypted transmission link, with the upload bandwidth controlled within 1Mbps to avoid consuming too many communication resources.
[0100] In the parameter aggregation module, the weighted aggregation strategy with Byzantine robustness specifically involves: the central server calculating the aggregation weight of each local update parameter based on the quality assessment indicators uploaded by the clients of each road segment, using a trimmed average algorithm to remove extreme abnormal parameters from the local update parameters, and then performing a weighted summation of the effective parameters in the local update parameters based on the aggregation weight.
[0101] The formula for calculating the aggregate weight is: ; in, This represents the aggregate weight of the client in the i-th road segment; This represents the number of valid samples from the client in the i-th road segment; This represents the total number of valid client samples across all road segments; This represents the Gini coefficient for inter-class balance of the i-th client; Let KL divergence represent the distribution stability of the i-th client. These are all weighting coefficients, with default values of 0.5, 0.3, and 0.2.
[0102] Experimental verification: To verify the technical effectiveness of this invention, three different types of railway sections (high-speed main line, conventional freight line, and mountain branch line) were selected for field testing. One client machine was deployed on each section, and six months of heterogeneous monitoring data (including 5,000+ fault samples, covering 8 fault types and 4 fault levels) were collected. The data was compared with existing technologies (single-section local model and centralized diagnostic model). The results are as follows: 1. Fault Diagnosis Accuracy: The average fault diagnosis accuracy of this invention is 95.2%, with 96.8% for high-speed main lines, 94.5% for conventional freight lines, and 94.3% for mountain branch lines. The average accuracy of the local model for a single section is 82.7% (only 78.1% for mountain branch lines), and the average accuracy of the centralized diagnosis model is 90.3% (due to data silos, data for mountain branch lines was not included in the training). This invention significantly improves the cross-section diagnosis accuracy through federated learning and collaborative training, especially improving the diagnosis performance of sections with scarce samples.
[0103] 2. Data privacy protection: All original data in this invention is stored locally, and only encrypted model parameters are transmitted (the amount of parameter data is only 0.1% of the original data). According to third-party security assessment, no data leakage or tampering has occurred, which complies with the privacy protection standards of the railway industry. The centralized diagnostic model has the risk of leakage during data transmission (the amount of transmitted data is 100% of the original data), and the single-segment model has no data interaction but cannot be collaboratively optimized.
[0104] 3. Real-time performance: The local diagnosis latency of this invention is an average of 800ms, and the parameter transmission and aggregation time is an average of 5 minutes per round; the data transmission latency of the centralized diagnosis model is an average of 30 seconds (limited by bandwidth), and the diagnosis latency is an average of 2 seconds; the local diagnosis latency of the single-segment model is 900ms. This invention ensures collaboration while also taking into account the real-time diagnosis requirements.
[0105] 4. Anti-interference capability: By simulating two Byzantine nodes (uploading abnormal parameters), the diagnostic accuracy of the present invention decreased by only 1.2%, while the accuracy of the centralized diagnostic model decreased by 8.7%. The single-segment model has no anti-interference capability (node failure directly leads to diagnostic failure), which verifies the effectiveness of the Byzantine robust design of the present invention.
[0106] While specific embodiments of the present invention have been described above, those skilled in the art should understand that the specific embodiments described are merely illustrative and not intended to limit the scope of the present invention. Equivalent modifications and variations made by those skilled in the art in accordance with the spirit of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A collaborative diagnosis method for multi-segment railway track faults based on federated learning, characterized in that: Includes the following steps: Step S1: Deploy the client on different sections of the railway track, collect heterogeneous monitoring data of the railway track in the section, including at least track vibration signals, deformation data and environmental parameters, and store them locally with encryption. Step S2: The clients of each road segment preprocess and label the heterogeneous monitoring data to construct a local fault feature dataset; Step S3: The central server initializes a global fault diagnosis module and a global output module, and encrypts and sends the initial global diagnosis parameters of the global fault diagnosis module and the initial global output parameters of the global output module to the clients of each road segment. Step S4: Each road segment client initializes a local fault diagnosis module based on the initial global diagnostic parameters, initializes a personalized output module based on the initial global output parameters, and constructs a local fault diagnosis model based on the local fault diagnosis module and the personalized output module. Step S5: Each road segment client trains its local fault diagnosis model locally based on the fault feature dataset. After training is completed, the quality assessment index of the training data in this round is calculated, and the local update parameters of the local fault diagnosis module and the quality assessment index are encrypted and uploaded to the central server. Step S6: The central server adopts a weighted aggregation strategy with Byzantine robustness, calculates the aggregation weight based on the quality assessment indicators uploaded by each road segment client, aggregates the local update parameters based on the aggregation weights, updates the global fault diagnosis module to obtain optimized global diagnosis parameters, encrypts and sends the optimized global diagnosis parameters to each road segment client, and repeats the local training and parameter aggregation steps until the global fault diagnosis module reaches the preset convergence condition or training round, and encrypts and sends the latest optimized global diagnosis parameters to the road segment client to update the local fault diagnosis module. Step S7: Each section client uses the latest local fault diagnosis model to identify and classify faults in the local real-time heterogeneous monitoring data, outputs real-time fault diagnosis results, and realizes collaborative diagnosis of track faults across multiple sections.
2. The multi-segment railway track fault collaborative diagnosis method based on federated learning as described in claim 1, characterized in that: Step S1 specifically involves: The client-side components deployed on different sections of the railway track collect heterogeneous monitoring data of the track section via sensor arrays, including at least track vibration signals, deformation data, and environmental parameters. The track vibration signals include at least the vertical vibration acceleration of the rail, the lateral vibration acceleration of the rail, the sleeper vibration amplitude, and the vibration frequency. The deformation data includes at least the vertical displacement of the rail, the lateral displacement of the rail, the gauge deviation, and the rail curvature. The environmental parameters include at least the ambient temperature, ambient humidity, precipitation, wind speed, and dust concentration. The road segment client will collect the heterogeneous monitoring data and store it in a local database encrypted with the AES-256 encryption algorithm.
3. The multi-segment railway track fault collaborative diagnosis method based on federated learning as described in claim 2, characterized in that: The database has a role-based hierarchical authorization mechanism, which only assigns corresponding operation permissions to authorized operation and maintenance personnel, and combines multi-factor authentication to prevent unauthorized access. At the same time, it records all data access, data modification and data export operation logs and retains the operation logs for at least 90 days. When the road segment client updates each of the databases, it automatically calculates the data fingerprint of the database based on the HASH-256 algorithm and stores the data fingerprint on the blockchain. Before updating the database, it performs an integrity check based on the data fingerprint stored on the blockchain in the previous time, and periodically performs cloud-encrypted backup of the data stored in the database. Based on the storage space of the road segment client and the preset clearing rules, it continuously clears the invalid data in the database.
4. The multi-segment railway track fault collaborative diagnosis method based on federated learning as described in claim 1, characterized in that: Step S2 specifically involves: Each road segment client performs preprocessing on the heterogeneous monitoring data, including outlier handling, noise filtering, data alignment and synchronization, normalization, and feature engineering, to obtain standardized data; Based on railway industry standards, the standardized data are labeled with fault types and fault levels to obtain corresponding labeling information; By employing SMOTE oversampling technology or random undersampling technology, the proportion of fault samples to normal samples in each standardized dataset is balanced based on the annotation information, thereby constructing a local fault feature dataset. The outlier handling specifically involves: using statistical methods based on interquartile range or the 3σ criterion to identify and remove abnormal data points in the heterogeneous monitoring data caused by momentary sensor failure or external interference; and for continuous data in the heterogeneous monitoring data, using linear interpolation or the mean of the effective values before and after the data to fill in the gaps. The noise filtering specifically involves: using a Butterworth bandpass filter or wavelet threshold denoising method to filter out high-frequency noise and power frequency interference from the track vibration signal in the heterogeneous monitoring data, while retaining the effective frequency band signal related to the characteristics of rail faults; using a moving average filter to smooth the deformation data; and using a sliding window value filter method for the environmental parameters, while introducing physical constraint-based adaptive filtering and multi-sensor data verification. By using a preset instantaneous rate of change threshold, physically impossible abnormal jumps are identified and processed, and cross-validation is performed using the correlation between parameters. The data alignment and synchronization specifically involves: based on a unified first timestamp, performing time alignment on the heterogeneous monitoring data from sensors with different sampling frequencies, and unifying the data to the same time series through an interpolation resampling method to ensure the temporal correlation between the data; The normalization process specifically involves performing Z-score standardization or Min-Max normalization on the heterogeneous monitoring data to eliminate the influence of dimensions. The feature engineering specifically involves extracting time-domain features, frequency-domain features, and time-frequency-domain features from the heterogeneous monitoring data.
5. The multi-segment railway track fault collaborative diagnosis method based on federated learning as described in claim 1, characterized in that: Step S3 specifically involves: The central server initializes a global fault diagnosis module and a global output module; The global fault diagnosis module is constructed based on a heterogeneous data adaptation unit, a multi-scale feature extraction unit, a cross-road segment feature fusion unit, and a feature calibration unit; the global output module is constructed based on a feature enhancement unit, a multi-task prediction unit, and a probability calibration unit. Obtain the initial global diagnostic parameters of the global fault diagnosis module and the initial global output parameters of the global output module; Obtain the current second timestamp, calculate the hash value of the initial global diagnostic parameters, initial global output parameters, and second timestamp using the SM3 algorithm, encrypt the initial global diagnostic parameters, initial global output parameters, second timestamp, and hash value into an encrypted parameter packet using a preset SM4 raw key, retrieve the public key of each road segment client through the auxiliary server, encrypt the SM4 raw key with each public key to obtain the corresponding encryption key, and send the encrypted parameter packet and encryption key to the corresponding road segment client through an SSL encrypted transmission link.
6. The multi-segment railway track fault collaborative diagnosis method based on federated learning as described in claim 5, characterized in that: The heterogeneous data adaptation unit is constructed based on vibration signal branch, deformation data branch, environmental parameter branch, and fusion subunit. The vibration signal branch is used to capture the time-frequency local features of non-stationary track vibration signals by using wavelet basis functions through wavelet convolution layers. The deformation data branch is used to capture the spatial correlation features of rail deformation from the deformation data through a spatial convolutional layer using a 3×3 convolutional kernel; The environmental parameter branch is used to capture the temporal variation trend characteristics of environmental parameters through a temporal convolutional layer using a 1×k convolutional kernel; The fusion subunit is used to dynamically calculate the attention weights of the time-frequency local features, spatial correlation features, and temporal change trend features through a multi-head attention weighting layer, and then weight and fuse them into a unified embedded feature. The multi-scale feature extraction unit is constructed based on multi-scale dilated convolutional blocks, channel attention subunits, and spatial attention subunits. The multi-scale dilated convolutional block is used to capture local and global features from the unified embedding features through four parallel dilated convolutional layers; The channel attention subunit is used to perform channel-dimensional "squeeze-excitement" on the output of the multi-scale dilated convolutional block through the squeeze-excitement module, thereby enhancing fault-sensitive channels and suppressing redundant channels. The spatial attention subunit is used to introduce position information into the output of the channel attention subunit through the coordinate attention module, focus on the spatial location-related features of the fault occurrence, improve the spatial discriminativeness of the features, and output multi-scale fusion features; The cross-segment feature fusion unit is constructed based on a graph construction subunit, a graph convolutional layer, and a knowledge distillation subunit. The graph construction subunit is used to extract a normalized adjacency matrix from the multi-scale fusion features through a dynamic federated graph generator; The graph convolutional layer is used to aggregate the features of adjacent or similar road segments in the normalized adjacency matrix through weighted graph convolution, and output a common fault feature matrix. The knowledge distillation subunit is used to take the strong road segment features with abundant fault samples and high feature quality in the common fault feature matrix as teacher signals and distill them into the weak road segment features with scarce fault samples, so as to improve the feature expression ability of the weak road segment and output cross-road segment common features. The feature calibration unit is constructed based on the feature confidence evaluation subunit, the adaptive filtering subunit, and the feature smoothing subunit; The feature credibility assessment subunit is used to calculate the credibility score matrix of the cross-segment common features of each railway segment using a Mahalanobis distance calculator, based on each of the multi-scale fusion features and the global feature distribution statistics; the global feature distribution statistics are the mean vector μ and covariance matrix Σ of the multi-scale fusion feature statistics of all railway segments. The adaptive filtering subunit is used to perform weighted suppression on the cross-segment common features through a soft threshold filtering layer, based on the confidence score matrix, to obtain a filtered common feature matrix. The feature smoothing subunit is used to perform Gaussian smoothing on the filtered common feature matrix through a Gaussian kernel smoothing layer to reduce feature fluctuations, improve the stability of common features, and output calibration common fault features. The feature enhancement unit is constructed based on residual connection blocks and task attention subunits; The residual connection block is used to enhance the common fault features of calibration through the ResNet bottleneck structure to obtain residual enhanced common fault features; The task attention subunit is used to dynamically adjust the feature enhancement weights of the residual-enhanced common fault features based on the task differences between fault type prediction and fault level prediction through the dynamic attention layer, thereby improving the adaptability of features to tasks and obtaining task-enhanced common features. The multi-task prediction unit is constructed based on a shared feature layer, a fault type prediction branch, a fault level prediction branch, and a task interaction subunit. The shared feature layer is used to extract common dual-task shared basic features for fault type prediction task and fault level prediction task from the common features of task enhancement through two fully connected layers. The fault type prediction branch is used to identify the fault type probability distribution defined by railway industry standards from the shared basic features of the dual tasks through a 3-layer MLP and Softmax activation. The fault level prediction branch is used to identify the fault level probability distribution from the shared basic features of the dual tasks through a 3-layer MLP and Sigmoid activation. The task interaction subunit is used to calculate the interaction loss value through the interaction loss function, and to backpropagate and update the network parameters of the shared feature layer, the fault type prediction branch and the fault level prediction branch based on the interaction loss value. The probability calibration unit is constructed based on the distribution adaptation subunit and the probability correction subunit; The distribution adaptation subunit is used to perform temperature scaling on the fault type probability distribution and the fault level probability distribution. The probability correction subunit is used to perform nonlinear correction on the temperature-scaled fault type probability distribution and fault level probability distribution through the Beta calibration layer, so as to output fault diagnosis results carrying fault type and fault level.
7. The multi-segment railway track fault collaborative diagnosis method based on federated learning as described in claim 1, characterized in that: Step S4 specifically involves: Each client section receives the encrypted parameter packet and encryption key sent by the central server, calls the locally stored private key, decrypts the received encryption key to obtain the SM4 original key, and then decrypts the received encrypted parameter packet using the SM4 original key to obtain the initial global diagnostic parameters, initial global output parameters, second timestamp, and hash value. After performing integrity verification on the initial global diagnostic parameters, initial global output parameters, and second timestamp based on the hash value, a timeliness verification is performed based on the second timestamp. If the verification passes, the parameter reception is completed; if the verification fails, a retransmission request is sent to the central server. The central server, in conjunction with the access logs of the auxiliary server, investigates transmission anomalies and retransmits the data. Each road segment client initializes a local fault diagnosis module with the same network architecture as the global fault diagnosis module based on the initial global diagnostic parameters, and initializes a personalized output module with the same architecture as the global output module based on the initial global output parameters. A local fault diagnosis model is then constructed based on the local fault diagnosis module and the personalized output module.
8. The multi-segment railway track fault collaborative diagnosis method based on federated learning as described in claim 1, characterized in that: Step S5 specifically involves: Each road segment client uses a 30-day base time window and initially divides the fault feature dataset into a training set and a validation set in a 7:3 ratio to ensure that the training set covers samples of all fault types. Calculate the fault distribution entropy of the validation set. If the fault distribution entropy is less than a preset entropy threshold, dynamically adjust the ratio of the training set to the validation set within the basic time window. The training and validation sets are subjected to fault scenario-driven online data augmentation, specifically as follows: For the sample of track vibration signal, preset amplitude scaling, time stretching and Gaussian noise superposition are applied based on physical simulation rules to simulate the changes of fault signal under different train loads and speeds. For the deformation data samples, combined with the historical distribution of local environmental parameters, deformation-derived samples under different environmental couplings are generated through linear transformation. For samples of environmental parameters, virtual samples of the synergistic effects of multiple environmental factors are generated by permutation and combination to make up for the scarcity of local composite scene samples. The local fault diagnosis model is trained locally using the online data-augmented training set, and the trained local fault diagnosis model is validated using the online data-augmented validation set. After every 5 rounds of training, the base time window is updated, and the training set and validation set are re-divided until the preset number of training rounds is completed. After training is completed, the quality assessment index of the training data for this round is calculated. The local update parameters of the local fault diagnosis module and the quality assessment index are encrypted using a homomorphic encryption algorithm and then uploaded to the central server through an SSL encrypted transmission link.
9. The multi-segment railway track fault collaborative diagnosis method based on federated learning as described in claim 1, characterized in that: In step S6, the weighted aggregation strategy with Byzantine robustness specifically involves: the central server calculating the aggregation weight of each local update parameter based on the quality assessment indicators uploaded by the clients of each road segment, using a trimmed average algorithm to remove extreme abnormal parameters from the local update parameters, and then performing a weighted summation of the effective parameters in the local update parameters based on the aggregation weight.
10. A multi-segment railway track fault collaborative diagnosis system based on federated learning, characterized in that: Includes the following modules: The heterogeneous monitoring data acquisition module is used to deploy on the client side of different railway sections to collect heterogeneous monitoring data of the railway section, including at least track vibration signals, deformation data and environmental parameters, and store them locally with encryption. The fault feature dataset construction module is used by clients of each road segment to preprocess and label the heterogeneous monitoring data in order to construct a local fault feature dataset. The server model initialization module is used by the central server to initialize a global fault diagnosis module and a global output module, and to encrypt and send the initial global diagnosis parameters of the global fault diagnosis module and the initial global output parameters of the global output module to the clients of each road segment. The local fault diagnosis model initialization module is used by each road segment client to initialize a local fault diagnosis module based on the initial global diagnosis parameters, initialize a personalized output module based on the initial global output parameters, and construct a local fault diagnosis model based on the local fault diagnosis module and the personalized output module. The local training module is used by each road segment client to train the local fault diagnosis model locally based on the fault feature dataset. After training is completed, the quality assessment index of the training data in this round is calculated, and the local update parameters of the local fault diagnosis module and the quality assessment index are encrypted and uploaded to the central server. The parameter aggregation module is used by the central server to adopt a weighted aggregation strategy with Byzantine robustness, calculate the aggregation weight based on the quality assessment indicators uploaded by each road segment client, aggregate each locally updated parameter based on the aggregation weight, update the global fault diagnosis module to obtain optimized global diagnosis parameters, encrypt and send the optimized global diagnosis parameters to each road segment client, repeat the local training and parameter aggregation steps until the global fault diagnosis module reaches the preset convergence condition or training round, and encrypt and send the latest optimized global diagnosis parameters to the road segment client to update the local fault diagnosis module; The rail fault diagnosis module is used by clients of each section to identify and classify faults in local real-time heterogeneous monitoring data using the latest local fault diagnosis model, and output real-time fault diagnosis results to achieve collaborative diagnosis of rail faults across multiple sections.