Multi-line bogie fault diagnosis method, device, equipment and storage medium based on federated learning

By using the step by step federal distillation architecture of knowledge distillation in the multi-line bogie data imbalance scenario of high-speed trains, the problems of data imbalance and privacy security are solved, and efficient and accurate fault diagnosis performance is improved.

CN119397386BActive Publication Date: 2025-06-10SOUTHWEST JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411267930.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-11
Publication Date
2025-06-10
Estimated Expiration
2044-09-11

AI Technical Summary

Technical Problem

How to efficiently and accurately realize fault diagnosis under the federated learning framework when the multi-line bogie data of high-speed trains is uneven, especially when data privacy and security are protected.

Method used

A step by step federal distillation architecture based on knowledge distillation is proposed. By sorting and prioritizing multi-line data, using shallow freezing networks and personalized modules for model training and knowledge migration, dynamically adjusting the weight of the distillation loss function to improve the fault diagnosis performance of the small sample line model.

Benefits of technology

On the premise of ensuring data security, the fault diagnosis performance of the small sample line model is improved, the generalization ability and diagnostic accuracy of the model are improved, and the cost of model training is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119397386B_ABST
    Figure CN119397386B_ABST
Patent Text Reader

Abstract

The present application provides a multi-line bogie fault diagnosis method, device, equipment and storage medium based on federated learning, which relates to the technical fields of fault diagnosis and federated learning. Through a hierarchical federated distillation architecture, this method utilizes the model knowledge of other lines to improve the diagnostic performance of the small-sample line model while ensuring data security. For multiple fault diagnosis tasks in the architecture, the loss is decoupled into TCKD and NCKD and further reconstructed to balance the loss between the target class and the non-target class. For multiple distillation losses output by multiple tasks, an adaptive adjustment strategy for multiple loss weights is proposed to cooperate with multiple distillation tasks to achieve efficient training. Each line shares a basic shallow network based on model migration and simultaneously customizes a personalized deep network to improve the adaptability of the fault diagnosis model and reduce the learning cost of the proposed architecture. The present invention can improve the fault modeling accuracy of small-sample lines while ensuring the privacy and security of multi-party data, providing a feasible solution for implementing federated learning among users with data imbalance in industrial scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical fields of fault diagnosis and federated learning, and particularly relates to a multi-line bogie fault diagnosis method, device, equipment and storage medium based on federated learning. Background Technique

[0002] As an important tool for railway transportation, high-speed trains must first ensure their safe operation. There are many factors affecting the safe operation of trains, and one of the most important factors is the operation and maintenance of bogies. As one of the most important components in the structure of high-speed trains, bogies are mainly composed of wheelsets, bogie frames, primary suspensions, secondary suspensions, and basic braking equipment. It has the characteristics of complex mechanical structure, large number of components, and diverse operating conditions. While providing power for the train, it also bears the weight of the car body, providing an important guarantee for the safe and stable operation of the train. Due to the influence of different operating environments and conditions such as long uphill slopes, the proportion of bridges and tunnels, and climate conditions on high-speed trains, as the operating time of high-speed trains on the track increases, these influences are coupled through the direct excitation of each component of the bogie and the indirect excitation of the track, often resulting in a series of problems such as an accelerated wear rate of the wheelset, a deeper degradation of the performance of the primary suspension and secondary suspension systems, and an increased impact vibration of the car body and bogie systems. As a system with extremely high complexity and strong coupling, any component failure in the bogie system may cause a chain reaction, affecting the reliability of the whole vehicle operation. Therefore, accurate bogie fault diagnosis is crucial for ensuring the safe and stable operation of high-speed railways.

[0003] The performance of fault diagnosis algorithms based on deep learning depends on a large amount of high-quality labeled data. However, in the actual train operation and maintenance site, it is often difficult to perform fine manual annotation on the data, making it difficult to obtain high-quality labeled data, resulting in insufficient individual data, unbalanced sample types, and incomplete information dimensions, and it is difficult to establish a complete model alone.

[0004] At the same time, in the actual industrial scenario, due to different railway operating environments and different road conditions of the lines, for high-speed trains on different operating lines, their bogies usually contain characteristic data reflecting their common faults, and the fault data is small-scale and fragmented. The model obtained by training only through a single line is only applicable to this line and has poor generalization ability. Therefore, if the bogie data of different lines is collected for unified training, a model with very strong generalization ability will be obtained. Theoretically, this model can accurately identify all fault conditions of the lines participating in the training.

[0005] However, if the data of high-speed train bogies on multiple different lines are directly aggregated, the economic cost will be greatly increased. In addition, for some fields with high data sensitivity and easy to leak, such as finance and medical care, data is usually strictly protected and cannot leave the local storage device. At the same time, there may be data monopoly between enterprises, which makes data sharing and cooperation difficult to implement. In particular, data security issues are now receiving more and more attention, and national policies to protect data privacy have also been introduced. This makes data-driven methods have certain limitations in diagnosing faults in actual industrial environments. Therefore, how to effectively mine the intelligence contained in the data without infringing the privacy and security of the data itself is a difficult problem that needs to be solved nowadays.

[0006] A feasible idea is to jointly model multiple lines with different operating environments and working conditions based on federated learning. By migrating and reusing model knowledge between multiple lines and ensuring data privacy and security, mutual benefit and win-win results can be achieved. However, in the actual operation and maintenance of high-speed trains, due to the limitations of the actual equipment operating status and data collection implementation, when multiple lines are deployed according to the distributed learning architecture of federated learning, the high-speed train bogie data of the line often presents the characteristics of non-independent and identically distributed. The imbalance of data will lead to the imbalance of the line models of all parties, which is not conducive to the cooperation of multiple parties to establish a model with strong generalization ability, and will eventually seriously affect the training and detection of the diagnostic model.

[0007] In particular, in the current classic federated learning framework, each line has an equal status with the same weight, and multiple local models are averaged during model aggregation. However, lines with a small amount of data may also contain more valuable fault information, and direct aggregation of the model may mask the differences in their own data. The model drift of small sample lines causes each client to be unable to effectively learn information from other clients. Therefore, when multi-party data is unbalanced, giving multi-party lines the same priority in the architectural process may not have a positive effect on model optimization. Innovating the existing model aggregation process will improve the overall training efficiency of federated learning in some scenarios.

[0008] Secondly, since federated learning combines multiple data holders to build models, it often faces the problem of high computing overhead, resulting in low training efficiency and difficulty in implementation in actual scenarios.

[0009] From the above background, it can be clearly seen that when implementing federated modeling for multi-line high-speed trains, how to efficiently and accurately realize fault diagnosis under the federated learning framework when the multi-line bogie data is unbalanced has become an urgent problem to be solved. Summary of the invention

[0010] The present application provides a multi-line bogie fault diagnosis method, device, equipment and storage medium based on federated learning to improve the existing federated learning framework, propose a new framework for data imbalance scenarios, and thus improve the model accuracy and generalization ability.

[0011] In a first aspect, the present application provides a multi-line bogie fault diagnosis method based on federated learning, including:

[0012] Sort the multi-lines according to the data volume size to obtain a multi-line sequence Client=(Client1, Client2...Clienti...Clientn) sorted by priority, where Client1 represents the line with the highest priority, Clienti represents the i-th line, Clientn represents the line with the lowest priority, n represents the total number of multi-lines, and n≥3;

[0013] Line Client1 constructs a deep neural network locally and trains the deep neural network using the local data of Client1. After freezing the shallow network of the trained deep neural network, a shallow frozen network is obtained, and the shallow frozen network and the Client1 model parameters are transmitted to the central server;

[0014] Line Client2 trains the Client2 model using the local data of Client2, and the trained Client2 model parameters are transmitted to the central server; wherein, the Client2 model includes the shallow frozen network received from the central server and the first personalized module;

[0015] Line Clienti trains the Clienti model using the local data of Clienti, and the trained Clienti model parameters are transmitted to the central server; wherein, the Clienti model includes the shallow frozen network received from the central server and the (i-1)-th personalized module, and i is a natural number and 3≤i≤n.

[0016] In a possible design, when line Client2 trains the Client2 model using the local data of Client2, it includes:

[0017] The shallow frozen network received from the central server does not participate in the training;

[0018] Construct a distillation loss function for the target class and non-target class with the Client1 model parameters, and perform knowledge distillation based on the distillation loss function to complete the local model training of Client2.

[0019] In a possible design, the distillation loss function is determined by the following method:

[0020] Calculate the classification probability p based on the target class using the following formula t :

[0021]

[0022] where C is the number of classes, t represents the target class, and z j represents the logarithm value of the j-th class, and z t represents the logarithm of the target class;

[0023] Calculate the classification probability p based on the non-target class using the following formula \t :

[0024]

[0025] where z k represents the logarithm value of the non-target class;

[0026] After removing the target class, the classification probabilities of the other classes are:

[0027]

[0028] Use the KL divergence to measure the distillation loss between circuits T and S:

[0029]

[0030] where KD represents the distillation loss, KL represents the KL divergence calculation, p T represents the classification probability distribution of circuit T, and p S represents the classification probability distribution of circuit S, represents the classification probability of the target class t in circuit T, represents the classification probability of the target class t in circuit S, represents the classification probability of the non-target class i in circuit T, represents the classification probability of the non-target class i in circuit S;

[0031] Let The distillation loss function is obtained as:

[0032]

[0033] where b represents the binary classification scenario, TCKD is the target class knowledge distillation, NCKD is the non-target class knowledge distillation, is the classification probability of the target class t in circuit T, is the sum of the classification probabilities of the non-target classes in circuit T, is the classification probability of the remaining class i after removing the target class t in circuit T, is the classification probability of the remaining category i after removing the target class t in line S, is the sum of the classification probabilities of non-target classes in line T, is the sum of the classification probabilities of non-target classes in line S, is the classification probability distribution of line T after removing the target class t, is the classification probability distribution of line S after removing the target class t, b T and b S respectively represent the binary classification scenarios of line T and line S;

[0034] An additional weight value is assigned to TCKD through the following formula:

[0035] DKD = α·TCKD + β·NCKD

[0036] where DKD represents the decoupled knowledge distillation loss, α is the adjustment parameter for balancing TCKD and NCKD, and β is used to represent

[0037] In a possible design, line Clienti trains the Clienti model using the local data of Clienti, including:

[0038] The shallow frozen network received from the central server does not participate in the training;

[0039] Obtain the line model parameters with a higher priority than line Clienti from the central server;

[0040] One-to-one corresponding distillation loss functions are constructed between the Clienti model and the line model parameters of each line with a higher priority than line Clienti, and the weights of the distillation loss functions are dynamically adjusted to obtain the overall loss function of the Clienti model, and the Clienti model is trained using the overall loss function.

[0041] In a possible design, the overall loss function is expressed as:

[0042]

[0043] where, for the i-th line (i = 1, 2, L, N) in the federated task, represents the value of the local multi-class cross-entropy loss function, represents the two-norm value of the DKD loss of the t-th distillation task, represents the average two-norm value of the DKD losses of all tasks, r tIt represents the loss training completion degree of the t-th distillation task. λ is a hyperparameter that adjusts the training of the local model and the knowledge transfer with other parties. The term where 1 - λ is located is used to achieve knowledge distillation between multiple lines. μ is a hyperparameter that controls the co-training rate between multiple tasks, and the term where it is located is used to balance the overall loss. M is the total number of distillation tasks;

[0044] In each distillation task, is set as a fixed constant and does not participate in the backpropagation of gradients during training;

[0045] The local cross-entropy loss value of the i-th line is:

[0046]

[0047] where C represents the number of categories, y k represents the category corresponding to the true label, represents the probability predicted by the softmax function;

[0048] The loss training completion degree r t of the t-th distillation task is:

[0049] r t = L t / L avg

[0050] where L t is the average cross-entropy loss between pairs of lines, L avg is the average loss of all lines, and the smaller the r t value, the higher the training completion degree.

[0051] In a possible design, before sorting the multiple lines according to the data volume size, the method further includes:

[0052] Obtain the normal and fault data of the high-speed train bogie for different lines; wherein, the normal and fault data of the high-speed train bogie includes vibration signals under normal working conditions and vibration signals under fault conditions, and the fault conditions include AS air leakage, LD oil leakage, AD oil leakage, LD oil leakage + AD oil leakage, AS air leakage + LD oil leakage, and AS air leakage + AD oil leakage.

[0053] In a second aspect, the present application provides a multi-line bogie fault diagnosis device based on federated learning, including a central server and multiple lines, and the central server is signal-connected to all multiple lines:

[0054] The central server is used to sort multiple lines according to the data volume size, and obtain a multi-line sequence Client=(Client1, Client2...Clienti...Clientn) sorted by priority, where Client1 represents the line with the highest priority, Clienti represents the i-th line, Clientn represents the line with the lowest priority, and n represents the total number of multi-lines, n≥3;

[0055] Line Client1 is used to build a deep neural network locally, train the deep neural network using the local data of Client1, freeze the shallow network of the trained deep neural network to obtain a shallow frozen network, and transmit the shallow frozen network and the Client1 model parameters to the central server;

[0056] Line Client2 is used to train the Client2 model using the local data of Client2, and transmit the trained Client2 model parameters to the central server; wherein, the Client2 model includes a shallow frozen network received from the central server and a first personalized module;

[0057] Line Clienti is used to train the Clienti model using the local data of Clienti, and transmit the trained Clienti model parameters to the central server; wherein, the Clienti model includes a shallow frozen network received from the central server and an (i-1)-th personalized module, and i is a natural number and 3≤i≤n.

[0058] In a third aspect, the present application provides an electronic device, including: at least one processor and a memory; the memory stores computer execution instructions; the at least one processor executes the computer execution instructions stored in the memory, so that the at least one processor executes the federated learning-based multi-line bogie fault diagnosis method described in the first aspect and various possible designs of the first aspect above.

[0059] In a fourth aspect, the present application provides a computer-readable storage medium, in which computer execution instructions are stored, and when a processor executes the computer execution instructions, the federated learning-based multi-line bogie fault diagnosis method described in the first aspect and various possible designs of the first aspect above is implemented.

[0060] In a fifth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the federated learning-based multi-line bogie fault diagnosis method described in the first aspect and various possible designs of the first aspect above is implemented.

[0061] The multi-line bogie fault diagnosis method, device, equipment and storage medium based on federated learning provided by this application have the following beneficial technical effects:

[0062] 1. Aiming at the data imbalance problem existing in multi-line federated modeling, this application proposes a hierarchical federated distillation architecture based on knowledge distillation, which uses the models of other lines to improve the fault diagnosis performance of the small-sample line model on the premise of ensuring data security.

[0063] 2. In each distillation task of this application, the distillation loss is decomposed into two parts, TCKD and NCKD, and then recoupled to flexibly exert the representation ability of knowledge distillation for the target class and non-target class of bogie faults.

[0064] 3. This application proposes an adaptive adjustment strategy for the loss weights of multiple knowledge distillation tasks based on grad-norm to achieve efficient synchronous collaborative training among multi-line bogie fault diagnosis tasks.

[0065] 4. In the multi-party training of this architecture, the shallow model of the highest-priority party is frozen and shared with other parties, reducing the model training cost. At the same time, the fault diagnosis model of each line is configured with a personalized module, which improves the adaptability of the local model and balances the generalization and personalization of the model. Description of the Drawings

[0066] The drawings here are incorporated into the specification and constitute a part of this specification, showing the embodiments in line with this application, and are used together with the specification to explain the principles of this application.

[0067] Figure 1 It is a schematic diagram of the application scenario in the prior art;

[0068] Figure 2 It is a flowchart of the multi-line bogie fault diagnosis method based on federated learning provided by the embodiment of this application;

[0069] Figure 3 It is a diagram of common bogie fault types provided by the embodiment of this application, where: (a) represents air spring (AirSpring, AS) air leakage, (b) represents lateral damper (Lateral Damper, LD) oil leakage, and (c) represents anti-yaw damper (Anti-yaw Damper, AD) oil leakage;

[0070] Figure 4 It is a flowchart of the bogie fault diagnosis for 4 lines provided by the embodiment of this application;

[0071] Figure 5 It is a schematic diagram of the construction of the local personalized model for the bogie fault diagnosis of high-speed train on each line provided by the embodiment of this application;

[0072] Figure 6 This is the overall framework diagram of step-by-step federated distillation with four lines as an example provided by the embodiments of the present application;

[0073] Figure 7 This is the structural schematic diagram of a multi-line bogie fault diagnosis device based on federated learning provided by the embodiments of the present application;

[0074] Figure 8 This is the structural schematic diagram of an electronic device provided by the embodiments of the present application.

[0075] Through the above-mentioned drawings, specific embodiments of the present application have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Specific Embodiments

[0076] Here, exemplary embodiments will be described in detail, and their examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0077] In the technical solution of the present application, the collection, storage, use, processing, transmission, provision, and disclosure of information such as financial data or user data comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0078] It should be noted that in the embodiments of the present application, some industry-existing solutions such as certain software, components, models, etc. may be mentioned. They should be regarded as exemplary, and their purpose is only to illustrate the feasibility in the implementation of the technical solution of the present application, but it does not mean that the applicant has already or necessarily used this solution.

[0079] In the technical solution of the present application, the collection, storage, use, processing, transmission, provision, and disclosure of information such as financial data or user data comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0080] It should be noted that in the embodiments of the present application, some industry-existing solutions such as certain software, components, models, etc. may be mentioned. They should be regarded as exemplary, and their purpose is only to illustrate the feasibility in the implementation of the technical solution of the present application, but it does not mean that the applicant has already or necessarily used this solution.

[0081] Figure 1It is an application scenario diagram of a multi-line bogie fault diagnosis method based on federated learning. As Figure 1 shown, the application scenario includes a central server 101 and multiple lines 102, and each line is respectively represented as Client1, Client2, Clienti, Clientn; where n represents the total number of lines, and the total number of lines is determined according to the actual situation, and it is at least two. The central server 101 is used to sort the multi-lines according to the data volume size to obtain a multi-line sequence Client=(Client1, Client2...Clienti...Clientn) sorted by priority, where Client1 represents the line with the highest priority, Clienti represents the i-th line, Clientn represents the line with the lowest priority, n represents the total number of multi-lines, n≥3; the line Client1 constructs a deep neural network locally and uses the local data of Client1 to train the deep neural network. After freezing the shallow network of the trained deep neural network, a shallow frozen network is obtained, and the shallow frozen network and the Client1 model parameters are transmitted to the central server; the line Client2 uses the local data of Client2 to train the Client2 model, and the trained Client2 model parameters are transmitted to the central server; among them, the Client2 model includes the shallow frozen network received from the central server and the first personalized module; the line Clienti uses the local data of Clienti to train the Clienti model, and the trained Clienti model parameters are transmitted to the central server; among them, the Clienti model includes the shallow frozen network received from the central server and the (i-1)th personalized module, and i is a natural number and 3≤i≤n.

[0082] It should be noted that the "line" refers to the existing track operation line. Each track operation line includes a data processing unit. The process of knowledge distillation for each of the above-described lines is executed by the data processing unit included in the corresponding track operation line. The local data of each line is the normal and fault data of the train bogie collected by that line. Different levels of federated distillation structures are sequentially built for each line according to the amount of data. For example, the line Client1 with the largest amount of data is responsible for transmitting the model information trained by itself to all other lines through the central server and transferring model knowledge through knowledge distillation technology, which constitutes the first-level federated distillation structure. The line Client2 with the second-largest amount of data performs local model training, constituting the second-level federated distillation structure. The line Client3 with the third-largest amount of data performs local model training, constituting the third-level federated distillation structure. The line Client4 with the fourth-largest amount of data performs local model training, constituting the fourth-level federated distillation structure, and so on, to implement the n-level federated distillation structure, and improve the fault diagnosis performance of the small-sample line model by means of the models of other lines while ensuring data security.

[0083] Figure 2 This is the flowchart of the multi-line bogie fault diagnosis method based on federated learning provided by the embodiments of the present application. As Figure 2 shown, the embodiments of the present application provide a multi-line bogie fault diagnosis method based on federated learning. The method includes the following steps:

[0084] S210. Sort the multi-lines according to the amount of data to obtain a multi-line sequence Client = (Client1, Client2... Clienti... Clientn) sorted by priority, where Client1 represents the line with the highest priority, Clienti represents the i-th line, Clientn represents the line with the lowest priority, and n represents the total number of multi-lines, and n≥3.

[0085] In this embodiment, n is taken as 4. Taking 4 lines as an example, the multi-line sequence is successively represented as Client1, Client2, Client3, and Client4.

[0086] Before step S201, in this embodiment, the normal and fault data of the high-speed train bogies on different lines are collected through the SIMPACK platform, and then the multi-lines are sorted according to the amount of normal and fault data of the high-speed train bogies on different lines, so as to determine the priority of each line, which is beneficial to the design of the subsequent multi-level federated distillation structure.

[0087] Due to the limitations of actual industrial conditions, the cost of obtaining data through real high-speed train bogie fault experiments is very high. Therefore, the SIMPACK platform can effectively simulate and collect bogie fault data. The State Key Laboratory of Traction Power of Southwest Jiaotong University provided a nonlinear dynamic model of the CRH380A high-speed train system for the experiment. The entire train model has 62 degrees of freedom. In the vertical, lateral and longitudinal directions, sensors installed at different positions of the vehicle bogie obtained vibration signals, including acceleration and displacement of different components. The experiment recorded a total of 58-dimensional vibration signals.

[0088] The vibration data of the embodiment of the present application is generated by the track spectrum excitation of the Wuhan-Guangzhou Line, Zhengzhou-Xi'an Line, Beijing-Tianjin Line and Jiaoji Line in SIMPACK. The train bogie consists of two air springs (ASs), two lateral dampers (LDs) and four anti-snaking dampers (ADs), which are common faults of the bogie. The fault state is as follows Figure 3 As shown in the figure, they are AS air leakage, LD oil leakage and AD oil leakage. These phenomena indicate that the bogie components have been degraded to a certain extent, which will aggravate the vibration of the bogie and even affect the comfort and safety of the train. The experiment studied 7 common working conditions on each line, corresponding to labels 0, 1, 2, ..., 6, as shown in Table 1.

[0089] Table 1 Common working conditions of bogies

[0090] Label Operating condition 0 Normal 1 LD 2 AD 3 AS 4 LD + AD 5 AS + LD 6 AS + AD

[0091] The vibration data sampling frequency is set to 243 Hz, and the train running speed is set to 200 km / h. For each working condition, a small amount of time-series vibration signals are collected, and the corresponding data volume corresponds to the priority of the line in the framework decreasing step by step. The specific details are shown in Table 2.

[0092] Table 2 Design of bogie data for different lines

[0093]

[0094] S220, line Client1 builds a deep neural network locally, and uses Client1 local data to train the deep neural network, freezes the shallow network of the trained deep neural network to obtain a shallow frozen network, and transmits the shallow frozen network and Client1 model parameters to the central server.

[0095] Figure 4 A bogie fault diagnosis flowchart for four lines is provided in an embodiment of the present application.

[0096] like Figure 4 As shown, step S220 specifically includes:

[0097] S221. Client1 constructs a deep neural network locally and trains the model using only its own local data.

[0098] S222. After freezing the shallow network of the model trained by Client1, encrypt the model parameters and upload them to the central server.

[0099] S223. The central server distributes the model parameters of Client1 to the remaining low-priority lines (Client2, Client3, Client4).

[0100] S230. Line Client2 trains the Client2 model using Client2's local data, and the trained Client2 model parameters are transmitted to the central server; where the Client2 model includes the shallow frozen network received from the central server and the first personalized module.

[0101] As Figure 4 shown, step S230 specifically includes:

[0102] S231. The line Client2 with the second-largest data volume, after receiving the shallow frozen network of Client1, connects it to its own personalized module to form the Client2 local model.

[0103] S232. Construct a distillation loss function for the target class (target class knowledge distillation, TCKD) and non-target class (non-target class knowledge distillation, NCKD) with the Client1 model, and perform knowledge distillation based on this to complete the training of the Client2 local model.

[0104] S233. Encrypt the Client2 model parameters and upload them to the central server.

[0105] S234. The central server distributes the Client2 model parameters to the remaining low-priority lines (Client3, Client4).

[0106] S240. Line Clienti trains the Clienti model using Clienti's local data, and the trained Clienti model parameters are transmitted to the central server; where the Clienti model includes the shallow frozen network received from the central server and the (i - 1)th personalized module, and i is a natural number and 3 ≤ i ≤ n.

[0107] Taking four lines as an example, for line Client3, step S240 specifically includes:

[0108] S241. After the line Client3 with the third largest amount of data receives the shallow frozen network of Client1, it is connected to its own personalized module to form the Client3 local model.

[0109] S242. Based on the models of Client1 and Client2, Client3 constructs a one-to-one corresponding distillation loss function with each of them, and implements model training through the multi-task loss weight adaptive adjustment strategy based on grad-norm, and transfers the knowledge to the Client3 model.

[0110] S243. Encrypt the Client3 model parameters and upload them to the central server.

[0111] S244. The central server distributes the Client3 model parameters to the remaining low-priority lines (Client4).

[0112] For the line Client3, step S240 specifically includes:

[0113] S245: After the line Client4 with the least amount of data receives the shallow frozen network of Client1, it is connected to its own personalized module to form the Client4 local model.

[0114] S246. Based on the models of Client1, Client2 and Client3, Client4 constructs a one-to-one corresponding distillation loss function with each of them, and implements model training through the multi-task loss weight adaptive adjustment strategy based on grad-norm, and transfers the knowledge to the Client4 model, finally realizing the step-by-step federated distillation learning of multi-line faults in the data imbalance scenario.

[0115] In this embodiment, in each round of learning, after the line serving as the teacher completes local training, it will share some of its frozen models with other lines. The frozen models do not participate in the local training of other lines, but only serve as benchmark models to provide support for local training, avoiding the low training effect caused by starting training from scratch.

[0116] At the same time, each line's own personalized model will be connected after the shared frozen model. According to the characteristics of its own data, based on the designed multi-loss weight calculation method, adaptive training is carried out through the distillation strategy, so that the model converges on the local dataset and ensures the performance of the overall model. The basic idea is as Figure 5 shown.

[0117] In some embodiments, the line Client2 uses the local data of Client2 to train the Client2 model, including:

[0118] The shallow freezing network received from the central server does not participate in the training; a distillation loss function is constructed for the target class and non-target class with the model parameters of Client1, and knowledge distillation is performed based on the distillation loss function to complete the local model training of Client2.

[0119] Each model training is different from the conventional distillation method. The framework distills the knowledge between the two lines each time, and two types of losses based on the target class and non-target class are respectively used to jointly represent the distillation loss.

[0120] Among them, the classification probability p based on the target class t is:

[0121]

[0122] where C is the number of classes, t represents the target class, and z j represents the logarithm value of the j-th class, and z t represents the logarithm of the target class.

[0123] The classification probability p based on the non-target class \t is:

[0124]

[0125] where z k represents the logarithm value of the non-target class.

[0126] After removing the target class, the classification probability of other classes is:

[0127]

[0128] Therefore, when the Kullback-Leibler (KL) divergence is used to measure the distillation loss between lines T and S, it is as follows:

[0129]

[0130] where KD represents the distillation loss (Knowledge Distillation), KL represents the KL (Kullback-Leibler) divergence calculation, p T represents the classification probability distribution of line T, and p S represents the classification probability distribution of line S, represents the classification probability of the target class t in line T, represents the classification probability of the target class t in line S, represents the classification probability of the non-target class i in line T, represents the classification probability of the non-target class i in line S;

[0131] Since The above formula can be rewritten to obtain the distilled loss function expressed as:

[0132]

[0133] where b represents the binary classification scenario, TCKD is target class knowledge distillation, and NCKD is non-target class knowledge distillation. is the classification probability of target class t in circuit T. is the sum of the classification probabilities of non-target classes in circuit T. is the classification probability of the remaining class i after removing target class t in circuit T. is the classification probability of the remaining class i after removing target class t in circuit S. is the sum of the classification probabilities of non-target classes in circuit T. is the sum of the classification probabilities of non-target classes in circuit S. is the classification probability distribution after removing target class t in circuit T. is the classification probability distribution after removing target class t in circuit S, b T and b S represent the binary classification scenarios of circuit T and circuit S respectively.

[0134] TCKD can provide key knowledge about the distilled target, while NCKD pays more attention to the potential connections between non-target classes. By combining the two, the knowledge transfer between multiple circuits can be effectively characterized.

[0135] In particular, to better improve the adaptability between the two, an additional weight value is assigned to TCKD as follows:

[0136] DKD = α·TCKD + β·NCKD

[0137] where DKD represents the decoupled knowledge distillation loss, α is the adjustment parameter for balancing TCKD and NCKD, and β is used to represent

[0138] In some embodiments, the line Clienti trains the Clienti model using the local data of Clienti, including: the shallow frozen network received from the central server does not participate in the training, and the line model parameters with a higher priority than the line Clienti are obtained from the central server; a one-to-one corresponding distillation loss function is constructed between the Clienti model and the line model parameters of each line with a higher priority than the line Clienti, and the weight of the distillation loss function is dynamically adjusted to obtain the overall loss function of the Clienti model, and the Clienti model is trained using the overall loss function.

[0139] It should be noted that the method of constructing a one-to-one corresponding distillation loss function between the lienti model and the line model parameters of each line with a higher priority than the line Clienti can use two types of losses based on the target class and the non-target class to jointly represent the distillation loss. The specific calculation process has been described in the previous embodiments and will not be repeated here.

[0140] In the above designed architecture, multiple distillations are required between lines to optimize and adjust the model, and the loss function obtained from each knowledge distillation will participate in the training of the target line, which will result in multiple distillation losses. By dynamically adjusting the weights of these loss functions and treating the effectiveness of distillation between different lines differently, the accuracy of the model is further improved as a whole. The overall framework is as Figure 6 .

[0141] The overall loss function is designed as follows:

[0142]

[0143] Among them, for the i-th line (i = 1, 2, L, N) in the federated task, represents the value of the local multi-class cross-entropy loss function, represents the two-norm value of the DKD loss of the t-th distillation task, represents the average two-norm value of the DKD losses of all its tasks, r t represents the completion degree of the loss training of the t-th distillation task, λ is a hyperparameter that adjusts the training of the local model and the knowledge transfer from other parties, and the term where 1 - λ is located can achieve knowledge distillation between multiple lines. μ is a hyperparameter that controls the co-training rate between multiple tasks, and the term where it is located can balance the overall loss and play a role in achieving an equal training rate.

[0144] In each distillation task, will be set as a fixed constant and does not participate in the backpropagation of the gradient in the training to prevent the loss function from falsely drifting towards zero, and at the same time, the training target can be focused on the loss function of each local party.

[0145] The local cross-entropy loss value of the i-th line is as follows:

[0146]

[0147] Among them, C represents the number of classifications, y k represents the category corresponding to the true label, represents the probability predicted by the softmax function;

[0148] The loss training completion degree r of the t-th distillation task t is as follows:

[0149] r t = L t / L avg

[0150] Among them, L t is the average cross-entropy loss between pairs of lines, and L avg is the average loss of all lines. The smaller the r t value, the higher the training completion degree.

[0151] In the experiment, the diagnostic accuracy of each line is shown in Table 3. It can be concluded that the diagnostic performance of the local model of each line can be further improved by leveraging the knowledge of other lines, and the leakage of local data is also avoided.

[0152] Table 3 Diagnostic accuracy of train bogie faults for each line

[0153]

[0154] The multi-line bogie fault diagnosis based on federated learning proposed in the embodiments of this application has a complete diagnostic process. Based on the hierarchical federated distillation architecture, it solves the problem of secure joint modeling of multi-line high-speed train bogies in the data imbalance scenario. This hierarchical federated distillation architecture is redesigned on the basis of the basic process of federated learning, and adopts a hierarchical distillation strategy to improve the performance of the small-sample line model. For the multiple fault diagnosis distillation tasks in the proposed architecture, the loss value is decoupled into the target class loss and the non-target class loss and further reconstructed. Secondly, for the multiple distillation losses output by multiple tasks, an adaptive adjustment strategy for multiple loss weights is proposed to achieve efficient training in coordination with multiple distillation tasks.

[0155] To reduce the training cost of the fault diagnosis model of the proposed framework, each line trains the local model based on the model transfer shared basic shallow network and uses the frozen module. Other lines can draw on the knowledge of the pre-trained shallow model, while customizing the personalized deep network to improve the adaptability of the fault diagnosis model and conduct personalized exploration of the local line fault data characteristics. This method not only maintains good diagnostic performance but also significantly reduces the computational cost of model training.

[0156] The proposed framework shows satisfactory fault diagnosis performance on high - speed train bogies. The main reason is that multiple levels of step - by - step distillation are implemented for each line in the framework, which can extract the effective feature knowledge of other lines to the greatest extent. The overall diagnostic accuracy of train bogies for each line is higher than 92%, making it more suitable for fault diagnosis tasks in data - imbalance scenarios and meeting the diagnostic requirements.

[0157] Figure 7 This is a schematic structural diagram of a multi - line bogie fault diagnosis device based on federated learning provided by an embodiment of the present application. An embodiment of the present application also provides a multi - line bogie fault diagnosis device based on federated learning. As Figure 7 shown, the multi - line bogie fault diagnosis device based on federated learning includes a central server 701 and multiple lines 702. The central server 701 is signal - connected to each of the multiple lines 702:

[0158] The central server 701 is used to sort the multiple lines according to the data volume size, obtaining a multi - line sequence Client=(Client1, Client2...Clienti...Clientn) sorted by priority, where Client1 represents the line with the highest priority, Clienti represents the i - th line, Clientn represents the line with the lowest priority, n represents the total number of multiple lines, and n≥3;

[0159] Line Client1 is used to build a deep neural network locally and train the deep neural network using the local data of Client1. After freezing the shallow network of the trained deep neural network, a shallow frozen network is obtained, and the shallow frozen network and the Client1 model parameters are transmitted to the central server;

[0160] Line Client2 is used to train the Client2 model using the local data of Client2, and the trained Client2 model parameters are transmitted to the central server. Among them, the Client2 model includes the shallow frozen network received from the central server and the first personalized module;

[0161] Line Clienti is used to train the Clienti model using the local data of Clienti, and the trained Clienti model parameters are transmitted to the central server. Among them, the Clienti model includes the shallow frozen network received from the central server and the (i - 1) - th personalized module, where i is a natural number and 3≤i≤n.

[0162] In some embodiments, the multiple lines are sorted according to the data volume size to obtain a sequence of multiple lines sorted by priority Client=(Client1, Client2...Clienti...Clientn), where Client1 represents the line with the highest priority, Clienti represents the i-th line, Clientn represents the line with the lowest priority, n represents the total number of multiple lines, and n≥3;

[0163] Line Client1 constructs a deep neural network locally and trains the deep neural network using the local data of Client1. After freezing the shallow network of the trained deep neural network, a shallow frozen network is obtained, and the shallow frozen network and the Client1 model parameters are transmitted to the central server;

[0164] Line Client2 trains the Client2 model using the local data of Client2, and the trained Client2 model parameters are transmitted to the central server; wherein, the Client2 model includes the shallow frozen network received from the central server and the first personalized module;

[0165] Line Clienti trains the Clienti model using the local data of Clienti, and the trained Clienti model parameters are transmitted to the central server; wherein, the Clienti model includes the shallow frozen network received from the central server and the (i-1)-th personalized module, and i is a natural number and 3≤i≤n.

[0166] In some embodiments, line Client2 trains the Client2 model using the local data of Client2, including:

[0167] The shallow frozen network received from the central server does not participate in the training;

[0168] A distillation loss function is constructed for the target class and non-target class with respect to the Client1 model parameters, and knowledge distillation is performed based on the distillation loss function to complete the local model training of Client2.

[0169] In some embodiments, line Client2 determines the distillation loss function by the following method:

[0170] Calculate the classification probability p for the target class through the following formula t :

[0171]

[0172] where C is the number of classes, t represents the target class, z j represents the logarithm value of the j-th class, z tRepresents the logarithm of the target class;

[0173] Calculate the classification probability p based on non - target classes through the following formula \t :

[0174]

[0175] where z k represents the logarithm value of non - target classes;

[0176] After removing the target class, the classification probabilities of other classes are:

[0177]

[0178] Use the KL divergence to measure the distillation loss between circuits T and S:

[0179]

[0180] where KD represents the distillation loss, KL represents the KL divergence calculation, p T represents the classification probability distribution of circuit T, p S represents the classification probability distribution of circuit S, represents the classification probability of the target class t in circuit T, represents the classification probability of the target class t in circuit S, represents the classification probability of non - target class i in circuit T, represents the classification probability of non - target class i in circuit S;

[0181] Let Get the distillation loss function expressed as:

[0182]

[0183] where b represents the binary classification scenario, TCKD is the target class knowledge distillation, NCKD is the non - target class knowledge distillation, is the classification probability of the target class t in circuit T, is the sum of the classification probabilities of non - target classes in circuit T, is the classification probability of the remaining class i except the target class t in circuit T, is the classification probability of the remaining class i except the target class t in circuit S, is the sum of the classification probabilities of non - target classes in circuit T, is the sum of the classification probabilities of non - target classes in circuit S, is the classification probability distribution of circuit T except the target class t, is the classification probability distribution of circuit S except the target class t, b T and b Srespectively represent the binary classification scenarios of line T and line S;

[0184] Give an additional weight value to TCKD through the following formula:

[0185] DKD = α·TCKD + β·NCKD

[0186] where DKD represents the coupled distillation loss, α is the adjustment parameter for balancing TCKD and NCKD, and β is used to represent

[0187] In some embodiments, line Clienti uses the local data of Clienti to train the Clienti model, including:

[0188] The shallow frozen network received from the central server does not participate in the training;

[0189] Obtain the line model parameters with a higher priority than line Clienti from the central server;

[0190] Construct a one-to-one corresponding distillation loss function between the Clienti model and the line model parameters of each priority higher than line Clienti, and dynamically adjust the weight of the distillation loss function to obtain the overall loss function of the Clienti model, and use the overall loss function to train the Clienti model.

[0191] In some embodiments, the overall loss function is expressed as:

[0192]

[0193] where, for the i-th line (i = 1, 2, L, N) in the federated task, represents the local multi-class cross-entropy loss function value, represents the second norm value of the DKD loss of the t-th distillation task, represents the average second norm value of the DKD losses of all tasks, r t represents the loss training completion degree of the t-th distillation task, λ is a hyperparameter for adjusting the training of the local model and the knowledge transfer from other parties, the term where 1 - λ is located is used to achieve knowledge distillation between multiple lines, μ is a hyperparameter for controlling the co-training rate between multiple tasks, and the term where it is located is used to balance the overall loss, and M is the total number of distillation tasks;

[0194] In each distillation task, is set as a fixed constant and does not participate in the backpropagation of the gradient during training;

[0195] The local cross-entropy loss value of the i-th line is:

[0196]

[0197] Among them, C represents the number of classifications, and y k represents the category corresponding to the true label, and represents the probability predicted by the softmax function;

[0198] The loss training completion degree r of the t-th distillation task t is:

[0199] r t = L t / L avg

[0200] Among them, L t is the average cross-entropy loss between pairs of lines, and L avg is the average loss of all lines. The smaller the r t value, the higher the training completion degree.

[0201] In some embodiments, the device further includes a data acquisition module, and the data acquisition module is used to:

[0202] acquire normal and fault data of the bogies of high-speed trains on different lines; among them, the normal and fault data of the bogies of high-speed trains include vibration signals under normal working conditions and vibration signals under fault conditions, and the fault conditions include AS air leakage, LD oil leakage, AD oil leakage, LD oil leakage + AD oil leakage, AS air leakage + LD oil leakage, and AS air leakage + AD oil leakage.

[0203] The multi-line bogie fault diagnosis device based on federated learning provided by the embodiments of the present application can be used to execute the technical solutions of the multi-line bogie fault diagnosis method based on federated learning in the above embodiments. The implementation principles and technical effects are similar and will not be elaborated here.

[0204] Figure 8 This is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 8 shown, the electronic device may include: a processor 81 and a memory 82. Among them, the processor 81 and the memory 82 can communicate; exemplarily, the processor 81 and the memory 82 communicate through a communication bus 83.

[0205] The processor 81 executes the computer-executable instructions stored in the memory 82, enabling the processor 81 to execute the solutions in the above embodiments. The processor 81 can be a general-purpose processor, including a Central Processing Unit (CPU), a network processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0206] The communication bus 83 can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The system bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only a thick line is shown in the figure, but it does not mean there is only one bus or one type of bus. The transceiver is used to implement communication between the database access device and other computers (such as clients, read-write libraries, and read-only libraries). The memory may include Random Access Memory (RAM) and may also include non-volatile memory.

[0207] The electronic device provided in the embodiments of the present application can be the terminal device in the above embodiments.

[0208] The embodiments of the present application also provide a computer-readable storage medium, in which computer instructions are stored. When the computer instructions run on a computer, the computer is enabled to execute the technical solutions of the above embodiments of the multi-line bogie fault diagnosis method based on federated learning.

[0209] The embodiments of the present application also provide a computer program product, which includes a computer program stored in a computer-readable storage medium. At least one processor can read the computer program from the computer-readable storage medium, and when at least one processor executes the computer program, the technical solutions of the above embodiments of the multi-line bogie fault diagnosis method based on federated learning can be implemented.

[0210] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of devices or modules can be in electrical, mechanical or other forms.

[0211] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to implement the solution of this embodiment.

[0212] In addition, each functional module in various embodiments of the present application can be integrated in a processing unit, or each module can exist physically alone, or two or more modules can be integrated in one unit. The unit formed by the above modules can be implemented in the form of hardware, or in the form of a hardware plus a software functional unit.

[0213] The above integrated modules implemented in the form of software functional modules can be stored in a computer-readable storage medium. The above software functional modules stored in a storage medium include several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods in various embodiments of the present application.

[0214] It should be understood that the above processor can be a central processing unit (Central Processing Unit, abbreviated as CPU), and can also be other general-purpose processors, digital signal processors (Digital Signal Processor, abbreviated as DSP), application specific integrated circuits (Application Specific Integrated Circuit, abbreviated as ASIC), etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.

[0215] The memory may include high-speed RAM memory, and may also include non-volatile storage NVM, such as at least one disk memory, and can also be a USB flash drive, a mobile hard disk, a read-only memory, a disk or an optical disc, etc.

[0216] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, the buses in the drawings of this application are not limited to only one bus or one type of bus.

[0217] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk or an optical disc. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0218] An exemplary storage medium is coupled to the processor, enabling the processor to read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an Application Specific Integrated Circuits (ASIC). Of course, the processor and the storage medium can also exist as discrete components in an electronic control unit or a master control device.

[0219] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes various media that can store program codes, such as ROM, RAM, magnetic disks, or optical discs.

[0220] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. A multi-line bogie fault diagnosis method based on federated learning, characterized in that: include: Sort multiple lines according to the amount of data, and obtain a multi-line sequence Client=(Client1, Client2...Clienti...Clientn) sorted by priority, where Client1 represents the line with the highest priority, Clienti represents the i-th line, Clientn represents the line with the lowest priority, and n represents the total number of multiple lines, n≥3; Line Client1 constructs a deep neural network locally, and trains the deep neural network using Client1 local data, freezes the shallow network of the trained deep neural network to obtain a shallow frozen network, and transmits the shallow frozen network and Client1 model parameters to the central server; Line Client2 trains the Client2 model using the Client2 local data, and transmits the trained Client2 model parameters to the central server; wherein the Client2 model includes the shallow frozen network received from the central server and the first personalized module; Line Clienti trains the Clienti model using Clienti local data, and transmits the trained Clienti model parameters to the central server; wherein the Clienti model includes a shallow frozen network received from the central server and an i-1th personalized module, where i is a natural number and 3≤i≤n.

2. The method according to claim 1, characterized in that Line Client2 uses Client2 local data to train the Client2 model, including: The shallow frozen network received from the central server does not participate in training; A distillation loss function is constructed for the target class and the non-target class using the model parameters of Client1, and knowledge distillation is performed based on the distillation loss function to complete the local model training of Client2.

3. The method according to claim 2, characterized in that The distillation loss function is determined by the following method: The classification probability p based on the target class is calculated by the following formula t : Among them, C is the number of categories, t represents the target category, z j represents the logarithmic value of the jth class, z t represents the logarithm of the target class; The classification probability p based on the non-target class is calculated by the following formula \t : Among them, z k Represents the logarithmic value of the non-target class; After removing the target class, the classification probability of other classes for: The KL divergence is used to measure the distillation loss between lines T and S: Among them, KD represents distillation loss, KL represents KL divergence calculation, and p T represents the classification probability distribution of route T, p S represents the classification probability distribution of route S, represents the classification probability of target class t in line T, represents the classification probability of target class t in line S, represents the classification probability of non-target class i in line T, represents the classification probability of non-target class i in route S; make The distillation loss function is expressed as: Among them, b represents the binary classification scenario, TCKD is the target class knowledge distillation, and NCKD is the non-target class knowledge distillation. is the classification probability of target class t in line T, is the sum of the classification probabilities of non-target classes in line T, is the classification probability of the remaining categories i after removing the target class t in line T, is the classification probability of the remaining categories i in line S after removing the target class t, is the sum of the classification probabilities of non-target classes in line T, is the sum of the classification probabilities of non-target classes in line S, is the classification probability distribution of line T excluding target class t, is the classification probability distribution of line S without target class t, b T and b S They represent the binary classification scenarios of route T and route S respectively; An additional weight value is given to TCKD through the following formula: DKD=α·TCKD+β·NCKD Among them, DKD represents the coupled distillation loss, α is the adjustment parameter for balancing TCKD and NCKD, and β is used to represent 4. The method according to claim 1, characterized in that: Line Clienti uses Clienti local data to train the Clienti model, including: The shallow frozen network received from the central server does not participate in training; Obtain line model parameters with a higher priority than line Clienti from the central server; A one-to-one corresponding distillation loss function is constructed between the Clienti model and each line model parameter having a higher priority than the line Clienti, and the weight of the distillation loss function is dynamically adjusted to obtain an overall loss function of the Clienti model, and the Clienti model is trained using the overall loss function.

5. The method according to claim 4, characterized in that The overall loss function is expressed as: Among them, for the i-th line in the federated task (i=1,2,L,N), represents the local multi-classification cross entropy loss function value, represents the bi-norm value of the DKD loss of the t-th distillation task, represents the average bi-norm value of DKD loss for all tasks, r t represents the loss training completion of the t-th distillation task, λ is a hyperparameter that adjusts the local model training and the knowledge transfer from other parties, the item where 1-λ is located is used to realize the knowledge distillation between multiple lines, μ is a hyperparameter that controls the joint training rate between multiple tasks, and the item where it is located is used to balance the overall loss, and M is the total number of distillation tasks; In each distillation task, Set to a fixed constant and do not participate in the back propagation of gradients during training; The local cross entropy loss value of the i-th line is: Among them, C represents the number of categories, y k Indicates the true label corresponding category, Represents the probability predicted by the softmax function; The loss training completion of the t-th distillation task r t for: r t =L t / L avg Among them, L t is the average cross entropy loss between lines, L avg is the average loss of all lines, r t The smaller the value, the higher the training completion.

6. The method according to claim 1, characterized in that Before sorting the multiple lines according to the data volume, the method further includes: Obtain normal and fault data of high-speed train bogies on different lines; wherein, the normal and fault data of high-speed train bogies include vibration signals under normal working conditions and vibration signals under fault conditions, and the fault conditions include AS air leakage, LD oil leakage, AD oil leakage, LD oil leakage + AD oil leakage, AS air leakage + LD oil leakage and AS air leakage + AD oil leakage.

7. A multi-line bogie fault diagnosis device based on federated learning, characterized in that: It includes a central server and multiple lines, and the central server and the multiple lines are all signal connected: The central server is used to sort the multiple lines according to the amount of data, and obtain a multi-line sequence Client=(Client1, Client2...Clienti...Clientn) sorted by priority, wherein Client1 represents the line with the highest priority, Clienti represents the i-th line, Clientn represents the line with the lowest priority, and n represents the total number of multiple lines, n≥3; Line Client1 is used to construct a deep neural network locally, and train the deep neural network using Client1 local data, freeze the shallow network of the trained deep neural network to obtain a shallow frozen network, and transmit the shallow frozen network and Client1 model parameters to the central server; Line Client2 is used to train the Client2 model using the local data of Client2, and the parameters of the Client2 model obtained by training are transmitted to the central server; wherein the Client2 model includes the shallow frozen network received from the central server and the first personalized module; Line Clienti is used to train the Clienti model using Clienti local data, and the trained Clienti model parameters are transmitted to the central server; wherein the Clienti model includes a shallow frozen network received from the central server and an i-1th personalized module, where i is a natural number and 3≤i≤n.

8. An electronic device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 6 when executed by a processor.

10. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 6 when being executed by a processor.