Fault diagnosis method and device, storage medium and computer equipment
The global model parameters are obtained through pre-training of source nodes in the federated learning network, and the fault diagnosis model of the target node is updated, which solves the problems of data privacy leakage and central server credibility in mechanical equipment fault diagnosis, and improves the accuracy and efficiency of fault diagnosis.
Patent Information
- Application Number
- CN202411936102.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art has the risk of data privacy leakage in mechanical equipment fault diagnosis, the lack of credibility of central servers and the prone to system crashes. At the same time, high communication and computing burdens limit the efficiency of fault diagnosis.
Pre-training of source nodes in the federated learning network is obtained, and global model parameters are obtained, which are used to update the fault diagnosis model of the target node, reduce dependence on the central server, protect data privacy, and reduce communication and computing burden.
It improves the accuracy of fault diagnosis and classification of mechanical equipment under new operating conditions, reduces the burden of communication and computing, and avoids the credibility problems of central servers and the risk of data privacy leakage.
Smart Images

Figure CN120066827A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of mechanical equipment fault diagnosis, and particularly relates to a fault diagnosis method, device, storage medium, and computer device. Background Art
[0002] Mechanical fault diagnosis is crucial for ensuring the safe and reliable operation of industrial production processes. In related technologies, most deep learning-based fault diagnosis methods require a sufficient amount of labeled training data, and obtaining a sufficient amount of labeled data is usually challenging and expensive. Moreover, collecting data from distributed machines to a central server is prone to privacy issues, and due to the different distributions of training data and real-world test data, domain shift occurs. Although transfer learning-based fault diagnosis methods can alleviate domain shift, they also need to collect data to a central server for training, posing a risk of privacy leakage, thus preventing users from contributing their data. Federated learning can protect data privacy, but in related technologies, the fault diagnosis method based on federated transfer learning requires a central server, and the central server may lack credibility. In addition, a failure of the central server may lead to system crashes, resulting in significant losses. And there are limitations of high communication and computational burdens. Summary of the Invention
[0003] In view of this, this application provides a fault diagnosis method, device, storage medium, and computer device. Without a central node server and while protecting privacy, the fault diagnosis accuracy of the first fault diagnosis model of decentralized target nodes is improved through the global model parameters pre-trained by source nodes, and the communication and computational burdens are reduced.
[0004] According to one aspect of this application, a fault diagnosis method is provided, including: obtaining global model parameters of a federated learning network, where the federated learning network includes multiple source nodes, and the global model parameters are iteratively updated by adjacent source nodes in the federated learning network; updating model parameters of a first fault diagnosis model of a target node according to the global model parameters; and inputting operation status data of a device to be tested associated with the target node into the first fault diagnosis model to determine a fault type of the device to be tested.
[0005] According to another aspect of this application, a fault diagnosis device is provided, including: an obtaining module, configured to obtain global model parameters of a federated learning network, where the federated learning network includes multiple source nodes, and the global model parameters are iteratively updated by adjacent source nodes in the federated learning network; an updating module, configured to update model parameters of a first fault diagnosis model of a target node according to the global model parameters; and a determining module, configured to input operation status data of a device to be tested associated with the target node into the first fault diagnosis model to determine a fault type of the device to be tested.
[0006] According to another aspect of the present application, there is provided a readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the steps of the above-mentioned fault diagnosis method are implemented.
[0007] According to yet another aspect of the present application, there is provided a computer device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor. When the processor executes the program, the steps of the above-mentioned fault diagnosis method are implemented.
[0008] By means of the above technical solutions, the present application provides a fault diagnosis method, device, storage medium, and computer device. Using the labeled data at the source nodes in the federated learning network, pre-training is performed among the source nodes to obtain the optimal model parameters (global model parameters) of the second fault diagnosis model of the source nodes. Further, the global model parameters and the data features at the source nodes associated with the target node are transmitted to the target node, so that the target node updates the model parameters of the first fault diagnosis model of the target node according to the global model parameters, and aligns the features of the unlabeled data at the target node with the features of the labeled data at the source nodes associated with the target node, so as to reduce the incorrect fault diagnosis caused by domain shift. Thus, in the case of protecting privacy and lacking a central node server, the problem that the model obtained by training the source nodes cannot be applied to the target node due to the difference in feature distributions between the data of the source nodes and the data of the target node is effectively solved, and the fault diagnosis classification accuracy of mechanical equipment under new working conditions is improved.
[0009] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of the present application more obvious and understandable, the specific embodiments of the present application are specifically given below. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0011] Figure 1 shows a schematic flow chart of the fault diagnosis method provided by an embodiment of the present application;
[0012] Figure 2 shows a structural block diagram of the fault diagnosis device provided by an embodiment of the present application;
[0013] Figure 3 shows a structural block diagram of the second fault diagnosis model of the source node provided by an embodiment of the present application;
[0014] Figure 4The structural block diagram of the first fault diagnosis model of the target node provided by the embodiment of the present application is shown. Detailed implementation manners
[0015] The present application will be described in detail below with reference to the accompanying drawings and in combination with embodiments. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other.
[0016] Embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals indicate the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be construed as a limitation of the present application.
[0017] Those skilled in the art of the present technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "including" used in the specification of the present application means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "joined" to another element, it can be directly connected or joined to other elements, or there may also be intermediate elements. In addition, the "connection" or "joining" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.
[0018] Now, exemplary embodiments according to the present application will be described in more detail with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many different forms and should not be construed as being limited only to the embodiments set forth herein. It should be understood that these embodiments are provided so that the disclosure of the present application is thorough and complete, and the concept of these exemplary embodiments is fully conveyed to those of ordinary skill in the art.
[0019] Fault diagnosis is crucial for ensuring the safe and reliable production of industrial processes. In the related art, deep learning is used to diagnose the faults of mechanical equipment. Most of these methods require a sufficient amount of labeled training data, and obtaining sufficient labeled data is usually challenging and expensive. The collection of training data is usually from distributed machines to a powerful central server with computing capabilities, and this collection method raises privacy issues. In addition, due to highly complex operating conditions, the training data and the real-world test data usually have different distributions, that is, domain shift. Therefore, a model well-trained on the training data may not be suitable for the test data.
[0020] In related technologies, fault diagnosis methods based on transfer learning solve the problem of domain shift by transferring and adapting the knowledge learned from labeled source training data to target test data. Although they are effective in mitigating domain transfer, the vast majority require collecting data stored on local computers (users) to a computationally powerful central server for training. However, this process may lead to privacy leakage problems because sensitive data may be exposed during transmission, preventing users from contributing their data.
[0021] In related technologies, fault diagnosis methods based on federated learning protect data privacy by exchanging model parameters without accessing the data of individual users. To protect privacy and mitigate domain shift. In related federated transfer learning fault diagnosis methods, distributed machines exchange model parameters and features with a central server instead of raw data. However, there are some limitations when applying these methods. One limitation is that these methods all require a central server, but the central server may lack credibility. In addition, a failure of the central server may cause the system to crash, resulting in significant losses. Another limitation is the high communication and computational burden.
[0022] In this embodiment, a fault diagnosis method is provided, as Figure 1 shown, the method includes:
[0023] Step 101, obtaining the global model parameters of the federated learning network.
[0024] Among them, the federated learning network includes multiple source nodes, and the global model parameters are obtained by iteratively updating adjacent source nodes in the federated learning network.
[0025] In this embodiment, source nodes are determined according to mechanical equipment (source equipment) with known fault types, so that each source node is associated with at least one source equipment, thereby establishing a federated learning network based on the source nodes and performing decentralized federated pre-training among the source nodes to obtain the global model parameters of the federated learning network, avoiding the problems of lack of credibility and system crashes caused by failures that the central server may bring, and protecting data privacy at the same time.
[0026] In one embodiment, the fault diagnosis method further includes: constructing a second fault diagnosis model for the source node; training the second fault diagnosis model of the source node according to the operation state data of the source device associated with the source node to obtain the first loss value of the second fault diagnosis model, where the operation state data of the source device carries pre-labeled fault type tags; sending the model parameters of the second fault diagnosis model of the source node to the neighbor nodes of the source node, so that the neighbor nodes update the model parameters of the second fault diagnosis model of the neighbor nodes according to the model parameters sent by the source node, and the neighbor nodes send the model parameters of the second fault diagnosis model of the neighbor nodes to the source node, and the neighbor nodes are the source nodes in the federated learning network that have a communication edge with the source node; updating the model parameters of the second fault diagnosis model of the source node according to the model parameters sent by the neighbor nodes and the first loss value; if the second fault diagnosis model of the source node meets the first preset stop condition, taking the model parameters of the second fault diagnosis model that meets the first preset stop condition as the global model parameters.
[0027] In this embodiment, there is no need for a central server to coordinate the communication between source nodes. Instead, each source node communicates on its own. Each source node exchanges model parameters with its adjacent source nodes through the federated learning network, so as to use the labeled data at the source node to train the second fault diagnosis model of the source node between source nodes and obtain the optimal model parameters of the second fault diagnosis model, greatly reducing the communication overhead and the risk of single point of failure, avoiding risks such as lack of credibility and system collapse caused by faults that may be brought by the central server, protecting data privacy, and being more scalable and robust at the same time.
[0028] Specifically, the mechanical vibration data of the source device is collected by sensors at preset positions, and the mechanical vibration data of the source device is used as the operation state data of the source device, so as to obtain the operation state data of the source devices associated with each source node. And preprocessing operations such as segmentation and normalization are performed on the operation state data at the source node. Then, the preprocessed operation state data is labeled according to the fault types of the source device to determine the specific fault types corresponding to each operation state data, and the fault type tags of the operation state data are obtained. Then, the labeled operation state data is divided into the training data set and the verification data set of the source node according to a preset ratio (for example, 8:2 or 7:3), so as to train the second fault diagnosis model of the source node according to the data sets (training data set and verification data set) at the source node.
[0029] Exemplarily, the data sets D of all source nodes s are represented as: D i is the data set of source node i and is represented as: where N is the number of source nodes, x i,l is the l-th sample of source node i, yi,l For sample x i,l The corresponding fault type label, X i Is the distribution space of the samples of source node i, Y i Is the distribution space of the fault type labels of source node i, n i Is the number of samples of source node i.
[0030] Specifically, for example, in a factory, similar mechanical equipment in different workshops can be used as source nodes, and the mechanical vibration data they collect respectively can be used to jointly train a general fault diagnosis model.
[0031] Furthermore, each source node establishes the same second fault diagnosis model, which includes a second feature extractor, a second bottleneck layer, and a second classification layer. Among them, the second feature extractor is used to extract the features of the samples input into the second fault diagnosis model, the second bottleneck layer is used to compress the features extracted by the second feature extractor and reduce the model complexity, as well as reduce the computational amount, and the second classification layer is used to learn the classification rules according to the fault type labels of the samples input into the second fault diagnosis model, and perform fault type classification according to the features output by the second bottleneck layer to determine the fault type of the samples input into the second fault diagnosis model.
[0032] Exemplarily, as Figure 3 Shown, the second fault diagnosis model of source node i includes a second feature extractor F i (ξ i ), a second bottleneck layer B i (φ i ), and a second classification layer The model parameters θ i of the second fault diagnosis model of source node i are expressed as: Among them, ξ i , φ i and are the model parameters corresponding to the second feature extractor F i , the second bottleneck layer B i and the second classification layer C i respectively.
[0033] Furthermore, each source node obtains the first loss value of its second fault diagnosis model through its dataset and its second fault diagnosis model, and then transmits the model parameters of its second fault diagnosis model to its neighbor nodes through the federated learning network.
[0034] Exemplarily, as Figure 3 Shown, the model parameters θ i of the second fault diagnosis model of source node i are initialized, and then the sample x i,lInput the second fault diagnosis model after initialization to obtain the predicted classification result C of the second fault diagnosis model for the sample x i,l of the sample x i (B i (F i (x i,l ))), and calculate the first loss value L of the second fault diagnosis model according to the predicted classification result C i (B i (F i (x i,l ))) and the fault type label y of the sample x i,l of the sample x i,l The first loss value L of the second fault diagnosis model is calculated i ,
[0035]
[0036] where is the expected value calculation symbol. C i (B i (F i (x i,l ))) is used to represent the data processing process of inputting the sample x of the source node i into the second fault diagnosis model of the source node i, that is, the sample x i,l is first input into the second feature extractor F i,l , and then the result F i output by the second feature extractor F i is input into the second bottleneck layer B i (x i,l ) is input into the second bottleneck layer B i , and finally the result B i output by the second bottleneck layer B i (F i (x i,l ) is input into the second classification layer C i , and the predicted classification result C i,l of the sample x i (B i (F i (x i,l ))) is obtained
[0037] Furthermore, the federated learning network can be used as a communication network, enabling each source node to exchange information through the federated learning network. The communication graph can be represented by a random graph, and the random graph can be generated through a preset function, thus abstracting the communication between source nodes as an undirected connected graph G=(V, E, A). Among them, V is the set of source nodes, is the edge set describing the communication between source nodes. The adjacency matrix of the undirected connected graph G p∈V is used to represent whether there is an edge between the source node i and the source node p. If {i, p}∈E, that is, there is a communication edge between the source node i and the source node p, then aip > 0, that is, the source node i and the source node p can communicate with each other, then the source node i and the source node p are neighbor nodes of each other, otherwise a ip = 0, that is, there is no communication edge between the source node i and the source node p and they cannot communicate with each other. Degree matrix is a diagonal matrix, and the diagonal element is used to represent the degree of the source node i (the number of all communication edges connected to the source node i). In this embodiment, the set of neighbor nodes of the source node i is represented as The source node j is a neighbor node of the source node i. Therefore, the model parameters when the first loss value L of the second fault diagnosis model of the source node i is obtained are sent to the neighbor node j of the source node i through the federated learning network. i
[0038] It is worth mentioning that when each source node sends its model parameters to its neighbor nodes, each source node also acts as a neighbor node and receives the model parameters sent by other source nodes.
[0039] Furthermore, as Figure 3 shown, according to the first loss value of the second fault diagnosis model of the source node and the model parameters sent by its neighbor nodes, the model parameters of the second fault diagnosis model of the source node are updated, so that all source nodes can be combined to make the second fault diagnosis model more robust without transmitting the local data of the second fault diagnosis model.
[0040] Exemplarily, the model parameters of the second fault diagnosis model of the source node are updated according to the following formula:
[0041]
[0042] where, θ i (k + 1) is the model parameter when the second fault diagnosis model of the source node i is updated for the (k + 1)-th time, θ j (k) is the model parameter sent by the neighbor node j of the source node i when the second fault diagnosis model of the source node i is updated for the k-th time, ∈ i (k) is the learning rate when the second fault diagnosis model of the source node i is updated for the k-th time, which can be directly obtained, and μ = 0.1 is a hyperparameter used to balance the two terms in the above formula. W ij is the weight matrix between the source node i and its neighbor node j.
[0043] Further, after the second fault diagnosis model of each source node updates its model parameters, a new second fault diagnosis model is obtained. Subsequently, each source node starts a new round of model training to iteratively update the model parameters of the second fault diagnosis model of each source node until the second fault diagnosis model of each source node converges or reaches a preset maximum number of iterative updates (the first preset stop condition).
[0044] Exemplarily, based on the characteristic that the undirected communication graph is connected and the weight matrix W ij is doubly stochastic, the formula is obtained. This formula indicates that when the number of iterative updates of the second fault diagnosis model of the source node is large enough, the optimal model parameters of the second fault diagnosis model of the source node can be obtained are the optimal model parameters corresponding to the second feature extractor, the second bottleneck layer, and the second classification layer in the second fault diagnosis model of the source node, respectively.
[0045] It can be understood that after sufficient decentralized pre-training, the model parameters of all source nodes are equal to the same optimal model parameter, enabling each source node to accurately classify the fault types of its dataset using this optimal model parameter.
[0046] It should be noted that Figure 3 In Figure 4 , forward propagation and backpropagation are two key steps in neural networks, used to calculate and update the output and parameters of the model. Among them, forward propagation refers to starting from the input data, calculating and transmitting data layer by layer through each layer of the neural network, and finally obtaining the output result of the model. During forward propagation, the input data undergoes a linear transformation through the weights and biases of each layer, and a non-linear transformation through the activation function, and then is output to the next layer until the output layer is reached. The purpose of forward propagation is to calculate the predicted value of the model. Backpropagation refers to calculating the gradient in reverse according to the difference between the predicted result of the model and the true label through the chain rule, and propagating the gradient from the output layer back to each layer of the network for updating the model parameters. During backpropagation, first calculate the error of the output layer, then propagate the error from the output layer to the hidden layer, and then to the shallower hidden layer until it reaches the input layer. Through backpropagation, the gradient information of each parameter with respect to the loss function can be obtained, thereby realizing the optimization and update of the parameters.
[0047] Finally, the model parameters of the second fault diagnosis model that meet the first preset stop condition are used as the global model parameters of the federated learning network.
[0048] Further, as a refinement and extension of the specific implementation manner of the above embodiment, in order to completely illustrate the specific implementation process of this embodiment, the model parameters of the second fault diagnosis model of the source node are updated according to the model parameters and the first loss value sent by the neighbor node, including: determining the weight matrix between the source node and the neighbor node according to the adjacency matrix between the source node and the neighbor node, and the degree matrix of the federated learning network; updating the model parameters of the source fault diagnosis model according to the weight matrix, the first loss value, and the model parameters sent by the neighbor node.
[0049] In the embodiment, all weight matrices between the source node i and its neighbor nodes are represented as W, where I N is the identity matrix, and ||D|| ∞ is the infinity norm of the degree matrix D. It can be understood that the adjacency matrix A takes different values according to different neighbor nodes of the source node.
[0050] Step 102, updating the model parameters of the first fault diagnosis model of the target node according to the global model parameters.
[0051] It should be noted that the target node is determined according to the mechanical equipment (equipment to be measured) with unknown fault types, so that the target node is associated with at least one equipment to be measured, and thus the operation status data of the equipment to be measured without labeled fault types is used as the dataset (data at the target node) of the first fault diagnosis model of the target node associated with it.
[0052] In this embodiment, federated transfer learning is performed, and the model parameters of the first fault diagnosis model of the target node are updated using the global model parameters, so as to transfer and adjust the knowledge learned from the labeled data at the source node to the first fault diagnosis model of the target node.
[0053] Further, as a refinement and extension of the specific implementation manner of the above embodiment, in order to completely illustrate the specific implementation process of this embodiment, step 102, that is, updating the model parameters of the first fault diagnosis model of the target node according to the global model parameters, includes: inputting the associated data of the associated node into the associated fault diagnosis model of the associated node to determine the features of the associated node, where the associated node is the source node associated with the target node, the associated data is the operation status data of the source equipment associated with the associated node, and the associated fault diagnosis model is the second fault diagnosis model of the associated node configured with the global model parameters; initializing the model parameters of the first fault diagnosis model of the target node according to the global model parameters; updating the model parameters of the initialized first fault diagnosis model according to the features of the associated node and the operation status data of the equipment to be measured associated with the target node until the first fault diagnosis model of the target node meets the second preset stop condition.
[0054] In this embodiment, the features of the unlabeled data at the target node are aligned with the features of the labeled data at the source node associated with the target node, so as to reduce the distribution difference between the data features at the target node and the data features at its associated source node, making the feature distribution of the data at the target node closer to the feature distribution of the data at its associated source node, so that the global model parameters can be better applied to the first fault diagnosis model of the target node, reducing the fault misclassification caused by domain shift, and thus being able to more accurately determine the true fault condition of the target node and accurately perform fault classification.
[0055] Specifically, for example, a newly installed mechanical device or a remote device monitoring point can be used as the target node. The fault type of the target node is unknown, that is, the data at the target node does not have a fault type label. Then, the following situation may occur: Although the fault types of the data at the target node and the data at its associated source node are the same, due to the difference in the feature distribution of the data at the target node and the data at its source node, for example, caused by factors such as load or temperature, the second fault diagnosis model that directly uses the global model parameters at the target node cannot accurately classify the data at the target node.
[0056] In this embodiment, the source node corresponding to the existing device similar to the target node is used as the associated node of the target node. Thus, the first fault diagnosis model of the target node is trained by using the pre-trained global model parameters of the associated node and the features of the labeled data at the associated node, so as to update the model parameters of the first fault diagnosis model of the target node and obtain the optimal model parameters of the first fault diagnosis model of the target node. When the data at the target node and the associated node have the same fault, the features of the data at the target node and the associated node are closer, so that the pre-trained global model parameters at the associated node can be better applied to the target node, effectively solving the problem that the model obtained by training at the source node cannot be applied to the target node due to the difference in feature distribution between the dataset at the source node and the dataset at the target node caused by different operating conditions, and improving the fault diagnosis classification accuracy of mechanical equipment under new working conditions.
[0057] Specifically, the mechanical vibration data of the device under test is collected by sensors at preset positions, and the mechanical vibration data of the device under test is used as the operating state data of the device under test, so as to obtain the operating state data of the device under test associated with the target node. The operating state data of the device under test associated with the target node is subjected to preprocessing operations such as segmentation and initialization to obtain the dataset at the target node.
[0058] Exemplarily, the dataset D at the target node t t is expressed as: where n tis the number of samples of the target node t, represents the l-th t sample of the target node t, represents the distribution space of the samples of the target node t.
[0059] Furthermore, the second fault diagnosis model pre-trained at the associated node (i.e., the second fault diagnosis model of the associated node configured with global model parameters) is used to extract the features of the associated node and establish a feature set of the associated node.
[0060] Exemplarily, the feature of the l-th sample of the associated node q of the target node t is represented as: That is is the data after the l-th sample of the associated node q passes through the second feature extractor in the second fault diagnosis model configured with global model parameters at the associated node.
[0061] Next, a feature set of the associated nodes of the target node t is established
[0062]
[0063] where S t is the set of associated nodes of the target node t, is the number of all features of the associated nodes of the target node t, n q is the number of features of the associated node q, represents the feature sample pair, is the l'-th feature among all the features of the associated nodes of the target node t, is the fault type label of the source node corresponding to the l'-th feature, is the distribution space of all the features of the associated nodes of the target node t, is the distribution space of the fault type labels of the associated nodes of the target node t.
[0064] It should be noted that in order to achieve knowledge transfer from the source node to the target node, the set S t must contain at least one source node.
[0065] Then, a first fault diagnosis model is established at the target node. The first fault diagnosis model includes a first feature extractor F t (ξ t ), a first bottleneck layer B t (φ t ), a first classification layer and a discriminator D t (ζ t ), ξ t 、φ t 、 ζ t are respectively the model parameters corresponding to the first feature extractor F t , the first bottleneck layer B t , the first classification layer C t and the discriminator D t . And initialize the model parameters of the first fault diagnosis model of the target node t according to the global model parameters , that is, overwrite the model parameters corresponding to the first feature extractor with the model parameters corresponding to the second feature extractor in the global model parameters, overwrite the model parameters corresponding to the first bottleneck layer with the model parameters corresponding to the second bottleneck layer, and overwrite the model parameters corresponding to the first classification layer with the model parameters corresponding to the second classification layer. Specifically, it is expressed as:
[0066]
[0067]
[0068] where ξ t is randomly initialized.
[0069] It is worth mentioning that the first feature extractor is used to extract the features of the input data, the first bottleneck layer is used to compress the features extracted by the first feature extractor and reduce the model complexity, as well as reduce the computational amount. Adversarial training is carried out between the discriminator and the first feature extractor, which is used to promote the features of the data at the target node to be close to the feature set of the associated node, so that the first fault diagnosis model trained at the target node can be more applicable to the data at the target node. The first classification layer is used to determine the fault type of the input data according to the features extracted by the first feature extractor and the first bottleneck layer, and the input data is the data input into the first fault diagnosis model.
[0070] Furthermore, as Figure 4 shown, update the model parameters of the initialized first fault diagnosis model according to the features of the associated node and the data set at the target node until the first fault diagnosis model of the target node converges or reaches the maximum preset iteration update times (the second preset stop condition), and finally train to obtain the optimal model parameters of the first fault diagnosis model of the target node, so that the first fault diagnosis model of the target node has a good fault classification effect on the data set at the target node.
[0071] Further, as a refinement and extension of the specific implementation manner of the above embodiment, in order to completely illustrate the specific implementation process of this embodiment, the steps of updating the model parameters of the initialized first fault diagnosis model according to the characteristics of the associated nodes and the operation status data of the device under test associated with the target node specifically include: inputting the characteristics of the associated nodes into the first bottleneck layer of the initialized first fault diagnosis model to determine the first characteristics of the associated nodes, and respectively inputting the first characteristics into the discriminator and the first classification layer of the initialized first fault diagnosis model to determine the second characteristics and the fault classification result of the associated nodes; inputting the target data into the first feature extractor of the initialized first fault diagnosis model to determine the third characteristics of the target data, and inputting the third characteristics into the first bottleneck layer of the initialized first fault diagnosis model to determine the fourth characteristics of the target data, where the target data is the operation status data of the device under test associated with the target node; inputting the fourth characteristics into the discriminator of the initialized first fault diagnosis model to determine the fourth characteristics of the target data; calculating the second loss value, the third loss value, and the fourth loss value of the first fault diagnosis model according to the first characteristics, the second characteristics, the third characteristics, the fourth characteristics, the fifth characteristics, and the fault classification result; and updating the model parameters corresponding to the first feature extractor, the first bottleneck layer, the first classification layer, and the discriminator in the first fault diagnosis model according to the second loss value, the third loss value, and the fourth loss value.
[0072] In this embodiment, by combining the output results of the first feature extractor, the first bottleneck layer, the first classification layer, and the discriminator for the characteristics of the associated nodes and the data set at the target node, three loss values of the first fault diagnosis model are calculated to evaluate the performance of the model from different perspectives. And according to the three loss values, the model parameters corresponding to the first feature extractor, the first bottleneck layer, the first classification layer, and the discriminator are updated respectively, so as to promote the cooperation and common optimization of these components, improve the overall performance and generalization ability of the first fault diagnosis model, and enhance the robustness of the first fault diagnosis model.
[0073] Exemplarily, as Figure 4 shown, the second loss value L MMD of the first fault diagnosis model, the third loss value L D and the fourth loss value L c are respectively expressed as:
[0074]
[0075] where k(·,·) is a Gaussian kernel, and the calculation rules for elements a and b in the Gaussian kernel are and both represent the first characteristics of the associated nodes output after inputting the characteristics of the associated nodes into the first bottleneck layer of the initialized first fault diagnosis model; It represents the second feature of the associated node output by inputting the first feature of the associated node into the discriminator in the first fault diagnosis model after initialization. It represents the fault classification result of the associated node output by inputting the first feature of the associated node into the first classification layer in the first fault diagnosis model after initialization; F t (x t,r )、F t (x t,r′ )、 All represent the third feature of the target data output by inputting the target data into the first feature extractor in the first fault diagnosis model after initialization; B t (F t (x t,r ))、B t (F t (x t,r′ ))、 All represent the fourth feature of the target data output by inputting the third feature into the first bottleneck layer in the first fault diagnosis model after initialization.
[0076] It represents the fifth feature of the target data output by inputting the fourth feature into the discriminator in the first fault diagnosis model after initialization; the target data is the operation state data (dataset at the target node) of the device under test associated with the target node.
[0077] In this embodiment, the second loss value L MMD represents the maximum mean difference (Maximum Mean Discrepancy, MMD) between the data features at the associated node and the features of the data at the target node after being extracted by the first feature extractor. The smaller the maximum mean difference, the closer the data features at the associated node are to the data features at the target node. The fourth loss value L c is calculated according to the fault classification result of the first fault diagnosis model for the data at the associated node and the fault type corresponding to the data at the associated node. The smaller the fourth loss value, the more accurate the fault classification of the first fault diagnosis model.
[0078] It can be understood that during the iterative update of the model parameters of the first fault diagnosis model, the first feature, the second feature, the third feature, the fourth feature, the fifth feature, and the fault classification result are obtained according to the model parameters corresponding to the first feature extractor, the first bottleneck layer, the first classification layer, and the discriminator in the current first fault diagnosis model.
[0079] Furthermore, according to the first feature, the second feature, the third feature, the fourth feature, the fifth feature, and the fault classification result, calculate the second loss value, the third loss value, and the fourth loss value of the first fault diagnosis model respectively. Further, according to the second loss value, the third loss value, and the fourth loss value, update the model parameters corresponding to the first feature extractor, the first bottleneck layer, the first classification layer, and the discriminator in the first fault diagnosis model respectively.
[0080] Furthermore, as a refinement and extension of the specific implementation manner of the above embodiment, in order to completely illustrate the specific implementation process of this embodiment, the steps of updating the model parameters corresponding to the first feature extractor, the first bottleneck layer, the first classification layer, and the discriminator in the first fault diagnosis model according to the second loss value, the third loss value, and the fourth loss value specifically include: updating the model parameters corresponding to the first feature extractor according to the second loss value and the third loss value; updating the model parameters corresponding to the first bottleneck layer according to the second loss value, the third loss value, and the fourth loss value; updating the model parameters corresponding to the first classification layer and the discriminator according to the fourth loss value respectively.
[0081] In this embodiment, the model parameters of different components are updated according to different loss values, making the update process more targeted. Each component has a specific role in fault diagnosis. In this way, the first fault diagnosis model can be adjusted more finely, improving the performance of the first fault diagnosis model.
[0082] Exemplarily, as Figure 4 shown, the update rules for the model parameters corresponding to the first feature extractor, the first bottleneck layer, the first classification layer, and the discriminator are as follows:
[0083]
[0084] where ξ t (k′ + 1) is the model parameter corresponding to the first feature extractor F t at the (k′ + 1)-th iteration update; ξ t (k′) is the model parameter corresponding to the first feature extractor F t at the k′-th iteration update; ∈ ξ (k′) is the learning rate corresponding to the first feature extractor F t at the k′-th iteration update, which can be directly obtained; φ t (k′ + 1) is the model parameter corresponding to the first bottleneck layer B t at the (k′ + 1)-th iteration update; φ t (k′) is the model parameter corresponding to the first bottleneck layer B t at the k′-th iteration update; φ t (k′) is the learning rate corresponding to the first bottleneck layer B t at the k′-th iteration update, which can be directly obtained; is the first classification layer C t The model parameters corresponding to the (k'+1)-th iteration update; is the first classification layer C t The model parameters corresponding to the k'-th iteration update; is the first classification layer C t The learning rate corresponding to the k'-th iteration update, which can be directly obtained;
[0085] ζ t (k'+1) is the discriminator D t The model parameters corresponding to the (k'+1)-th iteration update; ζ t (k') is the discriminator D t The learning rate corresponding to the k'-th iteration update, which can be directly obtained; both β and α are hyperparameters used to balance several terms in the formula; is the symbol for partial differential calculation.
[0086] Step 103: Input the operation status data of the device under test associated with the target node into the first fault diagnosis model to determine the fault type of the device under test.
[0087] It should be noted that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0088] Furthermore, as Figure 2 shown, as a specific implementation of the above fault diagnosis method, an embodiment of the present application provides a fault diagnosis device 200, which includes: an acquisition module 201, an update module 202, and a determination module 203.
[0089] Among them, the acquisition module 201 is used to acquire the global model parameters of the federated learning network. The federated learning network includes multiple source nodes, and the global model parameters are obtained by iterative update of adjacent source nodes in the federated learning network;
[0090] The update module 202 is used to update the model parameters of the first fault diagnosis model of the target node according to the global model parameters;
[0091] The determination module 203 is used to input the operation status data of the device under test associated with the target node into the first fault diagnosis model to determine the fault type of the device under test.
[0092] In an embodiment, the fault diagnosis device 200 further includes:
[0093] A pre-training module for constructing a second fault diagnosis model of a source node; training the second fault diagnosis model of the source node according to the operation status data of the source device associated with the source node to obtain a first loss value of the second fault diagnosis model, where the operation status data of the source device carries a pre-annotated fault type label; sending the model parameters of the second fault diagnosis model of the source node to the neighbor nodes of the source node, so that the neighbor nodes update the model parameters of the second fault diagnosis model of the neighbor nodes according to the model parameters sent by the source node, and enabling the neighbor nodes to send the model parameters of the second fault diagnosis model of the neighbor nodes to the source node, where the neighbor nodes are source nodes in the federated learning network that have a communication edge with the source node; updating the model parameters of the second fault diagnosis model of the source node according to the model parameters sent by the neighbor nodes and the first loss value; if the second fault diagnosis model of the source node meets the first preset stop condition, taking the model parameters of the second fault diagnosis model that meets the first preset stop condition as the global model parameters.
[0094] In one embodiment, the pre-training module is specifically configured to determine a weight matrix between the source node and the neighbor nodes according to the adjacency matrix between the source node and the neighbor nodes and the degree matrix of the federated learning network; update the model parameters of the source fault diagnosis model according to the weight matrix, the first loss value, and the model parameters sent by the neighbor nodes.
[0095] In one embodiment, the update module 202 is specifically configured to input the associated data of the associated node into the associated fault diagnosis model of the associated node to determine the features of the associated node, where the associated node is a source node associated with the target node, the associated data is the operation status data of the source device associated with the associated node, and the associated fault diagnosis model is the second fault diagnosis model of the associated node configured with the global model parameters; initialize the model parameters of the first fault diagnosis model of the target node according to the global model parameters; update the model parameters of the initialized first fault diagnosis model according to the features of the associated node and the operation status data of the device to be measured associated with the target node until the first fault diagnosis model of the target node meets the second preset stop condition.
[0096] In one embodiment, the update module 202 is specifically configured to the first fault diagnosis model includes: a first feature extractor, a first bottleneck layer, a first classification layer, and a discriminator; where the first feature extractor and the first bottleneck layer are used to extract the features of the input data, the discriminator is used to reduce the distribution difference between the features of the input data and the features of the associated node, and the first classification layer is used to determine the fault type of the input data according to the features extracted by the first feature extractor and the first bottleneck layer, and the input data is the data input into the first fault diagnosis model.
[0097] In one embodiment, the updating module 202 is specifically configured to input the features of the associated nodes into the first bottleneck layer of the initialized first fault diagnosis model to determine the first features of the associated nodes, and input the first features into the discriminator and the first classification layer of the initialized first fault diagnosis model respectively to determine the second features and the fault classification result of the associated nodes; input the target data into the first feature extractor of the initialized first fault diagnosis model to determine the third features of the target data, and input the third features into the first bottleneck layer of the initialized first fault diagnosis model to determine the fourth features of the target data, where the target data is the operation status data of the device under test associated with the target node; input the fourth features into the discriminator of the initialized first fault diagnosis model to determine the fifth features of the target data; calculate the second loss value, the third loss value, and the fourth loss value of the first fault diagnosis model respectively according to the first features, the second features, the third features, the fourth features, the fifth features, and the fault classification result; and update the model parameters corresponding to the first feature extractor, the first bottleneck layer, the first classification layer, and the discriminator in the first fault diagnosis model respectively according to the second loss value, the third loss value, and the fourth loss value.
[0098] In one embodiment, the updating module 202 is specifically configured to update the model parameters corresponding to the first feature extractor according to the second loss value and the third loss value; update the model parameters corresponding to the first bottleneck layer according to the second loss value, the third loss value, and the fourth loss value; and update the model parameters corresponding to the first classification layer and the discriminator respectively according to the fourth loss value.
[0099] For the specific limitations of the fault diagnosis device, reference may be made to the limitations on the fault diagnosis method in the foregoing text, which will not be elaborated here. Each module in the foregoing fault diagnosis device may be implemented in whole or in part by software, hardware, and their combination. The foregoing modules may be embedded in the processor of the computer device in hardware form or independent thereof, or may be stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the foregoing modules.
[0100] Based on the above as Figure 1 shown in the method, correspondingly, an embodiment of the present application further provides a readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above as Figure 1 shown in the fault diagnosis method is implemented.
[0101] Based on such an understanding, the technical solution of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which may be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), including several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute the methods described in various implementation scenarios of the present application.
[0102] Based on the method as described above, and Figure 1 the virtual device embodiments as shown in Figure 2 , in order to achieve the above object, an embodiment of the present application further provides a computer device, which may specifically be a personal computer, a server, a network device, etc. The computer device includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to implement the fault diagnosis method as shown in Figure 1 above.
[0103] Optionally, the computer device may further include a user interface, a network interface, a camera, a radio frequency (RF) circuit, sensors, an audio circuit, a WI-FI module, etc. The user interface may include a display screen (Display), an input unit such as a keyboard (Keyboard), etc. Optionally, the user interface may further include a USB interface, a card reader interface, etc. The network interface may optionally include a standard wired interface, a wireless interface (such as a Bluetooth interface, a WI-FI interface), etc.
[0104] Those skilled in the art can understand that the structure of a computer device provided in this embodiment does not limit the computer device, and it may include more or fewer components, or combine certain components, or have different component arrangements.
[0105] The storage medium may further include an operating system and a network communication module. The operating system is a program for managing and storing the hardware and software resources of the computer device, and supports the operation of information processing programs and other software and / or programs. The network communication module is used to implement communication between components inside the storage medium, and communication between other hardware and software in this entity device.
[0106] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform, or the embodiments of the present application can be implemented by hardware.
[0107] Those skilled in the art can understand that the drawings are only schematic diagrams of a preferred implementation scenario, and the modules or processes in the drawings are not necessarily essential for implementing the present application. Those skilled in the art can understand that the modules in the devices in the implementation scenarios can be distributed in the devices in the implementation scenarios according to the description of the implementation scenarios, or can be changed accordingly and located in one or more devices different from the present implementation scenario. The modules in the above implementation scenarios can be combined into one module, or further split into multiple sub-modules.
[0108] The above serial numbers of the present application are only for description and do not represent the advantages or disadvantages of the implementation scenarios. The above disclosure is only several specific implementation scenarios of the present application. However, the present application is not limited thereto, and any changes that can be conceived by those skilled in the art should fall within the protection scope of the present application.
Claims
1. A fault diagnosis method, characterized in that: The method comprises: Acquire global model parameters of a federated learning network, wherein the federated learning network includes a plurality of source nodes, and the global model parameters are obtained by iteratively updating adjacent source nodes in the federated learning network; Update the model parameters of the first fault diagnosis model of the target node according to the global model parameters; The operating status data of the device under test associated with the target node is input into the first fault diagnosis model to determine the fault type of the device under test.
2. The fault diagnosis method according to claim 1, characterized in that: The method further comprises: Constructing a second fault diagnosis model of the source node; Training a second fault diagnosis model of the source node according to the operating status data of the source device associated with the source node to obtain a first loss value of the second fault diagnosis model, wherein the operating status data of the source device carries a pre-labeled fault type label; Sending the model parameters of the second fault diagnosis model of the source node to the neighboring nodes of the source node, so that the neighboring nodes update the model parameters of the second fault diagnosis model of the neighboring nodes according to the model parameters sent by the source node, and causing the neighboring nodes to send the model parameters of the second fault diagnosis model of the neighboring nodes to the source node, wherein the neighboring nodes are the source nodes having a communication edge with the source node in the federated learning network; updating the model parameters of the second fault diagnosis model of the source node according to the model parameters sent by the neighbor node and the first loss value; If the second fault diagnosis model of the source node meets the first preset stop condition, the model parameters of the second fault diagnosis model that meets the first preset stop condition are used as the global model parameters.
3. The fault diagnosis method according to claim 2, characterized in that: The updating of the model parameters of the second fault diagnosis model of the source node according to the model parameters sent by the neighboring node and the first loss value includes: Determine a weight matrix between the source node and the neighbor node according to an adjacency matrix between the source node and the neighbor node and a degree matrix of the federated learning network; The model parameters of the source fault diagnosis model are updated according to the weight matrix, the first loss value and the model parameters sent by the neighboring node.
4. The fault diagnosis method according to claim 1, characterized in that: The updating of the model parameters of the first fault diagnosis model of the target node according to the global model parameters comprises: Inputting the associated data of the associated node into the associated fault diagnosis model of the associated node to determine the characteristics of the associated node, wherein the associated node is the source node associated with the target node, the associated data is the operating status data of the source device associated with the associated node, and the associated fault diagnosis model is a second fault diagnosis model of the associated node configured with the global model parameters; Initialize model parameters of a first fault diagnosis model of the target node according to the global model parameters; The model parameters of the initialized first fault diagnosis model are updated according to the characteristics of the associated node and the operating status data of the device under test associated with the target node, until the first fault diagnosis model of the target node meets the second preset stop condition.
5. The fault diagnosis method according to claim 4, characterized in that: The first fault diagnosis model includes: a first feature extractor, a first bottleneck layer, a first classification layer and a discriminator; Among them, the first feature extractor and the first bottleneck layer are used to extract features of input data, the discriminator is used to reduce the distribution difference between the features of the input data and the features of the associated nodes, and the first classification layer is used to determine the fault type of the input data based on the features extracted by the first feature extractor and the first bottleneck layer, and the input data is the data input into the first fault diagnosis model.
6. The fault diagnosis method according to claim 5, characterized in that: The updating of the model parameters of the initialized first fault diagnosis model according to the operating status data of the device under test associated with the target node and the characteristics of the associated node comprises: Input the feature of the associated node into the first bottleneck layer of the initialized first fault diagnosis model to determine the first feature of the associated node, and input the first feature into the discriminator and the first classification layer of the initialized first fault diagnosis model respectively to determine the second feature of the associated node and the fault classification result; Inputting target data into the first feature extractor in the initialized first fault diagnosis model to determine a third feature of the target data, and inputting the third feature into the first bottleneck layer in the initialized first fault diagnosis model to determine a fourth feature of the target data, wherein the target data is operating status data of the device under test associated with the target node; Inputting the fourth feature into the discriminator in the initialized first fault diagnosis model to determine the fifth feature of the target data; Calculate a second loss value, a third loss value, and a fourth loss value of the first fault diagnosis model respectively according to the first feature, the second feature, the third feature, the fourth feature, the fifth feature, and the fault classification result; According to the second loss value, the third loss value and the fourth loss value, model parameters corresponding to the first feature extractor, the first bottleneck layer, the first classification layer and the discriminator in the first fault diagnosis model are updated respectively.
7. The fault diagnosis method according to claim 6, characterized in that: The updating of model parameters corresponding to the first feature extractor, the first bottleneck layer, the first classification layer and the discriminator in the first fault diagnosis model according to the second loss value, the third loss value and the fourth loss value respectively includes: Update the model parameters corresponding to the first feature extractor according to the second loss value and the third loss value; Update the model parameters corresponding to the first bottleneck layer according to the second loss value, the third loss value and the fourth loss value; The model parameters corresponding to the first classification layer and the discriminator are respectively updated according to the fourth loss value.
8. A fault diagnosis device, characterized in that: The device comprises: An acquisition module, used to acquire global model parameters of a federated learning network, wherein the federated learning network includes a plurality of source nodes, and the global model parameters are obtained by iteratively updating adjacent source nodes in the federated learning network; An updating module, used for updating the model parameters of the first fault diagnosis model of the target node according to the global model parameters; The determination module is used to input the operating status data of the device under test associated with the target node into the first fault diagnosis model to determine the fault type of the device under test.
9. A readable storage medium having a program or instruction stored thereon, characterized in that: When the program or instruction is executed by a processor, the steps of the fault diagnosis method according to any one of claims 1 to 7 are implemented.
10. A computer device comprising a storage medium, a processor, and a computer program stored in the storage medium and executable on the processor, characterized in that: When the processor executes the program, the fault diagnosis method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Fault diagnosis method, system and device, electronic equipment and storage medium
CN120448178A