A wind turbine unit edge cloud cooperative fault diagnosis system based on knowledge federation

CN118188346BActive Publication Date: 2026-09-25YANSHAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410283500.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-13
Publication Date
2026-09-25
Estimated Expiration
2044-03-13

AI Technical Summary

Benefits of technology

[0067]由于采用了上述技术方案,本发明取得的技术进步是:通过采用联邦学习的服务器-客户端训练架构来实现原始数据隐私保护的分布式风电机组智能故障诊断。使用学习自每个边缘客户端的诊断知识代替传统联邦框架中的模型参数,并在中央服务器上利用诊断知识形成一个全局诊断知识仓库。由于该知识仓库覆盖了各客户端上所有可能发生的故障类型,所以可以显著缓解客户端之间故障标签极端异构所造成的不利影响。通过在全局诊断模型使用集成半监督训练学习算法可以便捷地集成到服务器上全局诊断模型的训练过程中,从而有效缓解有限标签数据的影响,有效应对客户端上故障数据标签不足的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118188346B_ABST
    Figure CN118188346B_ABST
Patent Text Reader

Abstract

The application discloses a wind turbine edge cloud cooperative fault diagnosis system based on knowledge federation, and belongs to the technical field of distributed wind turbine intelligent fault diagnosis. The system is constructed by taking distributed wind turbines as edge clients and a central cloud server. The edge clients use a space-time memory enhanced self-encoder to perform first-stage training by using local wind turbine monitoring data, and the edge clients upload learned knowledge to the central server. A comprehensive global diagnosis knowledge warehouse is formed on the server. Based on the global warehouse, the server selects a suitable paradigm according to actual conditions to perform second-stage training, i.e. global diagnosis model training. Finally, the server distributes the trained global diagnosis model to each edge client. The application can enhance the diagnosis performance of the diagnosis model in the edge client under an extreme heterogeneous scenario of fault labels while ensuring the privacy of the original data of the user, and can also relieve the computing burden on the edge client.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent fault diagnosis technology for distributed wind turbine generators, and in particular to a knowledge federation-based edge-cloud collaborative fault diagnosis system for wind turbine generators. Background Technology

[0002] Wind energy, as a green and clean renewable energy source, has achieved tremendous development in recent years, making a significant contribution to addressing the global energy crisis and climate change. However, wind turbines typically operate in harsh, remote environments. Furthermore, modern wind turbines generally consist of numerous sub-components with complex coupling relationships, making them highly susceptible to a wide variety of faults. These faults significantly impact the power generation efficiency of wind turbines, with some severe faults even leading to equipment shutdowns. This not only results in high maintenance costs and economic losses but, more importantly, reduces the continuity, stability, and reliability of the power system. Therefore, health monitoring and intelligent fault diagnosis of wind turbines are of significant practical importance.

[0003] Due to the tremendous success and widespread application of artificial intelligence algorithms, represented by deep learning, in recent years, coupled with the fact that modern wind turbines almost all have locally installed SCADA systems that provide massive and diverse monitoring data such as speed, power, temperature, wind speed, and yaw angle, intelligent fault diagnosis methods for wind turbines based on SCADA data have received widespread attention and research. Although various advanced deep learning technologies have continuously improved diagnostic performance, a large number of high-quality and diverse supervised datasets for training diagnostic models remain an essential foundation for these achievements. Traditional deep model-based wind turbine diagnostic methods typically centralize the operational data generated from distributed wind turbines, forming a centralized training dataset on a central server. However, as data plays an increasingly important role in the current information age, and given that the condition monitoring data generated from wind turbines often contains rich information about technical details and commercial interests, data owners are unwilling to transmit and share their data outside their local storage. Therefore, storing distributed data directly locally has become a more preferred option for users. Therefore, how to fully utilize the local data of all parties while protecting the privacy of the original data, and then train a general global wind turbine fault diagnosis model, has become an urgent problem to be solved.

[0004] Federated learning offers an effective solution to the aforementioned problems. Its distributed computing paradigm aligns with the distributed nature of wind turbine equipment and data, allowing participants to collaborate while ensuring users' original privacy data remains locally. This leads to a global diagnostic model on the server side with performance similar to centralized training. However, in practice, directly utilizing the traditional federated learning paradigm for collaborative intelligent fault diagnosis of distributed wind turbines still presents several challenges: different distributed wind turbines typically encounter different fault types. This difference manifests as severe heterogeneity in fault labels across different clients under the federated learning paradigm, ultimately resulting in a significant degradation in the diagnostic performance of the global model. Furthermore, during actual wind turbine operation, not all monitoring data possesses label information. More commonly, only a small portion of the local dataset contains labeled data, while the majority remains unlabeled. This label scarcity severely impacts the diagnostic performance of the model under the federated learning paradigm. Meanwhile, traditional federated learning uses an average aggregation scheme for uploaded model parameters on the server. This leads to the blurring of personalized information on each client during the averaging process, resulting in a severe degradation of the diagnostic performance of the global model. Furthermore, the server's focus on averaging also wastes the central server's abundant computing resources. Summary of the Invention

[0005] The technical problem this invention aims to solve is to provide a knowledge-federated edge-cloud collaborative fault diagnosis system for wind turbines. By employing a federated learning paradigm, it can effectively mitigate the severe impact of extremely heterogeneous fault types and limited labeled data between different turbines on the diagnostic model, while fully ensuring the privacy of users' original data. Simultaneously, the edge-cloud collaborative training mechanism used in this system can alleviate the computational burden on the client side and improve the overall training efficiency of the system.

[0006] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is: a knowledge federation-based wind turbine edge-cloud collaborative fault diagnosis system, the specific construction steps of which are as follows:

[0007] Step 1: Build a server-client architecture, with a single distributed wind turbine as an edge client, and a cloud server built on top of each edge client. Upload and download can be realized between the cloud server and each edge client;

[0008] Step 2: Local diagnostic knowledge learning. A diagnostic knowledge learning model is built in each edge client. The model is trained using local monitoring data to acquire diagnostic knowledge. After training, each edge client uploads the learned diagnostic knowledge to the cloud server.

[0009] Step 3: Server-side global diagnostic model training. The diagnostic knowledge uploaded by each edge client is merged and summarized in the cloud server to form a diagnostic knowledge repository. A global diagnostic model is built in the cloud server, and the diagnostic knowledge repository is used to train the global diagnostic model.

[0010] Step 4: After the edge-cloud collaborative training is completed, the cloud server will distribute the trained global diagnostic model to each edge client, and combine it with the encoder in the local diagnostic knowledge learning model in the edge client to form a local diagnostic model.

[0011] A further improvement of the technical solution of the present invention is that: in the server-client architecture constructed in step 1, the status monitoring data generated by the wind turbine in each edge client during actual operation is only stored in the local client and is not transmitted or shared outside the local client, and the local data of each client is not visible to the cloud server.

[0012] A further improvement to the technical solution of the present invention is as follows: Step 2 is specifically as follows:

[0013] Step 2.1: Each edge client will obtain the corresponding wind turbine's operating status monitoring data from its local private status monitoring dataset, divide a portion of the local private dataset into a test set to test the performance of the obtained model, and divide the remaining portion into a training set for model training and parameter tuning;

[0014] Step 2.2: Construct a diagnostic knowledge learning model. Train the diagnostic knowledge learning model in an unsupervised manner using local monitoring data. When the diagnostic knowledge learning model converges, the output of the diagnostic knowledge learning model is regarded as diagnostic knowledge that fully represents the original data.

[0015] Step 2.3: Each edge client uploads diagnostic knowledge to the cloud server. If the local monitoring data on the edge client has a corresponding fault tag, the diagnostic knowledge and its corresponding tag are uploaded at the same time. If the local monitoring data does not have a corresponding tag, only the learned diagnostic knowledge needs to be uploaded.

[0016] A further improvement to the technical solution of the present invention is as follows: Step 2.2 is as follows:

[0017] The diagnostic knowledge learning model is a spatiotemporal memory-enhanced autoencoder, consisting of an encoder and a decoder. The encoder learns diagnostic knowledge in the time and space dimensions, respectively.

[0018] The local monitoring data for wind turbine condition monitoring is input into the encoder and first undergoes a spatiotemporal convolution operation, as shown in equations (1) and (2):

[0019]

[0020]

[0021] in and These are the latent vectors calculated through convolution, c t and c s It is the dimension of the vector. and These represent convolutional networks in the time and space dimensions, respectively. The learnable parameters in this network are represented by... and express;

[0022] When performing convolution calculations in both the temporal and spatial dimensions, the resulting vectors are input into corresponding attention-based memory modules. Each memory module contains a memory matrix, and the temporal and spatial memory matrices can be represented as follows: and Each memory matrix contains N memory units, and the memory unit vector is... as well as The result obtained after convolution is input into the memory module;

[0023] When calculating the time dimension, the convolution result is... As query vector and memory matrix The cosine similarity of each memory unit vector in the equation is calculated as shown in equation (3):

[0024]

[0025] The attention weight vector can be calculated. Then, the softmax function is used to map all elements of the weight vector to a sum of 1, as shown in equation (4):

[0026]

[0027] Finally, we can obtain the normalized attention weight vector in the time dimension:

[0028]

[0029] When calculating spatial dimensions, the convolution result is... As query vector and memory matrix The cosine similarity of each memory unit vector in the equation is calculated as shown in equation (5):

[0030]

[0031] The attention weight vector can be calculated. Then, the softmax function is used to map all elements of the weight vector to a sum of 1, as shown in equation (6):

[0032]

[0033] Finally, we can obtain the normalized attention weight vector in the time dimension:

[0034]

[0035] The obtained attention weights are used to perform linear weighting operations on the memory units in the temporal and spatial memory matrices, as shown in equations (7) and (8):

[0036]

[0037]

[0038] in This is the final output of the memory module in the time dimension; This is the final output of the memory module in the spatial dimension;

[0039] To ensure that the encoder can learn the most essential features about the fault mode while maintaining its generalization ability to different fault types, a skip connection is introduced between the output of the convolution operation and the memory module to bring common fault information to the model; this process is shown in Equation (9) in the time dimension:

[0040]

[0041] In terms of spatial dimensions, it is as shown in equation (10):

[0042]

[0043] Finally, the temporal and spatial results are added together and fused to obtain the final output of the encoder, as shown in equation (11):

[0044]

[0045] in and First, the temporal and spatial computation results are converted into two matrices of the same shape that can be directly added using convolution, thus obtaining the final output of the encoder.

[0046] The decoder consists of multiple deconvolutional layers, and its overall calculation process is shown in equation (12):

[0047]

[0048] in It is the original input signal restored by the decoder. This represents the trainable parameters of the decoder;

[0049] The entire diagnostic knowledge learning model is trained and updated using the MSE loss function, and the loss on the j-th edge client is shown in equation (13):

[0050]

[0051] Each edge client trains its local diagnostic knowledge learning model using the process described above. Once the local diagnostic knowledge learning model converges, each edge client uses the encoder output as the learned diagnostic knowledge.

[0052] A further improvement to the technical solution of the present invention is as follows: Step 3 is specifically as follows:

[0053] Step 3.1: The cloud server receives diagnostic knowledge from different edge clients and uses this uploaded diagnostic knowledge to form a global diagnostic knowledge repository on the server. This knowledge repository can comprehensively cover all fault types on each client, effectively expanding the server's scope and providing the server with richer information than model parameters.

[0054] Step 3.2: Construct a global diagnostic model. The global diagnostic model adopts a multi-scale spatiotemporal convolution strategy, introducing a multi-scale strategy in the time dimension; input diagnostic knowledge data from the global repository into the global diagnostic model;

[0055] Step 3.3: Construct a training mechanism. When the diagnostic knowledge in the diagnostic knowledge base has corresponding labels, general supervised learning is used to complete the training of the global diagnostic model. If the local monitoring data of the edge client has only a small amount of labeled data and the majority of the remaining data is unlabeled data, and most of the diagnostic knowledge uploaded to the global repository of the server also does not have corresponding labels, then ensemble semi-supervised learning is used to complete the training of the global diagnostic model.

[0056] Step 3.4: The central server selects an appropriate training method based on the actual situation to train the constructed global diagnostic model.

[0057] The further improvement of the technical solution of the present invention is as follows: Step 3.2 is as follows: The global diagnostic model uses a spatiotemporal decoupling method to complete the further extraction of features. In the time dimension, a multi-scale convolution strategy is adopted, that is, the input data is computed in parallel through convolutional layers with convolutional kernels of 1, 3 and 5 respectively to achieve better feature learning. This process is shown in Equation (15):

[0058]

[0059] Among them OT This represents the output of three temporal multi-scale convolutional modules; Represents three different scales of convolution operations; I T Represents shared input data, I T This does not represent the initial input of the diagnostic model, but rather a feature vector that has undergone initial convolution and max pooling operations, as shown in Equation (16):

[0060]

[0061] Where K is the final output of the encoder in the local diagnostic knowledge learning model, that is, the knowledge data in the global knowledge warehouse; in the spatial dimension, the original input diagnostic knowledge is directly convolved in one dimension, as shown in Equation (17):

[0062] O S =Conv S (K) (17)

[0063] Then, the convolution results in the two dimensions are added and fused together and input into the fully connected layer to finally obtain the prediction output of the global diagnostic model, as shown in Equation (18):

[0064] O M =FC(O T +O S (18)

[0065] Where FC(·) represents the computation process of the fully connected layer, O M This is the final output of the global diagnostic model.

[0066] A further improvement of the technical solution of the present invention is that: in step 4, the edge-cloud collaborative training is a two-stage-single-round edge-cloud collaborative training mechanism, and the information interaction between each edge client and the cloud server only requires one upload and one download operation.

[0067] The technological advancements achieved by this invention, due to the adoption of the aforementioned technical solution, are as follows: Intelligent fault diagnosis of distributed wind turbines with original data privacy protection is achieved through a server-client training architecture using federated learning. Diagnostic knowledge learned from each edge client replaces the model parameters in the traditional federated framework, and a global diagnostic knowledge repository is formed on the central server using this diagnostic knowledge. Since this knowledge repository covers all possible fault types on each client, it can significantly mitigate the adverse effects caused by the extreme heterogeneity of fault labels among clients. Furthermore, by using an ensemble semi-supervised training learning algorithm in the global diagnostic model, it can be easily integrated into the training process of the global diagnostic model on the server, thereby effectively mitigating the impact of limited labeled data and effectively addressing the problem of insufficient fault data labels on the clients.

[0068] This paper proposes a two-stage, single-round edge-cloud collaborative training mechanism for this federated knowledge framework. This mechanism distributes the training task of the deep learning model across both the client edge device and the server platform, requiring only a single global communication process for uploading and downloading. Compared to the traditional federated learning training paradigm, which completes the entire deep model training process on the client and requires multiple global communications of model parameters between the client and server, this mechanism leverages the stronger computing power of the server, reduces the computational burden on the client, and improves the overall training efficiency of the framework.

[0069] For wind turbine condition monitoring SCADA data on edge clients, which is a typical multivariate time series with complex spatiotemporal coupling, a spatiotemporal memory-enhanced autoencoder is implemented to learn diagnostic knowledge from local private data. This autoencoder employs an encoder-decoder architecture, generally using a strategy of decoupling operations in the temporal and spatial dimensions. Specifically, convolutional operations, attention-based memory modules, and treaty connections are deployed in the encoder in both the temporal and spatial dimensions. The decoder then reconstructs the original input signal through multiple deconvolutional layers. Through the computation process in the encoder, the local model can acquire the most representative diagnostic knowledge while preserving generalization performance, thus laying a solid foundation for the subsequent training of the global diagnostic model. Attached Figure Description

[0070] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0071] Figure 1 This is a schematic diagram of the fault diagnosis system of the present invention;

[0072] Figure 2 This is a schematic diagram of the structure of the autoencoder that enhances the spatiotemporal memory of diagnostic knowledge learned by this invention;

[0073] Figure 3 This is a schematic diagram of the global diagnostic model structure on the cloud server; Detailed Implementation

[0074] The present invention will be further described in detail below with reference to embodiments:

[0075] like Figure 1 The image shows a knowledge-fed edge-cloud collaborative fault diagnosis system for wind turbines. The specific construction steps are as follows:

[0076] Step 1: Construct a server-client architecture. Wind turbines are typically located in remote areas, exhibiting significant distributed characteristics among numerous wind turbines. Each distributed wind turbine is treated as an edge client, and a cloud server is built on top of each edge client. Upload and download functionality is available between the cloud server and each edge client. Wind turbine operational status monitoring data is stored locally on each client and is not shared between different clients; each client's local data is invisible to the cloud server. This fully guarantees the invisibility of wind turbine status monitoring data and protects user privacy. The cloud server can coordinate and manage multiple edge clients to form the proposed federated learning system for edge-cloud collaborative fault diagnosis of wind turbines.

[0077] Step 2: Local diagnostic knowledge learning. A diagnostic knowledge learning model is built in each edge client. The model is trained using local monitoring data to acquire diagnostic knowledge. After training, each edge client uploads the learned diagnostic knowledge to the cloud server.

[0078] Step 2.1: Each edge client will obtain the corresponding wind turbine's operating status monitoring data from its local private status monitoring dataset. A portion of the local private dataset will be divided into a test set for testing the performance of the obtained model, and the remaining portion will be divided into a training set for model training and parameter tuning. In this embodiment, each edge client first obtains the corresponding wind turbine's operating status monitoring data from its local private dataset. The private local dataset on the j-th edge client can be represented as... Where N j N represents the total number of samples on the j-th client. client This indicates the number of wind turbine units participating in the fault diagnosis model training process. Let represent the i-th sample on client j, where m and n represent the length of the sample in the time dimension and the number of wind turbine condition monitoring sensors, respectively, i.e., the spatial dimension. The local dataset D of edge client j... j A local test set is partitioned from the data. The remaining portion is used to test the performance of the current model, while the rest is used to train a local diagnostic knowledge learning model.

[0079] Step 2.2: Wind turbine condition monitoring data is a multivariate time series with complex spatiotemporal coupling relationships. To learn representative diagnostic knowledge from it without losing generalization, a diagnostic knowledge learning model is constructed. This model is used to learn representative and generalizable diagnostic knowledge from the original condition monitoring data of the wind turbine to replace the model parameters in the traditional federated learning scheme, thereby improving the diagnostic performance of the global model in the face of extreme heterogeneous fault types.

[0080] The knowledge learning model employs a spatiotemporal memory-enhanced autoencoder, consisting of an encoder and a decoder. Considering that fault mode information is simultaneously embedded in the temporal and spatial dimensions of the monitoring data, the encoder learns diagnostic knowledge in both the temporal and spatial dimensions. The specific structure is as follows: Figure 2 As shown. The local monitoring data of the wind turbine condition monitoring is input to the encoder and first undergoes a spatiotemporal convolution operation, as shown in equations (1) and (2):

[0081]

[0082]

[0083] in and These are the latent vectors calculated through convolution, c t and c s It is the dimension of the vector. and These represent convolutional networks in the time and space dimensions, respectively. The learnable parameters in this network are represented by... and express;

[0084] After performing convolution calculations in both the temporal and spatial dimensions, the resulting vectors are input into corresponding attention-based memory modules. Through these memory modules, the model can learn the most essential features related to the corresponding fault modes. A memory module contains a memory matrix, and the temporal and spatial memory matrices can be represented as follows: and Each memory matrix contains N memory units, and the memory unit vector is... as well as The result obtained from the convolution calculation is input into the memory module. When calculating the time dimension, the convolution result is... As query vector and memory matrix The cosine similarity of each memory unit vector in the equation is calculated as shown in equation (3):

[0085]

[0086] The attention weight vector can be calculated. Then, the softmax function is used to map all elements of the weight vector to a sum of 1, as shown in equation (4):

[0087]

[0088] Finally, we can obtain the normalized attention weight vector in the time dimension.

[0089] In the spatial dimension, the convolution result is also... As query vector and memory matrix The cosine similarity is calculated for each memory unit vector in the dataset. Then, the softmax function is used to map all elements of the weight vector to a sum of 1, ultimately obtaining the attention weight vector in the spatial dimension. The calculations in the spatial dimension are shown in equations (5) and (6):

[0090]

[0091]

[0092] The obtained attention weights are used to perform linear weighting operations on the memory units in the temporal and spatial memory matrices, as shown in equations (7) and (8):

[0093]

[0094]

[0095] in This represents the final output of the memory module in the time dimension. This is the final output of the memory module in the spatial dimension;

[0096] To ensure that the encoder maintains its generalization ability to different fault types while learning the most essential features of the fault mode, and to ensure that the encoder has generalization ability in the process of learning diagnostic knowledge, a skip connection is introduced between the output of the convolution operation and the memory module to bring common fault information to the model; this process is shown in Equation (9) in the time dimension:

[0097]

[0098] The calculation process in the spatial dimension is shown in equation (10), and the result is expressed as follows:

[0099]

[0100] Finally, the temporal and spatial results are added together and fused to obtain the final output of the encoder, as shown in equation (11):

[0101]

[0102] in and First, the temporal and spatial computation results are converted into two matrices of the same shape that can be directly added using convolution, thus obtaining the final output of the encoder.

[0103] The decoder consists of multiple deconvolutional layers, used to completely reconstruct the original input data at its output; its overall calculation process is shown in equation (12):

[0104]

[0105] in It is the original input signal restored by the decoder. This represents the trainable parameters of the decoder;

[0106] The entire diagnostic knowledge learning model is trained and updated using the MSE loss function, and the loss on the j-th edge client is shown in equation (13):

[0107]

[0108] Each edge client trains its local diagnostic knowledge learning model using the process described above. When the local diagnostic knowledge learning model converges, each edge client uses the encoder output as the learned diagnostic knowledge that is considered to fully represent the original data.

[0109] Step 2.3: Each edge client uploads diagnostic knowledge to the cloud server. If the local monitoring data on the edge client has a corresponding fault tag, the diagnostic knowledge and its corresponding tag are uploaded at the same time. If the local monitoring data does not have a corresponding tag, only the learned diagnostic knowledge needs to be uploaded.

[0110] Step 3: Server-side global diagnostic model training. The diagnostic knowledge uploaded by each edge client is merged and summarized in the cloud server to form a diagnostic knowledge repository. A global diagnostic model is built in the cloud server, and the diagnostic knowledge repository is used to train the global diagnostic model.

[0111] Step 3.1: The cloud server receives diagnostic knowledge from different edge clients and uses this uploaded diagnostic knowledge to form a global diagnostic knowledge repository on the server, as shown in equation (14):

[0112]

[0113] in This represents the collection of all diagnostic knowledge learned on client j. This knowledge repository can comprehensively cover all fault types on each client; it effectively expands the server's perspective and provides the server with richer information than the model parameters.

[0114] Step 3.2: Construct a global diagnostic model. The specific structure of this global model is as follows: Figure 3 As shown, the spatiotemporal decoupling method is also used to further extract features. Furthermore, it employs a multi-scale convolution strategy in the time dimension, that is, it uses convolutional layers with kernels of 1, 3, and 5 to compute the input data in parallel to achieve better feature learning. This process is shown in equation (15):

[0115]

[0116] Among them O T This represents the output of three temporal multi-scale convolutional modules; Represents three different scales of convolution operations; I T This represents shared input data; it's important to note that I... T This does not represent the initial input of the diagnostic model, but rather a feature vector that has undergone initial convolution and max pooling operations, as shown in Equation (16):

[0117]

[0118] Where K is the final output of the encoder in the local diagnostic knowledge learning model, which is the knowledge data in the global knowledge warehouse. In terms of spatial dimension, the original input diagnostic knowledge is directly convolved in one dimension, as shown in Equation (17):

[0119] O S =Conv S (K) (17)

[0120] Then, the convolution results in the two dimensions are added and fused, and then input into the fully connected layer to finally obtain the prediction output of the global diagnostic model, as shown in Equation (18):

[0121] O M =FC(O T +O S (18)

[0122] Where FC(·) represents the computation process of the fully connected layer, O M This is the final output of the global diagnostic model;

[0123] Step 3.3: Construct a training mechanism. When the diagnostic knowledge in the diagnostic knowledge base has a corresponding label, general supervised learning is used to train the global diagnostic model, and the general cross-entropy loss function is used to update the model parameters. This loss function can be expressed by equation (19):

[0124]

[0125] Where N represents the number of sample pairs (K) in the global repository. i ,yi The total number of y i These are the corresponding actual labels, and C is the number of all data categories visible to the server in the global repository. That is An element in the equation represents the probability that the i-th input data is predicted to be of class c, i.e., the probability of the global model predicting the input K. i The prediction results.

[0126] If the edge client's local monitoring data contains only a small amount of labeled data while the majority is unlabeled, then the data uploaded to the server's global repository will consist of a portion of labeled data and a portion of unlabeled data. This scenario is a typical application of semi-supervised learning. Integrated semi-supervised learning can be used to train the global diagnostic model, thereby mitigating the performance degradation caused by data label scarcity and improving the global model's diagnostic capabilities under label-scarce conditions.

[0127] Step 3.4: The central server selects an appropriate training paradigm based on the actual situation to train the constructed global diagnostic model.

[0128] Step 4: Edge-cloud collaborative training is complete. The cloud server distributes the trained global diagnostic model to each edge client, where it combines with the encoder in the local diagnostic knowledge learning model to form a local diagnostic model. Edge-cloud collaboration divides the entire model training process into two phases, allocated to the edge clients and the central server respectively. This forms a two-phase, single-round edge-cloud collaborative training mechanism. In the first training phase, each edge client uses its private state monitoring data to train its local diagnostic knowledge learning model. In the second training phase, the central server uses the global knowledge repository formed by the uploaded diagnostic knowledge to train the global diagnostic model. The information exchange between the client and server requires only one upload and one download operation: the client uploads diagnostic knowledge and corresponding tags (if they exist) to the server, and the server transmits the trained global diagnostic model to each client. This mechanism reduces the computational burden on the client while improving overall training efficiency.

[0129] After the above steps, each edge client can obtain a globally shared diagnostic model. By combining this model with the encoder in the local model, intelligent fault diagnosis of distributed wind turbines can be achieved in a collaborative manner between the edge and the cloud while protecting the privacy of the original data.

[0130] Example 1:

[0131] This embodiment uses a SCADA dataset generated from a baseline wind turbine model. The sampling interval in this SCADA dataset is 0.0125s, and the sampling duration is 100s, resulting in a total of 8000 sampling points in each time series sample. The SCADA dataset uses 15 sensor variables and has 11 data types, including one type of normal data and 10 types of fault data. Each data type has 450 time series samples, for a total of 4950 samples. In this embodiment, three extremely heterogeneous fault types are set up, with 3, 5, and 10 clients in each scenario, and the fault types on the edge clients in each scenario are completely non-overlapping. This embodiment aggregates the local test sets from all clients into a comprehensive global test set, repeats the experiment three times in each scenario, and finally reports the average accuracy of the client diagnostic model on the global test set.

[0132] Table 1 compares the accuracy of the federated framework in this embodiment with other federated frameworks specifically designed to address client heterogeneity issues in three scenarios, while also considering the performance of the diagnostic model obtained in a centralized scenario. It can be seen that the performance of all baselines exhibits varying degrees of decline, as scenarios 1 through 3 represent an increasing degree of client data heterogeneity. However, the proposed diagnostic knowledge-based method demonstrates superior diagnostic capabilities. It achieves significant performance improvements even with increasing client data heterogeneity because when there is only one fault category on the client, the autoencoder can focus entirely on learning this specific data distribution, thereby providing purer local diagnostic knowledge to the global diagnostic model, thus further improving the diagnostic effect.

[0133] Table 1. Comparison of model performance between different federation frameworks and centralized methods in different distributed scenarios.

[0134]

[0135]

[0136] Table 2 shows the overall training efficiency and resource requirements of the framework proposed in this invention in different distributed scenarios. It can be seen that the federated learning framework based on diagnostic knowledge described in this invention, with the help of its two-stage-single-round edge-cloud collaborative training mechanism, not only has the highest training efficiency (i.e., requires the least training time), but also has the lowest demand for client computing resources, including communication resources between the client and the server, compared to other distributed federated frameworks. This means that the edge-cloud collaborative federated fault diagnosis framework proposed in this invention achieves the best diagnostic performance while simultaneously possessing higher training efficiency and lower resource requirements.

[0137] Table 2 Comparison of training efficiency and resource utilization of different federated frameworks in different distributed scenarios.

[0138]

[0139]

[0140] Table 3 shows the diagnostic accuracy of the final diagnostic models obtained by edge clients using local autoencoders with different structures in distributed scenario 1. It can be seen that among all the different local autoencoder structures, the final diagnostic performance obtained using the spatiotemporal memory-enhanced autoencoder proposed in this invention is the best. This indicates that the local autoencoder on each edge client can learn the most essential features about the original data, and the diagnostic knowledge learned using this model is most beneficial for the subsequent training of the global diagnostic model.

[0141] Table 3. Accuracy of diagnostic models obtained using local autoencoders with different structures

[0142]

[0143] Table 4 shows the diagnostic performance of different methods under distributed scenario 1, when different proportions of unlabeled data exist on each edge client. In this embodiment of the invention, these methods include a federated framework based on diagnostic knowledge that does not use semi-supervised algorithms on the server, a federated framework based on diagnostic knowledge that uses two different semi-supervised algorithms, and a centralized method that does not use semi-supervised algorithms. It can be seen that, without using any enhancement strategies, when the labeled data on each client is only 10%, it still performs better than all the previously mentioned baseline methods, which shows that the proposed framework has good robustness in the face of insufficient labeled data. Furthermore, if semi-supervised learning is added on the server side, the performance of the global model is improved to varying degrees, which shows the improvement brought by directly integrating semi-supervised learning on the server to the model's diagnostic effect.

[0144] Table 4. Diagnostic accuracy obtained by different methods under different proportions of labeled data.

[0145]

[0146]

[0147] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made by those skilled in the art to the technical solutions of the present invention without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.

Claims

1. A knowledge-federation-based edge-cloud collaborative fault diagnosis system for wind turbine units, characterized in that: The specific construction steps are as follows: Step 1: Build a server-client architecture, with a single distributed wind turbine as an edge client, and a cloud server built on top of each edge client. Upload and download can be realized between the cloud server and each edge client; Step 2: Local diagnostic knowledge learning. A diagnostic knowledge learning model is built in each edge client. The model is trained using local monitoring data to acquire diagnostic knowledge. After training, each edge client uploads the learned diagnostic knowledge to the cloud server. Step 3: Server-side global diagnostic model training. The diagnostic knowledge uploaded by each edge client is merged and aggregated into a diagnostic knowledge repository on the cloud server. A global diagnostic model is then built on the cloud server, and the diagnostic knowledge repository is used to train the global diagnostic model. The specific steps are as follows: Step 3.1: The cloud server receives diagnostic knowledge from different edge clients and uses this uploaded diagnostic knowledge to form a global diagnostic knowledge repository on the server. This knowledge repository comprehensively covers all fault types on each client. It effectively expands the server's perspective and provides the server with richer information than the model parameters; Step 3.2: Construct a global diagnostic model. The global diagnostic model adopts a multi-scale spatiotemporal convolution strategy, introducing a multi-scale strategy in the time dimension; input diagnostic knowledge data from the global repository into the global diagnostic model; the specific steps are as follows: The global diagnostic model uses a spatiotemporal decoupling approach to extract features. In the time dimension, it employs a multi-scale convolution strategy, that is, it uses convolutional layers with kernels of 1, 3, and 5 to compute the input data in parallel to achieve feature learning. This process is shown in Equation (15): in, This represents the output of three temporal multi-scale convolutional modules; This represents three different scales of convolution operations; Represents shared input data. This does not represent the initial input of the diagnostic model, but rather a feature vector that has undergone initial convolution and max pooling operations, as shown in Equation (16): in, This is the final output of the encoder in the local diagnostic knowledge learning model, which is also the knowledge data in the global knowledge warehouse; in the spatial dimension, the original input diagnostic knowledge is directly convolved in one dimension, as shown in Equation (17): Then, the convolution results in the two dimensions are added and fused, and then input into the fully connected layer to finally obtain the prediction output of the global diagnostic model, as shown in Equation (18): in, This represents the computation process of a fully connected layer. This is the final output of the global diagnostic model; Step 3.3: Construct a training mechanism. When the diagnostic knowledge in the diagnostic knowledge base has corresponding labels, supervised learning is used to complete the training of the global diagnostic model. If the local monitoring data of the edge client has only a small amount of labeled data and the majority of the remaining data is unlabeled data, and most of the diagnostic knowledge uploaded to the global repository of the server also does not have corresponding labels, then ensemble semi-supervised learning is used to complete the training of the global diagnostic model. Step 3.4: The central server selects an appropriate training method based on the actual situation to train the constructed global diagnostic model; Step 4: After the edge-cloud collaborative training is completed, the cloud server will distribute the trained global diagnostic model to each edge client, and combine it with the encoder in the local diagnostic knowledge learning model in the edge client to form a local diagnostic model.

2. The wind turbine edge-cloud collaborative fault diagnosis system based on knowledge federation according to claim 1, characterized in that: In the server-client architecture constructed in step 1, the status monitoring data generated by the wind turbines during actual operation in each edge client is only stored on the local client and is not transmitted or shared outside the local client. The local data of each client is not visible to the cloud server.

3. The wind turbine edge-cloud collaborative fault diagnosis system based on knowledge federation according to claim 1, characterized in that: Step 2 is detailed below: Step 2.1: Each edge client will obtain the corresponding wind turbine's operating status monitoring data from its local private status monitoring dataset, divide a portion of the local private dataset into a test set to test the performance of the obtained model, and divide the remaining portion into a training set for model training and parameter tuning; Step 2.2: Construct a diagnostic knowledge learning model. Train the diagnostic knowledge learning model in an unsupervised manner using local monitoring data. When the diagnostic knowledge learning model converges, the output of the diagnostic knowledge learning model is regarded as diagnostic knowledge that fully represents the original data. Step 2.3: Each edge client uploads diagnostic knowledge to the cloud server. If the local monitoring data on the edge client has a corresponding fault tag, the diagnostic knowledge and its corresponding tag are uploaded at the same time. If the local monitoring data does not have a corresponding tag, only the learned diagnostic knowledge needs to be uploaded.

4. The wind turbine edge-cloud collaborative fault diagnosis system based on knowledge federation according to claim 3, characterized in that: Step 2.2 The specific steps are as follows: The diagnostic knowledge learning model is a spatiotemporal memory-enhanced autoencoder, consisting of an encoder and a decoder. The encoder learns diagnostic knowledge in the time and space dimensions, respectively. The local monitoring data for wind turbine condition monitoring is input into the encoder and first undergoes a spatiotemporal convolution operation, as shown in equations (1) and (2): in, and These are the latent vectors calculated through convolution. and It is the dimension of the vector. and These represent convolutional networks in the time and space dimensions, respectively. The learnable parameters in this network are represented by... and express; When performing convolution calculations in both the temporal and spatial dimensions, the resulting vectors are input into corresponding attention-based memory modules. Each memory module contains a memory matrix, and the temporal and spatial memory matrices are represented as follows: and Each memory matrix contains There are 1 memory unit, and the memory unit vector is 1. as well as The result obtained from the convolution calculation is input into the memory module. When calculating the time dimension, the convolution result is... As query vector and memory matrix The cosine similarity of each memory unit vector in the equation is calculated as shown in equation (3): Calculated attention weight vector Then use The function maps all elements of the weight vector to a sum of 1, as shown in equation (4): Finally, we obtain the normalized attention weight vector along the time dimension: ; When calculating spatial dimensions, the convolution result is... As query vector and memory matrix The cosine similarity of each memory unit vector in the equation is calculated as shown in equation (5): Calculated attention weight vector Then use The function maps all elements of the weight vector to a sum of 1, as shown in equation (6): The final result is the attention weight vector in the normalized spatial dimension: ; The obtained attention weights are used to perform linear weighting operations on the memory units in the temporal and spatial memory matrices, as shown in equations (7) and (8): in, This is the final output of the memory module in the time dimension; This is the final output of the memory module in the spatial dimension; To ensure that the encoder can learn the most essential features about the fault mode while maintaining its generalization ability to different fault types, a skip connection is introduced between the output of the convolution operation and the memory module to bring common fault information to the model; this process is shown in Equation (9) in the time dimension: In terms of spatial dimensions, it is as shown in equation (10): Finally, the temporal and spatial results are added together and fused to obtain the final output of the encoder, as shown in equation (11): in, and First, the temporal and spatial computation results are converted into two matrices of the same shape by convolution, which are then directly added together to obtain the final output of the encoder. ; The decoder consists of multiple deconvolutional layers, and its overall calculation process is shown in equation (12): in, It is the original input signal restored by the decoder. This represents the trainable parameters of the decoder; The entire diagnostic knowledge learning model is trained and updated using the MSE loss function. The loss on each edge client is shown in equation (13): Each edge client trains its local diagnostic knowledge learning model using the process described above. Once the local diagnostic knowledge learning model converges, each edge client uses the encoder output as the learned diagnostic knowledge.

5. The wind turbine edge-cloud collaborative fault diagnosis system based on knowledge federation according to claim 1, characterized in that: In step 4, the edge-cloud collaborative training is a two-stage, single-round edge-cloud collaborative training mechanism. The information interaction between each edge client and the cloud server only requires one upload and one download operation.