Cross-company non-independent identically distributed fault diagnosis method based on heterogeneous federated large model
By adopting the heterogeneous federal big model in industrial fault diagnosis, the problem of heterogeneous non-independent and homogeneous fault diagnosis in cross-company scenarios is solved, fault diagnosis on low-computer terminal devices is realized, and privacy protection is ensured.
Patent Information
- Application Number
- CN202510180616.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art is difficult to achieve heterogeneous non-independent homodistribution fault diagnosis in cross-company scenarios, and large-model training and data transmission pose major challenges to privacy protection.
Using a method based on heterogeneous federal big model, a server model and a client model are built. Through bidirectional heterogeneous knowledge distillation and differential federated learning architecture, the upstream learning of cross-company federal big models is carried out, the optimal performance server model is determined, and the terminal private small model is fine-tuned through forward heterogeneous knowledge distillation to achieve fault diagnosis on low-computing terminal devices.
It realizes non-independent and homogeneous fault diagnosis across companies, ensuring that the model and data are retained in the private domain of their respective companies, and has the advantages of strong heterogeneity, strict privacy protection, low learning costs, and light deployment.
Smart Images

Figure CN120180118A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of industrial fault diagnosis, and in particular, to a cross-company non-independent and identically distributed fault diagnosis method based on a heterogeneous federated large model. Background Art
[0002] Industrial fault diagnosis is crucial for ensuring the reliability, safety, and efficiency of the manufacturing process. Timely and accurate fault diagnosis can not only improve production efficiency and operation and maintenance flexibility but also reduce expensive downtime. However, due to challenges such as diverse equipment attributes, significant environmental impacts, differences in monitoring sensors, and diverse working conditions, fault diagnosis tasks in real industrial scenarios, including data and targets, are often not independent and identically distributed. This requires fault diagnosis methods to have strong knowledge extraction and generalization capabilities. In addition, with the growth of model training costs and the importance of historical data, both models and data have become the core private assets of companies, subject to strict privacy protection and prohibited from being transmitted outside the company's private domain. This privacy protection phenomenon has become particularly significant with the rapid development of high-cost large models, further exacerbating the difficulty of cross-company fault diagnosis.
[0003] Federated learning methods and large model methods are considered potential solutions to the above problems. The former achieves cross-domain learning by deploying local models on each client to train with local data and uploading model parameters, but it highly depends on frequent cross-domain model transmissions, which are extremely likely to lead to the leakage of model parameters and damage the interests of the model-owning companies. In contrast, the latter achieves cross-domain learning by uploading the data of each client while keeping the model on the server, but it highly depends on a large amount of training on high-performance server clusters, especially the need for a large number of labeled samples. Therefore, the large amount of data transmitted from the data-providing companies to the large model companies poses a major challenge to privacy protection.
[0004] The federated large model method is to introduce large model technology within the federated learning framework to simultaneously address the challenges of insufficient data resources and meet the needs of privacy protection. However, this technology is limited by the research depth and complexity and is still in the early research stage, especially with almost no relevant research in the field of fault diagnosis.
[0005] Therefore, in the related art, there is an urgent need for a method that can achieve cross-company heterogeneous non-independent and identically distributed fault diagnosis. Summary of the Invention
[0006] Based on this, in view of the above technical problems, it is necessary to provide a cross-company non-independent and identically distributed fault diagnosis method based on a heterogeneous federated large model.
[0007] In a first aspect, this application provides a cross-company non-independent and identically distributed fault diagnosis method based on a heterogeneous federated large model. The method includes:
[0008] Build a large server model and a small client model;
[0009] Perform two-way heterogeneous knowledge distillation based on the large server model and the small client model;
[0010] Build a differential federated learning architecture across companies under privacy protection constraints, perform upstream learning of a cross-company federated large model, and determine the optimal performance large server model;
[0011] Based on the optimal performance large server model, determine the corresponding small client model, perform forward heterogeneous knowledge distillation to determine the terminal private small model, fine-tune the small client model, perform lightweight personalized deployment for downstream companies, and achieve fault diagnosis on low-computing power terminal devices.
[0012] Optionally, in an embodiment of the present application, the building of the large server model and the small client model includes:
[0013] Determine the small client model using an adaptive scaling technique based on the complexity of the large server model and the client task.
[0014] Optionally, in an embodiment of the present application, the complexity of the client task is evaluated using attention entropy and dynamic time warping distance.
[0015] Optionally, in an embodiment of the present application, the two-way heterogeneous knowledge distillation based on the large server model and the small client model includes forward heterogeneous knowledge distillation from the large server model to the small client model and reverse heterogeneous knowledge distillation from the small client model to the large server model.
[0016] Optionally, in an embodiment of the present application, the structure of the forward heterogeneous knowledge distillation is expressed as:
[0017]
[0018] where L BiHKD-F is the forward heterogeneous knowledge distillation, L KD is an improved knowledge distillation loss, p s and p t are the predicted values of the small client model and the large server model on the sample x respectively, the subscripts c and are the predicted class and the true class of x respectively, and the parameter γ≥1 is a modulation parameter used to enhance the true class information.
[0019] Optionally, in an embodiment of the present application, the structure of the reverse heterogeneous knowledge distillation is expressed as:
[0020]
[0021] Among them, L BiHKD-R is reverse heterogeneous knowledge distillation, and μ is a hyperparameter for controlling the bias constraint, and are the parameter matrices of the server large model after and before local update by the m-th client, respectively.
[0022] Optionally, in an embodiment of the present application, the method further includes:
[0023] Performing secondary encryption on the server large model and the client small model using homomorphic encryption technology.
[0024] In a second aspect, the present application also provides a cross-company non-independent and identically distributed fault diagnosis device based on a heterogeneous federated large model. The device includes:
[0025] A model building module for building a server large model and a client small model;
[0026] A bidirectional heterogeneous knowledge distillation module for performing bidirectional heterogeneous knowledge distillation based on the server large model and the client small model;
[0027] An upstream learning module for building a differential federated learning architecture across companies under privacy protection constraints, performing upstream learning of a cross-company federated large model, and determining an optimal performance server large model;
[0028] A downstream deployment fault diagnosis module for determining a corresponding client small model based on the optimal performance server large model, performing forward heterogeneous knowledge distillation to determine a terminal private small model, fine-tuning the client small model, and performing lightweight personalized deployment for downstream companies to achieve fault diagnosis on low-computing power terminal devices.
[0029] In a third aspect, the present application also provides a computer device. The computer device includes a memory and a processor, the memory stores a computer program, and the processor executes the steps of the method in each of the above embodiments.
[0030] In a fourth aspect, the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method in each of the above embodiments are implemented.
[0031] The above cross-company non-independent and identically distributed fault diagnosis method based on heterogeneous federated large models first constructs a server large model and client small models; then, performs bidirectional heterogeneous knowledge distillation based on the server large model and client small models; then, constructs a differential federated learning architecture across companies under privacy protection constraints to perform upstream learning of the cross-company federated large model and determine the optimal performance server large model; finally, determines the corresponding client small models based on the optimal performance server large model, performs forward heterogeneous knowledge distillation to determine the terminal private small models, fine-tunes the client small models, and performs lightweight personalized deployment for downstream companies to achieve fault diagnosis on low-computing power terminal devices. Aiming at the problem that the large model and data, as valuable assets of each company, are subject to strict privacy protection, which poses a major challenge to the training and application of large models, a two-stage heterogeneous federated large model architecture with upstream and downstream learning and deployment structures is proposed. This architecture supports cross-company heterogeneous non-independent and identically distributed fault diagnosis while ensuring that the models and data remain in the private domains of their respective companies. The proposed architecture has the advantages of strong heterogeneity, strict privacy protection, low learning cost, and lightweight deployment, and can meet the privacy protection requirements in large model training and application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 FIG. is a schematic diagram of an application scenario considering strict privacy protection constraints in an embodiment;
[0033] Figure 2 FIG. is a schematic flowchart of a cross-company non-independent and identically distributed fault diagnosis method based on heterogeneous federated large models in an embodiment;
[0034] Figure 3 FIG. is a schematic diagram of forward heterogeneous knowledge distillation from a server large model to a client small model in an embodiment;
[0035] Figure 4 FIG. is a schematic diagram of a differential federated learning strategy in an embodiment;
[0036] Figure 5 FIG. is a federated large model learning and deployment architecture with two stages in an embodiment;
[0037] Figure 6 FIG. is a structural block diagram of a cross-company non-independent and identically distributed fault diagnosis device based on heterogeneous federated large models in an embodiment;
[0038] Figure 7 FIG. is an internal structural diagram of a computer device in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] To make the objectives, technical solutions, and advantages of this application clearer, the following further elaborates on this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely used to explain this application and not to limit this application.
[0040] The cross-company non-independent and identically distributed fault diagnosis method based on heterogeneous federated large models provided by the embodiments of this application can be applied to a cross-company non-independent and identically distributed fault diagnosis scenario considering strict privacy protection constraints as Figure 1 shown.
[0041] In one embodiment, as Figure 2 shown, a cross-company non-independent and identically distributed fault diagnosis method based on heterogeneous federated large models is provided, including the following steps:
[0042] S201: Build a server large model and a client small model.
[0043] In the embodiments of this application, first, build a server large model LM and a client small model SM m . Among them, there are no requirements or specific restrictions on the specific type and structure of the server large model. For the client small model, due to the application scenario and computing power limitations of the client company, a large model is often not required. Therefore, a client small model adapted to the client is determined based on the large model structure.
[0044] Specifically, in one embodiment of this application, the building of the server large model and the client small model includes:
[0045] Determine the client small model using an adaptive scaling technique based on the complexity of the server large model and the client task.
[0046] In one embodiment of this application, the complexity of the client task is evaluated using attention entropy and dynamic time warping distance.
[0047] In one embodiment of this application, design an adaptive scaling technique to perform adaptive scaling based on the complexity of the client task based on the large model structure to determine the structure of the client small model. Specifically, for M upstream clients, each client Client m∈M has a private dataset Data m∈M and a private task Task m∈M . To evaluate the complexity of Client m∈M , two non-parametric metrics, attention entropy and dynamic time warping distance, are introduced. The former calculates the entropy value of the signal based on the attention mechanism, and the latter uses the dynamic time warping strategy to perform non-linear alignment of the time series in the time domain to calculate the similarity of the signal. Based on these two metrics, the scaling factor sf of SM m m Designed to:
[0048]
[0049] Among them, softmax(·) is the activation function, and T is used to smooth each SM m The hyperparameters of the scaling factor, functions DTW(·) and AttEn(·) are used to calculate the dynamic time warping distance and attention entropy respectively. In addition, each scaling factor sf m Must be normalized to the range [lim min ,lim max ]∈(0,1].
[0050] By using adaptive scaling techniques, SM m However, the tasks and data of different clients are heterogeneous and non-IID, which requires the SM of different clients to m have different input and output sizes. Therefore, SM m The overall structure is designed as:
[0051]
[0052] Among them, X m and Y m SM m The input and output of the function For the structural scaling algorithm, PreNet and PostNet are task-customized input and output modules, respectively. Specifically, PreNet consists of multiple convolutional layers, while PostNet consists of multi-layer perceptrons.
[0053] S203: Perform bidirectional heterogeneous knowledge distillation based on the server large model and the client small model.
[0054] In the embodiment of the present application, bidirectional heterogeneous knowledge distillation is performed based on the constructed server large model and client small model to solve the problem of model and data heterogeneity.
[0055] Specifically, in one embodiment of the present application, the bidirectional heterogeneous knowledge distillation based on the server large model and the client small model includes forward heterogeneous knowledge distillation from the server large model to the client small model and reverse heterogeneous knowledge distillation from the client small model to the server large model.
[0056] In one embodiment of the present application, due to privacy protection restrictions and the different model sizes of the server and the client, the teacher large model LM in the design node -T and Student Model SM -SBidirectional Heterogeneous Knowledge Distillation (BiHKD) architecture between. As Figure 3 shown, the forward heterogeneous knowledge distillation (BiHKD-F) is the knowledge distillation from the larger model LM -T to the smaller model SM -S .
[0057] In one embodiment of the present application, the structure of the forward heterogeneous knowledge distillation is represented as:
[0058]
[0059] where L BiHKD-F is the forward heterogeneous knowledge distillation, L KD is an improved knowledge distillation loss, p s and p t are the predicted values of the client small model and the server large model on the sample x respectively, the subscripts c and are the predicted class and the true class of x respectively, and the parameter γ≥1 is a modulation parameter used to enhance the true class information.
[0060] Through BiHKD-F, the knowledge of LM -T is propagated to SM -S for local learning for client-specific data and tasks. After client learning, the acquired knowledge needs to be transmitted back from SM -S to LM -T through the reverse heterogeneous knowledge distillation (BiHKD-R) to facilitate the cross-client federated aggregation of the large model in subsequent steps.
[0061] It should be noted that BiHKD-R has an approximate architecture to BiHKD-F, except that the direction of the knowledge flow is different. In addition, since the client's data and labels are only a part of the overall dataset, there will inevitably be knowledge bias in the learned SM -C . To solve this problem, a proximal term constraint is introduced in BiHKD-R to ensure that the local update does not deviate too much from the initial LM -T , thereby improving the tolerance of the heterogeneous federated large model architecture to heterogeneous client models and non-independent and identically distributed tasks.
[0062] In one embodiment of the present application, the structure of the reverse heterogeneous knowledge distillation is represented as:
[0063]
[0064] where L BiHKD-R is the reverse heterogeneous knowledge distillation, μ is a hyperparameter controlling the bias constraint, and They are the parameter matrices of the server large model before local update and local update through the m-th client respectively.
[0065] S205: Build a cross-company differential federated learning architecture under privacy protection constraints, perform upstream learning of the cross-company federated large model, and determine the optimal performance server large model.
[0066] In the embodiments of the present application, although the knowledge interpropagation between heterogeneous models is promoted through bidirectional heterogeneous knowledge distillation, the probability distributions of tasks including data and categories for each client are still non-independent and identically distributed. Therefore, as Figure 4 shown, a cross-company differential federated learning (Di-FL) architecture under privacy protection constraints is built based on the federated learning strategy. This architecture has a contribution-aware aggregation strategy and a logical value calibration technique on the global side and the local side respectively.
[0067] On the global side, the model LM learned from each client -T is constrained by the task-independent attributes and can be considered as a component of LM -G , resulting in the phenomenon of incomplete global information update. In addition, the models LM from different clients -T also show different update preferences due to the influence of their respective task difficulties. To solve this problem, the global loss function is designed as:
[0068]
[0069] where F(w) is the global loss function of LM with parameter matrix w -G , is the local loss function of LM from client m with parameter matrix w m . -T
[0070] Considering the information contribution awareness of each client, the global update can be written as:
[0071]
[0072] where the softmax(·) function is used to amplify the gradient of the LM -T with relatively poor local update performance to accelerate the convergence of this client.
[0073] In addition to the model, the tasks of the client including data and labels can also be considered as part of the global information, which may lead to bias in the update of the global model. Considering the class distribution skew between clients, by adding paired label intervals, the fine-grained calibrated cross-entropy loss is applied to the local update, and then the local update f m (w m ) of each client can be expressed as:
[0074]
[0075] Among them, τ is a hyperparameter, u y and u z respectively represent the sample sizes of the y-th and z-th classes.
[0076] Through the designed f m (w m ), SM -S / -C can simultaneously minimize the classification error and force the learning to focus on the margins of the minority classes to achieve optimal results.
[0077] In addition, in an embodiment of the present application, the method further includes:
[0078] Performing secondary encryption on the server large model and the client small model using homomorphic encryption technology.
[0079] In an embodiment of the present application, although bidirectional heterogeneous knowledge distillation can ensure that the model and data remain within the private domains of their respective companies and can be regarded as the first layer of encryption, knowledge may still be leaked during the cross-domain propagation process. Therefore, homomorphic encryption technology is used to perform ciphertext calculation on the nodes to achieve secondary encryption. The secondary encryption can be regarded as a plug-and-play module. A simple RSA algorithm can be used, which is expressed as follows:
[0080] Ψ * = RSA(Ψ)
[0081] Ψ = RSA -1 (Ψ * )
[0082] Among them, Ψ * is the ciphertext of knowledge Ψ, and RSA and RSA -1 are the encryption and decryption functions respectively.
[0083] The process of upstream learning of the cross-company federated large model includes:
[0084] 1) Initialization: Design the large model structure and randomly generate its initialization parameter matrix, and allocate the large model LM -T to all nodes. Calculate the corresponding model scaling factor sf m according to the complexity of each client task, generate the client small model SM -C and initialize its parameter matrix.
[0085] 2) Due to communication limitations and the intermittency of client online, newly online clients and randomly selected clients participate in the communication of this round.
[0086] 3) After local update and local-side correction without sharing privacy data, aggregate the updated weights into the client model SM -C in it.
[0087] 4) Through homomorphic encryption SM -C →SM -S then upload SM -S to the node and update LM through BiHKD-R -T .
[0088] 5) Check whether LM has been established for all clients -T . If so, execute the subsequent process; otherwise, return to execute process 2).
[0089] 6) Perform global aggregation and global-side correction of Di-FL on the homomorphically encrypted .
[0090] 7) Distribute the updated LM -G to each node and update each SM using BiHKD-F -S . After decrypting SM -S →SM -C , return to execute process 2).
[0091] Save the server large model with the best performance in the upstream learning stage to prepare for the downstream deployment stage.
[0092] S207: Determine the corresponding client small model based on the best performance server large model, perform forward heterogeneous knowledge distillation to determine the terminal private small model, fine-tune the client small model, and perform lightweight personalization deployment for downstream companies to achieve fault diagnosis on low-computing power terminal devices.
[0093] In the embodiments of the present application, based on the best performance server large model, determine the client small model structure through the adaptive scaling technology according to the terminal task difficulty, then perform forward heterogeneous knowledge distillation to obtain the terminal private small model, and use the small private dataset of the terminal to fine-tune the client small model. Finally, realize the lightweight personalization deployment of the large model to the low-computing power terminal devices of downstream companies to achieve fault diagnosis on low-computing power terminal devices. As Figure 5 shown, the structure of the small terminal model SM -C in the downstream deployment stage is first extracted from LM -G based on the characteristics of the specific terminal task through the adaptive scaling technology. Subsequently, the knowledge learned upstream is propagated from LM -T to SM -S through BiHKD-F. After that, through cross-company communication and model decryption SM -S →SM -CObtain the terminal private small model SM -C Finally, use the small private dataset of the terminal to fine-tune SM -C so as to achieve lightweight deployment and personalized customization of the large model on the low-computing-power terminals of downstream companies.
[0094] In an embodiment of the present application, the present invention is respectively applied to industrial fault diagnosis scenarios with different sample numbers and different sample lengths. The experimental results are shown in Tables 1 and 2. The devices to be diagnosed include: intermediate bearings of aero-engine shafts, synchronous motors, and planetary gearboxes of wind turbine units. The overall fault diagnosis accuracy of the method proposed by the present invention for three non-independent and identically distributed downstream tasks with different sample numbers is 93.66%, and the overall fault diagnosis accuracy for different sample lengths is 98.14%. These experimental results prove the effectiveness of the present invention, indicating that the present invention can be well applied to cross-company heterogeneous non-independent and identically distributed fault diagnosis under privacy protection constraints.
[0095] Table 1 Experimental results for different sample numbers (%)
[0096]
[0097] Table 2 Experimental results for different sample lengths (%)
[0098]
[0099] In the above cross-company non-independent and identically distributed fault diagnosis method based on a heterogeneous federated large model, first, build a server large model and a client small model; then, perform two-way heterogeneous knowledge distillation based on the server large model and the client small model; then, build a cross-company differential federated learning architecture under privacy protection constraints to perform upstream learning of the cross-company federated large model and determine the optimal performance server large model; finally, determine the corresponding client small model based on the optimal performance server large model, perform forward heterogeneous knowledge distillation to determine the terminal private small model, fine-tune the client small model, and perform lightweight personalized deployment for downstream companies to achieve fault diagnosis on low-computing-power terminal devices. Aiming at the problems that the large model and data, as valuable assets of each company, are subject to strict privacy protection, which poses major challenges to the training and application of the large model, a two-stage heterogeneous federated large model architecture with an upstream and downstream learning and deployment structure is proposed. This architecture supports cross-company heterogeneous non-independent and identically distributed fault diagnosis, while ensuring that the model and data remain in the private domains of their respective affiliated companies. The proposed architecture has the advantages of strong heterogeneity, strict privacy protection, low learning cost, and lightweight deployment, and can meet the privacy protection requirements in large model training and application scenarios.
[0100] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0101] Based on the same inventive concept, an embodiment of the present application also provides a cross-company non-independent and identically distributed fault diagnosis device based on a heterogeneous federated large model for implementing the above-mentioned cross-company non-independent and identically distributed fault diagnosis method based on a heterogeneous federated large model. The implementation solutions provided by this device to solve problems are similar to the implementation solutions recorded in the above method. Therefore, the specific limitations in one or more embodiments of the cross-company non-independent and identically distributed fault diagnosis device based on a heterogeneous federated large model provided below can refer to the limitations on the cross-company non-independent and identically distributed fault diagnosis method based on a heterogeneous federated large model in the above text, and will not be repeated here.
[0102] In one embodiment, as Figure 6 shown, a cross-company non-independent and identically distributed fault diagnosis device 600 based on a heterogeneous federated large model is provided, including: a model building module 601, a bidirectional heterogeneous knowledge distillation module 603, an upstream learning module 605, and a downstream deployment fault diagnosis module 607, where:
[0103] The model building module 601 is used to build a server large model and a client small model.
[0104] The bidirectional heterogeneous knowledge distillation module 603 is used to perform bidirectional heterogeneous knowledge distillation based on the server large model and the client small model.
[0105] The upstream learning module 605 is used to build a cross-company differential federated learning architecture under privacy protection constraints, perform upstream learning of the cross-company federated large model, and determine the optimal performance server large model.
[0106] The downstream deployment fault diagnosis module 607 is used to determine the corresponding client small model based on the optimal performance server large model, perform forward heterogeneous knowledge distillation to determine the terminal private small model, fine-tune the client small model, perform lightweight personalized deployment of the downstream company, and implement fault diagnosis on low-computing power terminal devices.
[0107] In one embodiment of the present application, the model building module is further configured to:
[0108] Determine the client small model by using an adaptive scaling technique based on the complexity of the server large model and the client task.
[0109] In one embodiment of the present application, the complexity of the client task is evaluated by using attention entropy and dynamic time warping distance.
[0110] In one embodiment of the present application, the bidirectional heterogeneous knowledge distillation based on the server large model and the client small model includes forward heterogeneous knowledge distillation from the server large model to the client small model and reverse heterogeneous knowledge distillation from the client small model to the server large model.
[0111] In one embodiment of the present application, the structure of the forward heterogeneous knowledge distillation is represented as:
[0112]
[0113]
[0114] where L BiHKD-F is the forward heterogeneous knowledge distillation, L KD is an improved knowledge distillation loss, p s and p t are the predicted values of the client small model and the server large model on the sample x respectively, the subscripts c and are the predicted class and the true class of x respectively, and the parameter γ≥1 is a modulation parameter used to enhance the true class information.
[0115] In one embodiment of the present application, the structure of the reverse heterogeneous knowledge distillation is represented as:
[0116]
[0117] where L BiHKD-R is the reverse heterogeneous knowledge distillation, μ is a hyperparameter for controlling the bias constraint, and are the parameter matrices of the server large model after and before local update by the m-th client respectively.
[0118] In one embodiment of the present application, the method further includes:
[0119] Performing secondary encryption on the server large model and the client small model by using homomorphic encryption technology.
[0120] Each module in the above-mentioned cross-company non-i.i.d. fault diagnosis device based on the heterogeneous federated large model can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.
[0121] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 7 shown. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a cross-company non-i.i.d. fault diagnosis method based on the heterogeneous federated large model. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, a touchpad, or a mouse, etc.
[0122] Those skilled in the art can understand that Figure 7 the structure shown in
[0123] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0124] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0125] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0126] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties.
[0127] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., and are not limited thereto.
[0128] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0129] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A cross-company non-independent and identically distributed fault diagnosis method based on a heterogeneous federated large model, characterized in that: The method comprises: Build a large server model and a small client model; Perform bidirectional heterogeneous knowledge distillation based on the server large model and the client small model; Build a cross-company differential federated learning architecture under privacy protection constraints, conduct upstream learning of cross-company federated big models, and determine the best performing server big model; Based on the optimal performance server large model, the corresponding client small model is determined, and forward heterogeneous knowledge distillation is performed to determine the terminal private small model. The client small model is fine-tuned to perform lightweight personalized deployment for downstream companies and realize fault diagnosis on low-computing power terminal devices.
2. The cross-company non-independent and identically distributed fault diagnosis method based on a heterogeneous federated large model according to claim 1 is characterized in that: The construction of the server large model and the client small model includes: Adaptive scaling technology is used to determine the client small model based on the server large model and the complexity of the client task.
3. The cross-company non-independent and identically distributed fault diagnosis method based on a heterogeneous federated large model according to claim 2 is characterized in that: The complexity of the client task is evaluated using attention entropy and dynamic time warping distance.
4. The cross-company non-independent and identically distributed fault diagnosis method based on a heterogeneous federated large model according to claim 1 is characterized in that: The bidirectional heterogeneous knowledge distillation based on the server large model and the client small model includes forward heterogeneous knowledge distillation from the server large model to the client small model and reverse heterogeneous knowledge distillation from the client small model to the server large model.
5. The cross-company non-independent and identically distributed fault diagnosis method based on a heterogeneous federated large model according to claim 4 is characterized in that: The structure of the forward isomerous knowledge distillation is expressed as: Among them, L BiHKD-F is forward heterogeneous knowledge distillation, L KD is an improved knowledge distillation loss, p s and p t are the predicted values of the client small model and the server large model on sample x, respectively. are the predicted category and the true category of x, respectively, and the parameter γ≥1 is the modulation parameter used to enhance the true category information.
6. The cross-company non-independent and identically distributed fault diagnosis method based on a heterogeneous federated large model according to claim 4 is characterized in that: The structure of the reverse heterogeneous knowledge distillation is expressed as: Among them, L BiHKD-R is the reverse heterogeneous knowledge distillation, μ is the hyperparameter controlling the bias constraint, and They are the parameter matrices of the server model after and before local update by the mth client respectively.
7. The cross-company non-independent and identically distributed fault diagnosis method based on a heterogeneous federated large model according to claim 1 is characterized in that: The method further comprises: Homomorphic encryption technology is used to perform secondary encryption of the server large model and the client small model.
8. A cross-company non-independent and identically distributed fault diagnosis device based on a heterogeneous federated large model, characterized in that: The device comprises: Model building module, used to build large server models and small client models; A bidirectional heterogeneous knowledge distillation module, used for performing bidirectional heterogeneous knowledge distillation based on the server large model and the client small model; The upstream learning module is used to build a cross-company differential federated learning architecture under privacy protection constraints, conduct upstream learning of cross-company federated large models, and determine the optimal performance server large model; The downstream deployment fault diagnosis module is used to determine the corresponding client small model based on the optimal performance server large model, and perform forward heterogeneous knowledge distillation to determine the terminal private small model, fine-tune the client small model, perform lightweight personalized deployment for downstream companies, and realize fault diagnosis on low-computing power terminal devices.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.