Power equipment data management model training method based on dynamic federated learning and layered distillation, data management method and related device

Through the power equipment data management model training method of dynamic federated learning and hierarchical distillation, the problems of difficult cross-regional sharing and low management efficiency in power equipment data management are solved, and efficient and secure management of power equipment data is achieved, which is suitable for scenarios with limited resources.

CN120653981APending Publication Date: 2025-09-16TACHENG POWER SUPPLY CO OF STATE GRID XINJIANG ELECTRIC POWER CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510737941.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

The existing power equipment data management has problems such as difficulty in cross-departmental and cross-regional data sharing and collaborative operations, low management efficiency, high risks of information consistency and compliance, and difficulty for inspectors to easily and in real time access data.

Method used

A power equipment data management model training method based on dynamic federated learning and hierarchical distillation is adopted. Through knowledge distillation training of teacher model and student model, combined with dynamic contribution evaluation, a lightweight student model is constructed to realize the automated management of power equipment data.

Benefits of technology

It achieves efficient and secure management of power equipment data, provides accurate real-time support, breaks the "experience island" effect under the traditional model, improves management efficiency and information consistency, and is suitable for scenarios with limited resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120653981A_ABST
    Figure CN120653981A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data management, in particular to a power equipment data management model training method based on dynamic federal learning and hierarchical distillation, a data management method and a related device, comprising the following steps: inputting a training set into an initial student model and a pre-trained teacher model; the teacher model is utilized to perform knowledge distillation training on the student model to obtain a trained student model, the teacher model is obtained by performing federated learning on a plurality of samples, global teacher model parameters updated by dynamic contribution evaluation are introduced in the federated learning, and the trained student model is optimized by utilizing a verification set and a test set, so that the training efficiency is improved. And obtaining a power equipment data management model. According to the method, the power equipment data management model is constructed through cloud federation training and knowledge distillation optimization, the power equipment data are automatically analyzed, the limitation that data search is tedious and information check depends on manpower in the past is changed, and a more reliable guarantee is provided for lean management and efficient application of the power equipment data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data management, and is a data management model training method for electric power equipment based on dynamic federated learning and hierarchical distillation, a data management method and related devices. Background Art

[0002] The safe and stable operation of the power system is highly dependent on the accurate, complete and efficient management of data throughout the life cycle of power equipment. This data includes but is not limited to equipment drawings, technical specifications, secondary safety measures tickets, standardized operation cards and various digital backups (such as telecontrol backups, automation equipment backups, SCD, ICD, CID files, etc.).

[0003] The current management of power equipment data mainly relies on manual review and the establishment of ledgers. However, due to the wide variety of power equipment data, this method has the following problems:

[0004] First, in power equipment data management, the entry, storage, retrieval, and version control of power equipment data are often limited by the physical isolation of internal and external networks and system barriers. This makes cross-departmental and cross-regional data sharing and collaborative work difficult. Accurately matching and associating drawings with actual equipment and related work documents is time-consuming and laborious.

[0005] Secondly, a large amount of power equipment data is managed offline or stored in a decentralized manner, lacking a unified management mechanism. This not only affects management effectiveness but also the efficiency of subsequent work based on power equipment data, and also brings risks to information consistency and compliance.

[0006] Furthermore, the chaotic management of power equipment data makes it difficult for inspectors to conveniently and in real time access required equipment data and drawings via mobile terminals during inspections or maintenance. This is especially true in areas with poor signal coverage, where information acquisition is hindered. This directly affects the accuracy of on-site operations and emergency response capabilities.

[0007] Finally, although some units have completed the labeling and coding of the physical ID of the equipment, convenient information query and related application scenarios based on the physical ID have not been fully developed, and the inspection work still relies heavily on the traditional ledger model. Summary of the Invention

[0008] The present invention provides a power equipment data management model training method, data management method and related devices based on dynamic federated learning and hierarchical distillation, which overcomes the shortcomings of the above-mentioned existing technologies and can effectively solve the problem of low management efficiency caused by the dependence of verification and management processes on manual labor in existing power equipment data management methods.

[0009] One of the technical solutions of the present invention is achieved through the following measures: a method for training a power equipment data management model based on dynamic federated learning and hierarchical distillation, comprising:

[0010] Obtain several samples and divide them into training, validation, and test sets. Each sample includes local power equipment data and corresponding data management and analysis result labels. The data management and analysis results include data matching results, completeness scores, association recommendation confidence, and consistency judgments.

[0011] The training set is input into the initial student model and the pre-trained teacher model. The teacher model is used to perform knowledge distillation training on the student model to obtain the trained student model. The teacher model is obtained by federated learning using several samples, and the global teacher model parameters are updated by dynamic contribution evaluation in federated learning.

[0012] The trained student model is optimized using the validation set and test set to obtain the power equipment data management model.

[0013] The following are further optimizations and / or improvements to the above technical solutions:

[0014] The above pre-trained teacher model construction process includes:

[0015] Establish a global teacher model on the server and randomly initialize the parameters of the global teacher model;

[0016] Distribute the global teacher model and randomly initialized global teacher model parameters to each client participating in federated learning;

[0017] After receiving the randomly initialized global teacher model parameters, each client uses local data to train the global teacher model and uploads the trained local model parameters to the server.

[0018] The server triggers dynamic contribution evaluation to obtain the contribution of each local model parameter, aggregates the local model parameters based on the contribution, and obtains the updated global teacher model parameters;

[0019] The above steps are repeated until the preset stopping condition is met, and the current global teacher model is used as the pre-trained teacher model.

[0020] The above server triggers dynamic contribution evaluation to obtain the contribution of each local model parameter. It aggregates the local model parameters based on the contribution to obtain the updated global teacher model parameters, including:

[0021] On the server side, each local model parameter is verified using a preset validation dataset to obtain the corresponding accuracy, diversity index, and difference norm;

[0022] Using the accuracy, diversity index, and difference norm corresponding to each local model parameter, calculate the contribution of each local model parameter to the update of the global teacher model parameters in the current iteration;

[0023]

[0024] in, is the contribution; is the accuracy rate; Div(D k ) is the diversity index; is the difference norm, is the global teacher model parameter updated in the previous iteration, are the local model parameters obtained in the current iteration; α, β, γ are hyperparameters; f norm (·) is the normalization function;

[0025] Aggregate local model parameters based on contribution to obtain updated global teacher model parameters;

[0026]

[0027] in, is the updated global teacher model parameter; n k 、n j The amount of local data on the client; is the contribution; The local model parameters of the client are obtained for the current iteration; k and j are the number of clients; n k 、n j The amount of local data on the client; is the contribution; is the normalization of contribution.

[0028] The training set is input into the initial student model and the pre-trained teacher model. The distillation loss in the knowledge distillation training of the student model using the teacher model is as follows:

[0029] L KD =L hard (M S (X),Y true )+λ1L soft (M S (X),M T (X))+λ2L HAFM (M S (X),M T (X)

[0030] Among them, L KD is the distillation loss; L HAFM is the hierarchical attention feature matching loss; Lsoft is the soft label loss; L hard is the supervised learning loss.

[0031] The second technical solution of the present invention is achieved through the following measures: a method for managing power equipment data based on dynamic federated learning and hierarchical distillation, comprising:

[0032] Obtain local power equipment data to be managed and analyzed;

[0033] The local power equipment information data to be managed and analyzed is input into the power equipment information management model to obtain the information management analysis results, wherein the power equipment information management model is trained by the power equipment information management model training method based on dynamic federated learning and hierarchical distillation as described in any one of claims 1 to 4.

[0034] The third technical solution of the present invention is achieved through the following measures: a power equipment data management model training device based on dynamic federated learning and hierarchical distillation, comprising:

[0035] A sample acquisition unit acquires a number of samples and divides them into a training set, a validation set, and a test set. Each sample includes local power equipment data and corresponding data management and analysis result labels. The data management and analysis results include data matching results, completeness scores, confidence levels for associated recommendations, and consistency judgments.

[0036] The model training unit inputs the training set into the initial student model and the pre-trained teacher model, and uses the teacher model to perform knowledge distillation training on the student model to obtain the trained student model. The teacher model is obtained by federated learning using several samples.

[0037] The model verification unit uses the verification set and the test set to optimize the trained student model to obtain the power equipment data management model.

[0038] The following are further optimizations and / or improvements to the above technical solutions:

[0039] The above also includes a teacher model training unit, including:

[0040] The server includes:

[0041] Initialization module, establishes a global teacher model and randomly initializes the parameters of the global teacher model;

[0042] The distribution module distributes the global teacher model and randomly initialized global teacher model parameters to each client participating in federated learning;

[0043] Dynamic Contribution Evaluation Module: Obtains the contribution of each local model parameter, aggregates the local model parameters based on the contribution, and obtains updated global teacher model parameters, including:

[0044] On the server side, each local model parameter is verified using a preset validation dataset to obtain the corresponding accuracy, diversity index, and difference norm;

[0045] Using the accuracy, diversity index, and difference norm corresponding to each local model parameter, calculate the contribution of each local model parameter to the update of the global teacher model parameters in the current iteration;

[0046]

[0047] in, is the contribution; is the accuracy rate; Div(D k ) is the diversity index; is the difference norm, is the global teacher model parameter updated in the previous iteration, are the local model parameters obtained in the current iteration; α, β, γ are hyperparameters; f norm (·) is the normalization function;

[0048] Aggregate local model parameters based on contribution to obtain updated global teacher model parameters;

[0049]

[0050] in, is the updated global teacher model parameter; n k 、n j The amount of local data on the client; is the contribution; The local model parameters of the client are obtained for the current iteration; k and j are the number of clients; n k 、n j The amount of local data on the client; is the contribution; is the normalization processing of contribution;

[0051] After receiving the randomly initialized global teacher model parameters, multiple clients use local data to train the global teacher model and upload the trained local model parameters to the server.

[0052] The fourth technical solution of the present invention is achieved through the following measures: a power equipment data management device based on dynamic federated learning and hierarchical distillation, comprising:

[0053] Data acquisition unit, which acquires local power equipment data to be managed and analyzed;

[0054] The data management and analysis unit inputs the local power equipment data to be managed and analyzed into the power equipment data management model to obtain the data management analysis results, wherein the power equipment data management model is trained through the power equipment data management model training method based on dynamic federated learning and hierarchical distillation.

[0055] The fifth technical solution of the present invention is achieved through the following measures: an electronic device, characterized in that it includes a processor and a memory, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the steps in the power equipment data management method based on dynamic federated learning and hierarchical distillation.

[0056] The sixth technical solution of the present invention is achieved through the following measures: a storage medium, characterized in that a computer program that can be read by a computer is stored on the storage medium, and the computer program is configured to execute the steps of the power equipment data management method based on dynamic federated learning and hierarchical distillation at runtime.

[0057] The beneficial effects of the present invention include:

[0058] Through cloud-based federated training (aggregation to generate teacher models) and knowledge distillation optimization (efficient transfer of knowledge to lightweight student models), a power equipment data management model is constructed. The power equipment data management model is used to automatically analyze power equipment data, changing the previous limitations of cumbersome data search and manual reliance on information verification. It enables operating personnel to obtain more accurate real-time support during equipment inspections, data approvals, drawing retrieval, and work card execution, providing more reliable guarantees for the lean management and efficient application of power equipment data.

[0059] Based on the introduction of a federated learning mechanism with dynamic contribution evaluation, each power production unit can safely and efficiently contribute the analysis methods and management experience of power equipment data to the teacher model without leaking local power equipment data. This mechanism dynamically adjusts the weight of each participant in the global teacher model aggregation based on factors such as model performance, data diversity and update effectiveness, ensuring the priority integration of high-quality knowledge, thereby constructing a teacher model for data management and analysis, effectively breaking the "experience island" effect under the traditional model.

[0060] Combined with hierarchical knowledge distillation to generate a lightweight student model, the trained lightweight student model can fully extract knowledge from the teacher model that is beneficial to the learning of the lightweight student model, thereby better improving the lightweight performance and making the lightweight small model have performance comparable to or even better than that of the large model. It is suitable for resource-constrained scenarios and provides more reliable intelligent guarantees for the lean management and efficient application of power equipment data. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Attachment Figure 1 A schematic diagram of an implementation environment provided by the present invention.

[0062] Attachment Figure 2 This is a flow chart of the power equipment data management model training method provided by the present invention.

[0063] Attachment Figure 3 This is a flow chart of the teacher model construction method provided by the present invention.

[0064] Attachment Figure 4 This is a schematic diagram of the dynamic contribution evaluation process provided by the present invention.

[0065] Attachment Figure 5 This is a flow chart of the power equipment data management method provided by the present invention.

[0066] Attachment Figure 6 This is a structural diagram of the power equipment data management model training device provided by the present invention.

[0067] Attachment Figure 7 This is a structural diagram of the power equipment data management device provided by the present invention. DETAILED DESCRIPTION

[0068] The present invention is not limited to the following embodiments, and specific implementation methods can be determined based on the technical solutions of the present invention and actual conditions.

[0069] Those skilled in the art will understand that, unless otherwise stated, in the embodiments of the present invention, a "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0070] In addition, in the embodiments of the present invention, “plurality” refers to two or more than two, and “first” and “second” are used for distinguishing descriptions and should not be understood as implying relative importance.

[0071] Since the existing power equipment data management method that relies on manual management is inefficient, a large amount of power equipment data is easily in offline management or decentralized storage, and there is a lack of a unified management mechanism, an embodiment of the present invention provides a power equipment data management method based on dynamic federated learning and hierarchical distillation.

[0072] An embodiment of the present invention provides a method for managing electric power equipment data based on dynamic federated learning and hierarchical distillation, wherein a training set is input into an initial student model and a pre-trained teacher model, and the teacher model is used to perform knowledge distillation training on the student model to obtain a trained student model, and the trained student model is optimized using a validation set and a test set to obtain an electric power equipment data management model, and local electric power equipment data to be managed and analyzed is input into the electric power equipment data management model to obtain data management analysis results, which include data matching results, integrity scores, associated recommendation confidences, consistency judgments, etc., providing stable and reliable analysis data support for electric power equipment data management.

[0073] Among them, the method provided by the embodiment of the present invention may involve artificial intelligence (AI) technology and can be implemented based on artificial intelligence technology, for example, using deep learning to obtain a corresponding model through sample training.

[0074] Machine Learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI.

[0075] Deep learning (DL) specifically refers to machine learning based on deep neural network models and methods. It is developed based on statistical machine learning, artificial neural networks, and other algorithmic models, combined with the development of modern big data and massive computing power. The most important technical feature of deep learning is its ability to automatically extract features.

[0076] The above-mentioned machine learning and deep learning usually include technologies such as neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.

[0077] The loss function is used during neural network training to ensure that the model network's output is as close as possible to the desired predicted value. This is done by comparing the current network's predicted value with the target value, and then updating the weight vectors of each layer of the neural network based on the difference between the two (this is usually preceded by an initialization process, where parameters are preconfigured for each layer of the neural network), until the neural network can predict the target value or a value very close to it. Therefore, deep learning requires pre-defining how to compare the difference between the predicted and target values, which is the loss function.

[0078] As attached Figure 1 FIG. 1 shows a schematic diagram of an implementation environment provided by an embodiment of the present invention. The implementation environment may include: training equipment and user equipment.

[0079] Both the training device and the usage device are computer devices; optionally, the computer device is a terminal device, such as a mobile phone, tablet computer, PC (Personal Computer) and other electronic devices; or, the computer device is a server, which can be a single server, a server cluster composed of multiple servers, or a cloud computing service center, which is not limited in this embodiment of the present invention.

[0080] A training device refers to a computer device capable of training and learning network models. Optionally, the training device has the ability to acquire network models and train and learn them according to application requirements. For example, the training device acquires network models from other devices via the network and then trains them using training samples according to application requirements, so that the network model is capable of obtaining data management and analysis results. Optionally, the training device has the ability to construct network models and can independently construct network models according to application requirements and then train and learn them. For example, in order to obtain data management and analysis results based on drilling power equipment data, the training device independently constructs a network model and then trains and learns it using samples according to application requirements.

[0081] The using device refers to a computer device that has the need to use the network model. Optionally, the using device obtains the network model from other devices through the network according to application requirements. For example, the using device has the need to obtain data management analysis results. It can obtain the network model that has completed training and learning to predict data management analysis results from other devices through the network, and use the network model to predict data management analysis results.

[0082] Based on this, the technical solution of the present invention will be introduced and explained with reference to several examples below.

[0083] Example 1: As shown in the attached Figure 2As shown, an embodiment of the present invention discloses a method for training a power equipment data management model based on dynamic federated learning and hierarchical distillation, comprising:

[0084] Step S110: obtain several samples and divide them into a training set, a validation set, and a test set. Each sample includes local power equipment data and corresponding data management analysis result labels. The data management analysis results include data matching results, integrity scores, associated recommendation confidence, and consistency judgments.

[0085] In this embodiment, samples can be obtained from the power equipment data and related metadata, flow records, and application data generated by each participating power production unit during daily operation, maintenance, and management, as follows:

[0086] (1) Obtaining data on a number of local power equipment to form corresponding samples, wherein the data on the local power equipment includes:

[0087] Text data, equipment technical manuals, various operation instruction cards, bill text content, and approval opinions;

[0088] Structured or semi-structured data, drawing metadata, parameters and logical relationships in configuration files such as SCD / ICD / CID, timestamps and status information of data approval flow, and equipment ledger parameters.

[0089] (2) Label each sample with label information, including: data matching results (data type identification results, version identification results, content anomaly detection results), integrity score (data integrity score), association recommendation confidence (confidence of the correlation between data, the greater the confidence, the more accurate the correlation between data, the correlation between data can be no correlation or correlation level), consistency judgment (whether the device and data are consistent).

[0090] In step S120, the training set is input into the initial student model and the pre-trained teacher model, and the teacher model is used to perform knowledge distillation training on the student model to obtain a trained student model. The teacher model is obtained by federated learning using several samples, and the global teacher model parameters updated by dynamic contribution evaluation are introduced in the federated learning.

[0091] Step S130: Optimize the trained student model using the validation set and the test set to obtain a power equipment data management model.

[0092] In this embodiment, the validation set and the test set are used to optimize the trained student model, and evaluation indicators such as accuracy and F1 score can be selected as optimization indicators.

[0093] This embodiment introduces a teacher-student model for deep learning and establishes a power equipment data management model. The structure of the student model and the teacher model specifically includes:

[0094] The network structure of the teacher model is selected according to the actual situation. In this embodiment, it can include, but is not limited to, three stages, namely, feature extraction stage, feature fusion stage, and analysis and output stage, as follows:

[0095] In the feature extraction stage, for text data, standard preprocessing (such as word segmentation, IDization, padding / truncation) is first performed on the text data, and then the preprocessed text data is input into the BERT model. The BERT model uses the self-attention mechanism of its multi-layer Transformer encoder to capture the complex contextual dependencies and deep semantic information in the text sequence, and outputs a high-dimensional feature vector representing the entire text sequence; for structured or semi-structured data, the structured or semi-structured data is first preprocessed (such as one-hot encoding and normalization), and then input into a multi-layer perceptron (MLP) composed of several fully connected layers. The high-order combination relationship between features is learned through nonlinear transformation, and the feature vector of the structured data is output.

[0096] In the feature fusion stage, the high-dimensional feature vectors extracted from text data and the feature vectors extracted from structured or semi-structured data are merged. Specifically, the two can be directly connected by concatenation to form a comprehensive feature vector containing multimodal information. Then, one or more multi-layer perceptrons (MLPs) are used to learn the interaction between different modal features, perform feature selection and dimensionality reduction, and output a deeply fused feature vector.

[0097] In the analysis output stage, the fused feature vector is sent to the multi-layer perceptron (MLP) for deeper feature abstraction, extracting the most discriminative features for the data management analysis task. Then, through an output layer (a single-layer fully connected network without nonlinear activation), these features are mapped to a space that matches the task output dimension to generate the data management analysis results.

[0098] The student's network structure is selected according to the actual situation, but it must be a lightweight model. In this embodiment, it can include but is not limited to three stages, namely, feature extraction stage, feature fusion stage, and analysis and output stage, as follows:

[0099] In the feature extraction stage, for text data, standard preprocessing (such as word segmentation, IDization, padding / truncation) is first performed on the text data, and then it is input into the LSTM model. The word ID sequence in the LSTM model is converted into a word vector sequence through a word embedding layer (which can load lightweight pre-trained word vectors or randomly initialized). The word vector sequence is processed by a 1-2 layer (optionally bidirectional) LSTM network, which captures the contextual dependencies of the text through its gating mechanism and outputs a feature vector representing the semantics of the text (usually the hidden state of the last time step or the sequence pooling result). For structured or semi-structured data, an MLP with a shallower and narrower branch than the corresponding branch of the teacher model is used for feature extraction to obtain the feature vector of the structured data.

[0100] In the feature fusion stage, the feature vector representing the text semantics and the feature vector of the structured data are spliced ​​together. The spliced ​​comprehensive feature vector is only passed through a very shallow and narrow multi-layer perceptron (MLP) for feature interaction and dimensionality reduction to generate a fused feature vector.

[0101] In the analysis output stage, the fused feature vector is sent to a shallow and narrow multi-layer perceptron (MLP) for deeper feature extraction, extracting the most discriminative features for the data management analysis task. Then, through an output layer (a single-layer fully connected network without nonlinear activation), these features are mapped to a space that matches the task output dimension to generate the data management analysis results.

[0102] In this embodiment, the lightweight student model is trained using a pre-trained teacher model and combined with knowledge distillation, so that the trained lightweight student model can fully extract knowledge from the teacher model that is beneficial to the learning of the lightweight student model, thereby better improving the performance of the lightweight model, so that the lightweight small model has performance comparable to or even better than that of the large model, and is suitable for resource-constrained scenarios. Furthermore, in this embodiment, the teacher model is obtained by federated learning using a number of samples. In this learning method, user data does not leave the local device, only model parameters are shared, which helps to protect data privacy. In addition, only model parameters are transmitted, which reduces a large amount of data transmission and saves bandwidth and energy. At the same time, since the model is trained locally, the data of each device may come from a different distribution, which helps to improve the generalization ability of the model. In addition, the global teacher model parameters updated by dynamic contribution evaluation are introduced in federated learning, making parameter updates more accurate.

[0103] The present invention discloses a method for training a power equipment data management model based on dynamic federated learning and hierarchical distillation, which automates the power equipment data management and changes the previous limitations of cumbersome data search and manual reliance on information verification. It enables operators to obtain more accurate real-time support during equipment inspections, data approvals, drawing retrievals, and job card executions, providing more reliable guarantees for the lean management and efficient application of power equipment data. In addition, based on the introduction of a federated learning method with dynamic contribution evaluation, a teacher model is constructed without leaking local original power equipment data, and a lightweight student model is further generated based on knowledge distillation. The student model has performance comparable to or even better than that of the teacher model, greatly reducing the demand for computing resources while ensuring prediction accuracy, thereby expanding the applicability of the model.

[0104] Example 2: As shown in the attached Figure 3 As shown, the embodiment of the present invention is a further optimization of the above embodiment, wherein the pre-trained teacher model construction process includes:

[0105] Step S210: Establish a global teacher model in the server and randomly initialize the parameters of the global teacher model.

[0106] It should be noted that the network structure of the global teacher model here is selected according to actual conditions. The structure in this embodiment can be as shown in the description of Example 1.

[0107] The randomly initialized global teacher model parameters mentioned above include weights and biases, etc. Random initialization usually involves randomly assigning parameter values ​​within a certain range or according to a certain distribution.

[0108] Step S220: Distribute the global teacher model and the randomly initialized global teacher model parameters to each client participating in federated learning.

[0109] In step S230, after receiving the randomly initialized global teacher model parameters, each client uses local data to train the global teacher model and uploads the trained local model parameters to the server.

[0110] Each client has local data, which is a unique data set of power equipment information in the local area. k Each power equipment data is labeled with the corresponding data management analysis result label. In each client, the global teacher model is trained using local data, that is, the general deep learning steps are used to train the global teacher model parameters after training.

[0111] In step S240, the server triggers dynamic contribution evaluation to obtain the contribution of each local model parameter, aggregates the local model parameters based on the contribution, and obtains updated global teacher model parameters.

[0112] The above steps, as shown in the attached Figure 4 As shown, specifically including:

[0113] Step S241: Verify each local model parameter on the server side using a preset verification data set to obtain the corresponding accuracy, diversity index, and difference norm;

[0114] Here, the accuracy is the ratio of correctly predicted samples to the total number of samples; the diversity index is the data type and application scenario diversity index contained in the local power equipment data, which is used to measure the ability of client data to cover diversified equipment data analysis and complex application scenarios for the global model; the difference norm is used to determine the effectiveness of local model parameter updates and to filter out invalid or abnormal amplitude updates.

[0115] Step S242, using the accuracy, diversity index, and difference norm corresponding to each local model parameter, calculate the contribution of each local model parameter to the update of the global teacher model parameters in the current iteration;

[0116]

[0117] in, For contribution; is the accuracy rate; Div(D k ) is the diversity index; is the difference norm, is the global teacher model parameter updated in the previous iteration, are the local model parameters obtained in the current iteration; α, β, γ are hyperparameters, i.e., the weights of accuracy, diversity index, and difference norm; f norm (·) is a normalization function to ensure that the contribution value is within a reasonable range;

[0118] Step S243, combining the contribution to aggregate the local model parameters to obtain updated global teacher model parameters;

[0119]

[0120] in, is the updated global teacher model parameter; n k 、n j The amount of local data on the client; For contribution; The local model parameters of the client are obtained for the current iteration; k and j are the number of clients; n k 、n j The amount of local data on the client; For contribution; is the normalization of contribution.

[0121] Step S250, loop the above steps until the preset stop condition is met, and use the current global teacher model as the pre-trained teacher model.

[0122] The above stopping conditions may include the maximum number of cycles and the loss limit. If one of them is met, learning can be stopped. The loss function used for loss value calculation adopts the local loss function L local,k Typically, it is the standard supervised learning loss L corresponding to the specific analysis task. task (M local,k (X k ),Y true,k ).

[0123] In the embodiment of the present invention, by introducing dynamic contribution evaluation, it is ensured that the global teacher model can focus more on absorbing knowledge from clients with high data quality, wide data coverage scenarios, and effective model updates, thereby effectively suppressing the negative impact of low-quality data or model updates on the global model, accelerating the convergence speed of federated learning, and ultimately learning a teacher model with better robustness and accuracy for use in device data management and analysis.

[0124] Example 3: This embodiment of the present invention is a further optimization of the above embodiment, in which the training set is input into the initial student model and the pre-trained teacher model, and the teacher model is used to perform knowledge distillation training on the student model to achieve efficient knowledge transfer and accurate preservation from the complex teacher model to the lightweight student model. During the training process, this embodiment performs knowledge distillation at the feature level, output level, and label level as the total loss of the student model training, and the distillation loss is shown as follows:

[0125] L KD =L hard (M S (X),Y true )+λ1L soft (M S (X),M T (X))+λ2L HAFM (M S (X),M T (X)

[0126] Where X is the input local power equipment data; M S For student models; M T is the teacher model; Y true is the true label; λ1,λ2 are the hyperparameters for balancing the losses; L KD is the distillation loss; L HAFM is the hierarchical attention feature matching loss; L soft is the soft label loss; L hardis the supervised learning loss.

[0127] The above supervised learning loss L hard Used to measure the difference between the student model prediction and the true label, which is determined by the type of data management analysis result, for example:

[0128] Classification tasks (such as data category judgment) are determined using cross entropy loss, as follows:

[0129]

[0130] Regression tasks (such as completeness scoring) are determined using mean squared error loss, as follows:

[0131]

[0132] The above soft label loss L soft To encourage the student model to imitate the teacher model M T The output probability distribution of is determined using the Kullback-Leibler (KL) divergence with a temperature coefficient T, as follows:

[0133] L soft =KL(P S ||P T )

[0134] Among them, P S =softmax(Z S / T), P T =softmax(Z T / T) are the softened output probabilities of the student model and the teacher model respectively.

[0135] The most critical improvement in the above loss is reflected in the layered attention feature matching loss L HAFM In this embodiment, the teacher model M is selected T The L intermediate feature layers that contribute the most to the intelligent analysis tasks of device data (such as data consistency verification, relevance recommendation, abnormal pattern recognition, etc.) have their activation outputs represented by F T,1 ,F T,2 ,...,F T,L , correspondingly, the student model M S There is also a corresponding intermediate feature layer, whose activation output is F S,1 ,F S,2 ,...,F S,L For each pair of selected feature layers (F T,l ,F S,l ), first through an attention module Att l (·) Calculate the teacher model feature F T,1Importance weight of each internal element A T,l =Att l (F T,l ), the attention weight A T,l It reflects which local features the teacher model pays more attention to when processing device data information; then the student model is guided to learn to imitate the teacher model features weighted by this attention. HAFM To guide the student model to learn the key representations of the teacher model in L intermediate feature layers, the core is the feature matching loss L with attention feat,l , the hierarchical attention feature matching loss is as follows:

[0136]

[0137] Among them, F T,l , F S,l are the feature activations of the teacher model and the student model at layer l, A T,l =Att l (F T,l ) is the attention module Att l (·) calculates the teacher model feature attention weight, ⊙ is the element-wise product, δ l The weight coefficient for controlling the importance of feature matching in the first layer.

[0138] Through this hierarchical, attention-based feature matching, the student model not only learns the "thinking results" (final output) of the teacher model, but more importantly, learns its "thinking process" (how to extract and focus on key analysis features and internal correlations from the input power equipment data). This ensures that even if the student model structure is greatly simplified, it can maximize the retention of the complex intelligent analysis logic and key feature extraction capabilities learned by the teacher model from massive power equipment data, effectively breaking through the "information bottleneck" or loss of key details that may exist in traditional knowledge distillation methods.

[0139] The student model training method disclosed in this embodiment, the student model M S In the effective inheritance teacher model M T While strengthening the equipment data analysis capabilities, the complexity of its own model is significantly reduced. Specifically, it is reflected in:

[0140] Model parameters Params(M S ) is much smaller than the number of parameters of the teacher model Params(M T ), namely Params(M S )< <Params(M T ), which directly reduces the space required for model storage;

[0141] The number of floating point operations FLOPs (MS) during inference is also much lower than the teacher model FLOPs (M T ), namely FLOPs(M S )< <FLOPs(M T ), ensuring fast response on devices with limited computing power. For example, the number of parameters in the student model can be compressed to 1 / N of the teacher model p times (where N p >1), the inference speed can be increased to N s times (where N s >1), while the decline in key device data intelligent analysis indicators (such as data integrity assessment accuracy and associated recommendation F1 score) is kept within a very small range. This balance of performance and efficiency enables the lightweight student model to be successfully deployed on mobile handheld devices, edge computing nodes, or specific embedded systems where computing power, storage space, and energy consumption are strictly limited.

[0142] Example 4: As shown in the attached Figure 5 As shown, an embodiment of the present invention discloses a method for managing power equipment data based on dynamic federated learning and hierarchical distillation, including:

[0143] Step S310, obtaining local power equipment data to be managed and analyzed;

[0144] In step S310 , the local power equipment data to be managed and analyzed is input into the power equipment data management model to obtain data management analysis results, wherein the power equipment data management model is trained using a power equipment data management model training method based on dynamic federated learning and hierarchical distillation.

[0145] In this embodiment, the power equipment data management model is trained by a power equipment data management model training method based on dynamic federated learning and hierarchical distillation. The specific method is as described in Examples 1 to 3 and will not be repeated here.

[0146] Example 5: As shown in the attached Figure 6 As shown, an embodiment of the present invention discloses a power equipment data management model training device based on dynamic federated learning and hierarchical distillation, comprising:

[0147] A sample acquisition unit acquires a number of samples and divides them into a training set, a validation set, and a test set. Each sample includes local power equipment data and corresponding data management and analysis result labels. The data management and analysis results include data matching results, completeness scores, confidence levels for associated recommendations, and consistency judgments.

[0148] The model training unit inputs the training set into the initial student model and the pre-trained teacher model, and uses the teacher model to perform knowledge distillation training on the student model to obtain the trained student model. The teacher model is obtained by federated learning using several samples.

[0149] The model verification unit uses the verification set and the test set to optimize the trained student model to obtain the power equipment data management model.

[0150] Furthermore, it also includes a teacher model training unit, including:

[0151] The server includes:

[0152] Initialization module, establishes a global teacher model and randomly initializes the parameters of the global teacher model;

[0153] The distribution module distributes the global teacher model and randomly initialized global teacher model parameters to each client participating in federated learning;

[0154] Dynamic Contribution Evaluation Module: Obtains the contribution of each local model parameter, aggregates the local model parameters based on the contribution, and obtains updated global teacher model parameters, including:

[0155] On the server side, each local model parameter is verified using a preset validation dataset to obtain the corresponding accuracy, diversity index, and difference norm;

[0156] Using the accuracy, diversity index, and difference norm corresponding to each local model parameter, calculate the contribution of each local model parameter to the update of the global teacher model parameters in the current iteration;

[0157]

[0158] in, is the contribution; is the accuracy rate; Div(D k ) is the diversity index; is the difference norm, is the global teacher model parameter updated in the previous iteration, are the local model parameters obtained in the current iteration; α, β, γ are hyperparameters; f norm (·) is the normalization function;

[0159] Aggregate local model parameters based on contribution to obtain updated global teacher model parameters;

[0160]

[0161] in, is the updated global teacher model parameter; nk 、n j The amount of local data on the client; For contribution; The local model parameters of the client are obtained for the current iteration; k and j are the number of clients; n k 、n j The amount of local data on the client; For contribution; is the normalization processing of contribution;

[0162] After receiving the randomly initialized global teacher model parameters, multiple clients use local data to train the global teacher model and upload the trained local model parameters to the server.

[0163] In this embodiment, the specific implementation steps are as described in Examples 1 to 3 and will not be repeated here.

[0164] Example 6: As shown in the attached Figure 7 As shown, an embodiment of the present invention discloses a power equipment data management device based on dynamic federated learning and hierarchical distillation, comprising:

[0165] Data acquisition unit, which acquires local power equipment data to be managed and analyzed;

[0166] The data management and analysis unit inputs the local power equipment data to be managed and analyzed into the power equipment data management model to obtain the data management analysis results, wherein the power equipment data management model is trained through the power equipment data management model training method based on dynamic federated learning and hierarchical distillation.

[0167] In this embodiment, the power equipment data management model is trained by a power equipment data management model training method based on dynamic federated learning and hierarchical distillation. The specific method is as described in Examples 1 to 3 and will not be repeated here.

[0168] Example 7: An embodiment of the present invention discloses a storage medium, on which a computer program that can be read by a computer is stored. The computer program is configured to execute a method for managing power equipment data by dynamic federated learning and hierarchical distillation at runtime.

[0169] The above storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory, a mobile hard disk, a magnetic disk, or an optical disk.

[0170] Example 8: An embodiment of the present invention discloses an electronic device, including a processor and a memory, wherein the memory stores a computer program, and the computer program is loaded and executed by the processor to implement a method for managing power equipment data based on dynamic federated learning and hierarchical distillation.

[0171] The processor may be a central processing unit (CPU), a general-purpose processor, a digital signal processor (DSP), an ASIC, an FPGA, or other programmable logic device, a transistor logic device, a hardware component, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the present disclosure. It may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and so on. Memory may include, but is not limited to, various media capable of storing computer programs, such as USB flash drives, read-only memories, removable hard drives, magnetic disks, or optical disks.

[0172] It will be understood by those skilled in the art that the embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present invention may be implemented in various computer languages, for example, the object-oriented programming language Java and the interpreted scripting language JavaScript.

[0173] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0174] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0175] The above content is only a specific implementation method of the present invention, which has strong adaptability and implementation effect, but the scope of protection of the present invention is not limited to this. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed by the present invention, which should be covered by the scope of protection of the present invention. Therefore, equivalent changes made according to the claims of the present invention are still within the scope covered by the present invention.

Claims

1. A power equipment data management model training method based on dynamic federated learning and hierarchical distillation, characterized in that: include: Obtain several samples and divide them into training, validation, and test sets. Each sample includes local power equipment data and corresponding data management and analysis result labels. The data management and analysis results include data matching results, completeness scores, association recommendation confidence, and consistency judgments. The training set is input into the initial student model and the pre-trained teacher model. The teacher model is used to perform knowledge distillation training on the student model to obtain the trained student model. The teacher model is obtained by federated learning using several samples, and the global teacher model parameters are updated by dynamic contribution evaluation in federated learning. The trained student model is optimized using the validation set and test set to obtain the power equipment data management model.

2. The power equipment data management model training method based on dynamic federated learning and hierarchical distillation according to claim 1 is characterized in that: The pre-trained teacher model construction process includes: Establish a global teacher model on the server and randomly initialize the parameters of the global teacher model; Distribute the global teacher model and randomly initialized global teacher model parameters to each client participating in federated learning; After receiving the randomly initialized global teacher model parameters, each client uses local data to train the global teacher model and uploads the trained local model parameters to the server. The server triggers dynamic contribution evaluation to obtain the contribution of each local model parameter, aggregates the local model parameters based on the contribution, and obtains the updated global teacher model parameters; The above steps are repeated until the preset stopping condition is met, and the current global teacher model is used as the pre-trained teacher model.

3. The power equipment data management model training method based on dynamic federated learning and hierarchical distillation according to claim 2 is characterized in that: The server triggers dynamic contribution evaluation to obtain the contribution of each local model parameter, aggregates the local model parameters based on the contribution, and obtains updated global teacher model parameters, including: On the server side, each local model parameter is verified using a preset validation dataset to obtain the corresponding accuracy, diversity index, and difference norm; Using the accuracy, diversity index, and difference norm corresponding to each local model parameter, calculate the contribution of each local model parameter to the update of the global teacher model parameters in the current iteration; in, For contribution; is the accuracy rate; Div(D k ) is the diversity index; is the difference norm, is the global teacher model parameter updated in the previous iteration, are the local model parameters obtained in the current iteration; α, β, γ are hyperparameters; f norm (·) is the normalization function; Aggregate local model parameters based on contribution to obtain updated global teacher model parameters; in, is the updated global teacher model parameter; n k 、n j The amount of local data on the client; is the contribution; The local model parameters of the client are obtained for the current iteration; k and j are the number of clients; n k 、n j The amount of local data on the client; is the contribution; is the normalization of contribution.

4. The power equipment data management model training method based on dynamic federated learning and hierarchical distillation according to any one of claims 1 to 3, characterized in that: The training set is input into the initial student model and the pre-trained teacher model, and the distillation loss in the knowledge distillation training of the student model using the teacher model is as follows: L KD =L hard (M S (X),Y true )+λ1L soft (M S (X),M T (X))+λ2L HAFM (M S (X),M T (X)) Among them, L KD is the distillation loss; L HAFM is the hierarchical attention feature matching loss; L soft is the soft label loss; L hard is the supervised learning loss.

5. A method for managing power equipment data based on dynamic federated learning and hierarchical distillation, characterized in that: include: Obtain local power equipment data to be managed and analyzed; The local power equipment information data to be managed and analyzed is input into the power equipment information management model to obtain the information management analysis results, wherein the power equipment information management model is trained by the power equipment information management model training method based on dynamic federated learning and hierarchical distillation as described in any one of claims 1 to 4.

6. A power equipment data management model training device based on dynamic federated learning and hierarchical distillation using the method according to any one of claims 1 to 4, characterized in that: include: A sample acquisition unit acquires a number of samples and divides them into a training set, a validation set, and a test set. Each sample includes local power equipment data and corresponding data management and analysis result labels. The data management and analysis results include data matching results, completeness scores, confidence levels for associated recommendations, and consistency judgments. The model training unit inputs the training set into the initial student model and the pre-trained teacher model, and uses the teacher model to perform knowledge distillation training on the student model to obtain the trained student model. The teacher model is obtained by federated learning using several samples. The model verification unit uses the verification set and the test set to optimize the trained student model to obtain the power equipment data management model.

7. The power equipment data management model training device based on dynamic federated learning and hierarchical distillation according to claim 6 is characterized in that: Also included is a teacher model training unit, including: The server includes: Initialization module, establishes a global teacher model and randomly initializes the parameters of the global teacher model; The distribution module distributes the global teacher model and randomly initialized global teacher model parameters to each client participating in federated learning; Dynamic Contribution Evaluation Module: Obtains the contribution of each local model parameter, aggregates the local model parameters based on the contribution, and obtains updated global teacher model parameters, including: On the server side, each local model parameter is verified using a preset validation dataset to obtain the corresponding accuracy, diversity index, and difference norm; Using the accuracy, diversity index, and difference norm corresponding to each local model parameter, calculate the contribution of each local model parameter to the update of the global teacher model parameters in the current iteration; in, is the contribution; is the accuracy rate; Div(D k ) is the diversity index; is the difference norm, is the global teacher model parameter updated in the previous iteration, are the local model parameters obtained in the current iteration; α, β, γ are hyperparameters; f norm (·) is the normalization function; Aggregate local model parameters based on contribution to obtain updated global teacher model parameters; in, is the updated global teacher model parameter; n k 、n j The amount of local data on the client; is the contribution; The local model parameters of the client are obtained for the current iteration; k and j are the number of clients; n k 、n j The amount of local data on the client; is the contribution; is the normalization processing of contribution; After receiving the randomly initialized global teacher model parameters, multiple clients use local data to train the global teacher model and upload the trained local model parameters to the server.

8. A power equipment data management device based on dynamic federated learning and hierarchical distillation using the method of claim 5, characterized in that: include: Data acquisition unit, which acquires local power equipment data to be managed and analyzed; The data management and analysis unit inputs the local power equipment data to be managed and analyzed into the power equipment data management model to obtain the data management analysis results, wherein the power equipment data management model is trained by the power equipment data management model training method based on dynamic federated learning and hierarchical distillation as described in any one of claims 6 to 7.

9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the steps in the method according to any one of claims 6 to 7.

10. A storage medium, characterized in that: The storage medium stores a computer program that can be read by a computer, and the computer program is configured to execute the steps of the method according to any one of claims 6 to 7 when run.