Non-intrusive load monitoring method based on hybrid federated learning and cross Transform

By adopting hybrid federated learning and cross-transformer models in NILM, combined with graph neural network and layer sensitivity pruning technology, the shortcomings of existing NILM models in accuracy and privacy protection are solved, and efficient and accurate load monitoring and data privacy protection are achieved.

CN119962582AActive Publication Date: 2025-05-09JILIN UNIVERSITY

Patent Information

Application Number
CN202510436461.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-05-09
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

The existing NILM model is insufficient in handling diversified load working characteristics, and centralized data processing methods have data privacy protection and security risks, and federated learning is insufficient in scalability and privacy protection in NILM.

Method used

A non-invasive load monitoring method based on hybrid federated learning and cross-transformer is adopted. By dynamically switching centralized and decentralized federated learning modes, combined with graph neural network and layer sensitivity pruning technology, an intersection Transformer model is designed for load power prediction.

Benefits of technology

It improves the accuracy and adaptability of load monitoring, ensures user data privacy, enhances the scalability and attack resistance of the system, and reduces communication overhead and computing complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119962582A_ABST
    Figure CN119962582A_ABST
Patent Text Reader

Abstract

The invention discloses a non-intrusive load monitoring method based on hybrid federated learning and cross Transform, belongs to the technical field of power system management and energy efficiency, and solves the problems of insufficient accuracy, expansibility and privacy protection in the prior art. The method comprises the following steps of: combining a mixed federated learning framework of centralized federated learning and decentralized federated learning; a communication topology of decentralized federated learning is constructed through a graph neural network, and deep graph structure information is encoded into embedded representation of a client by using the graph neural network; the privacy protection of federal learning is enhanced through a layer sensitivity pruning and robust aggregation method; an input bus power sequence is processed through a one-dimensional convolutional layer and a shared Transform encoder, and prediction of a plurality of load working modes is realized by adopting a plurality of load branch structures; and in the decentralized federated learning training, judging whether the model is converged or not by adopting a voting mechanism. The method is suitable for monitoring the energy consumption of various load devices in a monitoring environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of power system management and energy efficiency technology, and more particularly to non-intrusive load monitoring in home and industrial environments. Background Art

[0002] With the continuous growth of global energy demand, especially in the field of building energy management, how to effectively reduce energy consumption and optimize energy use has become a hot issue in academia and industry. Non-Intrusive Load Monitoring (NILM) technology provides a low-cost and convenient solution by analyzing bus data to estimate the energy consumption of each load. NILM technology enables users and energy managers to achieve more sophisticated energy consumption monitoring and management, reduce energy waste, improve the level of energy automation management, and improve energy efficiency. Although NILM technology has broad application prospects, it still faces many challenges in practical applications in terms of accuracy, scalability, and privacy protection. Early NILM methods mainly relied on hardware devices to collect high-frequency voltage and current, and the high data acquisition cost limited its large-scale application. With the rapid development of machine learning and deep learning, NILM methods based on technologies such as recurrent neural networks (RNNs), convolutional neural networks (CNNs), generative adversarial networks (GANs), and attention mechanisms have achieved certain success. However, these methods often have difficulty maintaining high accuracy when dealing with diverse load behaviors in home and commercial environments. The models are still insufficient in accuracy and generalization ability, and a more effective load forecasting and monitoring method is needed.

[0003] In addition, the existing NILM model usually relies on a central server to centrally process a large amount of raw power data, which brings huge data privacy and security risks. The traditional centralized NILM data processing method can no longer meet the strict privacy protection requirements. Therefore, it is necessary to find a method that can effectively monitor the load and protect the privacy of user data.

[0004] As an emerging technical paradigm, federated learning allows models to be trained locally on multiple clients, thus avoiding data centralization and protecting user privacy. However, in the actual deployment of NILM, federated learning still faces problems such as insufficient personalization, high communication overhead, and complex privacy protection mechanisms. Although existing privacy protection technologies such as homomorphic encryption and secure multi-party computing can solve privacy problems in theory, they have high requirements for system resources and are not conducive to actual deployment; and the noise added by differential privacy will affect the accuracy and convergence speed of the model.

[0005] In summary, the prior art has the following problems: Insufficient accuracy of NILM model: Existing NILM technology has difficulty in handling diverse load working characteristics, and a method that can improve accuracy and adaptability is needed.

[0006] Data privacy protection and security risks: Centralized data processing methods are prone to expose user data privacy and pose high security risks. There is an urgent need for an NILM technology solution that protects user data privacy.

[0007] Insufficient scalability and privacy protection of federated learning: The practicality of traditional federated learning methods in NILM tasks is limited, and they face problems such as high communication overhead and difficulty in balancing privacy protection. It is necessary to improve the scalability, communication efficiency and anti-attack capabilities of federated learning in NILM while ensuring privacy protection. Summary of the invention

[0008] The purpose of the present invention is to solve the above technical problems and provide a non-intrusive load monitoring method based on hybrid federated learning and cross Transformer.

[0009] The present invention is implemented by the following technical solutions. On the one hand, the present invention provides a non-intrusive load monitoring method based on hybrid federated learning and cross Transformer, and the method includes: Step 1: Preprocess the power signal data; Step 2: Determine the federated learning training mode based on the monitoring network bandwidth, latency, and data privacy requirements. The federated learning training mode includes centralized federated learning and decentralized federated learning. If it is determined to be centralized federated learning, execute step 3, otherwise execute step 4. Step 3: Centralized federated learning training, deploying the CrossTransNILM model on the client and the central server, and centralizing all client model updates participating in the training to the central server for aggregation; the CrossTransNILM model includes a multi-layer Transformer encoder and a load branching module, and predicts the power consumption of multiple loads through bus power; Step 4: Decentralized federated learning training, each client independently trains the CrossTransNILM model and exchanges model parameters through point-to-point communication, optimizes the communication topology through the graph attention network, and judges the convergence of the global model through the voting mechanism; Step 5: Non-intrusive load monitoring and real-time feedback. The trained global CrossTransNILM model is deployed to the monitoring terminal to collect the power signal of the main line in real time. After preprocessing, it is input into the model for end-to-end monitoring. The model outputs the power decomposition results of each electrical device through the load branch module, and maps the load power to the start and stop status of the specific electrical device. The monitoring terminal regularly triggers the model parameter update process according to the federated learning mode selection mechanism, and maintains the model's adaptability to new electrical appliances and electricity consumption behaviors through federated training.

[0010] Furthermore, in step 3, the CrossTransNILM model specifically includes: a one-dimensional convolutional layer, a positional encoding, a shared Transformer encoder, and multiple load branches, each load branch processing a load type; each load branch includes a branched Transformer encoder layer, a cross Transformer layer, and two fully connected layers; The input bus sequence of the model is ,in is the length of the sequence, and is an odd number, used to locate the center point; the learning goal of the model is to predict the midpoint of the sequence Each load Power , the learning objective is expressed as: , in, is the hypothesis space, represents the total number of sampling points, is the mean square error loss function, is the total number of load types; The one-dimensional convolution layer is used to extract the local features of the sequence. The convolution operation formula is expressed as: , in, represents a one-dimensional convolution operation, is the input bus power sequence, is the output after convolution, representing the extracted local features; Positional encoding is used for each position Encoding the position in the sequence Generated by sine and cosine functions, the calculation formula is:

[0011] in, is the model dimension, Represents the sequence position index; The embedding layer is used to add the convolutional layer output to the positional encoding to obtain the embedded representation of the sequence. : , in, is the output after the convolutional layer. is the position information encoded by sine and cosine; The shared Transformer encoder is used to process and analyze the input bus power sequence. Through its multi-layer structure and multi-head self-attention mechanism, it captures long-distance dependencies in the sequence data.

[0012] Furthermore, the shared Transformer encoder consists of multiple encoding layers, each of which includes a multi-head self-attention mechanism and a feed-forward neural network.

[0013] Further, step 3 includes: Step 3.1: The central server first initializes the global model parameters , and then send it to all clients participating in the training. The mathematical expression is: ,in, Indicates the total number of clients; Step 3.2: Client Using local data Perform training and update model parameters , the update rule is: , in, Is the client In the Model parameters for round 1 local training, is the learning rate for model training, is the mean square error loss function Gradients with respect to parameters; Define each client In the The model trained locally is ,in, Represents the client In the In the training round The parameters of the layer, is the total number of layers of the CrossTransNILM model. Layer parameter mean change , called layer sensitivity, is defined as: ; Calculate the sensitivity of each layer and sort it, and select the one with the highest sensitivity. Layer, where is the transmission ratio coefficient, the value range is ; Step 3.3: The server collects all model parameters uploaded by the client ,in is the index of the round number of global training; for each client , calculate the Euclidean distance between it and all other client updates and : , in, Indicates that except client The index of other clients participating in the training. Indicates Client in round global training Model parameters of The smaller the score, the The closer the update is to other clients, the more likely it is to be considered a normal update; then Sort by removing the highest Layer update, where is the aggregation ratio coefficient, and its value range is ; Aggregate the remaining normal updates for each layer of the CrossTransNILM model , initialize the weight accumulator and counter , for each client upload weight , update the accumulator and counter , the average update of the calculation layer: ; Add the average update to the corresponding layer of the global model superior: ; Step 3.4: The central server updates the global model parameters Distribute to all clients to continue the next round of training: , .

[0014] Further, step 4 includes: Step 4.1: Use the graph attention network to learn the similarity between clients and select the most relevant edges based on the attention weights to obtain the communication topology. Step 4.2: The central server generates initial model parameters , the model parameters contain all the network layers and parameters that need to be trained for the NILM task; then, the central server initializes the global model parameters Distribute to all clients participating in the training so that each client has the same starting point before the training begins. The mathematical expression is as follows: , in, Represents the client The model parameters at initialization, is the total number of clients; Step 4.3: Perform local model training. Each client builds a communication topology based on the Share model parameters with neighboring clients; each client After receiving the updated parameters from its neighboring clients, the model parameters of each client are aggregated by weighted averaging; Step 4.4: Evaluate convergence and record the judgment result as a binary signal: , in, Indicates that the client believes that the current model has converged. Indicates that the model has not yet converged; Represents the client Validate the set locally The loss value on ; is the preset threshold, is the target accuracy, For Model Validate the set locally The accuracy of Step 4.5: After each round of training is completed, all clients send their own convergence signals Sent to the central server; the central server determines whether to stop training based on the statistical voting results. ; Set voting threshold is a static fixed value. , the system stops training, otherwise it enters the next round of training.

[0015] Further, step 4.1 includes: Step 4.1.1: Each client First, based on local data , extract key statistical features, including mean, standard deviation, minimum and maximum values, mathematically expressed as: , Among them, the characteristics of all clients are expressed as , ; make Represents the communication topology, client set , edge set Represents the potential communication connection relationship between clients, and initializes the edge set to a fully connected structure , does not include self-loops , that is, initially set to a completely directed graph ; Step 4.1.2: Use a three-layer graph attention network to build communication connections between clients. With its neighbor clients , the graph attention network first performs a learnable linear mapping on the client features ; Then, for each edge Calculate the attention coefficient , which is of the form: , in, is the activation function, is the learnable attention vector, represents vector concatenation, and Respectively represent the client and its neighbor clients The feature matrix of Then, after softmax normalization, the final attention coefficient is obtained : ; In the forward propagation process, the updated node representation is further obtained. And the set of attention coefficients of all edges ; Represented as client and The connection weight of Call , in, Embedding matrix for the updated client, is the set of attention coefficients; Step 4.1.3: Set the attention weight threshold for The median of The topological edges of: , and remove self-loops ; for Each directed edge in , only the undirected form is retained , the obtained undirected edge set ; For any client , Count its degree ,in, Representation and Connected neighbor client; if there is , indicating that the client is If there is no connecting edge in Pick a neighbor client and merge the new edge into it ; make , the final undirected graph This is the decentralized communication topology obtained based on GAT learning.

[0016] Furthermore, step 2 specifically includes: When network bandwidth ,Delay , and the data privacy requirements are low, the system chooses centralized federated learning, otherwise, it switches to decentralized federated learning; If the conditions for centralized federated learning training are met, the central server will aggregate and optimize the model updates of all clients; If the conditions for decentralized federated learning training are triggered, the client trains independently in a distributed environment and exchanges model updates with neighboring clients.

[0017] In a second aspect, the present invention provides a non-intrusive load monitoring system based on hybrid federated learning and cross Transformer, the system comprising: A preprocessing module, used for preprocessing power signal data; A federated learning training mode selection module is used to determine the federated learning training mode according to the monitoring network bandwidth, latency and data privacy requirements, wherein the federated learning training mode includes centralized federated learning and decentralized federated learning; if centralized federated learning is determined, the centralized federated learning module is executed, otherwise the decentralized federated learning module is executed; Centralized federated learning module, used for centralized federated learning training, deploys CrossTransNILM model on the client and central server, and centralizes all client model updates participating in the training to the central server for aggregation; CrossTransNILM model includes multi-layer Transformer encoder and load branch module, and predicts the power consumption of multiple loads through bus power; Decentralized federated learning module, used for decentralized federated learning training. Each client independently trains the CrossTransNILM model and exchanges model parameters through point-to-point communication. The communication topology is optimized through the graph attention network, and the convergence of the global model is judged through the voting mechanism. The monitoring module is used for non-intrusive load monitoring and real-time feedback. It deploys the trained CrossTransNILM model to the monitoring terminal, collects the power signal of the main line in real time, and inputs it into the model for end-to-end monitoring after preprocessing. The power decomposition results of each electrical equipment are output through the load branch module, and the load power is mapped to the start and stop status of the electrical equipment. The monitoring terminal regularly updates the model parameters according to the federated learning mode selection mechanism.

[0018] In a third aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor runs the computer program stored in the memory, the steps of a non-intrusive load monitoring method based on hybrid federated learning and cross Transformer as described above are executed.

[0019] In a fourth aspect, the present invention provides a computer-readable storage medium, wherein a plurality of computer instructions are stored in the computer-readable storage medium, and the plurality of computer instructions are used to enable a computer to execute a non-intrusive load monitoring method based on hybrid federated learning and cross Transformer as described above.

[0020] Beneficial effects of the present invention: The present invention provides a non-intrusive load monitoring method based on a hybrid federated learning framework and a cross-Transformer structure, aiming to solve the shortcomings of accuracy, scalability and privacy protection in the prior art. The present invention provides a flexible hybrid federated learning framework that combines the advantages of centralized and decentralized federated learning to adapt to different network conditions and privacy protection requirements, and designs a cross-Transformer model for simultaneously predicting the power consumption of multiple loads through bus power. The method of the present invention specifically includes: (1) A hybrid federated learning framework that combines centralized federated learning and decentralized federated learning: By dynamically switching between centralized and decentralized federated learning modes to adapt to different network conditions and privacy requirements, the adaptability and robustness of the model are enhanced.

[0021] (2) Constructing the communication topology of decentralized federated learning through graph neural networks: Using graph neural networks to encode deep graph structure information into the client's embedded representation, capture complex communication topologies and remote dependencies between clients, improve communication efficiency and stability, and enhance the system's resistance to Byzantine attacks.

[0022] (3) NILM privacy protection and anti-attack strategy: Enhance the privacy protection of federated learning through layer-sensitive pruning and robust aggregation methods. Layer-sensitive pruning reduces communication requirements and enhances data security by pruning key model parameters; the Krum algorithm robust aggregation method is used to resist Byzantine attacks and ensure the reliability of model updates.

[0023] (4) Cross-Transformer structure (CrossTransNILM) design: The input bus power sequence is processed through a one-dimensional convolutional layer and a shared Transformer encoder, and multiple load branch structures are used to realize the prediction of multiple load working modes. The input is the bus power sequence, and the output is the power values ​​of multiple loads at the midpoint of the sequence.

[0024] (5) Voting mechanism consensus for judging the convergence of the NILM model: In decentralized federated learning training, a voting mechanism is used to judge whether the model has converged, thereby increasing the adaptability of model training.

[0025] The present invention provides a new solution for the NILM field by introducing a hybrid federated learning framework, a cross-Transformer structure, a layer sensitivity pruning strategy, Krum robust aggregation, Graph Neural Networks (GNN) and a model convergence voting mechanism. It has significant technical advantages and has important application prospects in smart grids and household energy consumption management. The present invention uses a hybrid federated learning framework and a cross-Transformer model to monitor the energy consumption of various load devices in the monitoring environment without connecting sensors to each load separately.

[0026] The present invention is applicable to environments requiring privacy protection and data security, such as data centers and residential buildings, where the confidentiality of user data and the security of the system are of vital importance. In addition, the application of the present invention can also be extended to the construction of smart cities and smart grids to provide support for urban energy management. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the technical solution of the present application, the drawings required for use in the embodiments are briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0028] Figure 1 is a schematic diagram of the method of the present invention; Figure 2 Schematic diagram of the NILM model CrossTransNILM model of the present invention; Figure 3 A schematic diagram of a shared Transformer encoder of the present invention; Figure 4 Schematic diagram of the cross Transformer encoder layer; Figure 5 Loss curve of the CrossTransNILM model under different pruning ratios. DETAILED DESCRIPTION

[0029] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and should not be construed as limiting the present invention.

[0030] Specific implementation method: Figure 1 As shown in the figure, a non-intrusive load monitoring method based on hybrid federated learning and cross Transformer is described in detail below. The working principle, working process and steps of the non-intrusive load monitoring method are introduced in detail.

[0031] Step 1: Data Preprocessing Data preprocessing is a key step to ensure the effectiveness of subsequent model training, verification, and testing. Power signals collected by sensors or meters require a series of processing to improve data quality and consistency. The following are the detailed preprocessing steps: Step 1.1: Data cleaning and alignment First, the power of the bus and each load is visualized to check the integrity and consistency of the data. By observing the power waveform, the abnormal parts are identified and deleted, such as missing data lasting more than half an hour. For short-term missing data (a few seconds to a few minutes), since it has little impact on the overall data distribution, it is retained and used as a noise source for model regularization to enhance the robustness of the model.

[0032] Then, the sampling rates of the bus power and the power of each load are unified to achieve consistency in the time step. The constructed data set contains the power of each load as well as the bus power. Bus power It is calculated by adding the sum of the load powers and the random noise term: , in, is the total number of load types, Indicates load In time The power, is the noise term.

[0033] Step 1.2: Data Standardization In order to make the loads and bus powers comparable and eliminate the impact of dimension on model training, data standardization is required. Taking bus power as an example, the standardized power Calculated by the following formula: , in, is the mean power, is the standard deviation of power.

[0034] Step 2: Federated Learning Training Mode Selection Hybrid federated learning combines centralized and decentralized training modes, dynamically selecting the optimal mode to adapt to different network conditions and privacy requirements, thereby improving the robustness of the system. ,Delay As well as data privacy requirements, determine whether the current CrossTransModel model training uses centralized federated learning (step 3) or decentralized federated learning (step 4).

[0035] Step 2.1: Conditions for switching to federated learning mode set up is the network bandwidth threshold, is the delay threshold, by monitoring , and privacy requirements, the system dynamically selects the training mode: .

[0036] Specifically, when the network bandwidth ,Delay , and when the data privacy requirements are low, the system chooses centralized federated learning to ensure higher communication efficiency and global synchronization, and accelerate the aggregation and optimization of the global model. Otherwise, the system switches to decentralized federated learning to reduce network communication burden and latency, and minimize the transmission of sensitive data.

[0037] Step 2.2: Processing flow after mode selection If the conditions for centralized federated learning training are met, proceed to step 3, centralized federated learning training, where the central server aggregates and optimizes the model updates of all clients to obtain higher model consistency and overall performance.

[0038] If the conditions for decentralized federated learning training are triggered, proceed to step 4, decentralized federated learning training. The client trains independently in a distributed environment and exchanges model updates with neighboring clients to improve the system's fault tolerance and privacy protection level.

[0039] Step 3: Centralized Federated Learning Training Centralized federated learning aggregates all client model updates participating in the training to a central server to ensure model consistency and efficiency. This process includes four steps: client model initialization, local model training and layer sensitivity pruning and uploading, server model parameter aggregation, and global model distribution: Step 3.1: Client model initialization Step 3.1.1. CrossTransModel model structure design The NILM model is deployed on the client and the central server. The NILM model of this embodiment is a prediction model based on a cross Transformer structure, denoted as CrossTransNILM. The model combines a multi-layer Transformer encoder and a load branching module to predict the power consumption of multiple loads through bus power.

[0040] The structure of the CrossTransNILM model is as follows Figure 2 As shown in the figure, it includes: one-dimensional convolutional layer, positional encoding, shared Transformer encoder and multiple load branches. Each load branch contains a branch Transformer encoder layer, a cross Transformer layer and two fully connected layers. The overall structure is designed to effectively capture the load operation law and perform sequence-to-point multi-load power prediction for NILM tasks.

[0041] Step 3.1.1.1. Model input In this implementation, a sequence-to-point learning strategy is adopted to improve the accuracy of prediction and reduce the computational complexity. First, the input bus sequence is defined as ,in is the sequence length, to facilitate the location of the center point, is chosen as an odd number. The learning goal of the model is to predict the midpoint of the sequence Each load Power . Therefore, the learning objective can be expressed as: , in, is the hypothesis space, represents the total number of sampling points, is the mean squared error (MSE) loss function, is the total number of load types.

[0042] Step 3.1.1.2, one-dimensional convolutional layer The one-dimensional convolution layer is used to extract local features of the sequence to reduce the dimension of the data and retain the significant feature patterns in the sequence. The convolution operation formula is expressed as follows: , in, represents a one-dimensional convolution operation, is the input bus power sequence, is the output after convolution, representing the extracted local features.

[0043] Step 3.1.1.3: Position encoding In order to enable the NILM model to identify the position relationship in the sequence, each position Encoding the position in the sequence Generated by the sine and cosine functions, the calculation formula is as follows: , , in, is the model dimension, Represents the sequence position index, and improves the model's perception of time series through sine and cosine position encoding.

[0044] Step 3.1.1.4. Embedding layer Add the convolutional layer output to the positional encoding to get the embedded representation of the sequence : , in, is the output after the convolutional layer. It is the position information of sine and cosine encoding. This representation method combines position encoding and local feature information, enhancing the representation ability of the model.

[0045] Step 3.1.1.5. Shared Transformer Encoder The shared Transformer encoder used in this embodiment is designed to process and analyze the input bus power sequence. Through its multi-layer structure and multi-head self-attention mechanism, it can capture the long-distance dependencies in the sequence data and achieve a global understanding of the entire sequence. The schematic diagram of the shared Transformer encoder in this embodiment is shown in Figure 3 The encoder consists of multiple encoding layers, each of which includes a multi-head self-attention mechanism and a feedforward neural network.

[0046] (1) Multi-head attention mechanism The multi-head attention mechanism allows the model to learn the internal structure of the input data in different representation spaces in parallel, with each head performing a different linear transformation on the data and applying the attention mechanism independently. In each encoder layer, the position-encoded sequence is mapped into a query matrix ( ), key matrix( ) and the value matrix ( ): , , , in, , , is the parameter matrix learned during training.

[0047] The attention mechanism helps the model extract important information from the input. The calculation method is as follows: , in, is the key matrix The dimension of Adjust the calculation stability. Through multiple heads in parallel, the attention output corresponding to each head is calculated as follows: , in, , , is the independent transformation matrix corresponding to each head, is the number of heads in each encoder layer. Then, the results of all heads are concatenated and the output weight matrix is Perform a linear transformation: , This design enables the model to integrate multiple feature expressions and improve information integration capabilities.

[0048] By introducing the multi-head self-attention mechanism, the shared Transformer encoder of this embodiment can not only effectively process and analyze data, but also enhance the expressiveness and prediction accuracy of the model by processing various features in parallel.

[0049] (2) Residual connection and layer normalization In order to improve the efficiency and effectiveness of training deep networks, a residual connection is added after each attention output: , Residual connections help gradients flow directly to deeper network layers, preventing the gradient vanishing problem during training.

[0050] Then, layer normalization is applied after each residual connection to improve the stability of model training: , in, The columns of the matrix The mean value is obtained by taking the unit as is the dimension of the matrix rows, Indicates the matrix Line Elements of a column. The columns of the matrix The variance is obtained with unit. A very small number to prevent the denominator from being 0 during the calculation process.

[0051] (3) Feedforward neural network, residual connection and layer normalization Build a basic neural network structure, which contains two linear layers and a ReLU (Rectified Linear Unit) activation function. By combining the linear layer and the ReLU activation function, the input Perform feature transformation and nonlinear mapping to obtain the output of the feedforward neural network : , Among them, the ReLU activation function , which helps to introduce nonlinear characteristics into the network, alleviate the gradient vanishing problem, and speed up the convergence of the network.

[0052] Then, the residual connection and layer normalization are performed in the same way as (2): , Get the output of the shared Transformer .

[0053] Step 3.1.1.6: Multiple load branches In order to simultaneously predict the power of different loads, the CrossTransNILM model of this implementation is designed with multiple load branches, each branch processing a load type. This design allows the model to process the unique characteristics of each load more carefully and improve the prediction efficiency. Each load branch contains an independent Transformer encoder layer, a cross attention layer, and two fully connected layers.

[0054] 1. Transformer encoder layer Each load branch contains an independent Transformer encoder layer, which has a similar structure to the shared Transformer encoder, including multi-head self-attention and feedforward networks. Feature extraction and learning for each load can refine the characteristics of each load for a deeper analysis and understanding.

[0055] (II) Cross Transformer Encoder Layer The cross-attention mechanism is used in the cross-Transformer encoder layer to make each load branch refer to the global information from the shared Transformer encoder when making power predictions. The design of the cross-attention mechanism in this embodiment is based on the view that although each load has its own unique load characteristics, they often have some connections in practical applications. Through cross-attention, the model can take into account the interaction and influence with other loads when predicting the power of one load. The schematic diagram of the cross-Transformer encoder layer is shown in Figure 2. Figure 4 shown.

[0056] Specifically, the query matrix of the cross attention Generated by the output of the Transformer encoder layer of the current load branch, and the key matrix Sum Matrix Output from the shared Transformer encoder, which enables each load branch to take into account the interaction and impact with other loads in power prediction.

[0057] (III) Fully connected layer and output The output of each load branch passes through two linear layers and a (Gaussian Error Liner Unit) activation function combination for the final feature transformation and nonlinear mapping. Among them, Activation Function , is the input, yes The standard normal cumulative distribution function of is a nonlinear error function used to smoothly transition input values. It provides a smoother nonlinear transformation, so the output of the activation function can more finely adjust the transmission of the input signal, especially in the weak change area of ​​the signal. This helps to capture more subtle pattern changes and improve the sensitivity and accuracy of the model in processing data details.

[0058] Finally, the load is predicted At the midpoint of the input sequence The corresponding power value : , in, is the output of the cross Transformer encoder layer.

[0059] Step 3.1.2: Model initialization and synchronization Before training begins, the central server first initializes the global model parameters , and then send it to all clients participating in the training, so as to ensure that each client starts training from the same initial state of the NILM model, which can be expressed mathematically as: , in, Represents the total number of clients. This ensures that in the 0th round of training, the initial state of all client models is consistent.

[0060] Step 3.2: Local model training, layer sensitivity pruning and uploading In each round of centralized federated learning global training, each client uses local data to perform independent model training. The specific steps are as follows: First, the client Using local data Perform training and update model parameters The update rule is: , in, Is the client In the Model parameters for round 1 local training, is the learning rate for model training, is the mean square error loss function The gradient of the parameter.

[0061] Then, due to the problem of weight parameter redundancy in the deep neural network model, this implementation method proposes to perform layer sensitivity pruning when uploading parameters, and only upload some sensitive layer parameters to reduce the amount of transmitted data, improve communication efficiency and enhance privacy protection. Define each client In the The model trained locally is ,in, Represents the client In the In the training round The parameters of the layer, is the total number of layers of the CrossTransNILM model. Layer parameter mean change , called layer sensitivity, is defined as: , The greater the layer sensitivity, the more important changes in this layer may be to improve model performance.

[0062] Finally, the sensitivity of each layer is calculated and sorted, and the top layer with the largest sensitivity is selected. Layer, where is the transmission ratio coefficient, the value range is Through this pruning strategy, the client only needs to upload some layer parameters with high sensitivity, which reduces communication overhead, protects sensitive information, and improves model security.

[0063] Step 3.3: Server model parameter aggregation In order to deal with malicious attacks and abnormal parameter updates in centralized federated learning, this implementation adopts the Krum aggregation algorithm on the server side.

[0064] First, the server collects all model parameters uploaded by the client. ,in is the index of the round number of global training. , calculate the Euclidean distance between it and all other client updates and : , in, Indicates that except client The index of other clients participating in the training. Indicates Client in round global training The model parameters. The smaller the score, the The closer the update is to other clients, the more likely it is to be considered a normal update. Sort by removing the highest Layer update, where is the aggregation ratio coefficient, and its value range is .

[0065] Finally, the remaining normal updates are aggregated. Since each client performs layer sensitivity pruning, the central server only receives the weights of some layers of each client. , initialize the weight accumulator and counter . For each client upload weight , update the accumulator and counter . Compute the average update of the layer: .

[0066] Add the average update to the corresponding layer of the global model superior: .

[0067] Step 3.4: Global model distribution After completing the aggregation of model parameters, the central server will update the global model parameters Distribute to all clients to continue the next round of training: , , Ensure that client model parameters in centralized federated learning are effectively optimized and updated.

[0068] Step 4: Decentralized Federated Learning Training When model training switches to decentralized federated learning mode, each client independently trains the CrossTransNILM model and exchanges model parameters through peer-to-peer communication to enhance privacy protection. The communication topology is optimized through the GraphAttention Network (GAT), and the convergence of the global model is judged through a voting mechanism. The decentralized structure improves the system's ability to resist Byzantine attacks and the scalability of communication, and reduces reliance on single-point failures of central servers. The specific steps are as follows: Step 4.1: Communication topology construction In a decentralized federated learning setting, the communication topology determines the direction in which model parameters flow during model training. In order to optimize communication efficiency and enhance network robustness in decentralized federated learning, this implementation uses GAT to learn the similarity between clients and select the most relevant edges based on attention weights to obtain the communication topology. The system dynamically adjusts the communication structure after a certain number of training rounds to cope with changes in network conditions and failures in communication links. The following are the detailed steps for building the communication topology: Step 4.1.1: Initialize node characteristics and communication topology Each client First, based on local data , extract key statistical features. These features include mean, standard deviation, minimum and maximum values, thus forming a feature vector to describe the data characteristics of each client. The mathematical expression is as follows: , Among them, the characteristics of all clients are expressed as , .

[0069] make Represents the communication topology, client set , edge set In order to form an effective learning and information exchange topology network and avoid isolated clients, this implementation initializes the edge set to a fully connected structure. , does not include self-loops , that is, initially set to a completely directed graph .

[0070] Step 4.1.2: Communication topology construction based on GAT model Three-layer GAT is used to build communication connections between clients. With its neighbor clients , GAT first performs a learnable linear mapping on the client features Then, for each edge Calculate the attention coefficient , which is generally in the form of: , in, is the activation function, is the learnable attention vector, represents vector concatenation, and Respectively represent the client and its neighbor clients The feature matrix of . Then it is normalized by softmax to get the final attention coefficient : .

[0071] In the forward propagation process, the updated node representation is further obtained. And the set of attention coefficients of all edges In this paper’s decentralized federated learning Represented as client and The connection weight of .

[0072] Call ,in, Embedding matrix for the updated client, is the set of attention coefficients.

[0073] Step 4.1.3: Communication topology edge screening First, set the attention weight threshold for The median of The topological edges of: , and remove self-loops .

[0074] Then, for Each directed edge in , only the undirected form is retained , the obtained undirected edge set .

[0075] Next, for any client , count its degree ,in, Representation and The connected neighbor client. , indicating that the client is If there is no connecting edge in Pick a neighbor client and merge the new edge into it , to avoid its isolation.

[0076] Finally, let The final undirected graph This is the decentralized communication topology obtained based on GAT learning.

[0077] To adapt to changing network conditions and client privacy protection requirements, the system will periodically re-evaluate and optimize the communication topology. After a round of training, the system will reuse GAT to calculate client embeddings and rebuild the communication topology based on these new embeddings to ensure an efficient information exchange path, improve training efficiency and overall performance.

[0078] Step 4.2: Client initialization Before training begins, you first need to ensure that all clients start training from the same model state to maintain consistency during training. Specifically: First, the central server generates the initial model parameters , which contains all the network layers and parameters that need to be trained for the NILM task. Then, the central server initializes the global model parameters Distribute to all clients participating in the training so that each client has the same starting point before training begins. The mathematical representation is as follows: , in, Represents the client The model parameters at initialization, The total number of clients.

[0079] Step 4.3: Client local model training and update In the decentralized federated learning model, each client follows a similar process to the centralized model, first training the model locally. After each round of local training, each client builds a communication topology. The model parameters are shared with neighboring clients to reduce the amount of data transmission and enhance data privacy. Then, each client After receiving the updated parameters from its neighboring clients, the model parameters of each client are aggregated by weighted averaging.

[0080] Step 4.4: Convergence evaluation and voting mechanism In decentralized federated learning, convergence evaluation and voting mechanisms are designed to ensure that training is stopped when sufficient model performance is achieved to save computing resources while preventing overfitting.

[0081] Step 4.4.1. Convergence Assessment Convergence evaluation in decentralized federated learning includes the following two aspects: On the one hand, after each client completes local training, the model is obtained , when the local validation set The loss change on (like ), it means that the model has stabilized after multiple rounds of training. On the other hand, when the model Validate the set locally The accuracy reaches or exceeds the target accuracy When , the model is considered to have achieved the expected performance and can be considered to have converged.

[0082] Each client is evaluated based on the above criteria and the judgment results are Recorded as a binary signal: , in, Indicates that the client believes that the current model has converged. Indicates that the model has not yet converged. Represents the client Validate the set locally The loss value on .

[0083] To evaluate In this implementation, the mean absolute error (MAE), normalized signal aggregate error (SAE) and energy per day (EpD) are selected to evaluate the algorithm: 1) MAE: Evaluation Load At every point in time Power prediction value With actual value The mean absolute error between . The smaller the value, the more accurate the prediction result of the model. Represents the total number of test samples.

[0084] .

[0085] 2) SAE: Represents the relative error of total energy. The smaller the SAE, the smaller the prediction error of the model and the better the model performance.

[0086] , in, Indicates load The total energy consumption, Indicates the predicted load The total energy consumption, is the sampling interval.

[0087] 3) EpD: represents the absolute error of energy prediction within one day. The smaller the EpD, the smaller the absolute error of energy prediction within one day, and the better the model performance.

[0088] , in, , indicating the sample The total number of days included in . , indicating the number of sampling points contained in a day. , indicating load The sum of actual energy consumed during a day cycle, , indicating load The total amount of energy predicted over a one-day period.

[0089] Step 4.4.2, Voting Mechanism After each round of training, all clients send their own convergence signals The central server determines whether to stop training based on the statistical voting results. . The system sets the voting threshold is a static fixed value. When , the system stops training, otherwise it enters the next round of training. In summary, the training process will terminate when any of the following conditions are met: 1) Reach the maximum number of rounds .

[0090] 2) Exceeding the voting threshold (For example ).

[0091] 3) The global model performance reaches the target (such as the loss change on the global validation set is less than the preset threshold) hour).

[0092] Through convergence evaluation and voting mechanism, this implementation effectively avoids overfitting and ensures that the training process stops after reaching the set goal. At the same time, the voting mechanism also improves the robustness of the system, avoiding model deviations caused by improper training of a few clients or communication problems, thereby improving the overall performance and stability of the system.

[0093] Conventionally, model convergence is based on the transformation of loss or reaching a specified number of rounds. However, in decentralized federated learning, due to the aggregation of models from multiple clients, the loss may fluctuate greatly. Using only the loss function or specifying the number of training rounds may make it difficult to judge convergence or increase the number of training rounds for convergence. Therefore, a convergence judgment and voting mechanism is added here.

[0094] In order to verify the effectiveness of the hybrid federated learning framework and CrossTransNILM model proposed in this paper, this paper designed and conducted a number of experiments to evaluate the performance of this paper under different conditions. This paper selected two public household electricity datasets, namely REFIT and UK-DALE: 1) REFIT: REFIT data is collected from 20 buildings in a certain country from 2013 to 2015. The sampling period is once every 8 seconds, including the power data of the bus and each load.

[0095] 2) UK-DALE: The UK-DALE dataset contains power data of five households in a country from 2013 to 2015. The sampling period of the total power supply is once every 1 second, and the sampling period of each load is every 6 seconds.

[0096] In the simulation design, five typical household loads (microwave oven, refrigerator, dishwasher, washing machine and electric kettle) are selected to verify the performance of the NILM model of the present invention. In the present invention, the format of the data is defined as: (washing machine power, dishwasher power, microwave oven power, refrigerator power, electric kettle power, bus power), where the bus power is the sum of the powers of the above five loads.

[0097] Before applying the NILM algorithm, the data set needs to be uniformly sampled to achieve time step alignment. In the present invention, 1 / 8 Hz is selected as the standard sampling rate. In addition, the two data sets are divided into federated learning clients, and the specific division method is shown in Table 1.

[0098] Table 1 Client data division

[0099] Finally, REFIT and UK-DALE are standardized. In the simulation of the present invention, the mean and standard deviation values ​​used for standardization are shown in Table 2.

[0100] Table 2 Standardization parameters of data

[0101] It should be noted that the parameters in Table 2 are only used to standardize model training, validation and test data, and are not used to determine the actual mean and variance of these loads. The following are specific experimental results and their explanation of the effects of the invention.

[0102] 1. Performance comparison between the CrossTransNILM model of the present invention and other NILM models In order to evaluate the effectiveness of the CrossTransNILM model proposed in this paper, two currently most advanced NILM models were selected for comparative experiments using local training: a sequence-to-point model based on a convolutional neural network (sequence-to-point S2P) and a BERT4NILM model that first applied Transformer. The specific model descriptions are as follows: Sequence-to-point model (sequence-to-point S2P) based on convolutional neural network: This model contains multiple one-dimensional convolutional layers and ReLU activation functions, as well as two fully connected layers, which are used to process the input active power sequence. The features of the input bus power sequence are extracted through the convolutional layer, and the power estimates of the five loads corresponding to the center position of the sequence are output through the fully connected layer. Transformer-based BERT4NILM model: BERT4NILM combines the encoder-decoder architecture and the attention mechanism, and uses the advantages of Transformer in processing time series data to achieve the key capture of sequence information through the attention mechanism, especially when processing power data with complex patterns. Compared with the model without attention mechanism, it has more advantages. In order to ensure fairness, the hyperparameter settings of the above two comparison models are consistent with the CrossTransNILM model of the present invention, and all three models are trained until convergence. The general hyperparameter settings of the algorithm are shown in Table 3, and the data partitioning used is shown in Table 4.

[0103] Table 3 General hyperparameter settings of the algorithm

[0104] Table 4. Local model training data division

[0105] Table 5 is the comparison results of the three groups of experiments. It can be seen from Table 5 that the CrossTransNILM model shows significant advantages in most load types and three indicators. Especially in the loads of washing machines, microwave ovens and electric kettles, the performance of CrossTransNILM is significantly better than the other two models. In addition, CrossTransNILM also showed the best results in MAE of dishwashers and SAE and EpD of refrigerators. This shows that CrossTransNILM can more accurately capture the energy consumption characteristics of different loads and effectively balance the prediction accuracy and energy error. It is suitable for the sequence-to-point multi-load prediction scenario of NILM, which reflects the advanced nature of the CrossTransNILM model of the present invention.

[0106] Table 5 Performance comparison of CrossTransNILM local training and other models

[0107] 2. Pruning ratio experiment In order to test the impact of different pruning ratios on model performance, we set four pruning ratios of 0%, 5%, 10%, and 20%. The experimental results are shown in Table 6. The results of each evaluation index under different pruning ratios of different load types are compared.

[0108] Table 6 CrossTransNILM model performance under different pruning ratios

[0109] The results in Table 6 show that at a pruning ratio of 5% to 10%, the performance of the CrossTransNILM model on most load types is comparable to that of the unpruned model, and it is able to reduce the amount of communication while maintaining a high prediction accuracy. In particular, at a pruning ratio of 5%, in the prediction of electric kettles, its MAE and SAE indicators are even better than the unpruned results. This may be because pruning introduces a certain regularization effect, which helps prevent overfitting. However, when the pruning ratio reaches 20%, the model performance decreases significantly, indicating that the effect of pruning is not infinitely increased, but will have a negative impact on the model's prediction ability to a certain extent. Therefore, in the choice of pruning, it is necessary to balance model performance and communication efficiency and set the pruning ratio reasonably.

[0110] To evaluate the impact of layer sensitivity pruning on model training, the Loss curves of centralized federated learning training of the CrossTransNILM model under different layer sensitivity pruning ratios are compared, as shown in Figure 5.

[0111] from Figure 5 The experimental results show that the unpruned model (0% pruning ratio) has the best performance, and the loss value drops rapidly and remains at the lowest level. In contrast, the model with a 5% pruning ratio can also converge well, but there are some fluctuations during the training process, and the final loss is slightly higher than the unpruned model. As the pruning ratio increases, the 10% pruned model can converge, but the loss value is slightly higher and the fitting ability is weakened. The 20% pruned model fluctuates greatly during the convergence process, and the final loss is the highest, which significantly affects the learning ability of the model.

[0112] Figure 5 The results further show that the choice of pruning ratio has a direct impact on model performance. Moderate pruning (such as 5% pruning ratio) can reduce the computational complexity while retaining model performance, while excessive pruning (such as 20% pruning ratio) may weaken the model's fitting ability. Therefore, through moderate pruning processing, the present invention can maintain high model accuracy and good convergence speed while reducing communication overhead, which shows obvious advantages under limited bandwidth conditions.

[0113] 3. Federated Learning Performance Comparison Experiment In order to evaluate the effectiveness of the federated learning of various models of the present invention, the performance of local training, centralized federated learning and decentralized federated learning was compared. Each model was tested using the data of the client numbered 1. The experimental results are shown in Table 7.

[0114] Table 7 Performance comparison of CrossTransNILM local training and other models

[0115] By comparing the performance of local training, 5% pruned centralized federated learning, and decentralized federated learning under five load types, namely washing machine, dishwasher, microwave oven, refrigerator, and electric kettle, in Table 7, the results show that local training outperforms the federated learning method in all performance indicators. The performance of centralized federated learning is better than that of decentralized federated learning on all load types, which may be because the centralized method can better coordinate the information of each client and reduce model inconsistency. However, the performance degradation of these federated learning methods is not significant, indicating that after the introduction of federated learning, the model can still maintain a high prediction accuracy. This verifies the effectiveness of the various federated learning models proposed in the present invention in processing different load types. Although there is a certain gap compared to local training, its performance is sufficient to support its feasibility and advantages in practical applications.

[0116] IV. Conclusion Through the above experiments, the present invention effectively demonstrates the performance of the hybrid federated learning framework and the CrossTransNILM model in various environments. The results of various experiments verify the significant advantages of the present invention in communication efficiency, model accuracy and anti-attack capabilities. The invention can not only achieve efficient load monitoring in limited bandwidth and privacy protection scenarios, but also show strong adaptability and stability when facing limited communication conditions and security threats. These characteristics make the present invention suitable for a wide range of application scenarios, including smart grids, home energy consumption management and other fields.

[0117] It should be noted that the above design is only an example of the implementation of the present invention, and can be adjusted and optimized according to specific needs in actual applications. For example: (1) Communication topology and federated learning mode selection: When designing a decentralized federated learning system, the present invention adopts a communication topology structure inferred by a graph neural network and a hybrid federated learning mode. However, depending on network conditions and data privacy requirements, different communication topologies (such as fully connected or rule-based connections) and federated learning modes (such as centralized, decentralized, or hybrid modes) can be flexibly selected.

[0118] (2) Pruning strategy and pruning ratio: When transmitting model parameters, the present invention adopts layer-sensitive pruning technology to reduce the amount of data transmitted and improve data security. However, the pruning ratio and strategy can be flexibly adjusted according to the performance requirements of the actual application scenario. For applications with high requirements for communication efficiency, a higher pruning ratio can be used to significantly reduce communication overhead; for scenarios that pursue higher model accuracy, the pruning ratio can be appropriately reduced to maintain model performance.

[0119] (3) Flexibility of the cross-Transformer structure: The cross-Transformer structure (CrossTransNILM) of the present invention extracts features from the bus power sequence through a one-dimensional convolutional layer and a shared Transformer encoder, and uses multiple load branch structures to achieve accurate prediction of complex load working modes. In practical applications, the number of Transformer layers, convolution kernel size, and branch structure can be adjusted according to the characteristics of different loads to adapt to different load characteristics and ensure the prediction accuracy and computational efficiency of the model in complex scenarios.

[0120] It is worth noting that although federated learning and graph neural network technologies have been successfully applied in many fields, such as privacy computing, recommendation systems, and network analysis, their application in the NILM field is still in the exploratory stage. Therefore, the hybrid federated learning framework proposed in this invention, which combines centralized and decentralized networks, can provide a new and flexible solution for the NILM field by dynamically switching to cope with different network conditions and privacy requirements.

[0121] In addition, the cross-Transformer structure (CrossTransNILM) in this invention uses the powerful feature extraction capability of the Transformer model to efficiently predict the complex patterns of power loads. This innovative architecture design, by combining the client embedding optimization of the graph neural network and the cross-attention mechanism, can not only effectively improve the load forecasting accuracy, but also better capture the potential connections between loads, providing a new approach for deep learning modeling of NILM tasks.

[0122] In general, the present invention is based on the innovative combination of hybrid federated learning framework, cross Transformer structure, graph neural network and privacy protection technology, and is expected to bring new technological breakthroughs and application prospects for non-intrusive load monitoring in the fields of smart grid and household energy consumption management.

Claims

1. A non-intrusive load monitoring method based on hybrid federated learning and cross-Transformer, characterized in that: The method comprises: Step 1: Preprocess the power signal data; Step 2: Determine the federated learning training mode based on the monitoring network bandwidth, latency, and data privacy requirements. The federated learning training mode includes centralized federated learning and decentralized federated learning. If it is determined to be centralized federated learning, execute step 3, otherwise execute step 4. Step 3: Centralized federated learning training, deploying the CrossTransNILM model on the client and the central server, and centralizing all client model updates participating in the training to the central server for aggregation; the CrossTransNILM model includes a multi-layer Transformer encoder and a load branching module, and predicts the power consumption of multiple loads through bus power; Step 4: Decentralized federated learning training, each client independently trains the CrossTransNILM model and exchanges model parameters through point-to-point communication, optimizes the communication topology through the graph attention network, and judges the convergence of the global model through the voting mechanism; Step 5: Non-intrusive load monitoring and real-time feedback. The trained CrossTransNILM model is deployed to the monitoring terminal to collect the power signal of the main line in real time. After preprocessing, it is input into the model for end-to-end monitoring. The power decomposition results of each electrical equipment are output through the load branch module, and the load power is mapped to the start and stop status of the electrical equipment. The monitoring terminal regularly updates the model parameters according to the federated learning mode selection mechanism.

2. According to claim 1, a non-intrusive load monitoring method based on hybrid federated learning and cross Transformer is characterized in that: In step 3, the CrossTransNILM model specifically includes: a one-dimensional convolutional layer, a positional encoding, a shared Transformer encoder, and multiple load branches, each of which processes a load type; each load branch contains a branched Transformer encoder layer, a cross Transformer layer, and two fully connected layers; The input bus sequence of the model is ,in is the length of the sequence, and is an odd number, used to locate the center point; the learning goal of the model is to predict the midpoint of the sequence Each load Power , the learning objective is expressed as: , in, is the hypothesis space, represents the total number of sampling points, is the mean square error loss function, is the total number of load types; The one-dimensional convolution layer is used to extract the local features of the sequence. The convolution operation formula is expressed as: , in, represents a one-dimensional convolution operation, is the input bus power sequence, is the output after convolution, representing the extracted local features; Positional encoding is used for each position Encoding the position in the sequence Generated by sine and cosine functions, the calculation formula is: in, is the model dimension, Represents the sequence position index; The embedding layer is used to add the convolutional layer output to the positional encoding to obtain the embedded representation of the sequence. : , in, is the output after the convolutional layer. is the position information encoded by sine and cosine; The shared Transformer encoder is used to process and analyze the input bus power sequence. Through its multi-layer structure and multi-head self-attention mechanism, it captures long-distance dependencies in the sequence data.

3. According to claim 2, a non-intrusive load monitoring method based on hybrid federated learning and cross Transformer is characterized in that: The shared Transformer encoder consists of multiple encoding layers, each of which includes a multi-head self-attention mechanism and a feed-forward neural network.

4. According to claim 3, a non-intrusive load monitoring method based on hybrid federated learning and cross Transformer is characterized in that: Step 3 includes: Step 3.1: The central server first initializes the global model parameters , and then send it to all clients participating in the training. The mathematical expression is: ,in, Indicates the total number of clients; Step 3.2: Client Using local data Perform training and update model parameters , the update rule is: , in, Is the client In the Model parameters for round 1 local training, is the learning rate for model training, is the mean square error loss function Gradients with respect to parameters; Define each client In the The model trained locally is ,in, Represents the client In the In the training round The parameters of the layer, is the total number of layers of the CrossTransNILM model. Layer parameter mean change , called layer sensitivity, is defined as: ; Calculate the sensitivity of each layer and sort it, and select the one with the highest sensitivity. Layer, where is the transmission ratio coefficient, the value range is ; Step 3.3: The server collects all model parameters uploaded by the client ,in is the index of the round number of global training; for each client , calculate the Euclidean distance between it and all other client updates and : , in, Indicates that except client The index of other clients participating in the training. Indicates Client in round global training Model parameters of The smaller the score, The closer the update is to other clients; then, Sort by removing the highest Layer update, where is the aggregation ratio coefficient, and its value range is ; Aggregate the remaining normal updates for each layer of the CrossTransNILM model , initialize the weight accumulator and counter , for each client upload weight , update the accumulator and counter , the average update of the calculation layer: ; Add the average update to the corresponding layer of the global model superior: ; Step 3.4: The central server updates the global model parameters Distribute to all clients to continue the next round of training: , 。 5. According to claim 1, a non-intrusive load monitoring method based on hybrid federated learning and cross Transformer is characterized in that: Step 4 includes: Step 4.1: Use the graph attention network to learn the similarity between clients and select the most relevant edges based on the attention weights to obtain the communication topology. Step 4.2: The central server generates initial model parameters , the model parameters contain all the network layers and parameters that need to be trained for the NILM task; then, the central server initializes the global model parameters Distribute to all clients participating in the training so that each client has the same starting point before the training begins. The mathematical expression is as follows: , in, Represents the client The model parameters at initialization, is the total number of clients; Step 4.3: Perform local model training. Each client builds a communication topology based on the Share model parameters with neighboring clients; each client After receiving the updated parameters from its neighboring clients, the model parameters of each client are aggregated by weighted averaging; Step 4.4: Evaluate convergence and record the judgment result as a binary signal: , in, Indicates that the client believes that the current model has converged. Indicates that the model has not yet converged; Represents the client Validate the set locally The loss value on ; is the preset threshold, is the target accuracy, For Model Validate the set locally The accuracy of Step 4.5: After each round of training is completed, all clients send their own convergence signals Sent to the central server; the central server determines whether to stop training based on the statistical voting results. ; Set voting threshold is a static fixed value. , the system stops training, otherwise it enters the next round of training.

6. A non-intrusive load monitoring method based on hybrid federated learning and cross-Transformer according to claim 5, characterized in that: Step 4.1 includes: Step 4.1.1: Each client First, based on local data , extract key statistical features, including mean, standard deviation, minimum and maximum values, mathematically expressed as: , Among them, the characteristics of all clients are expressed as , ; make Represents the communication topology, client set , edge set Represents the potential communication connection relationship between clients, and initializes the edge set to a fully connected structure , does not include self-loops , that is, initially set to a completely directed graph ; Step 4.1.2: Use a three-layer graph attention network to build communication connections between clients. With its neighbor clients , the graph attention network first performs a learnable linear mapping on the client features ; Then, for each edge Calculate the attention coefficient , which is of the form: , in, is the activation function, is the learnable attention vector, represents vector concatenation, and Respectively represent the client and its neighbor clients The feature matrix of Then, after softmax normalization, the final attention coefficient is obtained : ; In the forward propagation process, the updated node representation is further obtained. And the set of attention coefficients of all edges ; Represented as client and The connection weight of Call , in, Embedding matrix for the updated client, is the set of attention coefficients; Step 4.1.3: Set the attention weight threshold for The median of The topological edges of: , and remove self-loops ; for Each directed edge in , only the undirected form is retained , the obtained undirected edge set ; For any client , Count its degree ,in, Representation and Connected neighbor client; if there is , indicating that the client is If there is no connecting edge in Pick a neighbor client and merge the new edge into it ; make , the final undirected graph This is the decentralized communication topology obtained based on GAT learning.

7. The non-intrusive load monitoring method based on hybrid federated learning and cross-Transformer according to claim 1, characterized in that: Step 2 specifically includes: When network bandwidth ,Delay , and the data privacy requirements are low, the system chooses centralized federated learning, otherwise, it switches to decentralized federated learning; If the conditions for centralized federated learning training are met, the central server will aggregate and optimize the model updates of all clients; If the conditions for decentralized federated learning training are triggered, the client trains independently in a distributed environment and exchanges model updates with neighboring clients.

8. A non-intrusive load monitoring system based on hybrid federated learning and cross-Transformer, characterized in that: The system comprises: A preprocessing module, used for preprocessing power signal data; A federated learning training mode selection module is used to determine the federated learning training mode according to the monitoring network bandwidth, latency and data privacy requirements, wherein the federated learning training mode includes centralized federated learning and decentralized federated learning; if centralized federated learning is determined, the centralized federated learning module is executed, otherwise the decentralized federated learning module is executed; Centralized federated learning module, used for centralized federated learning training, deploys CrossTransNILM model on the client and central server, and centralizes all client model updates participating in the training to the central server for aggregation; CrossTransNILM model includes multi-layer Transformer encoder and load branch module, and predicts the power consumption of multiple loads through bus power; Decentralized federated learning module, used for decentralized federated learning training. Each client independently trains the CrossTransNILM model and exchanges model parameters through point-to-point communication. The communication topology is optimized through the graph attention network, and the convergence of the global model is judged through the voting mechanism. The monitoring module is used for non-intrusive load monitoring and real-time feedback. It deploys the trained CrossTransNILM model to the monitoring terminal, collects the power signal of the main line in real time, and inputs it into the model for end-to-end monitoring after preprocessing. The power decomposition results of each electrical equipment are output through the load branch module, and the load power is mapped to the start and stop status of the electrical equipment. The monitoring terminal regularly updates the model parameters according to the federated learning mode selection mechanism.

9. A computer device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor runs the computer program stored in the memory, the steps of the method according to any one of claims 1 to 7 are performed.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a plurality of computer instructions, and the plurality of computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • A federated learning-based non-intrusive enterprise load decomposition method

    CN113962314A

  • Load monitoring model construction method and device, computer equipment, medium and product

    CN118839988A

  • Incremental clustering classifier and predictor

    US20020059202A1

Cited By

  • Genomics-oriented anti-quantum robust parameter aggregation federated learning method and device

    CN120146158A

  • Quantum-resistant robust parameter aggregation federated learning method and device for genomics

    CN120146158B

  • Transform model based on transcription factor cross attention prediction enhancer-promoter interaction

    CN120183487A

  • Non-intrusive load decomposition method and device for energy scene

    CN120337161A

  • Short-term power load prediction method based on KNN federated distillation learning and Seq2Seq

    CN120709957A